October 2025 arXiv papers — page 119
Showing 11,801–11,900 of 25,213 papers
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
cs.CVMattia Segu, Marta Tintore Gazulla, Yongqin Xian, Luc Van Gool
Scaling up model size and training data has advanced foundation models for instance-level perception, achieving state-of-the-art in-domain and zero-shot performance across object detection and segmentation. However, their high computational cost limits adoption on resource-constrained platforms. We first examine the limitations of existing architectures in e
Alessandro Della Croce
Black holes (BHs) play a major role in the structural and dynamical evolutions of Globular Clusters (GCs). Several recent works searched for BHs in Galactic GCs using scaling relations derived from numerical simulations. However, the conclusions drawn by such approaches are strongly dependent on the specific prescriptions adopted in numerical simulations. Th
An Investigation into the Low-Mass Fundamental Metallicity Relation in the Local and High-z Universe
astro-ph.GAIsaac H. Laseter, Michael V. Maseda, Andrew J. Bunker, Alex J. Cameron
Recent JWST/NIRSpec observations have revealed high-$z$ star-forming galaxies depart from the Fundamental Metallicity Relation (FMR), yet the $z = 0$ FMR has not been well-characterized in the low-mass regime ($\rm log(M_{\star}/M_{\odot}) \lesssim 9$) for an appropriate comparison of low- and high-$z$ systems. We attempt to rectify this limitation through a
Radiative Correction from Secret Neutrino Interactions and Implications for Neutrino-Scattering Experiments
hep-phSaeid Foroughi-Abari, Kevin J. Kelly, Yue Zhang
New, neutrinophilic mediators are one potential extension beyond the Standard Model of particle physics. Often, studies of neutrinophilic mediator consist of searching for direct evidence of its production and/or its tree-level virtual effect for generating strong neutrino self-interaction. In this work, we focus instead on the fact that such new mediators \
Hadi Alzayer, Yunzhi Zhang, Chen Geng, Jia-Bin Huang
We present an inference-time diffusion sampling method to perform multi-view consistent image editing using pre-trained 2D image editing models. These models can independently produce high-quality edits for each image in a set of multi-view images of a 3D scene or object, but they do not maintain consistency across views. Existing approaches typically addres
Wenqian Zhang, Yangyi Huang, Weiyang Liu, Zhen Liu
Large language models (LLMs) have shown strong abilities in writing and revising programs, yet many program-synthesis benchmarks still evaluate programs in symbolic or digital environments. We introduce compositional machine design, a physically grounded form of program synthesis where machines are written as programs that compose standardized parts, and suc
Haiwen Diao, Mingxuan Li, Silei Wu, Linjun Dai
The edifice of native Vision-Language Models (VLMs) has emerged as a rising contender to typical modular VLMs, shaped by evolving model architectures and training paradigms. Yet, two lingering clouds cast shadows over its widespread exploration and promotion: (-) What fundamental constraints set native VLMs apart from modular ones, and to what extent can the
Nupur Kumari, Sheng-Yu Wang, Nanxuan Zhao, Yotam Nitzan
Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Current workarounds use synthetic training pairs that leverage the
Yuanhui Huang, Weiliang Chen, Wenzhao Zheng, Xin Tao
World models have garnered increasing attention for comprehensive modeling of the real world. However, most existing methods still rely on pixel-aligned representations as the basis for world evolution, neglecting the inherent 3D nature of the physical world. This could undermine the 3D consistency and diminish the modeling efficiency of world models. In thi
Shaowei Liu, Chuan Guo, Bing Zhou, Jian Wang
Close-proximity human-human interactive poses convey rich contextual information about interaction dynamics. Given such poses, humans can intuitively infer the context and anticipate possible past and future dynamics, drawing on strong priors of human behavior. Inspired by this observation, we propose Ponimator, a simple framework anchored on proximal intera
Hengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang
Identity-consistent generation has become an important focus in text-to-image research, with recent models achieving notable success in producing images aligned with a reference identity. Yet, the scarcity of large-scale paired datasets containing multiple images of the same individual forces most approaches to adopt reconstruction-based training. This relia
Hansheng Chen, Kai Zhang, Hao Tan, Leonidas Guibas
Few-step diffusion or flow-based generative models typically distill a velocity-predicting teacher into a student that predicts a shortcut towards denoised data. This format mismatch has led to complex distillation procedures that often suffer from a quality-diversity trade-off. To address this, we propose policy-based flow models ($\pi$-Flow). $\pi$-Flow mo
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
This work studies how to adaptively recompute key-value (KV) caches for diffusion large language models (DLMs) to maximize prediction accuracy while minimizing decoding latency. Prior methods' decoders recompute QKV for all tokens at every denoising step and layer, despite KV states changing little across most steps, especially in shallow layers, leading to
Mert Sonmezer, Matthew Zheng, Pinar Yanardag
Low-rank Adaptation (LoRA) models have revolutionized the personalization of pre-trained diffusion models by enabling fine-tuning through low-rank, factorized weight matrices specifically optimized for attention layers. These models facilitate the generation of highly customized content across a variety of objects, individuals, and artistic styles without th
Yinxi Li, Yuntian Deng, Pengyu Nie
Large language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but driven by statistics rather than grammar. As a result, semantically identical code snippets can be tokenized differently depending on superficial factors such as whitespace or identifie
Christopher A. Schroeder, Hung P. Tong-Viet
In this paper, we investigate structural properties of finite groups that are detected by certain group invariants arising from Dijkgraaf--Witten theory, a topological quantum field theory, in one space and one time dimension. In this setting, each finite group $G$ determines a family of numerical invariants associated with closed orientable surfaces, expres
Biology-informed neural networks learn nonlinear representations from omics data to improve genomic prediction and interpretability
cs.LGKatiana Kontolati, Rini Jasmine Gladstone, Ian Davis, Ethan Pickering
We extend biologically-informed neural networks (BINNs) for genomic prediction (GP) and selection (GS) in crops by integrating thousands of single-nucleotide polymorphisms (SNPs) with multi-omics measurements and prior biological knowledge. Traditional genotype-to-phenotype (G2P) models depend heavily on direct mappings that achieve only modest accuracy, for
Yiming Wang, Da Yin, Yuedong Cui, Ruichen Zheng
Digital agents require diverse, large-scale UI trajectories to generalize across real-world tasks, yet collecting such data is prohibitively expensive in both human annotation, infra and engineering perspectives. To this end, we introduce $\textbf{UI-Simulator}$, a scalable paradigm that generates structured UI states and transitions to synthesize training t
Mingxuan Yan, Yuping Wang, Zechun Liu, Jiachen Li
To tackle long-horizon tasks, recent hierarchical vision-language-action (VLAs) frameworks employ vision-language model (VLM)-based planners to decompose complex manipulation tasks into simpler sub-tasks that low-level visuomotor policies can easily handle. Typically, the VLM planner is finetuned to learn to decompose a target task. This finetuning requires
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
cs.CLGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan
Large language model (LLM)-based agents are increasingly trained with reinforcement learning (RL) to enhance their ability to interact with external environments through tool use, particularly in search-based settings that require multi-turn reasoning and knowledge acquisition. However, existing approaches typically rely on outcome-based rewards that are onl
Jiaxin Ge, Grace Luo, Heekyung Lee, Nishant Malpani
Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community perceptions of progress and formal evaluation. To address thi
Zachary Robertson
Pairwise comparisons of large language models using total variation distance mutual information (TVD-MI) produce binary critic decisions per pair. We show that averaging TVD-MI's binary trials yields centered-probability scores with additive structure suitable for item-response theory (IRT) without nonlinear link functions. Maximum-likelihood approaches to I
Miao Hu, Zhiwei Huang, Tai Wang, Jiangmiao Pang
Real-world robots localize objects from natural-language instructions while scenes around them keep changing. Yet most of the existing 3D visual grounding (3DVG) method still assumes a reconstructed and up-to-date point cloud, an assumption that forces costly re-scans and hinders deployment. We argue that 3DVG should be formulated as an active, memory-driven
Daniel R. Reynolds, Sylvia Amihere, Dashon Mitchell, Vu Thai Luan
In this work we present two new families of multirate time step adaptivity controllers, that are designed to work with embedded multirate infinitesimal (MRI) time integration methods for adapting time steps when solving problems with multiple time scales. We compare these controllers against competing approaches on two benchmark problems, showing that the pr
Orders matter: tight bounds on the precision of sequential quantum estimation for multiparameter models
quant-phGabriele Fazio, Jiayu He, Matteo G. A. Paris
In multiparameter quantum metrology, the ultimate precision of joint estimation is dictated by the Holevo Cram\'er-Rao bound. In this paper, we discuss and analyze in detail an alternative approach: the stepwise estimation strategy. In this approach, parameters are estimated sequentially, using an optimized fraction of the total available resources allocated
Thao Nguyen, Jiaqi Ma, Fahad Shahbaz Khan, Souhaib Ben Taieb
Precipitation nowcasting, predicting future radar echo sequences from current observations, is a critical yet challenging task due to the inherently chaotic and tightly coupled spatio-temporal dynamics of the atmosphere. While recent advances in diffusion-based models attempt to capture both large-scale motion and fine-grained stochastic variability, they of
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
cs.LGJonas Geiping, Xinyu Yang, Guinan Su
Language models with recurrent depth, also referred to as universal or looped when considering transformers, are defined by the capacity to increase their computation through the repetition of layers. Recent efforts in pretraining have demonstrated that these architectures can scale to modern language modeling tasks while exhibiting advantages in reasoning t
Shizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao Wang
Recovering 4D from monocular video, which jointly estimates dynamic geometry and camera poses, is an inevitably challenging problem. While recent pointmap-based 3D reconstruction methods (e.g., DUSt3R) have made great progress in reconstructing static scenes, directly applying them to dynamic scenes leads to inaccurate results. This discrepancy arises becaus
Weikang Shi, Aldrich Yu, Rongyao Fang, Houxing Ren
While Large Language Models (LLMs) have excelled in textual reasoning, they struggle with mathematical domains like geometry that intrinsically rely on visual aids. Existing approaches to Visual Chain-of-Thought (VCoT) are often limited by rigid external tools or fail to generate the high-fidelity, strategically-timed diagrams necessary for complex problem-s
Rayne Liu, Yijie Zhu, Wayne Hu, Vivian Miranda
Supernova (SN) and baryon acoustic oscillation (BAO) distance measures have recently provided hints that the dark energy is not only dynamical but apparently evolves from normal to phantom dark energy between redshifts $0<z<1$. A normal axion dark energy component in the mass range just below the Hubble scale can mimic a phantom component by appearing as dar
Ryota Inagaki, Dimana Pramatarova
We define a weighted analog for the multidimensional Catalan numbers, obtain matrix-based recurrences for some of them, and give conditions under which they are periodic. Building on this framework, we introduce two new sequences of triangular arrays: the first one enumerates the $k$-dimensional Balanced ballot paths of exact height $s$; the second one is a
Guo Cheng, Danni Yang, Ziqi Huang, Jianlou Si
Video generative models have recently achieved notable advancements in synthesis quality. However, generating complex motions remains a critical challenge, as existing models often struggle to produce natural, smooth, and contextually consistent movements. This gap between generated and real-world motions limits their practical applicability. To address this
Zhe Li, Weihao Yuan, Weichao Shen, Siyu Zhu
Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous methods that usually employ discrete masked modeling or autoregressive modeling, we develop a continuous masked autoregre
Luka Vujeva, Jose María Ezquiaga, Daniel Gilman, Srashti Goyal
Gravitational lensing is an invaluable probe of the nature of dark matter, and the structures it forms. Lensed gravitational waves in particular allow for unparalleled sensitivity to small scale structures within the lenses, due to the precise time resolution in combination with the continuous monitoring of the entire sky. In this work, we show two distinct
Zhe Li, Cheng Chi, Yangyang Wei, Boan Zhu
Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone to cumulative errors, introduces hig
A universal description of Mott insulators: Characterizing quantum phases beyond broken symmetries
cond-mat.str-elMatheus de Sousa, Zhiyu Fan, Wei Ku
Using Mott insulators as a prototypical example, we demonstrate a dynamics-based characterization of quantum phases of matter through a general N-body renormalization group framework. The essential "Mott-ness" turns out to be characterized by a change of size-scaling of the effective intra- momentum repulsions between long-lived emergent "eigen-particles" th
Mark Dominique Dalipe Muñoz
Model misspecification of formative indicators remains a widely documented issue across academic literature, yet scholars lack a clear consensus on pragmatic, prescriptive approaches to manage this gap. This ambiguity forces researchers to rely on psychometric frameworks primarily intended for reflective models, and thus risks misleading findings. This artic
Yu Zhou, Sohyun An, Haikang Deng, Da Yin
Contact languages like English exhibit rich regional variations in the form of dialects, which are often used by dialect speakers interacting with generative models. However, can multimodal generative models effectively produce content given dialectal textual input? In this work, we study this question by constructing a new large-scale benchmark spanning six
S. Reimann, H. Albers, R. W. Assmann, P. Gasik
The international Facility for Antiproton and Ion Research (FAIR) is under construction at the GSI Helmholtz Centre in Darmstadt. The first project stage includes the superconducting 100 Tm heavy-ion synchrotron SIS100, the Super Fragment Separator, and associated beam transport lines. Part of GSI's existing accelerator chain, comprising UNILAC and SIS18, wi
Blake Werner, Lizhi Yang, Aaron D. Ames
Robust humanoid locomotion in unstructured environments requires architectures that balance fast low-level stabilization with slower perceptual decision-making. We show that a simple layered control architecture (LCA), a proprioceptive stabilizer running at high rate, coupled with a compact low-rate perceptual policy, enables substantially more robust perfor
Romina Aalishah, Mozhgan Navardi, Tinoosh Mohsenin
Deployment of efficient and accurate Deep Learning models has long been a challenge in autonomous navigation, particularly for real-time applications on resource-constrained edge devices. Edge devices are limited in computing power and memory, making model efficiency and compression essential. In this work, we propose EdgeNavMamba, a reinforcement learning-b
JoungBin Lee, Jaewoo Jung, Jisang Han, Takuya Narihira
We present 3DScenePrompt, a framework that generates the next video chunk from arbitrary-length input while enabling precise camera control and preserving scene consistency. Unlike methods conditioned on a single image or a short clip, we employ dual spatio-temporal conditioning that reformulates context-view referencing across the input video. Our approach
Yuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi
Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics presents unique challenges with its complex biochemical pathways, heterogeneous identifier systems, and fragmented databa
Wenkai Yang, Weijie Liu, Ruobing Xie, Yiju Guo
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a core paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs). To address the lack of verification signals at test time, prior studies incorporate the training of model's self-verification capability into the standard RLVR process, thereby unifying reason
Yao Zhang, Yu Wu, Haowei Zhang, Weiguo Li
Process Reward Models (PRMs) aim to improve multi-step reasoning in Large Language Models (LLMs) by supervising intermediate steps and identifying errors. However, building effective PRMs remains challenging due to the lack of scalable, high-quality annotations. Existing approaches rely on costly human labeling, LLM-based self-evaluation that is prone to hal
Fan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi
Language models demonstrate remarkable abilities when pre-trained on large text corpora and fine-tuned for specific tasks, but how and why pre-training shapes the success of the final model remains poorly understood. Notably, although pre-training success is often quantified by cross-entropy loss, cross-entropy can be a poor predictor of downstream performan
Samuel Sánchez López, Alexandros Karam, Dhiraj Kumar Hazra
We analyze a model of quintessence governed by an exponential potential and non-minimally coupled to gravity, in light of recent datasets, including cosmic microwave background, baryon acoustic oscillations, and supernovae distance moduli observations. Mainly focusing on the Palatini formulation of gravity, a phase space analysis reveals the existence of a l
J. C. Helo, M. Hirsch, T. Ota
In Standard Model Effective Field Theory (SMEFT), invisible neutron decay arises from d = 12 operators. Adding new light particles to the field content of the SM, such as right-handed neutrinos, allows one to construct operators for invisible neutron decay at much lower dimensions. Observing invisible neutron decay, if nucleon decays with charged leptons rem
Junliang Ye, Shenghao Xie, Ruowen Zhao, Zhengyi Wang
3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient, inconsistent, and often fail to preserve unedited regions. Most methods rely on editing multi-view renderings followed by reconstruction, which introduces artifacts and limits practicality. To address these challenges
Ken R. Duffy, Moritz Grundei, Jane A. Millward, Muralidhar Rangaswamy
Inter symbol interference (ISI), which occurs in a wide variety of channels, is a result of time dispersion. It can be mitigated by equalization, which results in noise coloring. Inspired by the development of Approximate Independence in statistical physics, for such colored noise we propose a decoder called Ordered Reliability Bits Guessing Random Additive
Rashid Sunyaev, Ildar Khabibullin, Eugene Churazov, Marat Gilfanov
The Galactic microquasar SS 433 and the radio nebula W50 surrounding it present a prototypical example of a hyper-Eddington binary system shaping its ambient interstellar medium via energetic outflows. In this paper, we present X-ray observations of the SS 433/W50 complex by the eROSITA telescope onboard the \textit{SRG} space observatory. These data provide
Jianfeng Zhu, Julina Maharjan, Xinyu Li, Karin G. Coifman
Mental health disorders remain among the leading cause of disability worldwide, yet conditions such as depression, anxiety, and Post-Traumatic Stress Disorder (PTSD) are frequently underdiagnosed or misdiagnosed due to subjective assessments, limited clinical resources, and stigma and low awareness. In primary care settings, studies show that providers misid
Elena Golimblevskaia, Aakriti Jain, Bruno Puri, Ammar Ibrahim
The fields of explainable AI and mechanistic interpretability aim to uncover the internal structure of neural networks, with circuit discovery as a central tool for understanding model computations. Existing approaches, however, rely on manual inspection and remain limited to toy tasks. Automated interpretability offers scalability by analyzing isolated feat
Abraar Chaudhry, Katya Scheinberg
In many applications of mathematical optimization, one may wish to optimize an objective function without access to its derivatives. These situations call for derivative-free optimization (DFO) methods. Among the most successful approaches in practice are model-based trust-region methods, such as those pioneered by M.J.D Powell. While relatively complex to i
Ming-Hao Hsu, Liang-Hsuan Tseng, Hung-yi Lee, Zhizheng Wu
We propose Text-Aligned Speech Tokens with Multiple Layer-Aggregation (TASLA), which is a text-aligned speech tokenization framework that aims to address the problem that under a low-frame-rate and text-aligned regime, single-source speech tokens may lose acoustic details during reconstruction. On the other hand, this paper further explains how different enc
Dip Sarker, Abdoulaye Ndao
Precise control of plasmonic resonances across a broad spectral range is central to the development of tunable optical devices. Yet, achieving both redshifts and blueshifts within a single nanostructure has remained elusive. Here we introduce a metal-dielectric-metal (MDM) nanodisk array that enables bidirectional tuning of resonance wavelengths throughout t
Further Results on Safety-Critical Stabilization of Force-Controlled Nonholonomic Mobile Robots
eess.SYBo Wang, Tianyu Han, Guangwei Wang
In this paper, we address the stabilization problem for force-controlled nonholonomic mobile robots under safety-critical constraints. We propose a continuous, time-invariant control law based on the gamma m-quadratic programming (gamma m-QP) framework, which unifies control Lyapunov functions (CLFs) and control barrier functions (CBFs) to enforce both stabi
Mingxuan Liu, Honglin He, Elisa Ricci, Wayne Wu
Urban embodied AI agents, ranging from delivery robots to quadrupeds, are increasingly populating our cities, navigating chaotic streets to provide last-mile connectivity. Training such agents requires diverse, high-fidelity urban environments to scale, yet existing human-crafted or procedurally generated simulation scenes either lack scalability or fail to
Binghao Huang, Jie Xu, Iretiayo Akinola, Wei Yang
Humans excel at bimanual assembly tasks by adapting to rich tactile feedback -- a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy learning framework that combines real-world demonstratio
ChenYu Wu, Yi Wang, Yang Liao
Large language models (LLMs) are increasingly vulnerable to multi-turn jailbreak attacks, where adversaries iteratively elicit harmful behaviors that bypass single-turn safety filters. Existing defenses predominantly rely on passive rejection, which either fails against adaptive attackers or overly restricts benign users. We propose a honeypot-based proactiv
Yingtian Chen, Oleg Y. Gnedin, Adrian M. Price-Whelan, Colin Holm-Hansen
The Gaia mission has led to the discovery of over 100 stellar streams in the Milky Way, most of which likely originated from globular clusters (GCs). As the upcoming wide-field surveys can potentially continue to increase the number of known streams, there is a growing need to shift focus from manual detection of individual streams to automated detection met
Eric Christopher, Kevin Crossan, Wolff Dobson, Chris Kennelly
Migrating codebases from one instruction set architecture (ISA) to another is a major engineering challenge. A recent example is the adoption of Arm (in addition to x86) across the major Cloud hyperscalers. Yet, this problem has seen limited attention by the academic community. Most work has focused on static and dynamic binary translation, and the tradition
NIRPS and TESS reveal a peculiar system around the M dwarf TOI-756: A transiting sub-Neptune and a cold eccentric giant
astro-ph.EPLéna Parc, François Bouchy, Neil J. Cook, Nolan Grieves
The Near InfraRed Planet Searcher (NIRPS) joined HARPS on the 3.6-m ESO telescope at La Silla Observatory in April 2023, dedicating part of its Guaranteed Time Observations (GTO) program to the radial velocity follow-up of TESS planet candidates to confirm and characterize transiting planets around M dwarfs. We report the first results of this program with t
Yingtian Chen, Oleg Y. Gnedin, Adrian M. Price-Whelan
We apply the automatic stellar stream detection algorithm StarStream to Gaia Data Release 3 and identify 87 stellar streams associated with Galactic globular clusters (GCs), including 34 high-quality cases with median completeness and purity both exceeding 50%, as estimated from modeling mock streams. These detections double the number of known GC streams, a
Matisse De Lescluze, Michal P. Heller, Aleksas Mazeliauskas, Bruno Scheihing-Hitschfeld
Nonthermal attractors govern the emergent self-similar dynamics of far-from-equilibrium quantum systems, from ultrarelativistic nuclear collisions to cold-atom experiments. Within the framework of adiabatic hydrodynamization, the approach to a nonthermal attractor is described by the decay of excited states of an effective Hamiltonian. Using an exactly solva
Aaron Baier-Reinio, Patrick E. Farrell, Charles W. Monroe
We present a broad family of high-order finite element algorithms for simulating the flow of electroneutral electrolytes. The governing partial differential equations that we solve are the electroneutral Navier--Stokes--Onsager--Stefan--Maxwell (NSOSM) equations, which model momentum transport, multicomponent diffusion and electrical effects within the elect
Annisaa Fitri Nurfidausi, Eleonora Mancini, Paolo Torroni
Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging complementary signals. However, existing studies are limited in scope, lack systematic comparisons of features, and suffer from inconsistent evalu
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
cs.CVMor Ventura, Michael Toker, Or Patashnik, Yonatan Belinkov
Text-to-Image (T2I) models have advanced rapidly, yet they remain vulnerable to semantic leakage, the unintended transfer of semantically related features between distinct entities. Existing mitigation strategies are often optimization-based or dependent on external inputs. We introduce DeLeaker, a lightweight, optimization-free inference-time approach that
Justin Faber, Alexandros C Alampounti, Marcos Georgiades, Joerg T Albert
The use of auditory masking has long been of interest in psychoacoustics and for engineering purposes, in order to cover sounds that are disruptive to humans or to species whose habitats overlap with ours. In most cases, we seek to minimize the disturbances to the communication of wildlife. However, in the case of pathogen-carrying insects, we may want to ma
Sumit Singh, Sivaram Ambikasaran
Kernel functions are frequently encountered in differential equations and machine learning applications. In this work, we study the rank of matrices arising out of the kernel function $K: X \times Y \mapsto \mathbb{R}$, where the sets $X, Y \in \mathbb{R}^d$ are hypercubes that share a boundary. The main contribution of this work is the analysis of the rank
Kyle Montgomery, David Park, Jianhong Tu, Michael Bendersky
Scaling laws have transformed our understanding of large language models by linking upstream metrics like cross-entropy loss to design factors such as model size, training data, and compute. However, these conventional laws fail to capture downstream task performance, where context plays a critical role. In this work, we propose a straightforward, interpreta
Hung Le, Lazar Milenković, Shay Solomon, Cuong Than
Sparse shortcuttings of trees -- equivalently, sparse 1-spanners for tree metrics with bounded hop-diameter -- have been studied extensively (under different names and settings), since the pioneering works of [Yao82, Cha87, AS87, BTS94], initially motivated by applications to range queries, online tree product, and MST verification, to name a few. These cons
Hasan Ahmed, Deena Goodgold, Khushali Kothari, Rustom Antia
Cumulants and moments are closely related to the basic mathematics of continuous and discrete selection (respectively). These relationships generalize Fisher's fundamental theorem of natural selection and also make clear some of its limitation. The relationship between cumulants and continuous selection is especially intuitive and also provides an alternativ
Xujun Peng, Anoop Kumar, Jingyu Wu, Parker Glenn
Retrieval-Augmented Generation (RAG) systems leverage Large Language Models (LLMs) to generate accurate and reliable responses that are grounded in retrieved context. However, LLMs often generate inconsistent outputs for semantically equivalent inputs, a problem compounded by the scarcity of consistency-focused training data and the limitations of current fi
Ruhan Yang, Ellen Yi-Luen Do
Building robots is an engaging activity that provides opportunities for hands-on learning. However, traditional robot-building kits are usually costly with limited functionality due to material and technology constraints. To improve the accessibility and flexibility of such kits, we take paper as the building material and extensively explore the versatility
Kyle Montgomery, Sijun Tan, Yuqi Chen, Siyuan Zhuang
Test-time scaling is a powerful strategy for boosting the performance of large language models on complex reasoning tasks. While state-of-the-art approaches often employ generative verifiers to select the best solution from a pool of candidates, this method incurs prohibitive computational costs, limiting its practicality. In this work, we shift the focus to
Decoherence-Aware Entangling and Swapping Strategy Optimization for Entanglement Routing in Quantum Networks
quant-phShao-Min Huang, Cheng-Yang Cheng, Ming-Huang Chien, Jian-Jhih Kuo
Quantum teleportation enables high-security communications through end-to-end quantum entangled pairs. End-to-end entangled pairs are created by using swapping processes to consume short entangled pairs and generate long pairs. However, due to environmental interference, entangled pairs decohere over time, resulting in low fidelity. Thus, generating entangle
Dude, Where's My (Autonomous) Car? Defining an Accessible Description Logic for Blind and Low Vision Travelers Using Autonomous Vehicles
cs.HCPaul D. S. Fink, Justin R. Brown, Rachel Coombs, Emily A. Hamby
Purpose: Autonomous vehicles (AVs) are becoming a promising transportation solution for blind and low-vision (BLV) travelers, offering the potential for greater independent mobility. This paper explores the information needs of BLV users across multiple steps of the transportation journey, including finding and navigating to, entering, and exiting vehicles i
Carlos Román, Etienne Sandier, Sylvia Serfaty
We complete our study of the three dimensional Ginzburg--Landau functional with magnetic field, in the asymptotic regime of a small inverse Ginzburg--Landau parameter $\varepsilon$, and near the first critical field $H_{c_1}$ for which the first vortex filaments appear in energy minimizers. Under a nondegeneracy condition, we show a next order asymptotic exp
The Impact of Medicaid Coverage on Mental Health, Why Insurance Makes People Happier in OHIE: by Spending Less or by Spending More?
econ.GNYangyang Li
The Oregon Health Insurance Experiment (OHIE) offers a unique opportunity to examine the causal relationship between Medicaid coverage and happiness among low-income adults, using an experimental design. This study leverages data from comprehensive surveys conducted at 0 and 12 months post-treatment. Previous studies based on OHIE have shown that individuals
Fermi Bubbles Without AGN: Gamma-Ray Bubbles in MHD Galaxy Formation Simulations with Full Cosmic Ray Spectra
astro-ph.HEIsabel S. Sands, Philip F. Hopkins, Sam B. Ponnada
For the first time, we show in MHD simulations with cosmological initial conditions that bi-lobed gamma-ray outflows similar to the Fermi bubbles can form from star formation and supernova feedback, without involvement from active galactic nuclei (AGN). We use simulations run with full MHD and dynamical, on-the-fly multi-species cosmic ray transport in MeV-T
Algorithms for dynamic scheduling in manufacturing, towards digital factories Improving Deadline Feasibility and Responsiveness via Temporal Networks
cs.AIIoan Hedea
Modern manufacturing systems must meet hard delivery deadlines while coping with stochastic task durations caused by process noise, equipment variability, and human intervention. Traditional deterministic schedules break down when reality deviates from nominal plans, triggering costly last-minute repairs. This thesis combines offline constraint-programming (
Zixuan Liu, Yi Zhao, Zhuotao Liu, Qi Li
Machine Learning (ML)-based malicious traffic detection is a promising security paradigm. It outperforms rule-based traditional detection by identifying various advanced attacks. However, the robustness of these ML models is largely unexplored, thereby allowing attackers to craft adversarial traffic examples that evade detection. Existing evasion attacks typ
L. Basseto, N. P. Vizarim, J. C. Bellizotti Souza, P. A. Venegas
We investigate the driven dynamics of a single skyrmion in a square lattice of mixed pinning sites, where attractive and repulsive defects coexist using a particle-based model. The mixed landscape yields directional locking at $\theta_{\rm sk}=-45^\circ$ and flow at locked angles near the intrinsic skyrmion Hall angle. By mapping defect strengths, we show th
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
cs.ROHan Zhao, Jiaxuan Zhang, Wenxuan Song, Pengxiang Ding
Current vision-language-action (VLA) models, pre-trained on large-scale robotic data, exhibit strong multi-task capabilities and generalize well to variations in visual and language instructions for manipulation. However, their success rate drops significantly when faced with object concepts outside the training data, such as unseen object descriptions and t
Aayush Karan, Yilun Du
Frontier reasoning models have exhibited incredible capabilities across a wide array of disciplines, driven by posttraining large language models (LLMs) with reinforcement learning (RL). However, despite the widespread success of this paradigm, much of the literature has been devoted to disentangling truly novel behaviors that emerge during RL but are not pr
Mapping Smarter, Not Harder: A Test-Time Reinforcement Learning Agent That Improves Without Labels or Model Updates
cs.AIWen-Kwang Tsao, Yao-Ching Yu, Chien-Ming Huang
The Enterprise Intelligence Platform must integrate logs from numerous third-party vendors in order to perform various downstream tasks. However, vendor documentation is often unavailable at test time. It is either misplaced, mismatched, poorly formatted, or incomplete, which makes schema mapping challenging. We introduce a reinforcement learning agent that
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
cs.CVFurkan Mumcu, Michael J. Jones, Anoop Cherian, Yasin Yilmaz
Existing semi-supervised video anomaly detection (VAD) methods often struggle with detecting complex anomalies involving object interactions and generally lack explainability. To overcome these limitations, we propose a novel VAD framework leveraging Multimodal Large Language Models (MLLMs). Unlike previous MLLM-based approaches that make direct anomaly judg
Alexander Cowtan, Zhiyang He, Dominic J. Williamson, Theodore J. Yoder
Quantum code surgery is a promising technique to perform fault-tolerant computation on quantum low-density parity-check codes. Recent developments have significantly reduced the space overhead of surgery. However, generic surgery operations still require $O(d)$ rounds of repeated syndrome extraction to be made fault-tolerant. In this work, we focus on reduci
Secure Sparse Matrix Multiplications and their Applications to Privacy-Preserving Machine Learning
cs.CRMarc Damie, Florian Hahn, Andreas Peter, Jan Ramon
To preserve data privacy, multi-party computation (MPC) enables executing Machine Learning (ML) algorithms on private data. However, MPC frameworks do not include optimized operations on sparse data. This absence makes them unsuitable for ML applications involving sparse data; e.g., recommender systems or genomics. Even in plaintext, such applications involv
Shubham Varma, Ananya Warior, Avani Sakhapara, Dipti Pawade
The Indian judicial system faces a critical challenge with approximately 52 million pending cases, causing significant delays that impact socio-economic stability. This study proposes a cloud-based software framework to classify and prioritize court cases using algorithmic methods based on parameters such as severity of crime committed, responsibility of par
Andrew Welbaum, Wanli Qiao
In a mixture of linear regression model, the regression coefficients are treated as random vectors that may follow either a continuous or discrete distribution. We propose two Expectation-Maximization (EM) algorithms to estimate this prior distribution. The first algorithm solves a kernelized version of the nonparametric maximum likelihood estimation (NPMLE)
Detecting Early and Implicit Suicidal Ideation via Longitudinal and Information Environment Signals on Social Media
cs.SISoorya Ram Shimgekar, Ruining Zhao, Agam Goyal, Violeta J. Rodriguez
On social media, several individuals experiencing suicidal ideation (SI) do not disclose their distress explicitly. Instead, signs may surface indirectly through everyday posts or peer interactions. Detecting such implicit signals early is critical but remains challenging. We frame early and implicit SI as a forward-looking prediction task and develop a comp
João Rebouças, Victoria Lloyd, Jonathan Gordon, Guilherme Brando
Upcoming galaxy surveys will bring a wealth of information about the clustering of matter, but modeling small-scale structure beyond $\Lambda$CDM remains computationally challenging. While accurate N-body emulators exist to model the matter power spectrum for $\Lambda$CDM and some limited extensions, it's unfeasible to generate N-body simulation suites for a
Sizhe Li, Nicolas Christianson, Tongxin Li
Algorithms with predictions} has emerged as a powerful framework to combine the robustness of traditional online algorithms with the data-driven performance benefits of machine-learned (ML) predictions. However, most existing approaches in this paradigm are overly conservative, {as they do not leverage problem structure to optimize performance in a predictio
Jerónimo Duarte, Ignacio García-Mata, Diego A. Wisniacki
The out-of-time-order correlator (OTOC) quantifies information scrambling in quantum systems and serves as a key diagnostic of quantum chaos. In one-body systems with a classical counterpart, the relaxation of the OTOC is governed by Ruelle-Pollicott resonances. For many-body systems lacking a semiclassical limit, recent studies have identified an analogous
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
cs.CVLogan Lawrence, Oindrila Saha, Megan Wei, Chen Sun
Despite the renewed interest in zero-shot visual classification due to the rise of Multimodal Large Language Models (MLLMs), the problem of evaluating free-form responses of auto-regressive models remains a persistent challenge. Most existing works focus on language-only tasks or don't consider Multiple Choice Questions (MCQs) beyond 5-way options, both of w
Sarah Liaw, Benjamin Plaut
In high-stakes AI applications, even a single action can cause irreparable damage. However, nearly all of sequential decision-making theory assumes that all errors are recoverable (e.g., by bounding rewards). Standard bandit algorithms that explore aggressively may cause irreparable damage when this assumption fails. Some prior work avoids irreparable errors
J. Mateos Guilarte
Low energy dynamics of Kinks and Kink-AntiKink configurations in the Jackiw-Rebbi model is fully described. The strategy is based in the Collective Coordinates adiabatic approach. The necessary solution of Quantum Mechanical spectral problems, both for scalar and spinorial wave functions, is unveiled as an intermediate step.
ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
cs.CVKeli Liu, Zhendong Wang, Wengang Zhou, Shaodong Xu
Text-to-image generation with visual autoregressive~(VAR) models has recently achieved impressive advances in generation fidelity and inference efficiency. While control mechanisms have been explored for diffusion models, enabling precise and flexible control within VAR paradigm remains underexplored. To bridge this critical gap, in this paper, we introduce