December 2025 arXiv papers — page 101
Showing 10,001–10,100 of 21,731 papers
Sungnyun Kim
The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference. This dissertation posits that a systematic, hierarchical approach is essential to overcome these challenges, achieving
Siran Liu, Zane Cao, Yongchao He
Efficient long-context understanding and reasoning are increasingly vital for large language model (LLM) applications such as multi-turn dialogue and program analysis. However, the core self-attention mechanism scales quadratically with sequence length, creating a fundamental computational bottleneck. Existing sparse attention methods alleviate this issue bu
Yang Su, Xin Liu, Shiyu Zhang, Ji Yang
The origin of the multiphase gas within the Fermi/eROSITA bubbles is crucial for understanding Galactic Center (GC) feedback. We use HI4PI data to investigate the kinematics and physical properties of high-velocity clouds (HVCs) toward the GC. Our results reveal that the HVCs exhibit a distinct asymmetric distribution, closely associated with the bar-driven
Justin Jung
We propose discrete diffusion guidance for constraint satisfaction problems (CSPs) and demonstrate its ability to solve Sudoku puzzles without supervision.
Wentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica
Mixture of Experts (MoE) models have emerged as the de facto architecture for scaling up language models without significantly increasing the computational cost. Recent MoE models demonstrate a clear trend towards high expert granularity (smaller expert intermediate dimension) and higher sparsity (constant number of activated experts with a higher number of
Igor Halperin
This volume, \textbf{Physicists Are Still Joking}, serves as a definitive almanac of scientific humor spanning sixty years. It traces the evolution of professional folklore across geopolitical divides and technological eras. \textbf{Part I} restores the classic 1966 anthology \textbf{Physicists Joke}, which originally served as a window for Soviet scientists
Mayank Singh, Vikas Yadav, Shiva Krishna Reddy Malay, Shravan Nayak
Automatic search for Multi-Agent Systems has recently emerged as a key focus in agentic AI research. Several prior approaches have relied on LLM-based free-form search over the code space. In this work, we propose a more structured framework that explores the same space through a fixed set of simple, composable components. We show that, despite lacking the g
Da Zhang, Bingyu Li, Zhiyuan Zhao, Feiping Nie
Time series analysis plays a vital role in fields such as finance, healthcare, industry, and meteorology, underpinning key tasks including classification, forecasting, and anomaly detection. Although deep learning models have achieved remarkable progress in these areas in recent years, constructing an efficient, multi-task compatible, and generalizable unifi
Transcendence and algebraic independence of a family of $p$-adic valuation generating functions
math.NTKelvin Lam
We show that $T_p(z)=\prod_{j=1}^{\infty}(1-z^{p^{j}})^{-1/p^{j}}$ is transcendental over $\overline{\mathbb{Q}}(z)$, and establish the transcendence of its values at nonzero algebraic points inside the unit disk. Furthermore, we obtain an algebraic independence result for multiplicatively independent algebraic arguments. In summary, this paper extends Mahle
Xun Cai, Xiucheng Yang, Peter A. Raymond
Wetlands are significant carbon sinks, yet methane emissions partially offset this function due to its high global warming potential. Coastal tidal wetlands, unlike non-tidal wetlands, are regulated by oceanic drivers like salinity gradients and tidal inundation, which strongly influence methane production and release but remain poorly represented in regiona
Correlation functions at the topological quantum phase transition in the S=1 XXZ chain with single-ion anisotropy
cond-mat.str-elToshiya Hikihara, Akira Furusaki
We study the one-dimensional S=1 XXZ spin model with single-ion anisotropy. It is known that at the transition points between the Haldane and large-D phases, the model exhibits a quantum criticality described by the Gaussian theory, i.e., a conformal field theory with the central charge c=1. Using the bosonization approach, we investigate various correlation
Xiumei Li, Xiaotong Sun, Min Sha
In this paper, for an odd prime power $q$, we extend the construction of Xie et al. \cite{XOYM2023} to propose two classes of linear codes $\mathcal{C}_{Q}$ and $\mathcal{C}_{Q}'$ over the finite field $\mathbb{F}_{q}$ with at most four nonzero weights. These codes are derived from quadratic forms through a bivariate construction. We completely determine the
Shun Okumura, Moritz M. Hirschmann, Yukitoshi Motome
Spin spirals represent a fundamental class of noncollinear yet coplanar magnetic structures that give rise to diverse emergent phenomena reflecting spin chirality. We investigate metallic systems hosting commensurate spin spirals and uncover an unconventional anomalous Hall effect (AHE) induced by spiral magnetism. The spin spiral introduces odd-parity spin
From Obfuscated to Obvious: A Comprehensive JavaScript Deobfuscation Tool for Security Analysis
cs.CRDongchao Zhou, Lingyun Ying, Huajun Chai, Dongbin Wang
JavaScript's widespread adoption has made it an attractive target for malicious attackers who employ sophisticated obfuscation techniques to conceal harmful code. Current deobfuscation tools suffer from critical limitations that severely restrict their practical effectiveness. Existing tools struggle with diverse input formats, address only specific obfuscat
Junjie Ma, Jinlong Li, Jiajun Luo
Inference with modern Large Language Models (LLMs) is expensive and slow, and speculative sampling has emerged as an effective solution to this problem. However, the number of calls to the draft model for generating candidate tokens in speculative sampling is a preset hyperparameter, lacking flexibility. To generate and utilize the candidate tokens more effe
Shuang Cheng, Yuhua Jiang, Zineng Zhou, Dawei Liu
Block-wise discrete diffusion offers an attractive balance between parallel generation and causal dependency modeling, making it a promising backbone for vision-language modeling. However, its practical adoption has been limited by high training cost, slow convergence, and instability, which have so far kept it behind strong autoregressive (AR) baselines. We
Yonggan Fu, Lexington Whalen, Zhifan Ye, Xin Dong
Diffusion language models (dLMs) have emerged as a promising paradigm that enables parallel, non-autoregressive generation, but their learning efficiency lags behind that of autoregressive (AR) language models when trained from scratch. To this end, we study AR-to-dLM conversion to transform pretrained AR models into efficient dLMs that excel in speed while
Saiyang Zhang, Boyuan Liu, Volker Bromm, Florian Kühnel
The James Webb Space Telescope (JWST) has recently identified Abell 2744-QSO1 as a compact, metal-poor, black hole (BH) dominated galaxy at $z\simeq 7$. This system exhibits an extreme black-hole-to-stellar mass ratio and unusually low metallicity, posing significant challenges to BH seeding models. Motivated by these discoveries, we perform high-resolution
Integrability Breaking and Coherent Dynamics in Hermitian and Non-Hermitian Spin Chains with Long-Range Coupling
quant-phY. S. Liu, X. Z. Zhang
Unraveling the mechanisms of ergodicity breaking in complex quantum systems is a central pursuit in nonequilibrium physics. In this work, we investigate a one-dimensional spin model featuring a tunable long-range hopping term, $H_{n}$, which introduces nonlocal interactions and bridges the gap between Hermitian and non-Hermitian regimes. Through a systematic
Yi Hu, Cai Zhou, Muhan Zhang
The scaling of large language models (LLMs) emphasizes increasing depth, yet performance gains diminish with added layers. Prior work introduces the concept of "effective depth", arguing that deeper models fail to fully utilize their layers for meaningful computation. Building on this, we systematically study how effective depth varies with model scale, trai
Alessandro Casadei, Sreyoshi Bhaduri, Rohit Malshe, Pavan Mullapudi
Modern operational systems ranging from logistics and cloud infrastructure to industrial IoT, are governed by complex, interdependent processes. Understanding how interventions propagate through such systems requires causal inference methods that go beyond direct effects to quantify mediated pathways. Traditional mediation analysis, while effective in simple
Jiro Soda, Maki Takeuchi
We study the scattering of gravitational waves by axion domain walls in teleparallel gravity with the Nieh-Yan term. Since a domain wall causes the parity violation, the transmitted gravitational waves also exhibit the parity violation. We calculate the degree of circular polarization of gravitational waves. It turns out that gravitational waves after going
Matjaž Omladič, Martin Vuk, Aljaž Zalar
Copulas are the primary tool for dependence modeling in statistics, and quasi-copulas are their essential companions. The latter appear, say, as infima or suprema of sets of copulas; they form a huge class and have some unpleasant properties. Their statistical interpretation is challenged by the fact that they may lead to negative volumes of some boxes. So,
Hao Chen, Junyang Chen, Jinshan Pan, Jiangxin Dong
Recent diffusion-based one-step methods have shown remarkable progress in the field of image super-resolution, yet they remain constrained by three critical limitations: (1) inferior fidelity performance caused by the information loss from compression encoding of low-quality (LQ) inputs; (2) insufficient region-discriminative activation of generative priors;
First-order general constitutive equations for relativistic fluids using the projection method in the Chapman-Enskog expansion of the Boltzmann equation
gr-qcA. L. García-Perciante, A. R. Méndez, O. Sarbach
The first-order out of equilibrium correction to the distribution function, obtained by implementing the projection method for the perturbed relativistic Boltzmann equation using the Chapman-Enskog method, is generalized in order to explicitly include the freedom of choice for frame and representation. It is shown how this procedure leads to general constitu
Hiroyuki Fuji, Masahide Manabe, Yoshiyuki Watabiki
We propose a string field Hamiltonian formalism that associates a class of spectral curves and provides their quantization through the Chekhov-Eynard-Orantin topological recursion. As illustrative examples, we present Hamiltonians for the $(2,2m-1)$ minimal discrete and continuum dynamical triangulation (DT) models, the supersymmetric analogue of minimal con
Real-time prediction of workplane illuminance distribution for daylight-linked controls using non-intrusive multimodal deep learning
cs.CVZulin Zhuang, Yu Bian
Daylight-linked controls (DLCs) have significant potential for energy savings in buildings, especially when abundant daylight is available and indoor workplane illuminance can be accurately predicted in real time. Most existing studies on indoor daylight predictions were developed and tested for static scenes. This study proposes a multimodal deep learning f
Context Representation via Action-Free Transformer encoder-decoder for Meta Reinforcement Learning
cs.ROAmir M. Soufi Enayati, Homayoun Honari, Homayoun Najjaran
Reinforcement learning (RL) enables robots to operate in uncertain environments, but standard approaches often struggle with poor generalization to unseen tasks. Context-adaptive meta reinforcement learning addresses these limitations by conditioning on the task representation, yet they mostly rely on complete action information in the experience making task
Kim Sung-Bin, Joohyun Chang, David Harwath, Tae-Hyun Oh
Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying formulation, speech-conditional facial motion infilling. We explore facial motion infilling as a self-supervised pretext task that also serves as a unifying formulation of dynamic tal
Nicholas Lawson, William Bialek
It often is emphasized that gene expression is noisy. A seemingly contradictory view is that control mechanisms have been optimized to squeeze as much information as possible out of a limited number of molecules. Here we revisit these issues in a simple model where a single transcription factor (TF) controls a large number of target genes. We include only th
Humaira Tasnim, Ashik E Rasul, Bruce Jo, Hyung-Jin Yoon
Reliable helipad detection is essential for Autonomous Aerial Vehicle (AAV) landing, especially under GPS-denied or visually degraded conditions. While modern detectors such as YOLOv8 offer strong baseline performance, single-model pipelines struggle to remain robust across the extreme scale transitions that occur during descent, where helipads appear small
Sa Wang, Shuang Li, Jin-Wen Kang, Ben-Wei Zhang
Jet substructure provides a powerful probe of partonic interactions within the quark-gluon plasma (QGP) in heavy-ion collisions. In this paper, we present a systematic theoretical study of the groomed substructures for both inclusive jets and photon-tagged jets ($\gamma+$jets) utilizing the Dynamical and Soft-Drop Grooming algorithms in PbPb collisions by em
HyperAI Team, Yuchen Liu, Kaiyang Han, Zhiqiang Xia
Current multimodal large lanauge models possess strong perceptual and reasoning capabilities, however high computational and memory requirements make them difficult to deploy directly on on-device environments. While small-parameter models are progressively endowed with strong general capabilities, standard Vision Transformer (ViT) encoders remain a critical
Mengzhang Cai, Xin Gao, Yu Li, Honglin Lin
The rapid evolution of Large Language Models (LLMs) is predicated on the quality and diversity of post-training datasets. However, a critical dichotomy persists: while models are rigorously benchmarked, the data fueling them remains a black box--characterized by opaque composition, uncertain provenance, and a lack of systematic evaluation. This opacity hinde
Wenjun Liu, Qian Wu, Yifeng Hu, Yuke Li
We introduce SELECT (Scene tExt Label Errors deteCTion), a novel approach that leverages multi-modal training to detect label errors in real-world scene text datasets. Utilizing an image-text encoder and a character-level tokenizer, SELECT addresses the issues of variable-length sequence labels, label sequence misalignment, and character-level errors, outper
The influence of surface tension in thin-film hydrodynamics: gravity free planar hydraulic jumps
physics.flu-dynRajesh Kumar Bhagat
Hydraulic jumps in thin films are traditionally explained through gravity-driven shallow-water theory, with surface tension assumed to play only a secondary role via Laplace pressure. Recent experiments, however, suggest that surface tension can be the primary mechanism. In this work we develop a theoretical framework for surface tension driven hydraulic jum
Shen Li, Li Huang, Shaoxiong Zhan, Weifeng Sun
Large language models (LLMs) exhibit strong generative capabilities and have shown great potential in code generation. Existing chain-of-thought (CoT) prompting methods enhance model reasoning by eliciting intermediate steps, but suffer from two major limitations: First, their uniform application tends to induce overthinking on simple tasks. Second, they lac
Kaike Zhang, Qi Cao, Fei Sun, Xinran Liu
Sequential recommender systems have demonstrated strong capabilities in modeling users' dynamic preferences and capturing item transition patterns. However, real-world user behaviors are often noisy due to factors such as human errors, uncertainty, and behavioral ambiguity, which can lead to degraded recommendation performance. To address this issue, recent
Boyang Li, Zhongpeng Jin, Shuai Zhao, Jiahui Liao
The ability to adapt to changing environments is crucial for the autonomous navigation systems of Unmanned Aerial Vehicles (UAVs). However, existing navigation systems adopt fixed execution configurations without considering environmental dynamics based on available computing resources, e.g., with a high execution frequency and task workload. This static app
Omar Abusabha, Jiyong Uhm, Tamer Abuhmed, Hyungjoon Koo
A function inlining optimization is a widely used transformation in modern compilers, which replaces a call site with the callee's body in need. While this transformation improves performance, it significantly impacts static features such as machine instructions and control flow graphs, which are crucial to binary analysis. Yet, despite its broad impact, the
OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
cs.CVZhenguo Zhang, Haohan Zheng, Yishen Wang, Le Xu
The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability failures, most notably object hallucination. This failure stems from their reliance on ungrounded, text-based Chain-of-Thought (CoT) reasoning. While existing multi-modal CoT approaches attempt mitigation, they suffer
Enhong Liu, Haiyu Yang, Miel Hostens
Large Language Models (LLM) hold potential to support dairy scholars and farmers by supporting decision-making and broadening access to knowledge for stakeholders with limited technical expertise. However, the substantial computational demand restricts access to LLM almost exclusively through cloud-based service, which makes LLM-based decision support tools
Dynamic stacking ensemble learning with investor knowledge representations for stock market index prediction based on multi-source financial data
cs.CERuize Gao, Mei Yang, Yu Wang, Shaoze Cui
The patterns of different financial data sources vary substantially, and accordingly, investors exhibit heterogeneous cognition behavior in information processing. To capture different patterns, we propose a novel approach called the two-stage dynamic stacking ensemble model based on investor knowledge representations, which aims to effectively extract and i
Mingjia Yin, Junwei Pan, Hao Wang, Ximei Wang
Click-Through Rate (CTR) prediction, a core task in recommendation systems, aims to estimate the probability of users clicking on items. Existing models predominantly follow a discriminative paradigm, which relies heavily on explicit interactions between raw ID embeddings. However, this paradigm inherently renders them susceptible to two critical issues: emb
Boran Wang, Xinming Wang, Yi Chen, Xiang Li
With their high information density and intuitive readability, charts have become the de facto medium for data analysis and communication across disciplines. Recent multimodal large language models (MLLMs) have made notable progress in automated chart understanding, yet they remain heavily dependent on explicit textual annotations and the performance degrade
ASAP-Textured Gaussians: Enhancing Textured Gaussians with Adaptive Sampling and Anisotropic Parameterization
cs.CVMeng Wei, Cheng Zhang, Jianmin Zheng, Hamid Rezatofighi
Recent advances have equipped 3D Gaussian Splatting with texture parameterizations to capture spatially varying attributes, improving the performance of both appearance modeling and downstream tasks. However, the added texture parameters introduce significant memory efficiency challenges. Rather than proposing new texture formulations, we take a step back to
Martin R. Bridson, Timothy R. Riley
We exhibit novel geometric phenomena in the study of conjugacy problems for discrete groups. We prove that the snowflake groups $B_{pq}$, indexed by pairs of positive integers $p>q$, have conjugator length functions $\text{CL}(n)\simeq n$ and annular Dehn functions $\text{Ann}(n) \simeq n^{2\alpha}$, where $\alpha = \log_2(2p/q)$. Then, building on $B_{pq}$,
Yifan Shao, Peilin Zhou, Shoujin Wang, Weizhi Zhang
Inspired by advances in LLMs, reasoning-enhanced sequential recommendation performs multi-step deliberation before making final predictions, unlocking greater potential for capturing user preferences. However, current methods are constrained by static reasoning trajectories that are ill-suited for the diverse complexity of user behaviors. They suffer from tw
Confinement-Induced Nonlocality and Optical Nonlinearity of Transdimensional Titanium Nitride in the Epsilon-Near-Zero Region
physics.opticsFan-Ting Tseng, I-Hung Ho, Ting-Jui Kuo, Shangjr Gwo
Ultrathin plasmonic films that approach the trans-dimensional (TD) thickness limit provide a promising route for light_matter interaction control and manipulation, yet their nonlinear optical response near the epsilon_near_zero (ENZ) condition remains poorly understood. Here, we report the strongly enhanced optical nonlinearity for their typical representati
Yifan Shao, Peilin Zhou
Sequential recommendation systems aim to capture users' evolving preferences from their interaction histories. Recent reasoningenhanced methods have shown promise by introducing deliberate, chain-of-thought-like processes with intermediate reasoning steps. However, these methods rely solely on the next target item as supervision, leading to two critical issu
Jiajun Li, Yuekun Heng, Jiajie Ling, Zhi Wu
The Filling, Overflow, and Circulation (FOC) system is a critical subsystem of the Jiangmen Underground Neutrino Observatory (JUNO), responsible for the safe handling of the Liquid Scintillator (LS) and water throughout the detector's commissioning and operational lifetime. This paper details the design and operation of the FOC system, which accomplished the
Ignacio Alzugaray, Marwan Taher, Andrew J. Davison
We present a novel neural RGB-D Simultaneous Localization And Mapping (SLAM) system that learns an implicit map of the scene in real time. For the first time, we explore the use of Scene Coordinate Regression (SCR) as the core implicit map representation in a neural SLAM pipeline, a paradigm that trains a lightweight network to directly map 2D image features
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
cs.AILeo Lu, Jonathan Zhang, Sean Chua, Spencer Kim
Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of large language models (LLMs). While prior work focuses on improving model performance through internal reasoning strategies, little is known about the interchangeability of reasoning across different models. In this work, we explore whether a partially completed reasoni
Sample-Efficient Robot Skill Learning for Construction Tasks: Benchmarking Hierarchical Reinforcement Learning and Vision-Language-Action VLA Model
cs.ROZhaofeng Hu, Hongrui Yu, Vaidhyanathan Chandramouli, Ci-Jyun Liang
This study evaluates two leading approaches for teaching construction robots new skills to understand their applicability for construction automation: a Vision-Language-Action (VLA) model and Reinforcement Learning (RL) methods. The goal is to understand both task performance and the practical effort needed to deploy each approach on real jobs. The authors d
Piotr Jędrzejewicz, Mikołaj Marciniak
We present a formal version of the numbers of vertices, edges, and faces for infinite planar regular triangular meshes of degree r>6. These numbers are defined via Euler summation of sequences obtained from iterated expansions of a convex combinatorial disk. We prove that these formal quantities satisfy the classical Euler formula, providing a combinatorial
Nadia Abdolkhani, Walaa Hamouda
In cognitive Internet of Things (CIoT) networks, efficient spectrum sharing is essential to address increasing wireless demands. This paper presents a novel deep reinforcement learning (DRL)-based approach for joint cooperative caching and spectrum access coordination in CIoT networks, enabling the CIoT agents to collaborate with primary users (PUs) by cachi
Jiaheng Li, Qiyu Dai, Lihan Li, Praneeth Chakravarthula
We consider the problem of active 3D imaging using single-shot structured light systems, which are widely employed in commercial 3D sensing devices such as Apple Face ID and Intel RealSense. Traditional structured light methods typically decode depth correspondences through pixel-domain matching algorithms, resulting in limited robustness under challenging s
Seongjeong Kim
In \cite{Kim} it is shown that for an oriented surface $S_{g}$ of genus $g$ links in $S_{g} \times S^{1}$ can be presented by virtual diagrams with a decoration, called {\em double lines}. In this paper, first we define braids with double lines for links in $S_{g}\times S^{1}$. We denote the group of braids with double lines by $VB_{n}^{dl}$. The Alexander a
Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers
cs.CVYibing Fu, Yunpeng Zhao, Zhitao Zeng, Cheng Chen
Multi-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discriminative representations. However, existing SSL methods for imag
Kou Takahashi
In this paper, we study defining ideals of numerical semigroup rings. Let $H$ be a numerical semigroup with multiplicity $a_0$ and embedding dimension $n$. Assuming $a_0/2+1\leq n$, we prove that the defining ideal of $H$ is determinantal when the set of pseudo-Frobenius numbers forms an arithmetic sequence of length $n-1$. This partly resolves a conjecture
Ian Xu
Randomization-based inference commonly relies on grid search methods to construct confidence intervals by inverting hypothesis tests over a range of parameter values. While straightforward, this approach is computationally intensive and can yield conservative intervals due to discretization. We propose a novel method that exploits the algebraic structure of
Concentration of the truncated variation of fractional Brownian motions of any Hurst index, their $1/H$-variations and local times
math.PRWitold M. Bednorz, Rafał M. Łochowski
We obtain bounds for probabilities of deviations of the truncated variation functional of fractional Brownian motions (fBm) of any Hurst index $H \in (0,1)$ from their expected values. Obtained bounds are optimal for large values of deviations up to multiplicative constants depending on the parameter $H$ only. As an application, we give tight bounds for tail
Afia Maham, Dur E Nayab Tashfa
This paper provides a review of deep learning applications in scene understanding in autonomous robots, including innovations in object detection, semantic and instance segmentation, depth estimation, 3D reconstruction, and visual SLAM. It emphasizes how these techniques address limitations of traditional geometric models, improve depth perception in real ti
Juseung Yun, Sunwoo Yu, Sumin Ha, Jonghyun Kim
Cancer progression arises from interactions across multiple biological layers, especially beyond morphological and across molecular layers that remain invisible to image-only models. To capture this broader biological landscape, we present EXAONE Path 2.5, a pathology foundation model that jointly models histologic, genomic, epigenetic and transcriptomic mod
Jiuding Yang, Shengyao Lu, Hongxuan Liu, Shayan Shirahmad Gale Bagi
Large language models (LLMs) have achieved remarkable progress in automatic code generation, yet their ability to produce high-performance code remains limited--a critical requirement in real-world software systems. We argue that current LLMs struggle not only due to data scarcity but, more importantly, because they lack supervision that guides interpretable
Zongyao Li, Kengo Ishida, Satoshi Yamazaki, Xiaotong Ji
We propose KFS-Bench, the first benchmark for key frame sampling in long video question answering (QA), featuring multi-scene annotations to enable direct and robust evaluation of sampling strategies. Key frame sampling is crucial for efficient long-form video understanding. In long video QA, selecting informative frames enables multimodal large language mod
Wenjie Fu, Zhifei Zhu
We study the smallest area $A(M,g)$ of a 2-dimensional stationary integral varifold in a closed Einstein 4-manifold $(M^4,g)$ with $Ric_g = \lambda g, |\lambda|\leq 3, Vol(M,g)\geq v>0, diam(M,g)\leq D, H_1(M;\mathbb{Z})=0.$ Building on the previous work on homological filling functions, we show that for every $(M^4,g)$ in this Einstein class, there is an up
Frozen Gaussian sampling algorithms for simulating Markovian open quantum systems in the semiclassical regime
quant-phLimin Xu, Zhen Huang, Zhennan Zhou
Simulating Markovian open quantum systems in the semiclassical regime poses a grand challenge for computational physics, as the highly oscillatory nature of the dynamics imposes prohibitive resolution requirements on traditional grid-based methods. To overcome this barrier, this paper introduces an efficient Frozen Gaussian Sampling (FGS) algorithm based on
Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato
World models have shown great utility in improving the task performance of embodied agents. While prior work largely focuses on pixel-space world models, these approaches face practical limitations in GUI settings, where predicting complex visual elements in future states is often difficult. In this work, we explore an alternative formulation of world modeli
Hierarchical Deep Reinforcement Learning for Robust Access in Cognitive IoT Networks under Smart Jamming Attacks
eess.SPNadia Abdolkhani, Walaa Hamouda
In this paper, we address the challenge of dynamic spectrum access in a cognitive Internet of Things (CIoT) network where a secondary user (SU) operates under both energy constraints and adversarial interference from a smart jammer. The SU coexists with primary users (PUs) and must ensure that its transmissions do not exceed a predefined interference thresho
Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia
The rise of AI agents is transforming how software can be built. The promise of agents is that developers might write code quicker, delegate multiple tasks to different agents, and even write a full piece of software purely out of natural language. In reality, what roles agents play in professional software development remains in question. This paper investi
Yue Wan, Jiayi Yuan, Zhiwei Feng, Xiaowei Jia
Antigenic epitope presented by major histocompatibility complex II (MHC-II) proteins plays an essential role in immunotherapy. However, compared to the more widely studied MHC-I in computational immunotherapy, the study of MHC-II antigenic epitope poses significantly more challenges due to its complex binding specificity and ambiguous motif patterns. Consequ
Che-Chia Chang, Te-Sheng Lin, Ming-Chih Lai
The Stefan problem is a classical free-boundary problem that models phase-change processes and poses computational challenges due to its moving interface and nonlinear temperature-phase coupling. In this work, we develop a physics-informed neural network framework for solving two-phase Stefan problems. The proposed method explicitly tracks the interface moti
Aurélio Menegon
We develop a real-analytic framework, called perplex analysis, in which the complex, split-complex, and dual numbers arise as members of a single four-parameter family of two-dimensional commutative real algebras. Within this unified setting we define differentiability through a generalized Cauchy-Riemann structure, extending several features of complex geom
Denis Potapov, Fedor Sukochev, Anna Tomskova, Dmitriy Zanin
We extend the classical Fuglede commutativity theorem to the full scale of symmetrically normed operator ideals. Our main result provides a complete characterization: a symmetric ideal or symmetric operator space of $\tau$-measurable operators satisfies the Fuglede theorem if and only if its commutative core has non-trivial Boyd indices, or equivalently, if
Bitao Shen, Huajin Chang, Junhao Han, Yimeng Wang
Microcavity optical frequency combs (microcombs) are compact, coherent light sources whose chip-scale integrability is poised to drive advances in metrology, communications, and sensing. Among available microcomb generation methods, hybrid cavities uniquely co-locate gain and Kerr dynamics, where the lasing mode directly resonates in the nonlinear microcavit
Quantifying electron-nuclear spin entanglement dynamics in central-spin systems using one-tangles
quant-phIsabela Gnasso, Khadija Sarguroh, Dorian Gangloff, Sophia E. Economou
Optically-active solid-state systems such as self-assembled quantum dots, rare-earth ions, and color centers in diamond and SiC are promising candidates for quantum network, computing, and sensing applications. Although the nuclei in these systems naturally lead to electron spin decoherence, they can be repurposed, if they are controllable, as long-lived qua
Electrified EHL line contact with dielectric breakdown of lubricant -- a numerical model
physics.app-phYang Xu, Nick Morris, Yue Wu
With the rapid growth of the electric vehicles with drive systems with higher voltages, power outputs, frequencies, and speeds, mitigating electrically induced bearing damage (EIBD) in electric motors has become critical. In this study, a novel numerical model characterizing discharge-induced current density and voltage drop at the elastohydrodynamic lubrica
Chuanchao Gao, Arvind Easwaran
Vehicular Edge Computing (VEC) has emerged as a promising paradigm for enhancing the computational efficiency and service quality in intelligent transportation systems by enabling vehicles to wirelessly offload computation-intensive tasks to nearby Roadside Units. However, efficient task offloading and resource allocation for time-critical applications in VE
Zhuo Zhang, Yonghui Liu, Meijie Zhang, Feiyang Tan
In this paper, we unleash the potential of the powerful monodepth model in camera-LiDAR calibration and propose CLAIM, a novel method of aligning data from the camera and LiDAR. Given the initial guess and pairs of images and LiDAR point clouds, CLAIM utilizes a coarse-to-fine searching method to find the optimal transformation minimizing a patched Pearson c
Zheng He, Roman Pogodin, Yazhe Li, Namrata Deka
Tests of conditional independence (CI) underpin a number of important problems in machine learning and statistics, from causal discovery to evaluation of predictor fairness and out-of-distribution robustness. Shah and Peters (2020) showed that, contrary to the unconditional case, no universally finite-sample valid test can ever achieve nontrivial power. Whil
Arohee Bhoja
Vizing's theorem states that every simple undirected graph can be edge-colored using fewer than $\Delta + 1$ colors, where $\Delta$ is the graph's maximum degree. The original proof was given through a polynomial-time algorithmic procedure that iteratively extends a partial coloring until it becomes complete. In this work, I used the Lean theorem prover to p
Adamu Issifu, Julio C. M. Rocha, Francisco A. Brito, Tobias Frederico
We develop a unified framework in which the dynamics of a scalar glueball field, originating from phenomenological nonperturbative QCD confinement, simultaneously governs the deconfinement transition of strongly interacting matter and drives cosmological inflation. Starting from a temperature-dependent effective potential $V_{eff}(\phi, T)$, we show that the
Alisha Ukani, Katherine Izhikevich, Shambhavi Mittal, Manan Patel
Understanding where Internet services are hosted, and how users reach them, has captured the interest of government regulators and others concerned with the privacy of data flows. In this paper we focus on government websites -- services which arguably merit a higher expectation of protection against foreign surveillance or interference -- and seek to identi
Jin Sob Kim, Hyun Joon Park, Wooseok Shin, Sung Won Han
This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their position. Our system utilizes a ResNet-based model under a proposed robust framework for SELD. To guarantee the generali
Jyotishka Datta, Nick Polson, Vadim Sokolov
We propose a unified framework for global-local regularization that bridges the gap between classical techniques -- such as ridge regression and the nonnegative garotte -- and modern Bayesian hierarchical modeling. By estimating local regularization strengths via marginal likelihood under order constraints, our approach generalizes Stein's positive-part esti
Yao He, Youngjoong Kwon, Tiange Xiang, Wenxiao Cai
We present a framework that adapts 2D diffusion models for 3D shape completion from incomplete point clouds. While text-to-image diffusion models have achieved remarkable success with abundant 2D data, 3D diffusion models lag due to the scarcity of high-quality 3D datasets and a persistent modality gap between 3D inputs and 2D latent spaces. To overcome thes
Hiroshige Shiga
For each $\omega\in (0, 1)^{\mathbb N}$, we may construct a Cantor set $E(\omega)\subset [0, 1]$ called a generalized Cantor set for $\omega$. We study the moduli space of $\omega$ denoted by $\mathcal M(\omega)\subset (0, 1)^{\mathbb N}$. It is the set of $\omega'$ so that $E(\omega')$ is quasiconformally equivalent to $E(\omega)$. In this paper, we show th
Cindy Y. Zhang, Elif Ertekin, Peter Orbanz, Ryan P. Adams
Incorporating known symmetries in data into machine learning models has consistently improved predictive accuracy, robustness, and generalization. However, achieving exact invariance to specific symmetries typically requires designing bespoke architectures for each group, limiting scalability and preventing knowledge transfer across related symmetries. In th
First-return statistics in bounded radiative transport: A Motzkin polynomial framework
physics.opticsClaude Zeller, Robert Cordery
A photon entering a scattering medium executes a three-dimensional random walk determined by the Henyey-Greenstein phase function. The photon either reaches the boundary for a first passage or is absorbed. Projecting the walk onto the axial direction produces a one-dimensional alternating process whose peaks and valleys correspond to changes in the sign of t
A Generalized Formulation for Accurate and Robust Determination of Soil Shear Strength from Triaxial Tests
physics.geo-phAltamirano-Muñiz Emilio Fernando
This work presents an extended formulation of the Least Squares with Virtual Displacements (LSVD) method for estimating shear strength parameters from multiple soil samples under varying resistance conditions including cohesionless, frictional, and mixed types. LSVD is designed to identify a common tangent across n Mohr circles, even in the presence of measu
Carine Simo, Venceslas Nguefoue Meli, Patrick Louodop, Samuel Bowong
Pancreatic $\beta$-cells play a central role in maintaining glucose homeostasis through the pulsatile secretion of insulin. This essential function relies not only on intracellular regulatory mechanisms but also on coordinated interactions among $\beta$-cells within the islets of Langerhans. Disruptions in this intercellular coordination are increasingly imp
Dereje Shenkut, Vijayakumar Bhagavatula
Multi-agent collaborative perception (CP) is a promising paradigm for improving autonomous driving safety, particularly for vulnerable road users like pedestrians, via robust 3D perception. However, existing CP approaches often optimize for vehicle detection performance metrics, underperforming on smaller, safety-critical objects such as pedestrians, where d
Hossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari
Despite recent advances in robotics and human-robot collaboration in the AEC industry, trust has mostly been treated as a static factor, with little guidance on how it changes across events during collaboration. This paper investigates how a robot's task performance and its expressive responses after outcomes shape the dynamics of human trust over time. To t
Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models
cs.CLZhimin Qiu, Di Wu, Feng Liu, Yuxiao Wang
This paper proposes a structure-aware decoding method based on large language models to address the difficulty of traditional approaches in maintaining both semantic integrity and structural consistency in nested and overlapping entity extraction tasks. The method introduces a candidate span generation mechanism and structured attention modeling to achieve u
Ge Yan, Chung-En Sun, Linbo Liu, Tsui-Wei Weng
Large reasoning models achieve strong performance on diverse tasks by producing extended chains of thought. Self-reflection, the ability to review and revise prior reasoning steps, is widely regarded as a key contributor to this performance. However, self-reflection also incurs substantial inference cost, and its governing mechanism remains underexplored. In
Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms
cs.AIYang Cao, Yubin Chen, Xuyang Guo, Zhao Song
The rapid advancement of large language models (LLMs) has led to significant breakthroughs in automated mathematical reasoning and scientific discovery. Georgiev, G${\'o}$mez-Serrano, Tao, and Wagner [GGSTW+25] demonstrate that AI systems can explore new constructions and improve existing bounds, illustrating the growing potential of LLMs to accelerate mathe
XAI-Driven Diagnosis of Generalization Failure in State-Space Cerebrovascular Segmentation Models: A Case Study on Domain Shift Between RSNA and TopCoW Datasets
cs.CVYoussef Abuzeid, Shimaa El-Bana, Ahmad Al-Kabbany
The clinical deployment of deep learning models in medical imaging is severely hindered by domain shift. This challenge, where a high-performing model fails catastrophically on external datasets, is a critical barrier to trustworthy AI. Addressing this requires moving beyond simple performance metrics toward deeper understanding, making Explainable AI (XAI)
Marc Dambrine, Helmut Harbrecht
The present article is dedicated to the forward and backward solution of a transient one-phase Stefan problem. In the forward problem, we compute the evolution of the initial domain for a Stefan problem where the melting temperature varies over time. This occurs in practice, for example, when the pressure in the external space changes in time. In the corresp
Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
cs.ROHossein Naderi, Alireza Shojaei, Philip Agee, Kereshmeh Afsari
Construction safety inspection remains mostly manual, and automated approaches still rely on task-specific datasets that are hard to maintain in fast-changing construction environments due to frequent retraining. Meanwhile, field inspection with robots still depends on human teleoperation and manual reporting, which are labor-intensive. This paper aims to co