Skip to content

March 2026 arXiv papers — page 37

Showing 3,6013,700 of 25,974 papers

  1. Mitsumasa Wada

    We present a hybrid multi-phase page matching algorithm for automated comparison of Japanese building permit document sets. Building permit review in Japan requires cross-referencing large PDF document sets across revision cycles, a process that is labor-intensive and error-prone when performed manually. The algorithm combines longest common subsequence (LCS

  2. Gradwell Dzikanyanga, Weihao Yang, Hao Huang, Donglei Wu

    Key-value (KV) caching is critical for efficient inference in large language models (LLMs), yet its memory footprint scales linearly with context length, resulting in a severe scalability bottleneck. Existing approaches largely treat KV states as equally important across time, implicitly assuming uniform precision and accessibility. However, this assumption

  3. Jing Zhang, Bastien Bergere, Emilie Bollache, Jonas Leite

    Cardiac MRI late gadolinium enhancement (LGE) enables non-invasive identification of left atrial (LA) scar, whose spatial distribution is strongly associated with atrial fibrillation (AF) severity and recurrence. However, automatic LA scar segmentation remains challenging due to low contrast, annotation variability, and the lack of anatomical constraints, of

  4. Shu Hamanaka

    The non-Hermitian skin effect is an anomalous localization phenomenon induced by nonreciprocal dissipation and has attracted considerable attention in recent years both theoretically and experimentally. In this article, we review the multifractal aspects of the non-Hermitian skin effect. In particular, we discuss how the many-body skin effect exhibits multif

  5. Pan Zhao, Hui Yuan, Chang Sun, Chongzhen Tian

    Existing post-decoding quality enhancement methods for point clouds are designed for static data and typically process each frame independently. As a result, they cannot effectively exploit the spatiotemporal correlations present in point cloud sequences.We propose a unified geometry and attribute enhancement framework (DUGAE) for G-PCC compressed dynamic po

  6. Youngju Na, Jaeseong Yun, Soohyun Ryu, Hyunsu Kim

    While 3D Gaussian splatting has emerged as a powerful paradigm, it fundamentally fails to model transparency such as glass panels. The core challenge lies in decoupling the intertwined radiance contributions from transparent interfaces and the transmitted geometry observed through the glass. We present GLINT, a framework that models scene-scale transparency

  7. Eslam Badr, Takeshi Harui

    We consider smooth plane curves $\mathcal{X}$ of degree $d\geq4$, defined over an algebraically closed field of characteristic $0$, that possess a unique outer Galois point. This geometric condition forces the curve to be a cyclic covering of the projective line, and ensures that its automorphism group fits into a specific theoretical framework. For each pos

  8. Bozhao Li, Shaocong Wu, Tong Shao, Senqiao Yang

    Recent advances in open-vocabulary object detection focus primarily on two aspects: scaling up datasets and leveraging contrastive learning to align language and vision modalities. However, these approaches often neglect internal consistency within a single modality, particularly when background or environmental changes occur. This lack of consistency leads

  9. Jicheng Ma, Yunyan Yang, Juan Zhao, Liang Zhao

    We introduce the Geometric Evolution Graph Convolutional Network (GEGCN), a novel framework that enhances graph representation learning through explicit modeling of geometric evolution on graph structures. Specifically, GEGCN leverages a Long Short-Term Memory (LSTM) network to capture the dynamic structural sequence generated by discrete Ricci flow, and inf

  10. Tong Zhang, Hong Guo, Shuangzhou Yan, Dongkai Weng

    We present FatigueFormer, a semi-end-to-end framework that deliberately combines saliency-guided feature separation with deep temporal modeling to learn interpretable and generalizable muscle fatigue dynamics from surface electromyography (sEMG). Unlike prior approaches that struggle to maintain robustness across varying Maximum Voluntary Contraction (MVC) l

  11. Gilles Wainrib, Barbara Bodinier, Haitem Dakhli, Josep Monserrat

    Recent work has questioned whether large language models (LLMs) can perform genuine in-context learning (ICL) for scientific experimental design, with prior studies suggesting that LLM-based agents exhibit no sensitivity to experimental feedback. We shed new light on this question by carrying out 800 independently replicated experiments on iterative perturba

  12. Chonghuinan Wang, Zihan Chen, Yuxiang Wei, Tianyi Jiang

    Instruction-based multimodal image manipulation has recently made rapid progress. However, existing evaluation methods lack a systematic and human-aligned framework for assessing model performance on complex and creative editing tasks. To address this gap, we propose CREval, a fully automated question-answer (QA)-based evaluation pipeline that overcomes the

  13. Minsun Kim, Dawon Lee, Junyong Noh

    On general video-sharing platforms like YouTube, comments are displayed independently of video playback. As viewers often read comments while watching a video, they may encounter ones referring to moments unrelated to the current scene, which can reveal spoilers and disrupt immersion. To address this problem, we present ComVi, a novel system that displays co

  14. Evans M. Harrell, James B. Kennedy, Gabriel J. Ramos

    We study ratios of eigenvalues of the Laplacian on compact metric graphs. Our goals are threefold: First, we prove a sharp Ashbaugh--Benguria-type bound for the ratio of the first two eigenvalues on compact trees with Dirichlet conditions at all leaves, concretely showing that the ratio is maximized when the graph is an interval or an equilateral star. This

  15. Zefeng Tu, Rushuang Zhao, Hui Liu, Biping Gong

    Using two observations obtained with the Five-hundred-meter Aperture Spherical radio Telescope (FAST), we present a detailed single-pulse analysis of the high-nulling pulsar PSR J1820-0509. We measure an exceptionally high nulling fraction of approximately 81.78%, significantly exceeding previous estimates from Parkes observations. The single-pulse energy di

  16. Sergei Avdonin, Matti Lassas, Jinpeng Lu, Medet Nursultanov

    We study the inverse problem for a semilinear wave equation on metric tree graphs. From the Dirichlet-to-Neumann map defined at all but one of the boundary vertices, we recover unknown connectivity of the graph, lengths of the edges, the time-independent potential and the time-dependent coefficient of the nonlinear term of the equation.

  17. Shubhi Shukla, Pravin Nair

    Image restoration, the recovery of clean images from degraded measurements, has applications in various domains like surveillance, defense, and medical imaging. Despite achieving state-of-the-art (SOTA) restoration performance, existing convolutional and attention-based networks lack stability guarantees under minor shifts in input, exposing a robustness acc

  18. Yi Zhang, Hongbo Huang, Liang-Jie Zhang

    Diffusion models generate high-quality images but pose serious risks like copyright violation and disinformation. Watermarking is a key defense for tracing and authenticating AI-generated content. However, existing methods rely on threshold-based detection, which only supports fuzzy matching and cannot recover structured watermark data bit-exactly, making th

  19. Roberto Vila, Helton Saulo, Felipe Quintino

    We propose a new family of inequality indices that bridges the Hoover index and the Gini coefficient. The measure is defined as the normalized expected absolute value of a convex combination of deviations from the mean and pairwise differences, providing a continuous interpolation between these two classical indices. We establish key theoretical properties,

  20. Hao Liang, Zhengyang Zhao, Meiyi Qiang, Mingrui Chen

    Data-centric training has emerged as a promising direction for improving large language models (LLMs) by optimizing not only model parameters but also the selection, composition, and weighting of training data during optimization. However, existing approaches to data selection, data mixture optimization, and data reweighting are often developed in isolated c

  21. Rui-Qi Cui, Tong Liu

    Stellar-mass black holes (BHs) surrounded by neutrino-dominated accretion flows (NDAFs) are a leading central engine of gamma-ray bursts (GRBs). In this work, we investigate the electron fraction distribution in NDAFs with or without disk outflows for different accretion rates, BH spins, and outflow rates. As the results, for the cases of the massive disks a

  22. Markus Gahn, Tanja Lochner, Malte A. Peter

    Effective interface conditions for a periodically voided thin layer separating two homogeneous bulk regions are derived for the elastic wave equation by taking the simultaneous limit of vanishing layer periodicity and layer thickness. The limit problems are obtained using the unfolding method for thin perforated domains. We consider three different scalings

  23. Renjie Xu, Daowen Qiu, Ligang Xiao, Le Luo

    Solving the discrete logarithm problem (DLP) with quantum computers is a fundamental task with important implications. Beyond Shor's algorithm, many researchers have proposed alternative solutions in recent years. However, due to current hardware limitations, the scale of DLP instances that can be addressed by quantum computers remains insufficient. To overc

  24. Yuntao Shou, Jun Zhou, Tao Meng, Wei Ai

    Multimodal Emotion Recognition in Conversations (MERC) aims to predict speakers' emotional states in multi-turn dialogues through text, audio, and visual cues. In real-world settings, conversation scenarios differ significantly in speakers, topics, styles, and noise levels. Existing MERC methods generally neglect these cross-scenario variations, limiting the

  25. Wenhui Chen, Yan Liu, Manqing Luo

    This manuscript considers the Jordan-Moore-Gibson-Thompson (JMGT) equation and its linearized equation with an additional weak damping term (proposed by [B. Kaltenbacher, \emph{Inverse Problems} (2025)] firstly) in the whole space $\mathbb{R}^n$. We mainly study the unique existence and large time behavior, including optimal decay estimates and asymptotic pr

  26. Hiroyasu Izeki, Ran Ji, Anders Karlsson, Yunhui Wu

    We prove that finitely generated amenable groups acting on CAT(0) spaces satisfy the following alternative: either every action on a geodesically complete CAT(0) space with bounded geometry (or finite dimension) has a global fixed point, or the group admits a fixed-point-free action on $\mathbb{R}^n$. As a consequence, finitely generated amenable torsion gro

  27. Margherita Disertori, Javier Durán Fernández, Luca Fresta

    We consider a family of nonlinear sigma models on $\mathbb{Z}^{d}$ whose target space is the hyperbolic super manifold $H^{2|2n}$, $n >1$, introduced by Crawford as an extension of Zirnbauer's $H^{2|2}$ model for disordered systems. We prove exponential decay of the two-point correlation function in the high-temperature regime $\beta \leq C n^{-1}$, with $C>

  28. Vihang Jumle

    Framing continues to remain one of the most extensively applied theories in political communication. Developments in computation, particularly with the introduction of transformer architecture and more so with large language models (LLMs), have naturally prompted scholars to explore various novel computational approaches, especially for deductive frame detec

  29. Moritz Landwehr, Patrick Hoher, Johannes Reuter

    Existing approaches for battery health forecasting often rely on extensive cycling histories and continuously monitored cells. In contrast, many real-world scenarios provide only sparse information, e.g. a single diagnostic cycle. In our study, we investigate state of health (SoH)- and remaining useful life (RUL) estimation of previously unseen lithium-ion c

  30. Bang Huang, Shunyuan Shang, Mohamed-Slim Alouini

    This paper proposes a movable-antenna-based index modulation (MA-IM) framework that exploits the spatial mobility of a single reconfigurable antenna to create additional information-bearing dimensions for next-generation wireless systems. By discretizing the continuous movable region into a dense set of candidate sampling points and selecting representative

  31. Yu Zhai, Youhao Shang, Jian Liu

    The geometric phase (GP) is a fundamental quantum effect arising from conical intersections (CIs), with profound consequences for vibronic energy levels. Standard imaginary-time path integral molecular dynamics (PIMD) based on the Born-Oppenheimer approximation does not account for the GP, potentially leading to significant errors in low-temperature thermody

  32. Fuga Kobayashi, Takumi Takahashi, Shinsuke Ibi, Takanobu Doi

    This paper proposes a novel modulation and coding scheme (MCS) selection framework that integrates mutual information (MI) prediction based on vector similarity search (VSS) for massive multi-user multiple-input multiple-output orthogonal frequency-division multiplexing (MU-MIMO-OFDM) systems with advanced uplink multi-user detection (MUD). The framework per

  33. Yucheng Liu, Tak Shing Au Yeung, Eric T. Chung, Simon See

    Efficient simulation of Darcy flow in highly heterogeneous porous media requires iterative solvers that remain robust under large permeability contrasts and mixed boundary conditions. Spectral coarse spaces in two-level overlapping Schwarz methods provide such robustness, but their practical use is often limited by an expensive setup phase dominated by many

  34. Liyan Song, Qingchun Li, Chengyuan Qu

    This paper studies a fractional attraction-repulsion system with generalized logistic source and nonlinear productions: \begin{equation*} \left\{ \begin{aligned} &u_t = -(-\Delta)^\alpha u - \chi_1 \nabla \cdot (u \nabla v) + \chi_2 \nabla \cdot (u \nabla w) + au - bu^\gamma, &x \in \mathbb{R}^N, \, t > 0, \\ &0 = \Delta v - \lambda_1 v + \mu_1 u^k, &x \in \

  35. Akram Ben Ahmed, Takahiro Hirofuchi, Takaaki Fukai

    The rapid emergence of edge computing platforms and large-scale data centers has made power efficiency a primary design constraint, particularly for data-intensive and AI-driven workloads. Field-programmable gate arrays (FPGAs) are increasingly adopted due to their flexibility and potential for energy-efficient acceleration. However, FPGA supply voltages are

  36. Shuhei Tsuyuki, Reda Bensaid, Jérémy Morlier, Mathieu Léonardon

    Efficient and adaptable deep learning models are an important area of deep learning research, driven by the need for highly efficient models on edge devices. Few-shot learning enables the use of deep learning models in low-data regimes, a capability that is highly sought after in real-world applications where collecting large annotated datasets is costly or

  37. Jiwen Zhang, Xiangyu Shi, Siyuan Wang, Zerui Li

    Vision-and-Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero-shot navigation. While recent exploration-based zero-shot methods have shown promising results by leveraging global scene priors, they rely on high-quality human-crafted scene reconstructions, which are impractical for real-world robot dep

  38. Yilin Song, Ying Wang, Jiqiang Zheng, Ruihan Zhou

    In this paper, we establish a Paley-Wiener type uncertainty principle for Schr\"odinger equations with bounded electric and magnetic potentials, \begin{align*} i\partial_tu+\Delta_Au+V(t,x)u=0,\,\,u(0,x)=u_0(x), \end{align*} where $\Delta_A=(\nabla-iA)^2$ denotes the magnetic Schr\"odinger operator. Specifically, under suitable assumptions on $A$ and $V$, we

  39. Arpan Bairagi, Rakesh Dey, Siladittya Manna, Umapada Pal

    Remote photoplethysmography (rPPG) allows for the contactless estimation of physiological signals from facial videos by analyzing subtle skin color changes. However, rPPG signals are extremely susceptible to illumination changes, motion, shadows, and specular reflections, resulting in low-quality signals in unconstrained environments. To overcome these issue

  40. Amir Bouziane, Huseyin Arslan

    Standard periodic pilot patterns in orthogonal frequency division multiplexing (OFDM) systems induce severe delay-domain grating lobes, compromising radar sensing. This paper proposes a two-stage framework to design non-periodic pilot patterns that minimize the peak sidelobe level (PSL) while strictly enforcing communication anchor constraints. We black solv

  41. Jiajia Song, Zhihan Guo, Jionghao Lin

    Student simulation can support learning-by-teaching pedagogy where human students (as tutors) teach AI-simulated novice students (as tutees). Recent research often relies on prompt engineering with large language models (LLMs) to simulate novice student behaviour, but it is difficult to keep the AI-simulated student at a stable novice knowledge level. A key

  42. Shane D'Mello, Priya Rani

    We study the space of real rational curves of low degree in the quadric of signature $(3,2)$ and provides a classificaton of real rational knots and nodal curves. Apart from the classification, we also study the relationship between the real rational knots in the quadric and the real rational knots in $\mathbb{RP}^{3}$. Furthermore, a construction for the re

  43. Mostafa Haghir Chehreghani

    Graph Neural Networks (GNNs) face two fundamental challenges when scaled to deep architectures: oversmoothing, where node representations converge to indistinguishable vectors, and oversquashing, where information from distant nodes fails to propagate through bottlenecks. Both phenomena are intimately tied to the underlying graph structure, raising a natural

  44. Victor Deng, Aurèle Barrière, Clément Pit-Claudel

    Despite widespread use, the complexity class of modern regular expression matching was not well-understood. Previous work proved that regular expression matching with backreferences and lookarounds was PSPACE-complete, but the proof was not mechanized and applied to an abstract regex language. This paper clarifies the question for JavaScript regular expressi

  45. Humaira Kousar, Hasnain Irshad Bhatti, Jaekyun Moon

    Efficient data selection is crucial for enhancing the training efficiency of deep neural networks and minimizing annotation requirements. Traditional methods often face high computational costs, limiting their scalability and practical use. We introduce PruneFuse, a novel strategy that leverages pruned networks for data selection and later fuses them with th

  46. Xianpeng, Sun, Haonan Sun, Tian Yu

    Evaluation of repository-aware software engineering systems is often confounded by synthetic task design, prompt leakage, and temporal contamination between repository knowledge and future code changes. We present a time-consistent benchmark methodology that snapshots a repository at time T0, constructs repository-derived code knowledge using only artifacts

  47. Amar Almaini, Jakob Folz, Ghadeer Ashour

    Tiny Machine Learning enables real-time, energy-efficient data processing directly on microcontrollers, making it ideal for Internet of Things sensor networks. This paper presents a compact TinyML pipeline for detecting anomalies in environmental sound within IoT sensor networks. Acoustic monitoring in IoT systems can enhance safety and context awareness, ye

  48. Jintong Hu, Bin Chen, Zhenyu Hu, Jiayue Liu

    Video super-resolution (VSR) seeks to reconstruct high-resolution frames from low-resolution inputs. While diffusion-based methods have substantially improved perceptual quality, extending them to video remains challenging for two reasons: strong generative priors can introduce temporal instability, and multi-frame diffusion pipelines are often too expensive

  49. Rehan Ahmed, Pramod Kumar

    A sensor that can detect the direction of the incoming light plays a crucial role in further enhancing the versatility of the multifunction sensors for future applications, where the sensor can read multiple pieces of information, similar to the biological senses, like skin. A hybrid sensor based on an n-type ZnO micro-rod with p-type optically active organi

  50. Rikuya Ishikawa, Kyohei Takae, Daisuke Takegami, Yoshikazu Mizuguchi

    Multicomponent crystals are often assumed to form nearly random solid solutions when thermodynamically stable. However, crystal growth proceeds from structurally heterogeneous liquids, raising the possibility that the liquid state may influence which species are incorporated into the growing crystal. Here we demonstrate that liquid-state structural asymmetry

  51. Younghoon Ko, Hyemin Park, Hyuk-Jae Lee, Hyokeun Lee

    As the memory channel count is confined by physical dimensions, memory expanders appear to be a promising approach to extending memory capacity and channels by augmenting the existing I/O interface (e.g., PCIe) with memory-semantic protocols like CXL. Unfortunately, the physical constraints of a computing system restrict scalable capacity expansion with memo

  52. Deepak Kumar

    We introduce SWE-PRBench, a benchmark of 350 pull requests with human-annotated ground truth for evaluating AI code review quality. Evaluated against an LLM-as-judge framework validated at kappa=0.75, 8 frontier models detect only 15-31% of human-flagged issues on the diff-only configuration, demonstrating that AI code review remains far below human expert p

  53. Chi-Yeh Chen

    This paper addresses the scheduling problem for unrelated crowd workers in mobile social networks, where the required service time for each task varies among the assigned crowd workers. The goal is to minimize the total weighted completion time of all tasks. First, in an environment with identical crowd workers, we improve the approximation ratio of the Larg

  54. Mridul Khurana, Amin Karimi Monsefi, Justin Lee, Medha Sawhney

    Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the remarkable progress in text-to-image synthesis, existing models often fail to capture the fine-grained visual cues that define species identity, even when their outputs appear photo-re

  55. Samyak Rawlekar, Amitabh Swain, Yujun Cai, Yiwei Wang

    Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in [CLS] token attention maps of the final layer. However, these maps often contain spurious activations resulting in poor localization of objects. This is because the [CLS] token, trained on an image-level objective, summarizes the entire im

  56. Jinda Lu, Junkang Wu, Jinghan Li, Kexin Huang

    Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for multimodal large language models (MLLMs) have mainly focused on improving final answer correctness and strengthening visual grounding. However, a critical bottleneck remains: although models can attend to relevant visual regions, they often fail to effectively incorporate visual evi

  57. Yirun Wang, Yuyang Du, Soung Chang Liew, Yuchen Pan

    Achieving reliable communication has long been a fundamental challenge in networked systems. Semantic Error Correction (SEC) leverages the semantic understanding capabilities of language models (LMs) to perform application-layer error correction, complementing conventional channel decoding. While promising, existing SEC approaches rely solely on context capt

  58. Ardra Muriyankandathil, Parikshit Sahatiya, Brian Abbey, Nitish Kumar Gupta

    Second-harmonic generation in resonant structures is commonly evaluated in terms of intracavity field enhancement at the fundamental and harmonic frequencies. Here, we formulate nonlinear frequency conversion within a symmetry-resolved overlap framework that explicitly separates resonant field buildup from nonlinear mode projection. Using a simple and analyt

  59. Viktor Andersson, Ole Fredrik Brevig, Athanasios Kouroupis

    We present a new elementary proof of a theorem due to Harald Bohr, which states that an unbounded, analytic, and almost periodic function in a half-plane can be written as the sum of two analytic functions: the first is unbounded and periodic, while the second is bounded and almost periodic. The proof is based on a well-known arithmetical property of transla

  60. Zhangtianyi Chen, Yuhao Shen, Florensia Widjaja, Yan Xu

    While recent advancements in Large Language Models have significantly advanced dermatological diagnosis, monolithic LLMs frequently struggle with fine-grained, large-scale multi-class diagnostic tasks and rare skin disease diagnosis owing to training data sparsity, while also lacking the interpretability and traceability essential for clinical reasoning. Alt

  61. Ran Li, Zhong-Xiao Man, Jin Wang

    We investigate the decoherence of an Unruh-DeWitt detector coupled to scalar, electromagnetic, and spinor fields in four-dimensional Minkowski spacetime. By employing the Schwinger-Keldysh influence functional formalism, we derive a universal scaling law relating the decoherence rate to the proper acceleration $a$ and the scaling dimension $\Delta$ of the en

  62. Chul-Hwan Kim, Jeong-Eun Lee, Doug Johnstone, Gregory J. Herczeg

    Angular momentum removal is a fundamental requirement for star and planet formation, yet the mechanisms driving this process remain debated. Magnetohydrodynamic disk winds, launched along magnetic field lines from extended disk regions, offer a promising solution, particularly in regions where magnetorotational turbulence is weak. Here we present high-resolu

  63. Zunwei Fu, Ji Li, Chong-Wei Liang, Wei Wang

    In this paper, we introduce a class of twisted multiparameter singular integrals on $\mathbb{R}^{2m}$, motivated by the Cauchy--Szeg\H{o} projections and the solving operators for $\bar{\partial}_b$ on a broad family of quadratic surfaces of higher codimension in $\mathbb{C}^n$. These surfaces are represented as suitable quotients of products of Heisenberg g

  64. R. Vatré, G. Morettini, J. Beugnon, R. Lopes

    Ultracold Bose gases in one-dimensional optical lattices constitute an important benchmark problem in the study of strongly interacting many-body quantum phases. Here we present a combined experimental and theoretical study of their phase-coherence properties over a wide range of lattice depths. Experimentally, we extract the single-particle correlation func

  65. Namgyu Han, Seong Dae Yun, Chaeeun Lim, Sunghyun Seok

    Echo-planar imaging (EPI) remains the cornerstone of diffusion MRI, but it is prone to severe geometric distortions due to its rapid sampling scheme that renders the sequence highly sensitive to $B_{0}$ field inhomogeneities. While deep learning has helped improve MRI reconstruction, integrating robust geometric distortion correction into a self-supervised f

  66. Susmita Sett, Arash Bahramian, Kristen Dage, David Russell

    Following long periods of quiescence, low-mass X-ray binaries can exhibit intense X-ray outbursts triggered by instabilities within the accretion disk. These outbursts can sometimes be detected in optical wavelengths before being detected in X-rays, acting as an early onset warning and enabling a deep study of accretion disk properties informed by the lag be

  67. Magnus H. Strømme, Alex G. C. de Sá, David B. Ascher

    DPD-Cancer is a graph-attention deep learning framework for predicting small-molecule DPD-Cancer is a graph-attention deep learning framework for predicting small-molecule anti-cancer activity across the NCI-60 panel, trained and evaluated under a strict chemistry-aware data-partitioning scheme. On the hold-out test set, the classifier achieved an Area Under

  68. Kang Zhang, Suyeon Lee, Arda Senocak, Joon Son Chung

    Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS approaches are audio-only, overlooking the inherent audio-visual nature of films, where sounds often align with visual cues. We present the first framework for audio-visual CASS (

  69. B. Hariharan, S. K. Gupta, Y. Hayashi, P. Jagadeesan

    The electric fields inside thunderstorms can significantly modify the intensity of secondary cosmic ray muons at the ground level, producing measurable variations in their intensity ($\Delta$I$_{\mu}$). By utilizing the decade-long observations of thunderstorms (April 2011-December 2020) by the GRAPES-3 muon telescope (G3MT), a directional asymmetry in $\Del

  70. Yue Hu, Junqing Wang, Yingchao Liu

    The rapid scaling of Protein Language Models (PLMs) has unlocked unprecedented accuracy in protein structure prediction and design, but the quadratic memory growth of the Key-Value (KV) cache during inference remains a prohibitive barrier for single-GPU deployment and high-throughput generation. While 8-bit quantization is now standard, 3-bit quantization re

  71. Benedikt Dornauer, Mircea-Cristian Racasan

    This paper introduces RAGnaroX, a resource-efficient ChatOps assistant that operates entirely on commodity hardware. Unlike existing solutions that often rely on external providers such as Azure or OpenAI, RAGnaroX offers a fully auditable, on-premise stack implemented in Rust. Its architecture integrates modular data ingestion, hybrid retrieval, and functio

  72. Jiaming Liang, Yifeng Zhan, Chunlin Liu, Weihua Zheng

    Open-vocabulary object detection (OVOD) aims to detect known and unknown objects in the open world by leveraging text prompts. Benefiting from the emergence of large-scale vision--language pre-trained models, OVOD has demonstrated strong zero-shot generalization capabilities. However, when dealing with camouflaged objects, the detector often fails to disting

  73. Shuangliang Li, Siwei Li, Li Li, Weijie Zou

    Short-term (0-24 hours) precipitation forecasting is highly valuable to socioeconomic activities and public safety. However, the highly complex evolution patterns of precipitation events, the extreme imbalance between precipitation and non-precipitation samples, and the inability of existing models to efficiently and effectively utilize large volumes of mult

  74. Seunghwa Pyo, Donggun Lee, Jungwoo Rhee, Soobin Park

    People increasingly use multiple Multimodal Large Language Models (MLLMs) concurrently, selecting each based on its perceived strengths. This cross-platform practice creates coordination challenges: adapting prompts to different interfaces, calibrating trust against inconsistent behaviors, and navigating separate conversation histories. Prior HCI research fo

  75. Oucheng Liu, Lexing Xie, Jing Jiang

    Climate change is a major socio-scientific issue shapes public decision-making and policy discussions. As large language models (LLMs) increasingly serve as an interface for accessing climate knowledge, whether existing benchmarks reflect user needs is critical for evaluating LLM in real-world settings. We propose a Proactive Knowledge Behaviors Framework th

  76. Yuhang Ma, Jie Wang, Zheng Yan

    Large Language Models (LLMs) have advanced Graph Neural Networks (GNNs) by enriching node representations with semantic features, giving rise to LLM-enhanced GNNs that achieve notable performance gains. However, the robustness of these models against poisoning attacks, which manipulate both graph structures and textual attributes during training, remains une

  77. Tatsuya Yamaoka

    We investigate the Hamiltonian formulation of 1+1-dimensional staggered fermions and reconstruct the vector and axial charge operators, originally identified by Arkya Chatterjee et al., within the Wilson fermion formalism. These operators commute with the Hamiltonian and reduce, in the continuum limit, to the generators of the vector and axial $\mathrm{U}(1)

  78. Hideaki Hara, Riku Omoto, Noboru Sasao, Akihiro Yoshimi

    To achieve more controllable development of coherence in solids, we investigated the effect of a trigger laser tuned to the superradiance transition wavelength on periodic superradiance observed in an Er:YSO crystal. For period control, applying the trigger laser reduced both the superradiance period and its variance, demonstrating enhanced controllability o

  79. Ritwija Roy, Anindya Biswas

    We consider a non-contextual inequality in the sequential measurement scenario and derive the optimal quantum violation of it without assuming the dimension of the system. Since the measurement is dichotomic and the dimension of the quantum system is arbitrary, we formulate the concept of degeneracy-breaking (DB) measurement depending on how many projectors

  80. Jiayi Lei, Xidong Mu, Tiankui Zhang, Wenjun Xu

    A reconfigurable intelligent surface (RIS)-assisted non-orthogonal multiple access (NOMA) system is investigated, where the transmitter (Alice) is a dual functional radar communication (DFRC) base station (BS) that aims to sense the location of a potential warden (Willie), while simultaneously transmitting public and covert signals to the legitimate users, C

  81. Jinxin Hu, Hao Deng, Lingyu Mu, Hao Zhang

    Large-scale industrial recommenders typically use a fixed multi-stage pipeline (recall, ranking, re-ranking) and have progressed from collaborative filtering to deep and large pre-trained models. However, both multi-stage and so-called One Model designs remain essentially static: models are black boxes, and system improvement relies on manual hypotheses and

  82. Eunseo Oh, Suyoun Lee, Jae Young Choi, Soobin Park

    LLMs have become deeply embedded in knowledge work, raising concerns about growing dependency and the potential undermining of human skills. To investigate the pervasiveness of LLMs in work practices, we conducted a four-day diary study with frequent LLM users (N=10), observing how knowledge workers responded to a temporary withdrawal of LLMs. Our findings s

  83. Harunori Kawano, Takeshi Sasaki

    While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this bottleneck, we propose HEAR (Human-inspired Efficient Audio Representation), a novel decoupled architecture. Inspired by

  84. Yulun Wu, Sravan Kumar Ankireddy, Samuel Sharpe, Nikita Seleznev

    Efficiently aggregating spatial or temporal horizons to acquire compact representations has become a unifying principle in modern deep learning models, yet learning data-adaptive representations for long-horizon sequence data, especially continuous sequences like time series, remains an open challenge. While fixed-size patching has improved scalability and p

  85. Hyeongyu Kim, Geonhui Han, Dosik Hwang

    Test-time adaptation (TTA) aims to mitigate performance degradation under distribution shifts by updating model parameters during inference. Existing approaches have primarily framed adaptation around affine modulation, focusing on recalibrating normalization layers. This perspective, while effective, overlooks another influential component in representation

  86. Yehezkiel Darmadi, Thanh Thi Nguyen, Campbell Wilson

    Criminal networks, such as the Sicilian Mafia, pose substantial threats to public safety, national security, and economic stability. Outdated disruption methods with a focus on removing influential individuals or key players have proven ineffective due to the covertness of the network. Thus, researchers have been trying to apply Social Network Analysis (SNA)

  87. Muhammad Apriandito Arya Saputra, Andry Alamsyah, Dian Puteri Ramadhani, Thomhert Suprapto Siadari

    Determining whether a piece of text is relevant to a given topic is a fundamental task in natural language processing, yet it remains largely unexplored for Bahasa Indonesia. Unlike sentiment analysis or named entity recognition, relevancy classification requires the model to reason about the relationship between two inputs simultaneously: a topical context

  88. Suka Sriyansu Pattanaik, Sasmita Mishra

    We investigate the low-energy phenomenology of the Type-I seesaw mechanism within a 3+3 framework containing three active and three sterile neutrinos. Using the exact seesaw relation as a bridge between the high-scale sterile-sector parameters and the standard oscillation observables, we perform a comprehensive Monte Carlo scan of the 21-dimensional sterile

  89. Mohammed Elnawawy, Gargi Mitra, Shahrear Iqbal, Karthik Pattabiraman

    Safety-critical domains like healthcare rely on deep neural networks (DNNs) for prediction, yet DNNs remain vulnerable to evasion attacks. Anomaly detectors (ADs) are widely used to protect DNNs, but conventional ADs are trained indiscriminately on benign data from all patients, overlooking physiological differences that introduce noise, degrade robustness,

  90. Youngjun Song, Hyeongyu Kim, Dosik Hwang

    Test-Time Adaptation (TTA) enables real-time adaptation to domain shifts without off-line retraining. Recent TTA methods have predominantly explored additive approaches that introduce lightweight modules for feature refinement. Recently, a subtractive approach that removes domain-sensitive channels has emerged as an alternative direction. We observe that the

  91. Guoqing Wang, Zeyu Sun, Xiaofei Xie, Yizhou Chen

    Web-augmented large language models (LLMs) offer promising capabilities for automatic code generation. However, integrating live web search exposes models to unreliable or malicious content, leading to Search-Induced Issues (SII), a novel failure mode in which external pages mislead LLMs into producing incorrect code. This paper presents a comprehensive empi

  92. Christopher Ackerman

    The ability to represent oneself and others as agents with knowledge, intentions, and belief states that guide their behavior - Theory of Mind - is a human universal that enables us to navigate - and manipulate - the social world. It is supported by our ability to form mental models of ourselves and others. Its ubiquity in human affairs entails that LLMs hav

  93. Chen Liu, Qizhen Lan, Zhicheng Ding, Xinyu Chu

    As deep vision models grow increasingly complex to achieve higher performance, deployment efficiency has become a critical concern. Knowledge distillation (KD) mitigates this issue by transferring knowledge from large teacher models to compact student models. While many feature-based KD methods rely on spatial filtering to guide distillation, they typically

  94. Hiroki Iimori, Yuto Hama

    Massive multiple-input multiple-output (MIMO) has enabled substantial spatial multiplexing and array gains in real-world systems, while distributed MIMO (D-MIMO) improves macro-diversity over wide areas at the cost of deployment complexity. Repeater-assisted massive MIMO (RA-MIMO) is a lower-cost alternative that can recover key distributed-MIMO advantages.

  95. Asim D. Bakhshi

    Large language models (LLMs) exhibit systematic miscalibration with rhetorical intensity not proportionate to epistemic grounding. This study tests this hypothesis and proposes a framework for quantifying this decoupling by designing a triadic epistemic-rhetorical marker (ERM) taxonomy. The taxonomy is operationalized through composite metrics of form-meanin

  96. Shibo Liu

    Real-time 30-to-60 fps video frame interpolation on mobile neural processing units (NPUs) requires each synthesized frame within 33.3 ms. We show that mainstream flow-based video frame interpolation faces three structural deployment barriers on mobile NPUs: spatial sampling operators exceed the frame budget or lack hardware support, iterative flow refinement

  97. Farhan Fuad Abir, Sanjeda Sara Jennifer, Niloofar Yousefi, Laura J. Brattain

    We propose a hybrid diffusion-based augmentation framework to overcome the critical challenge of ultrasound data augmentation in breast ultrasound (BUS) datasets. Unlike conventional diffusion-based augmentations, our approach improves visual fidelity and preserves ultrasound texture by combining text-to-image generation with image-to-image (img2img) refinem

  98. Hao Zhang, Jinxin Hu, Hao Deng, Lingyu Mu

    AutoModel is an agent based architecture for the full lifecycle of industrial recommender systems. Instead of a fixed recall and ranking pipeline, AutoModel organizes recommendation as a set of interacting evolution agents with long term memory and self improvement capability. We instantiate three core agents along the axes of models, features, and resources

  99. Tetsuya Onogi, Tatsuya Yamaoka

    We study conserved charges of the staggered fermion Hamiltonian in 3+1 dimensions. By decomposing staggered fermions into Majorana components and exploiting lattice translation symmetries, we construct a set of conserved non-singlet charges. We analyze their algebra andshow that, although the charges exhibit nontrivial non-commutativity on the lattice, they

  100. Wei Zhong, Youjin Deng

    In this letter, we unveil a robust, pre-asymptotic scaling regime for the Binder cumulant $U_L$, a central finite-size scaling tool, demonstrating $U_L\sim N^{-1} |t|^{-d\nu}$ (disordered phase) and $\frac{2}{3}-U_L\sim N^{-1} |t|^{-d\nu}$ (ordered phase), with $t$ being the reduced control parameter, and $N$, $d$, $\nu$ represent the total number of sites,