Skip to content

March 2026 arXiv papers — page 114

Showing 11,30111,400 of 25,974 papers

  1. Yunshu Bai, RuiHao Li, Hao Zhang, Chien Her Lim

    Game UI implementation requires translating stylized mockups into interactive engine entities. However, current "Screenshot-to-Code" tools often struggle with the irregular geometries and deep visual hierarchies typical of game interfaces. To bridge this gap, we introduce SPRITE, a pipeline that transforms static screenshots into editable engine assets. By i

  2. Murali Haran, Bokgyeong Kang, Jaewoo Park

    In this paper we discuss a well known computing problem -- inference for models with intractable normalizing functions. Models with intractable normalizing functions arise in a wide variety of areas, for instance network models, models for spatial data on lattices, spatial point processes, flexible models for count data and gene expression, and models for pe

  3. Tianfu Li, Wenbo Chen, Haoxuan Xu, Xinhu Zheng

    In Vision-and-Language Navigation (VLN), an agent is required to plan a path to the target specified by the language instruction, using its visual observations. Consequently, prevailing VLN methods primarily focus on building powerful planners through visual-textual alignment. However, these approaches often bypass the imperative of comprehensive scene under

  4. Ratun Rahman, Dinh C. Nguyen

    Conventional federated learning (FL) frameworks often suffer from training degradation due to data uncertainty and heterogeneity across local clients. Probabilistic approaches such as Bayesian neural networks (BNNs) can mitigate this issue by explicitly modeling uncertainty, but they introduce additional runtime, latency, and bandwidth overhead that has rare

  5. Molla Basir Ahamed, Sanju Mandal

    This article investigates the Bohr phenomenon and sharp coefficient problems for the class $\mathcal{A}_{\beta}$, a subclass of analytic self-maps of the unit disk with the holomorphic generators of one-parameter continuous semigroups. By integrating concepts from complex dynamics and geometric function theory, we derive sharp improvements to the classical B

  6. Virginia Agostiniani, Riccarda Rossi, Giuseppe Savaré

    We consider singularly perturbed gradient flows in Hilbert spaces, driven by a time-dependent, nonconvex, and nonsmooth energy, and address the convergence of their solutions to curves of critical points of the driving energy functional. The degenerating nature of the estimates along the gradient-flow curves calls for novel compactness arguments, which we ca

  7. Riccardo Brasca, Gabriella Clemente

    This article is about the formalization of synthetic differential geometry with the Lean proof assistant and the mathematical library mathlib. The main result we prove and formalize is a Taylor theorem for functions of several variables, where the series expansion is around an infinitesimal neighborhood. Most of our proofs are in fact new. Our investigations

  8. Xinyuan Qian, Xinjia Zhu, Alessio Brutti, Dong Liang

    TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to missing visual data, neglecting the role of head orientation, and background noise. This study addresses these limitation

  9. Yijun Sun, Xudong Liao, Songrun Xie, Hao Chen

    Meeting stringent Time-To-First-Token (TTFT) requirements is crucial for LLM applications. To improve efficiency, modern LLM serving systems adopt disaggregated architectures with diverse parallelisms, introducing complex multi-stage workflows involving reusable KV-block retrieval, collective communication, and P2D transfer. Flows from dependent stages overl

  10. Weidong Chen, Cheng Ye, Zhendong Mao, Peipei Song

    Emotional Video Captioning (EVC) is an emerging task, which aims to describe factual content with the intrinsic emotions expressed in videos. Existing works perceive global emotional cues and then combine with video content to generate descriptions. However, insufficient factual and emotional cues mining and coordination during generation make their methods

  11. Leyan Li, Yuming Lin, Xiaohu Sun, Yajun Mao

    Future electron-positron colliders offer a unique opportunity for high-precision measurements of the top-quark mass, width, strong coupling constant, and top-quark Yukawa coupling via a scan of the $t\bar{t}$ threshold. We present the first prospect study of the simultaneous determination of these parameters, incorporating the latest reference detector desig

  12. Yuting Zheng, Zijian Chen, Qi Jia

    Unraveling the hierarchical structure-property relationships is the central challenge of materials science, necessitating the interpretation of data across vast physical scales from micro to macro. Despite the rapid integration of Large Multimodal Models (LMMs) into scientific workflows, existing scientific benchmarks primarily focus on general chart interpr

  13. Marc Damie, Florian Hahn, Andreas Peter, Jan Ramon

    Function Secret Sharing (FSS) schemes enable sharing efficiently secret functions. Schemes dedicated to point functions, referred to as Distributed Point Functions (DPFs), are the center of FSS literature thanks to their numerous applications including private information retrieval, anonymous communications, and machine learning. While two-party DPFs benefit

  14. Lucile Riaboff, Ingrid David

    Background Genetic parameters of feeding behaviours traits from electronic feeding stations in relation to feed efficiency have been widely explored. However, genetic determinism of the circadian rhythm of feed intake throughout the fattening phase in group-housed growing pigs fed ad libitum has never been investigated, despite the well-known relationships b

  15. Yu Nakayama, Tadashi Okazaki

    We investigate holographic spectral functions for general Sasaki-Einstein 5-manifolds dual to four-dimensional superconformal field theories, including supersymmetric indices, supersymmetric zeta functions, and supersymmetric determinants. The analytic structure of the supersymmetric zeta function, particularly its residue and special value, allows for the c

  16. Yong-Cheng Pan, Tommy Kotte, Toni Helm, Motoki Osada

    We report a systematic magnetotransport study on high-crystallinity La$_{1-x}$Sr$_{x}$NiO$_2$ (LSNO) thin films with $x=0.20-0.24$. By conducting pulsed-field transport experiment up to 62 T, we reveal two salient features of the normal-state transport in overdoped LSNO thin films: (1) the magnetoresistance does not follow the Kohler's rule but exhibits a $H

  17. Junyoung Kim, Woojoo Kim, Wonbin Kweon, Jaehyung Lim

    Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity and prevents Collaborative Filtering (CF) signals from being fully integrated into item representations. Inspired by the recent success of Large Language Models (LLMs) as high-capacity embedders, we investigate the use of

  18. Alex Shvets

    We study canonical one-step neighboring shuffle experiments for finite-output epsilon_0-LDP d-ary channels along growing alphabets, with frequency estimation and mechanism design under a pairwise chi-squared budget. The pairwise likelihood-ratio law nu_{ab,d} (pushforward of the row ratio under the null row) is the governing invariant: the canonical shuffled

  19. Tingcheng Bian, Jinchang Luo, Mingquan Cheng, Jinyu Zhang

    Large language models achieve breakthroughs in complex reasoning via long chain-of-thought sequences. However, this often leads to severe reasoning inflation, causing substantial computational redundancy. To maximize Intelligence per Token, we introduce a theoretical metric, MSL-Minimal Sufficient Length. MSL rigorously characterizes the shortest reasoning l

  20. Aayam Bansal, Ishaan Gangwani

    Cooperative multi-agent methods for embodied AI are almost universally evaluated under idealized communication: zero latency, no packet loss, and unlimited bandwidth. Real-world deployment on robots with wireless links, autonomous vehicles on congested networks, or drone swarms in contested spectrum offers no such guarantees. We introduce AgentComm-Bench, a

  21. Ziyi He, Yushi Feng, Shuangyu Yang, Yinghao Zhu

    Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiographic evidence) to determine complete referral plans. We present Dental-TriageBench, the first expert-annotated benchmark for reasoning-driven multimodal dental triage. Built from authentic outpatient workflo

  22. Dhivya Prabhu K, Sanjeev Singh, Antony Vijesh

    This paper develops an efficient iterative method for computing all zeros of solutions of second order ordinary differential equations. A third order Halleys method is first derived by approximating the solution of an associated Riccati differential equation. To improve computational efficiency, a modified Halleys method is proposed by fixing one of the func

  23. Anjan Daimari, Shivanee Borah, Diana Thongjaomayum

    We study a minimal model of disordered systems, the random field Ising model (RFIM) on a generalized Petersen Graph, GP(N,k). This graph has a connected inner and outer loop, where both the loops consist of N nodes constituting a total of 2N nodes. The parameter k satisfies the condition 1<=k<=N/2, such that any site i in the inner loop has i-k and i+k as it

  24. David Awad

    We present an empirical argument against the existence of single timeline backward time travel using the price behavior of prediction markets. If rational agents could travel backward in time, binary prediction contracts would converge to degenerate prices (0 or 1) immediately upon market formation. We observe no such behavior across large datasets of resolv

  25. Alexander V. Shenderuk-Zhidkov, Alexander E. Hramov

    This article introduces and substantiates the concept of Neuro-Linguistic Integration (NLI), a novel paradigm for human-technology interaction where Large Language Models (LLMs) act as a key semantic interface between raw neural data and their social application. We analyse the dual nature of LLMs in this role: as tools that augment human capabilities in com

  26. Nikolaos D. Tantaroudas, Ilias Karachalios

    H Infinity robust control synthesis for gust load alleviation of very flexible aircraft is presented. The controller is synthesised on a compact reduced-order model comprising 8 degrees of freedom for the UAV configuration and 9 for the flying-wing, obtained through nonlinear model order reduction of the coupled fluid-structure-flight dynamics system, and va

  27. Mankeun Jeong, Myungshin Im, Joonho Kim, Seo-Won Chang

    We present a comprehensive pipeline developed for the image processing of the KMTNet Synoptic Survey of the Southern Sky (KS4) Data Release 1. This pipeline encompasses several key processes, including data quality assurance, astrometry, photometric zero-point (ZP) calibration, bad pixel masking, image stacking, and difference image analysis (DIA). The astro

  28. Siqi Pei, Liang Tang, Tiaonan Duan, Long Chen

    GUI grounding is a critical capability for vision-language models (VLMs) that enables automated interaction with graphical user interfaces by locating target elements from natural language instructions. However, grounding on GUI screenshots remains challenging due to high-resolution images, small UI elements, and ambiguous user instructions. In this work, we

  29. Sofiya Karankova, Yeunjeong Lee, Seungmin Park, Kenji Watanabe

    Solid-state quantum emitters constitute an essential building blocks of integrated quantum photonic circuits. Among potential emitter platforms, hexagonal boron nitride (hBN) hosts single-photon emitters in an atomically thin lattice amenable to photonic integration. However, multi-step fabrication approaches, limited defect specificity, and poor emission wa

  30. Linxiao Yang, Xue Jiang, Gezheng Xu, Tian Zhou

    Transformers enable in-context learning (ICL) for rapid, gradient-free adaptation in time series forecasting, yet most ICL-style approaches rely on tabularized, hand-crafted features, while end-to-end sequence models lack inference-time adaptation. We bridge this gap with a unified framework, Baguan-TS, which integrates the raw-sequence representation learni

  31. Zhichun Yang, Li Jiang, Tianxiang Liu, Man-Chung Yue

    Classical analyses show that randomized coordinate descent (RCD) and gradient descent (GD) share the same convergence rates in terms of objective gap under convexity and specific global EB-type assumptions. However, this rate preservation phenomenon does not extend to iterate rates, almost-sure rates, or general local error bound conditions. In this paper, w

  32. Kehan Chen, Yan Huang, Dong An, Jiawei He

    Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to reason about spatial structures. Although human-readable spatial schematics (e.g., floor plans) are ubiquitous in real-world buildings, current agents lack the cognitive ability to

  33. Yue Hu, Jialiang Tang, Siwei Yu, Baosheng Yu

    Non-stationarity is a fundamental challenge in multivariate long-term time series forecasting, often manifested as rapid changes in amplitude and phase. These variations lead to severe distribution shifts and consequently degrade predictive performance. Existing normalization-based methods primarily rely on first- and second-order statistics, implicitly assu

  34. Ruibo Fan, Xiangrui Yu, Xinglin Pan, Zeyu Li

    Lossless model compression holds tremendous promise for alleviating the memory and bandwidth bottlenecks in bit-exact Large Language Model (LLM) serving. However, existing approaches often result in substantial inference slowdowns due to fundamental design mismatches with GPU architectures: at the kernel level, variable-length bitstreams produced by traditio

  35. Srikanth Cherukupally

    For number $n>1$, let $\mathcal{A}(n) = \{1\leq a<n: n|a^2-1, a|n^2-1 \}$. We show that the size of $\mathcal{A}(n)$ is connected to a property concerning integer evaluations of Fibonacci-like polynomials. In the process, we prove that $|\mathcal{A}(n)|< \log_2 n$, and establish the average value of $|\mathcal{A}(n)|$ to be a little above $2$, asymptotically

  36. Prince Zizhuang Wang, Shuli Jiang

    Large Language Model (LLM) agents have shown strong results on multi-turn tool-use tasks, yet they operate in isolation during training, failing to leverage experiences accumulated across episodes. Existing experience-augmented methods address this by organizing trajectories into retrievable libraries, but they retrieve experiences only once based on the ini

  37. Benjamin Ingimarson, Igor Kukavica

    Under the assumption that a solution to the 3D incompressible Euler equations blows up at a time $T_\ast$ and that $T_\ast $ is the first such time, we establish lower bounds on the rate of blow-up of the maximum norm of the vorticity. In particular, when the domain is $\mathbb{R}^3$ or $\mathbb{T}^3$, we provide lower bounds on $\int_{0}^{t}\Vert \omega\Ver

  38. J. Willingham, A. Hopkins, T. Zafar, J. Afonso

    We present a novel approach to correcting H$\alpha$ luminosity functions for dust extinction by calibrating against radio-based star formation rates (SFRs), using data from the Evolutionary Map of the Universe (EMU) and Galaxy and Mass Assembly (GAMA) surveys. Accurate dust correction is essential for deriving SFRs from rest-frame UV-optical emission lines,

  39. Tomochika Kurita

    As we are entering an early-FTQC era, circuit execution protocols with logical qubits and certain error-correcting codes are being discussed. Here, we propose a circuit execution protocol for the space-time efficient analog rotation (STAR) architecture. Gate operations within the STAR architecture is based on lattice surgery with surface codes, but it allows

  40. Xiangyu Kong, Xiaoyu Jin, Yihan Pan, Haoqin Sun

    In natural face-to-face interaction, participants seamlessly alternate between speaking and listening, producing facial behaviors (FBs) that are finely informed by long-range context and naturally exhibit contextual appropriateness and emotional rationality. Interactive Head Generation (IHG) aims to synthesize lifelike avatar head video emulating such capabi

  41. Zhenhai Pan, Yan Liu, Jia You

    Most automated electronic medical record (EMR) pipelines remain output-oriented: they transcribe, extract, and summarize after the consultation, but they do not explicitly model what is already known, what is still missing, which uncertainty matters most, or what question or recommendation should come next. We formulate doctor-patient dialogue as a proactive

  42. Akira Saito, Masato Tanaka

    Piecewise-linear nonlinear systems appear in many engineering disciplines. Prediction of the dynamic behavior of such systems is of great importance from practical and theoretical viewpoint. In this paper, a data-driven model order reduction method for piecewise-linear systems is proposed, which is based on dynamic mode decomposition (DMD). The overview of t

  43. Aaron Lau, Kouji Yano

    The weak and strong laws of large numbers for time-inhomogeneous Markov chains are studied under general conditions. First, under Drift Condition and Contraction Condition in total variation, we prove the weak law of large numbers. Then, assuming Drift Condition together with a time-inhomogeneous Doeblin minorization, we develop a Nummelin-type splitting and

  44. Jie Zheng, Dusit Niyato, Changyuan Zhao, Jiawen Kang

    The rapid evolution toward 6G and beyond communication systems is accelerating the convergence of digital twins and world models at the network edge. Traditional digital twins provide high-fidelity representations of physical systems and support monitoring, analysis, and offline optimization. However, in highly dynamic edge environments, they face limitation

  45. Saikat Maiti

    Autonomous AI agents powered by large language models are being deployed in production with capabilities including shell execution, file system access, database queries, and multi-party communication. Recent red teaming research demonstrates that these agents exhibit critical vulnerabilities in realistic settings: unauthorized compliance with non-owner instr

  46. Zichen Tang, Zirui Zhang, Qian Wang, Zhenheng Tang

    Current Large Language Models (LLMs) are gradually exploited in practically valuable agentic workflows such as Deep Research, E-commerce recommendation, and job recruitment. In these applications, LLMs need to select some optimal solutions from massive candidates, which we term as \textit{LLM-as-a-Recommender} paradigm. However, the reliability of using LLM

  47. Jinyu Miao, Pu Zhang, Rujun Yan, Yifei He

    Advanced autonomous driving systems require accurate vehicle dynamics modeling. However, identifying a precise dynamics model remains challenging due to strong nonlinearities and the coupled longitudinal and lateral dynamic characteristics. Previous research has employed physics-based analytical models or neural networks to construct vehicle dynamics represe

  48. J. A. Woodside, B. J. Coombes, A. E. Stuchbery, A. J. Mitchell

    The low-excitation states of atomic nuclei in the region around the $N = Z = 28$ shell closure are generally well described by the shell model. Most experimental observables in the iron isotopes $^{56}$Fe, $^{58}$Fe, and $^{60}$Fe ($Z = 26$; $N=30$, $32$, $34$) support a shell-model description. However, the lifetimes of the $4_1^+$ state in $^{58}$Fe in the

  49. Chaeyun Kim, Seunghoon Yi, Yejin Kim, Yohan Jo

    Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones. To address this, we first introduce an efficient data augmentation scheme that extracts motion-centric phrases from original caption

  50. Shiming Chen, Shuhuang Chen, Guo-Sen Xie, Xinge You

    Zero-shot learning (ZSL) aims to recognize the unseen classes in the open-world guided by the side-information (e.g., attributes). Its key task is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus conducting a desirable semantic knowledge transfer from seen classes to unseen ones. Prior works simply ut

  51. Thierry De Pauw

    We review recent results on Radon-Nikod\'ymification of abstract measure spaces, the particular case of integral geometric measure, and applications to the dual of SBV.

  52. Miaoqian Lu, Xinzhou Guan, Mohan Xia, Wenjuan Li

    We present an analytical framework for stabilizing second-order correlated tunneling of two spin-orbit-coupled bosons in a periodically driven non-Hermitian double-well potential. By combining Floquet theory with multiple-scale asymptotic analysis, we derive effective second-order dynamics and exact quasienergy spectra in the strongly interacting regime. Our

  53. Priyanka Aroda, Arup Chattopadhyay, Supratim Jana

    We introduce and systematically study a class of operators that arise naturally due to the Beurling decomposition of the Hardy space $H^2=K_\theta \oplus \theta H^2$. While the compressions of classical Toeplitz and Hankel operators to the Beurling subspace $\theta H^2$ and the model space $K_\theta$ account for the diagonal components of the decomposition,

  54. Runze Wang, Yuxuan Song, Youcheng Cai, Ligang Liu

    Online 3D reconstruction from streaming inputs requires both long-term temporal consistency and efficient memory usage. Although causal variants of VGGT address this challenge through a key-value (KV) cache mechanism, the cache grows linearly with the stream length, creating a major memory bottleneck. Under limited memory budgets, early cache eviction signif

  55. Xinning Chai, Zhengxue Cheng, Xin Li, Rong Xie

    Recent diffusion-based extreme image compression methods have demonstrated remarkable performance at ultra-low bitrates. However, most approaches require training separate diffusion models for each target bitrate, resulting in substantial computational overhead and hindering practical deployment. Meanwhile, recent studies have shown that joint super-resoluti

  56. Watanjeet Singh, Sumit Chandok

    This paper presents a modified iterative approach to solve the variational inequality problem using the double inertial technique in the context of a real Hilbert space. Our iterative technique involves a projection onto a generalized half-space and a self-adaptive step-size rule which works without prior knowledge of the Lipschitz constant of the operator.

  57. Moritz M. Hirschmann, Akira Furusaki, Max Hirschberger

    Owing to their relevance for spintronics, electronic band splitting and spin-polarization textures in magnets are active areas of research. In non-collinear magnets, alternating spin textures can arise both for isolated bands and for intersecting band pairs with nodal splitting. This raises the question of whether $p,f,...$-wave magnets should be defined by

  58. Shengjie Huang, Sijie Yang, Jianqiao Yi, Rui Zheng

    Replica exchange (REX) is one of the most widely used enhanced sampling methodologies, yet its efficiency is limited by the requirement for a large number of intermediate temperature replicas. Here we present Generative Replica Exchange (GREX), which integrates deep generative models into the REX framework to eliminate this temperature ladder. Drawing inspir

  59. Alireza Sadeghi, Wael AbdAlmageed

    Causal representation learning (CRL) models aim to transform high-dimensional data into a latent space, enabling interventions to generate counterfactual samples or modify existing data based on the causal relationships among latent variables. To facilitate the development and evaluation of these models, a variety of synthetic and real-world datasets have be

  60. Wenzhi Wang, Tianyu Li, Wei Yi

    Boundary conditions can have dramatic impact in non-Hermitian systems, as exemplified by the non-Hermitian skin effect. Focusing on one-dimensional non-Hermitian quasiperioidic lattices, we show that the interplay of quasiperiodicity and the non-Hermitian skin effect leads to counterintuitive localization properties. On the one hand, for Anderson localized s

  61. Yaozhong Shi, Grigorios Lavrentiadis, Konstantinos Tsalouchidis, Zachary E. Ross

    Earthquake hazard analysis and design of spatially distributed infrastructure, such as power grids and energy pipeline networks, require scenario-specific ground-motion time histories with realistic frequency content and spatiotemporal coherence. However, producing the large ensembles needed for uncertainty quantification with physics-based simulations is co

  62. Keiichi Shigechi

    We study four bijections, which are promotion, evacuation, rowmotion, and rowvacuation, on generalized Dyck paths in rational Catalan combinatorics. We define the maps on generalized Dyck paths, which have their origins in maps on Dyck paths and non-crossing partitions. They include rotation, Kreweras complement map, Simion--Ullman involution on non-crossing

  63. Yingpeng Qi, Nianke Chen, Zhihui Zhou, Qing Xu

    The intrinsic nature of glass states and glass transitions remain a fundamental open question in condensed-matter physics and materials science. The key to solving the glass transition problem lies in achieving a complete understanding of the physics governing the structural relaxation. Nonetheless, directly probing dynamic atomic-scale structural changes in

  64. Rui Hong, Shuxue Quan

    We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content uniformly, our method dynamically adjusts temporal attention receptive fields based on estimated motion content: high-motion sequences attend locally across frames to preserve rap

  65. Esli Diepenbroek, Leon A. Smook, Sissi de Beer

    With the ever-increasing digitization of society, the development of materials with low-power memory storage -similar to synapses- is becoming more relevant. The field of iontronic artificial synapses has gained traction, in particular with polymers as the memory-active material which allows for additional bio-compatibility, flexibility and tunability. Polye

  66. Rui Hong, Jana Kosecka

    Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture labels is available and show that gesture semantics can serve as a powerful inductive bias for 3D pose estimation. We present a two-stage f

  67. Takenori Kataoka, Manabu Ozaki

    Let $G$ be a finite $p$-group. We construct a $G$-extension $K/k$ of number fields such that the $p$-adic completion of the unit group of $K$ has a prescribed $\mathbb{Z}_p[G]$-module structure, up to free direct summands.

  68. Thomas Creutzig, Niklas Garner, Byeonggi Go, Heeyeon Kim

    We propose a three-dimensional field theory construction that realizes the vertex algebras associated with the intermediate Lie algebras and the related $C_2$-cofinite minimal $W$-algebras of the Deligne-Cvitanovi\'c (DC) series as boundary algebras. The construction is based on the minimal three-dimensional ${\mathcal N}=4$ superconformal field theory coupl

  69. Nobuaki Murase, Masaharu Isobe

    The phase diagram of self-propelled hard disk systems with Vicsek-type alignment interactions was investigated by event-driven molecular dynamics simulations. The model incorporates two competing order parameters: the polar order-disorder transition associated with collective velocity alignment (Vicsek model) and the orientational order arising from solid-fl

  70. Jiawen Kang, Kun Li, Dongrui Han, Jinchao Li

    Automated Alzheimer's Disease (AD) screening has predominantly followed the inductive paradigm of pattern recognition, which directly maps the input signal to the outcome label. This paradigm sacrifices construct validity of clinical protocol for statistical shortcuts. This paper proposes Agentic Cognitive Profiling (ACP), an agentic framework that realigns

  71. Ian Chen, Alfredo Alexander-Katz

    Free energies are fundamental quantities governing phase behavior and thermodynamic stability in polymer systems, yet their accurate computation often requires extensive simulations and post-processing techniques such as the Bennett Acceptance Ratio (BAR). While BAR provides reliable estimates when applied between closely related thermodynamic states, evalua

  72. Mathias Richerzhagen, Naidu Bezawada, Sebastian Elias Egner, Elizabeth George

    Large astronomical instruments using tens to hundreds of optical or infrared science detectors pose specific challenges for detector control, where, in addition to performance, other engineering aspects like scalability, power consumption, size, weight and programmatic aspects such as cost and sustainability need to be considered. In this paper we analyze th

  73. Rui Hong, Jana Kosecka

    Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of phonological attribute conditioning for sign language motion generation, using ASL-LEX 2.0 annotations such as hand shape, hand loc

  74. Guangzhi Wang, Yinghao Jiao, Zhi Liu

    The central challenge of reasoning-intensive retrieval lies in identifying implicitreasoning relationships between queries and documents, rather than superficial se-mantic or lexical similarity. The contrastive learning paradigm is fundamentallya static representation consolidation technique: during training, it encodes hier-archical relevance concepts into

  75. Guangzhi Wang, Xiaohui Yang, Kai Li, Jiawen He

    As retrieval models converge on generic benchmarks, the pressing question is no longer "who scores higher" but rather "where do systems fail, and why?" Person-job matching is a domain that urgently demands such diagnostic capability -- it requires systems not only to verify explicit constraints but also to perform skill-transfer inference and job-competency

  76. Rui Wu, Hong Xie, Yongjun Li

    Judea Pearl's do-calculus provides a foundation for causal inference, but its translation to continuous generative models remains fraught with geometric challenges. We establish the fundamental limits of such interventions. We define the Counterfactual Event Horizon and prove the Manifold Tearing Theorem: deterministic flows inevitably develop finite-time si

  77. Rui Wu, Hong Xie, Yongjun Li

    Current continuous generative models (e.g., Diffusion Models, Flow Matching) implicitly assume that locally consistent causal mechanisms naturally yield globally coherent counterfactuals. In this paper, we prove that this assumption fails fundamentally when the causal graph exhibits non-trivial homology (e.g., structural conflicts or hidden confounders). We

  78. Weixin Liu, Bowen Qu, Amy Stone, Maria E. Powell

    Velopharyngeal dysfunction (VPD) is characterized by inadequate velopharyngeal closure during speech and often causes hypernasality and reduced intelligibility. Although speech-based machine learning models can perform well under standardized clinical recording conditions, their performance often drops in real-world settings because of domain shift caused by

  79. Hongbo Lu, Liang Yao, Chenghao He, Fan Liu

    A fundamental bottleneck in Novel View Synthesis (NVS) for autonomous driving is the inherent supervision gap on novel trajectories: models are tasked with synthesizing unseen views during inference, yet lack ground truth images for these shifted poses during training. In this paper, we propose VisionNVS, a camera-only framework that fundamentally reformulat

  80. Zhenhai Pan, Yan Liu, Jia You

    Most dialogue-based electronic medical record (EMR) systems still behave as passive pipelines: transcribe speech, extract information, and generate the final note after the consultation. That design improves documentation efficiency, but it is insufficient for proactive consultation support because it does not explicitly address streaming speech noise, missi

  81. Shuizhou Chen, Lang Yu, Xueqin Lin, Xinjie Mao

    Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching.

  82. Kinya Guan, Hosho Katsura

    We introduce and study a disorder-free version of the quantum breakdown model with all-to-all interactions. The Hamiltonian factorizes into the product of the zero-momentum-mode occupation number and a quadratic Hamiltonian including only pairing terms. This structure makes the model exactly solvable and produces an extensive zero-energy degeneracy. We obtai

  83. Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu

    We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates reward and language models as choice data is received. The reward model is fit to the choice data, while the language model is updated by a variation of reinforce, with reinforcement

  84. Vadim Rozenfeld, Bracha Laufer Goldshtein

    Reliable Sound Source Localization (SSL) plays an essential role in many downstream tasks, where informed decision making depends not only on accurate localization but also on the confidence in each estimate. This need for reliability becomes even more pronounced in challenging conditions, such as reverberant environments and multi-source scenarios. However,

  85. Yang-Tian Sun, Zehuan Huang, Yifan Niu, Lin Ma

    We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry directly from disparity. To efficiently achieve consistent stere

  86. Mengyu Zhao, Di Fu, Yongyu Xie, Jiaxing Zhang

    Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when only a small number of frames can be retained, existing samplers often fail to balance broad video coverage with brief but critical events, which can lead to unreliable downstream

  87. Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva

    Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture whether a model is simultaneously pedagogically effective and safe across student-tutor interaction. We argue that tutoring safety is fundamentally different from conventional LLM

  88. Zhihua Wei, Qiang Li, Jian Ruan, Zhenxin Qin

    Large vision-language models (VLMs) often exhibit weakened safety alignment with the integration of the visual modality. Even when text prompts contain explicit harmful intent, adding an image can substantially increase jailbreak success rates. In this paper, we observe that VLMs can clearly distinguish benign inputs from harmful ones in their representation

  89. Liping Li

    We develop a unified representation theory for the categories of finite subsets and relation-preserving maps of highly homogeneous relational structures classified by Cameron. For any commutative coefficient ring $k$, we extend the classical Dold-Kan correspondence to this setting, with the sole exception of the category $\mathrm{FA}$, and prove that finitel

  90. Zhenxing Yan, Jidong Yuan, Yongqi Sun, Haiyang Liu

    Graph neural network (GNN)-based federated recommendation systems effectively capture user-item relationships while preserving data privacy. However, existing methods often face slow convergence on graph data and privacy leakage risks during collaboration. To address these challenges, we propose FastPFRec (Fast Personalized Federated Recommendation with Secu

  91. Umangi Jain, Vladimir Kim, Matheus Gadelha, Igor Gilitschenski

    We introduce the problem of material-aware part grouping in untextured meshes. Many real-world shapes, such as scales of pinecones or windows of buildings, contain repeated structures that share the same material but exhibit geometric variations. When assigning materials to such meshes, these repeated parts often require piece-by-piece manual identification

  92. Md Mahfuzur Rahman, Jareen Shuva, Nishith Tripathi, Lingjia Liu

    We propose a machine learning (ML) and smartphone-assisted framework for uplink performance prediction in a private, realistic 5G cellular system using real-time measurements in both indoor and outdoor settings. This work presents a comprehensive data-driven evaluation of 5G performance prediction using a controllable software-defined radio test environment.

  93. Jiaxin Zhang, Wenqian Shen, Kai Yang, Zhen Gao

    To address the challenges of high-dimensional channel estimation and underutilized spatial correlations among users in holographic MIMO (HMIMO) systems, this paper proposes a joint graph-cut algorithm for multi-user channel estimation in the wavenumber domain. The size of the conventional angular domain channel matrix increases with the number of antennas in

  94. Jianan Chen, Zhifang Zhang, Shuo He, Linan Yue

    Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of significantly degraded safety capabilities. In this paper, we reveal that LRMs' safety degradation occurs only after CoT is enabled, and this degradation is not observed when CoT i

  95. Zihan Yan, Denan Li, Xin Wu, Zhoulin Liu

    Machine-learned interatomic potentials have revolutionized molecular dynamics simulations by providing quantum-mechanical accuracy at empirical-potential speeds. The graphics processing unit molecular dynamics (GPUMD) package, featuring the highly efficient neuroevolution potential (NEP) framework, has emerged as a powerful tool in this domain. However, the

  96. Xishuo Wei, Handi Huang, Haotian Chen, Hongxuan Zhu

    The optimized stellarator is an attractive concept for which the averaged particle radial drift is zero, and the single particle loss can be significantly reduced. But for the reactor design, global physics such as turbulent transport also need to be optimized besides the confined single particle orbit, or properties estimated using local estimations and heu

  97. Ziran Liu

    Internal noise in deep networks is usually inherited from heuristics such as dropout, hard masking, or additive perturbation. We ask two questions: what correlation geometry should internal noise have, and is the implemented perturbation compatible with the representations it acts on? We answer these questions through Variational Kernel Design (VKD), a frame

  98. Samuel Laliberte, Reiko Toriumi

    We explore how matrix bootstrap techniques can be used to constrain matrix and tensor models at finite $N$, where $N$ is the dimension of the matrix/tensor, taking a Gaussian model with a quartic interaction as example. For matrix models, we find further evidence that bounds do not depend explicitly on $N$, but rather on properties of multi-trace expectation

  99. Guanghui Su, Timothy H. Nguyen, Balthazar Loglia, Aaron Weinstein

    Optical nanofibers with subwavelength diameters generate strong evanescent fields, enabling efficient light-matter interactions for optical sensing, spectroscopy, and cold-atom experiments. We report a heat-and-pull system for fabricating low-loss optical nanofibers with controllable waist dimensions and investigate the fabrication limits for achieving small

  100. Junzhuo Ma, Chenghuang Shen, Yi Yu, Xingyan Liu

    Technical-service LLM agents are entering production workflows, where value depends on whether engineers adopt generated replies. Service tickets hide decision logic, contain noisy single-reference responses, and make reward evaluation costly, making standard post-training brittle. Existing post-training and LLM-as-a-Judge approaches improve grounding or fee