Skip to content

May 2025 arXiv papers — page 70

Showing 6,9017,000 of 24,552 papers

  1. Felix Ahnefeld, Thomas Theurer, Martin B. Plenio

    Quantum phase estimation is a core task in quantum technologies ranging from metrology to quantum computing, where it appears as a key subroutine in various algorithms. Here, we quantitatively connect the performance of phase estimation protocols with quantum coherence. To achieve this, we construct and characterize resource theories of quantum networks that

  2. Baolei Zhang, Haoran Xin, Jiatong Li, Dongzhe Zhang

    Retrieval-Augmented Generation (RAG) has proven effective in mitigating hallucinations in large language models by incorporating external knowledge during inference. However, this integration introduces new security vulnerabilities, particularly to poisoning attacks. Although prior work has explored various poisoning strategies, a thorough assessment of thei

  3. Zhaoyuan Su, Zeyu Zhang, Tingfeng Lan, Zirui Wang

    Efficiently serving large language models (LLMs) under dynamic and bursty workloads remains a key challenge for real-world deployment. Existing serving frameworks and static model compression techniques fail to adapt to workload fluctuations, leading to either service-level objective (SLO) violations under full-precision serving or persistent accuracy degrad

  4. Yongjie Wang, Jonathan Leung, Zhiqi Shen

    Large Language Models (LLMs) have shown promise in character imitation, enabling immersive and engaging conversations. However, they often generate content that is irrelevant or inconsistent with a character's background. We attribute these failures to: (1) the inability to accurately recall character-specific knowledge due to entity ambiguity, and (2) a lac

  5. Peter Lichard

    We show that the $\psi(2\mathrm S)$ subthreshold pole influences the cross section of the electron-positron annihilation into the $\mathrm D^+\mathrm D^-$ and $\mathrm D^0\bar\mathrm D^0$ final states. We perform a fit to the merged BES \cite{bes2008} and BESIII \cite{besiii2024} data, providing the cross section for $e^+e^-$ annihilation into those final st

  6. Mrinmoy Samanta, Sudipta Mondal, Aditi Sen De

    We compare the multipartite entangling and disentangling powers of unitary operators by assessing their ability to generate or eliminate genuine multipartite entanglement. Our findings reveal that while diagonal unitary operators can exhibit equal entangling and disentangling powers, certain non-diagonal unitaries demonstrate an imbalance when acting on full

  7. Xin Wei, Huakun Liu, Yutaro Hirao, Monica Perusquia-Hernandez

    Refractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical oversight. This study explores a passive method for estimating refractive power using two eye movement recording techniques: electrooculography (EOG) and video-based eye tracking. Using a publicly available datas

  8. Siyu Lei, Ze-Huan Zheng, Qilin Duan, Feng Wu

    The evolutions of polarization singularities, including bound states in the continuum (BICs) and circularly polarized states (C points), are usually realized by tuning the geometric parameters of photonic crystal slabs. Here, we use the off-diagonal terms of permittivity tensor to manipulate polarization singularities without breaking the structural symmetry

  9. Haoyuan Sun, Jiaqi Wu, Bo Xia, Yifu Luo

    Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient appli

  10. Dmitry Dudukalov, Artem Logachov, Vladimir Lotov, Timofei Prasolov

    We study the convergence properties and escape dynamics of Stochastic Gradient Descent (SGD) in one-dimensional landscapes, separately considering infinite- and finite-variance noise. Our main focus is to identify the time scales on which SGD reliably moves from an initial point to the local minimum in the same ''basin''. Under suitable conditions on the noi

  11. Marziyeh Rezaei, Dan Sturm, Pengyu Zeng, Sajjad Moazeni

    Optical interconnects are becoming a major bottleneck in scaling up future GPU racks and network switches within data centers. Although 200 Gb/s optical transceivers using PAM-4 modulation have been demonstrated, achieving higher data rates and energy efficiencies requires high-order coherent modulations like 16-QAM. Current coherent links rely on energy-int

  12. Xiaobin Rong, Dahan Wang, Qinwen Hu, Yushi Wang

    Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a sep

  13. Mingyang Wu, Li Lin, Wenbin Zhang, Xin Wang

    The Area Under the ROC Curve (AUC) is a key metric for classification, especially under class imbalance, with growing research focus on optimizing AUC over accuracy in applications like medical image analysis and deepfake detection. This leads to fairness in AUC optimization becoming crucial as biases can impact protected groups. While various fairness mitig

  14. Jiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun

    Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, e.g., reinforcement learning from human feedback (RLHF). Generative reward models

  15. Pengyu Wang, Shuchang Ye, Usman Naseem, Jinman Kim

    Medical Large Vision-Language Models (Med-LVLMs) have been widely adopted for medical report generation. Despite Med-LVLMs producing state-of-the-art performance, they exhibit a bias toward predicting all findings as normal, leading to reports that overlook critical abnormalities. Furthermore, these models often fail to provide comprehensive descriptions of

  16. Yiqing Zhang, Xiaozhong Liu, Fabricio Murai

    Many existing models for clinical trial outcome prediction are optimized using task-specific loss functions on trial phase-specific data. While this scheme may boost prediction for common diseases and drugs, it can hinder learning of generalizable representations, leading to more false positives/negatives. To address this limitation, we introduce CLaDMoP, a

  17. Yunqin Zhu, Henry Shaowu Yuchi, Yao Xie

    Learning expressive kernels while retaining tractable inference remains a central challenge in scaling Gaussian processes (GPs) to large and complex datasets. We propose a scalable GP regressor based on deep basis kernels (DBKs). Our DBK is constructed from a small set of neural-network-parameterized basis functions with an explicit low-rank structure. This

  18. Haoyu Yang, Yutong Guan, Meixing Shi, Yuxiang Cai

    3D medical image segmentation is important for clinical diagnosis and treatment but faces challenges from high-dimensional data and complex spatial dependencies. Traditional single-modality networks, such as CNNs and Transformers, are often limited by computational inefficiency and constrained contextual modeling in 3D settings. To alleviate these limitation

  19. Guowei Xu, Mert Yuksekgonul, Carlos Guestrin, James Zou

    Large language models (LLMs) are increasingly used in learning algorithms, evaluations, and optimization tasks. Recent studies have shown that using LLM-based optimizers to automatically optimize model prompts, demonstrations, predictions themselves, or other components can significantly enhance the performance of AI systems, as demonstrated by frameworks su

  20. Sidra Malik, Muneera Bano, Didar Zowghi

    Growing awareness of social biases and inequalities embedded in Artificial Intelligence (AI) systems has brought increased attention to the integration of Diversity and Inclusion (D&I) principles throughout the AI lifecycle. Despite the rise of ethical AI guidelines, there is limited empirical evidence on how D&I is applied in real-world settings. This study

  21. Xin Lu, Yanyan Zhao, Si Wei, Shijin Wang

    Pre-trained language models represented by the Transformer have been proven to possess strong base capabilities, and the representative self-attention mechanism in the Transformer has become a classic in sequence modeling architectures. Different from the work of proposing sequence modeling architecture to improve the efficiency of attention mechanism, this

  22. Ritwik Murali, C Shunmuga Velayutham

    It is well known that anti-malware scanners depend on malware signatures to identify malware. However, even minor modifications to malware code structure results in a change in the malware signature thus enabling the variant to evade detection by scanners. Therefore, there exists the need for a proactively generated malware variant dataset to aid detection o

  23. Zhi-Yu Xiao, Zixiang Lu, Yixiao Chen, Tao Xiang

    We introduce an efficient approach to implement correlated many-body trial wave functions in auxiliary-field quantum Monte Carlo (AFQMC). To control the sign/phase problem in AFQMC, a constraint is derived from an exact gauge condition but is typically imposed approximately through a trial wave function or trial density matrix, whose quality can affect the a

  24. Litao Ye, Bin Chen, Chen Sun, Shuo Wang

    Current Wi-Fi authentication methods face issues such as insufficient security, user privacy leakage, high management costs, and difficulty in billing. To address these challenges, a Wi-Fi access control solution based on blockchain smart contracts is proposed. Firstly, semi-fungible Wi-Fi tokens (SFWTs) are designed using the ERC1155 token standard as crede

  25. Pooneh Mousavi, Shubham Gupta, Cem Subakan, Mirco Ravanelli

    Foundation models based on large language models (LLMs) have shown great success in handling various tasks and modalities. However, adapting these models for general-purpose audio-language tasks is challenging due to differences in acoustic environments and task variations. In this work, we introduce LiSTEN Learning Soft Token Embeddings for Neural Audio LLM

  26. Xiangyu Zhang, Fuming Fang, Peng Gao, Bin Qin

    Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the integration of speech into LLM frameworks, often by tokenizing continuous audio via neural speech codecs, enabling powerful speech language models. However, this dominant tokenizat

  27. Taeckyung Lee, Sorn Chottananurak, Junsu Kim, Jinwoo Shin

    Deep learning models perform poorly when domain shifts exist between training and test data. Test-time adaptation (TTA) is a paradigm to mitigate this issue by adapting pre-trained models using only unlabeled test samples. However, existing TTA methods can fail under severe domain shifts, while recent active TTA approaches requiring full-class labels are imp

  28. Weiwei Sun, Haokun Liu, Nikhil Kandpal, Colin Raffel

    Training data attribution (TDA) methods aim to measure how training data impacts a model's predictions. While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications. Representation-based approaches are far more scalable, but typically rely on h

  29. Soyoung Yoon, Gyuwan Kim, Gyu-Hwung Cho, Seung-won Hwang

    Listwise reranking with large language models (LLMs) enhances top-ranked results in retrieval-based applications. Due to the limit in context size and high inference cost of long context, reranking is typically performed over a fixed size of small subsets, with the final ranking aggregated from these partial results. This fixed computation disregards query d

  30. Promit Chakroborty, Michael D. Shields

    To ensure that real-world infrastructure is safe and durable, systems are designed to not fail for any but the most rarely occurring parameter values. By only happening deep in the tails of the parameter distribution, failure probabilities are kept small. At the same time, it is essential to understand the risk associated with the failure of a system, no mat

  31. Sayan Bagchi, Md Nurul Molla, Joydwip Singh

    This paper is devoted to the study of $L^{p_1} \times L^{p_2}$ to $L^{p}$ boundedness of the bilinear Bochner-Riesz mean $\mathcal{B}^{\alpha}$ associated with the Grushin operator $\mathcal{L} = -\Delta_{x'} - |x'|^2 \Delta_{x''}$ on $\mathbb{R}^{d_1} \times \mathbb{R}^{d_2}$. Our result almost resembles the corresponding Euclidean results, where the Euclid

  32. Kenneth M. Zick

    For the past 25 years, the Gset benchmark problems have challenged all manner of Ising and Max-Cut solvers. The largest of these problems have remained unsolved by any heuristic algorithm. In this report we provide data showing dramatically better speed and accuracy on these large sparse problems. Our newly discovered heuristic algorithm called Cosm reaches

  33. K. Medler, C. Ashall, M. Shahbandeh, J. M. DerKacy

    We present the first data release of the Hawaii Infrared Supernova Study (\textit{HISS}), consisting of a large sample of near-infrared (NIR) spectra, $0.7 - 2.5 \mathrm{\mu m}$, obtained with the Keck-II/NIRES and IRTF/SpeX spectrographs. This sample is comprised of 90 NIR spectra of 48 transient events, spanning from hours after explosion to $\geq + 350$ d

  34. Yongzheng Li, Wanchen Yang, Shuai S. A. Yuan, Zhitao Ye

    Theoretically, the three-dimensional (3D) array architecture provides a higher communication degree of freedom (DoF) compared to the planar arrays, allowing for greater capacity potential in multiple-input multiple-output (MIMO) systems. However, in practical implementations, the upper elements of 3D arrays significantly degrade the performance of the lower

  35. Yixuan Ma, Kai Yi, Pietro Lio, Shi Jin

    Hypergraphs effectively model higher-order relationships in natural phenomena, capturing complex interactions beyond pairwise connections. We introduce a novel hypergraph message passing framework inspired by interacting particle systems, where hyperedges act as fields inducing shared node dynamics. By incorporating attraction, repulsion, and Allen-Cahn forc

  36. Milo Bechtloff Weising

    We study symmetric function analogues of the higher order Bell numbers. Their construction involves iterated plethystic exponential towers mimicking the single variable exponential generating functions for the higher order Bell numbers. We derive explicit recurrence relations for the expansion coefficients of the Bell functions into the monomial and power su

  37. Aofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia

    Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A$^3$Tune, a nove

  38. Guodong Du, Zhuo Li, Xuanning Zhou, Junlin Li

    Cross-capability transfer is a key challenge in large language model (LLM) research, with applications in multi-task integration, model compression, and continual learning. Recent works like FuseLLM and FuseChat have demonstrated the potential of transferring multiple model capabilities to lightweight models, enhancing adaptability and efficiency, which moti

  39. Sanjay Roy, T. K. Samanta

    The concept of fixed point plays a crucial role in various fields of applied mathematics. The aim of this paper is to establish the existence of a unique fixed point of some type of functions which satisfy a new contraction principle, namely, TSR-contraction principle in various types of probabilistic metric spaces. The proposed contraction mapping is differ

  40. Xiaojun Guo, Ang Li, Yifei Wang, Stefanie Jegelka

    Although Large Language Models (LLMs) have demonstrated remarkable progress, their proficiency in graph-related tasks remains notably limited, hindering the development of truly general-purpose models. Previous attempts, including pretraining graph foundation models or employing supervised fine-tuning, often face challenges such as the scarcity of large-scal

  41. Jingguang Tian, Xinhui Hu, Xinkang Xu

    In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To address this issue, we propose multiple improvements to train speaker encoders to increase emotion robustness. Firstly, w

  42. Kefan Yu, Qingcheng Zeng, Weihao Xuan, Wanxin Li

    Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning, both of which require substantial pragmatic understanding. However, how LLMs acquire this pragmatic competence throughout the training process remains poorly understood. In this work, we int

  43. Zhejunyu Jin, Tianci Gong, Jie Liu, Huanhuan Yang

    Altermagnets recently are identified as a new class of magnets that break the time-reversal symmetry without exhibiting net magnetization. The role of the dipole-dipole interaction (DDI) on their dynamical properties however is yet to be addressed. In this work, we show that the DDI can induce the strong coupling between exchange magnons with opposite chiral

  44. Chen-Hao Chao, Wei-Fang Sun, Hanwen Liang, Chun-Yi Lee

    Masked diffusion models (MDM) are powerful generative models for discrete data that generate samples by progressively unmasking tokens in a sequence. Each token can take one of two states: masked or unmasked. We observe that token sequences often remain unchanged between consecutive sampling steps; consequently, the model repeatedly processes identical input

  45. Zihao Peng, Jiandian Zeng, Boyuan Li, Guo Li

    Federated Learning (FL) facilitates the fine-tuning of Foundation Models (FMs) using distributed data sources, with Low-Rank Adaptation (LoRA) gaining popularity due to its low communication costs and strong performance. While recent work acknowledges the benefits of heterogeneous LoRA in FL and introduces flexible algorithms to support its implementation, o

  46. Xiang Li, Yunai Li, Huiying Zhong, Lihua Lei

    Performativity of predictions refers to the phenomenon where prediction-informed decisions influence the very targets they aim to predict -- a dynamic commonly observed in policy-making, social sciences, and economics. In this paper, we initiate an end-to-end framework of statistical inference under performativity. Our contributions are twofold. First, we es

  47. Hao Gu, Lujun Li, Hao Wang, Lei Wang

    Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency. While recent sparsity-aware binarization achieves sub-1-bit compression via weight pruning, it faces critical challenges: performance degradation, mask-management overhead, and limited hardware compatibility. In this

  48. A. Nafis Arafat, Oleg L. Berman, Godfrey Gumbs, Peter B. Littlewood

    We develop a microscopic mean-field theory describing the coexistence of Bose-Einstein condensates of upper and lower polaritons (UP/LP) in a semiconductor microcavity. Incorporating interbranch scattering within a modified polariton Hamiltonian, we introduce a phenomenological population-split parameter $\alpha$ that quantifies the relative LP/UP occupation

  49. Xuan Xiao, Xiaotong Ren, Haitao Li

    Accurately estimating vehicle velocity via smartphone is critical for mobile navigation and transportation. This paper introduces a cutting-edge framework for velocity estimation that incorporates temporal learning models, utilizing Inertial Measurement Unit (IMU) data and is supervised by Global Navigation Satellite System (GNSS) information. The framework

  50. Jeehoon Park, Jaewon Yoo

    We provide a new $L^2$-Hodge theoretic construction of a Frobenius manifold structure on the cohomology of a Calabi-Yau smooth projective hypersurface $V$, using Li-Wen's $L^2$-Hodge theory [9] of a Landau-Ginzburg model with compact critical locus $V$. We also give a precise comparison result between the current construction and Barannikov-Kontsevich's cons

  51. Yanxiang Zhang, Zheng Xu, Shanshan Wu, Yuanbo Zhang

    Error correction is an important capability when applying large language models (LLMs) to facilitate user typing on mobile devices. In this paper, we use LLMs to synthesize a high-quality dataset of error correction pairs to evaluate and improve LLMs for mobile applications. We first prompt LLMs with error correction domain knowledge to build a scalable and

  52. Junlin Wang, Zhiyun Lin

    Learning effective visual representations for robotic manipulation remains a fundamental challenge due to the complex body dynamics involved in action execution. In this paper, we study how visual representations that carry body-relevant cues can enable efficient policy learning for downstream robotic manipulation tasks. We present $\textbf{I}$nter-token $\t

  53. Hong Jiao, Dan Song, Won-Chan Lee

    Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and induces least rater effects needs to be collected before the use of LLMs for automated scoring in practice. This study compared ten LLMs (Ch

  54. Timothy Do, Pranav Saran, Harshita Poojary, Pranav Prabhu

    In this paper, we address the persistent challenges that figurative language expressions pose for natural language processing (NLP) systems, particularly in low-resource languages such as Konkani. We present a hybrid model that integrates a pre-trained Multilingual BERT (mBERT) with a bidirectional LSTM and a linear classifier. This architecture is fine-tune

  55. Shengzhe Xu, Nikhil Muralidhar, Naren Ramakrishnan

    Numerous recent prompt optimization approaches like chain-of-thought, have been demonstrated to significantly improve the quality of content generated by large language models (LLMs). In-context learning (ICL), a recent paradigm where a few representative examples guide content generation has also led to strong improvements in generation quality of LLM gener

  56. Jule Valendo Halim, Siyi Wang, Hong Jia, Ting Dang

    Emotional intelligence in conversational AI is crucial across domains like human-computer interaction. While numerous models have been developed, they often overlook the complexity and ambiguity inherent in human emotions. In the era of large speech foundation models (SFMs), understanding their capability in recognizing ambiguous emotions is essential for th

  57. Ian McCulloh, Pedro Rodriguez, Srivaths Kumar, Manu Gupta

    The increasing demand for digital literacy and artificial intelligence (AI) fluency in the workforce has highlighted the need for scalable, efficient programming instruction. This study evaluates the effectiveness of integrating generative AI, specifically OpenAIs ChatGPT, into a self-paced Python programming module embedded within a sixteen-week professiona

  58. Hongjia Wu, Hongxin Zhang, Wei Chen, Jiazhi Xia

    Various industries have produced a large number of documents such as industrial plans, technical guidelines, and regulations that are structurally complex and content-wise fragmented. This poses significant challenges for experts and decision-makers in terms of retrieval and understanding. Although existing LLM-based Retrieval-Augmented Generation methods ca

  59. Ezra Brooker, Andrey Zhiglo, Tomasz Plewa

    The aim of this work is to characterize the thermodynamic state of fuel mixed into the turbulent flame brush in the context of the Zel'dovich deflagration-to-detonation transition (ZDDT) mechanism of Type Ia supernovae (SNe Ia). We perform a series of three-dimensional computer simulations of thermonuclear deflagrations subject to the Rayleigh-Taylor instabi

  60. James MacLaurin, Pedro Vilanova

    The theory of Balanced Neural Networks is a very popular explanation for the high degree of variability and stochasticity in the brain's activity. Roughly speaking, it entails that typical neurons receive many excitatory and inhibitory inputs. The network-wide mean inputs cancel, and one is left with the stochastic fluctuations about the mean. In this paper

  61. Lise-Marie Imbert-Gerard

    Trefftz-type of Galerkin methods for numerical PDEs use discrete spaces of problem-dependent functions. While Trefftz methods leverage discrete spaces of local exact solutions to the governing PDE, Taylor-based quasi-Trefftz methods leverage discrete spaces of local approximate solutions to the governing PDE. This notion of approximate solution, understood i

  62. Li-Syun Hsiung, Jun-Kai Tu, Kuan-Wu Chu, Yu-Hsuan Chiu

    This study aims to investigate the challenge of insufficient three-dimensional context in synthetic datasets for scene text rendering. Although recent advances in diffusion models and related techniques have improved certain aspects of scene text generation, most existing approaches continue to rely on 2D data, sourcing authentic training examples from movie

  63. Lucas Tecot, Di Luo, Cho-Jui Hsieh

    Advancements in quantum computing have spurred significant interest in harnessing its potential for speedups over classical systems. However, noise remains a major obstacle to achieving reliable quantum algorithms. In this work, we present a provably noise-resilient training theory and algorithm to enhance the robustness of parameterized quantum circuit clas

  64. Fukun Liu, Adam T. Greer, Gengchen Mai, Jin Sun

    Plankton are small drifting organisms found throughout the world's oceans and can be indicators of ocean health. One component of this plankton community is the zooplankton, which includes gelatinous animals and crustaceans (e.g. shrimp), as well as the early life stages (i.e., eggs and larvae) of many commercially important fishes. Being able to monitor zoo

  65. Colin P. Folsom, Christiana Erba, Veronique Petit, Shaquann Seadrow

    Spectropolarimetry, the observation of polarization and intensity as a function of wavelength, is a powerful tool in stellar astrophysics. It is particularly useful for characterizing stars and circumstellar material, and for tracing the influence of magnetic fields on a host star and its environment. Maintaining modern, flexible, and accessible computationa

  66. Mengran Li, Pengyu Zhang, Wenbin Xing, Yijia Zheng

    Graphs are a widely used paradigm for representing non-Euclidean data, with applications ranging from social network analysis to biomolecular prediction. While graph learning has achieved remarkable progress, real-world graph data presents a number of challenges that significantly hinder the learning process. In this survey, we focus on four fundamental data

  67. Zhiyuan Zhang, Zhengtong Xu, Jai Nanda Lakamsani, Yu She

    Visual Imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy archite

  68. Sebastian Gutierrez Hernandez, Peng Chen, Haomin Zhou

    We introduce Parametric Density Path Optimization (PDPO), a novel method for computing action-minimizing paths between probability densities. The core idea is to represent the target probability path as the pushforward of a reference density through a parametric map, transforming the original infinite-dimensional optimization over densities to a finite-dimen

  69. Quan Khanh Luu, Pokuang Zhou, Zhengtong Xu, Zhiyuan Zhang

    Supervised visuomotor policies have shown strong performance in robotic manipulation but often struggle in tasks with limited visual inputs, such as operations in confined spaces and dimly lit environments, or tasks requiring precise perception of object properties and environmental interactions. In such cases, tactile feedback becomes essential for manipula

  70. Guoheng Sun, Ziyao Wang, Xuandong Zhao, Bowei Tian

    Modern large language model (LLM) services increasingly rely on complex, often abstract operations, such as multi-step reasoning and multi-agent collaboration, to generate high-quality outputs. While users are billed based on token consumption and API usage, these internal steps are typically not visible. We refer to such systems as Commercial Opaque LLM Ser

  71. Christopher J. Mungall, Adnan Malik, Daniel R. Korn, Justin T. Reese

    Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring

  72. Jingkai Wang, Wu Miao, Jue Gong, Zheng Chen

    Face restoration has achieved significant advancements through the years of development. However, maintaining high fidelity and authenticity while avoiding artifacts remains challenging, especially in extreme degradation scenarios. This highlights the need for models that are more ``honest'' in their reconstruction from low-quality inputs, accurately

  73. Paulo E. Faria Junior, Daniel Hernangómez-Pérez, Tomer Amit, Jaroslav Fabian

    Magneto-optics of low dimensional semiconductors, such as monolayer transition metal dichalcogenides, offers a vast playground for exploring complex quantum phenomena. However, current ab initio approaches fail to capture important experimental observations related to brightening of excitonic levels and their g-factor dependence. Here, we develop a robust an

  74. Unggi Lee, Jaeyong Lee, Jiyeong Bae, Yeil Jeong

    Recent advances in large reasoning models (LRMs) show strong performance in structured domains such as mathematics and programming; however, they often lack pedagogical coherence and realistic teaching behaviors. To bridge this gap, we introduce Pedagogy-R1, a framework that adapts LRMs for classroom use through three innovations: (1) a distillation-based pi

  75. Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, Ziwei Zhu

    Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. This work addresses these gaps by conducting a multilingual and intersectional analysis of LLM outputs across 10 Indo-Aryan and Dravidian languages, identifying how cultural stigmas i

  76. Ugur Kursuncu, Trilok Padhi, Gaurav Sinha, Abdulkadir Erol

    The growing demand for accessible mental health support, compounded by workforce shortages and logistical barriers, has led to increased interest in utilizing Large Language Models (LLMs) for scalable and real-time assistance. However, their use in sensitive domains such as anxiety support remains underexamined. This study presents a systematic evaluation of

  77. Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang, Mohan Shi

    Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, leading to high transcription error rates. While ASR error correction (AEC) methods have improved adult speech transcription, their effectiveness on child speech remains largely unexp

  78. Xiaofei Huang, Kai Wei, Yang Rui, Dinghui Gong

    Atomic spin sensors are essential for beyond-the-standard-model exploration, biomagnetic measurement, and quantum navigation. While the traditional DC mode spin-exchange relaxation-free (SERF) comagnetometer achieves ultrahigh sensitivity, further improvements require suppressing technical noise and surpassing standard quantum limit. In this work, we develop

  79. Morteza Karimzadeh, Zhongying Wang, James L. Crooks

    Deep learning has shown strong performance in geospatial prediction tasks, but the role of geolocation information in improving accuracy and generalizability remains underexamined. Recent work has introduced location encoders that aim to represent spatial context in a transferable way, yet most evaluations have focused on static mapping tasks. Here, we study

  80. Viacheslav Tsaran, Francesco Marino, Sonia Bacca, Francesca Bonaiti

    We extend the pion-nucleus multiple-scattering framework to include detailed second-order rescattering dynamics for nuclei with non-zero isospin. To account for intermediate charge-exchange and nucleon spin-flip effects, we develop a scattering potential that depends on the one- and two-body densities of the target nucleus. We compute one-body densities from

  81. Xuanhe Zhou, Junxuan He, Wei Zhou, Haodong Chen

    The integration of large language model (LLM) and data management (DATA) is rapidly redefining both domains. In this survey, we comprehensively review the bidirectional relationships. On the one hand, DATA4LLM, spanning large-scale data processing, storage, and serving, feeds LLMs with high quality, diversity, and timeliness of data required for stages like

  82. Abir Ray

    This paper introduces EdgeAgentX, a novel framework integrating federated learning (FL), multi-agent reinforcement learning (MARL), and adversarial defense mechanisms, tailored for military communication networks. EdgeAgentX significantly improves autonomous decision-making, reduces latency, enhances throughput, and robustly withstands adversarial disruption

  83. Litu Rout, Constantine Caramanis, Sanjay Shakkottai

    Diffusion Language Models (DLMs) promise parallel generation and bidirectional context, yet they underperform autoregressive (AR) models in both likelihood modeling and generated text quality. We identify that this performance gap arises when important tokens (e.g., key words or low-frequency words that anchor a sentence) are masked early in the forward proc

  84. Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian

    The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its popularity and relevance, the theoretical properties of the softmax-

  85. Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang

    Recent advances in large language models (LLMs) have introduced latent reasoning as a promising alternative to autoregressive reasoning. By performing internal computation with hidden states from previous steps, latent reasoning benefit from more informative features rather than sampling a discrete chain-of-thought (CoT) path. Yet latent reasoning approaches

  86. Zhichao Wu, Yueteng Kang, Songjun Cao, Long Ma

    Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on multi-modal prompt. The system disentangles speech into the content, timbre, emotion and prosody, allowing emotion promp

  87. Heyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, Mark Dredze

    While Large Language Models (LLMs) can generate fluent and convincing responses, they are not necessarily correct. This is especially apparent in the popular decompose-then-verify factuality evaluation pipeline, where LLMs evaluate generations by decomposing the generations into individual, valid claims. Factuality evaluation is especially important for medi

  88. Toshiaki Koike-Akino, Jing Liu, Ye Wang

    To tackle the huge computational demand of large foundation models, activation-aware compression techniques without retraining have been introduced. However, since these rely on calibration data, domain shift may arise for unknown downstream tasks. With a computationally efficient calibration, activation-aware pruning can be executed for every prompt adaptiv

  89. Ainulla Khan, Yamada Moyuru, Srinidhi Akella

    Retrieval-Augmented Generation (RAG) has emerged as a promising technique to enhance the quality and relevance of responses generated by large language models. While recent advancements have mainly focused on improving RAG for text-based queries, RAG on multi-modal documents containing both texts and images has not been fully explored. Especially when fine-t

  90. Sarasija Sudharsan, Anupam Sharma

    This paper presents a numerical demonstration of the real-time application of two dynamic stall onset criteria for identifying and mitigating stall. These criteria - based on the leading-edge suction parameter (LESP) and boundary enstrophy flux (BEF) - are derived from prior research. The present work establishes a proof of concept for the practical use of t

  91. Yucheng Guo, Qinxin Yan

    We introduce a family of particle systems on sparse graphs where local interactions occur via hitting times, providing a dynamic and tractable model for default cascades in large sparsely-connected financial networks. Building on the framework of Lacker, Ramanan and Wu (2023), we extend convergence theory to systems with singular interactions, capturing the

  92. Chi Zhang, Ziying Jia, George K. Atia, Sihong He

    Transfer reinforcement learning aims to derive a near-optimal policy for a target environment with limited data by leveraging abundant data from related source domains. However, it faces two key challenges: the lack of performance guarantees for the transferred policy, which can lead to undesired actions, and the risk of negative transfer when multiple sourc

  93. Hojun Son, Asma Almutairi, Arpan Kusari

    Context bias refers to the association between the foreground objects and background during the object detection training process. Various methods have been proposed to minimize the context bias when applying the trained model to an unseen domain, known as domain adaptation for object detection (DAOD). But a principled approach to understand why the context

  94. J. -F. Wang, G. Qin

    In this article, momentum transport generated by the combined effects of pitch-angle diffusion and Background Flow Velocity Inhomogeneities (BFVIs) is proposed to obtain a cosmic rays acceleration mechanism, starting from the well-known focusing equation describing particle diffusion and acceleration. The inhomogeneities of background flow velocity is ubiqui

  95. Yiren Song, Cheng Liu, Mike Zheng Shou

    Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, composition, and fine details, and (2) preventing style degradation in image-to-image pipelines with style LoRAs. GPT-4o's exceptional stylization consistency highlights the performa

  96. Christian D. Newman, Anthony Peruma, Eman Abdullah AlOmar, Mahie Crabbe

    Identifier names are crucial components of code, serving as primary clues for developers to understand program behavior. This paper investigates the linguistic structure of identifier names by extending the concept of grammar patterns, which represent the part-of-speech (PoS) sequences underlying identifier phrases. The specific focus is on closed syntactic

  97. Serkan Hoşten

    Toric ideals are everywhere. They have been in the commutative algebra lexicon since about 1990 when Bernd Sturmfels used the term. The early days of toric ideals and their Gr\"obner bases were full of new results and promising developments in their applications. Bernd has been consistently their biggest promoter through his own work and that of his collabor

  98. Zhining Liu, Ze Yang, Xiao Lin, Ruizhong Qiu

    Time-series forecasting plays a critical role in many real-world applications. Although increasingly powerful models have been developed and achieved superior results on benchmark datasets, through a fine-grained sample-level inspection, we find that (i) no single model consistently outperforms others across different test samples, but instead (ii) each mode

  99. Romeo Valentin, Sydney M. Katz, Vincent Vanhoucke, Mykel J. Kochenderfer

    Dictionary learning has recently emerged as a promising approach for mechanistic interpretability of large transformer models. Disentangling high-dimensional transformer embeddings requires algorithms that scale to high-dimensional data with large sample sizes. Recent work has explored sparse autoencoders (SAEs) for this problem. However, SAEs use a simple l

  100. Zhaoyang Wang, Jinqi Jiang, Tian Qiu, Hui Liu

    Recent large reasoning models such as DeepSeek-R1 exhibit strong complex problems solving abilities by generating long chain-of-thought (CoT) reasoning steps. It is challenging to directly train small language models (SLMs) to emerge long CoT. Thus, distillation becomes a practical method to enable SLMs for such reasoning ability. However, the long CoT often