May 2025 arXiv papers — page 70
Showing 6,901–7,000 of 24,552 papers
Felix Ahnefeld, Thomas Theurer, Martin B. Plenio
Quantum phase estimation is a core task in quantum technologies ranging from metrology to quantum computing, where it appears as a key subroutine in various algorithms. Here, we quantitatively connect the performance of phase estimation protocols with quantum coherence. To achieve this, we construct and characterize resource theories of quantum networks that
Baolei Zhang, Haoran Xin, Jiatong Li, Dongzhe Zhang
Retrieval-Augmented Generation (RAG) has proven effective in mitigating hallucinations in large language models by incorporating external knowledge during inference. However, this integration introduces new security vulnerabilities, particularly to poisoning attacks. Although prior work has explored various poisoning strategies, a thorough assessment of thei
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
cs.DCZhaoyuan Su, Zeyu Zhang, Tingfeng Lan, Zirui Wang
Efficiently serving large language models (LLMs) under dynamic and bursty workloads remains a key challenge for real-world deployment. Existing serving frameworks and static model compression techniques fail to adapt to workload fluctuations, leading to either service-level objective (SLO) violations under full-precision serving or persistent accuracy degrad
Yongjie Wang, Jonathan Leung, Zhiqi Shen
Large Language Models (LLMs) have shown promise in character imitation, enabling immersive and engaging conversations. However, they often generate content that is irrelevant or inconsistent with a character's background. We attribute these failures to: (1) the inability to accurately recall character-specific knowledge due to entity ambiguity, and (2) a lac
Peter Lichard
We show that the $\psi(2\mathrm S)$ subthreshold pole influences the cross section of the electron-positron annihilation into the $\mathrm D^+\mathrm D^-$ and $\mathrm D^0\bar\mathrm D^0$ final states. We perform a fit to the merged BES \cite{bes2008} and BESIII \cite{besiii2024} data, providing the cross section for $e^+e^-$ annihilation into those final st
Disparity between multipartite entangling and disentangling powers of unitaries: Even vs Odd
quant-phMrinmoy Samanta, Sudipta Mondal, Aditi Sen De
We compare the multipartite entangling and disentangling powers of unitary operators by assessing their ability to generate or eliminate genuine multipartite entanglement. Our findings reveal that while diagonal unitary operators can exhibit equal entangling and disentangling powers, certain non-diagonal unitaries demonstrate an imbalance when acting on full
Mind Your Vision: Multimodal Estimation of Refractive Disorders Using Electrooculography and Eye Tracking
eess.IVXin Wei, Huakun Liu, Yutaro Hirao, Monica Perusquia-Hernandez
Refractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical oversight. This study explores a passive method for estimating refractive power using two eye movement recording techniques: electrooculography (EOG) and video-based eye tracking. Using a publicly available datas
Effects of off-diagonal permittivity terms on polarization singularities in anisotropic grating system
physics.opticsSiyu Lei, Ze-Huan Zheng, Qilin Duan, Feng Wu
The evolutions of polarization singularities, including bound states in the continuum (BICs) and circularly polarized states (C points), are usually realized by tuning the geometric parameters of photonic crystal slabs. Here, we use the off-diagonal terms of permittivity tensor to manipulate polarization singularities without breaking the structural symmetry
Haoyuan Sun, Jiaqi Wu, Bo Xia, Yifu Luo
Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient appli
Dmitry Dudukalov, Artem Logachov, Vladimir Lotov, Timofei Prasolov
We study the convergence properties and escape dynamics of Stochastic Gradient Descent (SGD) in one-dimensional landscapes, separately considering infinite- and finite-variance noise. Our main focus is to identify the time scales on which SGD reliably moves from an initial point to the local minimum in the same ''basin''. Under suitable conditions on the noi
A DSP-Free Carrier Phase Recovery System using 16-Offset-QAM Laser Forwarded Links for 400Gb/s and Beyond
eess.SPMarziyeh Rezaei, Dan Sturm, Pengyu Zeng, Sajjad Moazeni
Optical interconnects are becoming a major bottleneck in scaling up future GPU racks and network switches within data centers. Although 200 Gb/s optical transceivers using PAM-4 modulation have been demonstrated, achieving higher data rates and energy efficiencies requires high-order coherent modulations like 16-QAM. Current coherent links rely on energy-int
Xiaobin Rong, Dahan Wang, Qinwen Hu, Yushi Wang
Universal speech enhancement aims to handle input speech with different distortions and input formats. To tackle this challenge, we present TS-URGENet, a Three-Stage Universal, Robust, and Generalizable speech Enhancement Network. To address various distortions, the proposed system employs a novel three-stage architecture consisting of a filling stage, a sep
Mingyang Wu, Li Lin, Wenbin Zhang, Xin Wang
The Area Under the ROC Curve (AUC) is a key metric for classification, especially under class imbalance, with growing research focus on optimizing AUC over accuracy in applications like medical image analysis and deepfake detection. This leads to fairness in AUC optimization becoming crucial as biases can impact protected groups. While various fairness mitig
Jiayi Zhou, Jiaming Ji, Boyuan Chen, Jiapeng Sun
Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, e.g., reinforcement learning from human feedback (RLHF). Generative reward models
Pengyu Wang, Shuchang Ye, Usman Naseem, Jinman Kim
Medical Large Vision-Language Models (Med-LVLMs) have been widely adopted for medical report generation. Despite Med-LVLMs producing state-of-the-art performance, they exhibit a bias toward predicting all findings as normal, leading to reports that overlook critical abnormalities. Furthermore, these models often fail to provide comprehensive descriptions of
Yiqing Zhang, Xiaozhong Liu, Fabricio Murai
Many existing models for clinical trial outcome prediction are optimized using task-specific loss functions on trial phase-specific data. While this scheme may boost prediction for common diseases and drugs, it can hinder learning of generalizable representations, leading to more false positives/negatives. To address this limitation, we introduce CLaDMoP, a
Yunqin Zhu, Henry Shaowu Yuchi, Yao Xie
Learning expressive kernels while retaining tractable inference remains a central challenge in scaling Gaussian processes (GPs) to large and complex datasets. We propose a scalable GP regressor based on deep basis kernels (DBKs). Our DBK is constructed from a small set of neural-network-parameterized basis functions with an explicit low-rank structure. This
Haoyu Yang, Yutong Guan, Meixing Shi, Yuxiang Cai
3D medical image segmentation is important for clinical diagnosis and treatment but faces challenges from high-dimensional data and complex spatial dependencies. Traditional single-modality networks, such as CNNs and Transformers, are often limited by computational inefficiency and constrained contextual modeling in 3D settings. To alleviate these limitation
Guowei Xu, Mert Yuksekgonul, Carlos Guestrin, James Zou
Large language models (LLMs) are increasingly used in learning algorithms, evaluations, and optimization tasks. Recent studies have shown that using LLM-based optimizers to automatically optimize model prompts, demonstrations, predictions themselves, or other components can significantly enhance the performance of AI systems, as demonstrated by frameworks su
Sidra Malik, Muneera Bano, Didar Zowghi
Growing awareness of social biases and inequalities embedded in Artificial Intelligence (AI) systems has brought increased attention to the integration of Diversity and Inclusion (D&I) principles throughout the AI lifecycle. Despite the rise of ethical AI guidelines, there is limited empirical evidence on how D&I is applied in real-world settings. This study
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
cs.CLXin Lu, Yanyan Zhao, Si Wei, Shijin Wang
Pre-trained language models represented by the Transformer have been proven to possess strong base capabilities, and the representative self-attention mechanism in the Transformer has become a classic in sequence modeling architectures. Different from the work of proposing sequence modeling architecture to improve the efficiency of attention mechanism, this
Ritwik Murali, C Shunmuga Velayutham
It is well known that anti-malware scanners depend on malware signatures to identify malware. However, even minor modifications to malware code structure results in a change in the malware signature thus enabling the variant to evade detection by scanners. Therefore, there exists the need for a proactively generated malware variant dataset to aid detection o
Implementing advanced trial wave functions in fermion quantum Monte Carlo via stochastic sampling
cond-mat.str-elZhi-Yu Xiao, Zixiang Lu, Yixiao Chen, Tao Xiang
We introduce an efficient approach to implement correlated many-body trial wave functions in auxiliary-field quantum Monte Carlo (AFQMC). To control the sign/phase problem in AFQMC, a constraint is derived from an exact gauge condition but is typically imposed approximately through a trial wave function or trial density matrix, whose quality can affect the a
Litao Ye, Bin Chen, Chen Sun, Shuo Wang
Current Wi-Fi authentication methods face issues such as insufficient security, user privacy leakage, high management costs, and difficulty in billing. To address these challenges, a Wi-Fi access control solution based on blockchain smart contracts is proposed. Firstly, semi-fungible Wi-Fi tokens (SFWTs) are designed using the ERC1155 token standard as crede
Pooneh Mousavi, Shubham Gupta, Cem Subakan, Mirco Ravanelli
Foundation models based on large language models (LLMs) have shown great success in handling various tasks and modalities. However, adapting these models for general-purpose audio-language tasks is challenging due to differences in acoustic environments and task variations. In this work, we introduce LiSTEN Learning Soft Token Embeddings for Neural Audio LLM
Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
eess.ASXiangyu Zhang, Fuming Fang, Peng Gao, Bin Qin
Large Language Models (LLMs) have demonstrated remarkable success across diverse fields, establishing a powerful paradigm for complex information processing. This has inspired the integration of speech into LLM frameworks, often by tokenizing continuous audio via neural speech codecs, enabling powerful speech language models. However, this dominant tokenizat
Taeckyung Lee, Sorn Chottananurak, Junsu Kim, Jinwoo Shin
Deep learning models perform poorly when domain shifts exist between training and test data. Test-time adaptation (TTA) is a paradigm to mitigate this issue by adapting pre-trained models using only unlabeled test samples. However, existing TTA methods can fail under severe domain shifts, while recent active TTA approaches requiring full-class labels are imp
Weiwei Sun, Haokun Liu, Nikhil Kandpal, Colin Raffel
Training data attribution (TDA) methods aim to measure how training data impacts a model's predictions. While gradient-based attribution methods, such as influence functions, offer theoretical grounding, their computational costs make them impractical for large-scale applications. Representation-based approaches are far more scalable, but typically rely on h
Soyoung Yoon, Gyuwan Kim, Gyu-Hwung Cho, Seung-won Hwang
Listwise reranking with large language models (LLMs) enhances top-ranked results in retrieval-based applications. Due to the limit in context size and high inference cost of long context, reranking is typically performed over a fixed size of small subsets, with the final ranking aggregated from these partial results. This fixed computation disregards query d
Promit Chakroborty, Michael D. Shields
To ensure that real-world infrastructure is safe and durable, systems are designed to not fail for any but the most rarely occurring parameter values. By only happening deep in the tails of the parameter distribution, failure probabilities are kept small. At the same time, it is essential to understand the risk associated with the failure of a system, no mat
Sayan Bagchi, Md Nurul Molla, Joydwip Singh
This paper is devoted to the study of $L^{p_1} \times L^{p_2}$ to $L^{p}$ boundedness of the bilinear Bochner-Riesz mean $\mathcal{B}^{\alpha}$ associated with the Grushin operator $\mathcal{L} = -\Delta_{x'} - |x'|^2 \Delta_{x''}$ on $\mathbb{R}^{d_1} \times \mathbb{R}^{d_2}$. Our result almost resembles the corresponding Euclidean results, where the Euclid
Performance report of heuristic algorithm that cracked the largest Gset Ising problems (G81 cut=14060)
cs.DSKenneth M. Zick
For the past 25 years, the Gset benchmark problems have challenged all manner of Ising and Max-Cut solvers. The largest of these problems have remained unsolved by any heuristic algorithm. In this report we provide data showing dramatically better speed and accuracy on these large sparse problems. Our newly discovered heuristic algorithm called Cosm reaches
K. Medler, C. Ashall, M. Shahbandeh, J. M. DerKacy
We present the first data release of the Hawaii Infrared Supernova Study (\textit{HISS}), consisting of a large sample of near-infrared (NIR) spectra, $0.7 - 2.5 \mathrm{\mu m}$, obtained with the Keck-II/NIRES and IRTF/SpeX spectrographs. This sample is comprised of 90 NIR spectra of 48 transient events, spanning from hours after explosion to $\geq + 350$ d
Capacity Enhancement Analysis and Implementation of a 3D Array Based on Miniaturized Dipole Antennas
physics.app-phYongzheng Li, Wanchen Yang, Shuai S. A. Yuan, Zhitao Ye
Theoretically, the three-dimensional (3D) array architecture provides a higher communication degree of freedom (DoF) compared to the planar arrays, allowing for greater capacity potential in multiple-input multiple-output (MIMO) systems. However, in practical implementations, the upper elements of 3D arrays significantly degrade the performance of the lower
Yixuan Ma, Kai Yi, Pietro Lio, Shi Jin
Hypergraphs effectively model higher-order relationships in natural phenomena, capturing complex interactions beyond pairwise connections. We introduce a novel hypergraph message passing framework inspired by interacting particle systems, where hyperedges act as fields inducing shared node dynamics. By incorporating attraction, repulsion, and Allen-Cahn forc
Milo Bechtloff Weising
We study symmetric function analogues of the higher order Bell numbers. Their construction involves iterated plethystic exponential towers mimicking the single variable exponential generating functions for the higher order Bell numbers. We derive explicit recurrence relations for the expansion coefficients of the Bell functions into the monomial and power su
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
cs.CVAofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia
Medical Large Vision-Language Models (Med-LVLMs) often exhibit suboptimal attention distribution on visual inputs, leading to hallucinated or inaccurate outputs. Existing mitigation methods primarily rely on inference-time interventions, which are limited in attention adaptation or require additional supervision. To address this, we propose A$^3$Tune, a nove
Guodong Du, Zhuo Li, Xuanning Zhou, Junlin Li
Cross-capability transfer is a key challenge in large language model (LLM) research, with applications in multi-task integration, model compression, and continual learning. Recent works like FuseLLM and FuseChat have demonstrated the potential of transferring multiple model capabilities to lightweight models, enhancing adaptability and efficiency, which moti
Sanjay Roy, T. K. Samanta
The concept of fixed point plays a crucial role in various fields of applied mathematics. The aim of this paper is to establish the existence of a unique fixed point of some type of functions which satisfy a new contraction principle, namely, TSR-contraction principle in various types of probabilistic metric spaces. The proposed contraction mapping is differ
Xiaojun Guo, Ang Li, Yifei Wang, Stefanie Jegelka
Although Large Language Models (LLMs) have demonstrated remarkable progress, their proficiency in graph-related tasks remains notably limited, hindering the development of truly general-purpose models. Previous attempts, including pretraining graph foundation models or employing supervised fine-tuning, often face challenges such as the scarcity of large-scal
Jingguang Tian, Xinhui Hu, Xinkang Xu
In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To address this issue, we propose multiple improvements to train speaker encoders to increase emotion robustness. Firstly, w
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models
cs.CLKefan Yu, Qingcheng Zeng, Weihao Xuan, Wanxin Li
Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning, both of which require substantial pragmatic understanding. However, how LLMs acquire this pragmatic competence throughout the training process remains poorly understood. In this work, we int
Zhejunyu Jin, Tianci Gong, Jie Liu, Huanhuan Yang
Altermagnets recently are identified as a new class of magnets that break the time-reversal symmetry without exhibiting net magnetization. The role of the dipole-dipole interaction (DDI) on their dynamical properties however is yet to be addressed. In this work, we show that the DDI can induce the strong coupling between exchange magnons with opposite chiral
Chen-Hao Chao, Wei-Fang Sun, Hanwen Liang, Chun-Yi Lee
Masked diffusion models (MDM) are powerful generative models for discrete data that generate samples by progressively unmasking tokens in a sequence. Each token can take one of two states: masked or unmasked. We observe that token sequences often remain unchanged between consecutive sampling steps; consequently, the model repeatedly processes identical input
Zihao Peng, Jiandian Zeng, Boyuan Li, Guo Li
Federated Learning (FL) facilitates the fine-tuning of Foundation Models (FMs) using distributed data sources, with Low-Rank Adaptation (LoRA) gaining popularity due to its low communication costs and strong performance. While recent work acknowledges the benefits of heterogeneous LoRA in FL and introduces flexible algorithms to support its implementation, o
Xiang Li, Yunai Li, Huiying Zhong, Lihua Lei
Performativity of predictions refers to the phenomenon where prediction-informed decisions influence the very targets they aim to predict -- a dynamic commonly observed in policy-making, social sciences, and economics. In this paper, we initiate an end-to-end framework of statistical inference under performativity. Our contributions are twofold. First, we es
Hao Gu, Lujun Li, Hao Wang, Lei Wang
Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency. While recent sparsity-aware binarization achieves sub-1-bit compression via weight pruning, it faces critical challenges: performance degradation, mask-management overhead, and limited hardware compatibility. In this
A. Nafis Arafat, Oleg L. Berman, Godfrey Gumbs, Peter B. Littlewood
We develop a microscopic mean-field theory describing the coexistence of Bose-Einstein condensates of upper and lower polaritons (UP/LP) in a semiconductor microcavity. Incorporating interbranch scattering within a modified polariton Hamiltonian, we introduce a phenomenological population-split parameter $\alpha$ that quantifies the relative LP/UP occupation
Xuan Xiao, Xiaotong Ren, Haitao Li
Accurately estimating vehicle velocity via smartphone is critical for mobile navigation and transportation. This paper introduces a cutting-edge framework for velocity estimation that incorporates temporal learning models, utilizing Inertial Measurement Unit (IMU) data and is supervised by Global Navigation Satellite System (GNSS) information. The framework
$L^2$-Hodge theoretic construction of Frobenius manifolds for Calabi-Yau smooth projective hypersurfaces
math.AGJeehoon Park, Jaewon Yoo
We provide a new $L^2$-Hodge theoretic construction of a Frobenius manifold structure on the cohomology of a Calabi-Yau smooth projective hypersurface $V$, using Li-Wen's $L^2$-Hodge theory [9] of a Landau-Ginzburg model with compact critical locus $V$. We also give a precise comparison result between the current construction and Barannikov-Kontsevich's cons
Yanxiang Zhang, Zheng Xu, Shanshan Wu, Yuanbo Zhang
Error correction is an important capability when applying large language models (LLMs) to facilitate user typing on mobile devices. In this paper, we use LLMs to synthesize a high-quality dataset of error correction pairs to evaluate and improve LLMs for mobile applications. We first prompt LLMs with error correction domain knowledge to build a scalable and
Junlin Wang, Zhiyun Lin
Learning effective visual representations for robotic manipulation remains a fundamental challenge due to the complex body dynamics involved in action execution. In this paper, we study how visual representations that carry body-relevant cues can enable efficient policy learning for downstream robotic manipulation tasks. We present $\textbf{I}$nter-token $\t
Hong Jiao, Dan Song, Won-Chan Lee
Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and induces least rater effects needs to be collected before the use of LLMs for automated scoring in practice. This study compared ten LLMs (Ch
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
cs.CLTimothy Do, Pranav Saran, Harshita Poojary, Pranav Prabhu
In this paper, we address the persistent challenges that figurative language expressions pose for natural language processing (NLP) systems, particularly in low-resource languages such as Konkani. We present a hybrid model that integrates a pre-trained Multilingual BERT (mBERT) with a bidirectional LSTM and a linear classifier. This architecture is fine-tune
Shengzhe Xu, Nikhil Muralidhar, Naren Ramakrishnan
Numerous recent prompt optimization approaches like chain-of-thought, have been demonstrated to significantly improve the quality of content generated by large language models (LLMs). In-context learning (ICL), a recent paradigm where a few representative examples guide content generation has also led to strong improvements in generation quality of LLM gener
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
cs.SDJule Valendo Halim, Siyi Wang, Hong Jia, Ting Dang
Emotional intelligence in conversational AI is crucial across domains like human-computer interaction. While numerous models have been developed, they often overlook the complexity and ambiguity inherent in human emotions. In the era of large speech foundation models (SFMs), understanding their capability in recognizing ambiguous emotions is essential for th
Ian McCulloh, Pedro Rodriguez, Srivaths Kumar, Manu Gupta
The increasing demand for digital literacy and artificial intelligence (AI) fluency in the workforce has highlighted the need for scalable, efficient programming instruction. This study evaluates the effectiveness of integrating generative AI, specifically OpenAIs ChatGPT, into a self-paced Python programming module embedded within a sixteen-week professiona
Retrieval Augmented Decision-Making: A Requirements-Driven, Multi-Criteria Framework for Structured Decision Support
cs.AIHongjia Wu, Hongxin Zhang, Wei Chen, Jiazhi Xia
Various industries have produced a large number of documents such as industrial plans, technical guidelines, and regulations that are structurally complex and content-wise fragmented. This poses significant challenges for experts and decision-makers in terms of retrieval and understanding. Although existing LLM-based Retrieval-Augmented Generation methods ca
Mesoscale Turbulence in Type Ia Supernova Deflagrations: Buoyancy-Driven Fuel Heating and Prospects for Delayed-Detonations
astro-ph.SREzra Brooker, Andrey Zhiglo, Tomasz Plewa
The aim of this work is to characterize the thermodynamic state of fuel mixed into the turbulent flame brush in the context of the Zel'dovich deflagration-to-detonation transition (ZDDT) mechanism of Type Ia supernovae (SNe Ia). We perform a series of three-dimensional computer simulations of thermonuclear deflagrations subject to the Rayleigh-Taylor instabi
James MacLaurin, Pedro Vilanova
The theory of Balanced Neural Networks is a very popular explanation for the high degree of variability and stochasticity in the brain's activity. Roughly speaking, it entails that typical neurons receive many excitatory and inhibitory inputs. The network-wide mean inputs cancel, and one is left with the stochastic fluctuations about the mean. In this paper
Lise-Marie Imbert-Gerard
Trefftz-type of Galerkin methods for numerical PDEs use discrete spaces of problem-dependent functions. While Trefftz methods leverage discrete spaces of local exact solutions to the governing PDE, Taylor-based quasi-Trefftz methods leverage discrete spaces of local approximate solutions to the governing PDE. This notion of approximate solution, understood i
Li-Syun Hsiung, Jun-Kai Tu, Kuan-Wu Chu, Yu-Hsuan Chiu
This study aims to investigate the challenge of insufficient three-dimensional context in synthetic datasets for scene text rendering. Although recent advances in diffusion models and related techniques have improved certain aspects of scene text generation, most existing approaches continue to rely on 2D data, sourcing authentic training examples from movie
Lucas Tecot, Di Luo, Cho-Jui Hsieh
Advancements in quantum computing have spurred significant interest in harnessing its potential for speedups over classical systems. However, noise remains a major obstacle to achieving reliable quantum algorithms. In this work, we present a provably noise-resilient training theory and algorithm to enhance the robustness of parameterized quantum circuit clas
ZooplanktonBench: A Geo-Aware Zooplankton Recognition and Classification Dataset from Marine Observations
cs.CVFukun Liu, Adam T. Greer, Gengchen Mai, Jin Sun
Plankton are small drifting organisms found throughout the world's oceans and can be indicators of ocean health. One component of this plankton community is the zooplankton, which includes gelatinous animals and crustaceans (e.g. shrimp), as well as the early life stages (i.e., eggs and larvae) of many commercially important fishes. Being able to monitor zoo
Colin P. Folsom, Christiana Erba, Veronique Petit, Shaquann Seadrow
Spectropolarimetry, the observation of polarization and intensity as a function of wavelength, is a powerful tool in stellar astrophysics. It is particularly useful for characterizing stars and circumstellar material, and for tracing the influence of magnetic fields on a host star and its environment. Maintaining modern, flexible, and accessible computationa
Mengran Li, Pengyu Zhang, Wenbin Xing, Yijia Zheng
Graphs are a widely used paradigm for representing non-Euclidean data, with applications ranging from social network analysis to biomolecular prediction. While graph learning has achieved remarkable progress, real-world graph data presents a number of challenges that significantly hinder the learning process. In this survey, we focus on four fundamental data
Zhiyuan Zhang, Zhengtong Xu, Jai Nanda Lakamsani, Yu She
Visual Imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds, which provide geometry-aware, appearance-invariant representations, and by incorporating equivariance into policy archite
Sebastian Gutierrez Hernandez, Peng Chen, Haomin Zhou
We introduce Parametric Density Path Optimization (PDPO), a novel method for computing action-minimizing paths between probability densities. The core idea is to represent the target probability path as the pushforward of a reference density through a parametric map, transforming the original infinite-dimensional optimization over densities to a finite-dimen
Quan Khanh Luu, Pokuang Zhou, Zhengtong Xu, Zhiyuan Zhang
Supervised visuomotor policies have shown strong performance in robotic manipulation but often struggle in tasks with limited visual inputs, such as operations in confined spaces and dimly lit environments, or tasks requiring precise perception of object properties and environmental interactions. In such cases, tactile feedback becomes essential for manipula
Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services
cs.CRGuoheng Sun, Ziyao Wang, Xuandong Zhao, Bowei Tian
Modern large language model (LLM) services increasingly rely on complex, often abstract operations, such as multi-step reasoning and multi-agent collaboration, to generate high-quality outputs. While users are billed based on token consumption and API usage, these internal steps are typically not visible. We refer to such systems as Commercial Opaque LLM Ser
Christopher J. Mungall, Adnan Malik, Daniel R. Korn, Justin T. Reese
Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring
Jingkai Wang, Wu Miao, Jue Gong, Zheng Chen
Face restoration has achieved significant advancements through the years of development. However, maintaining high fidelity and authenticity while avoiding artifacts remains challenging, especially in extreme degradation scenarios. This highlights the need for models that are more ``honest'' in their reconstruction from low-quality inputs, accurately
Generalized many-body exciton g-factors: magnetic hybridization and non-monotonic Rydberg series in monolayer WSe$_2$
cond-mat.mes-hallPaulo E. Faria Junior, Daniel Hernangómez-Pérez, Tomer Amit, Jaroslav Fabian
Magneto-optics of low dimensional semiconductors, such as monolayer transition metal dichalcogenides, offers a vast playground for exploring complex quantum phenomena. However, current ab initio approaches fail to capture important experimental observations related to brightening of excitonic levels and their g-factor dependence. Here, we develop a robust an
Unggi Lee, Jaeyong Lee, Jiyeong Bae, Yeil Jeong
Recent advances in large reasoning models (LRMs) show strong performance in structured domains such as mathematics and programming; however, they often lack pedagogical coherence and realistic teaching behaviors. To bridge this gap, we introduce Pedagogy-R1, a framework that adapts LRMs for classroom use through three innovations: (1) a distillation-based pi
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
cs.CLMamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, Ziwei Zhu
Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. This work addresses these gaps by conducting a multilingual and intersectional analysis of LLM outputs across 10 Indo-Aryan and Dravidian languages, identifying how cultural stigmas i
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
cs.HCUgur Kursuncu, Trilok Padhi, Gaurav Sinha, Abdulkadir Erol
The growing demand for accessible mental health support, compounded by workforce shortages and logistical barriers, has led to increased interest in utilizing Large Language Models (LLMs) for scalable and real-time assistance. However, their use in sensitive domains such as anxiety support remains underexamined. This study presents a systematic evaluation of
Natarajan Balaji Shankar, Zilai Wang, Kaiyuan Zhang, Mohan Shi
Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, leading to high transcription error rates. While ASR error correction (AEC) methods have improved adult speech transcription, their effectiveness on child speech remains largely unexp
Xiaofei Huang, Kai Wei, Yang Rui, Dinghui Gong
Atomic spin sensors are essential for beyond-the-standard-model exploration, biomagnetic measurement, and quantum navigation. While the traditional DC mode spin-exchange relaxation-free (SERF) comagnetometer achieves ultrahigh sensitivity, further improvements require suppressing technical noise and surpassing standard quantum limit. In this work, we develop
Performance and Generalizability Impacts of Incorporating Location Encoders into Deep Learning for Dynamic PM2.5 Estimation
cs.LGMorteza Karimzadeh, Zhongying Wang, James L. Crooks
Deep learning has shown strong performance in geospatial prediction tasks, but the role of geolocation information in improving accuracy and generalizability remains underexamined. Recent work has introduced location encoders that aim to represent spatial context in a transferable way, yet most evaluations have focused on static mapping tasks. Here, we study
Viacheslav Tsaran, Francesco Marino, Sonia Bacca, Francesca Bonaiti
We extend the pion-nucleus multiple-scattering framework to include detailed second-order rescattering dynamics for nuclei with non-zero isospin. To account for intermediate charge-exchange and nucleon spin-flip effects, we develop a scattering potential that depends on the one- and two-body densities of the target nucleus. We compute one-body densities from
Xuanhe Zhou, Junxuan He, Wei Zhou, Haodong Chen
The integration of large language model (LLM) and data management (DATA) is rapidly redefining both domains. In this survey, we comprehensively review the bidirectional relationships. On the one hand, DATA4LLM, spanning large-scale data processing, storage, and serving, feeds LLMs with high quality, diversity, and timeliness of data required for stages like
Abir Ray
This paper introduces EdgeAgentX, a novel framework integrating federated learning (FL), multi-agent reinforcement learning (MARL), and adversarial defense mechanisms, tailored for military communication networks. EdgeAgentX significantly improves autonomous decision-making, reduces latency, enhances throughput, and robustly withstands adversarial disruption
Litu Rout, Constantine Caramanis, Sanjay Shakkottai
Diffusion Language Models (DLMs) promise parallel generation and bidirectional context, yet they underperform autoregressive (AR) models in both likelihood modeling and generated text quality. We identify that this performance gap arises when important tokens (e.g., key words or low-frequency words that anchor a sentence) are masked early in the forward proc
Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian
The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its popularity and relevance, the theoretical properties of the softmax-
Zhenrui Yue, Bowen Jin, Huimin Zeng, Honglei Zhuang
Recent advances in large language models (LLMs) have introduced latent reasoning as a promising alternative to autoregressive reasoning. By performing internal computation with hidden states from previous steps, latent reasoning benefit from more informative features rather than sampling a discrete chain-of-thought (CoT) path. Yet latent reasoning approaches
Zhichao Wu, Yueteng Kang, Songjun Cao, Long Ma
Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on multi-modal prompt. The system disentangles speech into the content, timbre, emotion and prosody, allowing emotion promp
MedScore: Generalizable Factuality Evaluation of Free-Form Medical Answers by Domain-adapted Claim Decomposition and Verification
cs.CLHeyuan Huang, Alexandra DeLucia, Vijay Murari Tiyyala, Mark Dredze
While Large Language Models (LLMs) can generate fluent and convincing responses, they are not necessarily correct. This is especially apparent in the popular decompose-then-verify factuality evaluation pipeline, where LLMs evaluate generations by decomposing the generations into individual, valid claims. Factuality evaluation is especially important for medi
Toshiaki Koike-Akino, Jing Liu, Ye Wang
To tackle the huge computational demand of large foundation models, activation-aware compression techniques without retraining have been introduced. However, since these rely on calibration data, domain shift may arise for unknown downstream tasks. With a computationally efficient calibration, activation-aware pruning can be executed for every prompt adaptiv
Ainulla Khan, Yamada Moyuru, Srinidhi Akella
Retrieval-Augmented Generation (RAG) has emerged as a promising technique to enhance the quality and relevance of responses generated by large language models. While recent advancements have mainly focused on improving RAG for text-based queries, RAG on multi-modal documents containing both texts and images has not been fully explored. Especially when fine-t
Sarasija Sudharsan, Anupam Sharma
This paper presents a numerical demonstration of the real-time application of two dynamic stall onset criteria for identifying and mitigating stall. These criteria - based on the leading-edge suction parameter (LESP) and boundary enstrophy flux (BEF) - are derived from prior research. The present work establishes a proof of concept for the practical use of t
Yucheng Guo, Qinxin Yan
We introduce a family of particle systems on sparse graphs where local interactions occur via hitting times, providing a dynamic and tractable model for default cascades in large sparsely-connected financial networks. Building on the framework of Lacker, Ramanan and Wu (2023), we extend convergence theory to systems with singular interactions, capturing the
Pessimism Principle Can Be Effective: Towards a Framework for Zero-Shot Transfer Reinforcement Learning
cs.LGChi Zhang, Ziying Jia, George K. Atia, Sihong He
Transfer reinforcement learning aims to derive a near-optimal policy for a target environment with limited data by leveraging abundant data from related source domains. However, it faces two key challenges: the lack of performance guarantees for the transferred policy, which can lead to undesired actions, and the risk of negative transfer when multiple sourc
Hojun Son, Asma Almutairi, Arpan Kusari
Context bias refers to the association between the foreground objects and background during the object detection training process. Various methods have been proposed to minimize the context bias when applying the trained model to an unseen domain, known as domain adaptation for object detection (DAOD). But a principled approach to understand why the context
A Cosmic Ray Acceleration Mechanism Based on Background Flow Velocity Inhomogeneities Yielding Power-Law Spectra
astro-ph.HEJ. -F. Wang, G. Qin
In this article, momentum transport generated by the combined effects of pitch-angle diffusion and Background Flow Velocity Inhomogeneities (BFVIs) is proposed to obtain a cosmic rays acceleration mechanism, starting from the well-known focusing equation describing particle diffusion and acceleration. The inhomogeneities of background flow velocity is ubiqui
Yiren Song, Cheng Liu, Mike Zheng Shou
Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, composition, and fine details, and (2) preventing style degradation in image-to-image pipelines with style LoRAs. GPT-4o's exceptional stylization consistency highlights the performa
Christian D. Newman, Anthony Peruma, Eman Abdullah AlOmar, Mahie Crabbe
Identifier names are crucial components of code, serving as primary clues for developers to understand program behavior. This paper investigates the linguistic structure of identifier names by extending the concept of grammar patterns, which represent the part-of-speech (PoS) sequences underlying identifier phrases. The specific focus is on closed syntactic
Serkan Hoşten
Toric ideals are everywhere. They have been in the commutative algebra lexicon since about 1990 when Bernd Sturmfels used the term. The early days of toric ideals and their Gr\"obner bases were full of new results and promising developments in their applications. Bernd has been consistently their biggest promoter through his own work and that of his collabor
Zhining Liu, Ze Yang, Xiao Lin, Ruizhong Qiu
Time-series forecasting plays a critical role in many real-world applications. Although increasingly powerful models have been developed and achieved superior results on benchmark datasets, through a fine-grained sample-level inspection, we find that (i) no single model consistently outperforms others across different test samples, but instead (ii) each mode
Romeo Valentin, Sydney M. Katz, Vincent Vanhoucke, Mykel J. Kochenderfer
Dictionary learning has recently emerged as a promising approach for mechanistic interpretability of large transformer models. Disentangling high-dimensional transformer embeddings requires algorithms that scale to high-dimensional data with large sample sizes. Recent work has explored sparse autoencoders (SAEs) for this problem. However, SAEs use a simple l
Zhaoyang Wang, Jinqi Jiang, Tian Qiu, Hui Liu
Recent large reasoning models such as DeepSeek-R1 exhibit strong complex problems solving abilities by generating long chain-of-thought (CoT) reasoning steps. It is challenging to directly train small language models (SLMs) to emerge long CoT. Thus, distillation becomes a practical method to enable SLMs for such reasoning ability. However, the long CoT often