Skip to content

May 2025 arXiv papers — page 14

Showing 1,3011,400 of 24,552 papers

  1. Qing Li, Jiahui Geng, Zongxiong Chen, Derui Zhu

    In recent years, large language models (LLMs) have made remarkable advancements, yet hallucination, where models produce inaccurate or non-factual statements, remains a significant challenge for real-world deployment. Although current classification-based methods, such as SAPLMA, are highly efficient in mitigating hallucinations, they struggle when non-factu

  2. Ishaan Rawal, Suryansh Kumar

    Text-conditioned diffusion models have emerged as powerful tools for high-quality video generation. However, enabling Interactive Video Generation (IVG), where users control motion elements such as object trajectory, remains challenging. Recent training-free approaches introduce attention masking to guide trajectory, but this often degrades perceptual qualit

  3. Yizhong Ding

    Frequent cyber-attacks have elevated WebShell exploitation and defense to a critical research focus within network security. However, there remains a significant shortage of publicly available, well-categorized malicious-code datasets organized by obfuscation method. Existing malicious-code generation methods, which primarily rely on prompt engineering, ofte

  4. Xiaoyu Li, Xiao Li, Li Gao, Yiding Liu

    The evolution of Large Language Models (LLMs) has significantly advanced multi-turn conversation systems, emphasizing the need for proactive guidance to enhance users' interactions. However, these systems face challenges in dynamically adapting to shifts in users' goals and maintaining low latency for real-time interactions. In the Baidu Search AI assistant,

  5. Ayush Jha, Abootaleb Shirvani, Ali Jaffri, Svetlozar T. Rachev

    This paper introduces a state-dependent momentum framework that integrates ESG regime switching with tail-risk-aware reward-risk metrics. Using a dynamic programming approach and solving a finite-horizon Bellman equation, we construct long-short momentum portfolios that adjust to changing ESG sentiment regimes. Unlike traditional momentum strategies based on

  6. Qingyao Tian, Huai Liao, Xinyan Huang, Bingyu Yang

    Vision-based 6-DOF bronchoscopy localization offers a promising solution for accurate and cost-effective interventional guidance. However, existing methods struggle with 1) limited generalization across patient cases due to scarce labeled data, and 2) poor robustness under visual degradation, as bronchoscopy procedures frequently involve artifacts such as oc

  7. Wei-Cheng Tseng, David Harwath

    Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving paradigm shifts across various speech processing tasks. Despite these advancements, their robustness in noisy environments rema

  8. Minchul Kim, Anil Jain, Xiaoming Liu

    Over the past five decades, automated face recognition (FR) has progressed from handcrafted geometric and statistical approaches to advanced deep learning architectures that now approach, and in many cases exceed, human performance. This paper traces the historical and technological evolution of FR, encompassing early algorithmic paradigms through to contemp

  9. Alice Qian, Ryland Shaw, Laura Dabbish, Jina Suh

    As AI systems are increasingly tested and deployed in open-ended and high-stakes domains, crowdworkers are often tasked with responsible AI (RAI) content work. These tasks include labeling violent content, moderating disturbing text, or simulating harmful behavior for red teaming exercises to shape AI system behaviors. While prior research efforts have highl

  10. Xin Kang, Zihan Zheng, Lei Chu, Yue Gao

    We present LTM3D, a Latent Token space Modeling framework for conditional 3D shape generation that integrates the strengths of diffusion and auto-regressive (AR) models. While diffusion-based methods effectively model continuous latent spaces and AR models excel at capturing inter-token dependencies, combining these paradigms for 3D shape generation remains

  11. Nir Endy, Idan Daniel Grosbard, Yuval Ran-Milo, Yonatan Slutzky

    This paper investigates the flow of factual information in Mamba State-Space Model (SSM)-based language models. We rely on theoretical and empirical connections to Transformer-based architectures and their attention mechanisms. Exploiting this relationship, we adapt attentional interpretability techniques originally developed for Transformers--specifically,

  12. Joohwan Ko, Justin Domke

    Variational inference often struggles with the posterior geometry exhibited by complex hierarchical Bayesian models. Recent advances in flow-based variational families and Variationally Inferred Parameters (VIP) each address aspects of this challenge, but their formal relationship is unexplored. Here, we prove that the combination of VIP and a full-rank Gaus

  13. Zhicheng He, Tinggui Wang, Gary J. Ferland

    Photoionized gases are prevalent throughout the universe. In such gases, the ion concentration typically exhibits two response modes to radiation: a positive response in the low-ionization state and a negative response in the high-ionization state. Here, we report the discovery of a widespread misalignment at the boundary between the above two response modes

  14. Naibin Gu, Yilong Chen, Zhenyu Zhang, Peng Fu

    Although scaling up the number of trainable parameters can effectively improve the training performance of large language models, it also leads to increased computational overhead. When delving into the parameter difference, we find that a subset of parameters, termed advantageous parameters, plays a crucial role in determining model performance. Further ana

  15. Tarak Chand, Saurabh Sharma, Koshvendra Singh, Jeewan Pandey

    We present a decade-long investigation of a poorly studied cluster, Berkeley 65 (Be 65), using deep optical data from the telescopes of ARIES, Nainital Observatory. We estimate its radius ($R_{cluster}$ = 1.6$^{'}$, aspect ratio of $\sim$1.1), distance (2.0 $\pm$ 0.1 kpc) and age ($\sim$160 Myrs). A clear turn-off point at $\sim$1.7 M$_\odot$ in the mass fun

  16. Sana Ebrahimi, Mohsen Dehghankar, Abolfazl Asudeh

    While multi-agent LLM systems show strong capabilities in various domains, they are highly vulnerable to adversarial and low-performing agents. To resolve this issue, in this paper, we introduce a general and adversary-resistant multi-agent LLM framework based on credibility scoring. We model the collaborative query-answering process as an iterative game, wh

  17. Peng Xie, Xingyuan Liu, Tsz Wai Chan, Yequan Bie

    Code-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition (ASR) systems, which are typically designed for a single language and struggle to handle multilingual inputs. The growing

  18. Bowen Dong, Minheng Ni, Zitong Huang, Guanglei Yang

    Multimodal hallucination in multimodal large language models (MLLMs) restricts the correctness of MLLMs. However, multimodal hallucinations are multi-sourced and arise from diverse causes. Existing benchmarks fail to adequately distinguish between perception-induced hallucinations and reasoning-induced hallucinations. This failure constitutes a significant i

  19. Xia Liao, Xiping Zhang

    Let $M$ be a complex manifold, $D\subset M$ a free divisor and $U=M\setminus D$ its complement. In this paper we study the characteristic cycle $\textup{CC}(\gamma\cdot \ind_U)$ of the restriction of a constructible function $\gamma$ on $U$. We globalise Ginzburg's local sharp construction and introduce the log transversality condition, which is a new transv

  20. Muralidharan K., Agniva Das, Shrey Pandya, Jong Min Kim

    The impact of statistical methodologies on studying groundwater has been significant in the last several decades, due to cheaper computational abilities and presence of technologies that enable us to extract and measure more and more data. This paper focuses on the validation of statistical methodologies that are in practice and continue to be at the earlies

  21. Yao Yu, Wen-Zhang Feng, Hong-Song Xie, Han Zhang

    Power-law inflation has stood as a classical model in inflationary cosmology since the early 1980s, prized for its exact analytical solutions and ability to naturally resolve the Big Bang theory's horizon and flatness problems through exponential expansion. However, its simplest form appears incompatible with modern precision observations, motivating increas

  22. Sumit Kumar Rai, Bhabani Prasad Mandal, Ronaldo Thibes

    We obtain the various forms of BRST symmetry by using the Batalin-Fradkin-Vilkovisky formalism in a prototypical first class system. We have shown that the various forms of symmetry can be obtained through canonical transformation in the ghost sector. The so called "dual-BRST" symmetry which is claimed to be an independent symmetry due to its roots in differ

  23. Haibo Jin, Peiyan Zhang, Peiran Wang, Man Luo

    Large foundation models (LFMs) are susceptible to two distinct vulnerabilities: hallucinations and jailbreak attacks. While typically studied in isolation, we observe that defenses targeting one often affect the other, hinting at a deeper connection. We propose a unified theoretical framework that models jailbreaks as token-level optimization and hallucinati

  24. Md Shahnawaz, Bishwajit Prasad Gond, Durga Prasad Mohapatra

    Malware detection and classification remains a topic of concern for cybersecurity, since it is becoming common for attackers to use advanced obfuscation on their malware to stay undetected. Conventional static analysis is not effective against polymorphic and metamorphic malware as these change their appearance without modifying their behavior, thus defying

  25. Murari Ambati

    We propose ProofNet++, a neuro-symbolic framework that enhances automated theorem proving by combining large language models (LLMs) with formal proof verification and self-correction mechanisms. Current LLM-based systems suffer from hallucinated logical steps and unverifiable reasoning. ProofNet++ mitigates these limitations by integrating symbolic proof tre

  26. Luong Ho, Khanh Le, Vinh Pham, Bao Nguyen

    Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integration of streaming ITN within streaming ASR remains largely unexplored due to challenges in accuracy, efficiency, and adaptability, particula

  27. A. Crisanti, A. Sarracino, M. Zannetti

    The realisation of Bose-Einstein condensation under grand-canonical conditions has provided the experimental evidence for the simultaneous occurrence of macroscopic fluctuations and phase coherence of the condensate. The observation of these two features, against a consolidated tradition which wants the fluctuations to be pathological (grand-canonical catast

  28. Ying Yang, Jie Zhang, Xiao Lv, Di Lin

    While adversarial attacks on vision-and-language pretraining (VLP) models have been explored, generating natural adversarial samples crafted through realistic and semantically meaningful perturbations remains an open challenge. Existing methods, primarily designed for classification tasks, struggle when adapted to VLP models due to their restricted optimizat

  29. Yibo Zhao, Jiapeng Zhu, Ye Guo, Kangkang He

    Graph-based RAG methods like GraphRAG have shown promising global understanding of the knowledge base by constructing hierarchical entity graphs. However, they often suffer from inefficiency and rely on manually pre-defined query modes, limiting practical use. In this paper, we propose E^2GraphRAG, a streamlined graph-based RAG framework that improves both E

  30. Haibo Jin, Peiyan Zhang, Man Luo, Haohan Wang

    Large Language Models (LLMs) have shown remarkable progress across domains, yet their ability to perform inductive reasoning - inferring latent rules from sparse examples - remains limited. It is often assumed that chain-of-thought (CoT) prompting, as used in Large Reasoning Models (LRMs), enhances such reasoning. We investigate this assumption with creating

  31. Chengxi Deng, Xurong Xie, Shujie Hu, Mengzhe Geng

    This paper proposes a novel Mixture of Prompt-Experts based Speaker Adaptation approach (MOPSA) for elderly speech recognition. It allows zero-shot, real-time adaptation to unseen speakers, and leverages domain knowledge tailored to elderly speakers. Top-K most distinctive speaker prompt clusters derived using K-means serve as experts. A router network is tr

  32. Jean-Benoit Delbrouck, Justin Xu, Johannes Moll, Alois Thomas

    Automated radiology report generation from chest X-ray (CXR) images has the potential to improve clinical efficiency and reduce radiologists' workload. However, most datasets, including the publicly available MIMIC-CXR and CheXpert Plus, consist entirely of free-form reports, which are inherently variable and unstructured. This variability poses challenges f

  33. Fangyikang Wang, Hubery Yin, Lei Qian, Yinan Li

    The diffusion models (DMs) have demonstrated the remarkable capability of generating images via learning the noised score function of data distribution. Current DM sampling techniques typically rely on first-order Langevin dynamics at each noise level, with efforts concentrated on refining inter-level denoising strategies. While leveraging additional second-

  34. Zhen Liu, Wenzhe Zhu, Yongkun Li, Yinlong Xu

    Persistent key-value (KV) stores are critical infrastructure for data-intensive applications. Leveraging high-performance Non-Volatile Memory (NVM) to enhance KV stores has gained traction. However, previous work has primarily focused on optimizing KV stores themselves, without adequately addressing their integration into applications. Consequently, existing

  35. Alain Albouy

    We present an elementary deduction of the Newtonian force from Kepler's laws. We relate it to a generalization by Jacobi of the Keplerian motion, where the Euclidean form in the plane is replaced by some function with the same homogeneity. We show how several convexity properties of the generalized Keplerian orbits appear in this context. We describe the gen

  36. Binbin Yan, Di Huang, Jianglai Liu, Xiaoying Lu

    The PandaX-4T liquid xenon detector uses Hamamatsu 3-inch R11410-23 photomultiplier tubes (PMTs) as the light sensors for ultra-low radioactivity, high quantum efficiency, and long-term stability at cryogenic temperature. Each PMT was thoroughly tested in a dedicated chamber before being installed in the detector to ensure compliance with experimental requir

  37. Lam Thanh Do, Aaditya Bodke, Pritom Saha Akash, Kevin Chen-Chuan Chang

    Unsupervised keyphrase prediction has gained growing interest in recent years. However, existing methods typically rely on heuristically defined importance scores, which may lead to inaccurate informativeness estimation. In addition, they lack consideration for time efficiency. To solve these problems, we propose ERU-KG, an unsupervised keyphrase generation

  38. Jeehoon Park, Jaewon Yoo

    Given a Calabi-Yau smooth projective complete intersection variety $V$ over $\mathbb{C}$, a hybrid Landau-Ginzburg (LG) model may be associated using the Cayley trick. This hybrid LG model comprises a non-compact Calabi-Yau manifold $X_{CY}$, and a holomorphic function $W$, defined on $X_{CY}$, such that the critical locus of $W$ is isomorphic to $V$. We con

  39. Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang, Chenyan Xiong

    Although Large Language Models (LLMs) have become capable reasoners, the problem of faithfulness persists: their reasoning can contain errors and omissions that are difficult to detect and that may obscure biases in model outputs. To address this issue, we introduce Semi-Structured Reasoning Models (SSRMs), which are trained to produce semi-structured repres

  40. Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard, Sohail Dianat

    This work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (SPM), and a novel reweighting strategy are introduced to enhance performance. SPM shuffles and blends image patches to generate diverse and challenging augmentations, while the rew

  41. Yusuke Nemoto

    In this paper, we give a transformation formula of Dwork's $p$-adic hypergeometric function between $t$ and $t^{-1}$. As an appendix, we introduce a finite analogue of this transformation formula, which implies the special case of the above transformation formula.

  42. Vincent Siu, Nicholas Crispino, Zihao Yu, Sam Pan

    Large Language Models (LLMs) encode behaviors such as refusal within their activation space, yet identifying these behaviors remains a significant challenge. Existing methods often rely on predefined refusal templates detectable in output tokens or require manual analysis. We introduce \textbf{COSMIC} (Cosine Similarity Metrics for Inversion of Concepts), an

  43. Redwan Sony, Parisa Farmanifard, Hamzeh Alzwairy, Nitish Shukla

    The advent of foundation models, particularly Vision-Language Models (VLMs) and Multi-modal Large Language Models (MLLMs), has redefined the frontiers of artificial intelligence, enabling remarkable generalization across diverse tasks with minimal or no supervision. Yet, their potential in biometric recognition and analysis remains relatively underexplored.

  44. Geng Li, Chunjiang Shi, Ying Chen, Wei Sun

    The $^1S_0$ and $^5S_2$ amplitudes of $J/ψ~J/ψ$ scattering are determined up to 6.6~GeV from $N_f=2$ lattice QCD simulations at $m_π\approx 420$~MeV and $250$~MeV. We observe an attractive interaction in the ${}^1S_0$ channel, which permits the existence of a near-threshold scalar structure. This structure may correspond to $X(6200)$, but the possible left-h

  45. Paolo Braccia, N. L. Diaz, Martin Larocca, M. Cerezo

    Sampling unitary Fermionic Linear Optics (FLO), or matchgate circuits, has become a fundamental tool in quantum information. Such capability enables a large number of applications ranging from randomized benchmarking of continuous gate sets, to fermionic classical shadows. In this work, we introduce optimal algorithms to sample over the non-particle-preservi

  46. Jiwan Chung, Janghan Yoon, Junhyeong Park, Sangeyl Lee

    Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models truly achieve cross-modal coherence, or is this coherence merely perceived? To explore this, we introduce ACON, a dataset of

  47. Zheng Tan, Weizhen Wang, Andrea L. Bertozzi, Ernest K. Ryu

    Diffusion models (DMs) and flow-matching models have demonstrated remarkable performance in image and video generation. However, such models require a significant number of function evaluations (NFEs) during sampling, leading to costly inference. Consequently, quality-preserving fast sampling methods that require fewer NFEs have been an active area of resear

  48. Sanghyeon Nam, Dongmin Kim, Seung-Hwan Choi, Chang-Hyun Kim

    Robotic manipulators are essential for precise industrial pick-and-place operations, yet planning collision-free trajectories in dynamic environments remains challenging due to uncertainties such as sensor noise and time-varying delays. Conventional control methods often fail under these conditions, motivating the development of Robust MPC (RMPC) strategies

  49. Wenhan Yang, Spencer Stice, Ali Payani, Baharan Mirzasoleiman

    Ensuring Vision-Language Models (VLMs) generate safe outputs is crucial for their reliable deployment. However, LVLMs suffer from drastic safety degradation compared to their LLM backbone. Even blank or irrelevant images can trigger LVLMs to generate harmful responses to prompts that would otherwise be refused in text-only contexts. The modality gap between

  50. Gang Wu, Junjun Jiang, Kui Jiang, Xianming Liu

    Unified image restoration models for diverse and mixed degradations often suffer from unstable optimization dynamics and inter-task conflicts. This paper introduces Self-Improved Privilege Learning (SIPL), a novel paradigm that overcomes these limitations by innovatively extending the utility of privileged information (PI) beyond training into the inference

  51. Takayuki Kobayashi, Ryosuke Nakasato

    We consider the initial value problem of the compressible Navier-Stokes-Korteweg equations in the whole space $\mathbb{R}^d$ ($d \ge 2$). The purposes of this paper are to obtain the global-in-time solution around the constant equilibrium states $(\rho_*,0)$ and investigate the $L^p$-$L^1$ type time-decay estimates in a scaling critical framework, where $\rh

  52. Mingze Wang, Weinan E

    Mixture-of-experts networks (MoEs) have demonstrated remarkable efficiency in modern deep learning. Despite their empirical success, the theoretical foundations underlying their ability to model complex tasks remain poorly understood. In this work, we conduct a systematic study of the expressive power of MoEs in modeling complex tasks with two common structu

  53. Tamojit Chakraborty, Anamitra Pal, Sam Maleki

    Low-frequency oscillations (LFOs) present a significant challenge to the stability and reliability of power systems, especially in grids with a high penetration of renewable energy sources. Traditional grid-following (GFL) inverters have proven less effective in damping such oscillations. This paper presents a GFL-power plant controller with an auxiliary pow

  54. Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, Aurélie Lozano

    Protein dynamics play a crucial role in protein biological functions and properties, and their traditional study typically relies on time-consuming molecular dynamics (MD) simulations conducted in silico. Recent advances in generative modeling, particularly denoising diffusion models, have enabled efficient accurate protein structure prediction and conformat

  55. Pooja Devi, Nat Gopalswamy, Seiji Yashiro, Sachiko Akiyama

    In this article, we present the relationship between prominence eruptions (PEs) and coronal mass ejections (CMEs) from May 2010 to December 2019 covering most of solar cycle 24. We used data from the Atmospheric Imaging Assembly (AIA) for PEs and the Large Angle and Spectrometric Coronagraph (LASCO) for CMEs. We identified 1225 PEs, with 67% being radial, 32

  56. Xu He, Di Wu, Yan Zhai, Kun Sun

    The rise of large language model (LLM)-based multi-agent systems (MAS) introduces new security and reliability challenges. While these systems show great promise in decomposing and coordinating complex tasks, they also face multi-faceted risks across prompt manipulation, unsafe tool usage, and emergent agent miscoordination. Existing guardrail mechanisms off

  57. Qingzheng Wang, Jiancheng Sun, Yifan Peng, Shinji Watanabe

    Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (ASR). However, these models struggle with limited resources during fine-tuning. This paper enhances multilingual LID and ASR on ML-SUPERB 2.0

  58. Yimin Du

    The quality of human preference data is crucial for training and evaluating large language models (LLMs), particularly in reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO) scenarios. Traditional side-by-side (SBS) annotation approaches often struggle with inherent uncertainty, annotator disagreement, and the complexit

  59. Yitang Li, Yuanhang Zhang, Wenli Xiao, Chaoyi Pan

    Can your humanoid walk up and hand you a full cup of beer, without spilling a drop? While humanoids are increasingly featured in flashy demos like dancing, delivering packages, traversing rough terrain, fine-grained control during locomotion remains a significant challenge. In particular, stabilizing a filled end-effector (EE) while walking is far from solve

  60. Bhrij Patel, Ashish Jagmohan, Aditya Vempaty

    Digital tool-based agents, powered by Large Language Models (LLMs), that invoke external Application Programming Interfaces (APIs) often rely on documentation to understand API functionality. However, such documentation is frequently missing, outdated, privatized, or inconsistent-hindering the development of reliable, general-purpose agents. In this work, we

  61. Longze Chen, Renke Shan, Huiming Wang, Lu Wang

    Speculative decoding (SD) is a promising method for accelerating the decoding process of Large Language Models (LLMs). The efficiency of SD primarily hinges on the consistency between the draft model and the verify model. However, existing drafting approaches typically require additional modules to be trained, which can be challenging to implement and ensure

  62. Zining Wang, Yuxuan Zhang, Dongwook Yoon, Nicholas Vincent

    With more than 11 times as many pageviews as the next largest edition, English Wikipedia dominates global knowledge access relative to other language editions. Readers are prone to assuming English Wikipedia as a superset of all language editions, leading many to prefer it even when their primary language is not English. Other language editions, however, com

  63. Yi-Qiang Wu, Xuan Liu, Hanlin Li, Fuqiang Wang

    Physics-informed neural networks (PINN) have been widely used in computational physics to solve partial differential equations (PDEs). In this study, we propose an energy-embedding-based physics-informed neural network method for solving the one-dimensional time-independent Schr\"{o}dinger equation to obtain ground- and excited-state wave functions, as well

  64. Ofir Schlisselberg, Tal Lancewicki, Peter Auer, Yishay Mansour

    We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from significant gaps to the known lower bounds, especially in the

  65. Daiki Goto, Kentaro Kuga, Kiyohisa Tanaka, Tsunehiro Takeuchi

    In the field of thermoelectric materials and devices, improving energy conversion efficiency remains a long-standing challenge. As a promising approach to address this issue, utilizing energy-dependent electron-scattering beyond the ordinary constant relaxation time approximation (CRTA) has been proposed. However, direct experimental evidence for an energy-d

  66. Tavis Bennett, Aidan Smith, Edric Matwiejew, Jingbo Wang

    We present benchmarking results for the non-variational Quantum Walk Optimisation Algorithm (non-variational QWOA) applied to the weighted maxcut problem, using classical simulations for problem sizes up to $n = 31$. The amplified quantum state, prepared using a quadratic number of alternating unitaries, achieves a constant average-case measurement probabili

  67. Lan-Cuong Nguyen, Quan Nguyen-Tri, Bang Tran Khanh, Dung D. Le

    Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between real and synthetic distributions. To address this limitation,

  68. Orlando Marquez Ayala, Patrice Bechard, Emily Chen, Maggie Baird

    Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that re

  69. Xinran Yu

    We study conformally compact metrics satisfying the Lovelock equations, which generalize the Einstein equation. We show that these metrics admit polyhomogeneous expansions, thereby naturally realizing the Fefferman-Graham expansion, which is an important tool in conformal geometry and the AdS/CFT correspondence. In even dimensions, we identify a boundary obs

  70. Mikhail L. Arbuzov, Sisong Bei, Ziwei Dong, Dmitri Kalaev

    The prevailing assumption of an exponential decay in large language model (LLM) reliability with sequence length, predicated on independent per-token error probabilities, posits an inherent limitation for long autoregressive outputs. Our research fundamentally challenges this view by synthesizing emerging evidence that LLM errors are not uniformly distribute

  71. Kun Yang, Jianxin Yu, Jia Zhang, Sheng Meng

    Two-dimensional ferroelectrics with robust polarization offer promising opportunities for non-volatile memory, field-effect transistors, and optoelectronic devices. However, the impact of lattice deformation on polarization and photoinduced structural response remains poorly understood. Here, we employ first-principles calculations to demonstrate photodoping

  72. Yipan Wei, Yuchen Zou, Yapeng Li, Bo Du

    Federated Multi-Task Learning (FMTL) enables multiple clients performing heterogeneous tasks without exchanging their local data, offering broad potential for privacy preserving multi-task collaboration. However, most existing methods focus on building personalized models for each client and unable to support the aggregation of multiple heterogeneous tasks i

  73. Alex Tao

    We prove a vanishing theorem of Betti numbers on compact, strictly pseudoconvex pseudohermitian manifolds with non-negative curvature operator. The proof is by an application of the Bochner technique to the setting of CR manifolds.

  74. Yaoyu Zhu, Di Huang, Hanqi Lyu, Xiaoyun Zhang

    Large language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like

  75. Zhuobai Dong, Junchao Yi, Ziyuan Zheng, Haochen Han

    Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GPT-4o demonstrate impressive perceptual and reasoning capabilities, our investigation reveals these models struggle profoundly with visual ph

  76. Yuyang Lai, Sina Heydari, On Shun Pak, Yi Man

    Motile microorganisms develop effective swimming gaits to adapt to complex biological environments. Translating this adaptability to smart microrobots presents significant challenges in motion planning and stroke design. In this work, we explore the use of reinforcement learning (RL) to develop stroke patterns for targeted navigation in a three-link swimmer

  77. Guanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li

    Chain of Thought (CoT) prompting improves the reasoning performance of large language models (LLMs) by encouraging step by step thinking. However, CoT-based methods depend on intermediate reasoning steps, which limits scalability and generalization. Recent work explores recursive reasoning, where LLMs reuse internal layers across iterations to refine latent

  78. Lisa Orloff Clark, Lynnel D. Naingue, Jocelyn P. Vilela

    We generalise recent results about quasi-Cartan, Cartan and diagonal subalgebras by introducing graded versions. We show that there is a correspondence between graded algebraic quasi-Cartan/ Cartan/ diagonal pairs and certain graded twisted Steinberg algebras and that the associated graded discrete twist is unique. Our results include all discrete group alge

  79. Xiaodong Ji, Hailin Zhang, Fangcheng Fu, Bin Cui

    Many advanced Large Language Model (LLM) applications require long-context processing, but the self-attention module becomes a bottleneck during the prefilling stage of inference due to its quadratic time complexity with respect to sequence length. Existing sparse attention methods accelerate attention computation by skipping less significant regions of the

  80. Katherine Tieu, Dongqi Fu, Jun Wu, Jingrui He

    In the era of foundation models, Out-of- Distribution (OOD) problems, i.e., the data discrepancy between the training environments and testing environments, hinder AI generalization. Further, relational data like graphs disobeying the Independent and Identically Distributed (IID) condition makes the problem more challenging, especially much harder when it is

  81. Jindiao Huang, Haifan Yin

    The Holographic Interference Surface (HIS) opens up a new prospect for building a more cost-effective wireless communication architecture by performing Radio Frequency (RF) domain signal processing. In this paper, we establish a wideband channel sensing architecture for electromagnetic wave reception and channel estimation based on the principle of holograph

  82. Zihao Yu, Xiang Li, Jing Zhang

    The rapid dissemination of rumors on social media highlights the urgent need for automatic detection methods to safeguard societal trust and stability. While existing multimodal rumor detection models primarily emphasize capturing consistency between intrinsic modalities (e.g., news text and images), they often overlook the intricate interplay between intrin

  83. Shirui Wei, Changhua Li, Yanxia Zhang, Chenzhou Cui

    Emission Line Galaxies (ELGs) are crucial for cosmological studies, particularly in understanding the large-scale structure of the Universe and the role of dark energy. ELGs form an essential component of the target catalogue for the Dark Energy Spectroscopic Instrument (DESI), a major astronomical survey. However, the accurate selection of ELGs for such sur

  84. Jiawei Hou, Xiangyang Xue, Taiping Zeng

    Autonomous operation of service robotics in human-centric scenes remains challenging due to the need for understanding of changing environments and context-aware decision-making. While existing approaches like topological maps offer efficient spatial priors, they fail to model transient object relationships, whereas dense neural representations (e.g., NeRF)

  85. Ryota Miyano, Yuki Arase

    This study proposes a simple yet effective LoRA merge method to achieve LLM adaptation for low-resource language generation tasks. The LoRA merge technique, which integrates multiple LoRA modules trained on different tasks, has gained attention as an effective and efficient approach for adapting LLMs to target tasks. However, previous methods are limited in

  86. Tianhong Zhou, Yin Xu, Yingtao Zhu, Chuxi Xiao

    Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the firs

  87. Lei Sang, Yu Wang, Yiwen Zhang

    Heterogeneous graph neural networks (HGNNs) have demonstrated their superiority in exploiting auxiliary information for recommendation tasks. However, graphs constructed using meta-paths in HGNNs are usually too dense and contain a large number of noise edges. The propagation mechanism of HGNNs propagates even small amounts of noise in a graph to distant nei

  88. Songtao He, Erfang Shan, Xinyu Sun

    B\'eal et al. (Int J Game Theory 54, 2025) introduce the Diversity Owen value for TU-games with diversity constraints, and provide axiomatic characterizations using the axioms of fairness and balanced contributions. However, there exist logical flaws in the proofs of the uniqueness of these characterizations. In this note we provide the corrected proofs of t

  89. Jordan Larson, Alexander J. Wagner

    Analytical solutions to the lattice Boltzmann Equation make it possible to study the method itself, explore the properties of its collision operator, and identify implementations of boundary conditions. In this paper, we propose a method to find analytical solutions where the macroscopic flow profile is known. We test this method on bulk Couette flow aligned

  90. Shuohao Ping, Naren Sathishkumar, Wan-Hsuan Lin, Hanyu Wang

    Quantum Layout Synthesis (QLS) is a critical compilation stage that adapts quantum circuits to hardware constraints with an objective of minimizing the SWAP overhead. While heuristic tools demonstrate good efficiency, they often produce suboptimal solutions, and exact methods suffer from limited scalability. In this work, we propose ML-SABRE, a high-performa

  91. Mingyao Cui, Qunsong Zeng, Minze Chen, Zhanwei Wang

    Harnessing multi-level electron transitions, Rydberg Atomic REceivers (RAREs) can detect wireless signals across a wide range of frequency bands, from Megahertz to Terahertz. This capability enables multi-band wireless communications and sensing (CommunSense). Existing research on multi-band RAREs primarily focuses on experimental demonstrations, lacking a t

  92. Junyu Chen, Shuwen Wei, Yihao Liu, Aaron Carass

    Recent advances in deep learning-based medical image registration have shown that training deep neural networks~(DNNs) does not necessarily require medical images. Previous work showed that DNNs trained on randomly generated images with carefully designed noise and contrast properties can still generalize well to unseen medical data. Building on this insight

  93. Junyu Chen, Zirui Jiang, Jennifer M. Coughlin, Ian Cheong

    Dynamic positron emission tomography (PET) imaging combined with radiotracer kinetic modeling is a powerful technique for visualizing biological processes in the brain, offering valuable insights into brain functions and neurological disorders such as Alzheimer's and Parkinson's diseases. Accurate kinetic modeling relies heavily on the use of a metabolite-co

  94. Yixuan Wang, Shiqi Zhou, Chuanzhe Guo, Qingfu Zhu

    Evol-Instruct has made significant improvements as a data synthesis method in several areas. Existing methods typically rely on a fixed set of strategies to evolve, which require manual design and are monolithic in form. In addition, iterative evolution also makes the acquisition of hard samples expensive. In view of this, we propose the Tag-Evol framework,

  95. Shilin Xu, Yanwei Li, Rui Yang, Tao Zhang

    Recent works on large language models (LLMs) have successfully demonstrated the emergence of reasoning capabilities via reinforcement learning (RL). Although recent efforts leverage group relative policy optimization (GRPO) for MLLMs post-training, they constantly explore one specific aspect, such as grounding tasks, math problems, or chart analysis. There a

  96. Jittarin Jetwiriyanon, Teo Susnjak, Surangika Ranathunga

    This study investigates zero-shot forecasting capabilities of Time Series Foundation Models (TSFMs) for macroeconomic indicators. We apply TSFMs to forecasting economic indicators under univariate conditions, bypassing the need for train bespoke econometric models using and extensive training datasets. Our experiments were conducted on a case study dataset,

  97. Shruti Kumar, Xiaoyu Chen, Xiaomei Wang

    Several papers have delved into the challenges of human-AI-robot co-learning and co-adaptation. It has been noted that the terminology used to describe this collaborative relationship in existing studies needs to be more consistent. For example, the prefix "co" is used interchangeably to represent both "collaborative" and "mutual," and the terms "co-learning

  98. Jiaqi Sun, Shiyou Qian, Zhangchi Han, Wei Li

    Knowledge Graphs (KGs) structure real-world entities and their relationships into triples, enhancing machine reasoning for various tasks. While domain-specific KGs offer substantial benefits, their manual construction is often inefficient and requires specialized knowledge. Recent approaches for knowledge graph construction (KGC) based on large language mode

  99. Jia Yuxin, Li Jinye, Jia Yudong, Li Futing

    Potato functional genomics lags due to unsystematic gene information curation, gene identifier inconsistencies across reference genome versions, and the increasing volume of research publications. To address these limitations, we developed the Potato Knowledge Hub (http://www.potato-ai.top), leveraging Large Language Models (LLMs) and a systematically curate

  100. Oliver Knill

    We define an evolution of multiple particles on a discrete manifold $G$. Each particle alone moves on geodesics and particles can interact if they are on the same facet. They move deterministically and reversibly on the frame bundle $P$ of the abstract simplicial complex $G$. Particles are signed and each is represented by a totally ordered maximal simplex $