Skip to content

May 2025 arXiv papers — page 59

Showing 5,8015,900 of 24,552 papers

  1. Yewon Han, Seoyun Yang, Taesup Kim

    Test-time adaptation (TTA) enhances model robustness by enabling adaptation to target distributions that differ from training distributions, improving real-world generalizability. However, most existing TTA approaches focus on adjusting the conditional distribution and therefore exhibit poor calibration, as they rely on uncertain predictions in the absence o

  2. Ryan Soh-Eun Shim, Domenico De Cristofaro, Chengzhi Martin Hu, Alessandro Vietti

    Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational similarity. However, prior work does not control for phonetic overlap between equivalent utterances, which may artificially suppor

  3. Aggrey Muhebwa, Khotso Selialia, Fatima Anwar, Khalid K. Osman

    Federated learning on heterogeneous (non-IID) client data experiences slow convergence due to client drift. To address this challenge, we propose Kuramoto-FedAvg, a federated optimization algorithm that reframes the weight aggregation step as a synchronization problem inspired by the Kuramoto model of coupled oscillators. The server dynamically weighs each c

  4. Ahan Prasannakumar Shetty

    Machine translation has become a critical tool in bridging linguistic gaps, especially between languages as diverse as English and Hindi. This paper comprehensively evaluates various machine translation models for translating between English and Hindi. We assess the performance of these models using a diverse set of automatic evaluation metrics, both lexical

  5. Ho Hin Lee, Quan Liu, Shunxing Bao, Yuankai Huo

    Large kernel convolutions offer a scalable alternative to vision transformers for high-resolution 3D volumetric analysis, yet naively increasing kernel size often leads to optimization instability. Motivated by the spatial bias inherent in effective receptive fields (ERFs), we theoretically demonstrate that structurally re-parameterized blocks induce spatial

  6. Kunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng Hwang

    Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. Nevertheless, the coarse-to-fine methodology inherent in VAR results in exponential growth of the KV cache during inference, causing consider

  7. Yeongmin Kim, Heesun Bae, Byeonghu Na, Il-Chul Moon

    Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identificatio

  8. Anggiat Mora Simamora, Asep Denih, Mohamad Iqbal Suriansyah

    This paper presents the design, implementation, and evaluation of an IoT-based robotic system for mapping and monitoring indoor air quality. The primary objective was to develop a mobile robot capable of autonomously mapping a closed environment, detecting concentrations of CO$_2$, volatile organic compounds (VOCs), smoke, temperature, and humidity, and tran

  9. Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

    Typical methods for evaluating the performance of language models evaluate their ability to answer questions accurately. These evaluation metrics are acceptable for determining the extent to which language models can understand and reason about text in a general sense, but fail to capture nuanced capabilities, such as the ability of language models to recogn

  10. Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo

    Large Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection attacks remains underexplored. This study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hij

  11. Zheng Wang, Xiaobin Rong, Yu Sun, Tianchi Sun

    Although deep learning based multi-channel speech enhancement has achieved significant advancements, its practical deployment is often limited by constrained computational resources, particularly in low signal-to-noise ratio (SNR) conditions. In this paper, we propose a lightweight hybrid dual-channel speech enhancement system that combines independent vecto

  12. Shuoming Zhang, Jiacheng Zhao, Chunwei Xia, Zheng Wang

    Large language models (LLMs) have the potential to revolutionize how we design and implement compilers and code translation tools. However, existing LLMs struggle to handle long and complex programs. We introduce LEGO-Compiler, a novel neural compilation system that leverages LLMs to translate high-level languages into assembly code. Our approach centers on

  13. Michiaki Takiwaki

    One-parameter persistence modules are applied to various subjects as tools in data analysis. On the other hand, since the theoretical study of multi-parameter persistence modules is not enough and in progress, they have few applications. The sheaf theory is expected to elucidate detailed properties of persistence modules and give features of multi-parameter

  14. Ryota Mizuno, Kazuhiko Kuroki, Masayuki Ochi

    Estimating the local two-particle vertex functions, which are crucial for capturing the spatial fluctuation of the effective field beyond the single-site DMFT, is still challenging. In our previous work, we developed a computationally efficient method for estimating the local full-vertex in DMFT, where we can obtain the local two-particle full-vertex from th

  15. Jeongsoo Choi, Zhikang Niu, Ji-Hoon Kim, Chunhui Wang

    The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to the implicit guidance of diffusion models in learning complex intermediate representations. To address this, we propose

  16. A. V. Tsiganov

    We present some new Poisson bivectors that are invariants by the flow of the nonholonomic Suslov problem. Two rank four invariant Poisson bivectors have globally defined Casimir functions and, therefore, define cubic Poisson brackets on the five dimensional state space with standard symplectic leaves. For the Suslov gyrostat in the potential field we found r

  17. Belacel Amar, Bougoutaia Amar, Rueda Pilar

    We explore the procedure given by left-hand quotients in the context of weighted holomorphic ideals. On the one hand, we show that this procedure does not generate new ideals other than the ideal of weighted holomorphic mappings when considering the left-hand quotients induced by the ideals of $p$-compact, weakly $p$-compact, unconditionally $p$-compact, app

  18. Zewei Xiong, Meng-Ru Wu, Noshad Khosravi Largani, Tobias Fischer

    Core-collapse supernovae undergoing a first-order quantum chromodynamics (QCD) phase transition experience the collapse of the central proto-neutron star that leads to a second bounce. This event is accompanied by the release of a second neutrino burst. Unlike the first stellar core bounce neutrino burst which consists exclusively of electron neutrinos, the

  19. Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan

    Large language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organizational structures that struggle to adapt as task complexity and

  20. Christian Janos Lebeda, Mathieu Even, Aurélien Bellet, Julie Josse

    Estimating causal effects from observational data is essential in fields such as medicine, economics and social sciences, where privacy concerns are paramount. We propose a general, model-agnostic framework for differentially private estimation of average treatment effects (ATE) that avoids strong structural assumptions on the data-generating process or the

  21. Yanzhen Shen, Sihao Chen, Xueqiang Xu, Yunyi Zhang

    While significant progress has been made with dual- and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case that is often overlooked yet important in downstream applications. Current dense retrievers struggle with such queries, such that the retrieved results do not respect the logical constraints implied in the q

  22. Shadi Alijani, Homayoun Najjaran

    Conformal prediction (CP) provides a framework for constructing prediction sets with guaranteed coverage, assuming exchangeable data. However, real-world scenarios often involve distribution shifts that violate exchangeability, leading to unreliable coverage and inflated prediction sets. To address this challenge, we first introduce Reconstruction Loss-Scale

  23. Dingyu Yao, Bowen Shen, Zheng Lin, Wei Liu

    The Key-Value (KV) cache in generative large language models (LLMs) introduces substantial memory overhead. Existing works mitigate this burden by offloading or compressing the KV cache. However, loading the entire cache incurs significant latency due to PCIe bandwidth bottlenecks in CPU-GPU communication, while aggressive compression causes notable performa

  24. Jiameng Li, Teodora Popordanoska, Aleksei Tiulpin, Sebastian G. Gruber

    Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment planning. These biomarkers are typically estimated from segmentation outputs by computing region-wise ratios. Despite the high-stakes nature of clinical decision making, existing methods

  25. Zongguo Si, Hongxin Wang, Lei Wang, Yang Xiao

    We develop a framework based on the full one-loop finite-temperature effective potential model, within which the bubble wall velocity is calculated using the local thermal equilibrium (LTE) approximation, and the kinetic energy fraction $K$ is computed directly. In cosmological phase transitions, these quantities play a critical role in determining the resul

  26. Biju Saha, Suman Sarkar, Arunima Banerjee

    About 30\% of disk galaxies show lopsidedness in their stellar disk. Although such a large-scale asymmetry in the disk can be primarily looked upon as a long-lived mode ($m=1$), the physical origin of the lopsidedness in the disk continues to be a puzzle. In this work, we employ a transfer-learning approach for the automated identification of lopsided galaxi

  27. Kaiqing Lin, Zhiyuan Yan, Ke-Yue Zhang, Li Hao

    Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore the valuable prior knowledge of known facial identities, e.g.,

  28. Ritesh K. Singh, Souradeep Sasmal, S. Nautiyal, A. K. Pan

    We present a device-independent (DI) self-testing protocol in a constrained prepare-measure scenario, based on the $n-$bit parity-oblivious multiplexing (POM) task. In this scenario, a parity-oblivious constraint is imposed on the preparations, allowing us to define a classical bound derived from a preparation noncontextual ontological model. We derive the o

  29. Masaki Murooka, Kensuke Fukumitsu, Marwan Hamze, Mitsuharu Morisawa

    To enable humanoid robots to work robustly in confined environments, multi-contact motion that makes contacts not only at extremities, such as hands and feet, but also at intermediate areas of the limbs, such as knees and elbows, is essential. We develop a method to realize such whole-body multi-contact motion involving contacts at intermediate areas by a hu

  30. Zhanpeng Cui, Bo Hou

    We introduce the notion of quasi-triangular Novikov bialgebras, which constructed from solutions of the Novikov Yang-Baxter equation whose symmetric parts are invariant. Triangular Novikov bialgebras and factorizable Novikov bialgebras are important subclasses of quasi-triangular Novikov bialgebras. A factorizable Novikov bialgebra induces a factorization of

  31. Dan Peng, Zhihui Fu, Zewen Ye, Zhuoran Song

    Sparse attention methods exploit the inherent sparsity in attention to speed up the prefilling phase of long-context inference, mitigating the quadratic complexity of full attention computation. While existing sparse attention methods rely on predefined patterns or inaccurate estimations to approximate attention behavior, they often fail to fully capture the

  32. Yeonjoon Jung, Daehyun Ahn, Hyungjun Kim, Taesu Kim

    Low-Rank Adaptation (LoRA) is a popular method for parameter-efficient fine-tuning (PEFT) of generative models, valued for its simplicity and effectiveness. Despite recent enhancements, LoRA still suffers from a fundamental limitation: overfitting when the bottleneck is widened. It performs best at ranks 32-64, yet its accuracy stagnates or declines at highe

  33. Yu Xi, Haoyu Li, Xiaoyu Gu, Yidi Jiang

    Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explore the entire search space without explicitly prioritizing keyword detection, often leading to suboptimal performance. In this paper, we propose an effective keyword-specific KWS fr

  34. Yujie Yang, Bing Yang, Xiaofei Li

    Online multichannel speech enhancement has been intensively studied recently. Though Mel-scale frequency is more matched with human auditory perception and computationally efficient than linear frequency, few works are implemented in a Mel-frequency domain. To this end, this work proposes a Mel-scale framework (namely Mel-McNet). It processes spectral and sp

  35. Kai Li, Conggai Li, Xin Yuan, Shenghong Li

    This paper focuses on Zero-Trust Foundation Models (ZTFMs), a novel paradigm that embeds zero-trust security principles into the lifecycle of foundation models (FMs) for Internet of Things (IoT) systems. By integrating core tenets, such as continuous verification, least privilege access (LPA), data confidentiality, and behavioral analytics into the design, t

  36. Gilad Orr, Eliran Talker

    Coherence time of thermal photons in rubidium vapor cells with varying thicknesses, reveal that there is clear dependence of the photon correlation time on cell thickness. Standard theoretical models accurately predict the coherence time in centimeter-scale cells. In this study we demonstrated, that these models break down in micrometer and sub-micrometer re

  37. Alejandro Murillo-Gonzalez, Lantao Liu

    Autonomous robots operating in complex, unstructured environments face significant challenges due to latent, unobserved factors that obscure their understanding of both their internal state and the external world. Addressing this challenge would enable robots to develop a more profound grasp of their operational context. To tackle this, we propose a novel fr

  38. Vaishnavi Gupta, Hitesh Raundal

    The paper investigates biorderability of knot quandles of prime knots up to eight crossings. We prove that knot quandles of knots $6_3$, $8_7$, $8_8$, $8_{10}$ and $8_{16}$ can not be biorderable. However, we see that knot quandles of knots $4_1$, $6_1$, $6_2$, $7_6$, $7_7$, $8_1$, $8_2$, $8_3$, $8_4$, $8_5$, $8_6$, $8_9$, $8_{11}$, $8_{12}$, $8_{13}$, $8_{1

  39. Li Zeng, Zeming Liu, Chong Feng, Heyan Huang

    Model editing aims to correct errors and outdated knowledge in the Large language models (LLMs) with minimal cost. Prior research has proposed a variety of datasets to assess the effectiveness of these model editing methods. However, most existing datasets only require models to output short phrases or sentences, overlooks the widespread existence of documen

  40. Hu Xiaobin, Liang Yujie, Luo Donghao, Peng Xu

    While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three key reasons:(1) Current metrics inadequately reflect human perception, particularly in unpaired try-on settings;(2)Most existing test sets are limited to indoor scenarios, lacking c

  41. Modibo K. Camara, Nicole Immorlica, Brendan Lucier

    In many settings -- like market research and social choice -- people may be presented with unfamiliar options. Classical mechanisms may perform poorly because they fail to incentivize people to learn about these options, or worse, encourage counterproductive information acquisition. We formalize this problem in a model of robust mechanism design where agents

  42. Jianghang Lin, Yue Hu, Jiangtao Shen, Yunhang Shen

    Open vocabulary image segmentation tackles the challenge of recognizing dynamically adjustable, predefined novel categories at inference time by leveraging vision-language alignment. However, existing paradigms typically perform class-agnostic region segmentation followed by category matching, which deviates from the human visual system's process of recogniz

  43. E. C. I. Paterson, M. E. Tobar, M. Goryachev, J. Bourhill

    We report the experimental observation of two distinct Berry phases ($+\frac{2\pi}{3}$ and $-\frac{2\pi}{3}$) generated on the surface of a M\"{o}bius cavity resonator at microwave frequencies supporting the TE$_{1,0,n}$ mode family. This resonator consists of a twisted, mirror-asymmetric prism with a cross-section of the triangular $D_3$ symmetry group, ben

  44. Jiongchao Jin, Xiuju Fu, Xiaowei Gao, Tao Cheng

    Maritime transportation is the backbone of global trade, making ship inspection essential for ensuring maritime safety and environmental protection. Port State Control (PSC), conducted by national ports, enforces compliance with safety regulations, with ship detention being the most severe consequence, impacting both ship schedules and company reputations. T

  45. Rasoul Zahedifar, Sayyed Ali Mirghasemi, Mahdieh Soleymani Baghshah, Alireza Taheri

    This study presents the LLM-Agent-Controller, a multi-agent large language model (LLM) system developed to address a wide range of problems in control engineering (Control Theory). The system integrates a central controller agent with multiple specialized auxiliary agents, responsible for tasks such as controller design, model representation, control analysi

  46. Panos Pantidis, Lampros Svolos, Diab Abueidda, Mostafa E. Mobasher

    We present a novel formulation for modeling phase-field fracture propagation based on the Integrated Finite Element Neural Network (IFENN) framework. IFENN is a hybrid solver scheme that utilizes neural networks as PDE solvers within FEM, preserving accuracy via residual minimization while achieving speed-up via swift network predictions and reduction of the

  47. Juntong Wu, Zijing Liu, He Cao, Hao Li

    In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequen

  48. George Karantaidis, Athanasios Pantsios, Ioannis Kompatsiaris, Symeon Papadopoulos

    Synthetic aperture radar automatic target recognition (SAR-ATR) systems have rapidly evolved to tackle incremental recognition challenges in operational settings. Data scarcity remains a major hurdle that conventional SAR-ATR techniques struggle to address. To cope with this challenge, we propose a few-shot class-incremental learning (FSCIL) framework based

  49. Haofan Ren, Zunjie Zhu, Xiang Chen, Ming Lu

    Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple bu

  50. Shi-Yu Tian, Zhi Zhou, Wei Dong, Kun-Yang Yu

    Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning over tabular data in real-world applications has been overlooked. For instance, applications such as business intelligence demand not only multi-step numerical reasoning with tabl

  51. Ying Xiao, Jie Huang, Ruijuan He, Jing Xiao

    Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question

  52. Yuan Feng, Yukun Cao, Hairu Wang, Xike Xie

    Sketches, probabilistic structures for estimating item frequencies in infinite data streams with limited space, are widely used across various domains. Recent studies have shifted the focus from handcrafted sketches to neural sketches, leveraging memory-augmented neural networks (MANNs) to enhance the streaming compression capabilities and achieve better spa

  53. Jianan Lou, Rong Zhang

    Global Navigation Satellite System (GNSS) is essential for autonomous driving systems, unmanned vehicles, and various location-based technologies, as it provides the precise geospatial information necessary for navigation and situational awareness. However, its performance is often degraded by Non-Line-Of-Sight (NLOS) and multipath effects, especially in urb

  54. Vladimir Gol'dshtein, Reuven Segev

    We outline here a simple mathematical introduction to the notions of multipoles for a general extensive property $\Pi$ from the point of view of continuum mechanics. Classically, $\Pi$ is the electric charge, but the theory is not limited to electrostatics. The proposed framework allows a simple computation of the bound "charges" and bound multipoles of lowe

  55. Zhaowei Zhang, Xiaobo Wang, Minghua Yi, Mengmeng Wang

    Achieving political consensus is crucial yet challenging for the effective functioning of social governance. However, although frontier AI systems represented by large language models (LLMs) have developed rapidly in recent years, their capabilities in this scope are still understudied. In this paper, we introduce PoliCon, a novel benchmark constructed from

  56. Diogo Da Silva Machado

    In this paper, we provide formulas for the sum of residues of type Camacho-Sad of a holomorphic foliation with respect to an invariant analytic subvariety. As application, in context of projective foliations, we obtain a formula that relates the sum these residues with the degree and other characteristics of the invariant subvariety. Furthermore, we establis

  57. Shouqiao Wang, Davide Crapis, Ciamac C. Moallemi

    This paper presents a comprehensive framework for transaction posting and pricing in Layer 2 (L2) blockchain systems, focusing on challenges stemming from fluctuating Layer 1 (L1) gas fees and the congestion issues within L2 networks. Existing methods have focused on the problem of optimal posting strategies to L1 in isolation, without simultaneously conside

  58. Wei Su, Xi Zou

    Modelling rarefied gas flow via the Boltzmann equation plays a vital role in many areas. Due to the high dimensionality of this kinetic equation and the coexistence of multiple characteristic scales in the transport processes, conventional solution strategies incur prohibitively high computational costs and are inadequate for rapid response for parametric an

  59. Jiongchao Jin, Shengchu Zhao, Dajun Chen, Wei Jiang

    Time consumption and the complexity of manual layout design make automated layout generation a critical task, especially for multiple applications across different mobile devices. Existing graph-based layout generation approaches suffer from limited generative capability, often resulting in unreasonable and incompatible outputs. Meanwhile, vision based gener

  60. Subham Dutta, Pralay Kumar Karmakar

    The effective inductive (L), capacitive (C), and resistive (R) behavior of a plasma sheath in a conjoint coupled form is well familiar among plasma physics communities. A dynamic sheath instability in laboratory plasmas is systematically modelled herein as an electrical series-resonance LCR circuit of the above kind. It theoretically yields experimentally ob

  61. Minkyu Kim, Kiyoung Seong, Dongyeop Woo, Sungsoo Ahn

    We address the challenge of training diffusion models to sample from unnormalized energy distributions in the absence of data, the so-called diffusion samplers. Although these approaches have shown promise, they struggle to scale in more demanding scenarios where energy evaluations are expensive and the sampling space is high-dimensional. To address this lim

  62. Jochen L. Cremer

    The electricity system becomes more complex, connecting massive numbers of end-users and distributed generators. Adding or removing grid connections requires expert studies to align technical constraints with user requests. In times of labour shortages, carrying out these studies represents a significant amount of time that engineers at system operators spen

  63. Georgios Mappouras

    With the rise of artificial intelligence (A.I.) and large language models like ChatGPT, a new race for achieving artificial general intelligence (A.G.I) has started. While many speculate how and when A.I. will achieve A.G.I., there is no clear agreement on how A.G.I. can be detected in A.I. models, even when popular tools like the Turing test (and its modern

  64. Derong Xu, Yi Wen, Pengyue Jia, Yingyi Zhang

    Large Language Models (LLMs) have recently been widely adopted in conversational agents. However, the increasingly long interactions between users and agents accumulate extensive dialogue records, making it difficult for LLMs with limited context windows to maintain a coherent long-term dialogue memory and deliver personalized responses. While retrieval-augm

  65. Xufeng Duan, Zhaoqian Yao, Yunhao Zhang, Shaonan Wang

    Large language models (LLMs) have been found to develop surprising internal specializations: Individual neurons, attention heads, and circuits become selectively sensitive to syntactic structure, reflecting patterns observed in the human brain. While this specialization is well-documented, how it emerges during training and what influences its development re

  66. Haoyu Zhang, Wentao Zhang, Hao Miao, Xinke Jiang

    Spatio-Temporal Graph Neural Networks (STGNNs) have emerged as a powerful tool for modeling dynamic graph-structured data across diverse domains. However, they often fail to generalize in Spatio-Temporal Out-of-Distribution (STOOD) scenarios, where both temporal dynamics and spatial structures evolve beyond the training distribution. To address this problem,

  67. Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Mehrdad Noori

    Test-Time Training (TTT) has emerged as a promising solution to address distribution shifts in 3D point cloud classification. However, existing methods often rely on computationally expensive backpropagation during adaptation, limiting their applicability in real-world, time-sensitive scenarios. In this paper, we introduce SMART-PC, a skeleton-based framewor

  68. Guy F. de Teramond, Arpon Paul, Hans Gunter Dosch, Stanley J. Brodsky

    We extend our recent analytic study of the strong coupling $\alpha_{\rm eff}$ in the nonperturbative and near-perturbative regimes~\cite{deTeramond:2024ikl} by imposing rigorous renormalization-group results from asymptotically free gauge theories at $Q^2 \to \infty$. The asymptotic boundary conditions modify the scaling properties of $\alpha_{\rm eff}$ at l

  69. Jialei Chen, Yuanbo Xu, Yiheng Jiang

    In this paper, we focus on the often-overlooked issue of embedding collapse in existing diffusion-based sequential recommendation models and propose ADRec, an innovative framework designed to mitigate this problem. Diverging from previous diffusion-based methods, ADRec applies an independent noise process to each token and performs diffusion across the entir

  70. Yiyun Zhou, Zheqi Lv, Shengyu Zhang, Jingyuan Chen

    Knowledge Tracing (KT) is a core component of Intelligent Tutoring Systems, modeling learners' knowledge state to predict future performance and provide personalized learning support. Traditional KT models assume that learners' learning abilities remain relatively stable over short periods or change in predictable ways based on prior performance. However, in

  71. Anwesh Ray

    Motivated by the work of Greenberg-Vatsal and Emerton-Pollack-Weston, I investigate the extent to which Mazur's conjecture on the growth of Selmer ranks in $\mathbb{Z}_p$-extensions of an imaginary quadratic field persists under $p$-congruences between Galois representations. As a first step, I establish Mazur's conjecture for certain triples $(E, K, p)$ und

  72. Chen Jiang, Haidong Liu

    Let $X$ be a $\mathbb Q$-factorial weak Fano $3$-fold with at worst isolated canonical singularities. We show that the $\mathbb Q$-Fano index of $X$ is at most $61$.

  73. Junhyung Kim, Hokyun Lee, Jaeheung Park

    Advancements in optimization solvers and computing power have led to growing interest in applying whole-body model predictive control (WB-MPC) to bipedal robots. However, the high degrees of freedom and inherent model complexity of bipedal robots pose significant challenges in achieving fast and stable control cycles for real-time performance. This paper int

  74. Dong Liu, Yanxuan Yu, Jiayi Zhang, Yifan Li

    Diffusion Transformers (DiT) are powerful generative models but remain computationally intensive due to their iterative structure and deep transformer stacks. To alleviate this inefficiency, we propose \textbf{FastCache}, a hidden-state-level caching and compression framework that accelerates DiT inference by exploiting redundancy within the model's internal

  75. Zhongqin Wang, J. Andrew Zhang, Kai Wu, Y. Jay Guo

    Accurate water level sensing is essential for flood monitoring, agricultural irrigation, and water resource optimization. Traditional methods require dedicated sensor deployments, leading to high installation costs, vulnerability to interference, and limited resolution. This work proposes PMNs-WaterSense, a novel scheme leveraging Channel State Information (

  76. Yuxing Lu, Gecheng Fu, Wei Wu, Xukai Zhao

    Existing medical RAG systems mainly leverage knowledge from medical knowledge bases, neglecting the crucial role of experiential knowledge derived from similar patient cases -- a key component of human clinical reasoning. To bridge this gap, we propose DoctorRAG, a RAG framework that emulates doctor-like reasoning by integrating both explicit clinical knowle

  77. Yi Feng, Kaito Fujii, Stratis Skoulakis, Xiao Wang

    Since Polyak's pioneering work, heavy ball (HB) momentum has been widely studied in minimization. However, its role in min-max games remains largely unexplored. As a key component of practical min-max algorithms like Adam, this gap limits their effectiveness. In this paper, we present a continuous-time analysis for HB with simultaneous and alternating update

  78. Jintao Tong, Wenwei Jin, Pengda Qin, Anqi Li

    Large vision-language models (LVLMs) excel at multimodal understanding but suffer from high computational costs due to redundant vision tokens. Existing pruning methods typically rely on single-layer attention scores to rank and prune redundant visual tokens to solve this inefficiency. However, as the interaction between tokens and layers is complicated, thi

  79. Juntong Wang, Jiarui Wang, Huiyu Duan, Guangtao Zhai

    Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this critical gap, we introduce TDVE-DB, a large-scale benchmark dataset for text-driven video editing. TDVE-DB consists of 3,857

  80. Yongyi Zang, Jingyi Li, Qiuqiang Kong

    Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this work, we reveal that pretrained one-step audio source separation models can be leveraged for multi-step separation without additional traini

  81. Yachuan Liu, Xiaochun Wei, Lin Shi, Xinnuo Li

    Large language models (LLMs) face significant challenges in ex-ante reasoning, where analysis, inference, or predictions must be made without access to information from future events. Even with explicit prompts enforcing temporal cutoffs, LLMs often generate outputs influenced by internalized knowledge of events beyond the specified cutoff. This paper introd

  82. Jerry Yao-Chieh Hu, Xiwen Zhang, Maojiang Su, Zhao Song

    We study the computational limits of learning $k$-bit Boolean functions (specifically, $\mathrm{AND}$, $\mathrm{OR}$, and their noisy variants), using a minimalist single-head softmax-attention mechanism, where $k=\Theta(d)$ relevant bits are selected from $d$ inputs. We show that these simple $\mathrm{AND}$ and $\mathrm{OR}$ functions are unsolvable with a

  83. Amartya Purushottam, Jack Yan, Christopher Yu, Joao Ramos

    Humanoid robots can support human workers in physically demanding environments by performing tasks that require whole-body coordination, such as lifting and transporting heavy objects.These tasks, which we refer to as Dynamic Mobile Manipulation (DMM), require the simultaneous control of locomotion, manipulation, and posture under dynamic interaction forces.

  84. Tanjil Hasan Sakib, Md. Tanzib Hosain, Md. Kishor Morol

    Small Language Models (SLMs) have gained substantial attention due to their ability to execute diverse language tasks successfully while using fewer computer resources. These models are particularly ideal for deployment in limited environments, such as mobile devices, on-device processing, and edge systems. In this study, we present a complete assessment of

  85. Yejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub Han

    Implicit hate speech detection is challenging due to its subtlety and reliance on contextual interpretation rather than explicit offensive words. Current approaches rely on contrastive learning, which are shown to be effective on distinguishing hate and non-hate sentences. Humans, however, detect implicit hate speech by first identifying specific targets wit

  86. Mohammed Djameleddine Belgoumri, Mohamed Reda Bouadjenek, Hakim Hacid, Imran Razzak

    Training large neural networks (NNs) requires optimizing high-dimensional data-dependent loss functions. The optimization landscape of these functions is often highly complex and textured, even fractal-like, with many spurious local minima, ill-conditioned valleys, degenerate points, and saddle points. Complicating things further is the fact that these lands

  87. Robert Fraser, Kyle Hambrook, Donggeun Ryou

    We prove the optimality of the exponent in the Mockenhaupt-Mitsis-Bak-Seeger Fourier restriction theorem in all dimensions $d$ and the full parameter range $0 < a,b < d$. Our construction is deterministic and also yields Salem sets.

  88. Jie Wu

    A graph $G$ has the $k$-strong parity property if for any $X\subseteq V(G)$ with $|X|$ even, $G$ contains a spanning subgraph $F$ with $d_F(u)\equiv1$ (mod 2) for each $u\in X$ and $d_F(v)\in\{k,k+2,k+4,\ldots\}$ for each $v\in V(G)\setminus X$, where $k\geq2$ is an even integer. Kano and Matsumura proposed a characterization for a graph with the $k$-strong

  89. Liangwei Nathan Zheng, Wei Emma Zhang, Mingyu Guo, Olaf Maennel

    Effectively managing missing modalities is a fundamental challenge in real-world multimodal learning scenarios, where data incompleteness often results from systematic collection errors or sensor failures. Sparse Mixture-of-Experts (SMoE) architectures have the potential to naturally handle multimodal data, with individual experts specializing in different m

  90. Jun Lai, Jinrui Zhang

    Inverse scattering in layered media has a wide range of applications, examples including geophysical exploration, medical imaging, and remote sensing. In this paper, we develop a selective focusing method for identifying multiple unknown buried scatterers in a layered medium. The method is derived through the asymptotic analysis of the time reversal operator

  91. Cen Liu, Xiangyun Zhou, Nan Yang, Salman Durrani

    This work studies near-field secure communications through transmit beamfocusing. We examine the benefit of having a protected eavesdropper-free zone around the legitimate receiver, and we determine the worst-case secrecy performance against a potential eavesdropper located anywhere outside the protected zone. A max-min optimization problem is formulated for

  92. Jiyu Hu, Haijiang Zeng, Zhen Tian

    In recent years, image classification, as a core task in computer vision, relies on high-quality labelled data, which restricts the wide application of deep learning models in practical scenarios. To alleviate the problem of insufficient labelled samples, semi-supervised learning has gradually become a research hotspot. In this paper, we construct a semi-sup

  93. Dongzhe Zheng, Wenjie Mei

    Learning unknown dynamics under environmental (or external) constraints is fundamental to many fields (e.g., modern robotics), particularly challenging when constraint information is only locally available and uncertain. Existing approaches requiring global constraints or using probabilistic filtering fail to fully exploit the geometric structure inherent in

  94. Guanhao Li

    This paper provides a combinatorial proof to show that, in the study of maximal Condorcet domains, the class of peak-pit Condorcet domains, the class of connected Condorcet domains, and the class of directly connected Condorcet domains are all equivalent.

  95. Xiaoyu Sun, Yang Yang, Xunde Dong

    In the field of automatic Electrocardiogram (ECG) diagnosis, due to the relatively limited amount of labeled data, how to build a robust ECG pretrained model based on unlabeled data is a key area of focus for researchers. Recent advancements in contrastive learning-based ECG pretrained models highlight the potential of exploiting the additional patient-level

  96. Gihoon Kim, Hyungjin Park, Taesup Kim

    Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the model's original generative capabilities. However, this process often leads to overfitting, where the model ignores the user's prompt and merely replicates the reference images. We attribute this issue to a fundamen

  97. Nakul Poudel, Zixin Yang, Kelly Merrell, Richard Simon

    Intra-operative data captured during image-guided surgery lacks sub-surface information, where key regions of interest, such as vessels and tumors, reside. Image-to-physical registration enables the fusion of pre-operative information and intra-operative data, typically represented as a point cloud. However, this registration process struggles due to partial

  98. Pieter van Goor, Robert Mahony

    This paper introduces the concept of a synchronous model as an extension of the internal model concept used in observer design for dynamical systems. A system is said to contain a synchronous model of another if there is a suitable error function between the two systems that remains stationary for all of the trajectories of the two systems. A system is said

  99. Rui Zhao, Yuze Fan, Ziguo Chen, Fei Gao

    End-to-end learning has emerged as a transformative paradigm in autonomous driving. However, the inherently multimodal nature of driving behaviors and the generalization challenges in long-tail scenarios remain critical obstacles to robust deployment. We propose DiffE2E, a diffusion-based end-to-end autonomous driving framework. This framework first performs

  100. Lavanya Prahallad, Radhika Mamidi

    We present a critical discourse analysis of the 2024 U.S. presidential debates, examining Donald Trump's rhetorical strategies in his interactions with Joe Biden and Kamala Harris. We introduce a novel annotation framework, BEADS (Bias Enriched Annotation for Dialogue Structure), which systematically extends the DAMSL framework to capture bias driven and adv