May 2025 arXiv papers — page 59
Showing 5,801–5,900 of 24,552 papers
Yewon Han, Seoyun Yang, Taesup Kim
Test-time adaptation (TTA) enhances model robustness by enabling adaptation to target distributions that differ from training distributions, improving real-world generalizability. However, most existing TTA approaches focus on adjusting the conditional distribution and therefore exhibit poor calibration, as they rely on uncertain predictions in the absence o
Ryan Soh-Eun Shim, Domenico De Cristofaro, Chengzhi Martin Hu, Alessandro Vietti
Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational similarity. However, prior work does not control for phonetic overlap between equivalent utterances, which may artificially suppor
Kuramoto-FedAvg: Using Synchronization Dynamics to Improve Federated Learning Optimization under Statistical Heterogeneity
cs.LGAggrey Muhebwa, Khotso Selialia, Fatima Anwar, Khalid K. Osman
Federated learning on heterogeneous (non-IID) client data experiences slow convergence due to client drift. To address this challenge, we propose Kuramoto-FedAvg, a federated optimization algorithm that reframes the weight aggregation step as a synchronization problem inspired by the Kuramoto model of coupled oscillators. The server dynamically weighs each c
Ahan Prasannakumar Shetty
Machine translation has become a critical tool in bridging linguistic gaps, especially between languages as diverse as English and Hindi. This paper comprehensively evaluates various machine translation models for translating between English and Hindi. We assess the performance of these models using a diverse set of automatic evaluation metrics, both lexical
Ho Hin Lee, Quan Liu, Shunxing Bao, Yuankai Huo
Large kernel convolutions offer a scalable alternative to vision transformers for high-resolution 3D volumetric analysis, yet naively increasing kernel size often leads to optimization instability. Motivated by the spatial bias inherent in effective receptive fields (ERFs), we theoretically demonstrate that structurally re-parameterized blocks induce spatial
Kunjun Li, Zigeng Chen, Cheng-Yen Yang, Jenq-Neng Hwang
Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scalability, and zero-shot generalization. Nevertheless, the coarse-to-fine methodology inherent in VAR results in exponential growth of the KV cache during inference, causing consider
Yeongmin Kim, Heesun Bae, Byeonghu Na, Il-Chul Moon
Direct preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identificatio
Anggiat Mora Simamora, Asep Denih, Mohamad Iqbal Suriansyah
This paper presents the design, implementation, and evaluation of an IoT-based robotic system for mapping and monitoring indoor air quality. The primary objective was to develop a mobile robot capable of autonomously mapping a closed environment, detecting concentrations of CO$_2$, volatile organic compounds (VOCs), smoke, temperature, and humidity, and tran
Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo
Typical methods for evaluating the performance of language models evaluate their ability to answer questions accurately. These evaluation metrics are acceptable for determining the extent to which language models can understand and reason about text in a general sense, but fail to capture nuanced capabilities, such as the ability of language models to recogn
Guanyu Hou, Jiaming He, Yinhang Zhou, Ji Guo
Large Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection attacks remains underexplored. This study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack, Instruction Following Attack, Context Injection Attack, and Judgment Hij
Zheng Wang, Xiaobin Rong, Yu Sun, Tianchi Sun
Although deep learning based multi-channel speech enhancement has achieved significant advancements, its practical deployment is often limited by constrained computational resources, particularly in low signal-to-noise ratio (SNR) conditions. In this paper, we propose a lightweight hybrid dual-channel speech enhancement system that combines independent vecto
Shuoming Zhang, Jiacheng Zhao, Chunwei Xia, Zheng Wang
Large language models (LLMs) have the potential to revolutionize how we design and implement compilers and code translation tools. However, existing LLMs struggle to handle long and complex programs. We introduce LEGO-Compiler, a novel neural compilation system that leverages LLMs to translate high-level languages into assembly code. Our approach centers on
An isometry theorem induced by the Radon transform between the convolution and interleaving distances
math.ATMichiaki Takiwaki
One-parameter persistence modules are applied to various subjects as tools in data analysis. On the other hand, since the theoretical study of multi-parameter persistence modules is not enough and in progress, they have few applications. The sheaf theory is expected to elucidate detailed properties of persistence modules and give features of multi-parameter
Improvement of the simplification method for the local two-particle full-vertex towards precise frequency behavior
cond-mat.str-elRyota Mizuno, Kazuhiko Kuroki, Masayuki Ochi
Estimating the local two-particle vertex functions, which are crucial for capturing the spatial fluctuation of the effective field beyond the single-site DMFT, is still challenging. In our previous work, we developed a computationally efficient method for estimating the local full-vertex in DMFT, where we can obtain the local two-particle full-vertex from th
Jeongsoo Choi, Zhikang Niu, Ji-Hoon Kim, Chunhui Wang
The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to the implicit guidance of diffusion models in learning complex intermediate representations. To address this, we propose
A. V. Tsiganov
We present some new Poisson bivectors that are invariants by the flow of the nonholonomic Suslov problem. Two rank four invariant Poisson bivectors have globally defined Casimir functions and, therefore, define cubic Poisson brackets on the five dimensional state space with standard symplectic leaves. For the Suslov gyrostat in the potential field we found r
Belacel Amar, Bougoutaia Amar, Rueda Pilar
We explore the procedure given by left-hand quotients in the context of weighted holomorphic ideals. On the one hand, we show that this procedure does not generate new ideals other than the ideal of weighted holomorphic mappings when considering the left-hand quotients induced by the ideals of $p$-compact, weakly $p$-compact, unconditionally $p$-compact, app
Zewei Xiong, Meng-Ru Wu, Noshad Khosravi Largani, Tobias Fischer
Core-collapse supernovae undergoing a first-order quantum chromodynamics (QCD) phase transition experience the collapse of the central proto-neutron star that leads to a second bounce. This event is accompanied by the release of a second neutrino burst. Unlike the first stellar core bounce neutrino burst which consists exclusively of electron neutrinos, the
Yufan Dang, Chen Qian, Xueheng Luo, Jingru Fan
Large language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organizational structures that struggle to adapt as task complexity and
Christian Janos Lebeda, Mathieu Even, Aurélien Bellet, Julie Josse
Estimating causal effects from observational data is essential in fields such as medicine, economics and social sciences, where privacy concerns are paramount. We propose a general, model-agnostic framework for differentially private estimation of average treatment effects (ATE) that avoids strong structural assumptions on the data-generating process or the
Yanzhen Shen, Sihao Chen, Xueqiang Xu, Yunyi Zhang
While significant progress has been made with dual- and bi-encoder dense retrievers, they often struggle on queries with logical connectives, a use case that is often overlooked yet important in downstream applications. Current dense retrievers struggle with such queries, such that the retrieved results do not respect the logical constraints implied in the q
WQLCP: Weighted Adaptive Conformal Prediction for Robust Uncertainty Quantification Under Distribution Shifts
cs.LGShadi Alijani, Homayoun Najjaran
Conformal prediction (CP) provides a framework for constructing prediction sets with guaranteed coverage, assuming exchangeable data. However, real-world scenarios often involve distribution shifts that violate exchangeability, leading to unreliable coverage and inflated prediction sets. To address this challenge, we first introduce Reconstruction Loss-Scale
Dingyu Yao, Bowen Shen, Zheng Lin, Wei Liu
The Key-Value (KV) cache in generative large language models (LLMs) introduces substantial memory overhead. Existing works mitigate this burden by offloading or compressing the KV cache. However, loading the entire cache incurs significant latency due to PCIe bandwidth bottlenecks in CPU-GPU communication, while aggressive compression causes notable performa
Jiameng Li, Teodora Popordanoska, Aleksei Tiulpin, Sebastian G. Gruber
Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment planning. These biomarkers are typically estimated from segmentation outputs by computing region-wise ratios. Despite the high-stakes nature of clinical decision making, existing methods
The Bubble Wall Velocity in Local Thermal Equilibrium and Energy Budget with Full Effective Potential
hep-phZongguo Si, Hongxin Wang, Lei Wang, Yang Xiao
We develop a framework based on the full one-loop finite-temperature effective potential model, within which the bubble wall velocity is calculated using the local thermal equilibrium (LTE) approximation, and the kinetic energy fraction $K$ is computed directly. In cosmological phase transitions, these quantities play a critical role in determining the resul
Biju Saha, Suman Sarkar, Arunima Banerjee
About 30\% of disk galaxies show lopsidedness in their stellar disk. Although such a large-scale asymmetry in the disk can be primarily looked upon as a long-lived mode ($m=1$), the physical origin of the lopsidedness in the disk continues to be a puzzle. In this work, we employ a transfer-learning approach for the automated identification of lopsided galaxi
Kaiqing Lin, Zhiyuan Yan, Ke-Yue Zhang, Li Hao
Securing personal identity against deepfake attacks is increasingly critical in the digital age, especially for celebrities and political figures whose faces are easily accessible and frequently targeted. Most existing deepfake detection methods focus on general-purpose scenarios and often ignore the valuable prior knowledge of known facial identities, e.g.,
Ritesh K. Singh, Souradeep Sasmal, S. Nautiyal, A. K. Pan
We present a device-independent (DI) self-testing protocol in a constrained prepare-measure scenario, based on the $n-$bit parity-oblivious multiplexing (POM) task. In this scenario, a parity-oblivious constraint is imposed on the preparations, allowing us to define a classical bound derived from a preparation noncontextual ontological model. We derive the o
Whole-body Multi-contact Motion Control for Humanoid Robots Based on Distributed Tactile Sensors
cs.ROMasaki Murooka, Kensuke Fukumitsu, Marwan Hamze, Mitsuharu Morisawa
To enable humanoid robots to work robustly in confined environments, multi-contact motion that makes contacts not only at extremities, such as hands and feet, but also at intermediate areas of the limbs, such as knees and elbows, is essential. We develop a method to realize such whole-body multi-contact motion involving contacts at intermediate areas by a hu
Zhanpeng Cui, Bo Hou
We introduce the notion of quasi-triangular Novikov bialgebras, which constructed from solutions of the Novikov Yang-Baxter equation whose symmetric parts are invariant. Triangular Novikov bialgebras and factorizable Novikov bialgebras are important subclasses of quasi-triangular Novikov bialgebras. A factorizable Novikov bialgebra induces a factorization of
Dan Peng, Zhihui Fu, Zewen Ye, Zhuoran Song
Sparse attention methods exploit the inherent sparsity in attention to speed up the prefilling phase of long-context inference, mitigating the quadratic complexity of full attention computation. While existing sparse attention methods rely on predefined patterns or inaccurate estimations to approximate attention behavior, they often fail to fully capture the
Yeonjoon Jung, Daehyun Ahn, Hyungjun Kim, Taesu Kim
Low-Rank Adaptation (LoRA) is a popular method for parameter-efficient fine-tuning (PEFT) of generative models, valued for its simplicity and effectiveness. Despite recent enhancements, LoRA still suffers from a fundamental limitation: overfitting when the bottleneck is widened. It performs best at ranks 32-64, yet its accuracy stagnates or declines at highe
Yu Xi, Haoyu Li, Xiaoyu Gu, Yidi Jiang
Keyword spotting (KWS) is essential for voice-driven applications, demanding both accuracy and efficiency. Traditional ASR-based KWS methods, such as greedy and beam search, explore the entire search space without explicitly prioritizing keyword detection, often leading to suboptimal performance. In this paper, we propose an effective keyword-specific KWS fr
Yujie Yang, Bing Yang, Xiaofei Li
Online multichannel speech enhancement has been intensively studied recently. Though Mel-scale frequency is more matched with human auditory perception and computationally efficient than linear frequency, few works are implemented in a Mel-frequency domain. To this end, this work proposes a Mel-scale framework (namely Mel-McNet). It processes spectral and sp
Zero-Trust Foundation Models: A New Paradigm for Secure and Collaborative Artificial Intelligence for Internet of Things
cs.CRKai Li, Conggai Li, Xin Yuan, Shenghong Li
This paper focuses on Zero-Trust Foundation Models (ZTFMs), a novel paradigm that embeds zero-trust security principles into the lifecycle of foundation models (FMs) for Internet of Things (IoT) systems. By integrating core tenets, such as continuous verification, least privilege access (LPA), data confidentiality, and behavioral analytics into the design, t
Gilad Orr, Eliran Talker
Coherence time of thermal photons in rubidium vapor cells with varying thicknesses, reveal that there is clear dependence of the photon correlation time on cell thickness. Standard theoretical models accurately predict the coherence time in centimeter-scale cells. In this study we demonstrated, that these models break down in micrometer and sub-micrometer re
Alejandro Murillo-Gonzalez, Lantao Liu
Autonomous robots operating in complex, unstructured environments face significant challenges due to latent, unobserved factors that obscure their understanding of both their internal state and the external world. Addressing this challenge would enable robots to develop a more profound grasp of their operational context. To tackle this, we propose a novel fr
Vaishnavi Gupta, Hitesh Raundal
The paper investigates biorderability of knot quandles of prime knots up to eight crossings. We prove that knot quandles of knots $6_3$, $8_7$, $8_8$, $8_{10}$ and $8_{16}$ can not be biorderable. However, we see that knot quandles of knots $4_1$, $6_1$, $6_2$, $7_6$, $7_7$, $8_1$, $8_2$, $8_3$, $8_4$, $8_5$, $8_6$, $8_9$, $8_{11}$, $8_{12}$, $8_{13}$, $8_{1
Li Zeng, Zeming Liu, Chong Feng, Heyan Huang
Model editing aims to correct errors and outdated knowledge in the Large language models (LLMs) with minimal cost. Prior research has proposed a variety of datasets to assess the effectiveness of these model editing methods. However, most existing datasets only require models to output short phrases or sentences, overlooks the widespread existence of documen
Hu Xiaobin, Liang Yujie, Luo Donghao, Peng Xu
While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three key reasons:(1) Current metrics inadequately reflect human perception, particularly in unpaired try-on settings;(2)Most existing test sets are limited to indoor scenarios, lacking c
Modibo K. Camara, Nicole Immorlica, Brendan Lucier
In many settings -- like market research and social choice -- people may be presented with unfamiliar options. Classical mechanisms may perform poorly because they fail to incentivize people to learn about these options, or worse, encourage counterproductive information acquisition. We formalize this problem in a model of robust mechanism design where agents
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
cs.CVJianghang Lin, Yue Hu, Jiangtao Shen, Yunhang Shen
Open vocabulary image segmentation tackles the challenge of recognizing dynamically adjustable, predefined novel categories at inference time by leveraging vision-language alignment. However, existing paradigms typically perform class-agnostic region segmentation followed by category matching, which deviates from the human visual system's process of recogniz
E. C. I. Paterson, M. E. Tobar, M. Goryachev, J. Bourhill
We report the experimental observation of two distinct Berry phases ($+\frac{2\pi}{3}$ and $-\frac{2\pi}{3}$) generated on the surface of a M\"{o}bius cavity resonator at microwave frequencies supporting the TE$_{1,0,n}$ mode family. This resonator consists of a twisted, mirror-asymmetric prism with a cross-section of the triangular $D_3$ symmetry group, ben
Jiongchao Jin, Xiuju Fu, Xiaowei Gao, Tao Cheng
Maritime transportation is the backbone of global trade, making ship inspection essential for ensuring maritime safety and environmental protection. Port State Control (PSC), conducted by national ports, enforces compliance with safety regulations, with ship detention being the most severe consequence, impacting both ship schedules and company reputations. T
LLM-Agent-Controller: A Universal Multi-Agent Large Language Model System as a Control Engineer
cs.AIRasoul Zahedifar, Sayyed Ali Mirghasemi, Mahdieh Soleymani Baghshah, Alireza Taheri
This study presents the LLM-Agent-Controller, a multi-agent large language model (LLM) system developed to address a wide range of problems in control engineering (Control Theory). The system integrates a central controller agent with multiple specialized auxiliary agents, responsible for tasks such as controller design, model representation, control analysi
Integrated Finite Element Neural Network (IFENN) for Phase-Field Fracture with Minimal Input and Generalized Geometry-Load Handling
cs.CEPanos Pantidis, Lampros Svolos, Diab Abueidda, Mostafa E. Mobasher
We present a novel formulation for modeling phase-field fracture propagation based on the Integrated Finite Element Neural Network (IFENN) framework. IFENN is a hybrid solver scheme that utilizes neural networks as PDE solvers within FEM, preserving accuracy via residual minimization while achieving speed-up via swift network predictions and reduction of the
Juntong Wu, Zijing Liu, He Cao, Hao Li
In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequen
George Karantaidis, Athanasios Pantsios, Ioannis Kompatsiaris, Symeon Papadopoulos
Synthetic aperture radar automatic target recognition (SAR-ATR) systems have rapidly evolved to tackle incremental recognition challenges in operational settings. Data scarcity remains a major hurdle that conventional SAR-ATR techniques struggle to address. To cope with this challenge, we propose a few-shot class-incremental learning (FSCIL) framework based
Haofan Ren, Zunjie Zhu, Xiang Chen, Ming Lu
Neural fields are now the central focus of research in 3D vision and computer graphics. Existing methods mainly focus on various scene representations, such as neural points and 3D Gaussians. However, few works have studied the rendering process to enhance the neural fields. In this work, we propose a plug-in method named K-Buffers that leverages multiple bu
Shi-Yu Tian, Zhi Zhou, Wei Dong, Kun-Yang Yu
Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning over tabular data in real-world applications has been overlooked. For instance, applications such as business intelligence demand not only multi-step numerical reasoning with tabl
Ying Xiao, Jie Huang, Ruijuan He, Jing Xiao
Large language models (LLMs) are approaching expert-level performance in medical question answering (QA), demonstrating strong potential to improve public healthcare. However, underlying biases related to sensitive attributes such as sex and race pose life-critical risks. The extent to which such sensitive attributes affect diagnosis remains an open question
Yuan Feng, Yukun Cao, Hairu Wang, Xike Xie
Sketches, probabilistic structures for estimating item frequencies in infinite data streams with limited space, are widely used across various domains. Recent studies have shifted the focus from handcrafted sketches to neural sketches, leveraging memory-augmented neural networks (MANNs) to enhance the streaming compression capabilities and achieve better spa
LF-GNSS: Towards More Robust Satellite Positioning with a Hard Example Mining Enhanced Learning-Filtering Deep Fusion Framework
cs.ROJianan Lou, Rong Zhang
Global Navigation Satellite System (GNSS) is essential for autonomous driving systems, unmanned vehicles, and various location-based technologies, as it provides the precise geospatial information necessary for navigation and situational awareness. However, its performance is often degraded by Non-Line-Of-Sight (NLOS) and multipath effects, especially in urb
Vladimir Gol'dshtein, Reuven Segev
We outline here a simple mathematical introduction to the notions of multipoles for a general extensive property $\Pi$ from the point of view of continuum mechanics. Classically, $\Pi$ is the electric charge, but the theory is not limited to electrostatics. The proposed framework allows a simple computation of the bound "charges" and bound multipoles of lowe
Zhaowei Zhang, Xiaobo Wang, Minghua Yi, Mengmeng Wang
Achieving political consensus is crucial yet challenging for the effective functioning of social governance. However, although frontier AI systems represented by large language models (LLMs) have developed rapidly in recent years, their capabilities in this scope are still understudied. In this paper, we introduce PoliCon, a novel benchmark constructed from
Diogo Da Silva Machado
In this paper, we provide formulas for the sum of residues of type Camacho-Sad of a holomorphic foliation with respect to an invariant analytic subvariety. As application, in context of projective foliations, we obtain a formula that relates the sum these residues with the degree and other characteristics of the invariant subvariety. Furthermore, we establis
Shouqiao Wang, Davide Crapis, Ciamac C. Moallemi
This paper presents a comprehensive framework for transaction posting and pricing in Layer 2 (L2) blockchain systems, focusing on challenges stemming from fluctuating Layer 1 (L1) gas fees and the congestion issues within L2 networks. Existing methods have focused on the problem of optimal posting strategies to L1 in isolation, without simultaneously conside
Wei Su, Xi Zou
Modelling rarefied gas flow via the Boltzmann equation plays a vital role in many areas. Due to the high dimensionality of this kinetic equation and the coexistence of multiple characteristic scales in the transport processes, conventional solution strategies incur prohibitively high computational costs and are inadequate for rapid response for parametric an
Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation
cs.CVJiongchao Jin, Shengchu Zhao, Dajun Chen, Wei Jiang
Time consumption and the complexity of manual layout design make automated layout generation a critical task, especially for multiple applications across different mobile devices. Existing graph-based layout generation approaches suffer from limited generative capability, often resulting in unreasonable and incompatible outputs. Meanwhile, vision based gener
An electric circuital analysis of laboratory plasma sheath fluctuations and propagations
physics.plasm-phSubham Dutta, Pralay Kumar Karmakar
The effective inductive (L), capacitive (C), and resistive (R) behavior of a plasma sheath in a conjoint coupled form is well familiar among plasma physics communities. A dynamic sheath instability in laboratory plasmas is systematically modelled herein as an electrical series-resonance LCR circuit of the above kind. It theoretically yields experimentally ob
Minkyu Kim, Kiyoung Seong, Dongyeop Woo, Sungsoo Ahn
We address the challenge of training diffusion models to sample from unnormalized energy distributions in the absence of data, the so-called diffusion samplers. Although these approaches have shown promise, they struggle to scale in more demanding scenarios where energy evaluations are expensive and the sampling space is high-dimensional. To address this lim
Jochen L. Cremer
The electricity system becomes more complex, connecting massive numbers of end-users and distributed generators. Adding or removing grid connections requires expert studies to align technical constraints with user requests. In times of labour shortages, carrying out these studies represents a significant amount of time that engineers at system operators spen
Georgios Mappouras
With the rise of artificial intelligence (A.I.) and large language models like ChatGPT, a new race for achieving artificial general intelligence (A.G.I) has started. While many speculate how and when A.I. will achieve A.G.I., there is no clear agreement on how A.G.I. can be detected in A.I. models, even when popular tools like the Turing test (and its modern
From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
cs.CLDerong Xu, Yi Wen, Pengyue Jia, Yingyi Zhang
Large Language Models (LLMs) have recently been widely adopted in conversational agents. However, the increasingly long interactions between users and agents accumulate extensive dialogue records, making it difficult for LLMs with limited context windows to maintain a coherent long-term dialogue memory and deliver personalized responses. While retrieval-augm
Xufeng Duan, Zhaoqian Yao, Yunhao Zhang, Shaonan Wang
Large language models (LLMs) have been found to develop surprising internal specializations: Individual neurons, attention heads, and circuits become selectively sensitive to syntactic structure, reflecting patterns observed in the human brain. While this specialization is well-documented, how it emerges during training and what influences its development re
Haoyu Zhang, Wentao Zhang, Hao Miao, Xinke Jiang
Spatio-Temporal Graph Neural Networks (STGNNs) have emerged as a powerful tool for modeling dynamic graph-structured data across diverse domains. However, they often fail to generalize in Spatio-Temporal Out-of-Distribution (STOOD) scenarios, where both temporal dynamics and spatial structures evolve beyond the training distribution. To address this problem,
Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Mehrdad Noori
Test-Time Training (TTT) has emerged as a promising solution to address distribution shifts in 3D point cloud classification. However, existing methods often rely on computationally expensive backpropagation during adaptation, limiting their applicability in real-world, time-sensitive scenarios. In this paper, we introduce SMART-PC, a skeleton-based framewor
Guy F. de Teramond, Arpon Paul, Hans Gunter Dosch, Stanley J. Brodsky
We extend our recent analytic study of the strong coupling $\alpha_{\rm eff}$ in the nonperturbative and near-perturbative regimes~\cite{deTeramond:2024ikl} by imposing rigorous renormalization-group results from asymptotically free gauge theories at $Q^2 \to \infty$. The asymptotic boundary conditions modify the scaling properties of $\alpha_{\rm eff}$ at l
Unlocking the Power of Diffusion Models in Sequential Recommendation: A Simple and Effective Approach
cs.IRJialei Chen, Yuanbo Xu, Yiheng Jiang
In this paper, we focus on the often-overlooked issue of embedding collapse in existing diffusion-based sequential recommendation models and propose ADRec, an innovative framework designed to mitigate this problem. Diverging from previous diffusion-based methods, ADRec applies an independent noise process to each token and performs diffusion across the entir
Cuff-KT: Tackling Learners' Real-time Learning Pattern Adjustment via Tuning-Free Knowledge State Guided Model Updating
cs.LGYiyun Zhou, Zheqi Lv, Shengyu Zhang, Jingyuan Chen
Knowledge Tracing (KT) is a core component of Intelligent Tutoring Systems, modeling learners' knowledge state to predict future performance and provide personalized learning support. Traditional KT models assume that learners' learning abilities remain relatively stable over short periods or change in predictable ways based on prior performance. However, in
Anwesh Ray
Motivated by the work of Greenberg-Vatsal and Emerton-Pollack-Weston, I investigate the extent to which Mazur's conjecture on the growth of Selmer ranks in $\mathbb{Z}_p$-extensions of an imaginary quadratic field persists under $p$-congruences between Galois representations. As a first step, I establish Mazur's conjecture for certain triples $(E, K, p)$ und
Chen Jiang, Haidong Liu
Let $X$ be a $\mathbb Q$-factorial weak Fano $3$-fold with at worst isolated canonical singularities. We show that the $\mathbb Q$-Fano index of $X$ is at most $61$.
Real-time Whole-body Model Predictive Control for Bipedal Locomotion with a Novel Kino-dynamic Model and Warm-start Method
cs.ROJunhyung Kim, Hokyun Lee, Jaeheung Park
Advancements in optimization solvers and computing power have led to growing interest in applying whole-body model predictive control (WB-MPC) to bipedal robots. However, the high degrees of freedom and inherent model complexity of bipedal robots pose significant challenges in achieving fast and stable control cycles for real-time performance. This paper int
Dong Liu, Yanxuan Yu, Jiayi Zhang, Yifan Li
Diffusion Transformers (DiT) are powerful generative models but remain computationally intensive due to their iterative structure and deep transformer stacks. To alleviate this inefficiency, we propose \textbf{FastCache}, a hidden-state-level caching and compression framework that accelerates DiT inference by exploiting redundancy within the model's internal
Zhongqin Wang, J. Andrew Zhang, Kai Wu, Y. Jay Guo
Accurate water level sensing is essential for flood monitoring, agricultural irrigation, and water resource optimization. Traditional methods require dedicated sensor deployments, leading to high installation costs, vulnerability to interference, and limited resolution. This work proposes PMNs-WaterSense, a novel scheme leveraging Channel State Information (
Yuxing Lu, Gecheng Fu, Wei Wu, Xukai Zhao
Existing medical RAG systems mainly leverage knowledge from medical knowledge bases, neglecting the crucial role of experiential knowledge derived from similar patient cases -- a key component of human clinical reasoning. To bridge this gap, we propose DoctorRAG, a RAG framework that emulates doctor-like reasoning by integrating both explicit clinical knowle
Yi Feng, Kaito Fujii, Stratis Skoulakis, Xiao Wang
Since Polyak's pioneering work, heavy ball (HB) momentum has been widely studied in minimization. However, its role in min-max games remains largely unexplored. As a key component of practical min-max algorithms like Adam, this gap limits their effectiveness. In this paper, we present a continuous-time analysis for HB with simultaneous and alternating update
Jintao Tong, Wenwei Jin, Pengda Qin, Anqi Li
Large vision-language models (LVLMs) excel at multimodal understanding but suffer from high computational costs due to redundant vision tokens. Existing pruning methods typically rely on single-layer attention scores to rank and prune redundant visual tokens to solve this inefficiency. However, as the interaction between tokens and layers is complicated, thi
Juntong Wang, Jiarui Wang, Huiyu Duan, Guangtao Zhai
Text-driven video editing is rapidly advancing, yet its rigorous evaluation remains challenging due to the absence of dedicated video quality assessment (VQA) models capable of discerning the nuances of editing quality. To address this critical gap, we introduce TDVE-DB, a large-scale benchmark dataset for text-driven video editing. TDVE-DB consists of 3,857
Yongyi Zang, Jingyi Li, Qiuqiang Kong
Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this work, we reveal that pretrained one-step audio source separation models can be leveraged for multi-step separation without additional traini
Yachuan Liu, Xiaochun Wei, Lin Shi, Xinnuo Li
Large language models (LLMs) face significant challenges in ex-ante reasoning, where analysis, inference, or predictions must be made without access to information from future events. Even with explicit prompts enforcing temporal cutoffs, LLMs often generate outputs influenced by internalized knowledge of events beyond the specified cutoff. This paper introd
Jerry Yao-Chieh Hu, Xiwen Zhang, Maojiang Su, Zhao Song
We study the computational limits of learning $k$-bit Boolean functions (specifically, $\mathrm{AND}$, $\mathrm{OR}$, and their noisy variants), using a minimalist single-head softmax-attention mechanism, where $k=\Theta(d)$ relevant bits are selected from $d$ inputs. We show that these simple $\mathrm{AND}$ and $\mathrm{OR}$ functions are unsolvable with a
Amartya Purushottam, Jack Yan, Christopher Yu, Joao Ramos
Humanoid robots can support human workers in physically demanding environments by performing tasks that require whole-body coordination, such as lifting and transporting heavy objects.These tasks, which we refer to as Dynamic Mobile Manipulation (DMM), require the simultaneous control of locomotion, manipulation, and posture under dynamic interaction forces.
Tanjil Hasan Sakib, Md. Tanzib Hosain, Md. Kishor Morol
Small Language Models (SLMs) have gained substantial attention due to their ability to execute diverse language tasks successfully while using fewer computer resources. These models are particularly ideal for deployment in limited environments, such as mobile devices, on-device processing, and edge systems. In this study, we present a complete assessment of
Yejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub Han
Implicit hate speech detection is challenging due to its subtlety and reliance on contextual interpretation rather than explicit offensive words. Current approaches rely on contrastive learning, which are shown to be effective on distinguishing hate and non-hate sentences. Humans, however, detect implicit hate speech by first identifying specific targets wit
Mohammed Djameleddine Belgoumri, Mohamed Reda Bouadjenek, Hakim Hacid, Imran Razzak
Training large neural networks (NNs) requires optimizing high-dimensional data-dependent loss functions. The optimization landscape of these functions is often highly complex and textured, even fractal-like, with many spurious local minima, ill-conditioned valleys, degenerate points, and saddle points. Complicating things further is the fact that these lands
Robert Fraser, Kyle Hambrook, Donggeun Ryou
We prove the optimality of the exponent in the Mockenhaupt-Mitsis-Bak-Seeger Fourier restriction theorem in all dimensions $d$ and the full parameter range $0 < a,b < d$. Our construction is deterministic and also yields Salem sets.
Jie Wu
A graph $G$ has the $k$-strong parity property if for any $X\subseteq V(G)$ with $|X|$ even, $G$ contains a spanning subgraph $F$ with $d_F(u)\equiv1$ (mod 2) for each $u\in X$ and $d_F(v)\in\{k,k+2,k+4,\ldots\}$ for each $v\in V(G)\setminus X$, where $k\geq2$ is an even integer. Kano and Matsumura proposed a characterization for a graph with the $k$-strong
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
cs.LGLiangwei Nathan Zheng, Wei Emma Zhang, Mingyu Guo, Olaf Maennel
Effectively managing missing modalities is a fundamental challenge in real-world multimodal learning scenarios, where data incompleteness often results from systematic collection errors or sensor failures. Sparse Mixture-of-Experts (SMoE) architectures have the potential to naturally handle multimodal data, with individual experts specializing in different m
Jun Lai, Jinrui Zhang
Inverse scattering in layered media has a wide range of applications, examples including geophysical exploration, medical imaging, and remote sensing. In this paper, we develop a selective focusing method for identifying multiple unknown buried scatterers in a layered medium. The method is derived through the asymptotic analysis of the time reversal operator
Cen Liu, Xiangyun Zhou, Nan Yang, Salman Durrani
This work studies near-field secure communications through transmit beamfocusing. We examine the benefit of having a protected eavesdropper-free zone around the legitimate receiver, and we determine the worst-case secrecy performance against a potential eavesdropper located anywhere outside the protected zone. A max-min optimization problem is formulated for
Applications and Effect Evaluation of Generative Adversarial Networks in Semi-Supervised Learning
cs.CVJiyu Hu, Haijiang Zeng, Zhen Tian
In recent years, image classification, as a core task in computer vision, relies on high-quality labelled data, which restricts the wide application of deep learning models in practical scenarios. To alleviate the problem of insufficient labelled samples, semi-supervised learning has gradually become a research hotspot. In this paper, we construct a semi-sup
Dongzhe Zheng, Wenjie Mei
Learning unknown dynamics under environmental (or external) constraints is fundamental to many fields (e.g., modern robotics), particularly challenging when constraint information is only locally available and uncertain. Existing approaches requiring global constraints or using probabilistic filtering fail to fully exploit the geometric structure inherent in
Guanhao Li
This paper provides a combinatorial proof to show that, in the study of maximal Condorcet domains, the class of peak-pit Condorcet domains, the class of connected Condorcet domains, and the class of directly connected Condorcet domains are all equivalent.
Enhancing Contrastive Learning-based Electrocardiogram Pretrained Model with Patient Memory Queue
eess.SPXiaoyu Sun, Yang Yang, Xunde Dong
In the field of automatic Electrocardiogram (ECG) diagnosis, due to the relatively limited amount of labeled data, how to build a robust ECG pretrained model based on unlabeled data is a key area of focus for researchers. Recent advancements in contrastive learning-based ECG pretrained models highlight the potential of exploiting the additional patient-level
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
cs.CVGihoon Kim, Hyungjin Park, Taesup Kim
Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the model's original generative capabilities. However, this process often leads to overfitting, where the model ignores the user's prompt and merely replicates the reference images. We attribute this issue to a fundamen
Toward Patient-specific Partial Point Cloud to Surface Completion for Pre- to Intra-operative Registration in Image-guided Liver Interventions
cs.CVNakul Poudel, Zixin Yang, Kelly Merrell, Richard Simon
Intra-operative data captured during image-guided surgery lacks sub-surface information, where key regions of interest, such as vessels and tumors, reside. Image-to-physical registration enables the fusion of pre-operative information and intra-operative data, typically represented as a point cloud. However, this registration process struggles due to partial
Pieter van Goor, Robert Mahony
This paper introduces the concept of a synchronous model as an extension of the internal model concept used in observer design for dynamical systems. A system is said to contain a synchronous model of another if there is a suitable error function between the two systems that remains stationary for all of the trajectories of the two systems. A system is said
Rui Zhao, Yuze Fan, Ziguo Chen, Fei Gao
End-to-end learning has emerged as a transformative paradigm in autonomous driving. However, the inherently multimodal nature of driving behaviors and the generalization challenges in long-tail scenarios remain critical obstacles to robust deployment. We propose DiffE2E, a diffusion-based end-to-end autonomous driving framework. This framework first performs
Analyzing Biases in Political Dialogue: Tagging U.S. Presidential Debates with an Extended DAMSL Framework
cs.CLLavanya Prahallad, Radhika Mamidi
We present a critical discourse analysis of the 2024 U.S. presidential debates, examining Donald Trump's rhetorical strategies in his interactions with Joe Biden and Kamala Harris. We introduce a novel annotation framework, BEADS (Bias Enriched Annotation for Dialogue Structure), which systematically extends the DAMSL framework to capture bias driven and adv