Skip to content

May 2025 arXiv papers — page 24

Showing 2,3012,400 of 24,552 papers

  1. Hongcan Guo, Guoshun Nan, Yuan Yang, Diyang Zhang

    Scaling Low-Rank Adaptation (LoRA)-based Mixture-of-Experts (MoE) facilitates large language models (LLMs) to efficiently adapt to diverse tasks. However, traditional gating mechanisms that route inputs to the best experts may fundamentally hinder LLMs' scalability, leading to poor generalization and underfitting issues. We identify that the root cause lies

  2. Gabriele Sarti, Vilém Zouhar, Malvina Nissim, Arianna Bisazza

    Word-level quality estimation (WQE) aims to automatically identify fine-grained error spans in machine-translated outputs and has found many uses, including assisting translators during post-editing. Modern WQE techniques are often expensive, involving prompting of large language models or ad-hoc training on large amounts of human-labeled data. In this work,

  3. Srijith Nair, Michael Lin, Peizhong Ju, Amirreza Talebi

    Collaborative training methods like Federated Learning (FL) and Split Learning (SL) enable distributed machine learning without sharing raw data. However, FL assumes clients can train entire models, which is infeasible for large-scale models. In contrast, while SL alleviates the client memory constraint in FL by offloading most training to the server, it inc

  4. Tian Tian, Chunyan Miao, Hangwei Qian

    Contrastive learning has emerged as a competent approach for unsupervised representation learning. However, the design of an optimal augmentation strategy, although crucial for contrastive learning, is less explored for time series classification tasks. Existing predefined time-domain augmentation methods are primarily adopted from vision and are not specifi

  5. Ping Wang, Lishun Wang, Gang Qu, Xiaodong Wang

    Deep-unrolling and plug-and-play (PnP) approaches have become the de-facto standard solvers for single-pixel imaging (SPI) inverse problem. PnP approaches, a class of iterative algorithms where regularization is implicitly performed by an off-the-shelf deep denoiser, are flexible for varying compression ratios (CRs) but are limited in reconstruction accuracy

  6. Sungjune Park, Hyunjun Kim, Junho Kim, Seongho Kim

    MLLMs have demonstrated significant visual understanding capabilities, yet their fine-grained visual perception in complex real-world scenarios, such as densely crowded public areas, remains limited. Inspired by the recent success of RL in both LLMs and MLLMs, in this paper, we explore how RL can enhance visual perception ability of MLLMs. Then we develop a

  7. Tonglin Liao, Youming Li

    In this paper, we consider discrete-time D-BMAP/G/\inf queueing model. We construct effective discrete-time Markovian dynamics for this model and utilize it to derive exact time-dependent distribution of customer number and the corresponding moments for the original queueing model. Numerical simulations are used to verify our results. Using our result, we pr

  8. Wenjing Xing, Wenke Lu, Yeheng Duan, Bing Zhao

    Traditional code instruction data synthesis methods suffer from limited diversity and poor logic. We introduce Infinite-Instruct, an automated framework for synthesizing high-quality question-answer pairs, designed to enhance the code generation capabilities of large language models (LLMs). The framework focuses on improving the internal logic of synthesized

  9. Shiwei Li, Xiandi Luo, Haozhao Wang, Xing Tang

    To improve the training efficiency of federated learning (FL), previous research has employed low-rank decomposition techniques to reduce communication overhead. In this paper, we seek to enhance the performance of these low-rank decomposition methods. Specifically, we focus on three key issues related to decomposition in FL: what to decompose, how to decomp

  10. Changyi Lin, Yuxin Ray Song, Boda Huo, Mingyang Yu

    Quadrupedal robots have demonstrated remarkable agility and robustness in traversing complex terrains. However, they struggle with dynamic object interactions, where contact must be precisely sensed and controlled. To bridge this gap, we present LocoTouch, a system that equips quadrupedal robots with tactile sensing to address a particularly challenging task

  11. Naman Ahuja, Fenil Bardoliya, Chitta Baral, Vivek Gupta

    Transforming dense, detailed, unstructured text into an interpretable and summarised table, also colloquially known as Text-to-Table generation, is an essential task for information retrieval. Current methods, however, miss out on how and what complex information to extract; they also lack the ability to infer data from the text. In this paper, we introduce

  12. Shohei Enomoto

    Deep learning models often struggle to maintain performance when deployed on data distributions different from their training data, particularly in real-world applications where environmental conditions frequently change. While Multi-source Domain Generalization (MDG) has shown promise in addressing this challenge by leveraging multiple source domains during

  13. Khac-Hoang Ngo, Diego Cuevas, Ruben de Miguel Gil, Victor Monzon Baeza

    Noncoherent communication is a promising paradigm for future wireless systems where acquiring accurate channel state information (CSI) is challenging or infeasible. It provides methods to bypass the need for explicit channel estimation in practical scenarios such as high-mobility networks, massive distributed antenna arrays, energy-constrained Internet-of-Th

  14. Liu Liu, Xiaofeng Wang, Guosheng Zhao, Keyu Li

    The goal of general-purpose robotics is to create agents that can seamlessly adapt to and operate in diverse, unstructured human environments. Imitation learning has become a key paradigm for robotic manipulation, yet collecting large-scale and diverse demonstrations is prohibitively expensive. Simulators provide a cost-effective alternative, but the sim-to-

  15. Jian Zhu, Farhan Samir, Eleanor Chodroff, David R. Mortensen

    We present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. We first curated IPAPack++, a large-scale multilingual speech corpus with 17,132 hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. With the large-sc

  16. Yusuke Ishigaki, Takayuki Kobayashi

    We study the large time behavior of solutions to the system of equations describing motion of compressible viscoelastic fluids. We focus on the linearized system around a motionless state in a three-dimensional exterior domain and derive the local energy decay estimate of its solution to give the diffusion wave phenomena caused by sound wave viscous diffusio

  17. Kengo Matsumoto, Taro Sogabe

    Combining the theory of extensions of C*-algebras and the Pimsner construction, we show that every countable infinite discrete group admits an ergodic action on arbitrary unital Kirchberg algebra. In the proof, we give a Pimsner construction realizing many unital subalgebras of a given unital Kirchberg algebra as the fixed point algebras of single automorphi

  18. Valentina Bais, Rafael Torres

    We introduce a simple cut-and-paste mechanism to construct both orientable and nonorientable four-manifolds from a given initial one. This mechanism alters the fundamental group while preserving other essential topological invariants. It avoids codimension two cut-and-paste fundamental group computations and fast tracks the search for fixed-point free involu

  19. Li Lucy, Camilla Griffiths, Sarah Levine, Jennifer L. Eberhardt

    Conventional bag-of-words approaches for topic modeling, like latent Dirichlet allocation (LDA), struggle with literary text. Literature challenges lexical methods because narrative language focuses on immersive sensory details instead of abstractive description or exposition: writers are advised to "show, don't tell." We propose Retell, a simple, accessible

  20. Le Yang, Vincent Y. F. Tan, Wang Chi Cheung

    We study the best arm identification (BAI) problem with potentially biased offline data in the fixed confidence setting, which commonly arises in real-world scenarios such as clinical trials. We prove an impossibility result for adaptive algorithms without prior knowledge of the bias bound between online and offline distributions. To address this, we propose

  21. Jonathan Husson, Guido Mazzuca, Alessandra Occelli

    In this paper, we study the asymptotic behaviour of plane partitions distributed according to a $q^{\text{Volume}}$-weighted Muttalib--Borodin ensemble and its associated discrete point process. We establish a Large Deviation Principle for the process, explicitly characterizing the rate function. A defining feature of our model is the emergence of a strict u

  22. Antonio D'Orazio, Maria Rosaria Briglia, Donato Crisostomi, Dario Loi

    CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space back to images. In this work, we show that image synthesis is nevertheless possible using CLIP alone -- without any decode

  23. Lorenzo Marinucci, Claudio Battiloro, Paolo Di Lorenzo

    This paper introduces a novel adaptive framework for processing dynamic flow signals over simplicial complexes, extending classical least-mean-squares (LMS) methods to high-order topological domains. Building on discrete Hodge theory, we present a topological LMS algorithm that efficiently processes streaming signals observed over time-varying edge subsets.

  24. Jonas Kulhanek, Marie-Julie Rakotosaona, Fabian Manhardt, Christina Tsalicoglou

    In this work, we present a novel level-of-detail (LOD) method for 3D Gaussian Splatting that enables real-time rendering of large-scale scenes on memory-constrained devices. Our approach introduces a hierarchical LOD representation that iteratively selects optimal subsets of Gaussians based on camera distance, thus largely reducing both rendering time and GP

  25. Ming Hsiao

    We establish a short-time existence theory for complete Ricci flows under scaling-invariant curvature bounds, starting from rotationally symmetric metrics on $\mathbb{R}^{n+1}$ that are noncollapsed at infinity, without assuming bounded curvature. As a consequence, we construct a complete Ricci flow solution coming out of a rotationally symmetric metric, whi

  26. Torsten Röper, Daniel Rosenbach, Achim Rosch, Alexey A. Taskin

    Quantum anomalous Hall (QAH) insulators exhibit chiral dissipationless edge states without an external magnetic field, making them a promising material for quantum metrology and microwave applications. However, the breakdown of the zero-resistance state at low currents hinders progress. We investigate and characterize this breakdown under microwave fields (1

  27. Xiao Yu, Yan Fang, Xiaojie Jin, Yao Zhao

    Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge model sizes, limiting their real-time applicability. We introduce Online Audio-Visual Event Parsing (On-AVEP), a novel paradigm for parsing audio, visual, and audio-visual events by se

  28. Fan Wang, Shaoshan Liu

    Collective Adaptive Intelligence (CAI) represent a transformative approach in embodied AI, wherein numerous autonomous agents collaborate, adapt, and self-organize to navigate complex, dynamic environments. By enabling systems to reconfigure themselves in response to unforeseen challenges, CAI facilitate robust performance in real-world scenarios. This artic

  29. Donghwa Kim, Jaewook Lee, Chulhee Yun

    We analyze the convergence rates of two popular variants of coordinate descent (CD): random CD (RCD), in which the coordinates are sampled uniformly at random, and random-permutation CD (RPCD), in which random permutations are used to select the update indices. Despite abundant empirical evidence that RPCD outperforms RCD in various tasks, the theoretical ga

  30. Qian-Yu An, Yang Huang, Wei-Min Gu, Yong Shao

    Binary systems consisting of an early type star and a black hole (BH) are crucial for understanding various astrophysical phenomena, particularly the origins of detected gravitational wave sources. Be binary systems are expected to represent a key evolutionary stage in hosting BHs. However, while hundreds of Be X-ray binaries are known, the only confirmed BH

  31. Michal Nauman, Marek Cygan, Carmelo Sferrazza, Aviral Kumar

    Recent advances in language modeling and vision stem from training large models on diverse, multi-task data. This paradigm has had limited impact in value-based reinforcement learning (RL), where improvements are often driven by small models trained in a single-task context. This is because in multi-task RL sparse rewards and gradient conflicts make optimiza

  32. Dragos-Patru Covei

    We extend the stochastic production planning framework to manufacturing systems, where the set of admissible production configurations is described by a general smooth convex domain $\omega $. In our setting, production operations continue as long as the production inventory $y(t)$ remains inside the capacity limits of $\omega $ and are halted once the state

  33. Minghao Xie, Sheng Jin, Dong-Hong Wu

    The observed exoplanet population exhibits a scarcity of short-period Saturn-mass planets, a phenomenon referred to as the ``hot Saturn desert". This observational scarcity can be utilized to validate the theories regarding the formation and evolution of gas planets. In this study, we conduct large-scale numerical simulations to explore how the initial condi

  34. Anke Fischer-Janzen, Thomas M. Wendt, Daniel Görlich, Kristof Van Laerhoven

    Advances in eye-tracking control for assistive robotic arms provide intuitive interaction opportunities for people with physical disabilities. Shared control has gained interest in recent years by improving user satisfaction through partial automation of robot control. We present an eye-tracking-guided shared control design based on insights from state-of-th

  35. Qiuyu Ding, Zhiqiang Cao, Hailong Cao, Tiejun Zhao

    Bilingual Lexicon Induction (BLI) is generally based on common domain data to obtain monolingual word embedding, and by aligning the monolingual word embeddings to obtain the cross-lingual embeddings which are used to get the word translation pairs. In this paper, we propose a new task of BLI, which is to use the monolingual corpus of the general domain and

  36. Youshen Xiao, Yiling Shi, Ruixi Sun, Hongjiang Wei

    Dynamic Photoacoustic Computed Tomography (PACT) is an important imaging technique for monitoring physiological processes, capable of providing high-contrast images of optical absorption at much greater depths than traditional optical imaging methods. However, practical instrumentation and geometric constraints limit the number of acoustic sensors available

  37. Jeongsol Kim, Yeobin Hong, Jonghyun Park, Jong Chul Ye

    Recent inversion-free, flow-based image editing methods such as FlowEdit leverages a pre-trained noise-to-image flow model such as Stable Diffusion 3, enabling text-driven manipulation by solving an ordinary differential equation (ODE). While the lack of exact latent inversion is a core advantage of these methods, it often results in unstable editing traject

  38. Yangyang Cao, Qian Huang, Julian Koellermeier, Alexander Kurganov

    We develop second-order path-conservative central-upwind (PCCU) schemes for the hyperbolic shallow water linearized moment equations (HSWLME), which are an extension of standard depth-averaged models for free-surface flows. The proposed PCCU schemes are constructed via flux globalization strategies adapted to the nonconservative form via a path-conservative

  39. Jinquan Guan, Qi Chen, Lizhou Liang, Yuhang Liu

    Artificial intelligence (AI)-based chest X-ray (CXR) interpretation assistants have demonstrated significant progress and are increasingly being applied in clinical settings. However, contemporary medical AI models often adhere to a simplistic input-to-output paradigm, directly processing an image and an instruction to generate a result, where the instructio

  40. Ian Langmore

    Positive semi-definite kernels are used to induce pseudo-metrics, or ``distances'', between measures. We write these as an expected quadratic variation of, or expected inner product between, a random field and the difference of measures. This alternate viewpoint offers important intuition and interesting connections to existing forms. Metric distances leadin

  41. Qiuyu Ding, Zhiqiang Cao, Hailong Cao, Tiejun Zhao

    Large language models have demonstrated exceptional performance across multiple crosslingual NLP tasks, including machine translation (MT). However, persistent challenges remain in addressing context-sensitive units (CSUs), such as polysemous words. These CSUs not only affect the local translation accuracy of LLMs, but also affect LLMs' understanding capabil

  42. Yosuke Kawamato, Genki Shibukawa

    The aim of this paper is to study intertwining relations for Laguerre process with inverse temperature $\beta \ge 1$ and parameter $\alpha >-1$. We introduce a Markov kernel that depends on both $\beta $ and $ \alpha $, and establish new intertwining relations for the $\beta$-Laguerre processes using this kernel. A key observation is that Jack symmetric poly

  43. Kosei Tsuji, Ichiro Maruta, Kenji Fujimoto, Tomoyuki Maeda

    Nonlinear Model Predictive Control (NMPC) offers a powerful approach for controlling complex nonlinear systems, yet faces two key challenges. First, accurately modeling nonlinear dynamics remains difficult. Second, variables directly related to control objectives often cannot be directly measured during operation. Although high-cost sensors can acquire these

  44. Yuval Samoilov-Kats, Matan Noach, Noam Beer, Yuval Efrati

    In recent years, there is a growing need and opportunity to use online platforms for psychophysics research. Online experiments make it possible to evaluate large and diverse populations remotely and quickly, complementing laboratory-based research. However, developing and running online psychophysics experiments poses several challenges: i) a high barrier-t

  45. Xuwen Chen, Jiahao Wu, Zhifei Zhang

    We consider a dilute Fermi gas in the thermodynamic limit with interaction potential scattering length $\mathfrak{a}_0$ at temperature $T>0$. We prove the 2nd order Huang-Yang approximation for the Fermi pressure of the system, in which there is a 2nd order term carrying the positive temperature efffect.Our formula is valid up to the temperature $T<\rho^{\fr

  46. Wuhao Wang, Zhiyong Chen

    Partially Observable Markov Decision Processes (POMDPs) remain a core challenge in reinforcement learning due to incomplete state information. We address this by reformulating POMDPs as fully observable processes with fixed-length observation histories as augmented states. To efficiently encode these histories, we propose a lightweight temporal encoder based

  47. Zhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel

    Large language models (LLMs) are increasingly integrated in software development, but ensuring correctness in LLM-generated code remains challenging and often requires costly manual review. Verifiable code generation -- jointly generating code, specifications, and proofs of code-specification alignment -- offers a promising path to address this limitation an

  48. Tongtong Su, Chengyu Wang, Jun Huang, Dongming Lu

    Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspects of objects. To overcome these limitations, this paper introduces a novel approach named {Zero-to-Hero}, which focuses on reference-based

  49. Shi Heng Zhang, Zhengjie Miao, Jiannan Wang

    As enterprise data grows in size and complexity, column-level data lineage, which records the creation, transformation, and reference of each column in the warehouse, has been the key to effective data governance that assists tasks like data quality monitoring, storage refactoring, and workflow migration. Unfortunately, existing systems introduce overheads b

  50. Seung Gyu Jeong, Seong Eun Kim

    Auscultation is crucial for diagnosing lung diseases. The COVID-19 pandemic has revealed the limitations of traditional, in-person lung sound assessments. To overcome these issues, advancements in digital stethoscopes and artificial intelligence (AI) have led to the development of new diagnostic methods. In this context, our study aims to use smartphone micr

  51. Haoyu Chen, Keda Tao, Yizao Wang, Xinlei Wang

    Photo retouching is integral to photographic art, extending far beyond simple technical fixes to heighten emotional expression and narrative depth. While artists leverage expertise to create unique visual effects through deliberate adjustments, non-professional users often rely on automated tools that produce visually pleasing results but lack interpretative

  52. Bin Wang, Pingjun Li, Jinkun Liu, Jun Cheng

    End-to-end autonomous driving faces persistent challenges in both generating diverse, rule-compliant trajectories and robustly selecting the optimal path from these options via learned, multi-faceted evaluation. To address these challenges, we introduce HMAD, a framework integrating a distinctive Bird's-Eye-View (BEV) based trajectory proposal mechanism with

  53. Shengjin Ji, Balázs Patkós, Erfei Yue

    A family $\mathcal{G}$ of sets is a(n induced) copy of a poset $P=(P,\leqslant)$ if there exists a bijection $b:P\rightarrow \mathcal{G}$ such that $p\leqslant q$ holds if and only if $b(p)\subseteq b(q)$. The induced saturation number sat$^*(n,P)$ is the minimum size of a family $\mathcal{F}\subseteq 2^{[n]}$ that does not contain any copy of $P$, but for a

  54. Raúl Hidalgo-Sacoto, Thomas Busch, D. Blume

    While elementary particles obey either bosonic or fermionic exchange statistics, generalized exchange statistics that interpolate between bosons and fermions -- applicable to quasi-particles -- constitute an intriguing topic, both from the fundamental and practical points of view. This work develops a scattering framework for two identical 1D bosonic anyons

  55. Atharva Naik, Prakam, Yash Mathur, Darsh Agrawal

    Although many benchmarks evaluate the reasoning abilities of Large Language Models (LLMs) within domains such as mathematics, coding, or data wrangling, few abstract away from domain specifics to examine reasoning as a capability in and of itself. We contribute a novel type of benchmark evaluating the inductive reasoning capabilities of LLMs that is inspired

  56. Romeo Ortega, Alexey Bobtsov, Leyan Fang, Oscar Texis-Loaiza

    In this paper we address the problem of online detection of inter-turn short-circuit faults (ITSCFs) that occur in permanent magnet synchronous motors (PMSMs). We propose two solutions to this problem: (i) a very simple linear observer and (ii) a generalized parameter estimation based observer, that incorporates a high performance estimator -- with both obse

  57. Junyan Liu, Arnab Maiti, Artin Tajdini, Kevin Jamieson

    We initiate the study of a repeated principal-agent problem over a finite horizon $T$, where a principal sequentially interacts with $K\geq 2$ types of agents arriving in an adversarial order. At each round, the principal strategically chooses one of the $N$ arms to incentivize for an arriving agent of unknown type. The agent then chooses an arm based on its

  58. Ruilin Xu, Yuchen Song, Kaijie Li, Xitong Gao

    Offline map matching involves aligning historical trajectories of mobile objects, which may have positional errors, with digital maps. This is essential for applications in intelligent transportation systems (ITS), such as route analysis and traffic pattern mining. Existing methods have two main limitations: (i) they assume a uniform Localization Error Distr

  59. Vladimir Belavin, Juan Ramos Cabezas, Boris Runov

    We study the construction of correlation numbers in super minimal Liouville gravity. In particular, we construct the fundamental physical fields in the Ramond sector and compute the three-point correlation number involving two physical fields in the Ramond sector and one in the NS sector. Furthermore, we establish the relation between Ramond physical fields

  60. Yiming Lei, Zhizheng Yang, Zeming Liu, Haitao Leng

    Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQForme

  61. Siyuan Wang, Jiawei Liu, Wei Wang, Yeying Jin

    Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying solely on audio as the control signal often fails to capture large gesture movements in videos, resulting in more noticeab

  62. Keren Ye, Ignacio Garcia Dorado, Michalis Raptis, Mauricio Delbracio

    While recent advancements in Image Super-Resolution (SR) using diffusion models have shown promise in improving overall image quality, their application to scene text images has revealed limitations. These models often struggle with accurate text region localization and fail to effectively model image and multilingual character-to-shape priors. This leads to

  63. Zhongzhen Huang, Linjie Mu, Yakun Zhu, Xiangyu Zhao

    Effective clinical decision-making depends on iterative, multimodal reasoning across diverse sources of evidence. The recent emergence of multimodal reasoning models has significantly transformed the landscape of solving complex tasks. Although such models have achieved notable success in mathematics and science, their application to medical domains remains

  64. Yuatyong Chaichana, Thanapat Trachu, Peerat Limkonchotiwat, Konpat Preechakul

    In the era of large-scale training, model merging has evolved into a tool for creating multitasking models efficiently. It enables the knowledge of models to be fused, without the need for heavy computation as required in traditional multitask learning. Existing merging methods often assume that entries at identical positions in weight matrices serve the sam

  65. Pengfei Zhou, Yunlong Liu, Junli Liang, Qi Song

    Time series forecasting with exogenous variables is a critical emerging paradigm that presents unique challenges in modeling dependencies between variables. Traditional models often struggle to differentiate between endogenous and exogenous variables, leading to inefficiencies and overfitting. In this paper, we introduce CrossLinear, a novel Linear-based for

  66. Yunshen Wang, Yicheng Liu, Tianyuan Yuan, Yingshi Liang

    Accurately predicting 3D occupancy grids from visual inputs is critical for autonomous driving, but current discriminative methods struggle with noisy data, incomplete observations, and the complex structures inherent in 3D scenes. In this work, we reframe 3D occupancy prediction as a generative modeling task using diffusion models, which learn the underlyin

  67. Seohyeong Lee, Eunwon Kim, Hwaran Lee, Buru Chang

    Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs while preserving alignment effectiveness. To address this issue, we propose Alignment Data Map, a data analysis tool for

  68. Olivia McGough, Daniela Witten, Daniel Kessler

    Suppose that a data analyst wishes to report the results of a least squares linear regression only if the overall null hypothesis, $H_0^{1:p}: \beta_1= \beta_2 = \ldots = \beta_p=0$, is rejected. This practice, which we refer to as F-screening (since the overall null hypothesis is typically tested using an $F$-statistic), is in fact common across a number of

  69. Leyan Fang, Romeo Ortega, Robert Griñó

    We carry-out a detailed analysis of direct voltage control of a Boost converter feeding a simple resistive load. First, we prove that using a classical PI control to stabilize a desired equilibrium leads to a very complicated dynamic behavior consisting of two equilibrium points, one of them always unstable for all PI gains and circuit parameter values. Inte

  70. Alexander J. Elias, John T. Wen

    The ABB YuMi is a 7-DOF collaborative robot arm with a complex, redundant kinematic structure. Path planning for the YuMi is challenging, especially with joint limits considered. The redundant degree of freedom is parameterized by the Shoulder-Elbow-Wrist (SEW) angle, called the arm angle by ABB, but the exact definition must be known for path planning outsi

  71. Sahil Verma, Keegan Hines, Jeff Bilmes, Charlotte Siska

    The emerging capabilities of large language models (LLMs) have sparked concerns about their immediate potential for harmful misuse. The core approach to mitigate these concerns is the detection of harmful queries to the model. Current detection approaches are fallible, and are particularly susceptible to attacks that exploit mismatched generalization of mode

  72. Hideo Sugama

    A linearized Vlasov-Poisson system of equations is transformed into a Schr\"{o}dinger equation, which is used to demonstrate that the fluctuation theorem holds for the relative stochastic entropy, defined in terms of the probability density functional of the particle velocity distribution function in the Landau damping process. The difference between the ene

  73. Anusha A. S., Uma Ranjan, Medha Sharma, Siddharth Dutt

    Early detection of dementia is crucial to devise effective interventions. Comprehensive cognitive tests, while being the most accurate means of diagnosis, are long and tedious, thus limiting their applicability to a large population, especially when periodic assessments are needed. The problem is compounded by the fact that people have differing patterns of

  74. Zexuan Li, Hongliang Dai, Piji Li

    Using Large Language Models (LLMs) to generate training data can potentially be a preferable way to improve zero or few-shot NLP tasks. However, many problems remain to be investigated for this direction. For the task of Relation Extraction (RE), we find that samples generated by directly prompting LLMs may easily have high structural similarities with each

  75. Pushapdeep Singh, Jyoti Nigam, Medicherla Vamsi Krishna, Arnav Bhavsar

    While electroencephalography (EEG) has been a popular modality for neural decoding, it often involves task specific acquisition of the EEG data. This poses challenges for the development of a unified pipeline to learn embeddings for various EEG signal classification, which is often involved in various decoding tasks. Traditionally, EEG classification involve

  76. Ning Liu, Yue Yu

    Attention mechanisms have emerged as transformative tools in core AI domains such as natural language processing and computer vision. Yet, their largely untapped potential for modeling intricate physical systems presents a compelling frontier. Learning such systems often entails discovering operators that map between functional spaces using limited instances

  77. Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar

    When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network performance, is highly inefficient, requiring datacenters to reserve full racks of idle accelerators for fault tolerance. In this paper, we address this resource inefficiency by introduci

  78. Yu-Qi Chen, Hai-Shan Liu

    We consider Einstein-Bumblebee gravity and construct a novel Taub-NUT-like black hole solution within this theory. Different from the Taub-NUT black hole in Einstein gravity (which is Ricci-flat), our newly constructed Taub-NUT-like black hole is not Ricci-flat. Armed with the Wald formalism, we extensively study the thermodynamics of this black hole solutio

  79. Zao-Li Chen

    We consider stationary sequences whose marginal tail is subexponential and lies in the Gumbel Maximum domain of attraction. Due to the extremely strong dependence, their extreme values are caused by multiple big values and are clustered in the large scale with fractal features. We establish functional extremal limit theorems with non-Gumbel limit objects to

  80. Yuka Ogino, Takahiro Toizumi, Atsushi Ito

    Low-Light Image Enhancement (LLIE) is crucial for improving both human perception and computer vision tasks. This paper addresses two challenges in zero-reference LLIE: obtaining perceptually 'good' images using the Contrastive Language-Image Pre-Training (CLIP) model and maintaining computational efficiency for high-resolution images. We propose CLIP-Utiliz

  81. Sultan Malik, Felix M. Mayor, Wentao Jiang, Hyunseok Oh

    We demonstrate wavelength-scale phononic waveguides formed by transfer-printed thin-film lithium niobate (LN) on bulk diamond (LNOD), a material stack that combines the strong piezoelectricity of LN with the high acoustic velocity and color-center compatibility of diamond. We characterize a delay line based on a 100 micron long phononic waveguide at room and

  82. Chongjie Si, Xuankun Yang, Muqing Liu, Yadao Wang

    Large-scale foundation models have demonstrated remarkable versatility across a wide range of downstream tasks. However, fully fine-tuning these models incurs prohibitive computational costs, motivating the development of Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA, which introduces low-rank updates to pre-trained weights. Despite their empir

  83. Kuan Xu, Zhiguang Cao, Chenlong Zheng, Linong Liu

    In this study, we propose a reinforcement learning-based adaptive variable neighborhood search (RL-AVNS) method designed for effectively solving the Vehicle Routing Problem with Multiple Time Windows (VRPMTW). Unlike traditional adaptive approaches that rely solely on historical operator performance, our method integrates a reinforcement learning framework t

  84. Qianchao Wang, Leena Heistrene, Yoash Levron, Yuxuan Ding

    In practical resource-constrained environments, efficiently extracting the potential high-frequency fault-critical information is an inherent problem. To overcome this problem, this work suggests leveraging a bi-residual neural network named Bi-ResNet to extract the inner spatial-temporal high-frequency features using embedded spatial-temporal convolution bl

  85. Zhenming Zhang, Haowei Li, Wei Yi

    We study the long-time dynamics of a dissipative Ising chain with varying quantum correlation. Invoking an ensemble-average formalism, and assuming spatial translation symmetry, we show that the dynamics can be described by a Lindblad master equation with an interpolated coherent Hamiltonian. In the classical limit, the interpolation Hamiltonian leads to a s

  86. Cliff Sun, Alexey Bezryadin

    An ordinary superconducting quantum interference device (SQUID) contains two weak links connected in parallel. We model a multiple-wire SQUID (MW-SQUID), generalized in two ways. First, the number of weak links, which are provided by parallel superconducting nanowires, is larger than two. Second, the current-phase relationship of each nanowire is assumed lin

  87. Chongjie Si, Zhiyi Shi, Yadao Wang, Xiaokang Yang

    The rapid development of large language models has revolutionized natural language processing, but their fine-tuning remains computationally expensive, hindering broad deployment. Parameter-efficient fine-tuning (PEFT) methods, such as LoRA, have emerged as solutions. Recent work like DoRA attempts to further decompose weight adaptation into direction and ma

  88. Mian Muhammad Naeem Abid, Nancy Mehta, Zongwei Wu, Radu Timofte

    Lightweight semantic segmentation is essential for many downstream vision tasks. Unfortunately, existing methods often struggle to balance efficiency and performance due to the complexity of feature modeling. Many of these existing approaches are constrained by rigid architectures and implicit representation learning, often characterized by parameter-heavy d

  89. Martin Grant, Kyle Hambrook, Alex Rusterholtz

    We prove that, in an arbitrary ordered field, L'H\^{o}pital's Rule is true if and only if the Least Upper Bound Property is true. We do the same for Taylor's Theorem with Peano Remainder, and for one other property sometimes given as a corollary of L'H\^{o}pital's Rule.

  90. Zeyu Liu, Yuhang Liu, Guanghao Zhu, Congkai Xie

    Recent advancements in large language models (LLMs) have demonstrated substantial progress in reasoning capabilities, such as DeepSeek-R1, which leverages rule-based reinforcement learning to enhance logical reasoning significantly. However, extending these achievements to multimodal large language models (MLLMs) presents critical challenges, which are frequ

  91. Xiaoyu Chang, Fan Zhang, Kexue Fu, Carla Diana

    Dancers often prototype movements themselves or with each other during improvisation and choreography. How are these interactions altered when physically manipulable technologies are introduced into the creative process? To understand how dancers design and improvise movements while working with instruments capable of non-humanoid movements, we engaged dance

  92. Andrew Wood

    A CR-dynamical system is a pair $(X, G)$, where $X$ is a non-empty compact Hausdorff space with uniformity $\mathscr{U}$ and $G$ is a closed relation on $X$. In this paper we introduce the $(i, j)$-shadowing properties in CR-dynamical systems, which generalises the shadowing property from topological dynamical systems $(X, f)$. This extends previous work on

  93. Li Lai, Cezar Lupu, Johannes Sprang

    A famous theorem of Zudilin states that at least one of the Riemann zeta values $\zeta(5), \zeta(7), \zeta(9), \zeta(11)$ is irrational. In this paper, we establish the $p$-adic analogue of Zudilin's theorem. As a weaker form of our result, it is proved that for any prime number $p \geqslant 5$ there exists an odd integer $i$ in the interval $[3,p+p/\log p+5

  94. Takase Shimizu, Kensaku Chida, Gento Yamahata, Katsuhiko Nishiguchi

    We measured the energy efficiency of information erasure using silicon DRAM cells capable of counting charges on capacitors at the single-electron level. Our measurements revealed that the efficiency decreased as the erasure error probability decreased, and notably, the Landauer limit was not achieved even under effectively infinite-time bit erasure. By comp

  95. Junyi An, Xinyu Lu, Chao Qu, Yunfei Shi

    Equivariant Graph Neural Networks (GNNs) have significantly advanced the modeling of 3D molecular structure by leveraging group representations. However, their message passing, heavily relying on Clebsch-Gordan tensor product convolutions, suffers from restricted expressiveness due to the limited non-linearity and low degree of group representations. To over

  96. Gwanghyun Kim, Xueting Li, Ye Yuan, Koki Nagano

    Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations, we present GeoMan, a novel architecture designed to produce

  97. Chang Yu, Fang Liu, Jie Zhu, Shaobo Guo

    This paper proposes a hybrid framework combining LSTM (Long Short-Term Memory) networks with LightGBM and CatBoost for stock price prediction. The framework processes time-series financial data and evaluates performance using seven models: Artificial Neural Networks (ANNs), Convolutional Neural Networks (CNNs), Bidirectional LSTM (BiLSTM), vanilla LSTM, XGBo

  98. Naoki Nishimura, Takumi Funato, Mamoru Matsuo, Takeo Kato

    We theoretically explore the generation of spin current driven by a temperature gradient in a junction between a chiral insulator and a normal metal. Based on the gyromagnetic response induced by microscopic acoustic-phonon-mediated lattice rotation, we derive a formula for the spin current when a finite temperature difference is imposed between two ends of

  99. Haitian Shang, Wei Zhao, Xiaoyu Hong, Leonid I. Gurvits

    We present an investigation of the compact structure of the AGN 2021+317 based on multi-epoch Very Long Baseline Interferometry (VLBI) observations at 15, 22, and 43 GHz in the period from 2013 through 2024. The VLBI images show a core-jet structure extended to the south, with two stationary components in the northern region, one of which likely to be the co

  100. Wenzhi Gao, Ya-Chi Chu, Yinyu Ye, Madeleine Udell

    This paper establishes the theoretical foundations of the online scaled gradient methods (OSGM), a framework that utilizes online learning to adapt stepsizes and provably accelerate first-order methods. OSGM quantifies the effectiveness of a stepsize by a feedback function motivated from a convergence measure and uses the feedback to adjust the stepsize thro