May 2025 arXiv papers — page 52
Showing 5,101–5,200 of 24,552 papers
GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scenes
cs.CVXiao Chen, Tai Wang, Quanyi Li, Tao Huang
Generalizable active mapping in complex unknown environments remains a critical challenge for mobile robots. Existing methods, constrained by insufficient training data and conservative exploration strategies, exhibit limited generalizability across scenes with diverse layouts and complex connectivity. To enable scalable training and reliable evaluation, we
Yifan Sun, Danding Wang, Qiang Sheng, Juan Cao
Concept-based explainable approaches have emerged as a promising method in explainable AI because they can interpret models in a way that aligns with human reasoning. However, their adaption in the text domain remains limited. Most existing methods rely on predefined concept annotations and cannot discover unseen concepts, while other methods that extract co
Shenghai Yuan, Xianyi He, Yufan Deng, Yang Ye
Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose OpenS2V-Nexus, consisting of (i) OpenS2V-Eval, a fine-grained benchmark, and (ii) OpenS2V-5M, a million-scale dataset. In contrast to
Di Wu, Yixin Wan, Kai-Wei Chang
Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such as pose and viewpoint. We proposeVisualize-then-Retrieve (VisRet), a retrieval paradigm that mitigates this limitation of cross-modal similarity alignment. VisRet first projects te
Vincent Liu, Ademi Adeniji, Haotian Zhan, Siddhant Haldar
Despite recent progress in general purpose robotics, robot policies still lag far behind basic human capabilities in the real world. Humans interact constantly with the physical world, yet this rich data resource remains largely untapped in robot learning. We propose EgoZero, a minimal system that learns robust manipulation policies from human demonstrations
Zeyi Huang, Yuyang Ji, Anirudh Sundara Rajan, Zefan Cai
We introduce VisTA, a new reinforcement learning framework that empowers visual agents to dynamically explore, select, and combine tools from a diverse library based on empirical performance. Existing methods for tool-augmented reasoning either rely on training-free prompting or large-scale fine-tuning; both lack active tool exploration and typically assume
Guangting Zheng, Yehao Li, Yingwei Pan, Jiajun Deng
Autoregressive models have emerged as a powerful generative paradigm for visual generation. The current de-facto standard of next token prediction commonly operates over a single-scale sequence of dense image tokens, and is incapable of utilizing global context especially for early tokens prediction. In this paper, we introduce a new autoregressive design to
Zhongwei Zhang, Fuchen Long, Zhaofan Qiu, Yingwei Pan
Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failing to disentangle object and camera moving. To alleviate thes
Alita: Generalist Agent Enabling Scalable Agentic Reasoning with Minimal Predefinition and Maximal Self-Evolution
cs.AIJiahao Qiu, Xuan Qi, Tongcheng Zhang, Xinzhe Juan
Recent advances in large language models (LLMs) have enabled agents to autonomously perform complex, open-ended tasks. However, many existing frameworks depend heavily on manually predefined tools and workflows, which hinder their adaptability, scalability, and generalization across domains. In this work, we introduce Alita--a generalist agent designed with
Weiqi Wu, Xin Guan, Shen Huang, Yong Jiang
Retrieval-Augmented Language Models (RALMs) represent a classic paradigm where models enhance generative capabilities using external knowledge retrieved via a specialized module. Recent advancements in Agent techniques enable Large Language Models (LLMs) to autonomously utilize tools for retrieval, planning, and reasoning. While existing training-based metho
Kaushik Senthoor
In distributed quantum storage, physical qubits of a code will be stored across the network. When qubits in one of the nodes are lost i.e. when the node is erased, the remaining nodes need to communicate with a new node to replace the lost qubits. Here, we look at the problem of how much entanglement cost is needed to perform such a distributed quantum erasu
Guangzhao He, Chen Geng, Shangzhe Wu, Jiajun Wu
The motion of deformable 4D objects lies in a low-dimensional manifold. To better capture the low dimensionality and enable better controllability, traditional methods have devised several heuristic-based methods, i.e., rigging, for manipulating dynamic objects in an intuitive fashion. However, such representations are not scalable due to the need for expert
Zitian Gao, Lynx Chen, Haoming Luo, Joey Zhou
We trained 13,440 large language models and found that entropy minimization requires only a single unlabeled data and 10 steps optimization to achieve performance improvements comparable to or even greater than those obtained using thousands of data and carefully designed rewards in rule-based reinforcement learning. This striking result may prompt a rethink
A Roadmap for neutrino charge assignments in $U(2)_F$ Flavor Models: Implications for LFV processes and leptonic anomalous magnetic moments
hep-phA. Giarnetti, S. Marciano, D. Meloni, M. Rettaroli
We build upon a simple $U(2)_F$ model of flavor, in which all fermion masses and mixing hierarchies arise from powers of two small parameters controlling $U(2)_F$ breaking. In the original formulation, an isomorphism to the discrete $D_6\times U(1)_F$ symmetry was invoked to generate a Majorana neutrino mass term. Here, we retain the successful features of t
Jonas Spinner, Luigi Favaro, Peter Lippmann, Sebastian Pitz
Lorentz-equivariant neural networks are becoming the leading architectures for high-energy physics. Current implementations rely on specialized layers, limiting architectural choices. We introduce Lorentz Local Canonicalization (LLoCa), a general framework that renders any backbone network exactly Lorentz-equivariant. Using equivariantly predicted local refe
Zhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang
The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving deep spatial understanding comparable to human capabilities poses significant challenges in model encoding and data acquisition. Existing meth
Sijia Chen, Yanqiu Yu, En Yu, Wenbing Tao
Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on explicit visual-textual matching and consequently fail to generalize to complex instructions that require logical reasoning. To overcome this, we propose Reasoning-based Multi-Object Tracking (ReaMOT), a novel task
Hoyeon Chang, Jinho Park, Hanseul Cho, Sohee Yang
Despite impressive capabilities, LLMs' successes often rely on pattern-matching behaviors, yet these are also linked to OOD generalization failures in compositional tasks. However, behavioral studies commonly employ task setups that allow multiple generalization sources (e.g., algebraic invariances, structural repetition), obscuring a precise and testable ac
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
cs.CLHaonan Zhang, Run Luo, Xiong Liu, Yuchuan Wu
Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the role's voice traits (e.g., voice style and emotions) as playing a crucial effect in
Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska
Large language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency. Quantization can mitigate these costs, but may degrade performance. In this work, we present the first systematic evaluation of quantized LLMs on tasks with long inputs (>64K tokens) and long-form out
Yang Ye, Xianyi He, Zongjian Li, Bin Lin
Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks. To overcome these limitations, we introduce ImgEdit, a large-scale, high-quality image-editing dataset
Kejing Lu, Chuan Xiao, Yoshiharu Ishikawa
In this paper, we study the angle testing problem in the context of similarity search in high-dimensional Euclidean spaces and propose two projection-based probabilistic kernel functions, one designed for angle comparison and the other for angle thresholding. Unlike existing approaches that rely on random projection vectors drawn from Gaussian distributions,
Ke Yang, ChengXiang Zhai
The rapid rise of AI-based autonomous agents is transforming human society and economic systems, as these entities increasingly exhibit human-like or superhuman intelligence. From excelling at complex games like Go to tackling diverse general-purpose tasks with large language and multimodal models, AI agents are evolving from specialized tools into dynamic p
Meng Cao, Haoze Zhao, Can Zhang, Xiaojun Chang
Large Vision-Language Models (LVLMs) have become powerful general-purpose assistants, yet their predictions often lack reliability and interpretability due to insufficient grounding in visual evidence. The emerging thinking-with-images paradigm seeks to address this issue by explicitly anchoring reasoning to image regions. However, we empirically find that m
In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
cs.CVYu Xu, Fan Tang, You Wu, Lin Gao
Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly "brushes" user-specified objects into a given image guided by textual prompts. However, existing methods often struggle to insert customized subjects with high fidelity and align results with the user's intent through t
ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation
cs.CVJinsheng Quan, Qiaowei Miao, Yichao Xu, Zizhuo Lin
The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consistency in predicting the future. To overcome this issue, we pr
Does AI and Human Advice Mitigate Punishment for Selfish Behavior? An Experiment on AI ethics From a Psychological Perspective
cs.CYMargarita Leib, Nils Köbis, Ivan Soraperra
People increasingly rely on AI-advice when making decisions. At times, such advice can promote selfish behavior. When individuals abide by selfishness-promoting AI advice, how are they perceived and punished? To study this question, we build on theories from social psychology and combine machine-behavior and behavioral economic approaches. In a pre-registere
Levi Cordeiro Carvalho, Saulo A. F. Oliveira, Thiago Alves Rocha
Providing explanations for the outputs of artificial neural networks (ANNs) is crucial in many contexts, such as critical systems, data protection laws and handling adversarial examples. Logic-based methods can offer explanations with correctness guarantees, but face scalability challenges. Due to these issues, it is necessary to compare different encodings
Fan Chen, Zeyu Jia, Alexander Rakhlin, Tengyang Xie
Reinforcement learning with outcome-based feedback faces a fundamental challenge: when rewards are only observed at trajectory endpoints, how do we assign credit to the right actions? This paper provides the first comprehensive analysis of this problem in online RL with general function approximation. We develop a provably sample-efficient algorithm achievin
Changjian Jiang, Kerui Ren, Linning Xu, Jiong Chen
High fidelity 3D reconstruction and rendering hinge on capturing precise geometry while preserving photo realistic detail. Most existing methods either fuse these goals into a single cumbersome model or adopt hybrid schemes whose uniform primitives lead to a trade off between efficiency and fidelity. In this paper, we introduce HaloGS, a dual representation
Alexander Conway, Debadeepta Dey, Stefan Hackmann, Matthew Hausknecht
Retrieval-Augmented Generation (RAG) pipelines are central to applying large language models (LLMs) to proprietary or dynamic data. However, building effective RAG flows is complex, requiring careful selection among vector databases, embedding models, text splitters, retrievers, and synthesizing LLMs. The challenge deepens with the rise of agentic paradigms.
Alexander M. Dalzell, András Gilyén, Connor T. Hann, Sam McArdle
We present a protocol for fault-tolerantly implementing the logical quantum random access memory (QRAM) operation, given access to a specialized, noisy QRAM device. For coherently accessing classical memories of size $2^n$, our protocol consumes only $\mathrm{poly}(n)$ fault-tolerant quantum resources (logical gates, logical qubits, quantum error correction
Dong Nguyen, Esther Ploeger
Although diversity in NLP datasets has received growing attention, the question of how to measure it remains largely underexplored. This opinion paper examines the conceptual and methodological challenges of measuring data diversity and argues that interdisciplinary perspectives are essential for developing more fine-grained and valid measures.
Alessandra Venditti, Julian B. Munoz, Volker Bromm, Seiji Fujimoto
The nature of the first, so-called Population III (Pop III) stars has for long remained largely unconstrained. However, the James Webb Space Telescope (JWST) finally opened new concrete prospects for their detection during the Epoch of Reionization (EoR), notably providing promising observational constraints on the Pop III ultra-violet luminosity function (U
Arkady A. Tseytlin, Zihan Wang
As was shown in arXiv:1203.1054, expanding the Nambu action near the "long string" vacuum in the static gauge one finds that the one-loop 2 to 2 scattering amplitude of the $D-2=24$ transverse 2d fluctuations $X^i$ is given by a pure phase expression consistent with underlying integrability. Similar computation in the Green-Schwarz superstring was carried ou
Eric J. Kuehnke, Kyano Levi, Joschka Roffe, Jens Eisert
Quantum error correction is the art of protecting fragile quantum information through suitable encoding and active interventions. After encoding $k$ logical qubits into $n>k$ physical qubits using a stabilizer code, this amounts to measuring stabilizers, decoding syndromes, and applying an appropriate correction. Although quantum information can be protected
Dmitry Galakhov, Alexei Morozov
We continue the study of quantum A-polynomials -- equations for knot polynomials with respect to their coloring (representation-dependence) -- as the relations between different links, obtained by hanging additional ``simple'' components on the original knot. Depending on the choice of this ``decoration'', the knot polynomial is either multiplied by a number
Haoyu Wang, Zeyu Qin, Yifei Zhao, Chao Du
LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for unseen attacks that may arise during deployment. To address this, we propose a lifelong safety al
Siye Wu, Jian Xie, Yikai Zhang, Aili Chen
While large reasoning models demonstrate strong performance on complex tasks, they lack the ability to adjust reasoning token usage based on task difficulty. This often leads to the "overthinking" problem -- excessive and unnecessary reasoning -- which, although potentially mitigated by human intervention to control the token budget, still fundamentally cont
Qianhui Xu, Guangyue Ji, Yuzhu Wang, Ha Quang Trung
In fractional quantum Hall fluids, the quasiparticle excitations are anyons with fractional charges and statistics. Effective interactions among the anyons can be induced by either model or realistic electron-electron (e-e) interactions. Without losing the generality, we investigate such phenomena for the Laughlin 1/3 and Moore-Read (MR) non-Abelian phases.
Hao Zhong, Muzhi Zhu, Zongze Du, Zheng Huang
Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution inputs. We tackle this trade-off with a two-system architecture: a Global Reasoning System selects informative keyframes an
Simin Fan, Maria Ios Glarou, Martin Jaggi
The performance of large language models (LLMs) across diverse downstream applications is fundamentally governed by the quality and composition of their pretraining corpora. Existing domain reweighting algorithms primarily optimize data mixtures for a single target task, thereby resulting in models that overfit to specialized objectives while exhibiting subs
Xiangchen Song, Aashiq Muhamed, Yujia Zheng, Lingjing Kong
Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration to identify a canonical set of features is challenged by the observed inconsistency of learned SAE features across different training runs, undermining the reliability and efficie
Quantisation ideals, canonical parametrisations of the unipotent group and consistent integrable systems
nlin.SIM. A. Chirkov, A. V. Mikhailov, D. V. Talalaev
Using the methods of quantisation ideals, we construct a family of quantisations corresponding to Case alpha in Sergeev's classification of solutions to the tetrahedron equation. This solution describes transformations between special parametrisations of the space of unipotent matrices with noncommutative coefficients. We analyse the classical limit of this
Sophia Hager, Aleem Khan, Andrew Wang, Nicholas Andrews
Most successful applications of deep learning involve similar training and test conditions. However, tasks such as biological sequence design involve searching for sequences that improve desirable properties beyond previously known values, which requires novel hypotheses that \emph{extrapolate} beyond training data. In these settings, extrapolation may be ac
Hee-Seon Kim, Minbeom Kim, Wonjun Lee, Kihyun Kim
Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the model predict the next token of a toxic prompt. However, we find that the Toxic-Continuation paradigm is effective at conti
Chirag Garg, Sayeef Salahuddin
Ising Machines are emerging hardware architectures that efficiently solve NP-Hard combinatorial optimization problems. Generally, combinatorial problems are transformed into quadratic unconstrained binary optimization (QUBO) form, but this transformation often complicates the solution landscape, degrading performance, especially for multi-state problems. To
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
cs.CLYongan Yu, Qingchen Hu, Xianda Du, Jiayin Wang
Climate change adaptation requires the understanding of disruptive weather impacts on society, where large language models (LLMs) might be applicable. However, their effectiveness is under-explored due to the difficulty of high-quality corpus collection and the lack of available benchmarks. The climate-related events stored in regional newspapers record how
Amar Deo Chandra
RX J0209.6-7427 is an ultraluminous X-ray pulsar (ULXP) having spin period of about 9.3 s. To date, no cyclotron resonance scattering features have been detected in this source, which can enable direct measurement of the magnetic field of the pulsar. We estimate the surface magnetic field of the neutron star in this source using different models and find tha
Translation of Enterprise Architecture Concept to Facilitate Digital Transformation Initiatives in Vietnam: Processes, Mechanisms and Impacts
cs.OHDuong Dang, Quang Bui
Governments around the world have increasingly adopted digital transformation (DT) initiatives to increase their strategic competitiveness in the global market. To support successful DT, governments have to introduce new governance logics and revise IT strategies to facilitate DT initiatives. In this study, we report a case study of how Enterprise Architectu
Jiahao Qiu, Fulian Xiao, Yimin Wang, Yuchen Mao
Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents p
KnowTrace: Bootstrapping Iterative Retrieval-Augmented Generation with Structured Knowledge Tracing
cs.CLRui Li, Quanyu Dai, Zeyu Zhang, Xu Chen
Recent advances in retrieval-augmented generation (RAG) furnish large language models (LLMs) with iterative retrievals of relevant information to handle complex multi-hop questions. These methods typically alternate between LLM reasoning and retrieval to accumulate external information into the LLM's context. However, the ever-growing context inherently impo
Biaxial characterization of soft elastomers: experiments and data-adaptive configurational forces for fracture
cond-mat.softMiguel Angel Moreno-Mateos, Simon Wiesheier, Ali Esmaeili, Mokarram Hossain
Understanding the fracture mechanics of soft solids remains a fundamental challenge due to their complex, nonlinear responses under large deformations. While multiaxial loading is key to probing their mechanical behavior, the role of such loading in fracture processes is still poorly understood. Here, we present a combined experimental-computational framewor
Bhawna Piryani, Abdelrahman Abdallah, Jamshid Mozafari, Avishek Anand
Time plays a critical role in how information is generated, retrieved, and interpreted. In this survey, we provide a comprehensive overview of Temporal Question Answering (TQA), a research area that focuses on answering questions involving temporal constraints or context. As time-stamped content from sources like news articles, web archives, and knowledge ba
Nguyen Thach, Aida Riahifar, Nathan Huynh, Hau Chan
Solving NP-hard combinatorial optimization problems (COPs) (e.g., traveling salesman problems (TSPs) and capacitated vehicle routing problems (CVRPs)) in practice traditionally involves handcrafting heuristics or specifying a search space for finding effective heuristics. The main challenges from these approaches, however, are the sheer amount of domain know
Qi Cao, Ruiyi Wang, Ruiyi Zhang, Sai Ashish Somayajula
Reasoning has substantially improved the performance of large language models (LLMs) on complicated tasks. Central to the current reasoning studies, Process Reward Models (PRMs) offer a fine-grained evaluation of intermediate reasoning steps and guide the reasoning process. However, extending PRMs to multimodal large language models (MLLMs) introduces challe
Hierarchical Bayesian estimation for continual learning during model-informed precision dosing
stat.COFranziska Thoma, Niklas Hartung, Manfred Opper, Wilhelm Huisinga
Model informed precision dosing (MIPD) is a Bayesian framework to individualize drug therapy based on prior knowledge and patient-specific monitoring data. Typically, prior knowledge results from controlled clinical trials with a more homogeneous patient population compared to the real-world patient population underlying the data to be analysed. Thus, devisi
Unleashing 5G Seamless Integration with TSN for Industry 5.0: Frame Forwarding and QoS Treatment
cs.NIOscar Adamuz-Hinojosa, Felix Delgado-Ferro, Jorge Navarro-Ortiz, Pablo Muñoz
Integrating Time-Sensitive Networking (TSN) and 5th Generation (5G) systems is key for providing wireless low-latency services in industry. Despite research efforts, challenges remain. Due to the lack of commercial 5G modems supporting Ethernet-based sessions, tunneling mechanisms must be used to enable Layer 2 connectivity between TSN islands via IP-based 5
Yasmin Moslem
Efficient deployment of large audio-language models for speech translation remains challenging due to their significant computational requirements. In this paper, we address this challenge through our system submissions to the "Model Compression" track at the International Conference on Spoken Language Translation (IWSLT 2025). We experiment with a combinati
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
cs.CVWeihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang
Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged as a lightweight and interpretable solution in large language models (LLMs). However, its effectiveness in vision-languag
Jonathan Wenger, Beau Coker, Juraj Marusic, John P. Cunningham
Modern deep learning models generalize remarkably well in-distribution, despite being overparametrized and trained with little to no explicit regularization. Instead, current theory credits implicit regularization imposed by the choice of architecture, hyperparameters, and optimization procedure. However, deep neural networks can be surprisingly non-robust,
Tamas Gombor
We derive a universal formula for the overlaps between integrable matrix product states (MPS) and Bethe eigenstates in $\mathfrak{gl}_{N}$ symmetric spin chains. This formula expresses the normalized overlap as a product of a MPS-independent Gaudin-determinant ratio and a MPS-dependent scalar factor constructed from eigenvalues of commuting operators, define
Eliran Sherzer, Yehezkel Resheff, Miklos Telek
Phase type (PH) distributions are widely used in modeling and simulation due to their generality and analytical properties. In such settings, it is often necessary to construct a PH distribution that aligns with real-world data by matching a set of prescribed moments. Existing approaches provide either exact closed-form solutions or iterative procedures that
Takaki Matsumoto, Kanta Nakano, Ryosuke Suda, Kentaroh Yoshida
Oscillons are classical oscillatory solutions with very long but finite lifetimes in real scalar field theories with appropriate potentials. An interesting feature is that resonances appear in the lifetimes of the oscillon for the initial size of the oscillon core $R_0$, which was discovered by Honda and Choptuik in the case of Minkowski space. In a previous
Pranav Poudel, Aavash Chhetri, Prashnna Gyawali, Georgios Leontidis
Multimodal federated learning holds immense potential for collaboratively training models from multiple sources without sharing raw data, addressing both data scarcity and privacy concerns, two key challenges in healthcare. A major challenge in training multimodal federated models in healthcare is the presence of missing modalities due to multiple reasons, i
Yiming Du, Bingbing Wang, Yang He, Bin Liang
Modern task-oriented dialogue (TOD) systems increasingly rely on large language model (LLM) agents, leveraging Retrieval-Augmented Generation (RAG) and long-context capabilities for long-term memory utilization. However, these methods are primarily based on semantic similarity, overlooking task intent and reducing task coherence in multi-session dialogues. T
Carlos J. Fernandez-Candel, Anthony Cleve, Jesus J. Garcia-Molina
In this paper, we present a static code analysis strategy to extract logical schemas from NoSQL applications. Our solution is based on a model-driven reverse engineering process composed of a chain of platform-independent model transformations. The extracted schema conforms to the U-Schema unified metamodel, which can represent both NoSQL and relational sche
Maximilian Dreyer, Lorenz Hufe, Jim Berend, Thomas Wiegand
Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. While recent works show that Sparse Autoencoders (SAEs) yield interpretable latent components, they focus on what these encode and miss how they drive predictions. We introduce a scala
Theoretical Study of Charge Transport Properties of Curved PAH Organic Semiconductors
physics.chem-phHengyu Jin, Xiaoqi Sun, Guiya Qin, Zhipeng Tong
Curved polycyclic aromatic hydrocarbons (PAHs) exhibit distinctive geometric and electronic structures, rendering them highly promising in addressing issues of solubility and air stability, which are faced for large linear arene $\pi$-conjugated organic semiconductors. In this study, a series of surface-curved PAHs and the heteroatom doped derivatives are se
Yi Wen, Yue Liu, Derong Xu, Huishi Luo
Multi-Domain Recommendation (MDR) achieves the desirable recommendation performance by effectively utilizing the transfer information across different domains. Despite the great success, most existing MDR methods adopt a single structure to transfer complex domain-shared knowledge. However, the beneficial transferring information should vary across different
Passive Cavitation Mitigation on Hydrofoils via Porous Media: A Comparative Study of LES and RANS Models
physics.flu-dynAli Alavi, Maziyar Ghasemnezhad, Ali Sangtarash, Ehsan Roohi
This study numerically investigates the use of porous media as a passive strategy for mitigating cavitation on a NACA 66 (MOD) hydrofoil subjected to unsteady two-phase flow.
Hao Kang, Zichun Yu, Chenyan Xiong
Recent large language models such as Gemini-1.5, DeepSeek-V3, and Llama-4 increasingly adopt Mixture-of-Experts (MoE) architectures, which offer strong efficiency-performance trade-offs by activating only a fraction of the model per token. Yet academic researchers still lack a fully open, end-to-end MoE platform for investigating scaling, routing, and expert
Shripad M. Garge, Deep H. Makadiya
The existence of triangular and unitriangular factorizations has been extensively studied for untwisted Chevalley groups, as well as for twisted Chevalley groups of types other than ${}^2A_{2n} \ (n \geq 1)$. However, the case of twisted Chevalley groups of type ${}^2A_{2n} \ (n \geq 1)$, has remained unresolved in the general setting of commutative rings. P
Yixin Cui, Haotian Lin, Shuo Yang, Yixiao Wang
The rapid evolution of large language models in natural language processing has substantially elevated their semantic understanding and logical reasoning capabilities. Such proficiencies have been leveraged in autonomous driving systems, contributing to significant improvements in system performance. Models such as OpenAI o1 and DeepSeek-R1, leverage Chain-o
FT-Boosted SV: Towards Noise Robust Speaker Verification for English Speaking Classroom Environments
eess.ASSaba Tabatabaee, Jing Liu, Carol Espy-Wilson
Creating Speaker Verification (SV) systems for classroom settings that are robust to classroom noises such as babble noise is crucial for the development of AI tools that assist educational environments. In this work, we study the efficacy of finetuning with augmented children datasets to adapt the x-vector and ECAPA-TDNN to classroom environments. We demons
Xiao Shou, Yanna Ding, Jianxi Gao
Training deep neural networks remains computationally intensive due to the itera2 tive nature of gradient-based optimization. We propose Gradient Flow Matching (GFM), a continuous-time modeling framework that treats neural network training as a dynamical system governed by learned optimizer-aware vector fields. By leveraging conditional flow matching, GFM ca
Marek Kosiek, Krzysztof Rudol
Using a description of the spectrum of bidual algebra $A^{**}$ of a uniform algebra $A$ we obtain abstract corona theorem for certain uniform algebras. It asserts the density of a specific Gleason part in the spectrum of an $H^\infty$ -- type subalgebra of $A^{**}$. There is an isometric isomorphism of the latter subalgebra with $H^\infty(G)$ for a wide clas
Francesco Orabona, Ryan D'Orazio
The Polyak stepsize has been proven to be a fundamental stepsize in convex optimization, giving near optimal gradient descent rates across a wide range of assumptions. The universality of the Polyak stepsize has also inspired many stochastic variants, with theoretical guarantees and strong empirical performance. Despite the many theoretical results, our unde
From one-dimensional diffusion processes metastable behaviour to parabolic equations asymptotics
math.PRClaudio Landim, Christian Maura
Consider the one-dimensional elliptic operator given by \begin{equation*} (L_\epsilon f)(x) \;=\; b (x) \, f'(x) \,+\, \epsilon\, a (x)\, f''(x) \;, \end{equation*} where the drift $b\colon R \to R$ and the diffusion coefficient $a\colon R \to R$ are periodic $C^1(R)$ functions satisfying further conditions, and $\epsilon>0$. Consider the initial-valued prob
Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence
eess.ASEdem Ahadzi, Vishwanath Pratap Singh, Tomi Kinnunen, Ville Hautamaki
In this work, we present the first study addressing automatic speech recognition (ASR) for children in an online learning setting. This is particularly important for both child-centric applications and the privacy protection of minors, where training models with sequentially arriving data is critical. The conventional approach of model fine-tuning often suff
Paolo Gajo, Domenic Rosati, Hassan Sajjad, Alberto Barrón-Cedeño
Dependency parsing is the task of inferring natural language structure, often approached by modeling word interactions via attention through biaffine scoring. This mechanism works like self-attention in Transformers, where scores are calculated for every pair of words in a sentence. However, unlike Transformer attention, biaffine scoring does not use normali
Sitong Fang, Wenjing Cao, Jiahao Li, Xuyao Wang
Reasoning models have attracted increasing attention for their ability to tackle complex tasks, embodying the System II (slow thinking) paradigm in contrast to System I (fast, intuitive responses). Yet a key question remains: Does slower reasoning necessarily lead to more truthful answers? Our findings suggest otherwise. We conduct the first systematic study
Qi Liang, RuGway Wu, Pradyumna Paranjape, Ben Schittenkopf
Spatio-temporal scaling dynamics connected to non-thermal fixed points has been suggested as a universal framework to describe the relaxation of isolated far-from-equilibrium systems. Experimental studies in weakly-interacting cold atom systems have found scaling dynamics connected to specific attractors. In our experiments, we study a quantum gas of strongl
Xiaoyuan Wu, Weiran Lin, Omer Akgul, Lujo Bauer
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have been proposed to mitigate such hallucinations and fragility, one of which is to measure the consistency of LLM responses -- the model's confidence in the response or likelihood of
Richard H. D. Townsend, Rianna V. Kuenzi, Jørgen Christensen-Dalsgaard
Stellar oscillation codes are software instruments that evaluate the normal-mode frequencies of an input stellar model. While inter-code comparisons are often used to confirm the correctness of calculations, they are not suitable for characterizing the numerical error of an individual code. To address this issue, we introduce a set of tools -- 'error measure
Junseo Hwang, Wonguk Cho, Taesup Kim
Fine-tuning large foundation models is essential for building expert models tailored to specialized tasks and domains, but fully updating billions of parameters is computationally prohibitive. Reducing the number of trainable parameters using Parameter-Efficient Fine-Tuning (PEFT), such as Low-Rank Adaptation (LoRA), is therefore crucial not only to reduce t
A structure-preserving multiscale solver for particle-wave interaction in non-uniform magnetized plasmas
math.NAKun Huang, Irene M. Gamba, Chi-Wang Shu
Particle-wave interaction is of fundamental interest in plasma physics, especially in the study of runaway electrons in magnetic confinement fusion. Analogous to the concept of photons and phonons, wave packets in plasma can also be treated as quasi-particles, called plasmons. To model the ``mixture" of electrons and plasmons in plasma, a set of ``collisiona
Joe Stacey, Lisa Alazraki, Aran Ubhi, Beyza Ermis
We investigate the robustness of fine-tuned Large Language Models (LLMs) for the task of Natural Language Inference (NLI), finding that the in-distribution gains from fine-tuning correspond to a large drop in out-of-distribution (OOD) performance. Despite the widespread use of closed-source LLMs, there are no robustness mitigation methods that work under the
Kyrylo Simonov, Rafael Wagner, Ernesto Galvão
Bargmann invariants of order $n$, defined as multivariate traces of quantum states $\text{Tr}[\rho_1\rho_2 \ldots \rho_n]$, are useful in applications ranging from quantum metrology to certification of nonclassicality. A standard quantum circuit used to estimate Bargmann invariants is the cycle test. In this work, we propose generalizations of the cycle test
Soham Chakraborty, S. Krishna, Andreas Pavlogiannis, Omkar Tuppe
GPU computing is embracing weak memory concurrency for performance improvement. However, compared to CPUs, modern GPUs provide more fine-grained concurrency features such as scopes, have additional properties like divergence, and thereby follow different weak memory consistency models. These features and properties make concurrent programming on GPUs more co
Umut Cihan, Arda İçöz, Vahid Haratian, Eray Tüzün
Context: Code reviews are crucial for software quality. Recent AI advances have allowed large language models (LLMs) to review and fix code; now, there are tools that perform these reviews. However, their reliability and accuracy have not yet been systematically evaluated. Objective: This study compares different LLMs' performance in detecting code correctne
Pei-Hao Fu, Sayan Mondal, Jun-Feng Liu, Yukio Tanaka
We consider unconventional magnets with and without spin-singlet $s$-wave superconductivity and demonstrate the emergence of spin triplet states due to light drives. In particular, we find that a high-frequency linearly polarized light drive induces a spin-triplet density in $d$-wave altermagnets which does not exist in the static regime and can directly rev
Mechanism of defect formation in the quantum annealing of the random transverse-field Ising chain
cond-mat.stat-mechRóbert Juhász
Based on the strong-disorder renormalization group method, a microscopic mechanism of defect formation in the quantum annealing of the random transverse-field Ising chain is proposed, which represents the annealing process as a gradual aggregation of strongly coupled spin clusters. The ferromagnetic ground state of clusters is either preserved or get excited
PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology
cs.CVJiabo Ma, Yingxue Xu, Fengtao Zhou, Yihui Wang
The emergence of pathology foundation models has revolutionized computational histopathology, enabling highly accurate, generalized whole-slide image analysis for improved cancer diagnosis, and prognosis assessment. While these models show remarkable potential across cancer diagnostics and prognostics, their clinical translation faces critical challenges inc
Mohit Chandra, Siddharth Sriraman, Harneet Singh Khanuja, Yiqiao Jin
Limited access to mental healthcare, extended wait times, and increasing capabilities of Large Language Models (LLMs) has led individuals to turn to LLMs for fulfilling their mental health needs. However, examining the multi-turn mental health conversation capabilities of LLMs remains under-explored. Existing evaluation frameworks typically focus on diagnost
Dawn Virginillo, Asja Derviškadić, Mario Paolone
The expected decrease in system inertia and frequency stability motivates the development and maintenance of dynamic system models by Transmission System Operators. However, some dynamic model parameters can be unavailable due to market unbundling, or inaccurate due to aging infrastructure, non-documented tuning of controllers, or other factors. In this pape
Pengxiang Li, Shilin Yan, Joey Tsai, Renrui Zhang
Classifier-Free Guidance (CFG) significantly enhances controllability in generative models by interpolating conditional and unconditional predictions. However, standard CFG often employs a static unconditional input, which can be suboptimal for iterative generation processes where model uncertainty varies dynamically. We introduce Adaptive Classifier-Free Gu
Mengjian Hua, Charles S. Peskin
We describe in this paper a crossbridge model in which an attached crossbridge behaves like a linear spring with a variable rest length. We assume in particular that the rest length has a linear force-velocity relation, and that the force and rest length are both zero at the moment of crossbridge attachment. Crossbridges that are not attached in our model ha
Fabian Lange, Johann Usovitsch, Zihao Wu
We present version 3 of Kira, a Feynman integral reduction program for high-precision calculations in quantum field theory and gravitational-wave physics. Building on previous versions, Kira 3 introduces optimized seeding and equation selection algorithms, significantly improving performance for multi-loop and multi-scale problems. New features include conve
Yuetai Li, Zhangchen Xu, Fengqing Jiang, Bhaskar Ramasubramanian
Fine-tuning large language models (LLMs) is intended to improve their reasoning capabilities, yet we uncover a counterintuitive effect: models often forget how to solve problems they previously answered correctly during training. We term this phenomenon temporal forgetting and show that it is widespread across model sizes, fine-tuning methods (both Reinforce