May 2025 arXiv papers — page 58
Showing 5,701–5,800 of 24,552 papers
Rong-Cheng Tu, Zhao Jin, Jingyi Liao, Xiao Luo
Existing Zero-Shot Composed Image Retrieval (ZS-CIR) methods typically train adapters that convert reference images into pseudo-text tokens, which are concatenated with the modifying text and processed by frozen text encoders in pretrained VLMs or LLMs. While this design leverages the strengths of large pretrained models, it only supervises the adapter to pr
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
cs.CLTej Deep Pala, Panshul Sharma, Amir Zadeh, Chuan Li
Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Reward Models (PRMs) score each intermediate step to steer generation toward coherent solutions. We introduce PathFinder-PRM, a novel hierarchic
Solving Euler equations with Multiple Discontinuities via Separation-Transfer Physics-Informed Neural Networks
physics.flu-dynChuanxing Wang, Hui Luo, Kai Wang, Guohuai Zhu
Despite the remarkable progress of physics-informed neural networks (PINNs) in scientific computing, they continue to face challenges when solving hydrodynamic problems with multiple discontinuities. In this work, we propose Separation-Transfer Physics Informed Neural Networks (ST-PINNs) to address such problems. By sequentially resolving discontinuities fro
Kaizhe Chen, Heng Zhang
We derive some existence results for the solutions of the Tzitz\'eica equation \begin{equation*} -\Delta u + h_1(x)e^{Au} + h_2(x)e^{-Bu}=0 \end{equation*} and the generalized Tzitz\'eica equation \begin{equation*} -\Delta u + h_1(x)e^{Au}(e^{Au}-1)+h_2(x)e^{-Bu}(e^{-Bu}-1)=0 \end{equation*} on any connected finite graph \(G=(V, E)\). Here, \(h_1(x)>0\), \(h
Model Predictive Online Monitoring of Dynamical Systems for Nested Signal Temporal Logic Specifications
math.OCTao Han, Shaoyuan Li, Xiang Yin
This paper investigates the online monitoring problem for cyber-physical systems under signal temporal logic (STL) specifications. The objective is to design an online monitor that evaluates system correctness at runtime based on partial signal observations up to the current time so that alarms can be issued whenever the specification is violated or will ine
Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin
Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains challenging due to inherent limitations in text-only CoT, such as visual hallucinations and insufficient multimodal integrati
Kodai Hasegawa, Shigeaki Okumura, Hirofumi Taki, Hironobu Sunadome
Radar-based respiratory measurement is a promising tool for the noncontact detection of sleep apnea. Our team has reported that apnea events can be accurately detected using the statistical characteristics of the amplitude of respiratory displacement. However, apnea and hypopnea events are often followed by irregular breathing, reducing the detection accurac
Yi Liu, Dianqing Liu, Mingye Zhu, Junbo Guo
The widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse applications. To address this limitation, we propose a novel \
Jing Yu Lim, Rushi Shah, Zarif Ikram, Samson Yu
Recently, Model-Based Reinforcement Learning (MBRL) have achieved super-human level performance on the Atari100k benchmark on average. However, we discover that conventional aggregates mask a major problem, Performance Asymmetry: MBRL agents dramatically outperform humans in certain tasks (Agent-Optimal tasks) while drastically underperform humans in other t
Swati Saha, Ranbir Singh, Bedangadas Mohanty
The transverse momentum-differential radial flow observable $v_0(p_\mathrm{T})$, recently proposed and measured by the ATLAS and ALICE collaborations, provides a novel tool to probe radial expansion dynamics in high-energy heavy-ion collisions. In this work, we conduct a detailed study of $v_0(p_\mathrm{T})$ using a blast-wave model that incorporates hydrody
Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
cs.CVMohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Nour Aburaed
This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend tha
Galina Weinstein
This paper critically examines the central thesis of Kieran Fox's "I Am a Part of Infinity: The Spiritual Journey of Albert Einstein"-namely, that Einstein's intellectual development constitutes a coherent spiritual path culminating in a form of pantheistic mysticism shaped by both Western and Eastern traditions. Fox presents Einstein as the modern heir to a
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
cs.CVWen Yin, Yong Wang, Guiduo Duan, Dongyang Zhang
Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recogni
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
cs.SDDeok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee
Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model that integrates spherical VAD region classification to guide VAD regression for improved emotion prediction. In our framework, VAD values are
DriveCamSim: Generalizable Camera Simulation via Explicit Camera Modeling for Autonomous Driving
cs.CVWenchao Sun, Xuewu Lin, Keyu Chen, Zixiang Pei
Camera sensor simulation serves as a critical role for autonomous driving (AD), e.g. evaluating vision-based AD algorithms. While existing approaches have leveraged generative models for controllable image/video generation, they remain constrained to generating multi-view video sequences with fixed camera viewpoints and video frequency, significantly limitin
Chon Man Sou
Since the original derivation of Hawking radiation, there have been lots of alternative approaches to show the same fact that black holes emit particles as hot bodies with a temperature. These alternative methods generally rely on different conditions and physical quantities to manifest the radiation, providing various points of view of this effect in the in
Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan
Despite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain. Existing evaluations primarily assess response-level safety, neglecting a critical issue we identify as \textbf{\textit{Superficial Safety Alignment} (SSA)} -- a phenomenon where m
ATLAS Collaboration
Jet flavour tagging enables the identification of jets originating from heavy-flavour quarks in proton-proton collisions at the Large Hadron Collider, playing a critical role in its physics programmes. This paper presents GN2, a transformer-based flavour tagging algorithm deployed by the ATLAS Collaboration that represents a different methodology compared to
GeoPF: Infusing Geometry into Potential Fields for Reactive Planning in Non-trivial Environments
cs.ROYuhe Gong, Riddhiman Laha, Luis Figueredo
Reactive intelligence remains one of the cornerstones of versatile robotics operating in cluttered, dynamic, and human-centred environments. Among reactive approaches, potential fields (PF) continue to be widely adopted due to their simplicity and real-time applicability. However, existing PF methods typically oversimplify environmental representations by re
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
cs.SDDeok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee
Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully separate speaker and emotion characteristics, causing speaker leakage and degraded synthesis quality. To address this, we prop
Victor M. Tenorio, Nicolas Zilberstein, Santiago Segarra, Antonio G. Marques
Diffusion models have emerged as powerful generative models for graph generation, yet their use for conditional graph generation remains a fundamental challenge. In particular, guiding diffusion models on graphs under arbitrary reward signals is difficult: gradient-based methods, while powerful, are often unsuitable due to the discrete and combinatorial natu
Bingrui Sima, Linhua Cong, Wenxuan Wang, Kun He
The emergence of Multimodal Large Language Models (MLRMs) has enabled sophisticated visual reasoning capabilities by integrating reinforcement learning and Chain-of-Thought (CoT) supervision. However, while these enhanced reasoning capabilities improve performance, they also introduce new and underexplored safety risks. In this work, we systematically invest
Pengfei Cao, Tianyi Men, Wencan Liu, Jingwen Zhang
Planning represents a fundamental capability of intelligent agents, requiring comprehensive environmental understanding, rigorous logical reasoning, and effective sequential decision-making. While Large Language Models (LLMs) have demonstrated remarkable performance on certain planning tasks, their broader application in this domain warrants systematic inves
Bahareh Tasdighi, Manuel Haussmann, Yi-Shan Wu, Andres R. Masegosa
Deep actor-critic algorithms have reached a level where they influence everyday life. They are a driving force behind continual improvement of large language models through user feedback. However, their deployment in physical systems is not yet widely adopted, mainly because no validation scheme fully quantifies their risk of malfunction. We demonstrate that
Coherent Control of Ion-Photoelectron Dynamics through Rabi Oscillations: An ab initio study
physics.atom-phBo-Ren Shen, Yi-Jia Mao, Zhao-Han Zhang, Yang Li
We present first-principles numerical simulations of photoionization in neon induced by bichromatic extreme ultraviolet pulses with frequencies $\omega$ and $2\omega$, specially chosen to make $\omega$ equal to the energy difference between the $2s$ and $2p$ subshells. This allows for the production of photoelectrons from the $2s$ shell by $2\omega$ pulse an
Xinrui Wang, Shao-yuan Li, Jiaqiang Zhang, Songcan Chen
Multi-Label Online Continual Learning (MOCL) requires models to learn continuously from endless multi-label data streams, facing complex challenges including persistent catastrophic forgetting, potential missing labels, and uncontrollable imbalanced class distributions. While existing MOCL methods attempt to address these challenges through various technique
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
cs.CRBaolin Zheng, Guanlin Chen, Hongqiong Zhong, Qingyang Teng
Despite their remarkable achievements and widespread adoption, Multimodal Large Language Models (MLLMs) have revealed significant security vulnerabilities, highlighting the urgent need for robust safety evaluation benchmarks. Existing MLLM safety benchmarks, however, fall short in terms of data quality and coverge, and modal risk combinations, resulting in i
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
cs.CLZhaolin Li, Yining Liu, Danni Liu, Tuan Nam Nguyen
This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-tr
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
cs.CLHao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao
Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding the visual information during decoding. To alleviate this i
Chengcheng Dong, Yuefeng Yang, Changchang Dong
For a graph $\Gamma=(V\Gamma,E\Gamma)$, a subset $D$ of $V\Gamma$ is a perfect code in $\Gamma$ if every vertex of $\Gamma$ is dominated by exactly one vertex in $D$. In this paper, we classify all connected quartic Cayley graphs on generalized dihedral groups admitting a perfect code, and determine all perfect codes in such graphs.
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
cs.AILachlan McGinness, Peter Baumgartner
Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of State of the Art models from December 2023 and August 2024 on PRONTOQA steamroller reasoning problems. For that, we develop methods for assessing LLM response accuracy and correct a
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
cs.CLLiqin Ye, Agam Shah, Chao Zhang, Sudheer Chava
The traditional process of creating labeled datasets is labor-intensive and expensive. Recent breakthroughs in open-source large language models (LLMs) have opened up a new avenue in generating labeled datasets automatically for various natural language processing (NLP) tasks, providing an alternative to such an expensive annotation process. However, the rel
Alfio Bonanno, Samuele Silveravalle
A quantum ghost that destabilizes the Schwarzschild solution, transforming it into a naked singularity, may seem like a physicist's worst nightmare. However, we argue that this scenario represents the natural evolution of a black hole within a conservative high-energy gravity framework and may, in fact, be a desirable outcome. Quadratic curvature terms typic
Chaoyi Xiang, Chunhua Liu, Simon De Deyne, Lea Frermann
As the impact of large language models increases, understanding the moral values they reflect becomes ever more important. Assessing the nature of moral values as understood by these models via direct prompting is challenging due to potential leakage of human norms into model training data, and their sensitivity to prompt formulation. Instead, we propose to
Ziyang Yu, Zhejunyu Jin, Qianjun Zheng, Peng Yan
Phononic frequency combs (PFCs) typically require nonlinear elastic media, limiting their frequency range and stability. Here, we propose a transformative approach to generate PFCs in purely linear elastic media by harnessing the magnon nonlinearities, offering a new paradigm for frequency comb physics. By tuning the magnon-phonon coupling confined in a magn
Belcour Laurent, Fichet Alban, Barla Pascal
Fluorescent materials are characterized by a spectral reradiation toward longer wavelengths. Recent work [Fichet et al. 2024] has shown that the rendering of fluorescence in a non-spectral engine is possible through the use of appropriate reduced reradiation matrices. But the approach has limited expressivity, as it requires the storage of one reduced matrix
Bowen Zhang, Nur Afiqah Abdul Latiff, Justin Kan, Rong Tong
Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ASR model, an objective metrics extraction stage, and a generative pre-trained transformer (GPT) networ
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
cs.CLHao Yang, Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari
Large Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions. However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safety-alignment. Despite advances in defence measures for text and vision LLMs, effective safety-alignment str
Haiyang Sun, Shujie Hu, Shujie Liu, Lingwei Meng
Zero-shot streaming text-to-speech is an important research topic in human-computer interaction. Existing methods primarily use a lookahead mechanism, relying on future text to achieve natural streaming speech synthesis, which introduces high processing latency. To address this issue, we propose SMLLE, a streaming framework for generating high-quality speech
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
cs.CVTengda Huang, Yu Zhang, Tianren Li, Yufu Qu
Multi-image super-resolution (MISR) can achieve higher image quality than single-image super-resolution (SISR) by aggregating sub-pixel information from multiple spatially shifted frames. Among MISR tasks, burst super-resolution (BurstSR) has gained significant attention due to its wide range of applications. Recent methods have increasingly adopted Transfor
Weikang Yuan, Kaisong Song, Zhuoren Jiang, Junjie Cao
Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of professionals. While recent advances in Large Language Models (LLMs) offer a promising path toward scalable, low-cost legal assistance, current systems fall short in handling the int
Hasan Al-Nashash, Jiajin Wei, Ke Yang, Ayman Alzaatreh
Sample size calculation is crucial in biomedical in vivo research investigations mainly for two reasons: to design the most resource-efficient studies and to safeguard ethical issues when alive animals are subjects of testing. In this context, power analysis has been widely applied to compute the sample size by predetermining the desired statistical power an
Andrea Cavagna, Guido Cimino, Javier Cristín, Matteo Fiorini
Collective turns in starling flocks propagate linearly with negligible attenuation, indicating the existence of an underdamped sector in the dispersion relation. Beside granting linear propagation of the phase perturbations, the real part of the frequency should also yield a spin-wave form of the unperturbed correlation function. However, new high-resolution
Zhongyuan Cao, Mathieu Laurière
Motivated by recent interest in graphon mean field games and their applications, this paper provides a comprehensive probabilistic analysis of graphon mean field control (GMFC) problems, where the controlled dynamics are governed by a graphon mean field stochastic differential equation with heterogeneous mean field interactions. We formulate the GMFC problem
A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
cs.SDYigitcan Özer, Woosung Choi, Joan Serrà, Mayank Kumar Singh
We introduce the Robust Audio Watermarking Benchmark (RAW-Bench), a benchmark for evaluating deep learning-based audio watermarking methods with standardized and systematic comparisons. To simulate real-world usage, we introduce a comprehensive audio attack pipeline with various distortions such as compression, background noise, and reverberation, along with
Tingjia Shen, Hao Wang, Chuan Qin, Ruijun Sun
Open-domain question answering (OpenQA) represents a cornerstone in natural language processing (NLP), primarily focused on extracting answers from unstructured textual data. With the rapid advancements in Large Language Models (LLMs), LLM-based OpenQA methods have reaped the benefits of emergent understanding and answering capabilities enabled by massive pa
LangDAug: Langevin Data Augmentation for Multi-Source Domain Generalization in Medical Image Segmentation
cs.CVPiyush Tiwary, Kinjawl Bhattacharyya, Prathosh A. P
Medical image segmentation models often struggle to generalize across different domains due to various reasons. Domain Generalization (DG) methods overcome this either through representation learning or data augmentation (DAug). While representation learning methods seek domain-invariant features, they often rely on ad-hoc techniques and lack formal guarante
Ali Nouri, Beatriz Cabrero-Daniel, Zhennan Fei, Krishna Ronanki
Software engineers in various industrial domains are already using Large Language Models (LLMs) to accelerate the process of implementing parts of software systems. When considering its potential use for ADAS or AD systems in the automotive context, there is a need to systematically assess this new setup: LLMs entail a well-documented set of risks for safety
Luqi Wang, Yan Ning, Hongming Chen, Peize Liu
Multirotors are usually desired to enter confined narrow tunnels that are barely accessible to humans in various applications including inspection, search and rescue, and so on. This task is extremely challenging since the lack of geometric features and illuminations, together with the limited field of view, cause problems in perception; the restricted space
Tianren Ma, Xiaosong Zhang, Boyu Yang, Junlan Feng
In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous counterparts, which we attribute to noise (absorbing state) design and sampling heuristics. In this study, we propose a rehashing noise approach for discrete diffusion transformer
Jiaxin He, Qinfeng Li, Juncheng Wei, Hang Yang
Via continuous deformations based on natural flow evolutions, we prove several novel monotonicity results for Riesz-type nonlocal energies on triangles and quadrilaterals. Some of these results imply new and simpler proofs for known theorems without relying on any symmetrization arguments.
Aishik Chattopadhyay
We obtain nontrivial bounds for multiplicative character sums over codimension-one sublattices of finite field extensions $\mathbb{F}_{p^d}$. This extends earlier results of Davenport--Lewis and Chang from full-dimensional settings to the codimension-one setting. As an application, we establish cancellation in character sums over binary cubic forms for inter
Ning Yang, Hai Lin, Yibo Liu, Baoliang Tian
Aligning Large Language Models (LLMs) with human preferences is crucial for safe and effective AI interactions. While popular methods like Direct Preference Optimization (DPO) have simplified alignment, they remain sensitive to data noise and overlook the differential importance of individual tokens. Existing token-level approaches often rely on probability
Hongbin Wang, Zhihong Jia, Yuanzhong Shen, Ziwei Wang
Speech disorders such as dysarthria and anarthria can severely impair the patient's ability to communicate verbally. Speech decoding brain-computer interfaces (BCIs) offer a potential alternative by directly translating speech intentions into spoken words, serving as speech neuroprostheses. This paper reports an experimental protocol for Mandarin Chinese spe
Greybody factors as robust gravitational observables: insights into post-merger signals and echoes from ultracompact object
gr-qcRomeo Felice Rosato
The quasinormal mode spectrum plays a central role in modeling the post-merger ringdown phase of binary coalescences of compact objects. However, its interpretation is subject to certain ambiguities. Motivated by a recently discovered connection between greybody factors and post-merger black hole signals, we investigate the robustness of greybody factors as
Fanheng Kong, Jingyuan Zhang, Yahui Liu, Hongzhi Zhang
Multimodal information retrieval (MIR) faces inherent challenges due to the heterogeneity of data sources and the complexity of cross-modal alignment. While previous studies have identified modal gaps in feature spaces, a systematic approach to address these challenges remains unexplored. In this work, we introduce UNITE, a universal framework that tackles t
Santiago Radi
We prove that the iterated monodromy group of the polynomial $z^2+i$ is just-infinite, regular branch and does not have the congruence subgroup property. This yields the first example of an iterated monodromy group of a polynomial with these properties. Additional information is provided about the congruence kernel, rigid kernel and branch kernel of this gro
Qiaolan Meng, Juhua Pu, Hongting Niu, Yuyi Wang
We study the model enumeration problem of the function-free, finite domain fragment of first-order logic with two variables ($FO^2$). Specifically, given an $FO^2$ sentence $\Gamma$ and a positive integer $n$, how can one enumerate all the models of $\Gamma$ over a domain of size $n$? In this paper, we devise a novel algorithm to address this problem. The de
Xiaochuan Liu, Ruihua Song, Xiting Wang, Xu Chen
Automatic related work generation (RWG) can save people's time and effort when writing a draft of related work section (RWS) for further revision. However, existing methods for RWG always suffer from shallow comprehension due to taking the limited portions of references papers as input and isolated explanation for each reference due to ineffective capturing
Dongyeop Woo, Minsu Kim, Minkyu Kim, Kiyoung Seong
We propose Energy-based generator matching (EGM), a modality-agnostic approach to train generative models from energy functions in the absence of data. Extending the recently proposed generator matching, EGM enables training of arbitrary continuous-time Markov processes, e.g., diffusion, flow, and jump, and can generate data from continuous, discrete, and a
Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu
Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional dense models, MoEs achieve better performance with less computation. Speculative decoding (SD) is a widely used technique to accelerate LLM inference without accuracy loss, but it
Lijun Zhang, Lin Li, Yajie Qi, Huizhong Song
When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typically introduce KL divergence to constrain deviations between th
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
cs.SDAnton Firc, Manasi Chhibber, Jagabandhu Mishra, Vishwanath Pratap Singh
A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introdu
Elena Fernandez, Sandi Klavzar, Dorota Kuziak, Manuel Muñoz-Marquez
Given a connected graph $G$, a set of vertices $X\subset V(G)$ is a weak $k$-resolving set of $G$ if for each two vertices $y,z\in V(G)$, the sum of the values $|d_G(y,x)-d_G(z,x)|$ over all $x\in X$ is at least $k$, where $d_G(u,v)$ stands for the length of a shortest path between $u$ and $v$. The cardinality of a smallest weak $k$-resolving set of $G$ is t
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
cs.AIJunteng Liu, Yuanxiang Fan, Zhuo Jiang, Han Ding
Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplor
Roy Xie, David Qiu, Deepak Gopinath, Dong Lin
Long chain-of-thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs). However, extensive reasoning traces lead to inefficiencies and increased time-to-first-token (TTFT). We propose a training paradigm that uses only reinforcement learning (RL) to guide reasoning LLMs to interleave thinking and answering for multi-hop
Jiabao He, Yueyue Xu, Yue Ju, Cristian R. Rojas
This contribution revisits the classical approximate realization problem, which involves determining matrices of a state-space model based on estimates of a truncated series of Markov parameters. A Hankel matrix built up by these Markov parameters plays a fundamental role in this problem, leveraging the fact that both its range space and left null space enco
Ming Meng, Qi Dong, Jiajie Li, Zhe Zhu
Virtual try-on technology has become increasingly important in the fashion and retail industries, enabling the generation of high-fidelity garment images that adapt seamlessly to target human models. While existing methods have achieved notable progress, they still face significant challenges in maintaining consistency across different poses. Specifically, g
Byunghyun Yoo, Younghwan Shin, Hyunwoo Kim, Euisok Chung
In standard reinforcement learning, an episode is defined as a sequence of interactions between agents and the environment, which terminates upon reaching a terminal state or a pre-defined episode length. Setting a shorter episode length enables the generation of multiple episodes with the same number of data samples, thereby facilitating an exploration of d
Longitudinal magnetoconductivity in chiral multifold semimetals exemplified by pseudospin-1 nodal points
cond-mat.mes-hallIpsita Mandal
We embark on computing the longitudinal magnetoconductivity within the semiclassical Boltzmann formalism, where an isotropic triple-point semimetal (TSM) is subjected to collinear electric ($\boldsymbol E $) and magnetic ($\boldsymbol B$) fields. Except for the Drude part, the $B$-dependence arises exclusively from topological properties like the Berry curva
Ivan Y. Tyukin, Bogdan Grechuk, Evgeny M. Mirkes, Alexander N. Gorban
Concentration of distances in high dimension is an important factor for the development and design of stable and reliable data analysis algorithms. In this paper, we address the fundamental long-standing question about the concentration of distances in high dimension for fractional quasi $p$-norms, $p\in(0,1)$. The topic has been at the centre of various the
Zili Wang, Tianyu Zhang, Haoli Bai, Lu Hou
Test-Time Scaling (TTS) has proven effective in improving the performance of Large Language Models (LLMs) during inference. However, existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. Through a latency-aware evaluation of representative TTS methods, we demonstrate that a compute-optimal TTS does not always result in
Martijn Hanegraaf, Savio Sciancalepore, Gabriele Oligeri
State-of-the-art solutions detect jamming attacks ex-post, i.e., only when jamming has already disrupted the wireless communication link. In many scenarios, e.g., mobile networks or static deployments distributed over a large geographical area, it is often desired to detect jamming at the early stage, when it affects the communication link enough to be detec
M. V. Takook
Krein space quantization and the ambient space formalism have been successfully applied to address challenges in quantum geometry (e.g., quantum gravity) and the axiomatic formulation of quantum Yang-Mills theory, including phenomena such as color confinement and the mass gap. Building on these advancements, we aim to extend these methods to develop novel qu
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
cs.CLZihong Zhang, Liqi He, Zuchao Li, Lefei Zhang
Word segmentation stands as a cornerstone of Natural Language Processing (NLP). Based on the concept of "comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs based on word segmentation. We employ curren
Yichun Feng, Jiawei Wang, Lu Zhou, Yikai Zheng
Large language models (LLMs) struggle in real-world clinical consultations. Single-turn consultation systems require patients to describe all symptoms at once, which often leads to unclear complaints and vague diagnoses. Traditional dialogue models, constrained by static supervised learning, are limited to superficially imitating existing dialogue patterns a
Hassan Sartaj, Shaukat Ali, Ana Cavalcanti, Lukas Esterle
Self-adaptive robotic systems operate autonomously in dynamic and uncertain environments, requiring robust real-time monitoring and adaptive behaviour. Unlike traditional robotic software with predefined logic, self-adaptive robots exploit artificial intelligence (AI), machine learning, and model-driven engineering to adapt continuously to changing condition
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices
cs.CLSilin Li, Yuhang Guo, Jiashu Yao, Zeming Liu
Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment. While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling st
In-depth Investigation of Conduction Mechanism on Defect-induced Proton-conducting Electrolytes BaHfO$_3$
cond-mat.mtrl-sciPeng Feng, Hang Ma, Kuan Yang, Yingjie Lv
This study utilizes first-principles computational methods to comprehensively analyze the impact of A-site doping on the proton conduction properties of BaHfO$_3$. The goal is to offer theoretical support for the advancement of electrolyte materials for solid oxide fuel cells. Our research has uncovered that BaHfO$_3$ demonstrates promising potential for pro
Jiaxin Chen, Yiming Wang, Ziyu Zhang, Jiayang Han
The same speech content produced by different speakers exhibits significant differences in pitch contour, yet listeners' semantic perception remains unaffected. This phenomenon may stem from the brain's perception of pitch contours being independent of individual speakers' pitch ranges. In this work, we recorded electroencephalogram (EEG) while participants
Search-Based Software Engineering and AI Foundation Models: Current Landscape and Future Roadmap
cs.SEHassan Sartaj, Shaukat Ali, Paolo Arcaini, Andrea Arcuri
Search-based software engineering (SBSE), which integrates metaheuristic search techniques with software engineering, has been an active area of research for about 25 years. It has been applied to solve numerous problems across the entire software engineering lifecycle and has demonstrated its versatility in multiple domains. With recent advances in Artifici
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat
cs.CVPusheng Xu, Xia Gong, Xiaolan Chen, Weiyi Zhang
Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions published between January 1, 2016, and December 31, 2024, were collected from WeChat Official Accounts. Based on these captions, bilingual question-answer (QA) pairs in Chinese and
Yu Shang, Peijie Liu, Yuwei Yan, Zijing Wu
The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approaches, agentic recommender systems can dynamically gather and
Emmanuel Humbert, Kilian Raschel
We consider multidimensional random walks in pyramidal cones (or multidimensional orthants), which are intersections of a finite number of half-spaces. We explore the connection between the existence of (positive) discrete harmonic polynomials for the random walks, with Dirichlet conditions on the boundary of the cone, and geometric properties of the cone, b
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
cs.AIGeorge Kour, Itay Nakash, Ateret Anaby-Tavor, Michal Shmueli-Scheuer
As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibit subjective preferences, opinions, and beliefs. These tendencies may stem from biases within the models, which may shape their behavior, influence the advice and recommendations t
Is GGAG:Ce@SiO$_2$-RB composite a prospective material for X-ray induced photodynamic therapy?
physics.med-phIveta Terezie Hošnová, Kristýna Havlinová, Jan Bárta, Karolína Mocová
Nanocomposite material ($\mathrm{GGAG:Ce^{3+}@SiO_2-RB}$) for potential use in X-ray induced photodynamic therapy (X-PDT) was developed, thoroughly characterized, and evaluated. It consists of a scintillating $\mathrm{Gd_3(Ga_{1-x}Al_x)_5O_{12}:Ce^{3+}}$ core encapsulated in silica layer and functionalized with the photosensitizer Rose Bengal (RB). Radiolumi
Jiawen Chen, Qi Shao, Duxin Chen, Wenwu Yu
Spatio-temporal prediction is a pivotal task with broad applications in traffic management, climate monitoring, energy scheduling, etc. However, existing methodologies often struggle to balance model expressiveness and computational efficiency, especially when scaling to large real-world datasets. To tackle these challenges, we propose STH-SepNet (Spatio-Tem
Jun Tian, He Wang, Jibo He, Yu Pan
Convolutional neural networks (CNNs) have become widely adopted in gravitational wave (GW) detection pipelines due to their ability to automatically learn hierarchical features from raw strain data. However, the physical meaning of these learned features remains underexplored, limiting the interpretability of such models. In this work, we propose a hybrid ar
Hanze Liu, Jiahong Fu, Qi Xie, Deyu Meng
Self-supervised image denoising methods have garnered significant research attention in recent years, for this kind of method reduces the requirement of large training datasets. Compared to supervised methods, self-supervised methods rely more on the prior embedded in deep networks themselves. As a result, most of the self-supervised methods are designed wit
Hybrid Models for Financial Forecasting: Combining Econometric, Machine Learning, and Deep Learning Models
q-fin.TRDominik Stempień, Robert Ślepaczuk
This research systematically develops and evaluates various hybrid modeling approaches by combining traditional econometric models (ARIMA and ARFIMA models) with machine learning and deep learning techniques (SVM, XGBoost, and LSTM models) to forecast financial time series. The empirical analysis is based on two distinct financial assets: the S&P 500 index a
Rui Cai, Bangzheng Li, Xiaofei Wen, Muhao Chen
Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as irrelevant text in vision understanding, or irrelevant visual content in question answering. At its core, modality interference refers to cases where spurious signals from non-esse
$b$-axis and $c$-axis Knight shift measurements in the superconducting state on ultraclean UTe$_2$ with $T_c$ = 2.1 K
cond-mat.supr-conHiroki Matsumura, Yuki Takahashi, Katsuki Kinjo, Shunsaku Kitagawa
Knight shifts along the $b$ and $c$ axes ($K_b$ and $K_c$) at two crystallographically distinct Te sites were measured down to 70 mK using $^{125}$Te nuclear magnetic resonance (NMR) on an ultraclean UTe$_2$ single crystal with a superconducting (SC) transition temperature $T_{\mathrm{c}}$ = 2.1 K. This was carried out to determine the $\boldsymbol{d}$-vecto
TESSER: Transfer-Enhancing Adversarial Attacks from Vision Transformers via Spectral and Semantic Regularization
cs.CVAmira Guesmi, Bassem Ouni, Muhammad Shafique
Adversarial transferability remains a critical challenge in evaluating the robustness of deep neural networks. In security-critical applications, transferability enables black-box attacks without access to model internals, making it a key concern for real-world adversarial threat assessment. While Vision Transformers (ViTs) have demonstrated strong adversari
Yui Tatsumi, Ziyue Zeng, Hiroshi Watanabe
Conventional methods for scalable image coding for humans and machines require the transmission of additional information to achieve scalability. A recent diffusion-based approach avoids this by generating human-oriented images from machine-oriented images without extra bitrate. However, it utilizes a single random seed, which may lead to suboptimal image qu
Optimal Intervention for Self-triggering Spatial Networks with Application to Urban Crime Analytics
cs.SIPramit Das, Moulinath Banerjee, Yuekai Sun
In many network systems, events at one node trigger further activity at other nodes, e.g., social media users reacting to each other's posts or the clustering of criminal activity in urban environments. These systems are typically referred to as self-exciting networks. In such systems, targeted intervention at critical nodes can be an effective strategy for
Ruolin Shen, Xiaozhong Ji, Kai WU, Jiangning Zhang
Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish concealed objects, demonstrating an inability to emulate human cognitive processes which effectively utilize foreground-back
Ziga Kovacic, Justin T. Chiu, Celine Lee, Wenting Zhao
Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into reusable components. This challenge becomes particularly relevant as code agents become used to solve isolated one-off programming problems. We investigate code agents' capacity to r
Jiaxin Song, Yixu Wang, Jie Li, Rui Yu
Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local o
Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang
Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution pos
Davide Parodi, Federico Benvenuto, Sara Garbarino, Michele Piana
Implicit inverse problems, in which noisy observations of a physical quantity are used to infer a nonlinear functional applied to an associated function, are inherently ill posed and often exhibit non uniqueness of solutions. Such problems arise in a range of domains, including the identification of systems governed by Ordinary and Partial Differential Equat