Skip to content

May 2025 arXiv papers — page 58

Showing 5,7015,800 of 24,552 papers

  1. Rong-Cheng Tu, Zhao Jin, Jingyi Liao, Xiao Luo

    Existing Zero-Shot Composed Image Retrieval (ZS-CIR) methods typically train adapters that convert reference images into pseudo-text tokens, which are concatenated with the modifying text and processed by frozen text encoders in pretrained VLMs or LLMs. While this design leverages the strengths of large pretrained models, it only supervises the adapter to pr

  2. Tej Deep Pala, Panshul Sharma, Amir Zadeh, Chuan Li

    Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Reward Models (PRMs) score each intermediate step to steer generation toward coherent solutions. We introduce PathFinder-PRM, a novel hierarchic

  3. Chuanxing Wang, Hui Luo, Kai Wang, Guohuai Zhu

    Despite the remarkable progress of physics-informed neural networks (PINNs) in scientific computing, they continue to face challenges when solving hydrodynamic problems with multiple discontinuities. In this work, we propose Separation-Transfer Physics Informed Neural Networks (ST-PINNs) to address such problems. By sequentially resolving discontinuities fro

  4. Kaizhe Chen, Heng Zhang

    We derive some existence results for the solutions of the Tzitz\'eica equation \begin{equation*} -\Delta u + h_1(x)e^{Au} + h_2(x)e^{-Bu}=0 \end{equation*} and the generalized Tzitz\'eica equation \begin{equation*} -\Delta u + h_1(x)e^{Au}(e^{Au}-1)+h_2(x)e^{-Bu}(e^{-Bu}-1)=0 \end{equation*} on any connected finite graph \(G=(V, E)\). Here, \(h_1(x)>0\), \(h

  5. Tao Han, Shaoyuan Li, Xiang Yin

    This paper investigates the online monitoring problem for cyber-physical systems under signal temporal logic (STL) specifications. The objective is to design an online monitor that evaluates system correctness at runtime based on partial signal observations up to the current time so that alarms can be issued whenever the specification is violated or will ine

  6. Minheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin

    Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains challenging due to inherent limitations in text-only CoT, such as visual hallucinations and insufficient multimodal integrati

  7. Kodai Hasegawa, Shigeaki Okumura, Hirofumi Taki, Hironobu Sunadome

    Radar-based respiratory measurement is a promising tool for the noncontact detection of sleep apnea. Our team has reported that apnea events can be accurately detected using the statistical characteristics of the amplitude of respiratory displacement. However, apnea and hypopnea events are often followed by irregular breathing, reducing the detection accurac

  8. Yi Liu, Dianqing Liu, Mingye Zhu, Junbo Guo

    The widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse applications. To address this limitation, we propose a novel \

  9. Jing Yu Lim, Rushi Shah, Zarif Ikram, Samson Yu

    Recently, Model-Based Reinforcement Learning (MBRL) have achieved super-human level performance on the Atari100k benchmark on average. However, we discover that conventional aggregates mask a major problem, Performance Asymmetry: MBRL agents dramatically outperform humans in certain tasks (Agent-Optimal tasks) while drastically underperform humans in other t

  10. Swati Saha, Ranbir Singh, Bedangadas Mohanty

    The transverse momentum-differential radial flow observable $v_0(p_\mathrm{T})$, recently proposed and measured by the ATLAS and ALICE collaborations, provides a novel tool to probe radial expansion dynamics in high-energy heavy-ion collisions. In this work, we conduct a detailed study of $v_0(p_\mathrm{T})$ using a blast-wave model that incorporates hydrody

  11. Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Nour Aburaed

    This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensitive human judgments to a single scalar, obscuring semantic failures, user intent, and the rationale behind quality decisions. We contend tha

  12. Galina Weinstein

    This paper critically examines the central thesis of Kieran Fox's "I Am a Part of Infinity: The Spiritual Journey of Albert Einstein"-namely, that Einstein's intellectual development constitutes a coherent spiritual path culminating in a form of pantheistic mysticism shaped by both Western and Eastern traditions. Fox presents Einstein as the modern heir to a

  13. Wen Yin, Yong Wang, Guiduo Duan, Dongyang Zhang

    Visual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models' cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recogni

  14. Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee

    Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model that integrates spherical VAD region classification to guide VAD regression for improved emotion prediction. In our framework, VAD values are

  15. Wenchao Sun, Xuewu Lin, Keyu Chen, Zixiang Pei

    Camera sensor simulation serves as a critical role for autonomous driving (AD), e.g. evaluating vision-based AD algorithms. While existing approaches have leveraged generative models for controllable image/video generation, they remain constrained to generating multi-view video sequences with fixed camera viewpoints and video frequency, significantly limitin

  16. Chon Man Sou

    Since the original derivation of Hawking radiation, there have been lots of alternative approaches to show the same fact that black holes emit particles as hot bodies with a temperature. These alternative methods generally rely on different conditions and physical quantities to manifest the radiation, providing various points of view of this effect in the in

  17. Baihui Zheng, Boren Zheng, Kerui Cao, Yingshui Tan

    Despite the remarkable proficiency of \textit{Large Reasoning Models} (LRMs) in handling complex reasoning tasks, their reliability in safety-critical scenarios remains uncertain. Existing evaluations primarily assess response-level safety, neglecting a critical issue we identify as \textbf{\textit{Superficial Safety Alignment} (SSA)} -- a phenomenon where m

  18. ATLAS Collaboration

    Jet flavour tagging enables the identification of jets originating from heavy-flavour quarks in proton-proton collisions at the Large Hadron Collider, playing a critical role in its physics programmes. This paper presents GN2, a transformer-based flavour tagging algorithm deployed by the ATLAS Collaboration that represents a different methodology compared to

  19. Yuhe Gong, Riddhiman Laha, Luis Figueredo

    Reactive intelligence remains one of the cornerstones of versatile robotics operating in cluttered, dynamic, and human-centred environments. Among reactive approaches, potential fields (PF) continue to be widely adopted due to their simplicity and real-time applicability. However, existing PF methods typically oversimplify environmental representations by re

  20. Deok-Hyeon Cho, Hyung-Seok Oh, Seung-Bin Kim, Seong-Whan Lee

    Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully separate speaker and emotion characteristics, causing speaker leakage and degraded synthesis quality. To address this, we prop

  21. Victor M. Tenorio, Nicolas Zilberstein, Santiago Segarra, Antonio G. Marques

    Diffusion models have emerged as powerful generative models for graph generation, yet their use for conditional graph generation remains a fundamental challenge. In particular, guiding diffusion models on graphs under arbitrary reward signals is difficult: gradient-based methods, while powerful, are often unsuitable due to the discrete and combinatorial natu

  22. Bingrui Sima, Linhua Cong, Wenxuan Wang, Kun He

    The emergence of Multimodal Large Language Models (MLRMs) has enabled sophisticated visual reasoning capabilities by integrating reinforcement learning and Chain-of-Thought (CoT) supervision. However, while these enhanced reasoning capabilities improve performance, they also introduce new and underexplored safety risks. In this work, we systematically invest

  23. Pengfei Cao, Tianyi Men, Wencan Liu, Jingwen Zhang

    Planning represents a fundamental capability of intelligent agents, requiring comprehensive environmental understanding, rigorous logical reasoning, and effective sequential decision-making. While Large Language Models (LLMs) have demonstrated remarkable performance on certain planning tasks, their broader application in this domain warrants systematic inves

  24. Bahareh Tasdighi, Manuel Haussmann, Yi-Shan Wu, Andres R. Masegosa

    Deep actor-critic algorithms have reached a level where they influence everyday life. They are a driving force behind continual improvement of large language models through user feedback. However, their deployment in physical systems is not yet widely adopted, mainly because no validation scheme fully quantifies their risk of malfunction. We demonstrate that

  25. Bo-Ren Shen, Yi-Jia Mao, Zhao-Han Zhang, Yang Li

    We present first-principles numerical simulations of photoionization in neon induced by bichromatic extreme ultraviolet pulses with frequencies $\omega$ and $2\omega$, specially chosen to make $\omega$ equal to the energy difference between the $2s$ and $2p$ subshells. This allows for the production of photoelectrons from the $2s$ shell by $2\omega$ pulse an

  26. Xinrui Wang, Shao-yuan Li, Jiaqiang Zhang, Songcan Chen

    Multi-Label Online Continual Learning (MOCL) requires models to learn continuously from endless multi-label data streams, facing complex challenges including persistent catastrophic forgetting, potential missing labels, and uncontrollable imbalanced class distributions. While existing MOCL methods attempt to address these challenges through various technique

  27. Baolin Zheng, Guanlin Chen, Hongqiong Zhong, Qingyang Teng

    Despite their remarkable achievements and widespread adoption, Multimodal Large Language Models (MLLMs) have revealed significant security vulnerabilities, highlighting the urgent need for robust safety evaluation benchmarks. Existing MLLM safety benchmarks, however, fall short in terms of data quality and coverge, and modal risk combinations, resulting in i

  28. Zhaolin Li, Yining Liu, Danni Liu, Tuan Nam Nguyen

    This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-tr

  29. Hao Fang, Changle Zhou, Jiawei Kong, Kuofeng Gao

    Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding the visual information during decoding. To alleviate this i

  30. Chengcheng Dong, Yuefeng Yang, Changchang Dong

    For a graph $\Gamma=(V\Gamma,E\Gamma)$, a subset $D$ of $V\Gamma$ is a perfect code in $\Gamma$ if every vertex of $\Gamma$ is dominated by exactly one vertex in $D$. In this paper, we classify all connected quartic Cayley graphs on generalized dihedral groups admitting a perfect code, and determine all perfect codes in such graphs.

  31. Lachlan McGinness, Peter Baumgartner

    Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of State of the Art models from December 2023 and August 2024 on PRONTOQA steamroller reasoning problems. For that, we develop methods for assessing LLM response accuracy and correct a

  32. Liqin Ye, Agam Shah, Chao Zhang, Sudheer Chava

    The traditional process of creating labeled datasets is labor-intensive and expensive. Recent breakthroughs in open-source large language models (LLMs) have opened up a new avenue in generating labeled datasets automatically for various natural language processing (NLP) tasks, providing an alternative to such an expensive annotation process. However, the rel

  33. Alfio Bonanno, Samuele Silveravalle

    A quantum ghost that destabilizes the Schwarzschild solution, transforming it into a naked singularity, may seem like a physicist's worst nightmare. However, we argue that this scenario represents the natural evolution of a black hole within a conservative high-energy gravity framework and may, in fact, be a desirable outcome. Quadratic curvature terms typic

  34. Chaoyi Xiang, Chunhua Liu, Simon De Deyne, Lea Frermann

    As the impact of large language models increases, understanding the moral values they reflect becomes ever more important. Assessing the nature of moral values as understood by these models via direct prompting is challenging due to potential leakage of human norms into model training data, and their sensitivity to prompt formulation. Instead, we propose to

  35. Ziyang Yu, Zhejunyu Jin, Qianjun Zheng, Peng Yan

    Phononic frequency combs (PFCs) typically require nonlinear elastic media, limiting their frequency range and stability. Here, we propose a transformative approach to generate PFCs in purely linear elastic media by harnessing the magnon nonlinearities, offering a new paradigm for frequency comb physics. By tuning the magnon-phonon coupling confined in a magn

  36. Belcour Laurent, Fichet Alban, Barla Pascal

    Fluorescent materials are characterized by a spectral reradiation toward longer wavelengths. Recent work [Fichet et al. 2024] has shown that the rendering of fluorescence in a non-spectral engine is possible through the use of appropriate reduced reradiation matrices. But the approach has limited expressivity, as it requires the storage of one reduced matrix

  37. Bowen Zhang, Nur Afiqah Abdul Latiff, Justin Kan, Rong Tong

    Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ASR model, an objective metrics extraction stage, and a generative pre-trained transformer (GPT) networ

  38. Hao Yang, Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari

    Large Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions. However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safety-alignment. Despite advances in defence measures for text and vision LLMs, effective safety-alignment str

  39. Haiyang Sun, Shujie Hu, Shujie Liu, Lingwei Meng

    Zero-shot streaming text-to-speech is an important research topic in human-computer interaction. Existing methods primarily use a lookahead mechanism, relying on future text to achieve natural streaming speech synthesis, which introduces high processing latency. To address this issue, we propose SMLLE, a streaming framework for generating high-quality speech

  40. Tengda Huang, Yu Zhang, Tianren Li, Yufu Qu

    Multi-image super-resolution (MISR) can achieve higher image quality than single-image super-resolution (SISR) by aggregating sub-pixel information from multiple spatially shifted frames. Among MISR tasks, burst super-resolution (BurstSR) has gained significant attention due to its wide range of applications. Recent methods have increasingly adopted Transfor

  41. Weikang Yuan, Kaisong Song, Zhuoren Jiang, Junjie Cao

    Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of professionals. While recent advances in Large Language Models (LLMs) offer a promising path toward scalable, low-cost legal assistance, current systems fall short in handling the int

  42. Hasan Al-Nashash, Jiajin Wei, Ke Yang, Ayman Alzaatreh

    Sample size calculation is crucial in biomedical in vivo research investigations mainly for two reasons: to design the most resource-efficient studies and to safeguard ethical issues when alive animals are subjects of testing. In this context, power analysis has been widely applied to compute the sample size by predetermining the desired statistical power an

  43. Andrea Cavagna, Guido Cimino, Javier Cristín, Matteo Fiorini

    Collective turns in starling flocks propagate linearly with negligible attenuation, indicating the existence of an underdamped sector in the dispersion relation. Beside granting linear propagation of the phase perturbations, the real part of the frequency should also yield a spin-wave form of the unperturbed correlation function. However, new high-resolution

  44. Zhongyuan Cao, Mathieu Laurière

    Motivated by recent interest in graphon mean field games and their applications, this paper provides a comprehensive probabilistic analysis of graphon mean field control (GMFC) problems, where the controlled dynamics are governed by a graphon mean field stochastic differential equation with heterogeneous mean field interactions. We formulate the GMFC problem

  45. Yigitcan Özer, Woosung Choi, Joan Serrà, Mayank Kumar Singh

    We introduce the Robust Audio Watermarking Benchmark (RAW-Bench), a benchmark for evaluating deep learning-based audio watermarking methods with standardized and systematic comparisons. To simulate real-world usage, we introduce a comprehensive audio attack pipeline with various distortions such as compression, background noise, and reverberation, along with

  46. Tingjia Shen, Hao Wang, Chuan Qin, Ruijun Sun

    Open-domain question answering (OpenQA) represents a cornerstone in natural language processing (NLP), primarily focused on extracting answers from unstructured textual data. With the rapid advancements in Large Language Models (LLMs), LLM-based OpenQA methods have reaped the benefits of emergent understanding and answering capabilities enabled by massive pa

  47. Piyush Tiwary, Kinjawl Bhattacharyya, Prathosh A. P

    Medical image segmentation models often struggle to generalize across different domains due to various reasons. Domain Generalization (DG) methods overcome this either through representation learning or data augmentation (DAug). While representation learning methods seek domain-invariant features, they often rely on ad-hoc techniques and lack formal guarante

  48. Ali Nouri, Beatriz Cabrero-Daniel, Zhennan Fei, Krishna Ronanki

    Software engineers in various industrial domains are already using Large Language Models (LLMs) to accelerate the process of implementing parts of software systems. When considering its potential use for ADAS or AD systems in the automotive context, there is a need to systematically assess this new setup: LLMs entail a well-documented set of risks for safety

  49. Luqi Wang, Yan Ning, Hongming Chen, Peize Liu

    Multirotors are usually desired to enter confined narrow tunnels that are barely accessible to humans in various applications including inspection, search and rescue, and so on. This task is extremely challenging since the lack of geometric features and illuminations, together with the limited field of view, cause problems in perception; the restricted space

  50. Tianren Ma, Xiaosong Zhang, Boyu Yang, Junlan Feng

    In the visual generative area, discrete diffusion models are gaining traction for their efficiency and compatibility. However, pioneered attempts still fall behind their continuous counterparts, which we attribute to noise (absorbing state) design and sampling heuristics. In this study, we propose a rehashing noise approach for discrete diffusion transformer

  51. Jiaxin He, Qinfeng Li, Juncheng Wei, Hang Yang

    Via continuous deformations based on natural flow evolutions, we prove several novel monotonicity results for Riesz-type nonlocal energies on triangles and quadrilaterals. Some of these results imply new and simpler proofs for known theorems without relying on any symmetrization arguments.

  52. Aishik Chattopadhyay

    We obtain nontrivial bounds for multiplicative character sums over codimension-one sublattices of finite field extensions $\mathbb{F}_{p^d}$. This extends earlier results of Davenport--Lewis and Chang from full-dimensional settings to the codimension-one setting. As an application, we establish cancellation in character sums over binary cubic forms for inter

  53. Ning Yang, Hai Lin, Yibo Liu, Baoliang Tian

    Aligning Large Language Models (LLMs) with human preferences is crucial for safe and effective AI interactions. While popular methods like Direct Preference Optimization (DPO) have simplified alignment, they remain sensitive to data noise and overlook the differential importance of individual tokens. Existing token-level approaches often rely on probability

  54. Hongbin Wang, Zhihong Jia, Yuanzhong Shen, Ziwei Wang

    Speech disorders such as dysarthria and anarthria can severely impair the patient's ability to communicate verbally. Speech decoding brain-computer interfaces (BCIs) offer a potential alternative by directly translating speech intentions into spoken words, serving as speech neuroprostheses. This paper reports an experimental protocol for Mandarin Chinese spe

  55. Romeo Felice Rosato

    The quasinormal mode spectrum plays a central role in modeling the post-merger ringdown phase of binary coalescences of compact objects. However, its interpretation is subject to certain ambiguities. Motivated by a recently discovered connection between greybody factors and post-merger black hole signals, we investigate the robustness of greybody factors as

  56. Fanheng Kong, Jingyuan Zhang, Yahui Liu, Hongzhi Zhang

    Multimodal information retrieval (MIR) faces inherent challenges due to the heterogeneity of data sources and the complexity of cross-modal alignment. While previous studies have identified modal gaps in feature spaces, a systematic approach to address these challenges remains unexplored. In this work, we introduce UNITE, a universal framework that tackles t

  57. Santiago Radi

    We prove that the iterated monodromy group of the polynomial $z^2+i$ is just-infinite, regular branch and does not have the congruence subgroup property. This yields the first example of an iterated monodromy group of a polynomial with these properties. Additional information is provided about the congruence kernel, rigid kernel and branch kernel of this gro

  58. Qiaolan Meng, Juhua Pu, Hongting Niu, Yuyi Wang

    We study the model enumeration problem of the function-free, finite domain fragment of first-order logic with two variables ($FO^2$). Specifically, given an $FO^2$ sentence $\Gamma$ and a positive integer $n$, how can one enumerate all the models of $\Gamma$ over a domain of size $n$? In this paper, we devise a novel algorithm to address this problem. The de

  59. Xiaochuan Liu, Ruihua Song, Xiting Wang, Xu Chen

    Automatic related work generation (RWG) can save people's time and effort when writing a draft of related work section (RWS) for further revision. However, existing methods for RWG always suffer from shallow comprehension due to taking the limited portions of references papers as input and isolated explanation for each reference due to ineffective capturing

  60. Dongyeop Woo, Minsu Kim, Minkyu Kim, Kiyoung Seong

    We propose Energy-based generator matching (EGM), a modality-agnostic approach to train generative models from energy functions in the absence of data. Extending the recently proposed generator matching, EGM enables training of arbitrary continuous-time Markov processes, e.g., diffusion, flow, and jump, and can generate data from continuous, discrete, and a

  61. Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu

    Large Language Models (LLMs) have achieved remarkable success across many applications, with Mixture of Experts (MoE) models demonstrating great potential. Compared to traditional dense models, MoEs achieve better performance with less computation. Speculative decoding (SD) is a widely used technique to accelerate LLM inference without accuracy loss, but it

  62. Lijun Zhang, Lin Li, Yajie Qi, Huizhong Song

    When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typically introduce KL divergence to constrain deviations between th

  63. Anton Firc, Manasi Chhibber, Jagabandhu Mishra, Vishwanath Pratap Singh

    A key research area in deepfake speech detection is source tracing - determining the origin of synthesised utterances. The approaches may involve identifying the acoustic model (AM), vocoder model (VM), or other generation-specific parameters. However, progress is limited by the lack of a dedicated, systematically curated dataset. To address this, we introdu

  64. Elena Fernandez, Sandi Klavzar, Dorota Kuziak, Manuel Muñoz-Marquez

    Given a connected graph $G$, a set of vertices $X\subset V(G)$ is a weak $k$-resolving set of $G$ if for each two vertices $y,z\in V(G)$, the sum of the values $|d_G(y,x)-d_G(z,x)|$ over all $x\in X$ is at least $k$, where $d_G(u,v)$ stands for the length of a shortest path between $u$ and $v$. The cardinality of a smallest weak $k$-resolving set of $G$ is t

  65. Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Han Ding

    Recent advances such as OpenAI-o1 and DeepSeek R1 have demonstrated the potential of Reinforcement Learning (RL) to enhance reasoning abilities in Large Language Models (LLMs). While open-source replication efforts have primarily focused on mathematical and coding domains, methods and resources for developing general reasoning capabilities remain underexplor

  66. Roy Xie, David Qiu, Deepak Gopinath, Dong Lin

    Long chain-of-thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs). However, extensive reasoning traces lead to inefficiencies and increased time-to-first-token (TTFT). We propose a training paradigm that uses only reinforcement learning (RL) to guide reasoning LLMs to interleave thinking and answering for multi-hop

  67. Jiabao He, Yueyue Xu, Yue Ju, Cristian R. Rojas

    This contribution revisits the classical approximate realization problem, which involves determining matrices of a state-space model based on estimates of a truncated series of Markov parameters. A Hankel matrix built up by these Markov parameters plays a fundamental role in this problem, leveraging the fact that both its range space and left null space enco

  68. Ming Meng, Qi Dong, Jiajie Li, Zhe Zhu

    Virtual try-on technology has become increasingly important in the fashion and retail industries, enabling the generation of high-fidelity garment images that adapt seamlessly to target human models. While existing methods have achieved notable progress, they still face significant challenges in maintaining consistency across different poses. Specifically, g

  69. Byunghyun Yoo, Younghwan Shin, Hyunwoo Kim, Euisok Chung

    In standard reinforcement learning, an episode is defined as a sequence of interactions between agents and the environment, which terminates upon reaching a terminal state or a pre-defined episode length. Setting a shorter episode length enables the generation of multiple episodes with the same number of data samples, thereby facilitating an exploration of d

  70. Ipsita Mandal

    We embark on computing the longitudinal magnetoconductivity within the semiclassical Boltzmann formalism, where an isotropic triple-point semimetal (TSM) is subjected to collinear electric ($\boldsymbol E $) and magnetic ($\boldsymbol B$) fields. Except for the Drude part, the $B$-dependence arises exclusively from topological properties like the Berry curva

  71. Ivan Y. Tyukin, Bogdan Grechuk, Evgeny M. Mirkes, Alexander N. Gorban

    Concentration of distances in high dimension is an important factor for the development and design of stable and reliable data analysis algorithms. In this paper, we address the fundamental long-standing question about the concentration of distances in high dimension for fractional quasi $p$-norms, $p\in(0,1)$. The topic has been at the centre of various the

  72. Zili Wang, Tianyu Zhang, Haoli Bai, Lu Hou

    Test-Time Scaling (TTS) has proven effective in improving the performance of Large Language Models (LLMs) during inference. However, existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. Through a latency-aware evaluation of representative TTS methods, we demonstrate that a compute-optimal TTS does not always result in

  73. Martijn Hanegraaf, Savio Sciancalepore, Gabriele Oligeri

    State-of-the-art solutions detect jamming attacks ex-post, i.e., only when jamming has already disrupted the wireless communication link. In many scenarios, e.g., mobile networks or static deployments distributed over a large geographical area, it is often desired to detect jamming at the early stage, when it affects the communication link enough to be detec

  74. M. V. Takook

    Krein space quantization and the ambient space formalism have been successfully applied to address challenges in quantum geometry (e.g., quantum gravity) and the axiomatic formulation of quantum Yang-Mills theory, including phenomena such as color confinement and the mass gap. Building on these advancements, we aim to extend these methods to develop novel qu

  75. Zihong Zhang, Liqi He, Zuchao Li, Lefei Zhang

    Word segmentation stands as a cornerstone of Natural Language Processing (NLP). Based on the concept of "comprehend first, segment later", we propose a new framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) and evaluate the semantic understanding capabilities of LLMs based on word segmentation. We employ curren

  76. Yichun Feng, Jiawei Wang, Lu Zhou, Yikai Zheng

    Large language models (LLMs) struggle in real-world clinical consultations. Single-turn consultation systems require patients to describe all symptoms at once, which often leads to unclear complaints and vague diagnoses. Traditional dialogue models, constrained by static supervised learning, are limited to superficially imitating existing dialogue patterns a

  77. Hassan Sartaj, Shaukat Ali, Ana Cavalcanti, Lukas Esterle

    Self-adaptive robotic systems operate autonomously in dynamic and uncertain environments, requiring robust real-time monitoring and adaptive behaviour. Unlike traditional robotic software with predefined logic, self-adaptive robots exploit artificial intelligence (AI), machine learning, and model-driven engineering to adapt continuously to changing condition

  78. Silin Li, Yuhang Guo, Jiashu Yao, Zeming Liu

    Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home environment. While recent studies have explored integrating LLMs into smart home systems, they primarily focus on handling st

  79. Peng Feng, Hang Ma, Kuan Yang, Yingjie Lv

    This study utilizes first-principles computational methods to comprehensively analyze the impact of A-site doping on the proton conduction properties of BaHfO$_3$. The goal is to offer theoretical support for the advancement of electrolyte materials for solid oxide fuel cells. Our research has uncovered that BaHfO$_3$ demonstrates promising potential for pro

  80. Jiaxin Chen, Yiming Wang, Ziyu Zhang, Jiayang Han

    The same speech content produced by different speakers exhibits significant differences in pitch contour, yet listeners' semantic perception remains unaffected. This phenomenon may stem from the brain's perception of pitch contours being independent of individual speakers' pitch ranges. In this work, we recorded electroencephalogram (EEG) while participants

  81. Hassan Sartaj, Shaukat Ali, Paolo Arcaini, Andrea Arcuri

    Search-based software engineering (SBSE), which integrates metaheuristic search techniques with software engineering, has been an active area of research for about 25 years. It has been applied to solve numerous problems across the entire software engineering lifecycle and has demonstrated its versatility in multiple domains. With recent advances in Artifici

  82. Pusheng Xu, Xia Gong, Xiaolan Chen, Weiyi Zhang

    Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions published between January 1, 2016, and December 31, 2024, were collected from WeChat Official Accounts. Based on these captions, bilingual question-answer (QA) pairs in Chinese and

  83. Yu Shang, Peijie Liu, Yuwei Yan, Zijing Wu

    The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enable autonomous, adaptive decision-making. Unlike traditional recommendation approaches, agentic recommender systems can dynamically gather and

  84. Emmanuel Humbert, Kilian Raschel

    We consider multidimensional random walks in pyramidal cones (or multidimensional orthants), which are intersections of a finite number of half-spaces. We explore the connection between the existence of (positive) discrete harmonic polynomials for the random walks, with Dirichlet conditions on the boundary of the cone, and geometric properties of the cone, b

  85. George Kour, Itay Nakash, Ateret Anaby-Tavor, Michal Shmueli-Scheuer

    As Large Language Models (LLMs) become deeply integrated into human life and increasingly influence decision-making, it's crucial to evaluate whether and to what extent they exhibit subjective preferences, opinions, and beliefs. These tendencies may stem from biases within the models, which may shape their behavior, influence the advice and recommendations t

  86. Iveta Terezie Hošnová, Kristýna Havlinová, Jan Bárta, Karolína Mocová

    Nanocomposite material ($\mathrm{GGAG:Ce^{3+}@SiO_2-RB}$) for potential use in X-ray induced photodynamic therapy (X-PDT) was developed, thoroughly characterized, and evaluated. It consists of a scintillating $\mathrm{Gd_3(Ga_{1-x}Al_x)_5O_{12}:Ce^{3+}}$ core encapsulated in silica layer and functionalized with the photosensitizer Rose Bengal (RB). Radiolumi

  87. Jiawen Chen, Qi Shao, Duxin Chen, Wenwu Yu

    Spatio-temporal prediction is a pivotal task with broad applications in traffic management, climate monitoring, energy scheduling, etc. However, existing methodologies often struggle to balance model expressiveness and computational efficiency, especially when scaling to large real-world datasets. To tackle these challenges, we propose STH-SepNet (Spatio-Tem

  88. Jun Tian, He Wang, Jibo He, Yu Pan

    Convolutional neural networks (CNNs) have become widely adopted in gravitational wave (GW) detection pipelines due to their ability to automatically learn hierarchical features from raw strain data. However, the physical meaning of these learned features remains underexplored, limiting the interpretability of such models. In this work, we propose a hybrid ar

  89. Hanze Liu, Jiahong Fu, Qi Xie, Deyu Meng

    Self-supervised image denoising methods have garnered significant research attention in recent years, for this kind of method reduces the requirement of large training datasets. Compared to supervised methods, self-supervised methods rely more on the prior embedded in deep networks themselves. As a result, most of the self-supervised methods are designed wit

  90. Dominik Stempień, Robert Ślepaczuk

    This research systematically develops and evaluates various hybrid modeling approaches by combining traditional econometric models (ARIMA and ARFIMA models) with machine learning and deep learning techniques (SVM, XGBoost, and LSTM models) to forecast financial time series. The empirical analysis is based on two distinct financial assets: the S&P 500 index a

  91. Rui Cai, Bangzheng Li, Xiaofei Wen, Muhao Chen

    Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as irrelevant text in vision understanding, or irrelevant visual content in question answering. At its core, modality interference refers to cases where spurious signals from non-esse

  92. Hiroki Matsumura, Yuki Takahashi, Katsuki Kinjo, Shunsaku Kitagawa

    Knight shifts along the $b$ and $c$ axes ($K_b$ and $K_c$) at two crystallographically distinct Te sites were measured down to 70 mK using $^{125}$Te nuclear magnetic resonance (NMR) on an ultraclean UTe$_2$ single crystal with a superconducting (SC) transition temperature $T_{\mathrm{c}}$ = 2.1 K. This was carried out to determine the $\boldsymbol{d}$-vecto

  93. Amira Guesmi, Bassem Ouni, Muhammad Shafique

    Adversarial transferability remains a critical challenge in evaluating the robustness of deep neural networks. In security-critical applications, transferability enables black-box attacks without access to model internals, making it a key concern for real-world adversarial threat assessment. While Vision Transformers (ViTs) have demonstrated strong adversari

  94. Yui Tatsumi, Ziyue Zeng, Hiroshi Watanabe

    Conventional methods for scalable image coding for humans and machines require the transmission of additional information to achieve scalability. A recent diffusion-based approach avoids this by generating human-oriented images from machine-oriented images without extra bitrate. However, it utilizes a single random seed, which may lead to suboptimal image qu

  95. Pramit Das, Moulinath Banerjee, Yuekai Sun

    In many network systems, events at one node trigger further activity at other nodes, e.g., social media users reacting to each other's posts or the clustering of criminal activity in urban environments. These systems are typically referred to as self-exciting networks. In such systems, targeted intervention at critical nodes can be an effective strategy for

  96. Ruolin Shen, Xiaozhong Ji, Kai WU, Jiangning Zhang

    Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish concealed objects, demonstrating an inability to emulate human cognitive processes which effectively utilize foreground-back

  97. Ziga Kovacic, Justin T. Chiu, Celine Lee, Wenting Zhao

    Maintainable and general software allows developers to build robust applications efficiently, yet achieving these qualities often requires refactoring specialized solutions into reusable components. This challenge becomes particularly relevant as code agents become used to solve isolated one-off programming problems. We investigate code agents' capacity to r

  98. Jiaxin Song, Yixu Wang, Jie Li, Rui Yu

    Vision-Language Models (VLMs) exhibit impressive performance, yet the integration of powerful vision encoders has significantly broadened their attack surface, rendering them increasingly susceptible to jailbreak attacks. However, lacking well-defined attack objectives, existing jailbreak methods often struggle with gradient-based strategies prone to local o

  99. Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang

    Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution pos

  100. Davide Parodi, Federico Benvenuto, Sara Garbarino, Michele Piana

    Implicit inverse problems, in which noisy observations of a physical quantity are used to infer a nonlinear functional applied to an associated function, are inherently ill posed and often exhibit non uniqueness of solutions. Such problems arise in a range of domains, including the identification of systems governed by Ordinary and Partial Differential Equat