May 2025 arXiv papers — page 113
Showing 11,201–11,300 of 24,552 papers
Hybridized and Localized 4f Electronic States of Nd-based Intermetallic Compounds in Cubic Symmetry Probed by High-Energy Photoemission
cond-mat.str-elM. Sakaguchi, A. Enomoto, H. Fujiwara, G. Nozue
We have performed soft and hard X-ray photoemission spectroscopies on NdTi2Al20 and NdBe13 which show antiferromagnetic ordering at low temperatures. The Nd 3d core-level photoemission and Nd 3d-4f valence-band resonant photoemission spectra show finite 4f4 initial-state components in addition to the 4f3 configurations attributed to the c-f hybridization eff
Empirical Validation of Functional Multidimensional Scaling via Numerical Simulation and Real-World Application
stat.APLiting Li
This article presents an empirical validation of the functional multidimensional scaling model, a novel approach that improves the smoothness of time-varying dissimilarities in a low-dimensional space, embedding a modified Adam stochastic gradient descent method. We conduct a numerical simulation study to evaluate the feasibility of the functional multidimen
StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning
cs.CVHuaijie Wang, De Cheng, Guozhang Li, Zhipeng Xu
Video Class-Incremental Learning (VCIL) seeks to develop models that continuously learn new action categories over time without forgetting previously acquired knowledge. Unlike traditional Class-Incremental Learning (CIL), VCIL introduces the added complexity of spatiotemporal structures, making it particularly challenging to mitigate catastrophic forgetting
Akanksha Agrawal, Fedor V. Fomin, Daniel Lokshtanov, Saket Saurabh
A graph $G$ is contractible to a graph $H$ if there is a set $X \subseteq E(G)$, such that $G/X$ is isomorphic to $H$. Here, $G/X$ is the graph obtained from $G$ by contracting all the edges in $X$. For a family of graphs $\cal F$, the $\mathcal{F}$-\textsc{Contraction} problem takes as input a graph $G$ on $n$ vertices, and the objective is to output the la
Jie Xiang, Huijie Qiao
In this paper, we investigate a class of multiscale McKean-Vlasov stochastic systems, where the entire system depends on the distributions of both fast and slow components. First of all, by applying the Poisson equation method, we prove that the slow component converges to the solution of the averaging equation in the $L^p$ ($p\geq 2$) space with the optimal
Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe
LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth. This fails to capture broader forms of sycophancy such as affirming a user's self-image or other implicit beliefs. To a
Ruiyi Yang, Hao Xue, Imran Razzak, Shirui Pan
Retrieval-Augmented Generation (RAG) systems empower large language models (LLMs) with external knowledge, yet struggle with efficiency-accuracy trade-offs when scaling to large knowledge graphs. Existing approaches often rely on monolithic graph retrieval, incurring unnecessary latency for simple queries and fragmented reasoning for complex multi-hop questi
Payne Hennigan
I develop a dynamic model of how internal capital markets in conglomerates respond to liquidity shocks when affiliated firms vary in innovation potential. A two-stage framework defines cutoff rules for when the conglomerate should liquidate low-productivity firms, coerce intermediate types into short-termist strategies, or preserve high-potential firms for l
Azad Ch. Izmailov
The paper theoretically establishes and studies the sub-Doppler linear optical Ramsey resonances that arise under certain conditions near centers of optical transitions in the absorption of a sufficiently weak monochromatic light beam during its stationary propagation in the normal direction through an ultrathin gas cell whose internal thickness is less than
Luis Villegas-Aguilar, Farzad Ghafari, Matthew S. Winnel, Varun B. Verma
Large-scale quantum networking systems will inevitably require methods to overcome photon loss. While the no-cloning theorem forbids perfect and deterministic amplification of unknown quantum states, probabilistic heralded amplification schemes offer a viable path forward. Yet, for over a decade, successful multi-photon state amplification has remained out o
Jared Duker Lichtman
Let ${\rm rad}(n)$ denote the product of distinct prime factors of an integer $n\geq 1$. The celebrated $abc$ conjecture asks whether every solution to the equation $a+b=c$ in triples of coprime integers $(a,b,c)$ must satisfy ${\rm rad}(abc) > K_\varepsilon\, c^{1-\varepsilon}$, for some constant $K_\varepsilon>0$. In this expository note, we present a clas
Tingfeng Hui, Pengyu Zhu, Bowen Ping, Ling Tang
Instruction-following has emerged as a crucial capability for large language models (LLMs). However, existing approaches often rely on pre-existing documents or external resources to synthesize instruction-following data, which limits their flexibility and generalizability. In this paper, we introduce DecIF, a fully autonomous, meta-decomposition guided fram
Time Series Similarity Score Functions to Monitor and Interact with the Training and Denoising Process of a Time Series Diffusion Model applied to a Human Activity Recognition Dataset based on IMUs
cs.LGHeiko Oppel, Andreas Spilz, Michael Munz
Denoising diffusion probabilistic models are able to generate synthetic sensor signals. The training process of such a model is controlled by a loss function which measures the difference between the noise that was added in the forward process and the noise that was predicted by the diffusion model. This enables the generation of realistic data. However, the
Yanzhe Wen, Xunkai Li, Qi Zhang, Zhu Lei
Recently, large language models (LLMs) have significantly advanced text-attributed graph (TAG) learning. However, existing methods inadequately handle data uncertainty in open-world scenarios, especially concerning limited labeling and unknown-class nodes. Prior solutions typically rely on isolated semantic or structural approaches for unknown-class rejectio
Linxin Song, Taiwei Shi, Jieyu Zhao
Reinforcement finetuning (RFT) has become a standard approach for enhancing the reasoning capabilities of large language models (LLMs). However, its impact on model trustworthiness remains underexplored. In this work, we identify and systematically study a critical side effect of RFT, which we term the hallucination tax: a degradation in refusal behavior cau
Raghavendra Kaushal, Bhaghyesh
The mass spectrum of beauty hadrons ($b\overline{b}$ and $bbb$ baryons) and $bb$-diquarks are computed in a non-relativistic phenomenological potential model. The potential comprises of a short-range Coulomb potential, a screened confinement potential, and $O(1/m)$ corrections predicted from lattice and pNRQCD studies. Among the spin-dependent interactions,
Michael Dorner, Daniel Mendez
Background: Code review, a core practice in software engineering, has been widely studied as a collaborative process, with prior work suggesting it functions as a communication network. However, this theory remains untested, limiting its practical and theoretical significance. Objective: This study aims to (1) formalize the theory of code review as a communi
Joakim Arnlind, Victor Hildebrandsson
We study the existence of Levi-Civita connections, i.e torsion free connections compatible with a hermitian form, in the setting of derivation based noncommutative differential calculi over $\ast$-algebras. We prove a necessary and sufficient condition for the existence of Levi-Civita connections in terms of the image of an operator derived from the hermitia
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
cs.SDHao Shi, Xugang Lu, Kazuki Shimada, Tatsuya Kawahara
Diffusion-based speech enhancement (SE) models need to incorporate correct prior knowledge as reliable conditions to generate accurate predictions. However, providing reliable conditions using noisy features is challenging. One solution is to use features enhanced by deterministic methods as conditions. However, the information distortion and loss caused by
Jinzhou Li, Tianhao Wu, Jiyao Zhang, Zeyuan Chen
Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain comprehensively fused features but often ignore the fact that each modality requires different levels of attention at different manip
Deciphering the Formation and Dynamics of Double-decker Filament Through Component Magnetic Reconnection
astro-ph.SRDongxu Liu, Yuandeng Shen, Yi Bi, Zehao Tang
The formation of double-decker filaments has long been an enigma in the field of solar physics. Using stereoscopic observations from the Solar Dynamics Observatory and the Solar Terrestrial Relations Observatory, we show that the double-decker filament formed on 2013 August 30 resulted from the splitting of a braided magnetic flux rope. The splitting was dri
Grace Younes, Alban Quadrat, Fabrice Rouillier
The computation of the $L_\infty $-norm is an important issue in $H_{\infty}$ control, particularly for analyzing system stability and robustness. This paper focuses on symbolic computation methods for determining the $L_{\infty} $-norm of finite-dimensional linear systems, highlighting their advantages in achieving exact solutions where numerical methods of
Maya Srikanth, Run Chen, Julia Hirschberg
Multimodal models play a key role in empathy detection, but their performance can suffer when modalities provide conflicting cues. To understand these failures, we examine cases where unimodal and multimodal predictions diverge. Using fine-tuned models for text, audio, and video, along with a gated fusion model, we find that such disagreements often reflect
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
cs.SDYuan Gao, Hao Shi, Yahui Fu, Chenhui Chu
This study investigates the interaction between personality traits and emotion expression, exploring how personality information can improve speech emotion recognition (SER). We collect the personality annotation for the IEMOCAP dataset, making it the first speech dataset that contains both emotion and personality annotations (PA-IEMOCAP), and enabling direc
Upasana Sarmah, Parthajit Borah, D. K. Bhattacharyya
Applications over the Web primarily rely on the HTTP protocol to transmit web pages to and from systems. There are a variety of application layer protocols, but among all, HTTP is the most targeted because of its versatility and ease of integration with online services. The attackers leverage the fact that by default no detection system blocks any HTTP traff
Investigation of Martian UV Dayglow Emissions in the Southern Hemisphere during Solar Quiet-time Conditions: Insights from Multi-year MAVEN/IUVS Observations
physics.space-phAadarsh Raj Sharma, Lot Ram, Sumanta Sarkhel
The southern hemisphere of Mars possesses concentrated region of strong crustal magnetic fields (CMF), which generate localized magnetic anomalies that can influence atmospheric dynamics and energy deposition in the Martian thermospheric-ionospheric system. Although their effects on the atmosphere (>200 km) in the southern hemisphere are well documented, how
Taewoo Kim, Guisik Kim, Choongsang Cho, Young Han Lee
Recent advances in speech deepfake detection (SDD) have significantly improved artifacts-based detection in spoofed speech. However, most models overlook speech naturalness, a crucial cue for distinguishing bona fide speech from spoofed speech. This study proposes naturalness-aware curriculum learning, a novel training framework that leverages speech natural
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
cs.CLYuxuan Jiang, Dawei Li, Francis Ferraro
While Large Reasoning Models (LRMs) have demonstrated success in complex reasoning tasks through long chain-of-thought (CoT) reasoning, their inference often involves excessively verbose reasoning traces, resulting in substantial inefficiency. To address this, we propose Distilled Reasoning Pruning (DRP), a hybrid framework that combines inference-time pruni
Hye-young Kim, Minjin Choi, Sunkyung Lee, Ilwoong Baek
Side-information Integrated Sequential Recommendation (SISR) benefits from auxiliary item information to infer hidden user preferences, which is particularly effective for sparse interactions and cold-start scenarios. However, existing studies face two main challenges. (i) They fail to remove noisy signals in item sequence and (ii) they underutilize the pote
Wenhui Zhu, Xuanzhao Dong, Xin Li, Peijie Qiu
Recently, reinforcement learning (RL)-based tuning has shifted the trajectory of Multimodal Large Language Models (MLLMs), particularly following the introduction of Group Relative Policy Optimization (GRPO). However, directly applying it to medical tasks remains challenging for achieving clinically grounded model behavior. Motivated by the need to align mod
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
cs.CLQianli Wang, Van Bach Nguyen, Nils Feldhus, Luis Felipe Villa-Arenas
Counterfactual examples are widely employed to enhance the performance and robustness of large language models (LLMs) through counterfactual data augmentation (CDA). However, the selection of the judge model used to evaluate label flipping, the primary metric for assessing the validity of generated counterfactuals for CDA, yields inconsistent results. To dec
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
cs.SDMing Gao, Shilong Wu, Hang Chen, Jun Du
Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on multi-modal, multi-device meeting transcription by incorporating video modality alongside audio. The tasks include Audio-Visual Speaker Diariza
Émilie Fabre, Katie Seaborn, Shuta Koiwai, Mizuki Watanabe
Longitudinal engagement with generative AI (GenAI) storytelling agents is a timely but less charted domain. We explored multi-generational experiences with "Dreamsmithy," a daily dream-crafting app, where participants (N = 28) co-created stories with AI narrator "Makoto" every day. Reflections and interactions were captured through a two-week diary study. Re
Micromagnetic Study of the Dipolar-Exchange Spin Waves in Antiferromagnetic Thin Films
cond-mat.mes-hallJiongjie Wang, Jiang Xiao
In antiferromagnets, dipolar coupling is often disregarded due to the cancellation of magnetic moments between the two sublattices, leaving spin-wave dispersion predominantly determined by exchange interactions. However, antiferromagnetic spin waves typically involve a slight misalignment of the magnetic moments on the sublattices, giving rise to a small net
Hypothesis on the Functional Advantages of the Selection-Broadcast Cycle Structure: Global Workspace Theory and Dealing with a Real-Time World
cs.ROJunya Nakanishi, Jun Baba, Yuichiro Yoshikawa, Hiroko Kamide
This paper discusses the functional advantages of the Selection-Broadcast Cycle structure proposed by Global Workspace Theory (GWT), inspired by human consciousness, particularly focusing on its applicability to artificial intelligence and robotics in dynamic, real-time scenarios. While previous studies often examined the Selection and Broadcast processes in
Shangyou Zhang
We extend the $C^1$-$P_3$ Fraeijs de Veubeke-Sander finite element to two families of $C^1$-$P_k$ ($k>3$) macro finite elements on general quadrilateral meshes. On each quadrilateral, four $P_k$ polynomials are defined on the four triangles subdivided from the quadrilateral by its two diagonal lines. The first family of $C^1$-$P_k$ finite elements is the ful
Xu Yang, Xiao Yang, Shikai Fang, Yifei Zhang
Recent advances in AI and ML have transformed data science, yet increasing complexity and expertise requirements continue to hinder progress. Although crowd-sourcing platforms alleviate some challenges, high-level machine learning engineering (MLE) tasks remain labor-intensive and iterative. We introduce R&D-Agent, a comprehensive, decoupled, and extensible
A Quasi-Newton Method to Solve Uncertain Multiobjective Optimization Problems with Uncertainty Set of Finite Cardinality
math.OCK. Gupta, D. Ghosh, C. Tammer, X. Zhao
In this article, we derive an iterative scheme through a quasi-Newton technique to capture robust weakly efficient points of uncertain multiobjective optimization problems under the upper set less relation. It is assumed that the set of uncertainty scenarios of the problems being analyzed is of finite cardinality. We also assume that corresponding to each gi
Imran Ali Khan, Saif Khan Mohammed, Ronny Hadani, Ananthanarayanan Chockalingam
Across the world, there is growing interest in new waveforms, Zak-OTFS in particular, and over-the-air implementations are starting to appear. The choice between OFDM and Zak-OTFS is not so much a choice between waveforms as it is an architectural choice between preventing inter-carrier interference (ICI) and embracing ICI. In OFDM, once the Input-Output (I/
Jiamin Su, Yibo Yan, Zhuoran Gao, Han Zhang
Automated Essay Scoring (AES) is crucial for modern education, particularly with the increasing prevalence of multimodal assessments. However, traditional AES methods struggle with evaluation generalizability and multimodal perception, while even recent Multimodal Large Language Model (MLLM)-based approaches can produce hallucinated justifications and scores
Yuto Mandai, Katie Seaborn, Tomoyasu Nakano, Xin Sun
"Kawaii" is the Japanese concept of cute, which carries sociocultural connotations related to social identities and emotional responses. Yet, virtually all work to date has focused on the visual side of kawaii, including in studies of computer agents and social robots. In pursuit of formalizing the new science of kawaii vocalics, we explored what elements of
Taoran Li, Taobo Liao
We present a secure and efficient string-matching platform leveraging zk-SNARKs (Zero-Knowledge Succinct Non-Interactive Arguments of Knowledge) to address the challenge of detecting sensitive information leakage while preserving data privacy. Our solution enables organizations to verify whether private strings appear on public platforms without disclosing t
Haoyang Zhang, Hexin Liu, Xiangyu Zhang, Qiquan Zhang
The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact of frame rates on speech tokens remains underexplored. In this study, we investigate how varying frame rates affect speech tokenization by
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
cs.CLQianli Wang, Mingyang Wang, Nils Feldhus, Simon Ostermann
Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilities have been extensively studied, one critical area remains underexplored: factual knowledge recall (FKR), the process by which LLMs access stored knowledge. To this end, we condu
TDCOSMO XXI. Accurate stellar velocity dispersions of the SL2S lens sample and the fundamental plane of the lensing mass
astro-ph.GAPritom Mozumdar, Shawn Knabel, Tommaso Treu, Alessandro Sonnenfeld
We reanalyzed spectra that were taken as part of the SL2S lens galaxy survey with the goal to obtain the stellar velocity dispersion with a precision and accuracy sufficient for time-delay cosmography. In order to achieve this goal, we imposed stringent cuts on the signal-to-noise ratio (S/N), and employed recently developed methods to mitigate and quantify
WALLABY pilot survey: Spatially resolved gas scaling relations within the stellar discs of nearby galaxies
astro-ph.GASeona Lee, Barbara Catinella, Tobias Westmeier, Luca Cortese
The scatter in global atomic hydrogen (HI) scaling relations is partly attributed to differences in how HI and stellar properties are measured, with HI reservoirs typically extending beyond the inner regions of galaxies where star formation occurs. Using pilot observations from the WALLABY survey, we present the first measurements of HI mass enclosed within
Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
cs.HCTakao Fujii, Katie Seaborn, Madeleine Steeds, Jun Kato
Conversational agents that mimic people have raised questions about the ethics of anthropomorphizing machines with human social identity cues. Critics have also questioned assumptions of identity neutrality in humanlike agents. Recent work has revealed that intersectional Japanese pronouns can elicit complex and sometimes evasive impressions of agent identit
José A. Carrillo, Kam Fai Chan, Renjun Duan, Zongguang Li
In this paper, we study the spatially homogeneous inelastic Boltzmann equation for the angular cutoff pseudo-Maxwell molecules with an additional term of linear deformation. We establish the existence of non-Maxwellian self-similar profiles under the assumption of small deformation in the nearly elastic regime, and also obtain weak convergence to these self-
MultiDrive: A Co-Simulation Framework Bridging 2D and 3D Driving Simulation for AV Software Validation
cs.ROMarc Kaufeld, Korbinian Moller, Alessio Gambi, Paolo Arcaini
Scenario-based testing using simulations is a cornerstone of Autonomous Vehicles (AVs) software validation. So far, developers needed to choose between low-fidelity 2D simulators to explore the scenario space efficiently, and high-fidelity 3D simulators to study relevant scenarios in more detail, thus reducing testing costs while mitigating the sim-to-real g
Demonstrating Coherent Quantum Routers for Bucket-Brigade Quantum Random Access Memory on a Superconducting Processor
quant-phSheng Zhang, Yun-Jie Wang, Peng Wang, Ren-Ze Zhao
Quantum routers (QRouters) are essential components of bucket-brigade quantum random access memory (QRAM), enabling quantum applications such as Grover's search and quantum machine learning. Despite significant theoretical advances, achieving scalable and coherent QRouters experimentally remains challenging. Here, we demonstrate coherent quantum routers usin
Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng
Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities. While text-based RAG privacy risks have been studied, multimodal data presents unique challenges. We provide the first systematic analysis of MRAG privacy vulnerabilities across vision-language
Modulating Thermometric Performance via Dopant Concentration and Morphology in Luminescence Thermometer Exhibiting Dual Structural Phase Transitions
cond-mat.mtrl-sciMalgorzata Kubicka, Maja Szymczak, Maciej Ptak, Damian Szymanski
Expanding the operational range of luminescent thermometers that utilize thermally induced structural phase transitions in lanthanide-doped materials necessitates the exploration of novel host matrices with diverse thermal behaviors. In line with this objective, the present study offers a comprehensive analysis of the temperature-dependent spectroscopic prop
Paradigm Shift in Infrastructure Inspection Technology: Leveraging High-performance Imaging and Advanced AI Analytics to Inspect Road Infrastructure
cs.DCDu Wu, Enzhi Zhang, Isaac Lyngaas, Xiao Wang
Effective road infrastructure management is crucial for modern society. Traditional manual inspection techniques remain constrained by cost, efficiency, and scalability, while camera and laser imaging methods fail to capture subsurface defects critical for long-term structural integrity. This paper introduces ROVAI, an end-to-end framework that integrates hi
Jiahe Chen, Ziye Ma
Optimizing large-scale nonconvex problems, common in deep learning, demands balancing rapid convergence with computational efficiency. First-order (FO) optimizers, which serve as today's baselines, provide fast convergence and good generalization but often incur high computation and memory costs due to the large size of modern models. Conversely, zeroth-orde
Human Authenticity and Flourishing in an AI-Driven World: Edmund's Journey and the Call for Mindfulness
cs.HCSebastian Zepf, Mark Colley
Humans have always dreamed of possessing superpowers, and the rapid development of AI-based features promises to bring these dreams (closer) to reality. However, these advancements come with significant risks. This paper advocates for challenging existing methods and approaches in design and evaluation for more responsible AI. We stimulate reflection through
Yi-Xian Chen, Yinhao Wu, Ya-Ping Li, Douglas N. C. Lin
Mean-motion resonances (MMRs) form through convergent disc migration of planet pairs, which may be disrupted by dynamical instabilities after protoplanetary disc (PPD) dispersal. This scenario is supported by recent analysis of TESS data showing that neighboring planet pairs in younger planetary systems are closer to resonance. To study stability of MMRs dur
Yi-Xian Chen, Yan-Fei Jiang, Jeremy Goodman
Massive stars can form within or be captured by AGN disks, influencing both the thermal structure and metallicity of the disk environment. In a previous work, we investigated isotropic accretion onto massive stars from a gas-rich, high-entropy background. Here, we consider a more realistic scenario by incorporating the stratified geometry of the background d
Ziyang Zeng, Dun Zhang, Jiacheng Li, Panxiang Zou
This study investigates the position bias in information retrieval, where models tend to overemphasize content at the beginning of passages while neglecting semantically relevant information that appears later. To analyze the extent and impact of position bias, we introduce a new evaluation framework consisting of two position-aware retrieval benchmarks (SQu
Guochao Jiang, Guofeng Quan, Zepeng Ding, Ziqin Luo
Large Language Models (LLMs) have shown impressive performance in reasoning tasks. However, LLMs tend to generate excessively long reasoning content, leading to significant computational overhead. Our observations indicate that even on simple problems, LLMs tend to produce unnecessarily lengthy reasoning content, which is against intuitive expectations. Prel
Mingliang Zhai, Zhi Gao, Yuwei Wu, Yunde Jia
Embodied Question Answering (EQA) requires agents to autonomously explore and comprehend the environment to answer context-dependent questions. Typically, an EQA framework consists of four components: a planner, a memory module, a stopping module, and an answering module. However, the memory module is utilized inefficiently in existing methods, as the inform
Shirong Xu, Hengzhi He, Guang Cheng
In recent years, model collapse has become a critical issue in language model training, making it essential to understand the underlying mechanisms driving this phenomenon. In this paper, we investigate recursive parametric model training from a probabilistic perspective, aiming to characterize the conditions under which model collapse occurs and, crucially,
Changdae Oh, Jiatong Li, Shawn Im, Sharon Li
Despite widespread adoption, multimodal large language models (MLLMs) suffer performance degradation when encountering unfamiliar queries under distribution shifts. Existing methods to improve MLLM generalization typically require either more instruction data or larger advanced model architectures, both of which incur non-trivial human labor or computational
Siyuan Dong, Yuxuan Tian, Wenhan Ma, Tong Yang
Data stream monitoring is a crucial task which has a wide range of applications. The majority of existing research in this area can be broadly classified into two types, monitoring value sum and monitoring value cardinality. In this paper, we define a third type, monitoring value variation, which can help us detect flow gaps in data streams. To realize this
WAVE++: Capturing Within-Task Variance for Continual Relation Extraction with Adaptive Prompting
cs.CLBao-Ngoc Dao, Minh Le, Quang Nguyen, Luyen Ngo Dinh
Memory-based approaches have shown strong performance in Continual Relation Extraction (CRE). However, storing examples from previous tasks increases memory usage and raises privacy concerns. Recently, prompt-based methods have emerged as a promising alternative, as they do not rely on storing past samples. Despite this progress, current prompt-based techniq
Samee Arif, Sualeha Farid
This paper presents a comparative analysis of Large Language Models (LLMs) and traditional Optical Character Recognition (OCR) systems on Urdu newspapers, addressing challenges posed by complex multi-column layouts, low-resolution scans, and the stylistic variability of the Nastaliq script. To handle these challenges, we fine-tune YOLOv11x models for article
Diego Ortiz Barbosa, Luis Burbano, Carlos Hernandez, Zengxiang Lei
Intelligent mechanisms implemented in autonomous vehicles, such as proactive driving assist and collision alerts, reduce traffic accidents. However, verifying their correct functionality is difficult due to complex interactions with the environment. This problem is exacerbated in adversarial environments, where an attacker can control the environment surroun
Haoyang Fang, Boran Han, Nick Erickson, Xiyuan Zhang
Existing AutoML systems have advanced the automation of machine learning (ML); however, they still require substantial manual configuration and expert input, particularly when handling multimodal data. We introduce MLZero, a novel multi-agent framework powered by Large Language Models (LLMs) that enables end-to-end ML automation across diverse data modalitie
Development and Validation of Engagement and Rapport Scales for Evaluating User Experience in Multimodal Dialogue Systems
cs.CLFuma Kurata, Mao Saeki, Masaki Eguchi, Shungo Suzuki
This study aimed to develop and validate two scales of engagement and rapport to evaluate the user experience quality with multimodal dialogue systems in the context of foreign language learning. The scales were designed based on theories of engagement in educational psychology, social psychology, and second language acquisition.Seventy-four Japanese learner
Kun Li, Zhennan Wu, Shoupeng Wang, Jia Wu
Large language models (LLMs) integrated with autonomous agents hold significant potential for advancing scientific discovery through automated reasoning and task execution. However, applying LLM agents to drug discovery is still constrained by challenges such as large-scale multimodal data processing, limited task automation, and poor support for domain-spec
Ibrokhimbek Akramov, Dildora Ikromova
In this paper, we will consider $E$-type singularities which are Arnol'd type. We provide invariant conditions for a sufficiently smooth functions to have singularities of type $E_k (6\le k\le 8)$. We show the functions can be reduced to $E_k, k=6, 7, 8$ type normal form under some certain conditions. Moreover, we show that result on normal form for sufficie
Amitayush Thakur, Jasper Lee, George Tsoukalas, Meghana Sistla
We introduce ${\rm C{\small LEVER}}$, a high-quality, curated benchmark of 161 problems for end-to-end verified code generation in Lean. Each problem consists of (1) the task of generating a specification that matches a held-out ground-truth specification, and (2) the task of generating a Lean implementation that provably satisfies this specification. Unlike
Merina Aruja, Lisa Mathew, Jayakrishna Vijayakumar
Quantum computing is a relatively new field of computing, which utilises the fundamental concepts of quantum mechanics to process data. The seminal paper of Moore et al. [2000] introduced quantum grammars wherein a set of amplitudes was attached to each production. However they did not study the final probability of the derived word. Aruja et al. [2025] cons
Saydul Akbar Murad, Ashim Dahal, Nick Rahimi
With the rapid advancement of large language models like Gemini, GPT, and others, bridging the gap between the human brain and language processing has become an important area of focus. To address this challenge, researchers have developed various models to decode EEG signals into text. However, these models still face significant performance limitations. To
The Initial Mass Function of the Galactic Early-type Field Stars Based on the LAMOST Survey
astro-ph.GAQida Li, Jianping Xiong, Zhenwei Li, Dan Qiu
Research on the high-mass end of the initial mass function (IMF) has been limited due to a scarcity of samples. Recently, Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST), as the most efficient spectroscopic telescope, has provided new opportunities for related research. In this study, based on approximately 70,000 main-sequence early-type
Jialong Wu, Shaofeng Yin, Ningya Feng, Mingsheng Long
World models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likelihood estimation (MLE) often misalign with task-specific goals of world models, i.e., transition prediction metrics like accuracy or perceptual quality. In this paper, we present RL
Qingyu Li, Chiranjib Mukhopadhyay, Abolfazl Bayat, Ali Habibnia
Recent advances in quantum computing have demonstrated its potential to significantly enhance the analysis and forecasting of complex classical data. Among these, quantum reservoir computing has emerged as a particularly powerful approach, combining quantum computation with machine learning for modeling nonlinear temporal dependencies in high-dimensional tim
C. U. Angeliya, Arnab Char, T. Karthick
A class of graphs $\cal G$ is said to be \emph{near optimal colorable} if there exists a constant $c\in \mathbb{N}$ such that every graph $G\in \cal G$ satisfies $\chi(G) \leq \max\{c, \omega(G)\}$, where $\chi(G)$ and $\omega(G)$ respectively denote the chromatic number and clique number of $G$. The class of near optimal colorable graphs is an important sub
Sketch Interface for Teleoperation of Mobile Manipulator to Enable Intuitive and Intended Operation: A Proof of Concept
cs.ROYuka Iwanaga, Masayoshi Tsuchinaga, Kosei Tanada, Yuji Nakamura
Recent advancements in robotics have underscored the need for effective collaboration between humans and robots. Traditional interfaces often struggle to balance robot autonomy with human oversight, limiting their practical application in complex tasks like mobile manipulation. This study aims to develop an intuitive interface that enables a mobile manipulat
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
cs.SDYassine El Kheir, Tim Polzehl, Sebastian Möller
We propose BiCrossMamba-ST, a robust framework for speech deepfake detection that leverages a dual-branch spectro-temporal architecture powered by bidirectional Mamba blocks and mutual cross-attention. By processing spectral sub-bands and temporal intervals separately and then integrating their representations, BiCrossMamba-ST effectively captures the subtle
Jerome Droniou, Kim-Ngan Le, Huateng Zhu
We perform numerical analysis of a nonlinear gradient flow, which can be regarded as a parabolic minimal surface problem or a regularised total variation flow, using the gradient discretisation method (GDM). GDM is a unified convergence analysis framework that covers conforming and nonconforming numerical methods, for instance, conforming and nonconforming f
Qifeng Cai, Hao Liang, Zhaoyang Han, Hejun Dong
Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality captions, and coarse annotation granularity, which hinder the evaluation of advanced video-text retrieval methods. To address these limitations, we
Coupled Adaptable Backward-Forward-Backward Resolvent Splitting Algorithm (CABRA): A Matrix-Parametrized Resolvent Splitting Method for the Sum of Maximal Monotone and Cocoercive Operators Composed with Linear Coupling Operators
math.OCPeter Barkley, Robert L. Bassett
We present a novel matrix-parametrized frugal splitting algorithm which finds the zero of a sum of maximal monotone and cocoercive operators composed with linear selection operators. We also develop a semidefinite programming framework for selecting matrix parameters and demonstrate its use for designing matrix parameters which provide beneficial diagonal sc
An exploratory study of a tellurium-loaded liquid scintillator based on water and p-dioxane
physics.ins-detYe Liang, Haozhe Sun, Zhe Wang
Tellurium-loaded liquid scintillators are critical for neutrinoless double-beta decay experiments. However, conventional organic scintillators are constrained by the limited solubility of organic tellurium compounds compared with that of inorganic ones in water, whereas water-based scintillators are likely constrained by the destabilization of surfactants ca
Yunpeng Jiang, Jianshu Hu, Paul Weng, Yutong Ban
Symmetry is pervasive in robotics and has been widely exploited to improve sample efficiency in deep reinforcement learning (DRL). However, existing approaches primarily focus on spatial symmetries, such as reflection, rotation, and translation, while largely neglecting temporal symmetries. To address this gap, we explore time reversal symmetry, a form of te
Maicon R. Correa, Abimael F. D. Loula
Stable and accurate finite element methods are presented for Darcy flow in heterogeneous porous media with an interface of discontinuity of the hydraulic conductivity tensor. Accurate velocity fields are computed through global or local post-processing formulations that use previous approximations of the hydraulic potential. Stability is provided by combinin
An Explorative Analysis of SVM Classifier and ResNet50 Architecture on African Food Classification
cs.CVChinedu Emmanuel Mbonu, Kenechukwu Anigbogu, Doris Asogwa, Tochukwu Belonwu
Food recognition systems has advanced significantly for Western cuisines, yet its application to African foods remains underexplored. This study addresses this gap by evaluating both deep learning and traditional machine learning methods for African food classification. We compared the performance of a fine-tuned ResNet50 model with a Support Vector Machine
Gogulakrishnan Thiyagarajan, Vinay Bist, Prabhudarshi Nayak
Outdated software remains a potent and underappreciated menace in 2025's cybersecurity environment, exposing systems to a broad array of threats, including ransomware, data breaches, and operational outages that can have devastating and far-reaching impacts. This essay explores the unseen threats of cyberattacks by presenting robust statistical information,
Wanjing Huang, Weixiang Yan, Zhen Zhang, Ambuj Singh
Large Language Models (LLMs) demonstrate strong reasoning and task planning capabilities but remain fundamentally limited in physical interaction modeling. Existing approaches integrate perception via Vision-Language Models (VLMs) or adaptive decision-making through Reinforcement Learning (RL), but they fail to capture dynamic object interactions or require
Significant Enhancement of Carrier Mobility in Finite vs. Infinite Square Quantum Wells: A Comparative Study of GaAs/In$_x$Ga$_{1-x}$As/GaAs Heterostructures
cond-mat.mes-hallTruong Van Tuan, Nguyen Dung Chinh, Tran Trong Tai, Vo Van Tai
The geometry of quantum wells (QWs) critically influences carrier mobility, yet systematic comparisons between finite and infinite square QWs remain scarce. We present a comprehensive study of GaAs/In$_x$Ga$_{1-x}$As/GaAs heterostructures using a variational-subband-wave-function model, analyzing key scattering mechanisms: remote impurities (RI), alloy disor
Shivam Kumar Mishra, Jackson Levi Said, B. Mishra
Gravitational waves offer a key insight into the viability of classes of gravitational theories beyond general relativity. The observational constraints on their speed of propagation can provide strong constraints on generalized classes of broader gravitational frameworks. In this work, we reconsider the general class of Gauss-Bonnet theories in the context
Ruikun Li, Huandong Wang, Jingtao Ding, Yuan Yuan
Data-driven dynamics prediction often fails under environmental shifts, while traditional fine-tuning remains computationally prohibitive for hardware-constrained or data-scarce applications. We propose DynaDiff, a generative meta-learning framework that transitions the paradigm from gradient-based tuning or modulation to direct weight-space generation. Spec
Yixuan Huang, Kailai Wang, Jian Shi
The transition to hydrogen powered transportation requires regionally tailored yet scalable infrastructure planning. This study presents the first Texas specific, multi-period mixed integer optimization model for hydrogen transportation from 2025 to 2050, addressing challenges in infrastructure phasing, asset coordination, and multimodal logistics. The frame
Investigation of the neural origin of non-Euclidean visual space and analysis of visual phenomena using information geometry
q-bio.NCDebasis Mazumdar, Kuntal Ghosh, Soma Mitra, Late Kamales Bhaumik
The present paper aims to develop a mathematical model concerning the visual perception of spatial information. It is a challenging problem in theoretical neuroscience to investigate how the spatial information of the objects in the physical space is encoded and decoded in the neural processes in the brain. In the past, researchers conjectured the existence
Malakhi Hopkins, Alice Kate Li, Shobhita Kramadhati, Jackson Arnold
Common remote sensing modalities (RGB, multispectral, hyperspectral imaging or LiDAR) are often used to indirectly measure crop health and do not directly capture plant stress indicators. Commercially available direct leaf sensors are bulky, powered electronics that are expensive and interfere with crop growth. In contrast, low-cost, passive and bio-degradab
Chu Chen, Kangning Cui, Pasquale Cascarano, Wei Tang
Ultrasound imaging is widely applied in clinical practice, yet ultrasound videos often suffer from low signal-to-noise ratios (SNR) and limited resolutions, posing challenges for diagnosis and analysis. Variations in equipment and acquisition settings can further exacerbate differences in data distribution and noise levels, reducing the generalizability of p
Jake Chandler, Richard Booth
Despite efforts to better understand the constraints that operate on single-step parallel (aka "package", "multiple") revision, very little work has been carried out on how to extend the model to the iterated case. A recent paper by Delgrande & Jin outlines a range of relevant rationality postulates. While many of these are plausible, they lack an underlying
Hiram Ring
A fundamental concern in linguistics has been to understand how languages change, such as in relation to word order. Since the order of words in a sentence (i.e. the relative placement of Subject, Object, and Verb) is readily identifiable in most languages, this has been a productive field of study for decades (see Greenberg 1963; Dryer 2007; Hawkins 2014).
Qiaochu Ma, Xiang Tang, Hsian-Hua Tseng, Zhaoting Wei
We use flat antiholomorphic superconnections to study orbifold Chern character following the method introduced by Bismut, Shen, and Wei. We show the uniqueness of orbifold Chern character by proving a Riemann-Roch-Grothendieck theorem for orbifold embeddings.
Bronchovascular Tree-Guided Weakly Supervised Learning Method for Pulmonary Segment Segmentation
eess.IVRuijie Zhao, Zuopeng Tan, Xiao Xue, Longfei Zhao
Pulmonary segment segmentation is crucial for cancer localization and surgical planning. However, the pixel-wise annotation of pulmonary segments is laborious, as the boundaries between segments are indistinguishable in medical images. To this end, we propose a weakly supervised learning (WSL) method, termed Anatomy-Hierarchy Supervised Learning (AHSL), whic
Guangtao Zheng, Wenqian Ye, Aidong Zhang
Deep learning models often achieve high performance by inadvertently learning spurious correlations between targets and non-essential features. For example, an image classifier may identify an object via its background that spuriously correlates with it. This prediction behavior, known as spurious bias, severely degrades model performance on data that lacks