May 2025 arXiv papers — page 91
Showing 9,001–9,100 of 24,552 papers
Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
cs.CLVijeta Deshpande, Debasmita Ghose, John D. Patterson, Roger Beaty
Diverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training. We show that common diversity metrics, and even reward models used for preference optimization, systematically bias models toward shorter outputs, limiting expressiveness. To address this, we introduce Diverse, not Short (Diverse-NS), a leng
Masanari Kimura, Howard Bondell
The power prior is a class of informative priors designed to incorporate historical data alongside current data in a Bayesian framework. It includes a power parameter that controls the influence of historical data, providing flexibility and adaptability. A key property of the power prior is that the resulting posterior minimizes a linear combination of KL di
James A. Rossmanith, Christine Vaughan
Vlasov equations model the dynamics of plasma in the collisionless regime. A standard approach for numerically solving the Vlasov equation is to operator split the spatial and velocity derivative terms, allowing simpler time-stepping schemes to be applied to each piece separately (known as the Cheng-Knorr method). One disadvantage of such an operator split m
Runze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros
When applying offline reinforcement learning (RL) in healthcare scenarios, the out-of-distribution (OOD) issues pose significant risks, as inappropriate generalization beyond clinical expertise can result in potentially harmful recommendations. While existing methods like conservative Q-learning (CQL) attempt to address the OOD issue, their effectiveness is
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
cs.CLYuhan Ji, Song Gao, Ying Nie, Ivan Majić
Applying AI foundation models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of complex spatial relations. To address these issues, we investigate the extent to which a well-known-text (WKT) representati
Viet-Anh Nguyen, Shiqian Zhao, Gia Dao, Runyi Hu
Recently, Large Reasoning Models (LRMs) have demonstrated superior logical capabilities compared to traditional Large Language Models (LLMs), gaining significant attention. Despite their impressive performance, the potential for stronger reasoning abilities to introduce more severe security vulnerabilities remains largely underexplored. Existing jailbreak me
Camelia R. Walker, Md Nurul Anwar, Leandra Brauninger, Jack Richards
Plasmodium falciparum is responsible for the majority of malaria morbidity and mortality each year. Malaria transmission rates vary by location and time of year due to climate and environmental conditions. We show the impact of these factors by developing a stochastic spatiotemporal agent-based malaria model that captures the impact of spatially distributed
Zheng Chen, Zichen Zou, Kewei Zhang, Xiongfei Su
Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training o
Magnetic Charge State Controlled Spin-Wave Dynamics in Nanoscale Three-Dimensional Artificial Spin Ice
cond-mat.mes-hallChandan Kumar, Amrit Kumar Mondal, Sreya Pal, Sayan Mathur
Three-dimensional (3D) magnetic nanostructures offer a versatile platform for exploring complex spin textures and spin-wave (SW) dynamics, with implications in next-generation spintronic and magnonic technologies. Advances in 3D nanofabrication have allowed a wide-range of structures and phenomena to be realized. Whilst the study of simple cylindrical magnet
Align-GRAG: Anchor and Rationale Guided Dual Alignment for Graph Retrieval-Augmented Generation
cs.CLDerong Xu, Pengyue Jia, Xiaopeng Li, Yingyi Zhang
Despite the strong abilities, large language models (LLMs) still suffer from hallucinations and reliance on outdated knowledge, raising concerns in knowledge-intensive tasks. Graph-based retrieval-augmented generation (GRAG) enriches LLMs with knowledge by retrieving graphs leveraging relational evidence, but it faces two challenges: structure-coupled irrele
Base Station Placement Optimization for Networked Sensing Exploiting Target Location Distribution
cs.ITKaiyue Hou, Shuowen Zhang
This paper studies a networked sensing system with multiple base stations (BSs), which collaboratively sense the unknown and random three-dimensional (3D) location of a target based on the target-reflected echo signals received at the BSs. Considering a practical scenario where the target location distribution is known a priori for exploitation, we aim to de
Rashed Shelim, Shengzhe Xu, Walid Saad, Naren Ramakrishnan
Vector representations of contextual embeddings learned by pre-trained large language models (LLMs) are effective in various downstream tasks in numerical domains such as time series forecasting. Despite their significant benefits, the tendency of LLMs to hallucinate in such domains can have severe consequences in applications such as energy, nature, finance
Kosei Nakagawa, Yoshiko Kanada-En'yo
The $0_2^+$ state of ${}^8\mathrm{He}$ has been discovered by a recent experiment, which suggested a developed cluster structure of spatially correlated neutron pairs, called "dineutrons" (${}^2n$). We aim to investigate the structure of ${}^8\mathrm{He}(0^+_1)$ and ${}^8\mathrm{He}(0^+_2)$ to clarify the monopole excitation mode in ${}^8\mathrm{He}$ system
Wei Zhang, Zhenhong Zhou, Kun Wang, Junfeng Fang
While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions-e.g., write a 10,000-word novel. Additionally, models often generate far too short outputs, terminate prematurely, or even refuse the request. Existing benchmarks focus p
Rajesh Kumar, Suchi Kumari, Anubhav Mishra
Real-world complex systems exhibit intricate interconnections and dependencies, especially social networks, technological infrastructures, and communication networks. These networks are prone to disconnection due to random failures or external attacks on their components. Therefore, managing the security and resilience of such networks is a prime concern, pa
Ali Sarosh Bangash, Krish Veera, Ishfat Abrar Islam, Raiyan Abdul Baten
An objective, face-valid method for scoring idea originality is to measure each idea's statistical infrequency within a population -- an approach long used in creativity research. Yet, computing these frequencies requires manually bucketing idea rephrasings, a process that is subjective, labor-intensive, error-prone, and brittle at scale. We introduce MuseSc
Modern Earth-like Chemical Disequilibrium Biosignatures Are Challenging To Constrain Through Spectroscopic Retrievals
astro-ph.EPAmber Young, Tyler Robinson, Joshua Krissansen-Totton, Edward Schwieterman
Robust exoplanet characterization studies are underway, and the community is looking ahead toward developing observational strategies to search for life beyond our solar system. With the development of life detection approaches like searching for atmospheric chemical species indicative of life, chemical disequilibrium has also been proposed as a potentially
Shuo Zheng, Shuowen Zhang
Beyond diagonal intelligent reflecting surface (BD-IRS) is a new promising IRS architecture for which the reflection matrix is not limited to the diagonal structure as for conventional IRS. In this paper, we study a BD-IRS aided uplink integrated sensing and communication (ISAC) system where sensing is performed in a device-based manner. Specifically, we aim
Yuren Mao, Wenyi Xu, Yuyang Qin, Yunjun Gao
Computed Tomography (CT) scan, which produces 3D volumetric medical data that can be viewed as hundreds of cross-sectional images (a.k.a. slices), provides detailed anatomical information for diagnosis. For radiologists, creating CT radiology reports is time-consuming and error-prone. A visual question answering (VQA) system that can answer radiologists' que
Wei-Lun Huang, Joshua Liu, Davood Tashayyod, Jun Kang
Total Body Photography (TBP) is becoming a useful screening tool for patients at high risk for skin cancer. While much progress has been made, existing TBP systems can be further improved for automatic detection and analysis of suspicious skin lesions, which is in part related to the resolution and sharpness of acquired images. This paper proposes a novel sh
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
cs.CLBohao Wu, Qingyun Wang, Yue Guo
Personalizing jargon detection and explanation is essential for making technical documents accessible to readers with diverse disciplinary backgrounds. However, tailoring models to individual users typically requires substantial annotation efforts and computational resources due to user-specific finetuning. To address this, we present a systematic study of p
Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li
Tabular data, owing to its ubiquitous presence in real-world domains, has garnered significant attention in machine learning research. While tree-based models have long dominated tabular machine learning tasks, the recently proposed deep learning model TabPFN v2 has emerged, demonstrating unparalleled performance and scalability potential. Although extensive
Zihan Chen, Song Wang, Zhen Tan, Jundong Li
In-Context Learning (ICL) empowers Large Language Models (LLMs) to tackle diverse tasks by incorporating multiple input-output examples, known as demonstrations, into the input of LLMs. More recently, advancements in the expanded context windows of LLMs have led to many-shot ICL, which uses hundreds of demonstrations and outperforms few-shot ICL, which relie
Chuan-Ren Chen, Cheng-Wei Chiang, Leon M. G. de la Vega
In this work, we study a dark matter scenario where a dark Z boson possessing mass mixing with the SM Z boson couples to the DM candidate and serves as the portal to the SM. The UV origin of the mass mixing in the form of an extra dark Higgs doublet and a scalar dark singlet provides new exotic scalars which can constitute the final state of DM annihilation
Sangyong Lee, Subo Hwang, Dohoon Kim
In this paper, we propose MADCluster, a novel model-agnostic anomaly detection framework utilizing self-supervised clustering. MADCluster is applicable to various deep learning architectures and addresses the 'hypersphere collapse' problem inherent in existing deep learning-based anomaly detection methods. The core idea is to cluster normal pattern data into
Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang
With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While this offers scalability and flexibility, it also raises a critical, unresolved question: Can LLM judges fairly and robustly evaluate
Yifan Zhang, Xinkui Zhao, Zuxin Wang, Guanjie Cheng
The rapid advancement of large language models has unlocked remarkable capabilities across a diverse array of natural language processing tasks. However, the considerable differences among available LLMs-in terms of cost, performance, and computational demands-pose significant challenges for users aiming to identify the most suitable model for specific tasks
Liang-Yeh Shen, Shi-Xin Fang, Yi-Cheng Lin, Huang-Cheng Chou
This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a
Hidden-Charm Tetraquarks in a Mixture Model: Coupled-Channel Analysis with $c\bar{c}$ and Hadronic Molecular Components
hep-phKotaro Miyake, Yasuhiro Yamaguchi
The nature of the $X(3872)$ and other exotic hadrons has been a subject of extensive investigation since the first observation of the $X(3872)$ in 2003. While various theoretical models have been proposed, including hadronic molecular and compact tetraquark interpretations, some experimental evidence suggests that the $X(3872)$ may be a mixture state of a ha
Understanding the Ly{\alpha} Emission Observed by the Solar Disk Imager Aboard the Advanced Space-based Solar Observatory
astro-ph.SRYiliang Li, Ping Zhang, Zhengyuan Tian, Li Feng
The H I Lyman-alpha (Ly$\alpha$) emission, with a wavelength of 1216 \r{A}, is the brightest solar ultraviolet (UV) line. However, comprehensive observations of the Ly$\alpha$ emission line across the full solar disk remain limited. As part of the ASO-S mission, the Solar Disk Imager (SDI) has successfully captured full-disk images in the Ly$\alpha$ band. Ga
Hon Tik Tse, Siddarth Chandrasekar, Marlos C. Machado
In recent years, the successor representation (SR) has attracted increasing attention in reinforcement learning (RL), and it has been used to address some of its key challenges, such as exploration, credit assignment, and generalization. The SR can be seen as representing the underlying credit assignment structure of the environment by implicitly encoding it
Jisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi
Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. While recent studies have leveraged large language models (LLMs) to handle idioms across various tasks, e.g., idiom-containing sentence generation and idiomatic machine translation, little is known about the underlying mechanisms
Md Ashraf Uddin, Nam H. Chu, Reza Rafeh, Mutaz Barika
Due to its nature of dynamic, mobility, and wireless data transfer, the Internet of Vehicles (IoV) is prone to various cyber threats, ranging from spoofing and Distributed Denial of Services (DDoS) attacks to malware. To safeguard the IoV ecosystem from intrusions, malicious activities, policy violations, intrusion detection systems (IDS) play a critical rol
Henry X. Liu, Xintao Yan, Haowei Sun, Tinghan Wang
Autonomous vehicles (AVs) have significantly advanced in real-world deployment in recent years, yet safety continues to be a critical barrier to widespread adoption. Traditional functional safety approaches, which primarily verify the reliability, robustness, and adequacy of AV hardware and software systems from a vehicle-centric perspective, do not sufficie
Kazuyuki Yagasaki
We study the Kuramoto model (KM) having random natural frequencies and defined on uniform graphs that may be complete, random dense or random sparse. The natural frequencies are assumed to be independent and identically distributed on a bounded interval. In the previous work, the corresponding continuum limit (CL) was proven to approximate stable motions in
Anfeng Xu, Tiantian Feng, So Hyun Kim, Somer Bishop
Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In t
Kai Li, Can Shen, Yile Liu, Jirui Han
The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks are primarily designed for text and fail to capture vulnerabilities introduced by the acoustic properties of audio. We find that significant trustworthiness risks in ALLMs arise from
Zhihang Cai, Xingjun Zhang, Zhendong Tan, Zheng Wei
Large Language Models (LLMs) have demonstrated remarkable proficiency across a wide range of tasks. However, LLMs often require larger batch sizes to enhance throughput or longer context lengths to meet task demands, which significantly increases the memory resource consumption of the Key-Value (KV) cache during inference, becoming a major bottleneck in LLM
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
cs.CVShuchang Ye, Usman Naseem, Mingyuan Meng, Dagan Feng
Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses to clinicians' inquiries regarding medical images. Existing MedVQA models suffered from modality preference bias, where predictions are heavily dominated by one modality while overlooking the other (in MedVQA, us
Anton Erofeev, Balasubramanya T. Nadiga, Ilya Timofeyev
We apply Echo-State Networks to predict time series and statistical properties of the competitive Lotka-Volterra model in the chaotic regime. In particular, we demonstrate that Echo-State Networks successfully learn the chaotic attractor of the competitive Lotka-Volterra model and reproduce histograms of dependent variables, including tails and rare events.
Kentaro Onda, Yosuke Kashiwagi, Emiru Tsunoo, Hayato Futami
Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as intermediate representations for tasks such as automatic speech recognition (ASR). However, discrete tokens are typically obt
Finite temperatures and flat bands: the Hubbard model on three-dimensional Lieb lattices
cond-mat.str-elLucas O. Lima, Julián Faúndez, Natanael C. Costa, Raimundo R. dos Santos
We investigate some thermodynamic and magnetic properties of the Hubbard model on two three-dimensional extensions of the Lieb lattice: the perovskite Lieb lattice (PLL) and the layered Lieb lattice (LLL). Using determinant quantum Monte Carlo (DQMC) simulations alongside Hartree-Fock and cluster mean-field theory (CMFT) approaches, we analyze how flat-band
VIVID: A Novel Approach to Remediation Prioritization in Static Application Security Testing (SAST)
cs.CRNaeem Budhwani, Mohammad Faghani, Hayden Richard
Static Application Security Testing (SAST) enables organizations to detect vulnerabilities in code early; however, major SAST platforms do not include visual aids and present little insight on correlations between tainted data chains. We propose VIVID - Vulnerability Information Via Data flow - a novel method to extract and consume SAST insights, which is to
Directional Convergence, Benign Overfitting of Gradient Descent in leaky ReLU two-layer Neural Networks
cs.LGIchiro Hashimoto
In this paper, we provide sufficient conditions of benign overfitting of fixed width leaky ReLU two-layer neural network classifiers trained on mixture data via gradient descent. Our results are derived by establishing directional convergence of the network parameters and classification error bound of the convergent direction. Our classification error bound
Jesus Sanchez
We provide a recipe for building explicit representations of the real Clifford algebras once an explicit family is given in dimensions $1$ through $4$. We further give an explicit construction of spin coordinate systems for a given real spinor module and use it to explicitly compute the parallel transport of spinor fields. We further highlight some novelties
Indronil Bhattacharjee, Christabel Wayllace
KnowledgeTracing (KT) involves predicting students' knowledge states based on their interactions with Intelligent Tutoring Systems (ITS). A key challenge is the cold start problem, accurately predicting knowledge for new students with minimal interaction data. Unlike prior work, which typically trains KT models on initial interactions of all students and tes
Don N. Page
Alvarez-Dominguez, Garay, Martin-Martinez, and Polo-Gomez have suggested that ``it is not possible to concentrate enough light to precipitate the formation of an event horizon. We argue that the dissipative quantum effects coming from the self-interaction of light (such as vacuum polarization) are enough to prevent any meaningful buildup of energy that could
Ivan V. Morozov
The basis of this work is a simple, extended corollary of Wilson's theorem. This corollary generates many more quotients than those already generated by Wilson's theorem, and it was of interest to derive how they relate to each other and build on the established properties of the original quotients. The most important results that were found were expressions
Kavya Gaddipati
This technical article explores comprehensive strategies for integrating Electrostatic Discharge (ESD) protection diodes and termination resistors in LowVoltage Differential Signaling (LVDS) designs. The article examines critical aspects of protection mechanisms, design considerations, impedance matching, and placement optimization techniques. Through detail
Chaochen Gao, Xing Wu, Zijia Lin, Debing Zhang
High-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long-context instruction data remains proprietary. Human annotation is costly and challenging, while template-based synthesis methods limit scale, diversity, and quality. We introduce
Rikuhei Umemoto, Keisuke Fujii
In many real-world complex systems, the behavior can be observed as a collection of discrete events generated by multiple interacting agents. Analyzing the dynamics of these multi-agent systems, especially team sports, often relies on understanding the movement and interactions of individual agents. However, while providing valuable snapshots, event-based po
Effect of thermal conductivity on the simultaneous formation of a stable region at the top of Earth's core and magnetic field generation over four billion years
astro-ph.EPTakashi Nakagawa, Shin-ichi Takahero, Youhei Sasaki
The possibility of the emergence of a stratified region in the uppermost part of the Earth's outer core with long-term magnetic field generation is assessed, taking into account uncertainties in the thermal conductivity of the Earth's core and the present-day heat flow across the core-mantle boundary (CMB). The radial structures of the Earth's outer core are
Jesus Sanchez, Andres Franco Valiente
Given a contact sub-Riemannian manifold one obtains a non-integrable splitting of the tangent bundle into the directions along the contact distribution and the Reeb field. We generalize the construction of the Bismut superconnection to this non-integrable setting and show that although singularities appear within the superconnection, if one extracts the fini
Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie
A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short: 3D point cloud models lack semantic abstraction, while 2D image encoders struggle with spatial reasoning. To address this, we propose SEM
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
cs.SDZhi Zhong, Akira Takahashi, Shuyang Cui, Keisuke Toyama
Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative m
Pressure evolution of coplanar antiferromagnetism in heavy-fermion Ce$_{2}$CoAl$_{7}$Ge$_{4}$
cond-mat.str-elM. O. Ajeesh, A. O. Scheie, Yu Liu, L. Keller
Ce$_{2}$$M$Al$_{7}$Ge$_{4}$ ($M=$ Co, Ir, Ni or Pd) are heavy-fermion materials and host a variety of ground states ranging from magnetism to non-Fermi liquid behavior. The Co, Ir, and Ni members of the series undergo magnetic ordering with decreasing transition temperatures. In contrast, the Pd compound does not magnetically order down to 0.4 K and shows no
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
cs.CLDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma
The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
cs.CVChaoya Jiang, Yongrui Heng, Wei Ye, Han Yang
Recently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce \textbf{VLM-R$^3$} (\textbf{V
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
cs.SDKentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito
Recently, a method for synthesizing foreign-accented speech only with native speech data using discrete tokens obtained from self-supervised learning (SSL) models was proposed. Considering limited availability of accented speech data, this method is expected to make it much easier to simulate foreign accents. By using the synthesized accented speech as liste
Navid Seidi, Satyaki Roy, Sajal Das
Federated Learning (FL) holds great promise for digital health by enabling collaborative model training without compromising patient data privacy. However, heterogeneity across institutions, lack of sustained reputation, and unreliable contributions remain major challenges. In this paper, we propose a robust, peer-driven reputation mechanism for federated he
The Language of Interoception: Examining Embodiment and Emotion Through a Corpus of Body Part Mentions
cs.CLSophie Wu, Jan Philip Wahle, Saif M. Mohammad
This paper is the first investigation of the connection between emotion, embodiment, and everyday language in a large sample of natural language data. We created corpora of body part mentions (BPMs) in online English text (blog posts and tweets). This includes a subset featuring human annotations for the emotions of the person whose body part is mentioned in
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
cs.CLZirui He, Mingyu Jin, Bo Shen, Ali Payani
Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces. We employ s
Guanghe Li, Junming Zhao, Shengjie Wang, Yang Gao
Robotic insertion is a highly challenging task that requires exceptional precision in cluttered environments. Existing methods often have poor generalization capabilities. They typically function in restricted and structured environments, and frequently fail when the plug and socket are far apart, when the scene is densely cluttered, or when handling novel o
Kaiwen Zhou, Xuandong Zhao, Gaowen Liu, Jayanth Srinivasa
Large Reasoning Models (LRMs) introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs, supervised fine-tuning (SFT), improve safety performance, we fin
Gregoire Fournier, György Turán
Ehrenfeucht-Fra\"iss\'e (EF) games are a basic tool in finite model theory for proving definability lower bounds, with many applications in complexity theory and related areas. They have been applied to study various logics, giving insights on quantifier rank and other logical complexity measures. In this paper, we present an EF game to capture formula size
Pure nematic transition inside the superconducting dome of iron chalcogenide superconductor FeSe$_{1-x}$Te$_x$
cond-mat.supr-conK. Y. Liang, R . Z. Zhang, Z. F. Lin, Z. J. Li
Nematicity and magnetism are prevalent orders in high transition temperature (Tc) superconductors, coexisting in the parent compound of most material families. Quantum fluctuations of nematicity or spin orders are both plausible candidates for mediating unconventional Cooper pairing. Identifying the sole effect of a nematic quantum critical point (QCP) on th
Ultrafast charge-transfer dynamics in Ca$_2$CuO$_2$Cl$_2$ from time-resolved optical reflectivity
cond-mat.str-elHaiyun Huang, Xiu Zhang, Junzhi Zhu, Jianfa Zhao
We employ time-resolved optical reflectivity to investigate the ultrafast dynamics of the charge-transfer gap (CTG) in a parent cuprate compound Ca$_2$CuO$_2$Cl$_2$ (CCOC). We observe a persistent photoinduced red shift of the CTG that lasts up to 1000 ps. The red shift during the slow decay after 10 ps can be well modeled by the localized picture, whereas i
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
cs.SDKentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito
In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is observed, where non-native listeners who share the native language with the speaker understand the speech better compared eve
Mohammad Reza Taesiri, Brandon Collins, Logan Bolton, Viet Dac Lai
Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predicta
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
cs.CVAshim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi
Evaluating image captions requires cohesive assessment of both visual semantics and language pragmatics, which is often not entirely captured by most metrics. We introduce Redemption Score(RS), a novel hybrid framework that ranks image captions by triangulating three complementary signals: (1) Mutual Information Divergence (MID) for global image-text distrib
Ilya I. Bogdanov, Elizaveta Neustroeva, Georgy Sokolov, Alexei Volostnov
The paper is devoted to sufficient conditions for the existence of vertex cuts in simple graphs, where the induced subgraph on the cut vertices belongs to a specified graph class. In particular, we show that any connected graph with $n$ vertices and fewer than $(19n - 28)/8$ edges admits a forest cut. This result improves upon recent bounds, although it does
Shuai Wang, Yizhou Sun, Judea Pearl, Ang Li
Probabilities of causation play a central role in modern decision making. Tian and Pearl first introduced formal definitions and derived tight bounds for three binary probabilities of causation, such as the probability of necessity and sufficiency (PNS). However, estimating these probabilities requires both experimental and observational distributions specif
Linfeng Qi, Zhaoyang Jia, Jiahao Li, Bin Li
Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose \
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
cs.AIJun Rao, Xuebo Liu, Hexuan Deng, Zepeng Lin
In mathematical reasoning, data selection strategies predominantly rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. This misalignment limits the efficiency of Supervised Fine-Tuning and Reinforcement Learning. To bridge this gap, we introduce SAI-DPO (Self-Aware Iterative Data Persistent
Benjamin Schneider, Dongfu Jiang, Chao Du, Tianyu Pang
Long-video understanding has emerged as a crucial capability in real-world applications such as video surveillance, meeting summarization, educational lecture analysis, and sports broadcasting. However, it remains computationally prohibitive for VideoLLMs, primarily due to two bottlenecks: 1) sequential video decoding, the process of converting the raw bit s
Ping Liu, Chi Zhang
To what extent does concept erasure eliminate generative capacity in diffusion models? While prior evaluations have primarily focused on measuring concept suppression under specific textual prompts, we explore a complementary and fundamental question: do current concept erasure techniques genuinely remove the ability to generate targeted concepts, or do they
Exact Expansion Formalism for Transport Properties of Heterogeneous Materials Characterized by Arbitrary Continuous Random Fields
cond-mat.mtrl-sciLiyu Zhong, Sheng Mao
We derive an exact contrast-expansion formalism for the effective conductivity of heterogeneous materials (media) with local properties described by arbitrary continuous random fields, significantly generalizing the widely used binary-field models. The theory produces a rapidly convergent Neumann-series that, upon Gaussian closure via a Hermite expansion, yi
Automated Feedback Loops to Protect Text Simplification with Generative AI from Information Loss
cs.CLAbhay Kumara Sri Krishna Nandiraju, Gondy Leroy, David Kauchak, Arif Ahmed
Understanding health information is essential in achieving and maintaining a healthy life. We focus on simplifying health information for better understanding. With the availability of generative AI, the simplification process has become efficient and of reasonable quality, however, the algorithms remove information that may be crucial for comprehension. In
Mai Lee Chang, Kim Baraka, Greg Trafton, Zach Lalu Vazhekatt
When agents interact with people as part of a team, fairness becomes an important factor. Prior work has proposed fairness metrics based on teammates' capabilities for task allocation within human-agent teams. However, most metrics only consider teammate capabilities from a third-person point of view (POV). In this work, we extend these metrics to include ta
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
cs.SDHongfei Xue, Yufeng Tang, Jun Zhang, Xuelong Geng
Although multilingual automatic speech recognition (ASR) systems have significantly advanced, enabling a single model to handle multiple languages, inherent linguistic differences and data imbalances challenge SOTA performance across all languages. While language identification (LID) models can route speech to the appropriate ASR model, they incur high costs
Xiao Hu, Yang Ye
Robotic manipulation in industrial scenarios such as construction commonly faces uncertain observations in which the state of the manipulating object may not be accurately captured due to occlusions and partial observables. For example, object status estimation during pipe assembly, rebar installation, and electrical installation can be impacted by observati
Yuhao Xue, Zhifei Zhang, Xinyang Jiang, Yifei Shen
Adversarial attacks exploiting unrestricted natural perturbations present severe security risks to deep learning systems, yet their transferability across models remains limited due to distribution mismatches between generated adversarial features and real-world data. While recent works utilize pre-trained diffusion models as adversarial priors, they still e
Yechan Park, Gyuhyeon Pak, Euntai Kim
While most people associate LiDAR primarily with its ability to measure distances and provide geometric information about the environment (via point clouds), LiDAR also captures additional data, including reflectivity or intensity values. Unfortunately, when LiDAR is applied to Place Recognition (PR) in mobile robotics, most previous works on LiDAR-based PR
Tianlai Yang, Mo Xiong, Ming Xue, Xinwei Li
Integer factorization remains a significant challenge for classical computers and is fundamental to the security of RSA encryption. Adiabatic quantum algorithms present a promising solution, yet their practical implementation is limited by the short coherence times of current NISQ devices and quantum simulators. In this work, we apply the chopped random-basi
KNN-SSD: Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization
cs.CLMingbo Song, Heming Xia, Jun Zhang, Chak Tou Leong
Speculative Decoding (SD) has emerged as a widely used paradigm to accelerate the inference of large language models (LLMs) without compromising generation quality. It works by efficiently drafting multiple tokens using a compact model and then verifying them in parallel using the target LLM. Notably, Self-Speculative Decoding proposes skipping certain layer
Liyan Wang, Weixiang Zhou, Cong Wang, Kin-Man Lam
Ultra-high-definition (UHD) image restoration aims to specifically solve the problem of quality degradation in ultra-high-resolution images. Recent advancements in this field are predominantly driven by deep learning-based innovations, including enhancements in dataset construction, network architecture, sampling strategies, prior knowledge integration, and
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
cs.CLBin Xu, Yu Bai, Huashan Sun, Yiguan Lin
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable c
Tanqiu Jiang, Jiacheng Liang, Rongyi Zhu, Jiawei Zhou
Large vision-language models (VLMs) are highly vulnerable to multimodal jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails. In this paper, we present DTR, a novel inference-time defense that mitigates multimodal jailbreak attacks through optimizing the model's key-value (KV) caches. Rather than relying on curated safety-sp
Chongjie Si, Yidan Cui, Fuchao Yang, Xiaokang Yang
Learning from inaccurate annotations has gained significant attention due to the high cost of precise labeling. However, despite the presence of erroneous labels, models trained on noisy data often retain the ability to make accurate predictions. This intriguing phenomenon raises a fundamental yet largely unexplored question: why models can still extract cor
Confirming HSC strong lens candidates with DESI Spectroscopy. I. Project overview and first results
astro-ph.GAYiping Shu, Shen Li
Accurate redshift determinations of both lenses and sources are critical for confirming strong-lens systems and fully realizing their scientific value. However, the thousands of strong-lens candidates now routinely discovered in wide-field imaging surveys make one-by-one follow-up observations impractical. In this work, we investigate the capability and effi
Siu Lun Chau, Michele Caprio, Krikamol Muandet
Quantifying differences between probability distributions is fundamental to statistics and machine learning, primarily for comparing statistical uncertainty. In contrast, epistemic uncertainty -- due to incomplete knowledge -- requires richer representations than those offered by classical probability. Imprecise probability (IP) theory offers such models, ca
Rui Zhang, Na Zhang, Yapeng Zeng, Tao Yang
In this paper, Ore extensions of multiplier Hopf coquasigroups are studied. Necessary and sufficient conditions for the Ore extension of a regular multiplier Hopf coquasigroup to be a multiplier Hopf coquasigroup are given. Furthermore, the isomorphism between two such Ore extensions is discussed.
Ji Guo, Long Zhou, Zhijin Wang, Jiaming He
In recent years, deep learning-based Monocular Depth Estimation (MDE) models have been widely applied in fields such as autonomous driving and robotics. However, their vulnerability to backdoor attacks remains unexplored. To fill the gap in this area, we conduct a comprehensive investigation of backdoor attacks against MDE models. Typically, existing backdoo
B. Bahr-Kalus, D. Parkinson, K. Lodha, E. Mueller
The peak of the matter power spectrum, known as the turnover (TO) scale, is determined by the horizon size at the time of matter-radiation equality. This scale can serve as a standard ruler, independent of other features in the matter power spectrum, such as baryon acoustic oscillations (BAO). Here, we present the first detection of the turnover in the galax
Bolin Chen, Shanzhi Yin, Hanwei Zhu, Lingyu Zhu
In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body
Hongchen Wei, Zhenzhong Chen
Recent advances in Reasoning LLMs (e.g., DeepSeek-R1 and OpenAI-o1) have showcased impressive reasoning capabilities via reinforcement learning. However, extending these capabilities to Multimodal LLMs (MLLMs) is hampered by the prohibitive costs of retraining and the scarcity of high-quality, verifiable multimodal reasoning datasets. This paper introduces F
Cristian Lenart, Satoshi Naito, Daisuke Sagaki, Leonardo C. Mihalcea
We prove an identity for (torus-equivariant) 3-point, genus 0, $K$-theoretic Gromov-Witten invariants of flag manifolds $G/P$, which can be thought of as a replacement for the ``divisor axiom'' in their (torus-equivariant) quantum $K$-theory. This identity enables us to compute these invariants when two insertions are Schubert classes and the other a Schuber
Zirui Pang, Haosheng Tan, Yuhan Pu, Zhijie Deng
Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still suffer from pervasive noisy labels and often contain missing labels due to the co-existing image pattern where multiple classes appear in an image sample. This results in misleadi
Chongjie Si, Kangtao Lv, Jingjing Jiang, Yadao Wang
Model merging offers a training-free alternative to multi-task learning by combining independently fine-tuned models into a unified one without access to raw data. However, existing approaches often rely on heuristics to determine the merging coefficients, limiting their scalability and generality. In this work, we revisit model merging through the lens of l
Le Ma, Shirao Yang, Zihao Wang, Yinggui Wang
The proliferation of large models has intensified the need for efficient data valuation methods to quantify the contribution of individual data providers. Traditional approaches, such as game-theory-based Shapley value and influence-function-based techniques, face prohibitive computational costs or require access to full data and model training details, maki