October 2025 arXiv papers — page 174
Showing 17,301–17,400 of 25,213 papers
ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
cs.CLChung-En Sun, Ge Yan, Akshay Kulkarni, Tsui-Wei Weng
Recent advances in long chain-of-thought (CoT) reasoning have largely prioritized answer accuracy and token efficiency, while overlooking aspects critical to trustworthiness. We argue that usable reasoning systems must be trustworthy, characterized by three properties: interpretability, faithfulness, and reliability. To this end, we propose ReFIne, a new tra
Sang Hun Kim, Jongmin Lee, Dongkyu Park, So Young Lee
We propose a multi-agent framework for modeling artificial consciousness in large language models (LLMs), grounded in psychoanalytic theory. Our \textbf{Psychodynamic Model} simulates self-awareness, preconsciousness, and unconsciousness through agent interaction, guided by a Personalization Module combining fixed traits and dynamic needs. Using parameter-ef
Layout-Aware Parsing Meets Efficient LLMs: A Unified, Scalable Framework for Resume Information Extraction and Evaluation
cs.CLFanwei Zhu, Jinke Yu, Zulong Chen, Ying Zhou
Automated resume information extraction is critical for scaling talent acquisition, yet its real-world deployment faces three major challenges: the extreme heterogeneity of resume layouts and content, the high cost and latency of large language models (LLMs), and the lack of standardized datasets and evaluation tools. In this work, we present a layout-aware
Qian Wang, Xuheng Ding, John Silverman, J. Xavier Prochaska
During cosmic noon ($z\sim1-3$), when both star formation and black hole growth peaked, galaxy mergers are predicted to trigger dual active galactic nuclei (AGN) that eventually coalesce as supermassive black hole (SMBH) binaries. However, observations of dual quasars with sub-5 kpc separations-the critical phase preceding final coalescence-have remained cha
Model-Assisted and Human-Guided: Perceptions and Practices of Software Professionals Using LLMs for Coding
cs.SEItalo Santos, Cleyton Magalhaes, Ronnie de Souza Santos
Large Language Models have quickly become a central component of modern software development workflows, and software practitioners are increasingly integrating LLMs into various stages of the software development lifecycle. Despite the growing presence of LLMs, there is still a limited understanding of how these tools are actually used in practice and how pr
Ankit Yadav, Ritumoni Sarma, Anuj Kumar Bhagat
In \cite{shi2022few-weight}, Shi and Li studied $\mathcal{C}_D$-codes over the ring $\mathcal{R}:=\mathbb{F}_2[x,y]/\langle x^2, y^2, xy-yx\rangle$ and their binary Gray images, where $D$ is derived using certain simplicial complexes. We study the subfield codes $\mathcal{C}_{D}^{(2)}$ of $\mathcal{C}_{D}$-codes over $\mathcal{R},$ where $D$ is as in \cite{s
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
cs.SEJiale Guo, Suizhi Huang, Mei Li, Dong Huang
The integration of Large Language Models (LLMs) into software engineering has driven a transition from traditional rule-based systems to autonomous agentic systems capable of solving complex problems. However, systematic progress is hindered by a lack of comprehensive understanding of how benchmarks and solutions interconnect. This survey addresses this gap
Lesion-Aware Post-Training of Latent Diffusion Models for Synthesizing Diffusion MRI from CT Perfusion
cs.CVJunhyeok Lee, Hyunwoong Kim, Hyungjin Chung, Heeseong Eom
Image-to-Image translation models can help mitigate various challenges inherent to medical image acquisition. Latent diffusion models (LDMs) leverage efficient learning in compressed latent space and constitute the core of state-of-the-art generative image models. However, this efficiency comes with a trade-off, potentially compromising crucial pixel-level d
Haoran Sun, Zekun Zhang, Shaoning Zeng
One of the key factors influencing the reasoning capabilities of LLM-based agents is their ability to leverage long-term memory. Integrating long-term memory mechanisms allows agents to make informed decisions grounded in historical interactions. While recent advances have significantly improved the storage and retrieval components, by encoding memory into d
Sensing, Detection and Localization for Low Altitude UAV: A RF-Based Framework via Multiple BSs Collaboration
eess.SYTianhao Liang, Mu Jia, Tingting Zhang, Junting Chen
The rapid growth of the low-altitude economy has resulted in a significant increase in the number of Low, slow, and small (LLS) unmanned aerial vehicles (UAVs), raising critical challenges for secure airspace management and reliable trajectory planning. To address this, this paper proposes a cooperative radio-frequency (RF) detection and localization framewo
Song Li, Kelei Tian, Zhiwei Wu
The algebraic structures of integrable hierarchies play an important role in the study of soliton equations. In this paper, we use splitting theory to give a matrix representation of a constrained CKP hierarchy, which can be considered as a generalization of the $\hat{A}_{2n}^{(2)}$-KdV hierarchy and the constrained KP hierarchy. An equivalent construction i
Polarization Dependence of Excess Loss of Amorphous Coating Supermirror in Optical Region for Cavity Ringdown Spectroscopy
physics.opticsMitsunori Araki, Kohsuke Suma
A long optical path length is critical in achieving sensitive spectroscopy. For cavity ringdown spectroscopy, a cavity consisting of two supermirrors provides a long path length, where high reflectance of the supermirrors results from their slight excess loss. In the case of a crystal coating supermirror, the excess loss has been suggested to depend on polar
Chenxu Wang, Hao Li, Yiqun Zhang, Linyao Chen
Large language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a candidate model pool. However, routing performance relies on accurate model representations, and adding new models typically requires retraining, limiting scalability. To address t
Ce Xu
In this paper, we employ the theories and techniques of hypergeometric functions to provide two distinct proofs of the conjectured identities involving multiple Ap\'ery-like series with central binomial coefficients and multiple harmonic star sums, as recently proposed by Gen\v{c}ev and Rucki. Furthermore, we establish several more general identities for mul
Muhammad Ali Shafique, Kanwal Mehreen, Muhammad Arham, Maaz Amjad
Developing a high-performing large language models (LLMs) for low-resource languages such as Urdu, present several challenges. These challenges include the scarcity of high-quality datasets, multilingual inconsistencies, and safety concerns. Existing multilingual LLMs often address these issues by translating large volumes of available data. However, such tr
Dildar Ali, Rajibul Islam, Suman Banerjee
Billboard Advertisement has emerged as an effective out-of-home advertisement technique where the goal is to select a limited number of slots and play advertisement content over there with the hope that this will be observed by many people, and effectively, a significant number of them will be influenced towards the brand. Given a trajectory and a billboard
Joonghyuk Hahn, Soohan Lim, Yo-Sub Han
Predicting the complexity of source code is essential for software development and algorithm analysis. Recently, Baik et al. (2025) introduced CodeComplex for code time complexity prediction. The paper shows that LLMs without fine-tuning struggle with certain complexity classes. This suggests that no single LLM excels at every class, but rather each model sh
Xu Yang, Salvatore Rastelli, Alexander Jung
We study federated clustering, where interconnected devices collaboratively cluster the data points of private local datasets. Focusing on hard clustering via the k-means principle, we formulate federated k-means as an instance of generalized total variation minimization (GTVMin). This leads to a federated k-means algorithm in which each device updates its l
Spatio-Temporal Graph Convolutional Networks for EV Charging Demand Forecasting Using Real-World Multi-Modal Data Integration
cs.LGJose Tupayachi, Mustafa C. Camur, Kevin Heaslip, Xueping Li
Transportation remains a major contributor to greenhouse gas emissions, highlighting the urgency of transitioning toward sustainable alternatives such as electric vehicles (EVs). Yet, uneven spatial distribution and irregular utilization of charging infrastructure create challenges for both power grid stability and investment planning. This study introduces
Transfer Learning-Enabled Efficient Raman Pump Tuning under Dynamic Launch Power for C+L Band Transmission
eess.SPJiaming Liu, Rui Wang, JinJiang Li, Hong Lin
We propose a transfer learning-enabled Transformer framework to simultaneously realize accurate modeling and Raman pump design in C+L-band systems. The RMSE for modeling and peak-to-peak GSNR variation/deviation is within 0.22 dB and 0.86/0.1 dB, respectively.
Ting Yu, Hongyu Gong, Zhifu Gao, Zhongli Zhang
A systematic study of 80 known pulsars observed at 185 MHz has been conducted using archival incoherent-sum data from the Murchison Widefield Array (MWA). The dataset comprises 48 drift-scan observations from the MWA Voltage Capture System, covering approximately 30,000 square degrees of sky with sensitivities reaching about 8 mJy in the deepest regions. An
Manojit Chakraborty, Madhusudan Ghosh, Rishabh Gupta
In the domain of software development, LLMs have been utilized to automate tasks such as code translation, where source code from one programming language is translated to another while preserving its functionality. However, LLMs often struggle with long source codes that don't fit into the context window, which produces inaccurate translations. To address t
Imaging of Gate-Controlled Suppression of Superconductivity via the Meissner Effect
cond-mat.mes-hallP. J. Scheidegger, K. J. Knapp, U. Ognjanovic, L. Ruf
It was recently discovered that supercurrents flowing through thin superconducting nanowires can be quenched by a gate voltage. This gate control of supercurrents, known as the GCS effect, could enable superconducting transistor logic. Here, we report that the GCS also manifests in a suppression of Meissner screening, establishing the phenomenon as a genuine
Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
cs.AISang Hun Kim, Jongmin Lee, Dongkyu Park, So Young Lee
Human consciousness is still a concept hard to define with current scientific understanding. Although Large Language Models (LLMs) have recently demonstrated significant advancements across various domains including translation and summarization, human consciousness is not something to imitate with current upfront technology owing to so-called hallucination.
MAKO: Meta-Adaptive Koopman Operators for Learning-based Model Predictive Control of Parametrically Uncertain Nonlinear Systems
eess.SYMinghao Han, Kiwan Wong, Adrian Wing-Keung Law, Xunyuan Yin
In this work, we propose a meta-learning-based Koopman modeling and predictive control approach for nonlinear systems with parametric uncertainties. An adaptive deep meta-learning-based modeling approach, called Meta Adaptive Koopman Operator (MAKO), is proposed. Without knowledge of the parametric uncertainty, the proposed MAKO approach can learn a meta-mod
Atomistic origin of low thermal conductivity in quaternary chalcogenides Cu(Cd, Zn)$_2$InTe$_4$
cond-mat.mtrl-sciNirmalya Jana, Amit Agarwal, Koushik Pal
Crystalline semiconductors with intrinsically low lattice thermal conductivity ($\mathcal{K}$) are vital for device applications such as barrier coatings and thermoelectrics. Quaternary chalcogenide semiconductors such as CuCd$_2$InTe$_4$ and CuZn$_2$InTe$_4$ are experimentally shown to exhibit low $\mathcal{K}$, yet its microscopic origin remains poorly und
Low Complexity Detector for XL-MIMO Uplink: A Cross Splitting Based Information Geometry Approach
cs.ITWenjun Zhang, An-An Lu, Xiqi Gao
In this paper, we propose the cross splitting based information geometry approach (CS-IGA), a novel and low complexity iterative detector for uplink signal recovery in extralarge-scale MIMO (XL-MIMO) systems. Conventional iterative detectors, such as the approximate message passing (AMP) algorithm and the traditional information geometry algorithm (IGA), suf
Wenyi Wu, Kun Zhou, Ruoxin Yuan, Vivian Yu
We study how to endow GUI agents with scalable memory that help generalize across unfamiliar interfaces and long-horizon tasks. Prior GUI agents compress past trajectories into text tokens, which balloons context length and misses decisive visual cues (e.g., exact widget size and position). We propose a continuous memory that encodes each GUI trajectory into
Hansol Hong, Sangwon Lee, Jang-Ho Ha, Sung-June Chu
Untargeted metabolomics using LC-MS/MS offers the potential to comprehensively profile the chemical diversity of biological samples. However, the process is fundamentally limited by the "identification bottleneck," where only a small fraction of detected features can be annotated using existing spectral libraries, leaving the majority of data uncharacterized
Sicheol Sung, Joonghyuk Hahn, Yo-Sub Han
Regular expressions (regexes) are foundational to modern computing for critical tasks like input validation and data parsing, yet their ubiquity exposes systems to regular expression denial of service (ReDoS), a vulnerability requiring automated repair methods. Current approaches, however, are hampered by a trade-off. Symbolic, rule-based system are precise
Exploring Single Domain Generalization of LiDAR-based Semantic Segmentation under Imperfect Labels
cs.CVWeitong Kong, Zichao Zeng, Di Wen, Jiale Wei
Accurate perception is critical for vehicle safety, with LiDAR as a key enabler in autonomous driving. To ensure robust performance across environments, sensor types, and weather conditions without costly re-annotation, domain generalization in LiDAR-based 3D semantic segmentation is essential. However, LiDAR annotations are often noisy due to sensor imperfe
Jerome Bolte, Quoc-Tung Le, Edouard Pauwels
Ample empirical evidence in deep neural network training suggests that a variety of optimizers tend to find nearly global optima. In this article, we adopt the reversed perspective that convergence to an arbitrary point is assumed rather than proven, focusing on the consequences of this assumption. From this viewpoint, in line with recent advances on the edg
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
cs.CLChi Seng Cheang, Hou Pong Chan, Wenxuan Zhang, Yang Deng
Recent work suggests that LLMs "know what they don't know", positing that hallucinated and factually correct outputs arise from distinct internal processes and can therefore be distinguished using internal signals. However, hallucinations have multifaceted causes: beyond simple knowledge gaps, they can emerge from training incentives that encourage models to
Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language
cs.CLAdity Khisa, Nusrat Jahan Lia, Tasnim Mahfuz Nafis, Zarif Masud
As an Indo-Aryan language with limited available data, Chakma remains largely underrepresented in language models. In this work, we introduce a novel corpus of contextually coherent Bangla-transliterated Chakma, curated from Chakma literature, and validated by native speakers. Using this dataset, we fine-tune six encoder-based transformer models, including m
A Scalable, Privacy-Preserving Decentralized Identity and Verifiable Data Sharing Framework based on Zero-Knowledge Proofs
cs.CRHui Yuan
With the proliferation of decentralized applications (DApps), the conflict between the transparency of blockchain technology and user data privacy has become increasingly prominent. While Decentralized Identity (DID) and Verifiable Credentials (VCs) provide a standardized framework for user data sovereignty, achieving trusted identity verification and data s
Paul Bouchaud, Pedro Ramaciotti
Large language models rely on web-scraped text for training; concurrently, content creators are increasingly blocking AI crawlers to retain control over their data. We analyze crawler restrictions across the top one million most-visited websites since 2023 and examine their potential downstream effects on training data composition. Our analysis reveals growi
Keno Harada, Lui Yoshida, Takeshi Kojima, Yusuke Iwasawa
The performance of Large Language Models (LLMs) is highly sensitive to the prompts they are given. Drawing inspiration from the field of prompt optimization, this study investigates the potential for enhancing Automated Essay Scoring (AES) by refining the scoring rubrics used by LLMs. Specifically, our approach prompts models to iteratively refine rubrics by
A Davydov Ansatz approach to accurate system-bath dynamics in the presence of multiple baths with distinct temperatures
quant-phChenlin Ma, Fulu Zheng, Kewei Sun, Lu Wang
We perform benchmark simulations using the time-dependent variational approach with the multiple Davydov Ansatz (mDA) to study realtime nonequilibrium dynamics in a single qubit model coupled to two thermal baths with distinct temperatures. A broad region of the parameter space has been investigated, accompanied by a detailed analysis of the convergence beha
Shiyuan Guo, Henry Sleight, Fabien Roger
Detecting harmful AI actions is important as AI agents gain adoption. Chain-of-thought (CoT) monitoring is one method widely used to detect adversarial attacks and AI misalignment. However, attackers and misaligned models might evade CoT monitoring through ciphered reasoning: reasoning hidden in encrypted, translated, or compressed text. To assess this risk,
Katie Clinch, Serge Gaspers, Tao Zixu He, Simon Mackenzie
This work introduces two techniques for the design and analysis of branching algorithms, illustrated through the case study of the Vertex Cover problem. First, we present a method for automatically generating branching rules through a systematic case analysis of local structures. Second, we develop a new technique for analyzing randomized branching algorithm
Michele Bertolini, Silvia Tonghini, Marco Rossoni, Marina Carulli
The human nose exhibits a huge variation in shape among individuals. All these variants alter the airflow through the nasal cavity and can impact how we smell odors. To acquire a better understanding of physiological and pathological functioning, it is important to study the effects of these modifications. SSM, or Statistical Shape Modelling, is a widely use
On the torsion-free nilpotent fundamental groups of smooth quasi-projective varieties of rank up to seven
math.AGTaito Shimoji
Let $X$ be a smooth quasi-projective variety. Assume that the (topological) fundamental group $\pi_1(X, x)$ is torsion-free nilpotent. We show that if the first Betti number $b_1(X) \le 3$, then $\pi_1(X, x)$ is isomorphic to either $\mathbb{Z}^n$ for $n = 1, 2, 3$, a lattice in the Heisenberg group $H_3(\mathbb{R})$ or $\mathbb{R} \times H_3(\mathbb{R})$. M
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
cs.CYYuqi Bai, Tianyu Huang, Kun Sun, Yuting Chen
This research focuses on using large language models (LLMs) to simulate social experiments, exploring their ability to emulate human personality in virtual persona role-playing. The research develops an end-to-end evaluation framework, including individual-level analysis of stability and identifiability, as well as population-level analysis called progressiv
Louis Bahrman, Mathieu Fontaine, Gaël Richard
This paper introduces a new training strategy to improve speech dereverberation systems in an unsupervised manner using only reverberant speech. Most existing algorithms rely on paired dry/reverberant data, which is difficult to obtain. Our approach uses limited acoustic information, like the reverberation time (RT60), to train a dereverberation system. Expe
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
cs.LGMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful knowledge or remotely triggering malicious actions, respectively) are typically evaluated either against a static set of harmful attack strings, or against computationally weak op
Clément Morand, Anne-Laure Ligozat, Aurélie Névéol
Recent Machine Learning (ML) approaches have shown increased performance on benchmarks but at the cost of escalating computational demands. Hardware, algorithmic and carbon optimizations have been proposed to curb energy consumption and environmental impacts. Can these strategies lead to sustainable ML model training? Here, we estimate the environmental impa
Hamed Mahdavi, Pouria Mahdavinia, Samira Malek, Pegah Mohammadipour
State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 of 6 problems. Given this progress, we assess how well these models can grade proofs: detecting errors, judging their severity, and assigning fair scores beyond binary correctness.
From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
cs.CLRatna Kandala, Akshata Kishore Moharir, Divya Arvinda Nayak
Explainable Artificial Intelligence (XAI) has been presented as the critical component for unlocking the potential of machine learning in mental health screening (MHS). However, a persistent lab-to-clinic gap remains. Current XAI techniques, such as SHAP and LIME, excel at producing technically faithful outputs such as feature importance scores, but fail to
Zekai Chen, Xunkai Li, Sirui Zhang, Henan Sun
De novo ligand design is a fundamental task that seeks to generate protein or molecule candidates that can effectively dock with protein receptors and achieve strong binding affinity entirely from scratch. It holds paramount significance for a wide spectrum of biomedical applications. However, most existing studies are constrained by the \textbf{Pseudo De No
Measurement of the $e^+e^-\to\eta\gamma$ cross section near the $\phi(1020)$ resonance with the SND detector
hep-exSND Collaboration, M. N. Achasov, A. E. Alizzi, A. Yu. Barnyakov
In the experiment with the SND detector at the VEPP-2000 $e^+e^-$ collider, the $e^+e^-\to\eta\gamma$ cross section is measured in the energy range $E = 980 - 1060$ MeV. The measurement is carried out in the $\eta\to 2\gamma$ decay mode. Data with an integrated luminosity of 73 pb$^{-1}$ collected in 2018 and 2024 are used in the analysis. The measured cross
Ian Harshbarger, Calvin Chidambaram
Most neural network scheduling research focuses on optimizing static, end-to-end models of fixed width, overlooking dynamic approaches that adapt to heterogeneous hardware and fluctuating runtime conditions. We present Slim Scheduler, a hybrid scheduling framework that integrates a Proximal Policy Optimization (PPO) reinforcement learning policy with algorit
Rui Bu, Haofeng Zhong, Wenzheng Chen, Yangyan Li
Large models based on the Transformer architecture are susceptible to extreme-token phenomena, such as attention sinks and value-state drains. These issues, which degrade model performance, quantization fidelity, and interpretability, arise from a problematic mutual reinforcement mechanism where the model learns an inefficient 'no-op' behavior by focusing at
Zongcai Du, Guilin Deng, Xiaofeng Guo, Xin Gao
Recent progress in diffusion-based Singing Voice Synthesis (SVS) demonstrates strong expressiveness but remains limited by data scarcity and model scalability. We introduce a two-stage pipeline: a compact seed set of human-sung recordings is constructed by pairing fixed melodies with diverse LLM-generated lyrics, and melody-specific models are trained to syn
Shota Saito, Hamdi Joudeh
This paper considers the problem of soft guessing under a logarithmic loss distortion measure while allowing errors. We find an optimal guessing strategy, and derive single-shot upper and lower bounds for the minimal guessing moments as well as an asymptotic expansion for i.i.d. sources. These results are extended to the case where side information is availa
LitE-SQL: A Lightweight and Efficient Text-to-SQL Framework with Vector-based Schema Linking and Execution-Guided Self-Correction
cs.CLShengmin Piao, Jieun Lee, Sanghyun Park
The Text-to-SQL task translates natural language questions into SQL queries, enabling intuitive database interaction for non-experts. While recent methods leveraging Large Language Models (LLMs) achieve strong performance, their reliance on proprietary models raise concerns about deployment feasibility and data privacy. In this work, we introduce LitE-SQL, a
Daniel A. Williams, Airlie Chapman, Daniel R. Little, Chris Manzie
Advances in the control of autonomous systems have accompanied an expansion in the potential applications for autonomous robotic systems. The success of applications involving humans depends on the quality of interaction between the autonomous system and the human supervisor, which is particularly affected by the degree of trust that the supervisor places in
Xiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu
In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that facilitates higher autoregressive generation quality with faste
Yincen Qu, Huan Xiao, Feng Li, Gregory Li
Travel planning is a valuable yet complex task that poses significant challenges even for advanced large language models (LLMs). While recent benchmarks have advanced in evaluating LLMs' planning capabilities, they often fall short in evaluating feasibility, reliability, and engagement of travel plans. We introduce a comprehensive benchmark for travel planni
Yipu Zhang, Chaofang Ma, Jinming Ge, Lin Jiang
Neural Radiance Field (NeRF) has emerged as a promising 3D reconstruction method, delivering high-quality results for AR/VR applications. While quantization methods and hardware accelerators have been proposed to enhance NeRF's computational efficiency, existing approaches face crucial limitations. Current quantization methods operate without considering har
Leijie Wang, Kathryn Yurechko, Amy X. Zhang
While LLMs now enable users to create content classifiers easily through natural language, automatic prompt optimization techniques are often necessary to create performant classifiers. However, such techniques can fail to consider how social media users want to evolve their filters over the course of usage, including desiring to steer them in different ways
On Epistemic Uncertainty of Visual Tokens for Object Hallucinations in Large Vision-Language Models
cs.CVHoigi Seo, Dong Un Kang, Hyunjin Cho, Joohoon Lee
Large vision-language models (LVLMs), which integrate a vision encoder (VE) with a large language model, have achieved remarkable success across various tasks. However, there are still crucial challenges in LVLMs such as object hallucination, generating descriptions of objects that are not in the input image. Here, we argue that uncertain visual tokens withi
Changsheng Wang, Yihua Zhang, Dennis Wei, Jinghan Jia
Large language models (LLMs) exhibit remarkable generative capabilities but raise ethical and security concerns by memorizing sensitive data, reinforcing biases, and producing harmful content. These risks have spurred interest in LLM unlearning, the task of removing knowledge associated with undesirable data from pre-trained models. However, most existing me
Chandra Thapa, Surya Nepal
Future G network's new reality is a widespread cyber-physical environment created by Integrated Sensing and Communication (ISAC). It is a crucial technology that transforms wireless connections into ubiquitous sensors. ISAC unlocks transformative new capabilities, powering autonomous systems, augmented human sensing, and next-generation immersive application
Zikang Dong, Yutong Song, Ruihua Wang, Shengbo Zhao
In this article, we investigate the conditional large values of quadratic Dirichlet character sums. We prove an Omega result for quadratic character sums under the assumption of the generalized Riemann hypothesis.
Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
cs.CLYutao Mou, Xiaoling Zhou, Yuxiao Luo, Shikun Zhang
Safety alignment is essential for building trustworthy artificial intelligence, yet it remains challenging to enhance model safety without degrading general performance. Current approaches require computationally expensive searches for the optimal proportion of safety-critical and general-purpose data to balance safety and general performance, incurring high
D Ellis Hershkowitz, Richard Z Huang
In length-constrained minimum spanning tree (MST) we are given an $n$-node graph $G = (V,E)$ with edge weights $w : E \to \mathbb{Z}_{\geq 0}$ and edge lengths $l: E \to \mathbb{Z}_{\geq 0}$ along with a root node $r \in V$ and a length-constraint $h \in \mathbb{Z}_{\geq 0}$. Our goal is to output a spanning tree of minimum weight according to $w$ in which e
Jingyu Zhou, Lu Ma, Hao Liang, Chengyu Shen
Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group Relative Policy Optimization (GRPO) has emerged as the de facto approach for RLVR, inspiring numerous variants. However, our mathematical analysis reveals that these methods are f
Wen-Yuan Ai, Peisi Huang, Ke-Pan Xie
We propose a new nonthermal leptogenesis mechanism triggered by the cosmic first-order phase transition. The Standard Model is extended with two generations of TeV-scale vectorlike leptons. The lighter generation gives rise to an inverse electroweak phase transition of the Higgs field at $T\sim200~{\rm GeV}$, restoring the symmetry, and resulting in relativi
Ziyi Wang, Nan Jiang, Guang Lin, Qifan Song
Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization individually, often resulting in suboptimal compression rates to preserve acceptable performance drops. We introduce a unified framework for simultaneous pruning and low-bit quantization
Xunda Sun, Xin Wang, Fangzhou Jiang, Houjun Mo
We employ the high-redshift suite of FIRE-2 cosmological hydrodynamic zoom-in simulations to investigate the evolution of gas-phase metallicity radial gradients in galaxies in the epoch of reionization (EoR). Our sample consists of 22 galaxies spanning the redshift range $z \sim 10-5$. We find that galaxies at $z\sim10$ exhibit a median metallicity gradient
Spandan Garg, Benjamin Steenhoek, Yufan Huang
Current benchmarks for evaluating software engineering agents, such as SWE-Bench Verified, are predominantly derived from GitHub issues and fail to accurately reflect how developers interact with chat-based coding assistants in integrated development environments (IDEs). We posit that this mismatch leads to a systematic overestimation of agent's capabilities
CogniDir: Combating Cognitive Malicious Comments via Adaptive Distributional Learning for Robust Fake News Detection
cs.LGZhao Tong, Chunlin Gong, Yimeng Gu, Haichao Shi
The proliferation of Large Language Models (LLMs) has enabled a new class of psychologically grounded malicious comments, shifting fake news attacks from surface-level textual noise to deep cognitive and logical manipulation. This shift severely undermines existing detectors, which conventionally rely on static attack assumptions and fixed training distribut
Liang Zhang, Chenhao Pan, Jinze He, Danni Chen
Photonic time crystals (PTCs) - dielectric media whose permittivity is periodically modulated in time - map to a Dirac equation with an imaginary mass, opening a momentum gap (k-gap) where modes grow or decay exponentially. Here, we introduce a sequence of temporal Jackiw-Rebbi kinks that act as a programmable flip of the Dirac mass, exchanging the amplifyin
Yao Teng, Fuyun Wang, Xian Liu, Zhekai Chen
As a new paradigm of visual content generation, autoregressive text-to-image models suffer from slow inference due to their sequential token-by-token decoding process, often requiring thousands of model forward passes to generate a single image. To address this inefficiency, we propose Speculative Jacobi-Denoising Decoding (SJD2), a framework that incorporat
Yang Shi, Jingchao Wang, Liangsi Lu, Mingxuan Huang
Positron Emission Tomography (PET) is crucial in medicine, but its clinical use is limited due to high signal-to-noise ratio doses increasing radiation exposure. Lowering doses increases Poisson noise, which current denoising methods fail to handle, causing distortions and artifacts. We propose a Poisson Consistent U-Net (PC-UNet) model with a new Poisson Va
Xiaolong Tu, Dawei Chen, Kyungtae Han, Onur Altintas
Hardware-Aware Neural Architecture Search (HW-NAS) has emerged as a powerful tool for designing efficient deep neural networks (DNNs) tailored to edge devices. However, existing methods remain largely impractical for real-world deployment due to their high time cost, extensive manual profiling, and poor scalability across diverse hardware platforms with comp
Creation, Critique, and Consumption: Exploring Generative AI Descriptions for Supporting Blind and Low Vision Professionals with Visual Tasks
cs.HCLucy Jiang, Lotus Zhang, Leah Findlater
Many blind and low vision (BLV) people are excluded from professional roles that may involve visual tasks due to access barriers and persisting stigmas. Advancing generative AI systems can support BLV people through providing contextual and personalized visual descriptions for creation, critique, and consumption. In this workshop paper, we provide design sug
Mandira Roy, Novarun Deb, Nabendu Chaki, Agostino Cortesi
Software systems are a significant contributor to global sustainability concerns, demanding that environmental, social, technical, and economic factors be systematically addressed from the initial requirements engineering phase. Although existing research provides various sustainability requirements (SRs), these contributions are often fragmented, specific t
Liam Judd McClelland
Heat is a physical manifestation of entropy, where the removal of entropy from a thermal energy reservoir permits the conversion of heat into work. This entropy transfer is facilitated by the cold thermal energy reservoir in typical heat engines. Recent developments in quantum heat engines that operate between thermal energy and spin angular momentum reservo
Lan Zhang, Marco Valentino, André Freitas
Autoformalization serves a crucial role in connecting natural language and formal reasoning. This paper presents MASA, a novel framework for building multi-agent systems for autoformalization driven by Large Language Models (LLMs). MASA leverages collaborative agents to convert natural language statements into their formal representations. The architecture o
Qixiang Yin, Huanjin Yao, Jianghao Chen, Jiaxing Huang
Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, they encounter challenges in terms of reasoning efficiency, large model size and overthinking. However, existing lightweight MLLMs lack the capability to balance high efficiency and performance at a small scale. To this end, we propose Tiny-R1V,
Xuan Lu, Haohang Huang, Rui Meng, Yaohui Jin
Document reranking is a key component in information retrieval (IR), aimed at refining initial retrieval results to improve ranking quality for downstream tasks. Recent studies--motivated by large reasoning models (LRMs)--have begun incorporating explicit chain-of-thought (CoT) reasoning into LLM-based rerankers. However, the effectiveness of such reasoning
Jionghao Lou, Jian Zhang, Zhongmei Li, Lanlan Chen
The training of deep learning models in seizure prediction requires large amounts of Electroencephalogram (EEG) data. However, acquiring sufficient labeled EEG data is difficult due to annotation costs and privacy constraints. Federated Learning (FL) enables privacy-preserving collaborative training by sharing model updates instead of raw data. However, due
Keng Hao Ooi, Nguyen Cong Phuc
We verify a conjecture of D. R. Adams on a capacitary strong type inequality that generalizes the classical capacitary strong type inequality of V. G. Maz'ya. As a result, we characterize related function spaces as K\"othe duals to a class of Sobolev multiplier type spaces. Moreover, using tools from nonlinear potential theory, weighted norm inequalities, an
Mandira Roy, Novarun Deb, Nabendu Chaki, Agostino Cortesi
The rapid expansion of software development has significant environmental, technical, social, and economic impacts. Achieving the United Nations Sustainable Development Goals by 2030 compels developers to adopt sustainable practices. Existing methods mostly offer high-level guidelines, which are time-consuming to implement and rely on team adaptability. More
Mehmet Fatih Ozkan, Dennis Kibalama, Jacob Paugh, Marcello Canova
Connected and Automated Vehicles (CAVs) offer significant potential for improving energy efficiency and lowering vehicle emissions through eco-driving technologies. Control algorithms in CAVs leverage look-ahead route information and Vehicle-to-Everything (V2X) communication to optimize vehicle performance. However, existing eco-driving strategies often negl
Uncolorable Examples: Preventing Unauthorized AI Colorization via Perception-Aware Chroma-Restrictive Perturbation
cs.CVYuki Nii, Futa Waseda, Ching-Chun Chang, Isao Echizen
AI-based colorization has shown remarkable capability in generating realistic color images from grayscale inputs. However, it poses risks of copyright infringement -- for example, the unauthorized colorization and resale of monochrome manga and films. Despite these concerns, no effective method currently exists to prevent such misuse. To address this, we int
Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu
Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but are vulnerable to corpus poisoning and contamination attacks, which can compromise output integrity. Existing defenses often apply aggressive filtering, leading to unnecessary loss of valuable information and reduced reliability in generation. To add
Zichuan Wang, Bo Peng, Songlin Yang, Zhenchen Tang
Although recent text-to-image (T2I) models have significantly improved the overall visual quality of generated images, they still struggle in the generation of accurate details in complex local regions, especially human hands. Generated hands often exhibit structural distortions and unrealistic textures, which can be very noticeable even when the rest of the
Tengyang Liu, Yisong Yang
Motivated by the century-old problem of modeling the electron as a pointlike particle with finite self energy, we develop a new class of nonlinear perturbations of Maxwell's electrodynamics inspired by, but distinct from, the Born--Infeld theory. A hallmark of our construction is that the effective radius of an electric point charge can be reduced arbitraril
The Influence of Magnetic Complexity of Active Regions on Solar Wind Properties During Solar Cycles 23 and 24
astro-ph.SRXinzheng Shi, Hui Fu, Zhenghua Huang, Limei Yan
Linking solar wind properties to the activities and characteristics of its source regions can enhance our understanding of its origin and generation mechanisms. Using the Mount Wilson magnetic classification (MWMC), we categorize all active regions (ARs) between 1999 and 2020 into three groups: alpha, beta, and complex ARs. Subsequently, we classify the near
Bayesian Active Learning for Bayesian Model Updating: the Art of Acquisition Functions and Beyond
stat.COJingwen Song, Pengfei Wei
Estimating posteriors and the associated model evidences, with desired accuracy and affordable computational cost, is a core issue of Bayesian model updating, and can be of great challenge given expensive-to-evaluate models and posteriors with complex features such as multi-modalities of unequal importance, nonlinear dependencies and high sharpness. Bayesian
Fan Xia, K. C. Gary Chan, Emily Voldal, Avi Kenny
Stepped wedge designs (SWDs) are increasingly used to evaluate longitudinal cluster-level interventions but pose substantial challenges for valid inference. Because crossover times are randomized, intervention effects are intrinsically confounded with secular time trends, while heterogeneity across clusters, complex correlation structures, baseline covariate
Communication System Design using Synthetic Photoisomerizable Azobenzene-Regulated K+(SPARK) channel
q-bio.BMTaha Sajjad, Andrew W. Eckford
Biomolecules exhibit a remarkable property of transforming signals from their environment. This paper presents a communication system design using a light-modulated protein channel: Synthetic Photoisomerizable Azobenzene-regulated K+ (SPARK). Our approach involves a comprehensive design incorporating the SPARK-based receiver, encoding methods, modulation tec
Zhenyu Wang, Mahathir Monjur, Shahriar Nirjon
In mmWave-based pose estimation, sparse signals and weak reflections often cause models to infer body joints from statistical priors rather than sensor data. While prior knowledge helps in learning meaningful representations, over-reliance on it degrades performance in downstream tasks like gesture and activity recognition. In this paper, we introduce mmJoin
Suraj Kumar Sahoo, Narayanan C Krishnan
Learned Optimizers (LOs), a type of Meta-learning, have gained traction due to their ability to be parameterized and trained for efficient optimization. Traditional gradient-based methods incorporate explicit regularization techniques such as Sharpness-Aware Minimization (SAM), Gradient-norm Aware Minimization (GAM), and Gap-guided Sharpness-Aware Minimizati
Yeqing Yang, Le Xu, Lixia Tian
Accurate segmentation of 3D medical images is critical for clinical applications like disease assessment and treatment planning. While the Segment Anything Model 2 (SAM2) has shown remarkable success in video object segmentation by leveraging temporal cues, its direct application to 3D medical images faces two fundamental domain gaps: 1) the bidirectional an
Beyond Prefixes: Graph-as-Memory Cross-Attention for Knowledge Graph Completion with Large Language Models
cs.AIRuitong Liu, Boxu Lin, Peize Li, Siyuan Li
Fusing Knowledge Graphs with Large Language Models (LLMs) is crucial for knowledge-intensive tasks like knowledge graph completion. Existing LLM-based approaches typically inject graph information via prefix concatenation, resulting in shallow interactions that fail to support fine-grained evidence retrieval during generation. Beyond prefixes, we propose Gra
Junyu Xuan, Wenlong Chen, Yingzhen Li
Bayesian Optimisation (BO) is a powerful tool for optimising expensive blackbox functions but its effectiveness diminishes in highdimensional spaces due to sparse data and poor surrogate model scalability While Variational Autoencoder (VAE) based approaches address this by learning low-dimensional latent representations the reconstructionbased objective func
Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou
Recent advances in inference-time scaling, particularly those leveraging reinforcement learning with verifiable rewards, have substantially enhanced the reasoning capabilities of Large Vision-Language Models (LVLMs). Inspired by this success, similar strategies have been applied to multimodal reasoning, yet their impact on visual perception remains unclear.