October 2025 arXiv papers — page 165
Showing 16,401–16,500 of 25,213 papers
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
cs.AIYunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu
Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are largely confined to static, single-agent tasks. Inspired by Dixit, we introduce DixitWorld, a comprehensive evaluation suite designed to de
Xing Wei, Chunchun Chen, Rui Fan, Xiaofeng Cao
Graph neural networks (GNNs) can efficiently process text-attributed graphs (TAGs) due to their message-passing mechanisms, but their training heavily relies on the human-annotated labels. Moreover, the complex and diverse local topologies of nodes of real-world TAGs make it challenging for a single mechanism to handle. Large language models (LLMs) perform w
Kai Cao, Yucong Duan, Wensheng Gan
Incorporating utility into targeted pattern mining can address the practical limitations of traditional frequency-based approaches. However, utility-based methods often suffer from generating a large number of long and complicated sequences. To improve pattern relevance and interpretability, average utility provides a more balanced metric by considering both
Mixture of Inverse Gaussians for Hemodynamic Transport (MIGHT) in Multiple-Input Multiple-Output Vascular Networks
q-bio.QMTimo Jakumeit, Bastian Heinlein, Nunzio Tuccitto, Robert Schober
Synthetic molecular communication (MC) in the cardiovascular system is a key enabler for many envisioned medical applications inside the human body, such as targeted drug delivery, early disease detection, and continuous health monitoring. The design of synthetic MC systems for such applications requires suitable models for the signaling molecule propagation
Luyao Zhuang, Shengyuan Chen, Yilin Xiao, Huachi Zhou
Retrieval-Augmented Generation (RAG) is widely used to mitigate hallucinations of Large Language Models (LLMs) by leveraging external knowledge. While effective for simple queries, traditional RAG systems struggle with large-scale, unstructured corpora where information is fragmented. Recent advances incorporate knowledge graphs to capture relational structu
ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications
cs.CVYuxi Mi, Qiuyang Yuan, Zhizhou Zhong, Xuan Zhao
Recently, iris recognition is regaining prominence in immersive applications such as extended reality as a means of seamless user identification. This application scenario introduces unique challenges compared to traditional iris recognition under controlled setups, as the ocular images are primarily captured off-axis and less constrained, causing perspectiv
Thermal and Electrical Conductivities of Aluminum Up to 1000 eV: A First-Principles Prediction
physics.plasm-phQianrui Liu, Xiantu He, Mohan Chen
Accurate prediction of the thermal and electrical conductivities of materials under extremely high temperatures is essential in high-energy-density physics. These properties govern processes such as stellar core dynamics, planetary magnetic field generation, and laser-driven plasma evolution. However, first-principles methods like Kohn-Sham (KS) density func
Rui Chen, Bin Liu, Changtao Miao, Xinghao Wang
Advances in image tampering pose serious security threats, underscoring the need for effective image manipulation localization (IML). While supervised IML achieves strong performance, it depends on costly pixel-level annotations. Existing weakly supervised or training-free alternatives often underperform and lack interpretability. We propose the In-Context F
Hybrid Quantum Systems: Coupling Single-Molecule Magnet Qudits with Industrial Silicon Spin Qubits
cond-mat.mes-hallDaniel Schroller, Daniel Sitter, Thomas Koch, Viktor Adam
Molecular spin qudits offer an attractive platform for quantum memory, combining long coherence times with rich multi-level spin structures. Terbium bis(phthalocyaninato) (TbPc$_2$) exemplifies such systems, with demonstrated quantum control and chemical reproducibility. In hybrid quantum architectures, TbPc$_2$ can act as the primary memory element, with se
Integrating Structure-Aware Attention and Knowledge Graphs in Explainable Recommendation Systems
cs.IRShuangquan Lyu, Ming Wang, Huajun Zhang, Jiasen Zheng
This paper designs and implements an explainable recommendation model that integrates knowledge graphs with structure-aware attention mechanisms. The model is built on graph neural networks and incorporates a multi-hop neighbor aggregation strategy. By integrating the structural information of knowledge graphs and dynamically assigning importance to differen
Uncertainty-Aware Post-Detection Framework for Enhanced Fire and Smoke Detection in Compact Deep Learning Models
cs.CVAniruddha Srinivas Joshi, Godwyn James William, Shreyas Srinivas Joshi
Accurate fire and smoke detection is critical for safety and disaster response, yet existing vision-based methods face challenges in balancing efficiency and reliability. Compact deep learning models such as YOLOv5n and YOLOv8n are widely adopted for deployment on UAVs, CCTV systems, and IoT devices, but their reduced capacity often results in false positive
Yuxin Yang, Kun Liu, Xuefeng Zhang, Yi-Ming Hu
The measurement of Earth's free oscillations plays an important role in studying the Earth's large-scale structure. Space technology development presents a potential method to observe these normal modes by measuring inter-satellite distances. However, the disturbance from the Earth's low-degree gravity field makes it challenging for low Earth orbit gravity m
On the Propagation and Damping of Alfvenic Fluctuations in the Outer Solar Corona and Solar Wind
astro-ph.SRNikos Sioulas, Marco Velli, Chen Shi, Trevor A. Bowen
We analyze \textit{Parker Solar Probe} and \textit{Solar Orbiter} observations to investigate the propagation and dissipation of Alfv\'enic fluctuations from the outer corona to 1~AU. Conservation of wave-action flux provides the theoretical baseline for how fluctuation amplitudes scale with the Alfv\'en Mach number $M_a$, once solar-wind acceleration is acc
Lighter-X: An Efficient and Plug-and-play Strategy for Graph-based Recommendation through Decoupled Propagation
cs.LGYanping Zheng, Zhewei Wei, Frank de Hoog, Xu Chen
Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness in recommendation systems. However, conventional graph-based recommenders, such as LightGCN, require maintaining embeddings of size $d$ for each node, resulting in a parameter complexity of $\mathcal{O}(n \times d)$, where $n$ represents the total number of users and items. This scaling
Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
cs.CVMinbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li
Recent advances in large language models (LLMs) have demonstrated that reinforcement learning with verifiable rewards (RLVR) can significantly enhance reasoning abilities by directly optimizing correctness, rather than relying solely on supervised imitation. This paradigm has been extended to multimodal LLMs for complex video and image understanding tasks. H
Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen
Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN ($\underline{REF}$lective-$
Guilin Li, Yun Zhang, Xiuyuan Chen, Chengqi Li
Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world knowledge, they remain limited in modeling the behavioral knowledge contained within user interaction histories. User behavior forms a distinct modality, where each action, defined by
Kuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang
Unsupervised Federated Learning (UFL) aims to collaboratively train a global model across distributed clients without sharing data or accessing label information. Previous UFL works have predominantly focused on representation learning and clustering tasks. Recently, vision language models (e.g., CLIP) have gained significant attention for their powerful zer
Yuanche Liu, Yingxuan Xu, Yang Zhang
We introduce a machine-learning framework based on symbolic regression to extract the full symbol alphabet of multi-loop Feynman integrals. By targeting the analytic structure rather than reduction, the method is broadly applicable and interpretable across different families of integrals. It successfully reconstructs complete symbol alphabets in nontrivial e
Shreedhar Bhat, Achinta Kumar Nandi
The investigation of the dimension of Bergman spaces has long been a central topic in several complex variables, uncovering profound connections with potential theory and function theory since the pioneering work of Carleson, Wiegerinck, and others in the 1960s. We investigate the dimension of $p$-Bergman spaces associated with pseudoconvex domains in $\math
Jiahui Lu, Haihong Xiao, Xueyan Zhao, Wenxiong Kang
Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have advanced 3D reconstruction and novel view synthesis, but remain heavily dependent on accurate camera poses and dense viewpoint coverage. These requirements limit their applicability in sparse-view settings, where pose estimation becomes unreliable and supervision is insufficient. To overcome
Jacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud
OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep CoTs monitorable by training only against output monitors that do not have access to CoT. We show that such training can still cause obfuscated CoTs via two mechanisms. First, when
Weak solutions and weak-strong uniqueness for a compressible power-law-Oldroyd--B fluid model
math.APYong Lu, Milan Pokorny
We consider a model of a viscoelastic compressible flow in $R^{3}$ which is additionally shear thickening (the stress tensor corresponds to the power law model, however, the divergence of the velocity is due to the model bounded). We prove existence of a weak solution to this model provided the growth in the power law model is larger or equal than $\frac 52$
CardRewriter: Leveraging Knowledge Cards for Long-Tail Query Rewriting on Short-Video Platforms
cs.IRPeiyuan Gong, Feiran Zhu, Yaqi Yin, Chenglei Dai
Short-video platforms have rapidly become a new generation of information retrieval systems, where users formulate queries to access desired videos. However, user queries, especially long-tail ones, often suffer from spelling errors, incomplete phrasing, and ambiguous intent, resulting in mismatches between user expectations and retrieved results. While larg
Sumukha Sathyanarayana, N. Guru Sharan
In this paper, we establish Kronecker limit type formulas for the generalized Mordell--Tornheim zeta function $\Theta(r,s,t,x)$ as a function of the third variable, in terms of Riemann-zeta and Gamma values. We also give series evaluations of $\Theta(r,s,t,x)$ in terms of Herglotz-Zagier type functions, and their derivatives. As applications of this, we deri
Insights into Planet Formation from the Ages, Masses, and Elemental Abundances of Host Stars
astro-ph.SRXunzhou Chen, Tiancheng Sun, Lifei Ye
How planetary systems form and evolve is a key question in astronomy. Revealing how host star properties, such as elemental abundances, age, and mass, differ from those of non-host stars, and how they correlate with planetary characteristics such as radius, provides new insights into the formation and evolutionary pathways of planetary systems. We determine
Pan Zeming, Tan Naiming, Gao Chao, Yao Zhihai
The quantum Cheshire cat effect is an important phenomenon in quantum mechanics that reveals the separability of physical properties from their carriers. This effect transcends the classical framework whose attributes must be inherently attached to objects, providing new perspectives for quantum information and precision measurement. According to the quantum
Slim Ibrahim, Quyuan Lin, Lingjun Qian, Edriss S. Titi
The primitive equations (PEs) model planetary large-scale oceanic and atmospheric dynamics. While it has been shown that there are smooth solutions to the inviscid PEs (also called the hydrostatic Euler equations) with constant temperature (isothermal) that develop stable singularities in finite time, the effect of non-constant temperature on the singularity
Zixuan Gong, Yong Liu, Jiaye Teng
While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remains underexplored. In this paper, we explain this phenomenon through the lens of loss landscape geometry, inspired by empirical observations of their distinct dynamics at both sample
Matchmaker: An Open-source Library for Real-time Piano Score Following and Systematic Evaluation
cs.SDJiyun Park, Carlos Cancino-Chacón, Suhit Chiruthapudi, Juhan Nam
Real-time music alignment, also known as score following, is a fundamental MIR task with a long history and is essential for many interactive applications. Despite its importance, there has not been a unified open framework for comparing models, largely due to the inherent complexity of real-time processing and the language- or system-dependent implementatio
Beyond ADE and FDE: A Comprehensive Evaluation Framework for Safety-Critical Prediction in Multi-Agent Autonomous Driving Scenarios
cs.ROFeifei Liu, Haozhe Wang, Zejun Wei, Qirong Lu
Current evaluation methods for autonomous driving prediction models rely heavily on simplistic metrics such as Average Displacement Error (ADE) and Final Displacement Error (FDE). While these metrics offer basic performance assessments, they fail to capture the nuanced behavior of prediction modules under complex, interactive, and safety-critical driving sce
Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
cs.CRGuozhi Liu, Qi Mu, Tiansheng Huang, Xinhua Wang
Harmful fine-tuning issues present significant safety challenges for fine-tuning-as-a-service in large language models. Existing alignment-stage defenses, e.g., Vaccine, Repnoise, Booster, and T-Vaccine, mitigate harmful fine-tuning issues by enhancing the model's robustness during the alignment phase. While these methods have been proposed to mitigate the i
Tracking the Spatiotemporal Evolution of Landslide Scars Using a Vision Foundation Model: A Novel and Universal Framework
cs.CVMeijun Zhou, Gang Mei, Zhengjing Ma, Nengxiong Xu
Tracking the spatiotemporal evolution of large-scale landslide scars is critical for understanding the evolution mechanisms and failure precursors, enabling effective early-warning. However, most existing studies have focused on single-phase or pre- and post-failure dual-phase landslide identification. Although these approaches delineate post-failure landsli
Debabrata Mandal, Zhihan Peng, Yujie Wang, Praneeth Chakravarthula
We tackle the challenge of robust, in-the-wild imaging using ultra-thin nanophotonic metalens cameras. Meta-lenses, composed of planar arrays of nanoscale scatterers, promise dramatic reductions in size and weight compared to conventional refractive optics. However, severe chromatic aberration, pronounced light scattering, narrow spectral bandwidth, and low
Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects
cs.HCRaziyeh Zall, Alireza Kheyrkhah, Erik Cambria, Zahra Naseri
The development of agents with emotional intelligence is becoming increasingly vital due to their significant role in human-computer interaction and the growing integration of computer systems across various sectors of society. Affective computing aims to design intelligent systems that can recognize, evoke, and express human emotions, thereby emulating huma
Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
cs.CLParthiv Chatterjee, Shivam Sonawane, Amey Hengle, Aditya Tanna
Document summarization enables efficient extraction of user-relevant content but is inherently shaped by individual subjectivity, making it challenging to identify subjective salient information in multifaceted documents. This complexity underscores the necessity for personalized summarization. However, training models for personalized summarization has so f
Youshuai Tan, Zhanwei Zhang, Zishuo Ding, Lianyu Zheng
Floating-point program errors can lead to severe consequences, particularly in critical domains such as military applications. Only a small subset of inputs may induce substantial floating-point errors, prompting researchers to develop methods for identifying these error-inducing inputs. Although existing approaches have achieved some success, they still suf
Donghan Kim
We construct a multiset space $\mathbb{N}[X]$ over a metric space $X$ that simultaneously enjoys desirable topological properties and admits a natural matching metric $d_{\mathbb{N}[X]}$, making it a metrizable abelian topological monoid whose structure is compatible with the original metric on $X$. This framework extends naturally to the free abelian group
Angel Hsing-Chi Hwang, Fiona Li, Jacy Reese Anthis, Hayoun Noh
The quickly growing popularity of AI companions poses risks to mental health, personal wellbeing, and social relationships. Past work has identified many individual factors that can drive human-companion interaction, but we know little about how these factors interact and evolve over time. In Study 1, we surveyed AI companion users (N = 303) to map the psych
Chung-Soo Ahn, Rajib Rana, Sunil Sivadas, Carlos Busso
Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While generative data augmentation offers a promising solution, existing approaches often produce emotionally inconsistent sampl
Wenqing Wang, Muhammad Asif Ali, Ali Shoker, Ruohan Yang
Human preferences are diverse and dynamic, shaped by regional, cultural, and social factors. Existing alignment methods like Direct Preference Optimization (DPO) and its variants often default to majority views, overlooking minority opinions and failing to capture latent user intentions in prompts. To address these limitations, we introduce \underline{\textb
Interplay of choice and topology in percolation on mediation-driven attachment networks
cond-mat.stat-mechNilomber Roy, M. M. B. Sheraj, M. K. Hassan
We investigate bond percolation on mediation-driven attachment (MDA) networks under the generalized Achlioptas process, where $M>1$ candidate bonds are sampled and the one that minimizes the resulting cluster size is selected the best-of-$M$ rule. This framework offers a systematic approach to investigate how network topology and choice mechanisms jointly sh
Salomon Ibarra, Frida Cantu, Kaixiong Zhou, Li Zhang
Deep learning models have attracted lots of research attention in time series classification (TSC) task in the past two decades. Recently, deep neural networks (DNN) have surpassed classical distance-based methods and achieved state-of-the-art performance. Despite their promising performance, deep neural networks (DNNs) have been shown to rely on spurious co
Jiayi Mao, Liqun Li, Yanjie Gao, Zegang Peng
Effective incident management in large-scale IT systems relies on troubleshooting guides (TSGs), but their manual execution is slow and error-prone. While recent advances in LLMs offer promise for automating incident management tasks, existing LLM-based solutions lack specialized support for several key challenges, including managing TSG quality issues, inte
Zonghao Ying, Yangguang Shao, Jianle Gan, Gan Xu
Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed in real-world environments, they face serious security risks, motivating the design of security evaluation benchmarks. Existing benchmarks provide only partial coverage, typically restricted to narrow scenarios such a
Sanchar Palit, Subhasis Chaudhuri, Biplab Banerjee
Diffusion models have recently achieved significant success in various image manipulation tasks, including image super-resolution and perceptual quality enhancement. Pretrained text-to-image models, such as Stable Diffusion, have exhibited strong capabilities in synthesizing realistic image content, which makes them particularly attractive for addressing sup
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference
cs.CLHua Cai, Shuang Zhao, Liang Zhang, Xuli Shen
Reasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remains underexplored. In this paper, we introduce Unilaw-R1, a large language model tailored for legal reasoning. With a lightweight 7-billion parameter scale, Unilaw-R1 significantly reduces deployment cost w
Jinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao
Conventional continual pretraining (CPT) for large language model (LLM) domain adaptation often suffers from catastrophic forgetting and limited domain capacity. Existing strategies adopt layer expansion, introducing additional trainable parameters to accommodate new knowledge. However, the uniform expansion and updates still entangle general and domain lear
An efficient spectral Poisson solver for the nirvana-III code: the shearing-box case with vertical vacuum boundary conditions
astro-ph.IMS. Rendon Restrepo, O. Gressel
The stability of a differentially rotating fluid subject to its own gravity is a problem with applications across wide areas of astrophysics--from protoplanetary discs (PPDs) to entire galaxies. The shearing box formalism offers a conceptually simple framework for studying differential rotation in the local approximation. Aimed at self-gravitating, and impor
Zeyu Ling, Xiaodong Gu, Jiangnan Tang, Changqing Zou
We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with cross-modal contrastive alignment and employs three per-frame prompt tokens that explicitly encode the essential factor
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
cs.CVPîrvu Mihai-Cristian, Marius Leordeanu
The computer vision domain has greatly benefited from an abundance of data across many modalities to improve on various visual tasks. Recently, there has been a lot of focus on self-supervised pre-training methods through Masked Autoencoders (MAE) \cite{he2022masked,bachmann2022multimae}, usually used as a first step before optimizing for a downstream task,
Shan Jiang, Chenguang Zhu, Sarfraz Khurshid
JavaScript obfuscators are widely deployed to protect intellectual property and resist reverse engineering, yet their correctness has been largely overlooked compared to performance and resilience. Existing evaluations typically measure resistance to deobfuscation, leaving the critical question of whether obfuscators preserve program semantics unanswered. In
Broad nonlocal spectrum in the Pb-InSb hybrid three terminals for potential realization of Kitaev chains
quant-phGuoan Li, Xiaofan Shi, Ruixuan Zhang, Yuxiao Song
Hybrid superconductor-semiconductor(SC-SM) nanowires remain one of the foremost platforms for engineering topological superconductivity and Majorana zero modes(MZMs) towards fault-tolerant topological qubits, especially with the rapid development of artificial Kitaev chains. In contrast to the widely used aluminum(Al)-based hybrids, lead(Pb) offers a bulk su
Yibo Yang
Deep learning has advanced NLP, but interpretability remains limited, especially in healthcare and finance. Concept bottleneck models tie predictions to human concepts in vision, but NLP versions either use binary activations that harm text representations or latent concepts that weaken semantics, and they rarely model dynamic concept interactions such as ne
Adnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff
Comparing human and model performance offers a valuable perspective for understanding the strengths and limitations of embedding models, highlighting where they succeed and where they fail to capture meaning and nuance. However, such comparisons are rarely made, as human performance on embedding tasks is difficult to measure. To fill this gap, we introduce H
Hehe Fan, Yi Yang, Mohan Kankanhalli, Fei Wu
When modeling a given type of data, we consider it to involve two key aspects: 1) identifying relevant elements (e.g., image pixels or textual words) to a central element, as in a convolutional receptive field, or to a query element, as in self-attention, and 2) encoding these tokens effectively. Self-attention can adaptively identify these elements but reli
Ionospheric and Plasmaspheric Delay Characterization for Lunar Terrestrial GNSS Receivers with Global Core Plasma Model
cs.ROKeidai Iiyama, Grace Gao
Recent advancements in lunar positioning, navigation, and timing (PNT) have demonstrated that terrestrial GNSS signals, including weak sidelobe transmissions, can be exploited for lunar spacecraft positioning and timing. While GNSS-based navigation at the Moon has been validated recently, unmodeled ionospheric and plasmaspheric delays remain a significant er
Leonardo Rossetti, Stefano Mancini, Andreas Winter, Joseph Schindler
The use of coarse graining to connect physical and information theoretic entropies has recently been given a precise formulation in terms of ``observational entropy'', describing entropy for observers with respect to a measurement. Here we consider observers with various locality restrictions, including local measurements (LO), measurements based on local op
Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery
cs.CVZhi-Fen He, Ren-Dong Xie, Bo Li, Bin Liu
Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features and recovering missing labels. In this paper, we propose a Co-learning framework for Semantic-aware features and Label recovery (CSL), designed to address both challenges in a unifie
If $\sum_n n! c_n z^n$ is entire and $c_n$ does not terminate, then $\sum_n c_n z^n$ has infinitely many zeros
math.CVAlann Rosas
We prove that if $\sum_n n! c_n z^n$ is entire and $c_n$ does not terminate, then $\sum_n c_n z^n$ has infinitely many zeros. We then use this result to give alternative proofs that the Le Roy functions $f_r(z)=\sum_{n=0}^\infty \frac{z^n}{(n!)^r}$ for $r>1$ and Bessel functions $J_\alpha(z)=\sum_{m=0}^\infty \frac{(-1)^m}{m!\Gamma(m+\alpha+1)}\left(\frac{z}
Bo Peng, Zichuan Wang, Sheng Yu, Xiaochuan Jin
Deep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its comp
Kaitao Chen, Shaohao Rui, Yankai Jiang, Jiamin Wu
Medical vision-language models (VLMs) excel at image-text understanding but typically rely on a single-pass reasoning that neglects localized visual cues. In clinical practice, however, human experts iteratively scan, focus, and refine the regions of interest before reaching a final diagnosis. To narrow this machine-human perception gap, we introduce ViTAR,
Sitong Gong, Yunzhi Zhuge, Lu Zhang, Pingping Zhang
Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the segmentation accuracy and robustness. Traditional CNN approaches manage audio-visual interactions through basic operations lik
Takeshi Takahashi, Hiroki Saito
Bose-Einstein condensates of $^{87}\mathrm{Rb}$ atoms with a hyperfine spin of 2 are open quantum systems, where the atoms are lost through two-body inelastic collisions. In this dissipation process, a collision channel with total spin of 4 is forbidden by angular momentum conservation, which results in magnetization of the atoms remaining in the condensate.
Jiawen Li, Zheng Ning, Yuan Tian, Toby Jia-jun Li
Large language models (LLMs) enable end-users to delegate complex tasks to autonomous agents through natural language. However, prompt-based interaction faces critical limitations: Users often struggle to specify procedural requirements for tasks, especially those that don't have a factually correct solution but instead rely on personal preferences, such as
Between Knowledge and Care: Evaluating Generative AI-Based IUI in Type 2 Diabetes Management Through Patient and Physician Perspectives
cs.HCYibo Meng, Ruiqi Chen, Bingyi Liu, Yan Guan
Generative AI systems are increasingly used by patients seeking everyday health guidance, yet their appropriateness in chronic care contexts remains unclear. Focusing on Type 2 Diabetes Mellitus (T2DM), this paper presents a mixed-methods investigation into how AI-generated health information is interpreted by patients and evaluated by physicians in China. D
Ruohao Li, Hongjun Liu, Leyi Zhao, Zisu Li
Large language model (LLM) agents have shown remarkable reasoning abilities. However, existing multi-agent frameworks often rely on fixed roles or centralized control, limiting scalability and adaptability in long-horizon reasoning. We introduce SwarmSys, a closed-loop framework for distributed multi-agent reasoning inspired by swarm intelligence. Coordinati
LOMORO: Long-term Monitoring of Dynamic Targets with Minimum Robotic Fleet under Resource Constraints
cs.ROMingke Lu, Shuaikang Wang, Meng Guo
Long-term monitoring of numerous dynamic targets can be tedious for a human operator and infeasible for a single robot, e.g., to monitor wild flocks, detect intruders, search and rescue. Fleets of autonomous robots can be effective by acting collaboratively and concurrently. However, the online coordination is challenging due to the unknown behaviors of the
Qiaoyan Peng, Qingqing Wu, Guangji Chen, Wen Chen
In this paper, we investigate an intelligent reflecting surface (IRS) aided wireless communication system, where active IRSs (AIRSs) are deployed to assist communication between a base station (BS) and users of both the uplink (UL) and downlink (DL). We aim to maximize the weighted sum rate (WSR) of UL and DL communications through joint optimization of BS,
Rahul Vanukuri, Shafi Ullah Khan, Talip Tolga Sarı, Gokhan Secinti
The growing demand for effective spectrum management and interference mitigation in shared bands, such as the Citizens Broadband Radio Service (CBRS), requires robust radar detection algorithms to protect the military transmission from interference due to commercial wireless transmission. These algorithms, in turn, depend on large, diverse, and carefully lab
Tomasz Krajewski, Włodek Kluźniak
Nearly every galactic core contains a supermassive compact object, hypothesized to be a Kerr black hole. It was only with the advent of Event Horizon Telescope observations that the predictions of this hypothesis could be observationally tested for our own Galaxy, and the nearby elliptical M87, on spatial scales comparable to the gravitational radius. At the
Saleh Nikooroo, Thomas Engel
Belief systems are rarely globally consistent, yet effective reasoning often persists locally. We propose a novel graph-theoretic framework that cleanly separates credibility--external, a priori trust in sources--from confidence--an internal, emergent valuation induced by network structure. Beliefs are nodes in a directed, signed, weighted graph whose edges
Sahng-Min Han, Minjae Kim, Jinho Cha, Se-woon Choe
Deep learning in small and imbalanced biomedical datasets remains fundamentally constrained by unstable optimization and poor generalization. We present the first biomedical implementation of FOSSIL (Flexible Optimization via Sample-Sensitive Importance Learning), a regret-minimizing weighting framework that adaptively balances training emphasis according to
Shafi Ullah Khan, Michel Kulhandjian, Debashri Roy
Spectrum sharing is a critical strategy for meeting escalating user demands via commercial wireless services, yet its effective regulation and technological enablement, particularly concerning coexistence with incumbent systems, remain significant challenges. Federal organizations have established regulatory frameworks to manage shared commercial use alongsi
Enze Sun, Zhihao Gavin Tang, Yifan Wang
In online combinatorial allocation, agents arrive sequentially and items are allocated in an online manner. The algorithm designer only knows the distribution of each agent's valuation, while the actual realization of the valuation is revealed only upon her arrival. Against the offline benchmark, Feldman, Gravin, and Lucier (SODA 2015) designed an optimal $0
Oleksiy Dovgoshey, Olga Rovenska
Let $T$ be a tree of arbitrary finite or infinite order and let $U(T)$ be the set of all ultrametric spaces generated by vertex labelings of $T$. Let ${\bf US}$ denote the class of all ultrametric spaces generated by vertex labelings of star graphs. We prove that the inclusion $U(T)\subseteq {\bf US}$ holds if and only if the longest path in $T$ has a length
Automated Glaucoma Report Generation via Dual-Attention Semantic Parallel-LSTM and Multimodal Clinical Data Integration
cs.CECheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee
Generative AI for automated glaucoma diagnostic report generation faces two predominant challenges: content redundancy in narrative outputs and inadequate highlighting of pathologically significant features including optic disc cupping, retinal nerve fiber layer defects, and visual field abnormalities. These limitations primarily stem from current multimodal
Md. Nayeem, Md Shamse Tabrej, Kabbojit Jit Deb, Shaonti Goswami
Automatic Speech Recognition (ASR) has undergone a profound transformation over the past decade, driven by advances in deep learning. This survey provides a comprehensive overview of the modern era of ASR, charting its evolution from traditional hybrid systems, such as Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) and Deep Neural Network-HMMs (DNN-H
Jusheng Zhang, Kaitong Cai, Qinglin Zeng, Ningyuan Liu
Optimizing LLM-based workflows is typically formulated as a global search, where candidate workflows are evaluated based on a scalar metric. This paradigm, however, suffers from a critical flaw: information collapse. By reducing rich, multi-step execution traces to simple success/failure signals, existing methods are rendered blind to the underlying structur
Sebastian Gant, Ben Williams
Working over an algebraically closed field $k$ of characteristic $0$, we show that the motivic stable homotopy groups of the sphere spectrum can be determined entirely from the motivic homotopy groups of the $p$-completed sphere spectra and the motivic cohomology of the ground field, except possibly for the $0$ and $-1$-stems. Using this, we show that the co
Yiming Ren, Kwan Chuen Chan, Le Zhang, Yin Li
Accurate photometric redshift (photo-$z$) estimation is a key challenge in cosmology, as uncertainties in photo-$z$ directly limit the scientific return of large-scale structure and weak lensing studies, especially in upcoming Stage IV surveys. The problem is particularly severe for faint galaxies with sparse spectroscopic training data. In this work, we int
Huimin Ji, Zhihua Li, Wenzheng Cheng, Zheng Li
In the construction of High-Luminosity Large Hadron Collider (HL-LHC) and Future Circular Collider (FCC) experiments, 3D pixel sensors have become indispensable components due to their superior radiation hardness, fast response, and low power consumption. However, there are still significant challenges in the process of 3D sensors manufacturing. In this work
Henan Wang, Hanxin Zhu, Xinliang Gong, Tianyu He
3D Gaussian Splatting (3DGS) has garnered significant attention due to its superior scene representation fidelity and real-time rendering performance, especially for dynamic 3D scene reconstruction (\textit{i.e.}, 4D reconstruction). However, despite achieving promising results, most existing algorithms overlook the substantial temporal and spatial redundanc
Ruoxing Yang
We introduce PPOPT - Proximal Policy Optimization using Pretraining, a novel, model-free deep-reinforcement-learning algorithm that leverages pretraining to achieve high training efficiency and stability on very small training samples in physics-based environments. Reinforcement learning agents typically rely on large samples of environment interactions to l
Masakiyo Kitazawa, ShinIchi Esumi, Takafumi Niida, Toshihiro Nonaka
We derive analytic formulas to reconstruct particle-averaged quantities from experimental results that suffer from the efficiency loss of particle measurements. These formulas are derived under the assumption that the probabilities of observing individual particles are independent. The formulas do not agree with the conventionally used intuitive formulas.
Kazuki Sato
We study whether the norm one torus associated with a finite separable non-Galois field extension $K/k$ is $p$-retract rational over $k$ for a prime $p$, focusing on the case where the Galois group of the Galois closure of $K/k$ is either the symmetric or the alternating group.
Qi Liu, Yanbo Chen
This paper investigates McKean-Vlasov backward stochastic variational inequalities (BSVIs) whose generator depends on the joint law of the solution. We first establish the existence and uniqueness of the solution under globally Lipschitz and linear growth conditions. The analysis is then extended to the more general case of locally Lipschitz and non-linear g
Lightweight Baselines for Medical Abstract Classification: DistilBERT with Cross-Entropy as a Strong Default
cs.CLJiaqi Liu, Tong Wang, Su Liu, Xin Hu
The research evaluates lightweight medical abstract classification methods to establish their maximum performance capabilities under financial budget restrictions. On the public medical abstracts corpus, we finetune BERT base and Distil BERT with three objectives cross entropy (CE), class weighted CE, and focal loss under identical tokenization, sequence len
The eigentheory for nonlocal cooperative-advective system and its role in the study of free boundary system for directional epidemic models
math.APSoufiane Bentout, Hoang-Hung Vo
In this paper, we propose and analyze a nonlocal cooperative reaction--diffusion system with free boundaries and drift terms, motivated by directional epidemic spread. Lacking a variational structure but requiring sharper regularity of solutions, the model poses substantial analytical challenges compared with previous works~\cite{Du,Berestycki2016a,Berestyck
Yinghui He, Abhishek Panigrahi, Yong Lin, Sanjeev Arora
Language models often show little to no improvement (i.e., "saturation") when trained via vanilla supervised fine-tuning (SFT) on data similar to what they saw in their training set (e.g., MATH). We introduce a new fine-tuning strategy, STAT, to train such a student model by using the metacognition ability of a stronger large language model (LLM) as the teac
Junan Chen, Trung Thanh Nguyen, Takahiro Komamizu, Ichiro Ide
Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard "pre-training followed by fine-tuning" paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) appro
Euclid preparation. Cosmology Likelihood for Observables in Euclid (CLOE). 6: Impact of systematic uncertainties on the cosmological analysis
astro-ph.COEuclid Collaboration, L. Blot, K. Tanidis, G. Cañas-Herrera
Extracting cosmological information from the Euclid galaxy survey will require modelling numerous systematic effects during the inference process. This implies varying a large number of nuisance parameters, which have to be marginalised over before reporting the constraints on the cosmological parameters. This is a delicate process, especially with such a la
"Can I Decorate My Teeth With Diamonds?": Exploring Multi-Stakeholder Perspectives on Using VR to Reduce Children's Dental Anxiety
cs.HCYaxuan Mao, Yanheng Li, Duo Gong, Pengcheng An
Dental anxiety is prevalent among children, often leading to missed treatment and potential negative effects on their mental well-being. While several interventions (e.g., pharmacological and psychotherapeutic techniques) have been introduced for anxiety alleviation, the recently emerged virtual reality (VR) technology, with its immersive and playful nature,
Comprehensive X-ray Spectral-timing Analysis of GRS 1915+105 Based on Insight-HXMT Observations
astro-ph.HEXiao Chen, Weiping Liu, Wei Wang
GRS 1915+105 has been well studied since its discovery, and is well-known for its complex light curve variability. Using the full currently available Insight-HXMT dataset from July 2017 to June 2023, we make a comprehensive spectral-timing analysis of this source and report four main findings. First, we uncover a QPO frequency rising branch between MJD 58206
Thao Pham
As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has examined how AI systems scheme against human developers, LLM-to-LLM scheming remains underexplored. We investigate the scheming ability and propensity of frontier LLM agents through t
Hybrid Robotic Meta-gripper for Tomato Harvesting: Analysis of Auxetic Structures with Lattice Orientation Variations
cs.ROShahid Ansari, Vivek Gupta, Bishakh Bhattacharya
The agricultural sector is rapidly evolving to meet growing global food demands, yet tasks like fruit and vegetable handling remain labor-intensive, causing inefficiencies and post-harvest losses. Automation, particularly selective harvesting, offers a viable solution, with soft robotics emerging as a key enabler. This study introduces a novel hybrid gripper
End-to-end Compositional Verification of Program Safety through Verified and Verifying Compilation
cs.PLJinhua Wu, Yuting Wang, Liukun Yu, Linglong Meng
Program safety (i.e., absence of undefined behaviors) is critical for correct operation of computer systems. It is usually verified at the source level (e.g., by separation logics) and preserved to the target by verified compilers (e.g., CompCert), thereby achieving end-to-end verification of safety. However, modern safe programming languages like Rust pose
R. M. Albuquerque, S. Narison, D. Rabetiarivony
We review our estimations on the light scalar $\bar{q}q$, $(\bar{q}q')(\bar{q'}q)$ and $\overline{qq'}qq'$ ($q,q'\equiv u,d,s$) states from relativistic Laplace sum rule (LSR) within stability criteria and including higher order perturbative (PT) corrections up to the (estimated) N5LO. We evaluate the QCD spectral functions at Lowest Order (LO) of PT QCD and
Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao
As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models c
Multiscale Magnetic Correlations in La2Mn2-xNixO6: Role of Crystal Structure in Double Perovskites
cond-mat.mtrl-sciA. K. Bera, K. S. Chikara, B. Saha, S. M. Yusuf
The magnetic correlations in double perovskites La2Mn2-xNixO6 (x = 0.5, 0.75, 1.0, 1.25 and 1.5) have been systematically investigated across macroscopic, mesoscopic, and microscopic length scales using temperature-dependent bulk DC magnetization, neutron depolarization, and neutron powder diffraction measurements, respectivitly. The magnetic properties evol