Skip to content

October 2025 arXiv papers — page 165

Showing 16,40116,500 of 25,213 papers

  1. Yunxiang Mo, Tianshi Zheng, Qing Zong, Jiayu Liu

    Multimodal abductive reasoning--the generation and selection of explanatory hypotheses from partial observations--is a cornerstone of intelligence. Current evaluations of this ability in vision-language models (VLMs) are largely confined to static, single-agent tasks. Inspired by Dixit, we introduce DixitWorld, a comprehensive evaluation suite designed to de

  2. Xing Wei, Chunchun Chen, Rui Fan, Xiaofeng Cao

    Graph neural networks (GNNs) can efficiently process text-attributed graphs (TAGs) due to their message-passing mechanisms, but their training heavily relies on the human-annotated labels. Moreover, the complex and diverse local topologies of nodes of real-world TAGs make it challenging for a single mechanism to handle. Large language models (LLMs) perform w

  3. Kai Cao, Yucong Duan, Wensheng Gan

    Incorporating utility into targeted pattern mining can address the practical limitations of traditional frequency-based approaches. However, utility-based methods often suffer from generating a large number of long and complicated sequences. To improve pattern relevance and interpretability, average utility provides a more balanced metric by considering both

  4. Timo Jakumeit, Bastian Heinlein, Nunzio Tuccitto, Robert Schober

    Synthetic molecular communication (MC) in the cardiovascular system is a key enabler for many envisioned medical applications inside the human body, such as targeted drug delivery, early disease detection, and continuous health monitoring. The design of synthetic MC systems for such applications requires suitable models for the signaling molecule propagation

  5. Luyao Zhuang, Shengyuan Chen, Yilin Xiao, Huachi Zhou

    Retrieval-Augmented Generation (RAG) is widely used to mitigate hallucinations of Large Language Models (LLMs) by leveraging external knowledge. While effective for simple queries, traditional RAG systems struggle with large-scale, unstructured corpora where information is fragmented. Recent advances incorporate knowledge graphs to capture relational structu

  6. Yuxi Mi, Qiuyang Yuan, Zhizhou Zhong, Xuan Zhao

    Recently, iris recognition is regaining prominence in immersive applications such as extended reality as a means of seamless user identification. This application scenario introduces unique challenges compared to traditional iris recognition under controlled setups, as the ocular images are primarily captured off-axis and less constrained, causing perspectiv

  7. Qianrui Liu, Xiantu He, Mohan Chen

    Accurate prediction of the thermal and electrical conductivities of materials under extremely high temperatures is essential in high-energy-density physics. These properties govern processes such as stellar core dynamics, planetary magnetic field generation, and laser-driven plasma evolution. However, first-principles methods like Kohn-Sham (KS) density func

  8. Rui Chen, Bin Liu, Changtao Miao, Xinghao Wang

    Advances in image tampering pose serious security threats, underscoring the need for effective image manipulation localization (IML). While supervised IML achieves strong performance, it depends on costly pixel-level annotations. Existing weakly supervised or training-free alternatives often underperform and lack interpretability. We propose the In-Context F

  9. Daniel Schroller, Daniel Sitter, Thomas Koch, Viktor Adam

    Molecular spin qudits offer an attractive platform for quantum memory, combining long coherence times with rich multi-level spin structures. Terbium bis(phthalocyaninato) (TbPc$_2$) exemplifies such systems, with demonstrated quantum control and chemical reproducibility. In hybrid quantum architectures, TbPc$_2$ can act as the primary memory element, with se

  10. Shuangquan Lyu, Ming Wang, Huajun Zhang, Jiasen Zheng

    This paper designs and implements an explainable recommendation model that integrates knowledge graphs with structure-aware attention mechanisms. The model is built on graph neural networks and incorporates a multi-hop neighbor aggregation strategy. By integrating the structural information of knowledge graphs and dynamically assigning importance to differen

  11. Aniruddha Srinivas Joshi, Godwyn James William, Shreyas Srinivas Joshi

    Accurate fire and smoke detection is critical for safety and disaster response, yet existing vision-based methods face challenges in balancing efficiency and reliability. Compact deep learning models such as YOLOv5n and YOLOv8n are widely adopted for deployment on UAVs, CCTV systems, and IoT devices, but their reduced capacity often results in false positive

  12. Yuxin Yang, Kun Liu, Xuefeng Zhang, Yi-Ming Hu

    The measurement of Earth's free oscillations plays an important role in studying the Earth's large-scale structure. Space technology development presents a potential method to observe these normal modes by measuring inter-satellite distances. However, the disturbance from the Earth's low-degree gravity field makes it challenging for low Earth orbit gravity m

  13. Nikos Sioulas, Marco Velli, Chen Shi, Trevor A. Bowen

    We analyze \textit{Parker Solar Probe} and \textit{Solar Orbiter} observations to investigate the propagation and dissipation of Alfv\'enic fluctuations from the outer corona to 1~AU. Conservation of wave-action flux provides the theoretical baseline for how fluctuation amplitudes scale with the Alfv\'en Mach number $M_a$, once solar-wind acceleration is acc

  14. Yanping Zheng, Zhewei Wei, Frank de Hoog, Xu Chen

    Graph Neural Networks (GNNs) have demonstrated remarkable effectiveness in recommendation systems. However, conventional graph-based recommenders, such as LightGCN, require maintaining embeddings of size $d$ for each node, resulting in a parameter complexity of $\mathcal{O}(n \times d)$, where $n$ represents the total number of users and items. This scaling

  15. Minbin Huang, Runhui Huang, Chuanyang Zheng, Jingyao Li

    Recent advances in large language models (LLMs) have demonstrated that reinforcement learning with verifiable rewards (RLVR) can significantly enhance reasoning abilities by directly optimizing correctness, rather than relying solely on supervised imitation. This paradigm has been extended to multimodal LLMs for complex video and image understanding tasks. H

  16. Renliang Sun, Wei Cheng, Dawei Li, Haifeng Chen

    Chain-of-Thought (CoT) reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. However, excessive or redundant reasoning -- so-called overthinking -- can increase inference costs and lead LLMs toward incorrect conclusions. In this paper, we present REFRAIN ($\underline{REF}$lective-$

  17. Guilin Li, Yun Zhang, Xiuyuan Chen, Chengqi Li

    Large language models (LLMs) have shown that generative pretraining can distill vast world knowledge into compact token representations. While LLMs encapsulate extensive world knowledge, they remain limited in modeling the behavioral knowledge contained within user interaction histories. User behavior forms a distinct modality, where each action, defined by

  18. Kuangpu Guo, Lijun Sheng, Yongcan Yu, Jian Liang

    Unsupervised Federated Learning (UFL) aims to collaboratively train a global model across distributed clients without sharing data or accessing label information. Previous UFL works have predominantly focused on representation learning and clustering tasks. Recently, vision language models (e.g., CLIP) have gained significant attention for their powerful zer

  19. Yuanche Liu, Yingxuan Xu, Yang Zhang

    We introduce a machine-learning framework based on symbolic regression to extract the full symbol alphabet of multi-loop Feynman integrals. By targeting the analytic structure rather than reduction, the method is broadly applicable and interpretable across different families of integrals. It successfully reconstructs complete symbol alphabets in nontrivial e

  20. Shreedhar Bhat, Achinta Kumar Nandi

    The investigation of the dimension of Bergman spaces has long been a central topic in several complex variables, uncovering profound connections with potential theory and function theory since the pioneering work of Carleson, Wiegerinck, and others in the 1960s. We investigate the dimension of $p$-Bergman spaces associated with pseudoconvex domains in $\math

  21. Jiahui Lu, Haihong Xiao, Xueyan Zhao, Wenxiong Kang

    Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have advanced 3D reconstruction and novel view synthesis, but remain heavily dependent on accurate camera poses and dense viewpoint coverage. These requirements limit their applicability in sparse-view settings, where pose estimation becomes unreliable and supervision is insufficient. To overcome

  22. Jacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud

    OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep CoTs monitorable by training only against output monitors that do not have access to CoT. We show that such training can still cause obfuscated CoTs via two mechanisms. First, when

  23. Yong Lu, Milan Pokorny

    We consider a model of a viscoelastic compressible flow in $R^{3}$ which is additionally shear thickening (the stress tensor corresponds to the power law model, however, the divergence of the velocity is due to the model bounded). We prove existence of a weak solution to this model provided the growth in the power law model is larger or equal than $\frac 52$

  24. Peiyuan Gong, Feiran Zhu, Yaqi Yin, Chenglei Dai

    Short-video platforms have rapidly become a new generation of information retrieval systems, where users formulate queries to access desired videos. However, user queries, especially long-tail ones, often suffer from spelling errors, incomplete phrasing, and ambiguous intent, resulting in mismatches between user expectations and retrieved results. While larg

  25. Sumukha Sathyanarayana, N. Guru Sharan

    In this paper, we establish Kronecker limit type formulas for the generalized Mordell--Tornheim zeta function $\Theta(r,s,t,x)$ as a function of the third variable, in terms of Riemann-zeta and Gamma values. We also give series evaluations of $\Theta(r,s,t,x)$ in terms of Herglotz-Zagier type functions, and their derivatives. As applications of this, we deri

  26. Xunzhou Chen, Tiancheng Sun, Lifei Ye

    How planetary systems form and evolve is a key question in astronomy. Revealing how host star properties, such as elemental abundances, age, and mass, differ from those of non-host stars, and how they correlate with planetary characteristics such as radius, provides new insights into the formation and evolutionary pathways of planetary systems. We determine

  27. Pan Zeming, Tan Naiming, Gao Chao, Yao Zhihai

    The quantum Cheshire cat effect is an important phenomenon in quantum mechanics that reveals the separability of physical properties from their carriers. This effect transcends the classical framework whose attributes must be inherently attached to objects, providing new perspectives for quantum information and precision measurement. According to the quantum

  28. Slim Ibrahim, Quyuan Lin, Lingjun Qian, Edriss S. Titi

    The primitive equations (PEs) model planetary large-scale oceanic and atmospheric dynamics. While it has been shown that there are smooth solutions to the inviscid PEs (also called the hydrostatic Euler equations) with constant temperature (isothermal) that develop stable singularities in finite time, the effect of non-constant temperature on the singularity

  29. Zixuan Gong, Yong Liu, Jiaye Teng

    While looped transformers (termed as Looped-Attn) often outperform standard transformers (termed as Single-Attn) on complex reasoning tasks, the mechanism for this advantage remains underexplored. In this paper, we explain this phenomenon through the lens of loss landscape geometry, inspired by empirical observations of their distinct dynamics at both sample

  30. Jiyun Park, Carlos Cancino-Chacón, Suhit Chiruthapudi, Juhan Nam

    Real-time music alignment, also known as score following, is a fundamental MIR task with a long history and is essential for many interactive applications. Despite its importance, there has not been a unified open framework for comparing models, largely due to the inherent complexity of real-time processing and the language- or system-dependent implementatio

  31. Feifei Liu, Haozhe Wang, Zejun Wei, Qirong Lu

    Current evaluation methods for autonomous driving prediction models rely heavily on simplistic metrics such as Average Displacement Error (ADE) and Final Displacement Error (FDE). While these metrics offer basic performance assessments, they fail to capture the nuanced behavior of prediction modules under complex, interactive, and safety-critical driving sce

  32. Guozhi Liu, Qi Mu, Tiansheng Huang, Xinhua Wang

    Harmful fine-tuning issues present significant safety challenges for fine-tuning-as-a-service in large language models. Existing alignment-stage defenses, e.g., Vaccine, Repnoise, Booster, and T-Vaccine, mitigate harmful fine-tuning issues by enhancing the model's robustness during the alignment phase. While these methods have been proposed to mitigate the i

  33. Meijun Zhou, Gang Mei, Zhengjing Ma, Nengxiong Xu

    Tracking the spatiotemporal evolution of large-scale landslide scars is critical for understanding the evolution mechanisms and failure precursors, enabling effective early-warning. However, most existing studies have focused on single-phase or pre- and post-failure dual-phase landslide identification. Although these approaches delineate post-failure landsli

  34. Debabrata Mandal, Zhihan Peng, Yujie Wang, Praneeth Chakravarthula

    We tackle the challenge of robust, in-the-wild imaging using ultra-thin nanophotonic metalens cameras. Meta-lenses, composed of planar arrays of nanoscale scatterers, promise dramatic reductions in size and weight compared to conventional refractive optics. However, severe chromatic aberration, pronounced light scattering, narrow spectral bandwidth, and low

  35. Raziyeh Zall, Alireza Kheyrkhah, Erik Cambria, Zahra Naseri

    The development of agents with emotional intelligence is becoming increasingly vital due to their significant role in human-computer interaction and the growing integration of computer systems across various sectors of society. Affective computing aims to design intelligent systems that can recognize, evoke, and express human emotions, thereby emulating huma

  36. Parthiv Chatterjee, Shivam Sonawane, Amey Hengle, Aditya Tanna

    Document summarization enables efficient extraction of user-relevant content but is inherently shaped by individual subjectivity, making it challenging to identify subjective salient information in multifaceted documents. This complexity underscores the necessity for personalized summarization. However, training models for personalized summarization has so f

  37. Youshuai Tan, Zhanwei Zhang, Zishuo Ding, Lianyu Zheng

    Floating-point program errors can lead to severe consequences, particularly in critical domains such as military applications. Only a small subset of inputs may induce substantial floating-point errors, prompting researchers to develop methods for identifying these error-inducing inputs. Although existing approaches have achieved some success, they still suf

  38. Donghan Kim

    We construct a multiset space $\mathbb{N}[X]$ over a metric space $X$ that simultaneously enjoys desirable topological properties and admits a natural matching metric $d_{\mathbb{N}[X]}$, making it a metrizable abelian topological monoid whose structure is compatible with the original metric on $X$. This framework extends naturally to the free abelian group

  39. Angel Hsing-Chi Hwang, Fiona Li, Jacy Reese Anthis, Hayoun Noh

    The quickly growing popularity of AI companions poses risks to mental health, personal wellbeing, and social relationships. Past work has identified many individual factors that can drive human-companion interaction, but we know little about how these factors interact and evolve over time. In Study 1, we surveyed AI companion users (N = 303) to map the psych

  40. Chung-Soo Ahn, Rajib Rana, Sunil Sivadas, Carlos Busso

    Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While generative data augmentation offers a promising solution, existing approaches often produce emotionally inconsistent sampl

  41. Wenqing Wang, Muhammad Asif Ali, Ali Shoker, Ruohan Yang

    Human preferences are diverse and dynamic, shaped by regional, cultural, and social factors. Existing alignment methods like Direct Preference Optimization (DPO) and its variants often default to majority views, overlooking minority opinions and failing to capture latent user intentions in prompts. To address these limitations, we introduce \underline{\textb

  42. Nilomber Roy, M. M. B. Sheraj, M. K. Hassan

    We investigate bond percolation on mediation-driven attachment (MDA) networks under the generalized Achlioptas process, where $M>1$ candidate bonds are sampled and the one that minimizes the resulting cluster size is selected the best-of-$M$ rule. This framework offers a systematic approach to investigate how network topology and choice mechanisms jointly sh

  43. Salomon Ibarra, Frida Cantu, Kaixiong Zhou, Li Zhang

    Deep learning models have attracted lots of research attention in time series classification (TSC) task in the past two decades. Recently, deep neural networks (DNN) have surpassed classical distance-based methods and achieved state-of-the-art performance. Despite their promising performance, deep neural networks (DNNs) have been shown to rely on spurious co

  44. Jiayi Mao, Liqun Li, Yanjie Gao, Zegang Peng

    Effective incident management in large-scale IT systems relies on troubleshooting guides (TSGs), but their manual execution is slow and error-prone. While recent advances in LLMs offer promise for automating incident management tasks, existing LLM-based solutions lack specialized support for several key challenges, including managing TSG quality issues, inte

  45. Zonghao Ying, Yangguang Shao, Jianle Gan, Gan Xu

    Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed in real-world environments, they face serious security risks, motivating the design of security evaluation benchmarks. Existing benchmarks provide only partial coverage, typically restricted to narrow scenarios such a

  46. Sanchar Palit, Subhasis Chaudhuri, Biplab Banerjee

    Diffusion models have recently achieved significant success in various image manipulation tasks, including image super-resolution and perceptual quality enhancement. Pretrained text-to-image models, such as Stable Diffusion, have exhibited strong capabilities in synthesizing realistic image content, which makes them particularly attractive for addressing sup

  47. Hua Cai, Shuang Zhao, Liang Zhang, Xuli Shen

    Reasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remains underexplored. In this paper, we introduce Unilaw-R1, a large language model tailored for legal reasoning. With a lightweight 7-billion parameter scale, Unilaw-R1 significantly reduces deployment cost w

  48. Jinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao

    Conventional continual pretraining (CPT) for large language model (LLM) domain adaptation often suffers from catastrophic forgetting and limited domain capacity. Existing strategies adopt layer expansion, introducing additional trainable parameters to accommodate new knowledge. However, the uniform expansion and updates still entangle general and domain lear

  49. S. Rendon Restrepo, O. Gressel

    The stability of a differentially rotating fluid subject to its own gravity is a problem with applications across wide areas of astrophysics--from protoplanetary discs (PPDs) to entire galaxies. The shearing box formalism offers a conceptually simple framework for studying differential rotation in the local approximation. Aimed at self-gravitating, and impor

  50. Zeyu Ling, Xiaodong Gu, Jiangnan Tang, Changqing Zou

    We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked visual modeling with cross-modal contrastive alignment and employs three per-frame prompt tokens that explicitly encode the essential factor

  51. Pîrvu Mihai-Cristian, Marius Leordeanu

    The computer vision domain has greatly benefited from an abundance of data across many modalities to improve on various visual tasks. Recently, there has been a lot of focus on self-supervised pre-training methods through Masked Autoencoders (MAE) \cite{he2022masked,bachmann2022multimae}, usually used as a first step before optimizing for a downstream task,

  52. Shan Jiang, Chenguang Zhu, Sarfraz Khurshid

    JavaScript obfuscators are widely deployed to protect intellectual property and resist reverse engineering, yet their correctness has been largely overlooked compared to performance and resilience. Existing evaluations typically measure resistance to deobfuscation, leaving the critical question of whether obfuscators preserve program semantics unanswered. In

  53. Guoan Li, Xiaofan Shi, Ruixuan Zhang, Yuxiao Song

    Hybrid superconductor-semiconductor(SC-SM) nanowires remain one of the foremost platforms for engineering topological superconductivity and Majorana zero modes(MZMs) towards fault-tolerant topological qubits, especially with the rapid development of artificial Kitaev chains. In contrast to the widely used aluminum(Al)-based hybrids, lead(Pb) offers a bulk su

  54. Yibo Yang

    Deep learning has advanced NLP, but interpretability remains limited, especially in healthcare and finance. Concept bottleneck models tie predictions to human concepts in vision, but NLP versions either use binary activations that harm text representations or latent concepts that weaken semantics, and they rarely model dynamic concept interactions such as ne

  55. Adnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff

    Comparing human and model performance offers a valuable perspective for understanding the strengths and limitations of embedding models, highlighting where they succeed and where they fail to capture meaning and nuance. However, such comparisons are rarely made, as human performance on embedding tasks is difficult to measure. To fill this gap, we introduce H

  56. Hehe Fan, Yi Yang, Mohan Kankanhalli, Fei Wu

    When modeling a given type of data, we consider it to involve two key aspects: 1) identifying relevant elements (e.g., image pixels or textual words) to a central element, as in a convolutional receptive field, or to a query element, as in self-attention, and 2) encoding these tokens effectively. Self-attention can adaptively identify these elements but reli

  57. Keidai Iiyama, Grace Gao

    Recent advancements in lunar positioning, navigation, and timing (PNT) have demonstrated that terrestrial GNSS signals, including weak sidelobe transmissions, can be exploited for lunar spacecraft positioning and timing. While GNSS-based navigation at the Moon has been validated recently, unmodeled ionospheric and plasmaspheric delays remain a significant er

  58. Leonardo Rossetti, Stefano Mancini, Andreas Winter, Joseph Schindler

    The use of coarse graining to connect physical and information theoretic entropies has recently been given a precise formulation in terms of ``observational entropy'', describing entropy for observers with respect to a measurement. Here we consider observers with various locality restrictions, including local measurements (LO), measurements based on local op

  59. Zhi-Fen He, Ren-Dong Xie, Bo Li, Bin Liu

    Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features and recovering missing labels. In this paper, we propose a Co-learning framework for Semantic-aware features and Label recovery (CSL), designed to address both challenges in a unifie

  60. Alann Rosas

    We prove that if $\sum_n n! c_n z^n$ is entire and $c_n$ does not terminate, then $\sum_n c_n z^n$ has infinitely many zeros. We then use this result to give alternative proofs that the Le Roy functions $f_r(z)=\sum_{n=0}^\infty \frac{z^n}{(n!)^r}$ for $r>1$ and Bessel functions $J_\alpha(z)=\sum_{m=0}^\infty \frac{(-1)^m}{m!\Gamma(m+\alpha+1)}\left(\frac{z}

  61. Bo Peng, Zichuan Wang, Sheng Yu, Xiaochuan Jin

    Deep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes apart from real ones, in an objective way. On the other hand, the subjective perception of deepfakes, especially its comp

  62. Kaitao Chen, Shaohao Rui, Yankai Jiang, Jiamin Wu

    Medical vision-language models (VLMs) excel at image-text understanding but typically rely on a single-pass reasoning that neglects localized visual cues. In clinical practice, however, human experts iteratively scan, focus, and refine the regions of interest before reaching a final diagnosis. To narrow this machine-human perception gap, we introduce ViTAR,

  63. Sitong Gong, Yunzhi Zhuge, Lu Zhang, Pingping Zhang

    Audio-Visual Segmentation (AVS) aims to generate pixel-wise segmentation maps that correlate with the auditory signals of objects. This field has seen significant progress with numerous CNN and Transformer-based methods enhancing the segmentation accuracy and robustness. Traditional CNN approaches manage audio-visual interactions through basic operations lik

  64. Takeshi Takahashi, Hiroki Saito

    Bose-Einstein condensates of $^{87}\mathrm{Rb}$ atoms with a hyperfine spin of 2 are open quantum systems, where the atoms are lost through two-body inelastic collisions. In this dissipation process, a collision channel with total spin of 4 is forbidden by angular momentum conservation, which results in magnetization of the atoms remaining in the condensate.

  65. Jiawen Li, Zheng Ning, Yuan Tian, Toby Jia-jun Li

    Large language models (LLMs) enable end-users to delegate complex tasks to autonomous agents through natural language. However, prompt-based interaction faces critical limitations: Users often struggle to specify procedural requirements for tasks, especially those that don't have a factually correct solution but instead rely on personal preferences, such as

  66. Yibo Meng, Ruiqi Chen, Bingyi Liu, Yan Guan

    Generative AI systems are increasingly used by patients seeking everyday health guidance, yet their appropriateness in chronic care contexts remains unclear. Focusing on Type 2 Diabetes Mellitus (T2DM), this paper presents a mixed-methods investigation into how AI-generated health information is interpreted by patients and evaluated by physicians in China. D

  67. Ruohao Li, Hongjun Liu, Leyi Zhao, Zisu Li

    Large language model (LLM) agents have shown remarkable reasoning abilities. However, existing multi-agent frameworks often rely on fixed roles or centralized control, limiting scalability and adaptability in long-horizon reasoning. We introduce SwarmSys, a closed-loop framework for distributed multi-agent reasoning inspired by swarm intelligence. Coordinati

  68. Mingke Lu, Shuaikang Wang, Meng Guo

    Long-term monitoring of numerous dynamic targets can be tedious for a human operator and infeasible for a single robot, e.g., to monitor wild flocks, detect intruders, search and rescue. Fleets of autonomous robots can be effective by acting collaboratively and concurrently. However, the online coordination is challenging due to the unknown behaviors of the

  69. Qiaoyan Peng, Qingqing Wu, Guangji Chen, Wen Chen

    In this paper, we investigate an intelligent reflecting surface (IRS) aided wireless communication system, where active IRSs (AIRSs) are deployed to assist communication between a base station (BS) and users of both the uplink (UL) and downlink (DL). We aim to maximize the weighted sum rate (WSR) of UL and DL communications through joint optimization of BS,

  70. Rahul Vanukuri, Shafi Ullah Khan, Talip Tolga Sarı, Gokhan Secinti

    The growing demand for effective spectrum management and interference mitigation in shared bands, such as the Citizens Broadband Radio Service (CBRS), requires robust radar detection algorithms to protect the military transmission from interference due to commercial wireless transmission. These algorithms, in turn, depend on large, diverse, and carefully lab

  71. Tomasz Krajewski, Włodek Kluźniak

    Nearly every galactic core contains a supermassive compact object, hypothesized to be a Kerr black hole. It was only with the advent of Event Horizon Telescope observations that the predictions of this hypothesis could be observationally tested for our own Galaxy, and the nearby elliptical M87, on spatial scales comparable to the gravitational radius. At the

  72. Saleh Nikooroo, Thomas Engel

    Belief systems are rarely globally consistent, yet effective reasoning often persists locally. We propose a novel graph-theoretic framework that cleanly separates credibility--external, a priori trust in sources--from confidence--an internal, emergent valuation induced by network structure. Beliefs are nodes in a directed, signed, weighted graph whose edges

  73. Sahng-Min Han, Minjae Kim, Jinho Cha, Se-woon Choe

    Deep learning in small and imbalanced biomedical datasets remains fundamentally constrained by unstable optimization and poor generalization. We present the first biomedical implementation of FOSSIL (Flexible Optimization via Sample-Sensitive Importance Learning), a regret-minimizing weighting framework that adaptively balances training emphasis according to

  74. Shafi Ullah Khan, Michel Kulhandjian, Debashri Roy

    Spectrum sharing is a critical strategy for meeting escalating user demands via commercial wireless services, yet its effective regulation and technological enablement, particularly concerning coexistence with incumbent systems, remain significant challenges. Federal organizations have established regulatory frameworks to manage shared commercial use alongsi

  75. Enze Sun, Zhihao Gavin Tang, Yifan Wang

    In online combinatorial allocation, agents arrive sequentially and items are allocated in an online manner. The algorithm designer only knows the distribution of each agent's valuation, while the actual realization of the valuation is revealed only upon her arrival. Against the offline benchmark, Feldman, Gravin, and Lucier (SODA 2015) designed an optimal $0

  76. Oleksiy Dovgoshey, Olga Rovenska

    Let $T$ be a tree of arbitrary finite or infinite order and let $U(T)$ be the set of all ultrametric spaces generated by vertex labelings of $T$. Let ${\bf US}$ denote the class of all ultrametric spaces generated by vertex labelings of star graphs. We prove that the inclusion $U(T)\subseteq {\bf US}$ holds if and only if the longest path in $T$ has a length

  77. Cheng Huang, Weizheng Xie, Zeyu Han, Tsengdar Lee

    Generative AI for automated glaucoma diagnostic report generation faces two predominant challenges: content redundancy in narrative outputs and inadequate highlighting of pathologically significant features including optic disc cupping, retinal nerve fiber layer defects, and visual field abnormalities. These limitations primarily stem from current multimodal

  78. Md. Nayeem, Md Shamse Tabrej, Kabbojit Jit Deb, Shaonti Goswami

    Automatic Speech Recognition (ASR) has undergone a profound transformation over the past decade, driven by advances in deep learning. This survey provides a comprehensive overview of the modern era of ASR, charting its evolution from traditional hybrid systems, such as Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) and Deep Neural Network-HMMs (DNN-H

  79. Jusheng Zhang, Kaitong Cai, Qinglin Zeng, Ningyuan Liu

    Optimizing LLM-based workflows is typically formulated as a global search, where candidate workflows are evaluated based on a scalar metric. This paradigm, however, suffers from a critical flaw: information collapse. By reducing rich, multi-step execution traces to simple success/failure signals, existing methods are rendered blind to the underlying structur

  80. Sebastian Gant, Ben Williams

    Working over an algebraically closed field $k$ of characteristic $0$, we show that the motivic stable homotopy groups of the sphere spectrum can be determined entirely from the motivic homotopy groups of the $p$-completed sphere spectra and the motivic cohomology of the ground field, except possibly for the $0$ and $-1$-stems. Using this, we show that the co

  81. Yiming Ren, Kwan Chuen Chan, Le Zhang, Yin Li

    Accurate photometric redshift (photo-$z$) estimation is a key challenge in cosmology, as uncertainties in photo-$z$ directly limit the scientific return of large-scale structure and weak lensing studies, especially in upcoming Stage IV surveys. The problem is particularly severe for faint galaxies with sparse spectroscopic training data. In this work, we int

  82. Huimin Ji, Zhihua Li, Wenzheng Cheng, Zheng Li

    In the construction of High-Luminosity Large Hadron Collider (HL-LHC) and Future Circular Collider (FCC) experiments, 3D pixel sensors have become indispensable components due to their superior radiation hardness, fast response, and low power consumption. However, there are still significant challenges in the process of 3D sensors manufacturing. In this work

  83. Henan Wang, Hanxin Zhu, Xinliang Gong, Tianyu He

    3D Gaussian Splatting (3DGS) has garnered significant attention due to its superior scene representation fidelity and real-time rendering performance, especially for dynamic 3D scene reconstruction (\textit{i.e.}, 4D reconstruction). However, despite achieving promising results, most existing algorithms overlook the substantial temporal and spatial redundanc

  84. Ruoxing Yang

    We introduce PPOPT - Proximal Policy Optimization using Pretraining, a novel, model-free deep-reinforcement-learning algorithm that leverages pretraining to achieve high training efficiency and stability on very small training samples in physics-based environments. Reinforcement learning agents typically rely on large samples of environment interactions to l

  85. Masakiyo Kitazawa, ShinIchi Esumi, Takafumi Niida, Toshihiro Nonaka

    We derive analytic formulas to reconstruct particle-averaged quantities from experimental results that suffer from the efficiency loss of particle measurements. These formulas are derived under the assumption that the probabilities of observing individual particles are independent. The formulas do not agree with the conventionally used intuitive formulas.

  86. Kazuki Sato

    We study whether the norm one torus associated with a finite separable non-Galois field extension $K/k$ is $p$-retract rational over $k$ for a prime $p$, focusing on the case where the Galois group of the Galois closure of $K/k$ is either the symmetric or the alternating group.

  87. Qi Liu, Yanbo Chen

    This paper investigates McKean-Vlasov backward stochastic variational inequalities (BSVIs) whose generator depends on the joint law of the solution. We first establish the existence and uniqueness of the solution under globally Lipschitz and linear growth conditions. The analysis is then extended to the more general case of locally Lipschitz and non-linear g

  88. Jiaqi Liu, Tong Wang, Su Liu, Xin Hu

    The research evaluates lightweight medical abstract classification methods to establish their maximum performance capabilities under financial budget restrictions. On the public medical abstracts corpus, we finetune BERT base and Distil BERT with three objectives cross entropy (CE), class weighted CE, and focal loss under identical tokenization, sequence len

  89. Soufiane Bentout, Hoang-Hung Vo

    In this paper, we propose and analyze a nonlocal cooperative reaction--diffusion system with free boundaries and drift terms, motivated by directional epidemic spread. Lacking a variational structure but requiring sharper regularity of solutions, the model poses substantial analytical challenges compared with previous works~\cite{Du,Berestycki2016a,Berestyck

  90. Yinghui He, Abhishek Panigrahi, Yong Lin, Sanjeev Arora

    Language models often show little to no improvement (i.e., "saturation") when trained via vanilla supervised fine-tuning (SFT) on data similar to what they saw in their training set (e.g., MATH). We introduce a new fine-tuning strategy, STAT, to train such a student model by using the metacognition ability of a stronger large language model (LLM) as the teac

  91. Junan Chen, Trung Thanh Nguyen, Takahiro Komamizu, Ichiro Ide

    Recent advances in video captioning are driven by large-scale pretrained models, which follow the standard "pre-training followed by fine-tuning" paradigm, where the full model is fine-tuned for downstream tasks. Although effective, this approach becomes computationally prohibitive as the model size increases. The Parameter-Efficient Fine-Tuning (PEFT) appro

  92. Euclid Collaboration, L. Blot, K. Tanidis, G. Cañas-Herrera

    Extracting cosmological information from the Euclid galaxy survey will require modelling numerous systematic effects during the inference process. This implies varying a large number of nuisance parameters, which have to be marginalised over before reporting the constraints on the cosmological parameters. This is a delicate process, especially with such a la

  93. Yaxuan Mao, Yanheng Li, Duo Gong, Pengcheng An

    Dental anxiety is prevalent among children, often leading to missed treatment and potential negative effects on their mental well-being. While several interventions (e.g., pharmacological and psychotherapeutic techniques) have been introduced for anxiety alleviation, the recently emerged virtual reality (VR) technology, with its immersive and playful nature,

  94. Xiao Chen, Weiping Liu, Wei Wang

    GRS 1915+105 has been well studied since its discovery, and is well-known for its complex light curve variability. Using the full currently available Insight-HXMT dataset from July 2017 to June 2023, we make a comprehensive spectral-timing analysis of this source and report four main findings. First, we uncover a QPO frequency rising branch between MJD 58206

  95. Thao Pham

    As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has examined how AI systems scheme against human developers, LLM-to-LLM scheming remains underexplored. We investigate the scheming ability and propensity of frontier LLM agents through t

  96. Shahid Ansari, Vivek Gupta, Bishakh Bhattacharya

    The agricultural sector is rapidly evolving to meet growing global food demands, yet tasks like fruit and vegetable handling remain labor-intensive, causing inefficiencies and post-harvest losses. Automation, particularly selective harvesting, offers a viable solution, with soft robotics emerging as a key enabler. This study introduces a novel hybrid gripper

  97. Jinhua Wu, Yuting Wang, Liukun Yu, Linglong Meng

    Program safety (i.e., absence of undefined behaviors) is critical for correct operation of computer systems. It is usually verified at the source level (e.g., by separation logics) and preserved to the target by verified compilers (e.g., CompCert), thereby achieving end-to-end verification of safety. However, modern safe programming languages like Rust pose

  98. R. M. Albuquerque, S. Narison, D. Rabetiarivony

    We review our estimations on the light scalar $\bar{q}q$, $(\bar{q}q')(\bar{q'}q)$ and $\overline{qq'}qq'$ ($q,q'\equiv u,d,s$) states from relativistic Laplace sum rule (LSR) within stability criteria and including higher order perturbative (PT) corrections up to the (estimated) N5LO. We evaluate the QCD spectral functions at Lowest Order (LO) of PT QCD and

  99. Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao

    As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models c

  100. A. K. Bera, K. S. Chikara, B. Saha, S. M. Yusuf

    The magnetic correlations in double perovskites La2Mn2-xNixO6 (x = 0.5, 0.75, 1.0, 1.25 and 1.5) have been systematically investigated across macroscopic, mesoscopic, and microscopic length scales using temperature-dependent bulk DC magnetization, neutron depolarization, and neutron powder diffraction measurements, respectivitly. The magnetic properties evol