February 2026 arXiv papers — page 3
Showing 201–300 of 20,995 papers
Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification
cs.CLAnastasia Zhukova, Terry Ruas, Jan Philip Wahle, Bela Gipp
Research in CDCR remains fragmented due to heterogeneous dataset formats, varying annotation standards, and the predominance of the CDCR definition as the event coreference resolution (ECR). To address these challenges, we introduce uCDCR, a unified dataset that consolidates diverse publicly available English CDCR corpora across various domains into a consis
Qiushi Zhao, Zihan Feng, Ximing Xie, Hao Qin
The pinching-antenna systems (PASS) enable blockage mitigation in urban micro (UMi) networks through flexible antenna placement. However, the joint optimization of antenna positions and beamforming precoding is inherently nonconvex and becomes significantly more challenging under user mobility. To address this issue, we propose a bilevel optimization framewo
Li Sun, Zhenhao Huang, Silei Chen, Lanxu Yang
Multi-domain graph pre-training integrates knowledge from diverse domains to enhance performance in the target domains, which is crucial for building graph foundation models. Despite initial success, existing solutions often fall short of answering a fundamental question: how is knowledge integrated or transferred across domains? This theoretical limitation
Subramanyam Sahoo, Vinija Jain, Aman Chadha, Divya Chaudhary
The persistent militarization of large reasoning models stems not from technical necessity but from governance arrangements that strip researchers of meaningful authority to refuse harmful transfers and deployments. Existing accountability mechanisms such as model cards and responsible AI statements operate as reputational signals detached from decision maki
Debarpita Banerjee, Debasmita Lohar, Sumana Ghosh
Modern cyber-physical systems, such as automotive control, rely on feedback controllers that regulate the system towards desired a setpoint. In practice, however, the controller must also be scheduled efficiently on resource-constrained processors, where the choice of numerical precision for controller implementation directly affects both control quality and
Fanqi Pu, Lei Jiang, Wenming Yang
The performance of robotic imitation learning is fundamentally limited by data quality and training strategies. Prevalent sampling strategies on RLBench suffer from severe keyframe redundancy and imbalanced temporal distribution, leading to inefficient memory usage and unstable optimization. Moreover, reprojecting point clouds onto multi-view images with a b
Designing the Haystack: Programmable Chemical Space for Generative Molecular Discovery
physics.chem-phYuchen Zhu, Donghai Zhao, Yangyang Zhang, Yitong Li
Chemical space exploration underlies drug discovery, yet most generative models treat chemical space as a fixed, implicitly learned distribution, focusing on sampling molecules rather than deliberately designing the space itself. We introduce SpaceGFN, a generative framework that elevates chemical space to a programmable computational object: a controllable
Maeve Smyth, Jorge Kohanoff
We present a first-principles molecular dynamics study of an excess electron in condensed phase models of solvated DNA bases. Calculations on increasingly large microsolvated clusters taken from liquid phase simulations show that adiabatic electron affinities increase systematically upon solvation, as for optimized gas-phase geometries. Dynamical simulations
From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation
cs.CLRaneen Younis, Suvinava Basak, Lukas Chavez, Zahra Ahmadi
The rapid growth of biomedical literature and curated databases has made it increasingly difficult for researchers to systematically connect biomarker mechanisms to actionable drug combination hypotheses. We present AI Co-Scientist (CoDHy), an interactive, human-in-the-loop system for biomarker-guided drug combination hypothesis generation in cancer research
Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark
cs.CVLijing Cai, Zhan Shi, Chenglong Huang, Jinyao Wu
Recently, Spectral Compressive Imaging (SCI) has achieved remarkable success, unlocking significant potential for dynamic spectral vision. However, existing reconstruction methods, primarily image-based, suffer from two limitations: (i) Encoding process masks spatial-spectral features, leading to uncertainty in reconstructing missing information from single
Changxing Liu, Zichen Chao, Siheng Chen
Collaborative perception leverages data exchange among multiple agents to enhance overall perception capabilities. However, heterogeneity across agents introduces domain gaps that hinder collaboration, and this is further exacerbated by an underexplored issue: modality isolation. It arises when multiple agents with different modalities never co-occur in any
Mwayi Sonkhanani, Symon Chibaya, Clement N. Nyirenda
Student repetition in secondary education imposes significant resource burdens, particularly in resource-constrained contexts. Addressing this challenge, this study introduces a unified machine learning framework that simultaneously predicts pass/fail outcomes and continuous grades, a departure from prior research that treats classification and regression as
Ziheng Xi, Zihang Ao, Yitao Wang, Mingeze Gao
Accurate 3D hand pose and pressure sensing is essential for immersive human-computer interaction, yet simultaneously achieving both in mobile scenarios remains a significant challenge. We present WristPP, a camera-based wrist-worn system that estimates 3D hand pose and per-vertex pressure from a single wide-FOV RGB frame in real time. A Vision Transformer (V
$A_{\alpha}$-Spectra of $Q$- and $T$-Join Graphs with Applications to Cospectral Constructions
math.COMainak Basunia, Pratima Panigrahi
For $\alpha \in [0,1]$, the $A_{\alpha}$-matrix of a graph $G$ is defined by $A_{\alpha}(G) = \alpha D(G) + (1- \alpha) A(G)$, where $A(G)$ and $D(G)$ denote the adjacency matrix and the diagonal degree matrix of $G$, respectively. In this paper, we study the $A_{\alpha}$-characteristic polynomials and $A_{\alpha}$-spectra of graphs obtained via four recentl
Data-Centric Benchmark for Label Noise Estimation and Ranking in Remote Sensing Binary Building Segmentation
cs.CVKeiller Nogueira, Codrut-Andrei Diaconu, Dávid Kerekes, Jakob Gawlikowski
High-quality pixel-level annotations are essential for the semantic segmentation of remote sensing imagery. However, such labels are expensive to obtain and often affected by noise due to the labor-intensive and time-consuming nature of pixel-wise annotation, which makes it challenging for human annotators to label every pixel accurately. Annotation errors c
Jinkui Wan
We first introduce a new presentation for the mirabolic Hecke algebra $\mathscr{H}_{n,R}(q)$ over an arbitrary commutative ring $R$ and derive a new basis. Based on this presentation, specializing to the case of $\mathscr{H}_n(q)$ over the field $\mathbb{C}(q)$, we construct a basis for the cocenter of $\mathscr{H}_n(q)$, which facilitates the definition of
Li Sun, Lanxu Yang, Jiayu Tian, Bowen Fang
Detecting out-of-distribution (OOD) graphs is crucial for ensuring the safety and reliability of Graph Neural Networks. In unsupervised graph-level OOD detection, models are typically trained using only in-distribution (ID) data, resulting in incomplete feature space characterization and weak decision boundaries. Although synthesizing outliers offers a promi
Grigory Sapunov
AI code agents excel at isolated tasks yet struggle with multi-file software engineering requiring architectural understanding. We introduce Theory of Code Space (ToCS), a benchmark that evaluates whether agents can construct, maintain, and update coherent architectural beliefs during codebase exploration. Agents explore procedurally generated codebases unde
Yongxi Huang, Zhuohang Wang, Wenjing Tang, Xinyu He
Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant information - is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural langua
Li Sun, Ming Zhang, Wenxin Jin, Zhongtian Sun
Hypergraphs are the natural description of higher-order interactions among objects, widely applied in social network analysis, cross-modal retrieval, etc. Hypergraph Neural Networks (HGNNs) have become the dominant solution for learning on hypergraphs. Traditional HGNNs are extended from message passing graph neural networks, following the homophily assumpti
Daniel Mejer Christensen, Katja Stougård Jørgensen, Josefine Palsgaard Wyrtz, Jennie Torp Overgaard
Research in Human-Computer Interaction (HCI) has shown that caring for others, including both humans (e.g., close friends) and computers (e.g., Tamagotchi), can have a positive effect on people's wellbeing. However, we know less about the potential role of conversational AI in such settings. In this work, we explore how AI chatbots can support plant care and
Hassan Wasswa, Hussein Abbass, Timothy Lynar
The increasing incidence of IoT-based botnet attacks has driven interest in advanced learning models for detection. Recent efforts have focused on leveraging attention mechanisms to model long-range feature dependencies and Graph Neural Networks (GNNs) to capture relationships between data instances. Since GNNs require graph-structured input, tabular NetFlow
Jiahao Cui, Feng Yu, Linzuo Zhang, Yu Hu
Inertial Odometry (IO) has gained attention in quadrotor applications due to its sole reliance on inertial measurement units (IMUs), attributed to its lightweight design, low cost, and robust performance across diverse environments. However, most existing learning-based inertial odometry systems for quadrotors either use only IMU data or include additional d
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
cs.CLAditya Kamlesh Parikh, Cristian Tejedor-Garcia, Catia Cucchiarini, Helmer Strik
Reliable and interpretable automated assessment of second-language (L2) speech remains a central challenge, as large speech-language models (SpeechLLMs) often struggle to align with the nuanced variability of human raters. To address this, we introduce a rubric-guided reasoning framework that explicitly encodes multi-aspect human assessment criteria: accurac
Emmanuele Battista, Salvatore Capozziello, Stefano Pastore
We investigate Extended Geometric Trinity of Gravity at both classical and quantum cosmological levels using the minisuperspace approach. Adopting Noether symmetries to select viable models, we examine metric-affine theories of gravity, in particular the extensions of General Relativity, Teleparallel Equivalent General Relativity and Symmetric Teleparallel E
Chenggang Rong, Tao Han, Zhiyuan Zhao, Yaowu Fan
Counting is a core capability for multimodal large language models (MLLMs), yet there is no unified counting dataset to rigorously evaluate this ability across image, text, and audio. We present UNICBench, a unified multimodal, multi level counting benchmark and evaluation toolkit with accurate ground truth, deterministic numeric parsing, and stratified repo
Xianfa Hu, Fazhan Geng, Wansheng Wang
In this paper, we consider the integrating factor midpoint method for wave-type equations and derive optimal order a posteriori error estimates. We first introduce an integrating factor midpoint approximation defined by the piecewise linear approximate solutions, and derive suboptimal order residual-based error estimates using the energy technique. Hence the
András London
We study edge partitions of a bipartite graph into induced-$2K_2$-free bipartite graphs, i.e.\ into Ferrers (chain) graphs. We define $\fp(G)$ as the minimum number of parts in such a partition. We prove general lower and upper bounds in terms of induced matchings and Dilworth widths of neighborhood posets. We compute the parameter exactly for paths and even
Yuchen Hou, Lin Zhao
Vision-Language-Action (VLA) models achieve over 95% success on standard benchmarks. However, through systematic experiments, we find that current state-of-the-art VLA models largely ignore language instructions. Prior work lacks: (1) systematic semantic perturbation diagnostics, (2) a benchmark that forces language understanding by design, and (3) linguisti
Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking
cs.CLTianpeng Zheng, Zhehan Jiang, Jiayi Liu, Shicong Feng
The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data contamination, and lack calibrated measurement properties for fine-grained performance tracking. We propose and validate
Fair in Mind, Fair in Action? A Synchronous Benchmark for Understanding and Generation in UMLLMs
cs.AIYiran Zhao, Lu Zhou, Xiaogang Xu, Zhe Liu
As artificial intelligence (AI) is increasingly deployed across domains, ensuring fairness has become a core challenge. However, the field faces a "Tower of Babel'' dilemma: fairness metrics abound, yet their underlying philosophical assumptions often conflict, hindering unified paradigms-particularly in unified Multimodal Large Language Models (UMLLMs), whe
Cencen Liu, Dongyang Zhang, Wen Yin, Jielei Wang
Visual autoregressive (VAR) models have recently emerged as a promising alternative for image generation, offering stable training, non-iterative inference, and high-fidelity synthesis through next-scale prediction. This encourages the exploration of VAR for image super-resolution (ISR), yet its application remains underexplored and faces two critical challe
Chenhao Zhang, Muxing Li, Feng Liu, Weitong Chen
Evaluating machine unlearning remains challenging, as existing methods typically require retraining reference models or performing membership inference attacks, both of which rely on prior access to training configuration or supervision labels, making them impractical in realistic scenarios. Motivated by the fact that most unlearning algorithms remove a smal
Qin Guo, Tianyu Yang, Xuanhua He, Fei Shen
Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level consistency, or produce copy-paste artifacts where subjects a
Rongsheng Wang, Minghao Wu, Hongru Zhou, Zhihan Yu
Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds great promise for biomedical applications such as drug discovery, organ-on-chip systems, and disease mechanism studies, while also showing pot
S. Cocchi, F. Loi, M. Murgia, P. Marchegiani
Galaxy clusters imprint a distinctive signature on the cosmic microwave background through the thermal Sunyaev-Zel'dovich (SZ) effect, which enables to study the intracluster plasma distribution and makes them powerful cosmological probes. We present the first Sardinia Radio Telescope (SRT) detection of the SZ effect in the galaxy cluster MACS J1752+4440 at
On phase-isometries between the unit spheres of the Banach space of continuous real-valued functions
math.FAYuta Enami, Izuho Matsuzaki
For a locally compact Hausdorff space $L$, we denote by $C_0(L,\mathbb{R})$ the Banach space of all continuous real-valued functions on $L$ vanishing at infinity, endowed with the supremum norm. In this paper, we prove that every surjective phase-isometry $T\colon S(C_0(X,\mathbb{R}))\to S(C_0(Y,\mathbb{R}))$ between the unit spheres of $C_0(X,\mathbb{R})$ a
Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Research
cs.CLYubo Dong, Nianhao You, Yuxuan Hou, Zixun Sun
While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon planning, massive evidence gathering, and synthesis across heterogeneous sources-remains largely unexplored. We introduce Super Research, a task for complex autonomous research tasks t
High Resolution Microscopy and Raman Spectroscopic Studies on the Freshest Mukundpura Meteorite, Rajasthan, India: Presence of Nanodiamond
physics.geo-phD. Chandrasekharam, U. Govind, R. P. Tripathi, T. H. Amir
Carbonaceous Chondrites have special significance in the stellar evolution and in particular in the evolution of life on earth. The carbonaceous meteorite that fell in Mukundpura village, Jaipur, Rajasthan on 6th June 2017 is one such rare CM2 (Carbonaceous Chondrite) carbonaceous meteorite. We carried out high resolution scanning and transmission electron m
Orientational ordering and close packing properties of quasi-one-dimensional hard Gaussian overlap particles
cond-mat.softSakineh Mizani, Péter Gurin, Szabolcs Varga
We investigate the orientational ordering and close-packing behavior of hard Gaussian overlap (HGO) particles, which are confined into a quasi-one-dimensional (q1D) channel. In the channel, particles are allowed to move along the channel and to rotate in three dimensions. Using the transfer operator method, we show that oblate particles align with their shor
A Sensitivity Analysis of the Surrogate Index Approach for Estimating Long-Term Treatment Effects
econ.EMYanqin Fan, Carlos A. Manzanares, Hyeonseok Park, Yuan Qi
This paper develops a sensitivity analysis of the surrogacy assumption for the surrogate index approach in Athey et al. [2025b]. We introduce "Weighted Surrogate Indices (WSIs)," the analog of the surrogate index under the surrogacy assumption. We show that under comparability, the ATE on WSI identifies the ATE on the long-term outcome when a copula of the t
Jianheng Tang, Yajiang Huang, Kejia Fan, Feijiang Han
Federated Learning (FL) is a popular distributed learning paradigm to break down data silo. Traditional FL approaches largely rely on gradient-based updates, facing significant issues about heterogeneity, scalability, convergence, and overhead, etc. Recently, some analytic-learning-based work has attempted to handle these issues by eliminating gradient-based
Jie Cao, Tianwei Lin, Zhenxuan Fan, Bo Yuan
Long chain-of-thought~(CoT) has become a dominant paradigm for enhancing the reasoning capability of large reasoning models~(LRMs); however, the performance gains often come with a substantial increase in reasoning budget. Recent studies show that existing CoT paradigms tend to induce systematic overthinking, unnecessarily coupling reasoning capability with
Yoshiaki Goto
The Wirtinger integral is one of the integral representations of the Gauss hypergeometric function. Its integrand can be regarded as a multivalued function on an elliptic curve. In this paper, we study an analogue of the Wirtinger integral on a hyperelliptic curve of genus two, introduced by Mizutani and Watanabe. We investigate the associated twisted homolo
Jinhan Xu, Xing Tang, Houpeng Yang, Haoran Zhang
Symbolic music generation is a challenging task in multimedia generation, involving long sequences with hierarchical temporal structures, long-range dependencies, and fine-grained local details. Though recent diffusion-based models produce high quality generations, they tend to suffer from high training and inference costs with long symbolic sequences due to
Yucheng Zeng, Shupeng Li, Daxiang Dong, Ruijie Xu
Progress in software-engineering agents is increasingly constrained by the scarcity of executable, scalable, and realistic data for training and evaluation. This scarcity stems from three fundamental challenges in existing pipelines: environments are brittle and difficult to reproduce across languages; synthesizing realistic, system-level bugs at scale is co
Qi Hu, Jingyu Wang, Huriye Atilgan, Armin Lak
Three-photon (3-P) fluorescence microscopy enables deep in vivo imaging with subcellular resolution, but its performance is fundamentally constrained by the maximum permissible laser power required to avoid tissue heating and photodamage. Under these power-limited conditions, fluorescence signal generation, image contrast, and achievable imaging depth are st
The Bombieri--van der Poorten Formula for Partial Quotients of Higher Degree Algebraic Irrationals
math.NTKarsten Müller
The fundamental relationship between the partial quotients $b_{n+1}$ of an algebraic irrational $\alpha = \sqrt[m]{k}$ and its corresponding algebraic form $d_n = |p_n^m - k q_n^m|$ was elegantly proposed by Bombieri and van der Poorten. In this paper, we work out the explicit analytical details of the framework for any degree $m \geq 3$. We provide a closed
Physics-Based Seismic Hazard and Risk Assessment: A New Paradigm for Earthquake Forecasting
physics.geo-phDavide Zaccagnino, Didier Sornette
Epistemic uncertainty in probabilistic seismic hazard assessment (PSHA) is commonly addressed through a logic-tree framework that combines weighted alternative models to characterize the range of plausible hazard outcomes. Implicit in this approach is a critical assumption: that the available model class provides an adequate representation of the underlying
Haomin Qi, Bohan Liu, Zihan Dai, Yunkai Gao
TopoEdge is a topology-grounded, edge-deployable framework for end-to-end software-defined networking (SDN) configuration generation and repair, motivated by the brittleness of configuration artefacts under topology variation and by strict operational constraints on latency, privacy, and on-site execution. TopoEdge represents each target topology as a router
Yunqing Liu, Yi Zhou, Wenqi Fan
Molecule representation learning is crucial for understanding and predicting molecular properties. However, conventional atom-centric models, which treat chemical bonds merely as pairwise interactions, often overlook complex bond-level phenomena like resonance and stereoselectivity. This oversight limits their predictive accuracy for nuanced chemical behavio
Ray Telikani, Amir H. Gandomi
Neural contextual bandits are vulnerable to adversarial attacks, where subtle perturbations to rewards, actions, or contexts induce suboptimal decisions. We introduce AdvBandit, a black-box adaptive attack that formulates context poisoning as a continuous-armed bandit problem, enabling the attacker to jointly learn and exploit the victim's evolving policy. T
Modulating biodiversity through higher-order interactions and intraspecific competition in rock-paper-scissors dynamics
q-bio.PEChunpeng Du, Haoshu Wang, Yikang Lu, Lijuan Qin
Understanding the mechanisms that govern species coexistence and biodiversity represents a fundamental challenge in ecology. This study extends the classic rock-paper-scissors model by introducing a context-dependent higher-order interaction mechanism where intraspecific competition is dynamically regulated by local resource availability. Crucially, our quan
Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu
Multimodal Large Language Models (MLLMs) have achieved remarkable performance but remain vulnerable to jailbreak attacks that can induce harmful content and undermine their secure deployment. Previous studies have shown that introducing additional inference steps, which disrupt security attention, can make MLLMs more susceptible to being misled into generati
Sen Zhang, Jianguo Wei, Wenhuan Lu, Xianghu Yue
The Transformer-based Whisper model has achieved state-of-the-art performance in Automatic Speech Recognition (ASR). However, its Multi-Head Attention (MHA) mechanism results in significant GPU memory consumption due to the linearly growing Key-Value (KV) cache usage, which is problematic for many applications especially with long-form audio. To address this
Renkang Song, Junbo Xu, Yanzhen Yin, Yu Yin
Chip-scale nonlinear optics enables strong light-matter interactions within compact devices, serving as a fundamental platform for multifunctional integrated photonics from classical optical signal processing to quantum information technologies. Transition metal dichalcogenide (TMDC) waveguides have recently emerged as a highly promising platform owing to th
Second-order estimates for degenerate complex $k$-Hessian and Christoffel-Minkowski equations
math.APYasheng Lyu
It is known that the complex $k$-Hessian equation admits almost $C^{1,1}$ regularity (i.e., $\sup\Delta u<\infty$) and the Christoffel-Minkowski equation admits $C^{1,1}$ regularity under the sharp degenerate condition $f^{1/(k-1)}\in C^{1,1}$ for a nonnegative right-hand side $f$. Assuming instead the alternative sharp degenerate condition $f^{3/(2k-2)}\in
Yihua Shao, Kang Chen, Feng Xue, Siyu Chen
In operating rooms (OR), world-scale multi-view 3D tracking supports downstream applications such as surgeon behavior recognition, where physically meaningful quantities such as distances and motion statistics must be measured in meters. However, real clinical deployments rarely satisfy the geometric prerequisites for stable multi-view fusion and tracking: c
Heng Du, Yong Suk Moon, Koji Shimizu
We define the big crystalline site for a log scheme and prove the basic properties. In particular, we show the boundedness, base change, and perfectness theorems for the crystalline higher direct image of quasi-coherent crystals between fine log schemes. We also introduce the big absolute crystalline sites and discuss the Frobenius isogeny property of the cr
A Stable and General Quantum Fractional-Step Lattice Boltzmann Method for Incompressible Flows
quant-phYang Xiao, Liming Yang, Chang Shu, Yinjie Du
Quantum computing shows substantial potential in accelerating simulations and alleviating memory bottlenecks in computational fluid dynamics (CFD), owing to its inherent properties of superposition and entanglement. The lattice Boltzmann method (LBM), being largely algebraic in nature, has inspired the development of various quantum LBMs. However, most exist
Phase-Space Analysis of generalised Fractional Anharmonic and Ornstein-Uhlenbeck Semigroups on Weighted Modulation Spaces
math.FAAparajita Dasgupta, Uttam Kumar Dolai
We develop a phase-space framework for fractional generalised anharmonic oscillators and their heat semigroups on weighted modulation spaces. We consider operators of the form \[ \mathcal{H}_{k,l}=(-\Delta)^{l}+V(x), \] where $V$ is a strictly positive homogeneous potential of polynomial growth of order $2k$. By studying a H\"ormander metric adapted to the q
Andreas Gaugenrieder, Hari Hara Balasubramaniam, Jannik Möhrle, Rüdiger Daub
Skill-based programming of robots provides a flexible approach for automation. Existing solutions neglect the optimization of motion sequences, leading to inefficiencies in execution. This work introduces a planning method that enhances skill-based robot programming by integrating motion sequence optimization. This optimization leads to a new MoveContinuousS
Yuzo Maruyama
This paper is a follow-up to Maruyama and Strawderman (2006, Journal of Statistical Planning and Inference), which identified a new class of generalized Bayes estimators with a particularly simple form for estimating a normal variance under entropy loss. Although their previous work established the Bayesianity of these estimators, it did not provide a closed
Shiya Zhang, Yuhan Zhan, Ruixi Su, Ruihan Sun
Evaluating persona-aligned empathy in LLM-based dialogue agents remains challenging. User states are latent, feedback is sparse and difficult to verify in situ, and seemingly supportive turns can still accumulate into trajectories that drift from persona-specific needs. We introduce EMPA, a process-oriented framework that evaluates persona-aligned support as
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
cs.PFJiaqi Wang, Jingwei Sun, Jiyu Luo, Han Li
GPU architectural simulation is orders of magnitude slower than native execution, necessitating workload sampling for practical speedups. Existing methods rely on hand-crafted features with limited expressiveness, yielding either aggressive sampling with high errors or conservative sampling with constrained speedups. To address these issues, we propose GCL-S
Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
cs.CVYu Wang, Shengjie Zhao
Weakly supervised video anomaly detection (WS-VAD) involves identifying the temporal intervals that contain anomalous events in untrimmed videos, where only video-level annotations are provided as supervisory signals. However, a key limitation persists in WS-VAD, as dense frame-level annotations are absent, which often leaves existing methods struggling to l
Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi, Frank Eliassen
We present PM2Lat, a fast and generalized framework for accurately predicting the latency of deep neural network models on GPUs, with special focus on NVIDIA. Unlike prior methods that rely on deep learning models or handcrafted heuristics, PM2Lat leverages the Single-Instruction-Multiple-Thread architecture of GPUs to model execution time of DNN models. Fir
Laser-induced, blackbody-radiation-assisted rovibrational cooling of symmetric-top molecular ions: NH3+ and ND3+
physics.chem-phArchisman Sinha, Brianna R. Heazlewood, Nabanita Deb
Quantum-state preparation of molecular ions is a prerequisite for precision spectroscopy and controlled studies of cold ion-molecule dynamics. While such control has been extensively developed for diatomic ions and proposed for linear polyatomic ions, corresponding strategies for symmetric-top molecular ions remain largely unexplored. We present a theoretica
Planetary Desert around Compact Binaries: Dynamical Instability Triggered by Resonance-Induced Eccentricity Excitation
astro-ph.EPBin Liu, Dong Lai
Compact binaries with orbital periods shorter than about 7 days show an absence of transiting planets, a feature known as the ``circumbinary planet desert". The physical mechanism behind this desert remains unclear. We investigate its origin by simulating the long-term dynamics of multi-planet circumbinary systems with evolving inner binaries. Our simulation
Multiple Inputs and Mixwd data for Alzheimer's Disease Classification Based on 3D Vision Transformer
cs.CVJuan A. Castro-Silva, Maria N. Moreno Garcia, Diego H. Peluffo-Ordoñez
The current methods for diagnosing Alzheimer Disease using Magnetic Resonance Imaging (MRI) have significant limitations. Many previous studies used 2D Transformers to analyze individual brain slices independently, potentially losing critical 3D contextual information. Region of interest-based models often focus on only a few brain regions despite Alzheimer'
Aparna Gupte, Jiahui Liu, Luowen Qian, Justin Raizes
One-time programs (OTPs) aim to let a user evaluate a program on a single input while revealing nothing else. Classical OTPs require hardware assumptions, and even with quantum information, OTPs for deterministic functionalities remain impossible due to gentle-measurement attacks (Broadbent, Gutoski and Stebila, 2013). While recent works achieve positive res
Ke Cao, Xuanhua He, Xueheng Li, Lingting Zhu
Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-world, high-resolution scenarios. To bridge this gap, we syste
Adaptive Dynamic Dehazing via Instruction-Driven and Task-Feedback Closed-Loop Optimization for Diverse Downstream Task Adaptation
cs.CVYafei Zhang, Shuaitian Song, Huafeng Li, Shujuan Wang
In real-world vision systems,haze removal is required not only to enhance image visibility but also to meet the specific needs of diverse downstream tasks.To address this challenge,we propose a novel adaptive dynamic dehazing framework that incorporates a closed-loop optimization mechanism.It enables feedback-driven refinement based on downstream task perfor
Chenyu Zheng, Rongzhen Wang, Xinyu Zhang, Chongxuan Li
Generative foundation models are increasingly scaled in both width and depth, posing significant challenges for stable feature learning and reliable hyperparameter (HP) transfer across model sizes. While maximal update parameterization ($\mu$P) has provided a principled solution to both problems for width scaling, existing extensions to the joint width-depth
Yucheng Zeng, Weipeng Lu, Linyun Liu, Shupeng Li
The evolution of Large Language Models (LLMs) from static instruction-followers to autonomous agents necessitates operating within complex, stateful environments to achieve precise state-transition objectives. However, this paradigm is bottlenecked by data scarcity, as existing tool-centric reverse-synthesis pipelines fail to capture the rigorous logic of re
Are LLMs Reliable Code Reviewers? Systematic Overcorrection in Requirement Conformance Judgement
cs.SEHaolin Jin, Huaming Chen
Large language models (LLMs) have become essential tools in software development, widely used for requirements engineering, code generation and review tasks. Software engineers often rely on LLMs to verify if code implementation satisfy task requirements, thereby ensuring code robustness and accuracy. However, it remains unclear whether LLMs can reliably det
Ming Wen, Kun Yang, Xin Chen, Jingyu Zhang
Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal safety alignment via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is a primary mitigation strategy, current
Mathematical Foundations of Poisoning Attacks on Linear Regression over Cumulative Distribution Functions
cs.LGAtsuki Sato, Martin Aumüller, Yusuke Matsui
Learned indexes are a class of index data structures that enable fast search by approximating the cumulative distribution function (CDF) using machine learning models (Kraska et al., SIGMOD'18). However, recent studies have shown that learned indexes are vulnerable to poisoning attacks, where injecting a small number of poison keys into the training data can
Xianhao Zhou, Jianghao Wu, Lanfeng Zhong, Ku Zhao
Cone-beam CT (CBCT) is routinely acquired in radiotherapy but suffers from severe artifacts and unreliable Hounsfield Unit (HU) values, limiting its direct use for dose calculation. Synthetic CT (sCT) generation from CBCT is therefore an important task, yet paired CBCT--CT data are often unavailable or unreliable due to temporal gaps, anatomical variation, a
Cuiying Pei, Hongjoo Ha, Sen Shao, Shihao Zhu
Materials with honeycomb lattice structures exhibit unique electronic properties arising from their distinctive atomic arrangements. Their weakly coupled nature facilitates modulation by external stimuli, which leads to a diverse range of physical phenomena, particularly superconductivity. Here, we report the discovery of a pressure-induced M-shaped double-d
Shangda Wu, Ziya Zhou, Yongyi Zang, Yutong Zheng
We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190 multiple-choice questions through four stages - each followed by manual verification: 1) compiling a representative music list
Yandong Yan, Junwei Peng, Shijie Li, Chenxi Li
Autonomous agents are increasingly entrusted with complex, long-horizon tasks, ranging from mathematical reasoning to software generation. While agentic workflows facilitate these tasks by decomposing them into multi-step reasoning chains, reliability degrades significantly as the sequence lengthens. Specifically, minor interpretation errors in natural-langu
Osama A. Marzouk
The unsteady variations in the near wake of a moving cylinder induce lift and drag forces on it, which are customarily normalized and expressed in terms of nondimensional lift and drag coefficients. While there are already several wake oscillator models for either a fixed or moving cylinder, special attention was given to modeling the lift coefficient for th
Denis Blessing, Lorenz Richter, Julius Berner, Egor Malitskiy
Sampling from unnormalized densities using diffusion models has emerged as a powerful paradigm. However, while recent approaches that use least-squares `matching' objectives have improved scalability, they often necessitate significant trade-offs, such as restricting prior distributions or relying on unstable optimization schemes. By generalizing these metho
FedUAF: Uncertainty-Aware Fusion with Reliability-Guided Aggregation for Multimodal Federated Sentiment Analysis
cs.LGXianxun Zhu, Zezhong Sun, Imad Rida, Erik Cambria
Multimodal sentiment analysis in federated learning environments faces significant challenges due to missing modalities, heterogeneous data distributions, and unreliable client updates. Existing federated approaches often struggle to maintain robust performance under these practical conditions. In this paper, we propose FedUAF, a unified multimodal federated
Swapnil Parekh
Image captioning models are encoder-decoder architectures trained on large-scale image-text datasets, making them susceptible to adversarial attacks. We present CaptionFool, a novel universal (input-agnostic) adversarial attack against state-of-the-art transformer-based captioning models. By modifying only 7 out of 577 image patches (approximately 1.2% of th
Manuella Christelle Tossa, Fernando Madrigal, Ryan Blosser, Asma Jodeiri Akbarfam
Reliable grid operation depends on accurate and timely telemetry, making modern power systems vulnerable to communication layer cyberattacks. This paper evaluates how Denial of Service (DoS), Denial of Data (DoD), and False Data Injection (FDI) attacks disrupt the IEEE 14 bus system using a MATLAB only, time stepped simulation framework built on MATPOWER. Th
Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu
World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such models from distinct viewpoints of the same environment without any parameter sharing or coordination. After training, their internal representations exhibit a striking emergent property: the two latent spaces are rela
Wenjie Wei, Xiaolong Zhou, Malu Zhang, Ammar Belatreche
Spiking neural networks (SNNs) offer an energy-efficient alternative to traditional neural networks due to their event-driven computing paradigm. However, recent advancements in spiking transformers have focused on improving accuracy with large-scale architectures, which require significant computational resources and limit deployment on resource-constrained
Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation
cs.CVZhen Zhou, Jian Liu, Biwen Lei, Jing Xu
Reinforcement learning (RL) has demonstrated remarkable success in text and image generation, yet its potential in 3D generation remains largely unexplored. Existing attempts typically rely on offline direct preference optimization (DPO) method, which suffers from low training efficiency and limited generalization. In this work, we aim to enhance both the tr
Electron-positron Pair Production in Global GRMHD Simulations of Black Hole Accretion Flows
astro-ph.HEHo-Sang Chan, Jason Dexter, Mitchell C. Begelman
We present global, three-dimensional general relativistic magnetohydrodynamic simulations of accreting black holes that incorporate pair physics. Pairs are modeled as a passive scalar that maintains a constant temperature. For high accretion rate models, we observe a maximum pair fraction of $\sim \mathcal{O}(0.01)$, consistent with those inferred from some
Swapnil Parekh
Every mechanistic circuit carries an invisible asterisk: it reflects not just the model's computation, but the analyst's choice of pruning threshold. Change that choice and the circuit changes, yet current practice treats a single pruned subgraph as ground truth with no way to distinguish robust structure from threshold artifacts. We introduce CIRCUS, which
Zhimin Wang, Chenyu Gu, Feng Lu
Eye-hand coordinated interaction is becoming a mainstream interaction modality in Virtual Reality (VR) user interfaces.Current paradigms for this multimodal interaction require users to learn predefined gestures and memorize multiple gesture-task associations, which can be summarized as an ``Operation-to-Intent" paradigm. This paradigm increases users' learn
Lei Liu, Xiaoning Yu, Kang Chen, Jiahui Huang
Tropical cyclone (TC) forecasting is critical for disaster warning and emergency response. Deep learning methods address computational challenges but often neglect physical relationships between TC attributes, resulting in predictions lacking physical consistency. To address this, we propose Phys-Diff, a physics-inspired latent diffusion model that disentang
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
cs.SEBoxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals that one in five "solved" patches from the top-30 agents are semantically incorrect, passing only because weak test suites fail to expose their errors. We present SWE-ABS, an adversa
Yuyang Chen, Linqian Zeng, Yijin ZHou, Hengjie Li
Diffusion models have achieved remarkable success in generative AI, yet their computational efficiency remains a significant challenge, particularly for Diffusion Transformers (DiTs) requiring intensive full-attention computation. While existing acceleration approaches focus on content-agnostic uniform optimization strategies, we observe that different regio
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
cs.SDSwapnil Parekh
ASR systems exhibit persistent performance disparities across accents, but whether these gaps reflect superficial biases or deep structural vulnerabilities remains unclear. We introduce ACES, a three-stage audit that extracts accent-discriminative subspaces from ASR representations, constrains adversarial attacks to them, and tests whether removing them impr
Quan Kong, Yanru Xiao, Yuhao Shen, Cong Wang
Learning efficient and expressive visual representation has long been the pursuit of computer vision research. While Vision Transformers (ViTs) gradually replace traditional Convolutional Neural Networks (CNNs) as more scalable vision learners, their applications are plagued by the quadratic complexity of the self-attention mechanism. To address the challeng
Ziquan Wang, Haobo Wang, Ke Chen, Lei Feng
Machine Learning often involves various imprecise labels, leading to diverse weakly supervised settings. While recent methods aim for universal handling, they usually suffer from complex manual pre-work, ignore the relationships between associated labels, or are unable to batch process due to computational design flaws, resulting in long running times. To ad
Haodong Zhao, Jinming Hu, Zhaomin Wu, Zongru Wu
Federated Instruction Tuning (FIT) enables collaborative instruction tuning of large language models across multiple organizations (clients) in a cross-silo setting without requiring the sharing of private instructions. Recent findings on natural backdoors and the existing training data collection method suggest that poisoned samples may be pervasive and ina