March 2026 arXiv papers — page 67
Showing 6,601–6,700 of 25,974 papers
DecompGrind: A Decomposition Framework for Robotic Grinding via Cutting-Surface Planning and Contact-Force Adaptation
cs.ROShunsuke Araki, Takumi Hachimine, Yuki Saito, Kouhei Ohnishi
Robotic grinding is widely used for shaping workpieces in manufacturing, but it remains difficult to automate this process efficiently. In particular, efficiently grinding workpieces of different shapes and material hardness is challenging because removal resistance varies with local contact conditions. Moreover, it is difficult to achieve accurate estimatio
Abhinaba Basu
We introduce the Dual-View Pheromone Pathway Network (DPPN), an architecture that routes sparse attention through a persistent pheromone field over latent slot transitions, and use it to discover two independent requirements for persistent structural memory in neural networks. Through five progressively refined experiments using up to 10 seeds per condition
Retrieval-Guided Photovoltaic Inventory Estimation from Satellite Imagery for Distribution Grid Planning
eess.IVMuhao Guo, Lihao Mai, Erik Blasch, Jafarali Parol
The rapid expansion of distributed rooftop photovoltaic (PV) systems introduces increasing uncertainty in distribution grid planning, hosting capacity assessment, and voltage regulation. Reliable estimation of rooftop PV deployment from satellite imagery is therefore essential for accurate modeling of distributed generation at feeder and service-territory sc
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
cs.ARHyunwoo Oh, SungHeon Jeong, Suyeon Jang, Hanning Chen
Task-oriented object detection (TOOD) atop CLIP offers open-vocabulary, prompt-driven semantics, yet dense per-window computation and heavy memory traffic hinder real-time, power-limited edge deployment. We present \emph{TorR}, a brain-inspired \textbf{algorithm--architecture co-design} that \textbf{replaces CLIP-style dense alignment with a hyperdimensional
Avoiding Over-smoothing in Social Media Rumor Detection with Pre-trained Propagation Tree Transformer
cs.CLChaoqun Cui, Caiyan Jia
Deep learning techniques for rumor detection typically utilize Graph Neural Networks (GNNs) to analyze post relations. These methods, however, falter due to over-smoothing issues when processing rumor propagation structures, leading to declining performance. Our investigation into this issue reveals that over-smoothing is intrinsically tied to the structural
Haiyue Zhang, Yi Nian, Yue Zhao
What should a developer inspect before deploying an LLM agent: the model, the tool code, the deployment configuration, or all three? In practice, many security failures in agent systems arise not from model weights alone, but from the surrounding software stack: tool functions that pass untrusted inputs to dangerous operations, exposed credentials in deploym
Chengxin Lv, Yihui Li, Hongyu Yang, YunHong Wang
3D semantic occupancy prediction is crucial for autonomous driving. While multi-modal fusion improves accuracy over vision-only methods, it typically relies on computationally expensive dense voxel or BEV tensors. We present Gau-Occ, a multi-modal framework that bypasses dense volumetric processing by modeling the scene as a compact collection of semantic 3D
Chensheng Peng, Quentin Herau, Jiezhi Yang, Yichen Xie
We present UniQueR, a unified query-based feedforward framework for efficient and accurate 3D reconstruction from unposed images. Existing feedforward models such as DUSt3R, VGGT, and AnySplat typically predict per-pixel point maps or pixel-aligned Gaussians, which remain fundamentally 2.5D and limited to visible surfaces. In contrast, UniQueR formulates rec
Xiang Zhang, Wen Jiang, Fei Peng, Wenbin Huang
Video steganography based on block structure, which embeds secret information by modifying Coding Unit (CU) block structure of I-frames, is currently a research hotspot. However, the existing algorithms still suffer from the limitation of poor anti-steganalysis, which results from significantly disrupting the original CU block structure after embedding secre
Fangyuan Ma, Mengzhao Sun, Xuejian Gong, Jun Cai
Platinum oxides are vital catalysts, but their limited thermal stability hinders applications. Recent studies have uncovered a structural transition in two-dimensional platinum oxides that significantly enhances their thermal resilience by several hundred Kelvin. Herein, we demonstrate that this enhanced stability stems from the mechanical robustness of the
Yunheng Li, Hangyi Kuang, Hengrui Zhang, Jiangxia Cao
Multimodal Chain-of-Thought (CoT) reasoning requires large vision-language models to construct reasoning trajectories that interleave perceptual grounding with multi-step inference. However, existing Reinforcement Learning with Verifiable Rewards (RLVR) methods typically optimize reasoning at a coarse granularity, treating CoT uniformly without distinguishin
Youzhi Liu, Li Gao, Liu Liu, Mingyang Lv
Embodied Visual Tracking (EVT), a core dynamic task in embodied intelligence, requires an agent to precisely follow a language-specified target. Yet most existing methods rely on single-agent imitation learning, suffering from costly expert data and limited generalization due to static training environments. Inspired by competition-driven capability evolutio
Canruo Shen, Xintong Ji, Qiong Li, Wenzhi Yang
Gaussian Graphical Models (GGMs) are widely used to infer conditional dependence structures in high-dimensional data. However, standard precision matrix estimators are highly sensitive to data contamination, such as extreme outliers and heavy-tailed noise. In this paper, we propose DROP (Distributionally Robust Optimization), a robust estimation method formu
PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal
cs.AIZining Fang, Cheng Xue, Chunhui Liu, Bin Xu
Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-based desmoking approaches rely on scarce paired supervision and deterministic restoration pipelines, making it difficult to perform exploration or reinforcement-driven refinement under real surgical conditions. W
Takumi Jimbo, Tomomi Matsui
In this research, we address the problem of computing the Shapley value in minimum-cost spanning tree (MCST) games. We introduce the saving game as a key framework for approximating the Shapley value. By reformulating MCST games into their saving-game counterparts, we obtain structural properties that enable multiplicative (relative-error) approximation. Bui
Shuting Sun, Lin Mu, Lizhe Wang, Peng Liu
Change detection of high-resolution remote sensing images is an important task in earth observation and was extensively investigated. Recently, deep learning has shown to be very successful in plenty of remote sensing tasks. The current deep learning-based change detection method is mainly based on conventional long short-term memory (Conv-LSTM), which does
Jun Yang, Dong Wang, Hongxu Yin, Hongpeng Li
Drone detection is pivotal in numerous security and counter-UAV applications. However, existing deep learning-based methods typically struggle to balance robust feature representation with computational efficiency. This challenge is particularly acute when detecting miniature drones against complex backgrounds under severe environmental interference. To addr
URA-Net: Uncertainty-Integrated Anomaly Perception and Restoration Attention Network for Unsupervised Anomaly Detection
cs.CVWei Luo, Peng Xing, Yunkang Cao, Haiming Yao
Unsupervised anomaly detection plays a pivotal role in industrial defect inspection and medical image analysis, with most methods relying on the reconstruction framework. However, these methods may suffer from over-generalization, enabling them to reconstruct anomalies well, which leads to poor detection performance. To address this issue, instead of focusin
MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known Objects
cs.CVShiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam
Multi-camera dynamic Augmented Reality (AR) applications require a camera pose estimation to leverage individual information from each camera in one common system. This can be achieved by combining contextual information, such as markers or objects, across multiple views. While commonly cameras are calibrated in an initial step or updated through the constan
Tao Shen, Wanjie Wang
We study layer-specific community detection in an $L$-layer network $\{A^{(l)}\}_{l\in[L]}$ on a common set of $n$ nodes. Because modern networks are constructed from multi-modal data or with different contexts, the community labels $\pi^{(l)}\in[K]^n$ are layer-dependent and the degree heterogeneity parameters $\theta_i^{(l)}$ vary widely across nodes and l
Analysing LLM Persona Generation and Fairness Interpretation in Polarised Geopolitical Contexts
cs.CLMaida Aizaz, Quang Minh Nguyen
Large language models (LLMs) are increasingly utilised for social simulation and persona generation, necessitating an understanding of how they represent geopolitical identities. In this paper, we analyse personas generated for Palestinian and Israeli identities by five popular LLMs across 640 experimental conditions, varying context (war vs non-war) and ass
The Benjamin-Feir instability in KdV-like equations with general dispersion and monomial nonlinearity
math.APBhavna Kaushik, Bernard Deconinck
Nonlinear waves in dispersive media can be succeptible to modulational instabilities. We examine a category of scalar equations, with general dispersion and monomial nonlinearity, including a large variety of KdV-like equations. For small-amplitude traveling wave solutions, we provide a complete characterization of the spectrum near the origin of the linear
Lars Winkelmann, Wenying Yao
This paper examines how regulatory interventions in high-frequency financial markets affect price discovery. We focus on Breaking news, where dynamic circuit breakers trigger trading halts immediately after the release of macroeconomic fundamentals. Within a high-frequency signal-in-noise model, we show that triggering rules complicate statistical inference
Jing-Bin Cai, Bing Wang
Based on the framework of Koch-Lamm and tensor heat kernel estimates, we obtain a uniform proof of the short-time existence, uniqueness, and continuous dependence for Ricci flows starting from a complete Riemannian metric with bounded curvature. A new ingredient is an effective continuous dependence estimate without the assumption of injectivity radius lower
Po-Rong Lai, Jhen-Dong Lin, Yi-Te Huang, Po-Chen Kuo
The Hierarchical equations of motion (HEOM) method is an important non-perturbative technique, allowing numerically exact treatment of open quantum systems with strong coupling and non-Markovian memory. However, its encoding of bath memory into auxiliary density operators often limits direct access to detailed bath information. In contrast, the reaction-coor
Yu-Tong Su, Zhengxiang Li
Large-scale structure (LSS) and tracer bias connect observable populations to the cosmic matter distribution. While galaxies are standard tracers, transient events such as gravitational-wave sources can also probe LSS despite large localization uncertainties. Fast radio bursts (FRBs), owing to their cosmological distances and dispersion-measure information,
Ziting Pei, Xingye Yue, Xiaotao Zheng
G-expectation, as a sublinear expectation, provides a powerful framework for modeling uncertainty in financial markets. Motivated by the need for robust valuation under model uncertainty, this work develops a unified risk-neutral valuation approach within the G-expectation environment, yielding a nonlinear generalization of the Black-Scholes model, termed th
Siqi Xu, Qilong Cui, Shaowen Xu, Xianbo Chenwei
Two-dimensional altermagnets exhibit exceptional potential for low-power spintronics via nonrelativistic spin splitting and zero net magnetization. Here, we systematically investigate the influence of interlayer interactions on the electronic, magnetic and quantum transport properties of bilayer vanadium oxysulfide (V2S2O), a prototypical layered altermagnet
Shiji Zhao, Mengyang Wang, Shukun Xiong, Fangzhou Chen
With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely use
A. D. Barbour, Gesine Reinert
In directed random graphs, in which edges can be assigned to have one of two directions, or perhaps both, the distance between two vertices $v$ and $v'$ can be computed along paths that are directed from $v$ to $v'$, or along paths that are directed from $v'$ to $v$. These two distances are in general dependent. Here, we approximate their joint distribution
Wafer-to-Wafer Bonding: Part: I -- The Coupled Physics Problem and the 2D Finite Element Implementation
physics.comp-phKamalendu Ghosh, Bhavesh Shrimali, Subin Jeong
Wafer-to-wafer (WxW) bonding is a key enabler for three-dimensional integration, including hybrid bonding for fine-pitch Cu--Cu interconnects. During bonding, wafer deformation and the air entrapped between the wafers interact through a strongly coupled, time-dependent fluid--structure interaction (FSI) that can produce non-intuitive bonding dynamics and pro
MVRD-Bench: Multi-View Learning and Benchmarking for Dynamic Remote Photoplethysmography under Occlusion
cs.CVZuxian He, Xu Cheng, Zhaodong Sun, Haoyu Chen
Remote photoplethysmography (rPPG) is a non-contact technique that estimates physiological signals by analyzing subtle skin color changes in facial videos. Existing rPPG methods often encounter performance degradation under facial motion and occlusion scenarios due to their reliance on static and single-view facial videos. Thus, this work focuses on tackling
Ragib F. Hasan, Matthew Cummins, Waseem Kamleh, Dale Lawlor
A quantitative investigation into the modification of ground-state field structures in two-color QCD (QC$_2$D) is presented at finite chemical potential. Using lattice simulations with Wilson gauge and fermion actions, we explore the chromo-electromagnetic field strengths under varying matter densities. To ensure accurate measurements, we develop and calibra
Shengping Xie, Zekun Wu, Quan Chen, Kaixu Tang
Implicit bias induced by gradient-based algorithms is essential to the generalization of overparameterized models, yet its mechanisms can be subtle. This work leverages the Normalized Steepest Descent} (NSD) framework to investigate how optimization geometry shapes solutions on multiclass separable data. We introduce NucGD, a geometry-aware optimizer designe
Ivan Dobrovolskyi
Context. The problem of comparative evaluation of communication protocols for task orchestration by large language model (LLM) agents is considered. The object of study is the process of interaction between LLM agents and external tools, as well as between autonomous LLM agents, during task orchestration. Objective. The goal of this work is to develop a syst
Ji Eun Song, Eunchae Lee, Juhee Im, Hyunsoo Jang
Account sharing is common in subscription services and is now extending to generative AI platforms, which are still primarily designed for individual use. Sharing often requires workarounds that create new tensions. This study examines how LLM subscriptions are shared and the norms that develop. We combined a survey of 245 users with interviews of 36 partici
Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference
cs.CVZhiceng Shi, Changmiao Wang, Jun Wan, Wenwen Min
While spatial transcriptomics (ST) has advanced our understanding of gene expression in tissue context, its high experimental cost limits its large-scale application. Predicting ST from pathology images is a promising, cost-effective alternative, but existing methods struggle to capture complex cross-slide spatial relationships. To address the challenge, we
Sitong Zhou, Meliha Yetisgen, Mari Ostendorf
Tracking findings in longitudinal radiology reports is crucial for accurately identifying disease progression, and the time-consuming process would benefit from automatic summarization. This work introduces a structured summarization task, where we frame longitudinal report summarization as a timeline generation task, with dated findings organized in columns
TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
cs.CVChunxia Qin, Chenyu Liu, Pengcheng Xia, Jun Du
Tables are pervasive in diverse documents, making table recognition (TR) a fundamental task in document analysis. Existing modular TR pipelines separately model table structure and content, leading to suboptimal integration and complex workflows. End-to-end approaches rely heavily on large-scale TR data and struggle in data-constrained scenarios. To address
Tesshu Hanaka, Daisuke Tsuru
This paper investigates the complexity of finding secluded paths in graphs. We focus on the \textsc{Short Secluded Path} problem and a natural new variant we introduce, \textsc{Shortest Secluded Path}. Formally, given an undirected graph $G=(V, E)$, two vertices $s,t\in V$, and two integers $k,l$, the \textsc{Short Secluded Path} problem asks whether there e
Ji Eun Song, Hyunsoo Jang, Juhee Im, Joongseek Lee
On algorithmic social platforms, exchanging memes via direct messages (DMs) serves as phatic communication that affirms relationships, yet users often interpret these exchanges as signals shaping personalized recommendations, creating tension between relational practice and algorithmic control. This study examines how users perceive DM meme exchanges on Inst
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
cs.CLAbhinaba Basu, Pavan Chakraborty
Language models increasingly show their work by writing step-by-step reasoning before answering. But are these steps genuinely used, or is the answer rigid - fixed before reasoning begins? We introduce the Step-Level Reasoning Capacity (SLRC) metric and prove it is a consistent causal estimator (Theorem 1). We propose LC-CoSR, a training method with Lyapunov
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
cs.CVMincheol Kwon, Minseung Lee, Seonga Choi, Miso Choi
Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images, such as infographics or document layouts, requires these models to generate a large number of visual tokens, leading to s
Scaling atom-by-atom inverse design with nano-topology optimization and diffusion models
physics.app-phChun-Teh Chen, Denvid Lau
The mechanical properties of metallic nanostructures are governed not only by topology but also by crystal symmetry and face-specific surface physics, which are typically absent from continuum topology optimization. We develop an atom-by-atom inverse design framework that combines Nano-Topology Optimization (Nano-TO) with conditional denoising diffusion prob
Xianwei Cao, Dou Quan, Zhenliang Zhang, Shuang Wang
Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem
Hierarchical Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
cs.CLFilippo Morbiato, Markus Keller, Priya Nair, Luca Romano
Mapping Cyber Threat Intelligence (CTI) text to MITRE ATT\&CK technique IDs is a critical task for understanding adversary behaviors and automating threat defense. While recent Retrieval-Augmented Generation (RAG) approaches have demonstrated promising capabilities in this domain, they fundamentally rely on a flat retrieval paradigm. By treating all techniqu
Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration
cs.CLQiyao Sun, Xingming Li, Xixiang He, Ao Cheng
Large language models (LLMs) have achieved remarkable success in various natural language processing tasks, yet they remain prone to generating factually incorrect outputs known as hallucinations. While recent approaches have shown promise for hallucination detection by repeatedly sampling from LLMs and quantifying the semantic inconsistency among the genera
Gongcheng Yue, Xuqiang Wang, Yihan Miao, Bowen Chen
Ferroelectric materials are an ideal platform for high-speed reconfigurable photonic integrated circuits (PICs) for classical and quantum photonic computations, communications, and sensing. Most reconfigurable PIC devices achieve their functionalities via interference and are therefore highly sensitive to phase errors. Under static bias, carrier drift in fer
Universal and efficient graph neural networks with dynamic attention for machine learning interatomic potentials
cs.LGShuyu Bi, Zhede Zhao, Qiangchao Sun, Tao Hu
The core of molecular dynamics simulation fundamentally lies in the interatomic potential. Traditional empirical potentials lack accuracy, while first-principles methods are computationally prohibitive. Machine learning interatomic potentials (MLIPs) promise near-quantum accuracy at linear cost, but existing models still face challenges in efficiency and sta
Yongheng Han
In this paper, using heat kernel estimates and contraction mapping principle, we give a new proof of the existence and uniqueness of mean curvature flow starting from hypersurface with bounded second fundamental form. Moreover, we show the continuous dependence of mean curvature flow on initial data.
Praneeth Vepakomma
We introduce PolyVeil, a protocol for private Boolean summation across $k$ clients that encodes private bits as permutation matrices in the Birkhoff polytope. A two-layer architecture gives the server perfect simulation-based security (statistical distance zero) while a separate aggregator faces \#P-hard likelihood inference via the permanent and mixed discr
Dane Wachs
We prove that over function fields F_q(t), the Tate-Shafarevich group |Sha| is an invariant of the cyclotomic type of the L-polynomial, so that |Sha|-stratified murmuration densities reduce to type-weighted densities with no within-type zero displacement (Theorem A). Over Q, the obstruction vanishes because Satake parameters are continuous: conditioning on L
Thermalization of Weakly Nonintegrable FPUT and Toda Dynamics: A Lyapunov Spectrum Perspective
nlin.CDAniket Patra, Sergej Flach
We study the thermalization slowing down of Fermi-Past-Ulam-Tsingou (FPUT) chains and of Toda chains with nonintegrable boundaries. We focus on the transition from FPUT to harmonic chains, from FPUT to Toda chains with fixed boundaries, and from nonintegrable open boundary Toda to integrable fixed boundary Toda. We compute the Lyapunov spectrum and analyze i
Sidney Xiang, Nicholas David, Dallas Card, Wenhao Sun
Scientific innovation often comes from researchers who pivot across disciplines. However, prior work found that established researchers face productivity penalties when pivoting. Here, we investigate the consequences of pivoting at the beginning of a research career -- doctoral admissions -- when the benefits of importing new ideas might outweigh the switchi
Search for the radiative decays $D^0\to \gamma \bar K_1(1270)^0$ and $D^+\to \gamma K_1(1270)^+$
hep-exBESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson
A search for the radiative decays $D^0\to \gamma \bar K_1(1270)^0$ and $D^+\to \gamma K_1(1270)^+$ is conducted using $20.3~\mathrm{fb}^{-1}$ of $e^+e^-$ annihilation data collected at the center-of-mass energy $\sqrt{s}=3.773$ GeV by the BESIII detector operating at the BEPCII collider. No significant signals are observed, and upper limits on the branching
A Residual-Attention Physics-Informed Neural Network for Irregular Interfaces and Multi-Peak Transport Fields
physics.comp-phBaitong Zhou, Ze Tao, Fujun Liu, Xuan Fang
In complex engineering systems such as electro-thermal-fluid coupling, rapid and accurate prediction of multi-physics fields is essential for advanced applications like digital twins and real-time condition monitoring. Traditional numerical methods often suffer from high computational latency, whereas standard Physics-Informed Neural Networks (PINNs) frequen
Kunal Garg
In this paper, we present new results on finite- and fixed-time convergence for dynamical systems using LaSalle-like invariance principles. In particular, we provide first and second-order non-smooth Lyapunov-like results for finite- and fixed-time convergence, thereby relaxing the requirement of existence a differentiable, positive definite Lyapunov functio
Chenyang Zhang, Qingyue Zhao, Quanquan Gu, Yuan Cao
Transformers have achieved great success across a wide range of applications, yet the theoretical foundations underlying their success remain largely unexplored. To demystify the strong capacities of transformers applied to versatile scenarios and tasks, we theoretically investigate utilizing transformers as students to learn from a class of teacher models.
Aditya Potnis, Francisco Affonso, Shreya Gummadi, Naveen Kumar Uppalapati
Navigating unstructured environments requires assessing traversal risk relative to a robot's physical capabilities, a challenge that varies across embodiments. We present CATNAV, a cost-aware traversability navigation framework that leverages multimodal LLMs for zero-shot, embodiment-aware costmap generation without task-specific training. We introduce a vis
Blake Matheny, Phuong Minh Nguyen, Minh Le Nguyen
The category of figurative language contains many varieties, some of which are non-compositional in nature. This type of phrase or multi-word expression (MWE) includes idioms, which represent a single meaning that does not consist of the sum of its words. For language models, this presents a unique problem due to tokenization and adjacent contextual embeddin
Carlos Ortiz Marrero, Rui Jie Tang, Nathan Wiebe
Entangled quantum probes can achieve Heisenberg-limited measurement precision, but this advantage is typically destroyed by noise. We address this issue by introducing a framework that we call encoded quantum signal processing, which unifies quantum error detection and quantum signal processing into an effective single-qubit framework, and provides a paradig
Enhancing cosmological constraints with nonlinear tanh transformations of Hermite-Gaussian Derivative fields
astro-ph.COZhiwei Min, Ye Ma, Zhujun Jiang, Jiacheng Ding
A key goal in large-scale structure analysis is to extract multi-scale information to improve cosmological parameter constraints. In particular, higher-order derivative fields are especially valuable as they capture the geometric and topological information of the cosmic web that is highly sensitive to cosmological parameters. Traditional derivative-based me
Lirong Che, Zhenfeng Gan, Yanbo Chen, Junbo Tan
Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that achieves this by integrating Large Multimodal Models (LMMs) reasoning with a novel control paradigm. PhotoAgent first translates subjective aesthetic goals into solvable geometric
Guangxu Yang, Jiapeng Zhang
Numbers-on-Forehead (NOF) communication model is a central model in communication complexity. As a restricted variant, one-way NOF model is of particular interest. Establishing strong one-way NOF lower bounds would imply circuit lower bounds, resolve well-known problems in additive combinatorics, and yield wide-ranging applications in areas such as cryptogra
Seungju Han, Konwoo Kim, Chanwoo Park, Benjamin Newman
Synthetic data augmentation helps language models learn new knowledge in data-constrained domains. However, naively scaling existing synthetic data methods by training on more synthetic tokens or using stronger generators yields diminishing returns below the performance of RAG. To break the RAG ceiling, we introduce Synthetic Mixed Training, which combines s
Lishen Qu, Shihao Zhou, Jie Liang, Hui Zeng
Flicker artifacts, arising from unstable illumination and row-wise exposure inconsistencies, pose a significant challenge in short-exposure photography, severely degrading image quality. Unlike typical artifacts, e.g., noise and low-light, flicker is a structured degradation with specific spatial-temporal patterns, which are not accounted for in current gene
Instrument-Splatting++: Towards Controllable Surgical Instrument Digital Twin Using Gaussian Splatting
cs.ROShuojue Yang, Zijian Wu, Chengjiaao Liao, Qian Li
High-quality and controllable digital twins of surgical instruments are critical for Real2Sim in robot-assisted surgery, as they enable realistic simulation, synthetic data generation, and perception learning under novel poses. We present Instrument-Splatting++, a monocular 3D Gaussian Splatting (3DGS) framework that reconstructs surgical instruments as a fu
ABSTRAL: Automatic Design of Multi-Agent Systems Through Iterative Refinement and Topology Optimization
cs.AIWeijia Song, Jiashu Yue, Zhe Pang
How should multi-agent systems be designed, and can that design knowledge be captured in a form that is inspectable, revisable, and transferable? We introduce ABSTRAL, a framework that treats MAS architecture as an evolving natural-language document, an artifact refined through contrastive trace analysis. Three findings emerge. First, we provide a precise me
Kamil Khadiev, Liliya Safina
The Random Forest model is one of the popular models of Machine learning. We present a quantum algorithm for testing (forecasting) process of the Random Forest machine learning model for the Regression problem. The presented algorithm is more efficient (in terms of query complexity or running time) than the classical counterpart.
Krishna Upadhyay, Moshood Fakorede, Umar Farooq
Quantum simulators are a foundational component of the quantum software ecosystem. They are widely used to develop and debug quantum programs, validate compiler transformations, and support empirical claims about correctness and performance. In the absence of large-scale quantum hardware, simulator outputs are often treated as ground truth for algorithm deve
Ning Wang, Cheng Li, Hong Yao
The synthesis of super-heavy nuclei (SHN) through fusion reactions is a critical area of nuclear physics, offering insights into nuclear stability and the limits of the periodic table. However, theoretical predictions of evaporation residue cross sections $ σ_{\rm {ER} }$ remain challenging due to large uncertainties arising from complex reaction mechanisms
Andy Wang, Xu Yan, Brandon McMahan, Michael Zhou
Shared autonomy combines human user and AI copilot actions to control complex systems such as robotic arms. When a task is challenging, requires high dimensional control, or is subject to corruption, shared autonomy can significantly increase task performance by using a trained copilot to effectively correct user actions in a manner consistent with the user&
Paolo Gabriel, Peter Rehani, Zack Drumm, Tyler Troy
This retrospective cohort study used continuous AI monitoring to estimate fall rates by exposure time rather than occupied bed-days. From August 2024 to December 2025, 3,980 eligible monitoring units contributed 292,914 hourly rows, yielding probability-weighted rates of 17.8 falls per 1,000 chair exposure-hours and 4.3 per 1,000 bed exposure-hours. Within t
Amir Azarmehr, Soheil Behnezhad, Alma Ghafari
Large language models (LLMs) can often produce substantially better outputs when allowed to use additional test-time computation, such as sampling, chain of thought, backtracking, or revising partial solutions. Despite the growing empirical success of such techniques, there is limited theoretical understanding of how inference time computation should be stru
Krishna Pada Das, Juan M. Z. Pretel
Dark energy stars (DESs), described by the modified Chaplygin gas (MCG), can be dynamically stable and fall within different observational measurements. In this work, we employ diverse macroscopic properties, such as compactness $C$, moment of inertia $I$, tidal deformability $\Lambda$, gravitational binding energy $E_g$ and $f$-mode nonradial pulsation freq
Wenyue Chen, Wenjue Chen, Peng Li, Qinghe Wang
Recent advances in 3D generation have improved the fidelity and geometric details of synthesized 3D assets. However, due to the inherent ambiguity of single-view observations and the lack of robust global structural priors caused by limited 3D training data, the unseen regions generated by existing models are often stochastic and difficult to control, which
Manognya Lokesh Reddy, Zheng Liu
Accurate inter-vehicle distance estimation is a cornerstone of advanced driver assistance systems and autonomous driving. While LiDAR and radar provide high precision, their cost prohibits widespread adoption in mass-market vehicles. Monocular vision offers a low-cost alternative but suffers from scale ambiguity and sensitivity to environmental disturbances.
Yongjia Weng, Lufeng Liu, Zhonggui Chen, Xuan Zhou
High-order quadrilateral meshes offer superior accuracy and computational efficiency in numerical simulations. However, existing methods struggle to simultaneously preserve boundary/interface features, ensure high quality, and achieve efficient generation, particularly for complex geometries where degenerate and inverted elements frequently occur. To address
Zhi Sun, Wenming Zhang, Yi Wei, Liren Yu
Large Language Models (LLMs) are equipped with profound semantic knowledge, making them a natural choice for injecting semantic generalization into personalized search systems. However, in practice we find that directly fine-tuning LLMs on industrial personalized tasks (e.g. next item prediction) often yields suboptimal results. We attribute this bottleneck
Shota Kanasugi, Riki Toshio, Kazunori Maruyama, Hirotaka Oshima
Quantum simulation of molecular electronic structure is one of the most promising applications of quantum computing. However, achieving chemically accurate predictions for strongly correlated systems requires quantum phase estimation (QPE) on fault-tolerant quantum computing (FTQC) devices. Existing resource estimates for typical FTQC architectures suggest t
AgriPestDatabase-v1.0: A Structured Insect Dataset for Training Agricultural Large Language Model
cs.AIYagizhan Bilal Durak, Ahsan Ul Islam, Shahidul Islam, Ashley Morgan-Olvera
Agricultural pest management increasingly relies on timely and accurate access to expert knowledge, yet high quality labeled data and continuous expert support remain limited, particularly for farmers operating in rural regions with unstable/no internet connectivity. At the same time, the rapid growth of AI and LLMs has created new opportunities to deliver p
Jingwei Liao, Bo Chen, Klara Nahrstedt, Zhisheng Yan
Given the popularity of 360{\deg} images on social media platforms, 360{\deg} image compression becomes a critical technology for media storage and transmission. Conventional 360{\deg} image compression pipeline projects the spherical image into a single 2D plane, leading to issues of oversampling and distortion. In this paper, we propose a novel viewport-ba
Artur Kawalec
In this article, we develop a k-free zeta Dirichlet series into a Laurent series with a simple pole, and prove a Stieltjes like formula for the expansion coefficients of the regular part. We also investigate another analytical continuation of these series and develop a formula for $\zeta(\tfrac{1}{k})$ for positive integer $k\geq 2$ in terms of the k-free in
Distributed Hybrid Feedback for Global Pose Synchronization of Multiple Rigid Body Systems on $SE(3)$
eess.SYFengyu Lin, Miaomiao Wang, Housheng Su, Abdelhamid Tayebi
This paper investigates the problem of pose synchronization for multiple rigid body systems evolving on the matrix Lie group $\SE(3)$. We propose a distributed hybrid feedback control scheme with global asymptotic stability guarantees using relative pose and group velocity measurements. The key idea consists of constructing a new potential function on $\SE(3
J. P. Velasquez-Rodriguez
Let $p$ be a prime number, and let $\mathbb{G}$ be a compact $p$-adic Lie group. This work provides multiplier theorems for invariant operators on $\mathbb{G}$ acting on $L^r_\alpha(\mathbb{G})$, $1<r<\infty$, $\alpha>0$, in terms of the Ruzhansky-Turunen difference operators and Saloff-Coste's condition. As an application, a Littlewood-Paley decomposition i
Explainable Threat Attribution for IoT Networks Using Conditional SHAP and Flow Behavior Modelling
cs.CRSamuel Ozechi, Jennifer Okonkwoabutu
As the Internet of Things (IoT) continues to expand across critical infrastructure, smart environments, and consumer devices, securing them against cyber threats has become increasingly vital. Traditional intrusion detection models often treat IoT threats as binary classification problems or rely on opaque models, thereby limiting trust. This work studies mu
Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases
cs.AIDubai Li, Yuxiang He, Yan Hu, Yu Tian
Observational studies can yield clinically actionable evidence at scale, but executing them on real-world databases is open-ended and requires coherent decisions across cohort construction, analysis, and reporting. Prior evaluations of LLM agents emphasize isolated steps or single answers, missing the integrity and internal structure of the resulting evidenc
DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona
cs.CLJanghyeok Choi, Jaewon Lee, Sungzoon Cho
Data scarcity remains a persistent challenge in low-resource domains. While existing data augmentation methods leverage the generative capabilities of large language models (LLMs) to produce large volumes of synthetic data, these approaches often prioritize quantity over quality and lack domain-specific strategies. In this work, we introduce DALDALL, a perso
Yuanyuan Sun, Tiexin Guo, Qiang Tu
By making full use of the inherent connection between the theory of random conjugate spaces and the theory of classical conjugate spaces, in this paper we establish a random demiclosedness principle for a random asymptotically nonexpansive mapping, which generalizes Xu's classical demiclosedness principle from a uniformly convex Banach space to a complete ra
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
cs.CVAo Cheng, Xingming Li, Xuanyu Ji, Xixiang He
Electronic Navigational Charts (ENCs) are the safety-critical backbone of modern maritime navigation, yet it remains unclear whether multimodal large language models (MLLMs) can reliably interpret them. Unlike natural images or conventional charts, ENCs encode regulations, bathymetry, and route constraints via standardized vector symbols, scale-dependent ren
Matrix-Free Stabilized BDF Schemes for Semilinear Parabolic Equations with Unconditional Maximum Bound Principle Preservation and Energy Stability
math.NAHaishen Dai, Huan Lei, Bin Zheng
We develop a family of stabilized backward differentiation formula (sBDF) schemes of orders one through four for semilinear parabolic equations. The proposed methods are designed to achieve three properties that are rarely available simultaneously in high-order time discretizations: unconditional preservation of the maximum bound principle (MBP), uncondition
Po-Rong Lai, Hsien-Chao Jan, Jhen-Dong Lin, Yueh-Nan Chen
Enhancement of quantum battery performance is a popular subject in quantum thermodynamics. An interesting phenomenon is the quick charging effect [Phys. Rev. Res. 6, 023136 (2024)], which has been explored by utilizing a quantum interferometric technique known as superposition of trajectories. A similar technique used to boost quantum battery performance is
Ruisen Tu, Arth Shukla, Sohyun Yoo, Xuanlin Li
Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained geometry, and high-dimensional continuous actions, making standard imitation learning insufficient. We introduce a framework for learning sp
Jiayi Qin, Jingwei Li, Chuan Wu
Screen-shooting robust watermarking aims to imperceptibly embed extractable information into host images such that the watermark survives the complex distortion pipeline of screen display and camera recapture. However, achieving high extraction accuracy while maintaining satisfactory visual quality remains an open challenge, primarily because the screen-shoo
Human vs. NAO: A Computational-Behavioral Framework for Quantifying Social Orienting in Autism and Typical Development
cs.HCVartika Narayani Srinet, Anirudha Bhattacharjee, Braj Bhushan, Bishakh Bhattacharya
Responding to one's name is among the earliest-emerging social orienting behaviors and is one of the most prominent aspects in the detection of Autism Spectrum Disorder (ASD). Typically developing children exhibit near-reflexive orienting to their name, whereas children with ASD often demonstrate reduced frequency, increased latency, or atypical patterns of
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
cs.CVWonJun Moon, Hyun Seok Seong, Jae-Pil Heo
Video Object-Centric Learning seeks to decompose raw videos into a small set of object slots, but existing slot-attention models often suffer from severe over-fragmentation. This is because the model is implicitly encouraged to occupy all slots to minimize the reconstruction objective, thereby representing a single object with multiple redundant slots. We ta
Min Li, Jinghui He, Gang Li, Jiachen Li
The purpose of multimodal industrial anomaly detection is to detect complex geometric shape defects such as subtle surface deformations and irregular contours that are difficult to detect in 2D-based methods. However, current multimodal industrial anomaly detection lacks the effective use of crucial geometric information like surface normal vectors and 3D sh
Purui Bai, Tao Wu, Jiayang Sun, Xinyue Liu
The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however, are limited to static images or single videos, overlooking the complex interactions across multiple videos. To address t
KALAVAI: Predicting When Independent Specialist Fusion Works -- A Quantitative Model for Post-Hoc Cooperative LLM Training
cs.CLRamchand Kumaresan
Independently trained domain specialists can be fused post-hoc into a single model that outperforms any individual specialist, and the gain is predictable: gain = 0.82 x divergence - 2.72 (R^2 = 0.856, n=6, 3-26% divergence). This enables practitioners to estimate cooperative value before committing compute. Below ~3.3% divergence, gains approach zero.In the
Ruidi Chang, Jiawei Zhou, Hanjie Chen
Large language models (LLMs) solve complex problems by generating multi-step reasoning traces. Yet these traces are typically analyzed from only one of two perspectives: the sequence of tokens across different reasoning steps in the generated text, or the hidden-state vectors across model layers within one step. We introduce PRISM (Probabilistic Reasoning In