November 2025 arXiv papers — page 54
Showing 5,301–5,400 of 22,271 papers
Jinming Gao, Yijing Wang, Wentao Zhang, Rui Zhao
This paper investigates the problem of resilient control for multi-agent systems in the presence of Byzantine adversaries via an active secure neighbor selection framework. A pre-discriminative graph is first constructed to characterize the admissible set of candidate neighbors for each agent. Based on this graph, a dynamic in-neighbor selection strategy is
Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing
cs.CVCheng Jiang, Yihe Yan, Yanxiang Wang, Chun Tung Chou
While Wi-Fi sensing offers a compelling, privacy-preserving alternative to cameras, its practical utility has been fundamentally undermined by a lack of robustness across domains. Models trained in one setup fail to generalize to new environments, hardware, or users, a critical "domain shift" problem exacerbated by modest, fragmented public datasets. We shif
Mitchell Keegan, Michael Forbes, Paul Corry, Mahdi Abolghasemi
Decision trees are a popular machine learning model which are traditionally trained by heuristic methods. Massive improvements in computing power and optimisation techniques has led to renewed interest in learning globally optimal decision trees. Empirical evidence shows that optimal classification trees (OCTs) have better out-of-sample performance than heur
Dan Li, Hye-Bin Shin, Yeon-Woo Choi
Due to the significant variability in electroencephalo-gram (EEG) signals across individuals, knowledge acquired from previous subjects is often overwritten as new subjects are introduced in continual EEG decoding tasks. Existing methods mainly rely on storing historical data from seen subjects as replay buffers to mitigate forgetting, which is impractical u
Designing Gamified Social Interaction for Gen Z in the Metaverse: A Framework-Oriented Systematic Literature Review
cs.HCBaitong Xie, Mohd Fairuz Shiratuddin, Mostafa Hamadi, Joo Yeon Park
Gamification plays a pivotal role in enhancing user engagement in the Metaverse, particularly among Generation Z users who value autonomy, immersion, and identity expression. However, current research lacks a cohesive framework tailored to designing gamified social experiences in immersive virtual environments. This study presents a framework-oriented system
Exploring the changes in brain network SC-FC coupling patterns of partial sleep deprivation based on DTI-fMRI fusion analysis
q-bio.NCMengyuan Liu, Jing Hu, Zhenzhen Ru, Ruomeng Quan
Sleep disorder is a serious global public health issue, with cognitive-emotional dysfunction being a core symptom. The analysis of multimodal MRI data provides an effective method for detecting sleep deprivation-induced neural network abnormalities. The structure-function coupling (SC-FC) integrates functional connectivity with white matter structural inform
Perturbing the Derivative: Doubly Wild Refitting for Model-Free Evaluation of Opaque Machine Learning Predictors
cs.LGHaichen Hu, David Simchi-Levi
We study the problem of excess risk evaluation for empirical risk minimization (ERM) under convex losses. We show that by leveraging the idea of wild refitting, one can upper bound the excess risk through the so-called "wild optimism," without relying on the global structure of the underlying function class but only assuming black box access to the training
Shiyi Mu, Zichong Gu, Zhiqi Ai, Anqi Liu
Compared to monocular 3D object detection, stereo-based 3D methods offer significantly higher accuracy but still suffer from high computational overhead and latency. The state-of-the-art stereo 3D detection method achieves twice the accuracy of monocular approaches, yet its inference speed is only half as fast. In this paper, we propose StereoDETR, an effici
Bhuvan Sachdeva, Karan Uppal, Abhinav Java, Vineeth N. Balasubramanian
Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like depth estimation or object counting. Finetuning on one task can unpredictably affect performance on others, making task-specific finetuning challenging. In this paper, we address this challenge through a systematic
STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
cs.CVJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang
We present STCDiT, a video super-resolution framework built upon a pre-trained video diffusion model, aiming to restore structurally faithful and temporally stable videos from degraded inputs, even under complex camera motions. The main challenges lie in maintaining temporal stability during reconstruction and preserving structural fidelity during generation
Karen Gunderson, Karen Meagher, Joy Morris, Venkata Raghu Tej Pantangi
We give a Hilton-Milner Theorem for the $r$-independent sets in the graph that is the union of copies of $K_k$. That is, we determine the maximum intersecting families of $r$-independent sets in this graph, subject to the condition that the sets in a family do not all share a common element. As a by-product, we also find a tight upper bound for the sum of si
TDLight: A Framework for Incremental Light-Curve Management and Continuous Classification Optimization
astro-ph.IMXinghang Yu, Ce Yu, Zeguang Shao, Chen Wang
Time-series observations of celestial objects are fundamental to studying variable and transient phenomena. As time-domain surveys continue to scale up, timely automated classification is essential for rapid follow-up. Many existing approaches treat storage and analysis as separate stages, archiving data first and analyzing it later. Yet each source is a con
Renchu Guan, Xuyang Li, Yachao Zhang, Wei Pang
Hypergraphs, as a generalization of traditional graphs, naturally capture high-order relationships. In recent years, hypergraph neural networks (HNNs) have been widely used to capture complex high-order relationships. However, most existing hypergraph neural network methods inherently rely on the homophily assumption, which often does not hold in real-world
Lukas Twist, Jie M. Zhang
LLMs can generate useful code, but their outputs often contain small implementation-level bugs with large behavioural effects. In this paper, we ask whether natural-language code summaries provide useful diagnostic context for repairing these errors. We use summary-mediated repair as a simple prompt-only test of their diagnostic value: an LLM first summarise
A Novel Dual-Stream Framework for dMRI Tractography Streamline Classification with Joint dMRI and fMRI Data
cs.CVHaotian Yan, Bocheng Guo, Jianzhong He, Nir A. Sochen
Streamline classification is essential to identify anatomically meaningful white matter tracts from diffusion MRI (dMRI) tractography. However, current streamline classification methods rely primarily on the geometric features of the streamline trajectory, failing to distinguish between functionally distinct fiber tracts with similar pathways. To address thi
ConceptGuard: Proactive Safety in Text-and-Image-to-Video Generation through Multimodal Risk Detection
cs.CVRuize Ma, Minghong Cai, Yilei Jiang, Jiaming Han
Recent progress in video generative models has enabled the creation of high-quality videos from multimodal prompts that combine text and images. While these systems offer enhanced controllability, they also introduce new safety risks, as harmful content can emerge from individual modalities or their interaction. Existing safety methods are often text-only, r
Sanjit Bhowmick, Deepak Kumar Dalai, Sihem Mesnager
The hull of a linear code is defined as the intersection of the code and its dual. This concept was initially introduced to classify finite projective planes. The hull plays a crucial role in determining the complexity of algorithms used to check the permutation equivalence of two linear codes and compute a linear code's automorphism group. Research has show
Shivanshu Mishra, Ruixue Li, Dongjea Seo, Anil Adhikari
WSe2 p-MOSFETs with Nb-doped WS2 contacts formed using atomic layer deposition are demonstrated. The devices are fabricated using a technique that aligns the contact metallization with the Nb-doped WS2 contacts using a selective oxidation process. Devices with source/drain spacing of 0.15 um have on-state current of 103 uA/um at VDS = -1 V at a channel carri
Chenhong Zhou, Jie Chen, Zaifeng Yang
Neural operators have shown great potential in solving a family of Partial Differential Equations (PDEs) by modeling the mappings between input and output functions. Fourier Neural Operator (FNO) implements global convolutions via parameterizing the integral operators in Fourier space. However, it often results in over-smoothing solutions and fails to captur
Jun Kimura
In the outer solar system beyond Jupiter, water ice is a dominant component of planetary bodies, and most solid objects in this region are classified as icy bodies. Icy bodies display a remarkable diversity of geological, geophysical, and atmospheric processes, which differ fundamentally from those of the rocky terrestrial planets. Evidence from past and ong
Bashar Talafha, Amin Abu Alhassan, Muhammad Abdul-Mageed
Zero-shot ASR for Arabic remains challenging: while multilingual models perform well on Modern Standard Arabic (MSA), error rates rise sharply on dialectal and accented speech due to linguistic mismatch and scarce labeled data. We study context-aware decoding as a lightweight test-time adaptation paradigm that conditions inference on external side informatio
Senmao Tian, Xiang Wei, Shunli Zhang
Class imbalance remains a critical challenge in semi-supervised learning (SSL), especially when distributional mismatches between labeled and unlabeled data lead to biased classification. Although existing methods address this issue by adjusting logits based on the estimated class distribution of unlabeled data, they often handle model imbalance in a coarse-
Zihan Wang, Zhongkui Ma, Xinguo Feng, Chuan Yan
Deep neural networks (DNNs) have become valuable intellectual property of model owners, due to the substantial resources required for their development. To protect these assets in the deployed environment, recent research has proposed model usage control mechanisms to ensure models cannot be used without proper authorization. These methods typically lock the
Soumya Adhikari, Abhinava Bhattacharjee, Amitabh Virmani
We study fully BPS and a broad class of half-BPS stationary configurations of four-dimensional Euclidean N=2 supergravity with higher-derivative interactions. Working within the off-shell conformal supergravity framework of de Wit and Reys (arXiv:1706.04973), we analyse the complete set of Killing spinor equations and obtain the corresponding algebraic and d
HOPPS: Hardware-Aware Optimal Phase Polynomial Synthesis with Blockwise Optimization for Quantum Circuits
quant-phXinpeng Li, Ji Liu, Shuai Xu, Paul Hovland
Blocks composed of {CNOT, Rz} are ubiquitous in modern quantum applications, notably in circuits such as QAOA ansatzes and quantum adders. After compilation, many of them exhibit large CNOT counts or depths, which lowers fidelity. Therefore, we introduce HOPPS: a SAT-based hardware-aware optimal phase polynomial synthesis algorithm that could generate {CNOT,
Tesla Zhang, Asher Kornfeld, Rui Li, Sonya Simkin
Semantic typing has become a powerful tool for program verification, applying the technique of logical relations as not only a proof method, but also a device for prescribing program behavior. In recent work, Yao et al. scaled semantic typing to the verification of timed message-passing protocols, which are prevalent in, e.g., IoT and real-time systems appli
Meng Sun, Richard H. D. Townsend, Hongbo Xia, Jifeng Liu
We revisit the tidal evolution of the WASP-12 system using direct numerical calculations with the GYRE-tides code. WASP-12b is a hot Jupiter on a 1.1-day orbit around a slightly evolved F-type star. Its observed orbital decay rate, $|\dot{P}_{\rm orb}/P_{\rm orb}| \approx 3.2\,\mathrm{Myr}^{-1}$, provides a strong constraint on stellar tidal dissipation. We
Accelerated Transformer Energization Sequence for Inverter Based Resources in Black-Start Procedures with Active Flux Trajectory Manipulation in the Stationary Reference Frame
eess.SYJiyu Lee, Shenghui Cui
This paper proposes advanced soft-magnetization techniques to enable ultra-fast and reliable black-start of grid-forming (GFM) converters. Conventional hard-magnetization with well-established three-phase voltages during transformer energization induces severe inrush currents due to flux offset, which can damage power semiconductor devices. To overcome this
Qiang Hou, Cong Yu, Shu-ichiro Inutsuka
We investigate the impact of a low-mass planet on dust coagulation, and its consequent feedback on planetary migration, using a linear analysis of the coupled dust-gas hydrodynamic equations. Dust coagulation is incorporated via a single-size approximation. In the co-orbital region of the planet, we find that the growth of dust size is significantly suppress
Xintao Chen, Xiaohao Xu, Bozhong Zheng, Yun Liu
Unsupervised visual anomaly detection from multi-view images presents a significant challenge: distinguishing genuine defects from benign appearance variations caused by viewpoint changes. Existing methods, often designed for single-view inputs, treat multiple views as a disconnected set of images, leading to inconsistent feature representations and a high f
Modeling Bioelectric State Transitions in Glial Cells: An ASAL-Inspired Computational Approach to Glioblastoma Initiation
physics.bio-phWiktoria Agata Pawlak
Understanding how glioblastoma (GBM) emerges from initially healthy glial tissue requires models that integrate bioelectrical, metabolic, and multicellular dynamics. This work introduces an ASAL-inspired agent-based framework that simulates bioelectric state transitions in glial cells as a function of mitochondrial efficiency (Meff), ion-channel conductances
Xuanzhao Dong, Wenhui Zhu, Yujian Xiong, Xiwen Chen
Color fundus photography (CFP) is central to diagnosing and monitoring retinal disease, yet its acquisition variability (e.g., illumination changes) often degrades image quality, which motivates robust enhancement methods. Unpaired enhancement pipelines are typically GAN-based, however, they can distort clinically critical vasculature, altering vessel topolo
Tsogtgerel Gantumur
We give a short proof that for a bounded domain $\Omega\subset\mathbb{R}^n$ and continuous boundary data $g\in C(\partial\Omega)$ admitting a continuous finite-energy extension $\phi\in H^{1}(\Omega)\cap C(\bar\Omega)$, the minimizer of the Dirichlet energy \[ E(v) = \int_{\Omega} |\nabla v|^{2}\,dx, \qquad v-\phi\in H^{1}_{0}(\Omega), \] coincides with the
Hao Wu, Shoucheng Song, Chang Yao, Sheng Han
In multi-agent systems, explicit cognition of teammates' decision logic serves as a critical factor in facilitating coordination. Communication (i.e., ``\textit{Tell}'') can assist in the cognitive development process by information dissemination, yet it is inevitably subject to real-world constraints such as noise, latency, and attacks. Therefore, building
Xiyong Yan
We investigate the multisigns of Hamiltonian circles in the multisigned complete graph \(\Sigma_n := (K_n, \sigma, \mathbb{F}_2^m)\). The \emph{multisign} of a circle \(C\) is defined as the sum \[ \sigma(C) := \sum_{e \in E(C)} \sigma(e). \] For a fixed \(m\) and sufficiently large \(n\), we show that the set of multisigns of Hamiltonian circles \[ \{\sigma
From Features to Reference Points: Lightweight and Adaptive Fusion for Cooperative Autonomous Driving
cs.CVYongqi Zhu, Morui Zhu, Qi Chen, Deyuan Qu
We present RefPtsFusion, a lightweight and interpretable framework for cooperative autonomous driving. Instead of sharing large feature maps or query embeddings, vehicles exchange compact reference points, e.g., objects' positions, velocities, and size information. This approach shifts the focus from "what is seen" to "where to see", creating a sensor- and m
Xueyu Du, Lilian Zhang, Fuan Duan, Xincan Luo
Filter-based visual inertial navigation system (VINS) has attracted mobile-robot researchers for the good balance between accuracy and efficiency, but its limited mapping quality hampers long-term high-accuracy state estimation. To this end, we first propose a novel filter-based stereo VINS, differing from traditional simultaneous localization and mapping (S
Xiaotong Huang, He Zhu, Tianrui Ma, Yuxiang Xiong
3D Gaussian splatting (3DGS) has emerged as a promising direction for SLAM due to its high-fidelity reconstruction and rapid convergence. However, 3DGS-SLAM algorithms remain impractical for mobile platforms due to their high computational cost, especially for their tracking process. This work introduces Splatonic, a sparse and efficient real-time 3DGS-SLAM
Xin Liao
We study the stationary phi^6 model given by the equation -phi''(x) + 2 phi(x) - 8 phi(x)^3 + 6 phi(x)^5 = 0 for x in R, and establish sharp quantitative stability estimates for configurations close to two weakly interacting kinks. More precisely, there exist constants a > 0 and epsilon > 0 such that, for any function u in L-infinity satisfying || u - H_{0,1
Blinking Beyond EAR: A Stable Eyelid Angle Metric for Driver Drowsiness Detection and Data Augmentation
cs.CVMathis Wolter, Julie Stephany Berrio Perez, Mao Shan
Detecting driver drowsiness reliably is crucial for enhancing road safety and supporting advanced driver assistance systems (ADAS). We introduce the Eyelid Angle (ELA), a novel, reproducible metric of eye openness derived from 3D facial landmarks. Unlike conventional binary eye state estimators or 2D measures, such as the Eye Aspect Ratio (EAR), the ELA prov
Near-Field Sparse Bayesian Channel Estimation and Tracking for XL-IRS-Aided Wideband mmWave Systems
eess.SPXiaokun Tuo, Zijian Chen, Ming-Min Zhao, Changsheng You
The rapid development of 6G systems demands advanced technologies to boost network capacity and spectral efficiency, particularly in the context of intelligent reflecting surfaces (IRS)-aided millimeter-wave (mmWave) communications. A key challenge here is obtaining accurate channel state information (CSI), especially with extremely large IRS (XL-IRS), due t
Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion
cs.CLDaiqing Wu, Dongbao Yang, Yu Zhou, Can Ma
As posts on social media increase rapidly, analyzing the sentiments embedded in image-text pairs has become a popular research topic in recent years. Although existing works achieve impressive accomplishments in simultaneously harnessing image and text information, they lack the considerations of possible low-quality and missing modalities. In real-world app
Samya Praharaj, Koulik Khamaru
Statistical inference from data generated by multi-armed bandit (MAB) algorithms is challenging due to their adaptive, non-i.i.d. nature. A classical manifestation is that sample averages of arm rewards under bandit sampling may fail to satisfy a central limit theorem. Lai and Wei's stability condition provides a sufficient, and essentially necessary criteri
Large Language Models Require Curated Context for Reliable Political Fact-Checking -- Even with Reasoning and Web Search
cs.CLMatthew R. DeVerna, Kai-Cheng Yang, Harry Yaojun Yan, Filippo Menczer
Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. As mainstream chatbots increasingly ship with reasoning capabilities and web search tools -- and millions of users already rely on them for verification -- rigorous evaluation is urgent. We evaluate 15 recent LLMs from OpenAI, Google
Evaluation of Real-Time Mitigation Techniques for Cyber Security in IEC 61850 / IEC 62351 Substations
cs.CRAkila Herath, Chen-Ching Liu, Junho Hong, Kuchan Park
The digitalization of substations enlarges the cyber-attack surface, necessitating effective detection and mitigation of cyber attacks in digital substations. While machine learning-based intrusion detection has been widely explored, such methods have not demonstrated detection and mitigation within the required real-time budget. In contrast, cryptographic a
Alexander Clow, Hitesh Kumar, Shivaramakrishna Pragada
The ultimate independence ratio of a graph $G$ is defined as $\mathscr{I}(G) = \lim_{k\rightarrow\infty } \frac{\alpha(G^{\Box k})}{|V(G)|^k},$ where $\alpha(G^{\Box k})$ is the independence number of the Cartesian product of $k$ copies of $G$. For all graphs $G$, Hahn, Hell, and Poljak (1995) proved that $\frac{1}{\chi(G)} \leq \mathscr{I}(G) \leq \frac{1}{
Hao Li, Qiao Sun
While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-
Hierarchical Bayesian estimation of population-level torque law parameters from $68$ young radio pulsars observed with the Murriyang telescope
astro-ph.HEAndrés F. Vargas, Andrew Melatos, Julian B. Carlin, Marcus E. Lower
Abridged. The measured braking index, $n=\nu \ddot{\nu}/\dot{\nu}^2$, of a rotation-powered pulsar with spin frequency $\nu$ and braking torque $K \nu^{n_{\rm pl}}$, features secular and stochastic anomalies arising from $\dot{K} \neq 0$ and random torque noise respectively. Previous studies quantified the variance $\langle n^{2} \rangle = (n_{\rm pl}+\dot{K
RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
cs.CLYu Lei, Shuzheng Si, Wei Wang, Yifei Wu
Large language models are evolving from single-turn responders into tool-using agents capable of sustained reasoning and decision-making for deep research. Prevailing systems adopt a linear pipeline of plan to search to write to a report, which suffers from error accumulation and context rot due to the lack of explicit control over both model behavior and co
Zhenghan Fang, Jian Zheng, Qiaozi Gao, Xiaofeng Gao
Diffusion models have emerged as a dominant paradigm for generative modeling across a wide range of domains, including prompt-conditional generation. The vast majority of samplers, however, rely on forward discretization of the reverse diffusion process and use score functions that are learned from data. Such forward and explicit discretizations can be slow
On the Appropriateness of Linear Stress Recovery in Biomechanical Analysis of Abdominal Aortic Aneurysm
physics.med-phAlastair Catlin, Mostafa Jamshidian, Adam Wittek, Karol Miller
Abdominal aortic aneurysm (AAA) wall stress is a candidate rupture risk marker but is typically computed from single-phase images without known cardiac phase. Linear stress recovery methods, which solve a single geometrically linear equilibrium problem on the imaged, already-loaded geometry, have been validated for static stress estimation, but their robustn
Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
cs.IRYu Wang, Yonghui Yang, Le Wu, Yi Zhang
Recent advances in Large Language Models (LLMs) have opened new avenues for sequential recommendation by enabling natural language reasoning over user behavior sequences. A common approach formulates recommendation as a language modeling task, where interaction histories are transformed into prompts and user preferences are learned via supervised fine-tuning
Trust and Uncertainty in Strategic Interaction: Behavioural and Physiological Evidence from the Centipede Game
econ.GNDhiraj Jagadale, Kavita Vemuri
Mutual trust is a key determinant of decision-making in economic interactions, yet actual behavior often diverges from equilibrium predictions. This study investigates how emotional arousal, indexed by skin conductance responses,SCR, relates to trust behavior in a modified centipede game. To examine the impact of uncertainty, the game incorporated both fixed
Microscopic parameters of a type-II superconductor measured by small-angle neutron scattering
cond-mat.supr-conD. Alba Venero, A. -M. Valente-Feliciano, O. O. Bernal, V. Kozhevnikov
A necessary condition for understanding and predicting the properties of any material is knowledge of microscopic parameters which control these properties in a state of thermodynamic equilibrium. One can show (see, e.g., Ref.\,\cite{VK_book}), that in superconductors these parameters are the radius of the orbital motion of electrons bound in Cooper pairs $R
Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
cs.CVKeyang Lu, Sifan Zhou, Hongbin Xu, Gang Xu
Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we present Yo'City, a novel agentic framework that enables user-
Jeff Murugan
Information transmitted across modern communication platforms is degraded not only by intentional manipulation (disinformation) but also by intrinsic cognitive decay and topology-dependent social averaging (misinformation). We develop a continuous-fidelity field theory on multiplex networks with distinct layers representing private chats, group interactions,
Haoming Jia, Yi Han, Xiang Wang, Huizan Wang
Global ocean forecasting aims to predict key ocean variables such as temperature, salinity, and currents, which is essential for understanding and describing oceanic phenomena. In recent years, data-driven deep learning-based ocean forecast models, such as XiHe, WenHai, LangYa and AI-GOMS, have demonstrated significant potential in capturing complex ocean dy
Large-Scale In-Game Outcome Forecasting for Match, Team and Players in Football using an Axial Transformer Neural Network
cs.LGMichael Horton, Patrick Lucey
Football (soccer) is a sport that is characterised by complex game play, where players perform a variety of actions, such as passes, shots, tackles, fouls, in order to score goals, and ultimately win matches. Accurately forecasting the total number of each action that each player will complete during a match is desirable for a variety of applications, includ
Lin Liu, Caiyan Jia, Guanyi Yu, Ziying Song
Driving planning is a critical component of end-to-end (E2E) autonomous driving. However, prevailing Imitative E2E Planners often suffer from multimodal trajectory mode collapse, failing to produce diverse trajectory proposals. Meanwhile, Generative E2E Planners struggle to incorporate crucial safety and physical constraints directly into the generative proc
Maitreyi Chatterjee, Devansh Agarwal, Biplab Chatterjee
The transition to autonomous material systems necessitates adaptive control methodologies to maximize structural longevity. This study frames the self-healing process as a Reinforcement Learning (RL) problem within a Markov Decision Process (MDP), enabling agents to autonomously derive optimal policies that efficiently balance structural integrity maintenanc
LogSyn: A Few-Shot LLM Framework for Structured Insight Extraction from Unstructured General Aviation Maintenance Logs
cs.LGDevansh Agarwal, Maitreyi Chatterjee, Biplab Chatterjee
Aircraft maintenance logs hold valuable safety data but remain underused due to their unstructured text format. This paper introduces LogSyn, a framework that uses Large Language Models (LLMs) to convert these logs into structured, machine-readable data. Using few-shot in-context learning on 6,169 records, LogSyn performs Controlled Abstraction Generation (C
Kai-Feng Chen, Yi-An Chen, Cheng-Wei Chiang, Feng-Yang Hsieh
A reliable determination of the Higgs production mechanism in hadron collider experiments is essential in the program of the measurements of the Higgs couplings. We employ weak supervision, CWoLa in particular, to train deep neural networks using real data of the diphoton events, in the hope of reducing biases resulting from Monte Carlo simulations. Models b
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
cs.CVZhaoqi Xu, Yingying Zhang, Jian Li, Jianwei Guo
Recent advances in vision-language models (VLMs) have shown remarkable performance across multimodal tasks, yet their ever-growing scale poses severe challenges for deployment and efficiency. Existing compression methods often rely on heuristic importance metrics or empirical pruning rules, lacking theoretical guarantees about information preservation. In th
First Deep Learning Approach to Hammering Acoustics for Stem Stability Assessment in Total Hip Arthroplasty
eess.ASDongqi Zhu, Zhuwen Xu, Youyuan Chen, Minghao Jin
Audio event classification has recently emerged as a promising approach in medical applications. In total hip arthroplasty (THA), intra-operative hammering acoustics provide critical cues for assessing the initial stability of the femoral stem, yet variability due to femoral morphology, implant size, and surgical technique constrains conventional assessment
Neural B-Frame Coding: Tackling Domain Shift Issues with Lightweight Online Motion Resolution Adaptation
eess.IVSang NguyenQuang, Xiem HoangVan, Wen-Hsiao Peng
Learned B-frame codecs with hierarchical temporal prediction often encounter the domain-shift issue due to mismatches between the Group-of-Pictures (GOP) sizes for training and testing, leading to inaccurate motion estimates, particularly for large motion. A common solution is to turn large motion into small motion by downsampling video frames during motion
Longfei Wang, Junyan Liu, Fan Zhang, Jiangwen Wei
Parallelization has emerged as a promising approach for accelerating MILP solving. However, the complexity of the branch-and-bound (B&B) framework and the numerous effective algorithm components in MILP solvers make it difficult to parallelize. In this study, a scalable parallel framework, N2N (a node-to-node framework that maps the B&B nodes to distributed
Nicole F. Bell, Peter Cox, Jayden L. Newstead, Michael B. G. Verde
The neutron portal operator provides a theoretically motivated connection between the visible and dark sectors and features in several well-studied asymmetric dark matter models. This operator leads to dark matter induced nucleon decays that mimic the experimental signature of "ordinary" nucleon decays. In this work, we reinterpret Super-Kamiokande nucleon d
Adarsh Kumarappan, Ayushi Mehrotra
The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong assumption can limit the trustworthiness of the provided safety certificate. In this work, we address this limitation by introducing a more realistic probabilistic framewor
Toward Integrated Air-Ground Computing and Communications: A Synergy of Computing Power Networks and Low-Altitude Economy Network
cs.NIYan Sun, Yinqiu Liu, Shaoyong Guo, Ruichen Zhang
With the rapid rise of the Low-Altitude Economy (LAE), the demand for intelligent processing and real-time response in services such as aerial traffic, emergency communications, and environmental monitoring continues to grow. Meanwhile, the Computing Power Network (CPN) aims to integrate global computing resources and perform on-demand scheduling to efficien
Zeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu
We present LATTICE, a new framework for high-fidelity 3D asset generation that bridges the quality and scalability gap between 3D and 2D generative models. While 2D image synthesis benefits from fixed spatial grids and well-established transformer architectures, 3D generation remains fundamentally more challenging due to the need to predict both spatial stru
Omar Garib, Jayaprakash D. Kambhampaty, Olivia J. Pinon Fischer, Dimitri N. Mavris
We introduce AIRHILT (Aviation Integrated Reasoning, Human-in-the-Loop Testbed), a modular and lightweight simulation environment designed to evaluate multimodal pilot and air traffic control (ATC) assistance systems for aviation conflict detection. Built on the open-source Godot engine, AIRHILT synchronizes pilot and ATC radio communications, visual scene u
When and What to Recommend: Joint Modeling of Timing and Content for Active Sequential Recommendation
cs.IRJin Chai, Xiaoxiao Ma, Jian Yang, Jia Wu
Sequential recommendation models user preferences to predict the next target item. Most existing work is passive, where the system responds only when users open the application, missing chances after closure. We investigate active recommendation, which predicts the next interaction time and actively delivers items. Two challenges: accurately estimating the T
Adarsh Kumarappan, Ananya Mujoo
Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persistent threat to Large Language Models (LLMs). Progress in defending against these attacks is hindered by a reliance on manual, hard-to-scale d
GRIT-LP: Graph Transformer with Long-Range Skip Connection and Partitioned Spatial Graphs for Accurate Ice Layer Thickness Prediction
cs.LGZesheng Liu, Maryam Rahnemoonfar
Graph transformers have demonstrated remarkable capability on complex spatio-temporal tasks, yet their depth is often limited by oversmoothing and weak long-range dependency modeling. To address these challenges, we introduce GRIT-LP, a graph transformer explicitly designed for polar ice-layer thickness estimation from polar radar imagery. Accurately estimat
Shaoyin Ma, Chenggong Hu, Huiqiong Wang, Li Sun
Building effective LLM agents increasingly requires selecting appropriate AI models as tools from large open repositories (e.g., HuggingFace with > 2M models) based on natural language requests. Unlike invoking a fixed set of API tools, repository-scale model selection must handle massive, evolving candidates with incomplete metadata. Existing approaches inc
MAGMA-Edu: Multi-Agent Generative Multimodal Framework for Text-Diagram Educational Question Generation
cs.AIZhenyu Wu, Jian Li, Hua Huang
Educational illustrations play a central role in communicating abstract concepts, yet current multimodal large language models (MLLMs) remain limited in producing pedagogically coherent and semantically consistent educational visuals. We introduce MAGMA-Edu, a self-reflective multi-agent framework that unifies textual reasoning and diagrammatic synthesis for
Hongbin Lin, Yiming Yang, Chaoda Zheng, Yifan Zhang
In autonomous driving, vision-centric 3D object detection recognizes and localizes 3D objects from RGB images. However, due to high annotation costs and diverse outdoor scenes, training data often fails to cover all possible test scenarios, known as the out-of-distribution (OOD) issue. Training-free image editing offers a promising solution for improving mod
Liqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai
Visual grounding, the task of linking textual queries to specific regions within images, plays a pivotal role in vision-language integration. Existing methods typically rely on extensive task-specific annotations and fine-tuning, limiting their ability to generalize effectively to novel or out-of-distribution scenarios. To address these limitations, we intro
Tianyu Wang, Chunxiang Yan, Xuanhong Liao, Tao Zhang
Wheeled bipedal robots are emerging as flexible platforms for field exploration. However, head instability induced by uneven terrain can degrade the accuracy of onboard sensors or damage fragile payloads. Existing research primarily focuses on stabilizing the mobile platform but overlooks active stabilization of the head in the world frame, resulting in vert
Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
cs.CVShawn Young, Xingyu Zeng, Lijian Xu
This paper investigates the fundamental relationship between model capacity and the minimal number of visual tokens required to preserve image semantics. Inspired by the Minimum Description Length principle, we reinterpret image tokens as vectors in a visual semantic space and define the intrinsic semantic complexity of an image as the smallest set of basis
Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu
In this paper, we study the challenging task of Few-Shot Video Domain Adaptation (FSVDA). The multimodal nature of videos introduces unique challenges, necessitating the simultaneous consideration of both domain alignment and modality collaboration in a few-shot scenario, which is ignored in previous literature. We observe that, under the influence of domain
Autonomous Surface Selection For Manipulator-Based UV Disinfection In Hospitals Using Foundation Models
cs.ROXueyan Oh, Jonathan Her, Zhixiang Ong, Brandon Koh
Ultraviolet (UV) germicidal radiation is an established non-contact method for surface disinfection in medical environments. Traditional approaches require substantial human intervention to define disinfection areas, complicating automation, while deep learning-based methods often need extensive fine-tuning and large datasets, which can be impractical for la
Yanbin Li, Canran Xiao, Shenghai Yuan, Peilai Yu
Topological maps are more suitable than metric maps for robotic exploration tasks. However, real-time updating of accurate and detail-rich environmental topological maps remains a challenge. This paper presents a topological map updating method based on the Generalized Voronoi Diagram (GVD). First, the newly observed areas are denoised to avoid low-efficienc
Development of a projectile charge state analyzer and 10 kV bipolar power supply for MeV energy ion - atom/molecule collision experiments
physics.atom-phSandeep Bajrangi Bari, Sahan Raghava Sykam, Ranojit Das, Rohit Tyagi
We have developed a post-collision projectile charge state analyzer (CSA) for detecting the charge state of the projectile ion following ion-atom/molecule collision. The design of the analyzer, based on electrostatic parallel plate deflector was simulated using SIMION ion optics package. We have also developed a 10 kV bipolar programmable power supply to bia
Zhaoyang Jia, Zihan Zheng, Naifu Xue, Jiahao Li
Existing diffusion codecs typically build on text-to-image diffusion foundation models like Stable Diffusion. However, text conditioning is suboptimal from a compression perspective, hindering the potential of downstream diffusion codecs, particularly at ultra-low bitrates. To address it, we introduce \textbf{CoD}, the first \textbf{Co}mpression-oriented \te
Jie Jiang, Yang Wu, Qian Li, Yuling Xiong
Harnessing the reasoning power of Large Language Models (LLMs) for recommender systems is hindered by two fundamental challenges. First, current approaches lack a mechanism for automated, data-driven discovery of effective reasoning patterns, relying instead on brittle manual templates or unstable zero-shot prompting. Second, they employ structure-collapsing
Hai-Xiao Wang, Li Liang, Shuai Shao, Shiwei Tang
Crystalline symmetry offers a powerful tool to realize photonic topological phases, in which additional trivial claddings are typically required to confine topological boundary states. However, the utility of the trivial cladding in manipulating topological waves is often overlooked. Here, we demonstrate two topologically distinct kagome photonic crystals (K
Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
cs.CLSameeah Noreen Hameed, Surangika Ranathunga, Raj Prasanna, Kristin Stock
Large-scale disasters can often result in catastrophic consequences on people and infrastructure. Situation awareness about such disaster impacts generated by authoritative data from in-situ sensors, remote sensing imagery, and/or geographic data is often limited due to atmospheric opacity, satellite revisits, and time limitations. This often results in geo-
Luan C. V. Silva, Lívia Sancho, Mauricio S. Silva, Elisa Passos
The Southwestern South Atlantic (SWSA) is a key region for climate research and renewable energy assessment, yet high-resolution meteorological data are scarce. We present a multiresolution dataset spanning February 2017--November 2018, combining Weather Research and Forecasting (WRF) simulations with Sentinel-1A/B Synthetic Aperture Radar (SAR) wind fields
Ardalan Tajbakhsh, Augustinos Saravanos, James Zhu, Evangelos A. Theodorou
This paper addresses the challenge of coordinating multi-robot systems under realistic communication delays using distributed optimization. We focus on consensus ADMM as a scalable framework for generating collision-free, dynamically feasible motion plans in both trajectory optimization and receding-horizon control settings. In practice, however, these algor
CNN-Based Camera Pose Estimation and Localisation of Scan Images for Aircraft Visual Inspection
cs.ROXueyan Oh, Leonard Loh, Shaohui Foong, Zhong Bao Andy Koh
General Visual Inspection is a manual inspection process regularly used to detect and localise obvious damage on the exterior of commercial aircraft. There has been increasing demand to perform this process at the boarding gate to minimise the downtime of the aircraft and automating this process is desired to reduce the reliance on human labour. Automating t
Lizhe Hong
We learned the atomic deposition simulation of LAMMPS independently, referenced and optimized the modeling ideas of several papers, used the (1 1 1) crystalline surface of Cu atoms as a substrate, deposited C atoms produced by methane cleavage to obtain graphene flakes, and analyzed the deposition rate and deposition quality at three temperatures, obtaining
Mustafa Munir, Harsh Goel, Xiwen Wei, Minkyu Choi
Video editing and synthesis often introduce object inconsistencies, such as frame flicker and identity drift that degrade perceptual quality. To address these issues, we introduce ObjectAlign, a novel framework that seamlessly blends perceptual metrics with symbolic reasoning to detect, verify, and correct object-level and temporal inconsistencies in edited
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
cs.MMSiran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang
Existing video recommendation systems, relying mainly on ID-based embedding mapping and collaborative filtering, often fail to capture in-depth video content semantics. Moreover, most struggle to address biased user behaviors (e.g., accidental clicks, fast skips), leading to inaccurate interest modeling and frequent negative feedback in top recommendations w
Jiarui Xue, Dongjian Yang, Ye Sun, Gang Liu
In real-world scenarios of image recognition, there exists substantial noise interference. Existing works primarily focus on methods such as adjusting networks or training strategies to address noisy image recognition, and the anti-noise performance has reached a bottleneck. However, little is known about the exploration of anti-interference solutions from a
Aman Verma, Keshav Samdani, Mohd. Samiuddin Shafi
This paper presents the design, implementation, and evolution of a comprehensive multimodal room-monitoring system that integrates synchronized video and audio processing for real-time activity recognition and anomaly detection. We describe two iterations of the system: an initial lightweight implementation using YOLOv8, ByteTrack, and the Audio Spectrogram
Hongsheng Pang, Lixin He
In crystalline solids, the electronic polarization follows the \emph{generalized Neumann's principle}, under which all crystallographic point groups can, in principle, support ferroelectric polarization. However, in high-symmetry structures, polarization is constrained by symmetry operations and becomes quantized into discrete values. We demonstrate that thi
Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Models
cs.CLWangjiaxuan Xin
This report presents the Empathetic Cascading Networks (ECN) framework, a multi-stage prompting method designed to enhance the empathetic and inclusive capabilities of large language models. ECN employs four stages: Perspective Adoption, Emotional Resonance, Reflective Understanding, and Integrative Synthesis, to guide models toward generating emotionally re
Huanning Dong, Yinuo Huang, Fan Li, Ping Kuang
3D assets are essential in the digital age. While automatic 3D generation, such as image-to-3d, has made significant strides in recent years, it often struggles to achieve fast, detailed, and high-fidelity generation simultaneously. In this work, we introduce LatentDreamer, a novel framework for generating 3D objects from single images. The key to our approa
Changcai Li, Wenwei Lin, Zuoxun Hou, Gang Chen
In this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic pinhole-based 3D object detectors to fisheye imagery. To mitigate this, we then develop two methods that incorporate the u