February 2026 arXiv papers — page 4
Showing 301–400 of 20,995 papers
Jingyuan Xie, Wenjie Wang, Ji Wu, Jiandong Gao
Supervised fine-tuning (SFT) is essential for the development of medical large language models (LLMs), yet prior poisoning studies have mainly focused on the detectable backdoor attacks. We propose a novel poisoning attack targeting the reasoning process of medical LLMs during SFT. Unlike backdoor attacks, our method injects poisoned rationales into few-shot
Xingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun
Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM-4D, a compre
Enhanced Hydrogen Electrolyzer with Integrated Energy Storage to Provide Grid-Forming Services for Off-Grid ReP2H Application
math.OCYiwei Qiu, Jiahao Hu, Yi Zhou, Jie Zhu
This article proposes an energy storage-enhanced hydrogen electrolyzer (ESEHE) to provide grid-forming (GFM) services for off-grid renewable power to hydrogen (ReP2H) systems. Unlike conventional ReP2H systems that use a centralized energy storage (ES) plant, the proposed topology directly connects batteries to the DC buses of electrolysis rectifiers. A tail
Ram Milan Kumar Verma, Shashi Ranjan Kumar, Hemendra Arya
Precise motion control of underactuated surface vessels is a crucial task in various maritime applications. In this work, we develop a nonlinear motion control strategy for surface vessels inspired by the pursuit guidance philosophy. Any sufficiently smooth path can be seen as a continuum of virtual targets moving along a specified path, which the pursuer is
Wang Chen, Yuhui Zeng, Yongdong Luo, Tianyu Xie
Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting in a disjointed set of frames that disregard the narrative structure of video. In this paper, we introduce Wavelet-based
Ruoshuang Du, Xin Sun, Qiang Liu, Bowen Song
Visual Question Answering systems face reliability issues due to hallucinations, where models generate answers misaligned with visual input or factual knowledge. While Retrieval Augmented Generation frameworks mitigate this issue by incorporating external knowledge, static retrieval often introduces irrelevant or conflicting content, particularly in visual R
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
cs.CVYingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen
Multimodal large language models (MLLMs) project visual tokens into the embedding space of language models, yet the internal structuring and processing of visual semantics remain poorly understood. In this work, we introduce a two-fold analytical framework featuring a novel probing tool, $\textbf{EmbedLens}$, to conduct a fine-grained analysis. We uncover a
Ce Zhang, Cheng Xu, Haibo Hu, Jianliang Xu
Blockchain provides a decentralized and tamper-resistant ledger for securely recording transactions across a network of untrusted nodes. While its transparency and integrity are beneficial, the substantial storage requirements for maintaining a complete transaction history present significant challenges. For example, Ethereum nodes require around 23TB of sto
A Novel Monte Carlo Gradient Method Based on Meta-learning for Effective Step-size Selection in Active Noise Control
eess.SPLuyuan Li, Jisheng Bai, Xiruo Su, Xiaoyi Shen
Active noise control (ANC) is an effective approach to noise suppression, and the filtered-reference least mean square (FxLMS) algorithm is a widely adopted method in ANC systems, owing to its computational efficiency and stable performance. However, its convergence speed and noise reduction performance are highly dependent on the step size parameter. Common
Jiamin Shi, Haolin Zhang, Yuchen Yan, Shitao Chen
Navigating social robots in dense, dynamic crowds is challenging due to environmental uncertainty and complex human-robot interactions. While Model Predictive Control (MPC) offers strong real-time performance, its reliance on a fixed prediction horizon limits adaptability to changing environments and social dynamics. Furthermore, most MPC approaches treat pe
Hongquan Wang, Hanshu Chen, Ilia Marchevsky, Zhuojia Fu
DeepONet enables retraining-free inference across varying initial conditions or source terms at the cost of high computational requirements. This paper proposes a hybrid quantum operator network (Quantum AS-DeepOnet) suitable for solving 2D evolution equations. By combining Parameterized Quantum Circuits and cross-subnet attention methods, we can solve 2D ev
Kautuk Astu, Suman Raj, Priyanshu Pansari, Yogesh Simmhan
The increasing adoption of UAVs equipped with advanced sensors and GPU-accelerated edge computing has enabled real-time AI-driven applications in domains such as precision agriculture, wildfire monitoring, and environmental conservation. However, the integrated design and orchestration of navigation, sensing, and analytics, together with seamless real-time c
Keunho Byeon, Jinsol Song, Seong Min Hong, Yosep Chong
Whole-slide image analysis is essential for diagnostic tasks in pathology, yet existing deep learning methods primarily rely on flat classification, ignoring hierarchical relationships among class labels. In this study, we propose HiClass, a hierarchical classification framework for improved histopathology image analysis, that enhances both coarse-grained an
M$^2$: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight Retrieval
cs.CVDawei Yan, Haokui Zhang, Guangda Huzhang, Yang Li
Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on extensive data collection and model training, yet still struggle with high computational costs and insufficient reasoning cap
Wenhao Zheng, Wang Lu, Fangshuang Tang, Yiyang Lu
Early-stage users in a new scenario intensify cold-start challenges, yet prior works often address only parts of the problem through model architecture. Launching a new user experience to replace an established product involves sparse behavioral signals, low-engagement cohorts, and unstable model performance. We argue that effective recommendations require t
Jingwen Tong, Zijian Li, Fang Liu, Wei Guo
The integration of large language models (LLMs) into wireless networks has sparked growing interest in building autonomous AI agents for wireless tasks. However, existing approaches rely heavily on manually crafted prompts and static agentic workflows, a process that is labor-intensive, unscalable, and often suboptimal. In this paper, we propose WirelessAgen
Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation
cs.ROZilong Xie, Jingyu Gong, Xin Tan, Zhizhong Zhang
Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense reasoning, they struggle with geometric and spatial understanding required for pose prediction. In this paper, we propos
Zhang-nan Hu, Bing Li, YiJing Wang
In this paper, we study the uniform random covering problem in general metric space $(X,d)$. Let $\omega=(\omega_n)_{n\in\mathbb N}$ be a sequence of independent identically distributed random variables on $(X,\mu)$, and $\ell=(\ell_n)_{n\in\mathbb N}$ a sequence of positive real numbers. We analyze the size of the set \[\mathcal{U}(\omega,\ell)=\left\{y\in
Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
cs.LGQuoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui
Fine-tuning-as-a-service introduces a threat to Large Language Models' safety when service providers fine-tune their models on poisoned user-submitted datasets, a process known as harmful fine-tuning attacks. In this work, we show that by regularizing the gradient contribution of harmful samples encountered during fine-tuning, we can effectively mitigate the
Parameterizations of the Hubble Constant: Logarithmic vs Power-Law Expansion from the Binned Master Sample of SNe Ia
astro-ph.COMaria Giovanna Dainotti, Avik Banerjee, Andre' LeClair, Giovanni Montani
In view of the current and increasing evidence of a running Hubble constant, we investigate its redshift dependence within the flat $\Lambda$CDM framework using a 20-bin analysis of the Master SNe~Ia Sample \citep{2025JHEAp..4800405D}, considering cases with and without very low-redshift data. For each case, we obtain best-fitting values of $H_0$ and $\Omega
Kazuhiro Hiraki, Shinichi Ishihara, Takumi Kongo, Junnosuke Shino
In this paper, we provide a theoretically grounded and computationally efficient alternative to SHAP. To this end, we study feature attribution through the lens of cooperative game theory by formulating a class of XAI--TU games. Building on this formulation, we investigate equal-surplus-type and proportional-allocation-type attribution rules and propose a lo
Christopher Cruz
We introduce AI Runtime Infrastructure, a distinct execution-time layer that operates above the model and below the application, actively observing, reasoning over, and intervening in agent behavior to optimize task success, latency, token efficiency, reliability, and safety while the agent is running. Unlike model-level optimizations or passive logging syst
Jiahao Zeng, Zhenkui Shi, Chunpei Li, Mengkai Yan
With the advancement of vehicle-to-vehicle (V2V) ad hoc networks and wireless communication technologies, mobile edge caching has become a key enabler for enhancing network performance and user experience. However, traditional federated learning-based collaborative caching approaches in vehicular scenarios suffer from inadequate client selection mechanisms a
Emmanuel Pintelas, Ioannis E. Livieris
Excessive reliance on validation performance during model selection can lead to validation overfitting (VO), where models appear effective during development but fail at test time. This issue is further amplified in low-data regimes and under distribution shifts, where validation signals become unreliable. Although ensemble learning is widely used to improve
Van der Waals Antiferromagnets: From Early Discoveries to Future Directions in the 2D Limit
cond-mat.mtrl-sciRahul Kumar, Je-Geun Park
The emergence of a long-range magnetic order in the atomically thin, two-dimensional (2D) limit has long remained a fundamental question in condensed matter physics. The advent of exfoliable van der Waals (vdW) materials, particularly transition-metal phosphorus trisulfides (T MPS3; T M = Fe, Ni, and Mn), provided the first experimental access to this regime
COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
cs.CVYuchen Che, Jingtu Wu, Hao Zheng, Asako Kanezaki
Estimating the 6DoF pose of a novel object with a single reference view is challenging due to occlusions, view-point changes, and outliers. A core difficulty lies in finding robust cross-view correspondences, as existing methods often rely on discrete one-to-one matching that is non-differentiable and tends to collapse onto sparse key-points. We propose Conf
Riccardo de Lutio, Tobias Fischer, Yen-Yu Chang, Yuxuan Zhang
Per-scene optimization methods such as 3D Gaussian Splatting provide state-of-the-art novel view synthesis quality but extrapolate poorly to under-observed areas. Methods that leverage generative priors to correct artifacts in these areas hold promise but currently suffer from two shortcomings. The first is scalability, as existing methods use image diffusio
Xianchao Xiu, Shenghao Sun, Xinrong Li, Jiyuan Tao
Support matrix machine (SMM) is an emerging classification framework that directly handles matrix-structured observations, thereby avoiding the spatial correlations destroyed by vectorization. However, most existing SMM variants rely on convex or nonconvex surrogate loss functions, which may lead to high sensitivity to noise. To address this issue, we propos
Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen
The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely underexplored. Existing video benchmarks predominantly assess pas
Haoyu Gao, Hong Yi Lin, Christoph Treude, Gregory Gay
README files are critical for understanding and onboarding contributors to open-source software, yet they frequently become outdated. We formulate surgical documentation update recommendation as a task and present a Large Language Model-driven framework for use in a human-in-the-loop workflow. Given a pull request, the framework determines whether a README u
Dynamic Spatio-Temporal Graph Neural Network for Early Detection of Pornography Addiction in Adolescents Based on Electroencephalogram Signals
cs.LGAchmad Ardani Prasha, Clavino Ourizqi Rachmadi, Sabrina Laila Mutiara, Hilman Syachr Ramadhan
Adolescent pornography addiction requires early detection based on objective neurobiological biomarkers because self-report is prone to subjective bias due to social stigma. Conventional machine learning has not been able to model dynamic functional connectivity of the brain that fluctuates temporally during addictive stimulus exposure. This study proposes a
DRAFT: A Formally Verified Constructive Proof of the Consistency of Peano Arithmetic Using Ordinal Assignments
cs.LOAaron Bryce, Rajeev Gore'
Gentzen's 1936 proof of the consistency of Peano Arithmetic was a significant result in the foundations of mathematics. We provide here a modified version of the proof, based on G\"{o}del's reformulation, and including additional details and minor corrections which are necessary to definitively prove the well-foundedness of the cut-elimination argument in a
Qihang Fan, Yuang Ai, Huaibo Huang, Ran He
Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representative approach involves grouping tokens, performing self-attention calculations within each group, or pooling the tokens within each group into a single token. To this end, various c
"I Need to Find That One Chart": How Data Workers Navigate, Make Sense of, and Communicate Analytical Conversations
cs.HCKen Gu, Srishti Palani, Vidya Setlur
Conversational interfaces are increasingly used for data analysis, enabling data workers to express complex analytical intents in natural language. Yet, these interactions unfold as long, linear transcripts that are misaligned with the iterative, nonlinear nature of real-world analyses. Revisiting and summarizing conversations for different contexts is there
MergeDJD: A Fast Constructive Algorithm with Piece Merging for the Two-Dimensional Irregular Bin Packing Problem
cs.CGYi Zhou, Haocheng Fu, Yiping Liu, Jian Mao
The two-dimensional irregular bin packing problem (2DIBPP) aims to pack a given set of irregular polygons, referred to as pieces, into fixed-size rectangular bins without overlap, while maximizing bin utilization. Although numerous metaheuristic algorithms have been proposed for the 2DIBPP, many industrial applications favor simpler constructive heuristics d
Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu
Recent text-to-image (T2I) diffusion models achieve remarkable realism, yet faithful prompt-image alignment remains challenging, particularly for complex prompts with multiple objects, relations, and fine-grained attributes. Existing training-free inference-time scaling methods rely on fixed iteration budgets that cannot adapt to prompt difficulty, while ref
Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li
Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granularity, overlong visual token sequences, and inadequate cross-modal alignment. To overcome these challenges, we propose TaiChi, a novel VLM fr
Khaleque Md Aashiq Kamal, Surya Eada, Aayushi Verma, Subek Acharya
Developments in the machine learning voting domain have shown both promising results and risks. Trained models perform well on ballot classification tasks (> 99% accuracy) but are at risk from adversarial example attacks that cause misclassifications. In this paper, we analyze an attacker who seeks to deploy adversarial examples against machine learning ball
A. E. Petrova, S. Yu. Gavrilkin, D. Menzel, V. A. Stepanov
We studied the longitudinal and transverse magnetoresistance of helical magnets, MnSi and Mn$_{1-x}$Co$_x$Si, at temperatures between 1.8 and 100~K and in magnetic fields up to 9 Tesla. All substances exhibited negative longitudinal and transverse magnetoresistance at temperatures above 4~K, which is most likely related to the suppression of spin fluctuation
Pengcheng Shi, Minghui Zhang, Kehan Song, Jiaqi Liu
Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision-language models face two limitations: they do not leverage segmentation-pretrained encoders, and they inject visual features only at the input layer of
Xu Luo, Ji Zhang, Lianli Gao, Heng Tao Shen
Few-shot transfer has been revolutionized by stronger pre-trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is both challenging and realistic for real-world usage. In this work, we establish FEWTRANS, a comprehensive benchmark containing 10 diverse datasets, and propose the Hyperparameter Ense
Qiqi Gu, Chenpeng Wu, Heng Shi, Jianguo Yao
Stencil computation constitutes a cornerstone of scientific computing, serving as a critical kernel in domains ranging from fluid dynamics to weather simulation. While stencil computations are conventionally regarded as memory-bound and thus unsuitable for compute-centric Tensor Cores, recent empirical studies have demonstrated significant speedups after app
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
cs.CRLinxi Jiang, Zhijie Liu, Haotian Luo, Zhiqiang Lin
Browser-use agents are widely used for everyday tasks. They enable automated interaction with web pages through structured DOM based interfaces or vision language models operating on page screenshots. However, web pages often change between planning and execution, causing agents to execute actions based on stale assumptions. We view this temporal mismatch as
Harsha Desu, Niladri S. Satpathi, Lokesh Malik, Ashis K. Sen
Capillary-driven transport offers a simple, self-sustained alternative to externally pumped microfluidic systems, yet achieving precise control of such flows remains challenging. We experimentally and theoretically investigate capillary flow in rectangular microchannels when a liquid meniscus encounters a geometric step with varying channel width and height.
Jiacheng Wang, Yucheng Sheng, Le Liang, Hao Ye
This paper investigates the power control problem in wireless networks by repurposing pre-trained large language models (LLMs) as relational reasoning backbones. In hyper-connected interference environments, traditional optimization methods face high computational cost, while standard message passing neural networks suffer from aggregation bottlenecks that c
Detection and classification of astronomical sources with Astro-RetinaNet in crowded stellar fields
astro-ph.IMYibo Yan, Chao Liu, Jiadong Li, Feng Wang
Upcoming next-generation sky surveys will detect large number of faint objects with magnitudes larger than 25. When objects are crowded within a limited a field of view, blending becomes unavoidable. Blending leads to the omission of many sources during photometry in these fields, which cause an underestimates of tens of percent in crowded fields, and remain
Yijun Yu
Agentic AI systems exhibit numerous crosscutting concerns -- security, observability, cost management, fault tolerance -- that are poorly modularized in current implementations, contributing to the high failure rate of AI projects in reaching production. The goals-to-aspects methodology proposed at RE 2004 demonstrated that aspects can be systematically disc
Observation of Chiral Bound States in the Continuum in Self-biased Magneto-optical Photonic Crystals
physics.opticsMaohua Gong, Qiutong Zhen, Yujie Tang, Peng Hu
Chiral bound states in the continuum (BICs) are confined photonic modes with infinite quality factors and chiral response, offering significant potential for chiral optics. Although a novel type of spin-orbital-locking chiral BIC was recently predicted in magneto-optical (MO) photonic crystals (PhCs) that break time-reversal symmetry (TRS), its experimental
Seamless Handover in Direct-to-Device Satellite Networks: From an Interference-Aware Perspective
cs.NIYe Li, Yi Zhou, Yingdong Hu, Zhe Ji
The direct-to-device (D2D) satellite network is an important 6G evolution direction to enable seamless ubiquitous connectivity. However, the network faces critical handover challenges due to high satellite mobility and wide beam footprints. Conventional handover strategies, mostly designed for terrestrial networks, may encounter excessive co-channel interfer
Najeeb Khan
Operators of Earth observation satellites need justifications for scheduling decisions: why a request was selected, rejected, or what changes would make it schedulable. Existing approaches construct post-hoc reasoning layers independent of the optimizer, risking non-causal attributions, incomplete constraint conjunctions, and solver-path dependence. We take
Ethan Chu, Yiyang Guo, Jan Hoffmann
There exist many techniques for automatically deriving parametric resource (or cost) bounds by analyzing the source code of a program. These techniques work effectively for a large class of programs and language features. However, non-local transfer of control as needed for exception or effect handlers has remained a challenge. This paper presents the first
Multi-Agent Reinforcement Learning for Dynamic Pricing: Balancing Profitability,Stability and Fairness
cs.LGKrishna Kumar Neelakanta Pillai Santha Kumari Amma
Dynamic pricing in competitive retail markets requires strategies that adapt to fluctuating demand and competitor behavior. In this work, we present a systematic empirical evaluation of multi-agent reinforcement learning (MARL) approaches-specifically MAPPO and MADDPG-for dynamic price optimization under competition. Using a simulated marketplace environment
Yilun Wang, Guangba Yu, Haiyu Huang, Yujie Huang
LLM agents are increasingly explored for automating root cause analysis (RCA) in cloud-native systems, creating a need to evaluate both diagnostic correctness and the quality of the supporting investigation. Existing static benchmarks offer repeatable inputs but limited system-facing interaction, while live testbeds expose realistic tools but hinder controll
Pengju Sun, Banglei Guan, Jing Tao, Zhenbao Yu
High dynamic range (HDR) imaging under extreme illumination remains challenging for conventional cameras due to overexposure. Event cameras provide microsecond temporal resolution and high dynamic range, while spatially varying exposure (SVE) sensors offer single-shot radiometric diversity.We present a hardware--algorithm co-designed HDR imaging system that
Boming Tan, Xiangdong Zhang, Ning Liao, Yuqing Zhang
Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world-related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single wor
Yucheng Chu, Hang Li, Kaiqi Yang, Yasemin Copur-Gencturk
Automated assessment of open-ended student responses is a critical capability for scaling personalized feedback in education. While large language models (LLMs) have shown promise in grading tasks via in-context learning (ICL), their reliability is heavily dependent on the selection of few-shot exemplars and the construction of high-quality rationales. Stand
TAS-GNN: A Status-Aware Signed Graph Neural Network for Anomaly Detection in Bitcoin Trust Systems
cs.CRChang Xue, Fang Liu, Jiaye Wang, Jinming Xing
Decentralized financial platforms rely heavily on Web of Trust reputation systems to mitigate counterparty risk in the absence of centralized identity verification. However, these pseudonymous networks are inherently vulnerable to adversarial behaviors, such as Sybil attacks and camouflaged fraud, where malicious actors cultivate artificial reputations befor
Zhaolin Yu, Litao Yang, Ben Babicka, Ming Hu
Orthopantomograms (OPGs) are the standard panoramic radiograph in dentistry, used for full-arch screening across multiple diagnostic tasks. While Vision Language Models (VLMs) now allow multi-task OPG analysis through natural language, they underperform task-specific models on most individual tasks. Agentic systems that orchestrate specialized tools offer a
Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu
The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, which significantly increases KV cache memory usage and time-to-
Yucheng Chu, Haoyu Han, Shen Dong, Hang Li
Automated short answer grading (ASAG) is critical for scaling educational assessment, yet large language models (LLMs) often struggle with hallucinations and strict rubric adherence due to their reliance on generalized pre-training. While Rretrieval-Augmented Generation (RAG) mitigates these issues, standard "flat" vector retrieval mechanisms treat knowledge
Shuheng Chen, Namratha Patil, Haonan Pan, Angel Hsing-Chi Hwang
Clinical decision-making requires synthesizing heterogeneous evidence, including patient histories, clinical guidelines, and trajectories of comparable cases. While large language models (LLMs) offer strong reasoning capabilities, they remain prone to hallucinations and struggle to integrate long, structured medical documents. We present MED-COPILOT, an inte
Explainable Continuous-Time Mask Refinement with Local Self-Similarity Priors for Medical Image Segmentation
cs.CVRajdeep Chatterjee, Sudip Chakrabarty, Trishaani Acharjee
Accurate semantic segmentation of foot ulcers is essential for automated wound monitoring, yet boundary delineation remains challenging due to tissue heterogeneity and poor contrast with surrounding skin. To overcome the limitations of standard intensity-based networks, we present LSS-LTCNet:an ante-hoc explainable framework synergizing deterministic structu
Bin Chen, Weiqi Li, Shijie Zhao, Xuanyu Zhang
While many diffusion models have achieved impressive results in real-world video super-resolution (Real-VSR) by generating rich and realistic details, their reliance on multi-step sampling leads to slow inference. One-step networks like SeedVR2, DOVE, and DLoRAL alleviate this through condensing generation into one single step, yet they remain heavy, with bi
Hostility prevents the tragedy of the commons in metapopulation with asymmetric migration: A lesson from queenless ants
q-bio.PEJoy Das Bairagya, Sagar Chakraborty
A colony of the queenless ant species, \emph{Pristomyrmex punctatus}, can broadly be seen as consisting of small-body sized worker ants and relatively larger body-sized cheater ants. Hence, in the presence of inter-colony migration, a set of constituent colonies act as a metapopulation exclusively composed of cooperators and defectors. Such a set-up facilita
Yao Wu, Ziye Jia, Jingjing Zhao, Haoyang Wang
Unmanned aerial vehicle (UAV) networks are increasingly deployed for complex missions, including disaster response, intelligent logistics, and environmental monitoring. These missions generally require coordinated collaboration among multiple UAVs across distinct administrative domains. To support such cross-domain cooperation, service function chains (SFCs)
Shivanshu Tripathi, Reza Akbarian Bafghi, Maziar Raissi
In this work, we present a test-driven, agentic framework for synthesizing a deployable low-level robot controller for navigation tasks. Given a 2D map with an image of an ultrasonic sensor-based robot, or a 3D robotic simulation environment, our framework iteratively refines the generated controller code using diagnostic feedback from structured test suites
Quhura Fathima, Neda Moghim, Mostafa Taghizade Firouzjaee, Christo K. Thomas
The growing deployment of Internet of Things (IoT) devices in smart cities and industrial environments increases vulnerability to stealthy, multi-stage advanced persistent threats (APTs) that exploit wireless communication. Detection is challenging due to severe class imbalance in network traffic, which limits the effectiveness of traditional deep learning a
Jocelyn Shen, Nicolai Marquardt, Hugo Romat, Ken Hinckley
What if text could be sculpted and refined like clay -- or cultivated and pruned like a plant? Texterial reimagines text as a material that users can grow, sculpt, and transform. Current generative-AI models enable rich text operations, yet rigid, linear interfaces often mask such capabilities. We explore how the text-as-material metaphor can reveal AI-enabl
Yucheng Chu, Hang Li, Kaiqi Yang, Yasemin Copur-Gencturk
Accurate and unambiguous guidelines are critical for large language model (LLM) based graders, yet manually crafting these prompts is often sub-optimal as LLMs can misinterpret expert guidelines or lack necessary domain specificity. Consequently, the field has moved toward automated prompt optimization to refine grading guidelines without the burden of manua
Sam Deen, Derek Lam
2024 YR4 is a 40-100 meter-diameter asteroid and former Torino Scale 3 object which currently has a roughly 4% chance of impacting the Moon on 2032 December 22, an event which recent studies suggest could pose a hazard on Earth due to impact ejecta. We present a search for, and identification of, potential precovery observations of the virtual lunar impactor
Phokion G. Kolaitis
The semijoin operation is a fundamental operation of relational algebra that has been extensively used in query processing. Furthermore, semijoins have been used to formulate desirable properties of acyclic schemas; in particular, a schema is acyclic if and only if it has a full reducer, i.e., a sequence of semijoins that converts a given collection of relat
Xingtong Yu, Shenghua Ye, Ruijuan Liang, Chang Zhou
Graph foundation models (GFM) aim to acquire transferable knowledge by pre-training on diverse graphs, which can be adapted to various downstream tasks. However, domain shift in graphs is inherently two-dimensional: graphs differ not only in what they describe (topic domains) but also in how they are represented (format domains). Most existing GFM benchmarks
An Dang, Jayjun Lee, Mustafa Mukadam, X. Alice Wu
In this paper, we address the problem of tactile sim-to-real policy transfer for contact-rich tasks. Existing methods primarily focus on vision-based sensors and emphasize image rendering quality while providing overly simplistic models of force and shear. Consequently, these models exhibit a large sim-to-real gap for many dexterous tasks. Here, we present H
Modulational instability of small amplitude periodic traveling waves in the $b$-family of Novikov equation
math.APXin Zhao, Lin Lu, Aiyong Chen
We study the modulational instability of smooth, small-amplitude periodic traveling wave solutions to the $b$-family of Novikov equation with cubic nonlinearity with an arbitrary coefficient $b>0$. Our approach is based on applying spectral perturbation theory to the corresponding linearization process. We derive a modulation instability index dependent on t
Identifying and distinguishing quenching galaxies with spatially resolved star formation in TNG50
astro-ph.GACameron Lawlor-Forsyth, Michael L. Balogh, Elizaveta Sazonova, Cameron R. Morgan
Using the TNG50 simulation, we determine observationally motivated metrics that can distinguish quenching galaxies from star forming galaxies for $M_{*} \geqslant 10^{9.5}~M_{\odot}$, based on the spatial distribution of their stellar populations. Quenching galaxies are not fully quenched but have low levels of ongoing star formation that decreases over time
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
cs.CVZhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei
Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial components like arms or interacting objects. Generative models are
Y. Ishii, N. Sasabe, Y. Yamasaki
Altermagnets are compensated collinear magnets that break time-reversal symmetry without net magnetization, enabling unconventional magneto-optical responses. Here, altermagnetic X-ray magnetic circular dichroism (XMCD) is experimentally demonstrated in hematite $\alpha$-Fe$_2$O$_3$. By employing a symmetry-selective geometry in which the x-ray propagation v
Xudan Huang, Zifeng Yuan, Chon-Hei Lo, Huacong Sun
The selection of stacking order in a broad range of close-packed polymorphic materials remains a challenging enigma. Using in situ cryogenic transmission electron microscopy, we uncover the atomistic mechanisms governing the vapour deposition growth of ice. We find that the heterogeneous ice nucleation and growth undergoes recrystallization accompanied by bi
Xueyang Li, Yunzhong Lou, Yu Song, Xiangdong Zhou
Computer-Aided Design (CAD) generative modeling has a strong and long-term application in the industry. Recently, the parametric CAD sequence as the design logic of an object has been widely mined by sequence models. However, the industrial CAD models, especially in component objects, are fine-grained and complex, requiring a longer parametric CAD sequence t
Improving Automatic Summarization of Radiology Reports through Mid-Training of Large Language Models
cs.CLMengxian Lyu, Cheng Peng, Ziyi Chen, Mengyuan Zhang
Automatic summarization of radiology reports is an essential application to reduce the burden on physicians. Previous studies have widely used the "pre-training, fine-tuning" strategy to adapt large language models (LLMs) for summarization. This study proposed a subdomain adaptation through a mid-training method to improve summarization. We explored three ad
Zhiyi Zhao, Ye Guo, Zhenjia Lin, Yinliang Xu
This paper considers the flexibility degradation problem caused by excessive flexible ramping product (FRP) requirements with high variable energy resource (VER) penetration}. Based on the rolling-window co-optimization model of energy and FRP, theoretical analysis of this paper reveals a unit dispatch transfer effect, in which high FRP requirements under fo
Self-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language Models
cs.CVApril Fu
Although Large Vision-Language Models (LVLMs) have made substantial progress, hallucination, where generated text is not grounded in the visual input, remains a challenge. As LVLMs become stronger, previously reported hallucination patterns, such as linguistic bias and overthinking phenomenon, become far less consistent, making the corresponding mitigation t
Jinmyeong Shin, Joshua Tapia, Nicholas Ferreira, Gabriel Diaz
The need for machine unlearning is critical for data privacy, yet existing methods often cause Knowledge Contamination by unintentionally damaging related knowledge. Such a degraded model performance after unlearning has been recently leveraged for new inference and backdoor attacks. Most studies design adversarial unlearning requests that require poisoning
Changwen Xing, Yanfeng Lu, Lei Qi, Chenxu Niu
Industrial chip development is inherently iterative, favoring localized, intent-driven updates over rewriting RTL from scratch. Yet most LLM-Aided Hardware Design (LAD) work focuses on one-shot synthesis, leaving this workflow underexplored. To bridge this gap, we for the first time formalize $\Delta$Spec-to-RTL localization, a multi-positive problem mapping
TAP-SLF: Parameter-Efficient Adaptation of Vision Foundation Models for Multi-Task Ultrasound Image Analysis
cs.CVHui Wan, Libin Lan
Executing multiple tasks simultaneously in medical image analysis, including segmentation, classification, detection, and regression, often introduces significant challenges regarding model generalizability and the optimization of shared feature representations. While Vision Foundation Models (VFMs) provide powerful general representations, full fine-tuning
A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs
cs.CLAnna Feldman, Libby Barak, Jing Peng
We introduce a typology-aware diagnostic for multilingual masked language models that tests reliance on word order versus inflectional form. Using Universal Dependencies, we apply inference-time perturbations: full token scrambling, content-word scrambling with function words fixed, dependency-based head--dependent swaps, and sentence-level lemma substitutio
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
cs.CVHulingxiao He, Zhi Tan, Yuxin Peng
A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist. Large Multimodal Models (LMMs) have achieved remarkable progress in fine-grained visual recognition (FGVR) for known categories. However, t
Qing Luo, Fu Luo, Ke Li, Zhenkun Wang
Construction-based neural routing solvers, typically composed of an encoder and a decoder, have emerged as a promising approach for solving vehicle routing problems. While recent studies suggest that shifting parameters from the encoder to the decoder enhances performance, most works restrict the decoder size to 1-3M parameters, leaving the effects of scalin
Mohammad Amin Samadi, Nia Nixon
Collaborative problem solving and learning are shaped by who or what is on the team. As large language models (LLMs) increasingly function as collaborators rather than tools, a key question is whether AI teammates can be aligned to express personality in predictable ways that matter for interaction and learning. We investigate AI personality alignment throug
Zhenyu Ni, Dongquan Cheng, Jing Wang, Liying Kang
Given a graph $F$, the expansion $F^{(r)}$ of $F$ is defined as the $r$-uniform hypergraph obtained from $F$ by adding a set of $(r-2)$ distinct new vertices to each edge of $F$. In this paper, we investigate spectral stability results for hypergraphs and their applications.We first establish a spectral stability property: for any $r$-uniform hypergraph cont
A Unified Interpretation of Supernova, GRB, and QSO Time Dilation Signals in a Generalized Cosmological Time Framework
hep-phSeokcheon Lee
Cosmological time dilation (CTD) serves as a fundamental probe of cosmic expansion, historically verified through the characteristic (1+z) broadening of Type Ia supernova (SNe Ia) light curves. However, significant tensions arise when extending this test to other astrophysical regimes. While discrete, event-based transients such as Gamma-Ray Bursts (GRBs) ex
Yannian Gu, Zhongzhen Huang, Linjie Mu, Xizhuo Zhang
Multimodal large language models (MLLMs) demonstrate considerable potential in clinical diagnostics, a domain that inherently requires synthesizing complex visual and textual data alongside consulting authoritative medical literature. However, existing benchmarks primarily evaluate MLLMs in end-to-end answering scenarios. This limits the ability to disentang
Cunyuan Yang, Dejuan Song, Xiaotao Pang, Qianqian Shen
The automatic generation of medical reports utilizing Multimodal Large Language Models (MLLMs) frequently encounters challenges related to factual instability, which may manifest as the omission of findings or the incorporation of inaccurate information, thereby constraining their applicability in clinical settings. Current methodologies typically produce re
Dyah Adila, John Cooper, Alexander Yun, Avi Trost
Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices -- such as intervention location and parameterization -- that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions
Siqi Wu, Zhenqi Zhang, Xingyue Liu, Chuanshuai Zhu
Compact optical clocks with high stability are essential for next-generation frequency standard field applications, from navigation to geodesy, yet existing vapor cell clock systems have remained confined to fractional instabilities over $10^{-15}$. Here we report the breaking of this long standing barrier by demonstrating a molecular iodine optical clock th
Hyungi Min, Taeseung You, Hangyeul Lee, Yeongjae Cho
Counterfactual medical image generation have emerged as a critical tool for enhancing AI-driven systems in medical domain by answering "what-if" questions. However, existing approaches face two fundamental limitations: First, they fail to prevent unintended modifications, resulting collateral changes in demographic attributes when only disease features shoul
Harrison Katz
Tourism demand forecasting is methodologically mature, but it typically treats accommodation supply as fixed or exogenous. In platform-mediated short-term rentals, supply is elastic, decision-driven, and co-evolves with demand through pricing, information design, and interventions. I reframe the core issue as endogenous stock-out censoring: realized booked n
TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot
cs.RORuijie Tang, Chi Kit Ng, Kaixuan Wu, Long Bai
In-vivo environments, magnetically actuated soft robots offer advantages such as wireless operation and precise control, showing promising potential for painless detection and therapeutic procedures. We developed a trileg magnetically driven soft robot (TMR) whose multi-legged design enables more flexible gaits and diverse motion patterns. For the silicone m
Doyi Kim, Minseok Seo, Changick Kim
Precipitation forecasting relies on heterogeneous data. Weather radar is accurate, but coverage is geographically limited and costly to maintain. Weather stations provide accurate but sparse point measurements, while satellites offer dense, high-resolution coverage without direct rainfall retrieval. To overcome these limitations, we propose Query-Conditioned
Jeongho Bang, Kyoungho Cho
Beyond binary classification, learnability can become a logically fragile notion: in EMX, even the class of all finite subsets of $[0,1]$ is learnable in some models of ZFC and not in others. We argue the paradox is operational. The standard definitions quantify over arbitrary set-theoretic learners that implicitly assume non-operational resources (infinite