Skip to content

February 2026 arXiv papers — page 4

Showing 301400 of 20,995 papers

  1. Jingyuan Xie, Wenjie Wang, Ji Wu, Jiandong Gao

    Supervised fine-tuning (SFT) is essential for the development of medical large language models (LLMs), yet prior poisoning studies have mainly focused on the detectable backdoor attacks. We propose a novel poisoning attack targeting the reasoning process of medical LLMs during SFT. Unlike backdoor attacks, our method injects poisoned rationales into few-shot

  2. Xingyilang Yin, Chengzhengxu Li, Jiahao Chang, Chi-Man Pun

    Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Despite its importance, this capability remains a significant bottleneck for current multimodal large language models (MLLMs). To tackle this challenge, we introduce MLLM-4D, a compre

  3. Yiwei Qiu, Jiahao Hu, Yi Zhou, Jie Zhu

    This article proposes an energy storage-enhanced hydrogen electrolyzer (ESEHE) to provide grid-forming (GFM) services for off-grid renewable power to hydrogen (ReP2H) systems. Unlike conventional ReP2H systems that use a centralized energy storage (ES) plant, the proposed topology directly connects batteries to the DC buses of electrolysis rectifiers. A tail

  4. Ram Milan Kumar Verma, Shashi Ranjan Kumar, Hemendra Arya

    Precise motion control of underactuated surface vessels is a crucial task in various maritime applications. In this work, we develop a nonlinear motion control strategy for surface vessels inspired by the pursuit guidance philosophy. Any sufficiently smooth path can be seen as a continuum of virtual targets moving along a specified path, which the pursuer is

  5. Wang Chen, Yuhui Zeng, Yongdong Luo, Tianyu Xie

    Frame selection is crucial due to high frame redundancy and limited context windows when applying Large Vision-Language Models (LVLMs) to long videos. Current methods typically select frames with high relevance to a given query, resulting in a disjointed set of frames that disregard the narrative structure of video. In this paper, we introduce Wavelet-based

  6. Ruoshuang Du, Xin Sun, Qiang Liu, Bowen Song

    Visual Question Answering systems face reliability issues due to hallucinations, where models generate answers misaligned with visual input or factual knowledge. While Retrieval Augmented Generation frameworks mitigate this issue by incorporating external knowledge, static retrieval often introduces irrelevant or conflicting content, particularly in visual R

  7. Yingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu Shen

    Multimodal large language models (MLLMs) project visual tokens into the embedding space of language models, yet the internal structuring and processing of visual semantics remain poorly understood. In this work, we introduce a two-fold analytical framework featuring a novel probing tool, $\textbf{EmbedLens}$, to conduct a fine-grained analysis. We uncover a

  8. Ce Zhang, Cheng Xu, Haibo Hu, Jianliang Xu

    Blockchain provides a decentralized and tamper-resistant ledger for securely recording transactions across a network of untrusted nodes. While its transparency and integrity are beneficial, the substantial storage requirements for maintaining a complete transaction history present significant challenges. For example, Ethereum nodes require around 23TB of sto

  9. Luyuan Li, Jisheng Bai, Xiruo Su, Xiaoyi Shen

    Active noise control (ANC) is an effective approach to noise suppression, and the filtered-reference least mean square (FxLMS) algorithm is a widely adopted method in ANC systems, owing to its computational efficiency and stable performance. However, its convergence speed and noise reduction performance are highly dependent on the step size parameter. Common

  10. Jiamin Shi, Haolin Zhang, Yuchen Yan, Shitao Chen

    Navigating social robots in dense, dynamic crowds is challenging due to environmental uncertainty and complex human-robot interactions. While Model Predictive Control (MPC) offers strong real-time performance, its reliance on a fixed prediction horizon limits adaptability to changing environments and social dynamics. Furthermore, most MPC approaches treat pe

  11. Hongquan Wang, Hanshu Chen, Ilia Marchevsky, Zhuojia Fu

    DeepONet enables retraining-free inference across varying initial conditions or source terms at the cost of high computational requirements. This paper proposes a hybrid quantum operator network (Quantum AS-DeepOnet) suitable for solving 2D evolution equations. By combining Parameterized Quantum Circuits and cross-subnet attention methods, we can solve 2D ev

  12. Kautuk Astu, Suman Raj, Priyanshu Pansari, Yogesh Simmhan

    The increasing adoption of UAVs equipped with advanced sensors and GPU-accelerated edge computing has enabled real-time AI-driven applications in domains such as precision agriculture, wildfire monitoring, and environmental conservation. However, the integrated design and orchestration of navigation, sensing, and analytics, together with seamless real-time c

  13. Keunho Byeon, Jinsol Song, Seong Min Hong, Yosep Chong

    Whole-slide image analysis is essential for diagnostic tasks in pathology, yet existing deep learning methods primarily rely on flat classification, ignoring hierarchical relationships among class labels. In this study, we propose HiClass, a hierarchical classification framework for improved histopathology image analysis, that enhances both coarse-grained an

  14. Dawei Yan, Haokui Zhang, Guangda Huzhang, Yang Li

    Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on extensive data collection and model training, yet still struggle with high computational costs and insufficient reasoning cap

  15. Wenhao Zheng, Wang Lu, Fangshuang Tang, Yiyang Lu

    Early-stage users in a new scenario intensify cold-start challenges, yet prior works often address only parts of the problem through model architecture. Launching a new user experience to replace an established product involves sparse behavioral signals, low-engagement cohorts, and unstable model performance. We argue that effective recommendations require t

  16. Jingwen Tong, Zijian Li, Fang Liu, Wei Guo

    The integration of large language models (LLMs) into wireless networks has sparked growing interest in building autonomous AI agents for wireless tasks. However, existing approaches rely heavily on manually crafted prompts and static agentic workflows, a process that is labor-intensive, unscalable, and often suboptimal. In this paper, we propose WirelessAgen

  17. Zilong Xie, Jingyu Gong, Xin Tan, Zhizhong Zhang

    Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense reasoning, they struggle with geometric and spatial understanding required for pose prediction. In this paper, we propos

  18. Zhang-nan Hu, Bing Li, YiJing Wang

    In this paper, we study the uniform random covering problem in general metric space $(X,d)$. Let $\omega=(\omega_n)_{n\in\mathbb N}$ be a sequence of independent identically distributed random variables on $(X,\mu)$, and $\ell=(\ell_n)_{n\in\mathbb N}$ a sequence of positive real numbers. We analyze the size of the set \[\mathcal{U}(\omega,\ell)=\left\{y\in

  19. Quoc Minh Nguyen, Trung Le, Jing Wu, Anh Tuan Bui

    Fine-tuning-as-a-service introduces a threat to Large Language Models' safety when service providers fine-tune their models on poisoned user-submitted datasets, a process known as harmful fine-tuning attacks. In this work, we show that by regularizing the gradient contribution of harmful samples encountered during fine-tuning, we can effectively mitigate the

  20. Maria Giovanna Dainotti, Avik Banerjee, Andre' LeClair, Giovanni Montani

    In view of the current and increasing evidence of a running Hubble constant, we investigate its redshift dependence within the flat $\Lambda$CDM framework using a 20-bin analysis of the Master SNe~Ia Sample \citep{2025JHEAp..4800405D}, considering cases with and without very low-redshift data. For each case, we obtain best-fitting values of $H_0$ and $\Omega

  21. Kazuhiro Hiraki, Shinichi Ishihara, Takumi Kongo, Junnosuke Shino

    In this paper, we provide a theoretically grounded and computationally efficient alternative to SHAP. To this end, we study feature attribution through the lens of cooperative game theory by formulating a class of XAI--TU games. Building on this formulation, we investigate equal-surplus-type and proportional-allocation-type attribution rules and propose a lo

  22. Christopher Cruz

    We introduce AI Runtime Infrastructure, a distinct execution-time layer that operates above the model and below the application, actively observing, reasoning over, and intervening in agent behavior to optimize task success, latency, token efficiency, reliability, and safety while the agent is running. Unlike model-level optimizations or passive logging syst

  23. Jiahao Zeng, Zhenkui Shi, Chunpei Li, Mengkai Yan

    With the advancement of vehicle-to-vehicle (V2V) ad hoc networks and wireless communication technologies, mobile edge caching has become a key enabler for enhancing network performance and user experience. However, traditional federated learning-based collaborative caching approaches in vehicular scenarios suffer from inadequate client selection mechanisms a

  24. Emmanuel Pintelas, Ioannis E. Livieris

    Excessive reliance on validation performance during model selection can lead to validation overfitting (VO), where models appear effective during development but fail at test time. This issue is further amplified in low-data regimes and under distribution shifts, where validation signals become unreliable. Although ensemble learning is widely used to improve

  25. Rahul Kumar, Je-Geun Park

    The emergence of a long-range magnetic order in the atomically thin, two-dimensional (2D) limit has long remained a fundamental question in condensed matter physics. The advent of exfoliable van der Waals (vdW) materials, particularly transition-metal phosphorus trisulfides (T MPS3; T M = Fe, Ni, and Mn), provided the first experimental access to this regime

  26. Yuchen Che, Jingtu Wu, Hao Zheng, Asako Kanezaki

    Estimating the 6DoF pose of a novel object with a single reference view is challenging due to occlusions, view-point changes, and outliers. A core difficulty lies in finding robust cross-view correspondences, as existing methods often rely on discrete one-to-one matching that is non-differentiable and tends to collapse onto sparse key-points. We propose Conf

  27. Riccardo de Lutio, Tobias Fischer, Yen-Yu Chang, Yuxuan Zhang

    Per-scene optimization methods such as 3D Gaussian Splatting provide state-of-the-art novel view synthesis quality but extrapolate poorly to under-observed areas. Methods that leverage generative priors to correct artifacts in these areas hold promise but currently suffer from two shortcomings. The first is scalability, as existing methods use image diffusio

  28. Xianchao Xiu, Shenghao Sun, Xinrong Li, Jiyuan Tao

    Support matrix machine (SMM) is an emerging classification framework that directly handles matrix-structured observations, thereby avoiding the spatial correlations destroyed by vectorization. However, most existing SMM variants rely on convex or nonconvex surrogate loss functions, which may lead to high sensitivity to noise. To address this issue, we propos

  29. Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen

    The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective assistance in dynamic, real-world environments remains largely underexplored. Existing video benchmarks predominantly assess pas

  30. Haoyu Gao, Hong Yi Lin, Christoph Treude, Gregory Gay

    README files are critical for understanding and onboarding contributors to open-source software, yet they frequently become outdated. We formulate surgical documentation update recommendation as a task and present a Large Language Model-driven framework for use in a human-in-the-loop workflow. Given a pull request, the framework determines whether a README u

  31. Achmad Ardani Prasha, Clavino Ourizqi Rachmadi, Sabrina Laila Mutiara, Hilman Syachr Ramadhan

    Adolescent pornography addiction requires early detection based on objective neurobiological biomarkers because self-report is prone to subjective bias due to social stigma. Conventional machine learning has not been able to model dynamic functional connectivity of the brain that fluctuates temporally during addictive stimulus exposure. This study proposes a

  32. Aaron Bryce, Rajeev Gore'

    Gentzen's 1936 proof of the consistency of Peano Arithmetic was a significant result in the foundations of mathematics. We provide here a modified version of the proof, based on G\"{o}del's reformulation, and including additional details and minor corrections which are necessary to definitively prove the well-foundedness of the cut-elimination argument in a

  33. Qihang Fan, Yuang Ai, Huaibo Huang, Ran He

    Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representative approach involves grouping tokens, performing self-attention calculations within each group, or pooling the tokens within each group into a single token. To this end, various c

  34. Ken Gu, Srishti Palani, Vidya Setlur

    Conversational interfaces are increasingly used for data analysis, enabling data workers to express complex analytical intents in natural language. Yet, these interactions unfold as long, linear transcripts that are misaligned with the iterative, nonlinear nature of real-world analyses. Revisiting and summarizing conversations for different contexts is there

  35. Yi Zhou, Haocheng Fu, Yiping Liu, Jian Mao

    The two-dimensional irregular bin packing problem (2DIBPP) aims to pack a given set of irregular polygons, referred to as pieces, into fixed-size rectangular bins without overlap, while maximizing bin utilization. Although numerous metaheuristic algorithms have been proposed for the 2DIBPP, many industrial applications favor simpler constructive heuristics d

  36. Liyao Jiang, Ruichen Chen, Chao Gao, Di Niu

    Recent text-to-image (T2I) diffusion models achieve remarkable realism, yet faithful prompt-image alignment remains challenging, particularly for complex prompts with multiple objects, relations, and fine-grained attributes. Existing training-free inference-time scaling methods rely on fixed iteration budgets that cannot adapt to prompt difficulty, while ref

  37. Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li

    Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granularity, overlong visual token sequences, and inadequate cross-modal alignment. To overcome these challenges, we propose TaiChi, a novel VLM fr

  38. Khaleque Md Aashiq Kamal, Surya Eada, Aayushi Verma, Subek Acharya

    Developments in the machine learning voting domain have shown both promising results and risks. Trained models perform well on ballot classification tasks (> 99% accuracy) but are at risk from adversarial example attacks that cause misclassifications. In this paper, we analyze an attacker who seeks to deploy adversarial examples against machine learning ball

  39. A. E. Petrova, S. Yu. Gavrilkin, D. Menzel, V. A. Stepanov

    We studied the longitudinal and transverse magnetoresistance of helical magnets, MnSi and Mn$_{1-x}$Co$_x$Si, at temperatures between 1.8 and 100~K and in magnetic fields up to 9 Tesla. All substances exhibited negative longitudinal and transverse magnetoresistance at temperatures above 4~K, which is most likely related to the suppression of spin fluctuation

  40. Pengcheng Shi, Minghui Zhang, Kehan Song, Jiaqi Liu

    Automated radiology report generation is key for reducing radiologist workload and improving diagnostic consistency, yet generating accurate reports for 3D medical imaging remains challenging. Existing vision-language models face two limitations: they do not leverage segmentation-pretrained encoders, and they inject visual features only at the input layer of

  41. Xu Luo, Ji Zhang, Lianli Gao, Heng Tao Shen

    Few-shot transfer has been revolutionized by stronger pre-trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is both challenging and realistic for real-world usage. In this work, we establish FEWTRANS, a comprehensive benchmark containing 10 diverse datasets, and propose the Hyperparameter Ense

  42. Qiqi Gu, Chenpeng Wu, Heng Shi, Jianguo Yao

    Stencil computation constitutes a cornerstone of scientific computing, serving as a critical kernel in domains ranging from fluid dynamics to weather simulation. While stencil computations are conventionally regarded as memory-bound and thus unsuitable for compute-centric Tensor Cores, recent empirical studies have demonstrated significant speedups after app

  43. Linxi Jiang, Zhijie Liu, Haotian Luo, Zhiqiang Lin

    Browser-use agents are widely used for everyday tasks. They enable automated interaction with web pages through structured DOM based interfaces or vision language models operating on page screenshots. However, web pages often change between planning and execution, causing agents to execute actions based on stale assumptions. We view this temporal mismatch as

  44. Harsha Desu, Niladri S. Satpathi, Lokesh Malik, Ashis K. Sen

    Capillary-driven transport offers a simple, self-sustained alternative to externally pumped microfluidic systems, yet achieving precise control of such flows remains challenging. We experimentally and theoretically investigate capillary flow in rectangular microchannels when a liquid meniscus encounters a geometric step with varying channel width and height.

  45. Jiacheng Wang, Yucheng Sheng, Le Liang, Hao Ye

    This paper investigates the power control problem in wireless networks by repurposing pre-trained large language models (LLMs) as relational reasoning backbones. In hyper-connected interference environments, traditional optimization methods face high computational cost, while standard message passing neural networks suffer from aggregation bottlenecks that c

  46. Yibo Yan, Chao Liu, Jiadong Li, Feng Wang

    Upcoming next-generation sky surveys will detect large number of faint objects with magnitudes larger than 25. When objects are crowded within a limited a field of view, blending becomes unavoidable. Blending leads to the omission of many sources during photometry in these fields, which cause an underestimates of tens of percent in crowded fields, and remain

  47. Yijun Yu

    Agentic AI systems exhibit numerous crosscutting concerns -- security, observability, cost management, fault tolerance -- that are poorly modularized in current implementations, contributing to the high failure rate of AI projects in reaching production. The goals-to-aspects methodology proposed at RE 2004 demonstrated that aspects can be systematically disc

  48. Maohua Gong, Qiutong Zhen, Yujie Tang, Peng Hu

    Chiral bound states in the continuum (BICs) are confined photonic modes with infinite quality factors and chiral response, offering significant potential for chiral optics. Although a novel type of spin-orbital-locking chiral BIC was recently predicted in magneto-optical (MO) photonic crystals (PhCs) that break time-reversal symmetry (TRS), its experimental

  49. Ye Li, Yi Zhou, Yingdong Hu, Zhe Ji

    The direct-to-device (D2D) satellite network is an important 6G evolution direction to enable seamless ubiquitous connectivity. However, the network faces critical handover challenges due to high satellite mobility and wide beam footprints. Conventional handover strategies, mostly designed for terrestrial networks, may encounter excessive co-channel interfer

  50. Najeeb Khan

    Operators of Earth observation satellites need justifications for scheduling decisions: why a request was selected, rejected, or what changes would make it schedulable. Existing approaches construct post-hoc reasoning layers independent of the optimizer, risking non-causal attributions, incomplete constraint conjunctions, and solver-path dependence. We take

  51. Ethan Chu, Yiyang Guo, Jan Hoffmann

    There exist many techniques for automatically deriving parametric resource (or cost) bounds by analyzing the source code of a program. These techniques work effectively for a large class of programs and language features. However, non-local transfer of control as needed for exception or effect handlers has remained a challenge. This paper presents the first

  52. Krishna Kumar Neelakanta Pillai Santha Kumari Amma

    Dynamic pricing in competitive retail markets requires strategies that adapt to fluctuating demand and competitor behavior. In this work, we present a systematic empirical evaluation of multi-agent reinforcement learning (MARL) approaches-specifically MAPPO and MADDPG-for dynamic price optimization under competition. Using a simulated marketplace environment

  53. Yilun Wang, Guangba Yu, Haiyu Huang, Yujie Huang

    LLM agents are increasingly explored for automating root cause analysis (RCA) in cloud-native systems, creating a need to evaluate both diagnostic correctness and the quality of the supporting investigation. Existing static benchmarks offer repeatable inputs but limited system-facing interaction, while live testbeds expose realistic tools but hinder controll

  54. Pengju Sun, Banglei Guan, Jing Tao, Zhenbao Yu

    High dynamic range (HDR) imaging under extreme illumination remains challenging for conventional cameras due to overexposure. Event cameras provide microsecond temporal resolution and high dynamic range, while spatially varying exposure (SVE) sensors offer single-shot radiometric diversity.We present a hardware--algorithm co-designed HDR imaging system that

  55. Boming Tan, Xiangdong Zhang, Ning Liao, Yuqing Zhang

    Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world-related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single wor

  56. Yucheng Chu, Hang Li, Kaiqi Yang, Yasemin Copur-Gencturk

    Automated assessment of open-ended student responses is a critical capability for scaling personalized feedback in education. While large language models (LLMs) have shown promise in grading tasks via in-context learning (ICL), their reliability is heavily dependent on the selection of few-shot exemplars and the construction of high-quality rationales. Stand

  57. Chang Xue, Fang Liu, Jiaye Wang, Jinming Xing

    Decentralized financial platforms rely heavily on Web of Trust reputation systems to mitigate counterparty risk in the absence of centralized identity verification. However, these pseudonymous networks are inherently vulnerable to adversarial behaviors, such as Sybil attacks and camouflaged fraud, where malicious actors cultivate artificial reputations befor

  58. Zhaolin Yu, Litao Yang, Ben Babicka, Ming Hu

    Orthopantomograms (OPGs) are the standard panoramic radiograph in dentistry, used for full-arch screening across multiple diagnostic tasks. While Vision Language Models (VLMs) now allow multi-task OPG analysis through natural language, they underperform task-specific models on most individual tasks. Agentic systems that orchestrate specialized tools offer a

  59. Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu

    The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, which significantly increases KV cache memory usage and time-to-

  60. Yucheng Chu, Haoyu Han, Shen Dong, Hang Li

    Automated short answer grading (ASAG) is critical for scaling educational assessment, yet large language models (LLMs) often struggle with hallucinations and strict rubric adherence due to their reliance on generalized pre-training. While Rretrieval-Augmented Generation (RAG) mitigates these issues, standard "flat" vector retrieval mechanisms treat knowledge

  61. Shuheng Chen, Namratha Patil, Haonan Pan, Angel Hsing-Chi Hwang

    Clinical decision-making requires synthesizing heterogeneous evidence, including patient histories, clinical guidelines, and trajectories of comparable cases. While large language models (LLMs) offer strong reasoning capabilities, they remain prone to hallucinations and struggle to integrate long, structured medical documents. We present MED-COPILOT, an inte

  62. Rajdeep Chatterjee, Sudip Chakrabarty, Trishaani Acharjee

    Accurate semantic segmentation of foot ulcers is essential for automated wound monitoring, yet boundary delineation remains challenging due to tissue heterogeneity and poor contrast with surrounding skin. To overcome the limitations of standard intensity-based networks, we present LSS-LTCNet:an ante-hoc explainable framework synergizing deterministic structu

  63. Bin Chen, Weiqi Li, Shijie Zhao, Xuanyu Zhang

    While many diffusion models have achieved impressive results in real-world video super-resolution (Real-VSR) by generating rich and realistic details, their reliance on multi-step sampling leads to slow inference. One-step networks like SeedVR2, DOVE, and DLoRAL alleviate this through condensing generation into one single step, yet they remain heavy, with bi

  64. Joy Das Bairagya, Sagar Chakraborty

    A colony of the queenless ant species, \emph{Pristomyrmex punctatus}, can broadly be seen as consisting of small-body sized worker ants and relatively larger body-sized cheater ants. Hence, in the presence of inter-colony migration, a set of constituent colonies act as a metapopulation exclusively composed of cooperators and defectors. Such a set-up facilita

  65. Yao Wu, Ziye Jia, Jingjing Zhao, Haoyang Wang

    Unmanned aerial vehicle (UAV) networks are increasingly deployed for complex missions, including disaster response, intelligent logistics, and environmental monitoring. These missions generally require coordinated collaboration among multiple UAVs across distinct administrative domains. To support such cross-domain cooperation, service function chains (SFCs)

  66. Shivanshu Tripathi, Reza Akbarian Bafghi, Maziar Raissi

    In this work, we present a test-driven, agentic framework for synthesizing a deployable low-level robot controller for navigation tasks. Given a 2D map with an image of an ultrasonic sensor-based robot, or a 3D robotic simulation environment, our framework iteratively refines the generated controller code using diagnostic feedback from structured test suites

  67. Quhura Fathima, Neda Moghim, Mostafa Taghizade Firouzjaee, Christo K. Thomas

    The growing deployment of Internet of Things (IoT) devices in smart cities and industrial environments increases vulnerability to stealthy, multi-stage advanced persistent threats (APTs) that exploit wireless communication. Detection is challenging due to severe class imbalance in network traffic, which limits the effectiveness of traditional deep learning a

  68. Jocelyn Shen, Nicolai Marquardt, Hugo Romat, Ken Hinckley

    What if text could be sculpted and refined like clay -- or cultivated and pruned like a plant? Texterial reimagines text as a material that users can grow, sculpt, and transform. Current generative-AI models enable rich text operations, yet rigid, linear interfaces often mask such capabilities. We explore how the text-as-material metaphor can reveal AI-enabl

  69. Yucheng Chu, Hang Li, Kaiqi Yang, Yasemin Copur-Gencturk

    Accurate and unambiguous guidelines are critical for large language model (LLM) based graders, yet manually crafting these prompts is often sub-optimal as LLMs can misinterpret expert guidelines or lack necessary domain specificity. Consequently, the field has moved toward automated prompt optimization to refine grading guidelines without the burden of manua

  70. Sam Deen, Derek Lam

    2024 YR4 is a 40-100 meter-diameter asteroid and former Torino Scale 3 object which currently has a roughly 4% chance of impacting the Moon on 2032 December 22, an event which recent studies suggest could pose a hazard on Earth due to impact ejecta. We present a search for, and identification of, potential precovery observations of the virtual lunar impactor

  71. Phokion G. Kolaitis

    The semijoin operation is a fundamental operation of relational algebra that has been extensively used in query processing. Furthermore, semijoins have been used to formulate desirable properties of acyclic schemas; in particular, a schema is acyclic if and only if it has a full reducer, i.e., a sequence of semijoins that converts a given collection of relat

  72. Xingtong Yu, Shenghua Ye, Ruijuan Liang, Chang Zhou

    Graph foundation models (GFM) aim to acquire transferable knowledge by pre-training on diverse graphs, which can be adapted to various downstream tasks. However, domain shift in graphs is inherently two-dimensional: graphs differ not only in what they describe (topic domains) but also in how they are represented (format domains). Most existing GFM benchmarks

  73. An Dang, Jayjun Lee, Mustafa Mukadam, X. Alice Wu

    In this paper, we address the problem of tactile sim-to-real policy transfer for contact-rich tasks. Existing methods primarily focus on vision-based sensors and emphasize image rendering quality while providing overly simplistic models of force and shear. Consequently, these models exhibit a large sim-to-real gap for many dexterous tasks. Here, we present H

  74. Xin Zhao, Lin Lu, Aiyong Chen

    We study the modulational instability of smooth, small-amplitude periodic traveling wave solutions to the $b$-family of Novikov equation with cubic nonlinearity with an arbitrary coefficient $b>0$. Our approach is based on applying spectral perturbation theory to the corresponding linearization process. We derive a modulation instability index dependent on t

  75. Cameron Lawlor-Forsyth, Michael L. Balogh, Elizaveta Sazonova, Cameron R. Morgan

    Using the TNG50 simulation, we determine observationally motivated metrics that can distinguish quenching galaxies from star forming galaxies for $M_{*} \geqslant 10^{9.5}~M_{\odot}$, based on the spatial distribution of their stellar populations. Quenching galaxies are not fully quenched but have low levels of ongoing star formation that decreases over time

  76. Zhuoran Zhao, Xianghao Kong, Linlin Yang, Zheng Wei

    Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial components like arms or interacting objects. Generative models are

  77. Y. Ishii, N. Sasabe, Y. Yamasaki

    Altermagnets are compensated collinear magnets that break time-reversal symmetry without net magnetization, enabling unconventional magneto-optical responses. Here, altermagnetic X-ray magnetic circular dichroism (XMCD) is experimentally demonstrated in hematite $\alpha$-Fe$_2$O$_3$. By employing a symmetry-selective geometry in which the x-ray propagation v

  78. Xudan Huang, Zifeng Yuan, Chon-Hei Lo, Huacong Sun

    The selection of stacking order in a broad range of close-packed polymorphic materials remains a challenging enigma. Using in situ cryogenic transmission electron microscopy, we uncover the atomistic mechanisms governing the vapour deposition growth of ice. We find that the heterogeneous ice nucleation and growth undergoes recrystallization accompanied by bi

  79. Xueyang Li, Yunzhong Lou, Yu Song, Xiangdong Zhou

    Computer-Aided Design (CAD) generative modeling has a strong and long-term application in the industry. Recently, the parametric CAD sequence as the design logic of an object has been widely mined by sequence models. However, the industrial CAD models, especially in component objects, are fine-grained and complex, requiring a longer parametric CAD sequence t

  80. Mengxian Lyu, Cheng Peng, Ziyi Chen, Mengyuan Zhang

    Automatic summarization of radiology reports is an essential application to reduce the burden on physicians. Previous studies have widely used the "pre-training, fine-tuning" strategy to adapt large language models (LLMs) for summarization. This study proposed a subdomain adaptation through a mid-training method to improve summarization. We explored three ad

  81. Zhiyi Zhao, Ye Guo, Zhenjia Lin, Yinliang Xu

    This paper considers the flexibility degradation problem caused by excessive flexible ramping product (FRP) requirements with high variable energy resource (VER) penetration}. Based on the rolling-window co-optimization model of energy and FRP, theoretical analysis of this paper reveals a unit dispatch transfer effect, in which high FRP requirements under fo

  82. April Fu

    Although Large Vision-Language Models (LVLMs) have made substantial progress, hallucination, where generated text is not grounded in the visual input, remains a challenge. As LVLMs become stronger, previously reported hallucination patterns, such as linguistic bias and overthinking phenomenon, become far less consistent, making the corresponding mitigation t

  83. Jinmyeong Shin, Joshua Tapia, Nicholas Ferreira, Gabriel Diaz

    The need for machine unlearning is critical for data privacy, yet existing methods often cause Knowledge Contamination by unintentionally damaging related knowledge. Such a degraded model performance after unlearning has been recently leveraged for new inference and backdoor attacks. Most studies design adversarial unlearning requests that require poisoning

  84. Changwen Xing, Yanfeng Lu, Lei Qi, Chenxu Niu

    Industrial chip development is inherently iterative, favoring localized, intent-driven updates over rewriting RTL from scratch. Yet most LLM-Aided Hardware Design (LAD) work focuses on one-shot synthesis, leaving this workflow underexplored. To bridge this gap, we for the first time formalize $\Delta$Spec-to-RTL localization, a multi-positive problem mapping

  85. Hui Wan, Libin Lan

    Executing multiple tasks simultaneously in medical image analysis, including segmentation, classification, detection, and regression, often introduces significant challenges regarding model generalizability and the optimization of shared feature representations. While Vision Foundation Models (VFMs) provide powerful general representations, full fine-tuning

  86. Anna Feldman, Libby Barak, Jing Peng

    We introduce a typology-aware diagnostic for multilingual masked language models that tests reliance on word order versus inflectional form. Using Universal Dependencies, we apply inference-time perturbations: full token scrambling, content-word scrambling with function words fixed, dependency-based head--dependent swaps, and sentence-level lemma substitutio

  87. Hulingxiao He, Zhi Tan, Yuxin Peng

    A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist. Large Multimodal Models (LMMs) have achieved remarkable progress in fine-grained visual recognition (FGVR) for known categories. However, t

  88. Qing Luo, Fu Luo, Ke Li, Zhenkun Wang

    Construction-based neural routing solvers, typically composed of an encoder and a decoder, have emerged as a promising approach for solving vehicle routing problems. While recent studies suggest that shifting parameters from the encoder to the decoder enhances performance, most works restrict the decoder size to 1-3M parameters, leaving the effects of scalin

  89. Mohammad Amin Samadi, Nia Nixon

    Collaborative problem solving and learning are shaped by who or what is on the team. As large language models (LLMs) increasingly function as collaborators rather than tools, a key question is whether AI teammates can be aligned to express personality in predictable ways that matter for interaction and learning. We investigate AI personality alignment throug

  90. Zhenyu Ni, Dongquan Cheng, Jing Wang, Liying Kang

    Given a graph $F$, the expansion $F^{(r)}$ of $F$ is defined as the $r$-uniform hypergraph obtained from $F$ by adding a set of $(r-2)$ distinct new vertices to each edge of $F$. In this paper, we investigate spectral stability results for hypergraphs and their applications.We first establish a spectral stability property: for any $r$-uniform hypergraph cont

  91. Seokcheon Lee

    Cosmological time dilation (CTD) serves as a fundamental probe of cosmic expansion, historically verified through the characteristic (1+z) broadening of Type Ia supernova (SNe Ia) light curves. However, significant tensions arise when extending this test to other astrophysical regimes. While discrete, event-based transients such as Gamma-Ray Bursts (GRBs) ex

  92. Yannian Gu, Zhongzhen Huang, Linjie Mu, Xizhuo Zhang

    Multimodal large language models (MLLMs) demonstrate considerable potential in clinical diagnostics, a domain that inherently requires synthesizing complex visual and textual data alongside consulting authoritative medical literature. However, existing benchmarks primarily evaluate MLLMs in end-to-end answering scenarios. This limits the ability to disentang

  93. Cunyuan Yang, Dejuan Song, Xiaotao Pang, Qianqian Shen

    The automatic generation of medical reports utilizing Multimodal Large Language Models (MLLMs) frequently encounters challenges related to factual instability, which may manifest as the omission of findings or the incorporation of inaccurate information, thereby constraining their applicability in clinical settings. Current methodologies typically produce re

  94. Dyah Adila, John Cooper, Alexander Yun, Avi Trost

    Activation steering promises to be an extremely parameter-efficient form of adaptation, but its effectiveness depends on critical design choices -- such as intervention location and parameterization -- that currently rely on empirical heuristics rather than a principled foundation. We establish a first-order equivalence between activation-space interventions

  95. Siqi Wu, Zhenqi Zhang, Xingyue Liu, Chuanshuai Zhu

    Compact optical clocks with high stability are essential for next-generation frequency standard field applications, from navigation to geodesy, yet existing vapor cell clock systems have remained confined to fractional instabilities over $10^{-15}$. Here we report the breaking of this long standing barrier by demonstrating a molecular iodine optical clock th

  96. Hyungi Min, Taeseung You, Hangyeul Lee, Yeongjae Cho

    Counterfactual medical image generation have emerged as a critical tool for enhancing AI-driven systems in medical domain by answering "what-if" questions. However, existing approaches face two fundamental limitations: First, they fail to prevent unintended modifications, resulting collateral changes in demographic attributes when only disease features shoul

  97. Harrison Katz

    Tourism demand forecasting is methodologically mature, but it typically treats accommodation supply as fixed or exogenous. In platform-mediated short-term rentals, supply is elastic, decision-driven, and co-evolves with demand through pricing, information design, and interventions. I reframe the core issue as endogenous stock-out censoring: realized booked n

  98. Ruijie Tang, Chi Kit Ng, Kaixuan Wu, Long Bai

    In-vivo environments, magnetically actuated soft robots offer advantages such as wireless operation and precise control, showing promising potential for painless detection and therapeutic procedures. We developed a trileg magnetically driven soft robot (TMR) whose multi-legged design enables more flexible gaits and diverse motion patterns. For the silicone m

  99. Doyi Kim, Minseok Seo, Changick Kim

    Precipitation forecasting relies on heterogeneous data. Weather radar is accurate, but coverage is geographically limited and costly to maintain. Weather stations provide accurate but sparse point measurements, while satellites offer dense, high-resolution coverage without direct rainfall retrieval. To overcome these limitations, we propose Query-Conditioned

  100. Jeongho Bang, Kyoungho Cho

    Beyond binary classification, learnability can become a logically fragile notion: in EMX, even the class of all finite subsets of $[0,1]$ is learnable in some models of ZFC and not in others. We argue the paradox is operational. The standard definitions quantify over arbitrary set-theoretic learners that implicitly assume non-operational resources (infinite