Skip to content

April 2026 arXiv papers — page 112

Showing 11,10111,200 of 25,062 papers

  1. Jasper Lu, Zhenhao Shen, Yuanfei Wang, Shugao Liu

    Learning robust robot policies in real-world environments requires diverse data augmentation, yet scaling real-world data collection is costly due to the need for acquiring physical assets and reconfiguring environments. Therefore, augmenting real-world scenes into simulation has become a practical augmentation for efficient learning and evaluation. We prese

  2. Qwen Team

    In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-vis

  3. Jvbin Yao

    We study pair rapid decay for homogeneous spaces \(G/H\) and its applications to random walks and subgroup structure. The entropy framework for groups with rapid decay is extended to homogeneous spaces, proving that the asymptotic Shannon entropy on \(G/H\) agrees with a spectral-radius quantity \(c(G,H;\mu)\) for measures with finite entropy and suitable fi

  4. Hyunseok Park, Jihyeon Kim, Jongeun Kim, Dongsik Yoon

    Retrieval-Augmented Generation (RAG) systems lose retrieval accuracy when similar documents coexist in the vector database, causing unnecessary information, hallucinations, and factual errors. To alleviate this issue, we propose CHOP, a framework that iteratively evaluates chunk relevance with Large Language Models (LLMs) and progressively reconstructs docum

  5. L. Boscagli, G. Rigas, O. Marxen, P. J. K. Bruce

    The extreme heat fluxes characteristic of hypersonic flows significantly limit the flight envelope of hypersonic vehicles. The role of hydrodynamic instability and the onset of laminar to turbulent boundary layer transition is of notable importance. The effect of streaks on the suppression of planar (second Mack mode) instabilities has been previously invest

  6. Yueling Fan, Richard Lee Davis, Olga Viberg

    This study presents WriteFlow, an AI voice-based writing assistant designed to support reflective academic writing through goal-oriented interaction. Academic writing involves iterative reflection and evolving goal regulation, yet prior research and a formative study with 17 participants show that writers often struggle to articulate and manage changing goal

  7. Junpei Oba

    Enhanced local-excitation retention in atomic arrays allows to exploit cooperative radiative effects to suppress emission and prolong excited-state lifetimes. We consider an impurity-assisted setting involving a single storage atom being initially excited and study the survival of local excitation under neither write nor retrieval fields. Because the corresp

  8. Masanori Iwamoto, Kunihito Ioka

    We study induced (stimulated) scattering of linearly polarized, strong electromagnetic waves in pair plasmas, which is crucial for understanding the propagation of fast radio bursts (FRBs). Magnetars are the most promising progenitors of FRBs, and FRBs propagate through the magnetar wind and successfully escape before being significantly scattered. We revisi

  9. Michael Massoth

    This paper develops a physical framework for the prebiotic emergence of information and meaning. Building on Constructor Theory, we define information as a reproducible physical difference and meaning as a difference with stable functional consequences. Casimir-Lifshitz-coupled protocell clusters serve as a minimal model that exhibits reproducible attractors

  10. Masakazu Yoshimura, Zitang Sun, Yuiko Sakuma, Junji Otsuka

    Neural Architecture Search (NAS) aims to automatically discover high-performing deep neural network (DNN) architectures. However, conventional algorithm-driven NAS relies on carefully hand-crafted search spaces to ensure executability, which restricts open-ended exploration. Recent coding-based agentic approaches using large language models (LLMs) reduce man

  11. Biancamaria Sersante, Christos Georgiou, Nora Elisa Chisari

    Upcoming cosmological surveys will achieve increasingly precise constraints in cosmological parameter estimation. To guarantee the robustness of cosmological analyses, it is essential to account for and model systematic effects that can bias cosmological constraints, shifting the best fit parameters away from their fiducial values. It is possible to approxim

  12. Rui Guo, Ditza Auerbach, Yonina C. Eldar

    Quantitative speed-of-sound (SoS) and attenuation of tissues are closely related to pathology; however, conventional B-mode images are limited to qualitative visualization. Existing ultrasound full-waveform inversion (FWI) methods for quantitative SoS reconstruction are primarily developed under double-sided or ring-shaped arrays, which limits their applicab

  13. Suyan Dai, Chenxi Liu, Fazeng Li, Peican Lin

    3D object detection models trained in one server plays an important role in autonomous driving, robotics manipulation, and augmented reality scenarios. However, most existing methods face severe privacy concern when deployed on a multi-robot perception network to explore large-scale 3D scene. Meanwhile, it is highly challenging to employ conventional federat

  14. Chi Liu, Xin Chen, Xu Zhou, Fangbo Tu

    Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. However, they often suffer from performance degradation due to factors such as catastrophic forgetting during Supervised Fine-Tuning (SFT), quantization, and pruning. In this work, we introduce a performance recovery framework based on Self-Distillation Fine-

  15. Xiangkai Wang, Yun Zhao, Dongyi He, Qingling Xia

    Stroke patient cross-subject electroencephalography (EEG) decoding of motor imagery (MI) brain-computer interface (BCI) is essential for motor rehabilitation, yet lesion-related abnormal temporal dynamics and pronounced inter-patient heterogeneity often undermine generalization. Existing adaptation methods are easily misled by pathological slow-wave activity

  16. Liang Cheng, Zhenyu Lu

    In this paper, we provide a proof of Hamilton's extrinsic pinching theorem using the mean curvature flow approach.

  17. Miaoxuan Zhu, Yi Yu, Yuyang Li, Wei Li

    The quantification of uncertainty in prediction models is crucial for reliable decision-making, yet remains a significant challenge. Interval time series forecasting offers a principled solution to this problem by providing prediction intervals (PIs), which indicates the probability that the true value falls within the predicted range. We consider a recently

  18. Nikesh Lilani, Manus R. Visser

    We compute the efficiency of the reversible Stirling engine, with and without regeneration, for a broad class of working substances including Van der Waals fluids, quantum ideal gases (Bose and Fermi), Bose-Einstein condensates, thermal conformal field theories (CFTs), and holographic CFTs. Regeneration acts as an internal heat recycling mechanism that enhan

  19. Michael Massoth

    Casimir-Lifshitz forces generate an unavoidable, long-range attraction between protocells under prebiotically realistic conditions. This interaction stabilizes mesoscale clusters such as tetrahedra, octahedra, and 13-cell icosahedra. These highly symmetric assemblies act as persistent macrostates whose transitions remain reproducible despite microscopic nois

  20. Wai Man Si, Mingjie Li, Michael Backes, Yang Zhang

    As Large Language Models (LLMs) receive increasing attention and are being deployed across various domains, their potential risks, including generating harmful or biased content, producing unsupported claims, and exhibiting vulnerabilities to adversarial attacks, have drawn significant attention. To enable quick and low-cost adaptation, training-free methods

  21. He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang

    Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hyp

  22. David Berghaus

    We introduce EVIL (\textbf{EV}olving \textbf{I}nterpretable algorithms with \textbf{L}LMs), an approach that uses LLM-guided evolutionary search to discover simple, interpretable algorithms for dynamical systems inference. Rather than training neural networks on large datasets, EVIL evolves pure Python/NumPy programs that perform zero-shot, in-context infere

  23. Advait Sarkar

    Filter Babel is a thought experiment about a near future in which everything we read, watch, and even whom we "meet" is privately generated for each of us. If we each recede into a world of purely private experience, we may each develop a Wittgensteinian private language that remains intelligible to others only because an AI translator sits in the middle. Th

  24. Sarah Perez, Florian Doster, Hannah Menke, Ahmed ElSheikh

    Flow and transport in fractured geological media are strongly controlled by aperture heterogeneity and uncertainty in subsurface characterisation, yet most upscaling approaches rely on deterministic representations of fracture permeability. This study presents a scalable probabilistic workflow that bridges image-based fracture geometry and uncertainty-aware

  25. M. Tsaqif Wismadi, Oluwaleke Yusuf, Adil Rasheed, Yngve Karl Frøyen

    The strategic placement of bike-sharing infrastructure shapes urban accessibility and mobility outcomes. However, station-allocation approaches vary in their assumptions and decision logic. This study examines how alternative modelling paradigms prioritise urban space when applied to the same planning problem in Trondheim, Norway. We developed a unified anal

  26. Oluwaleke Yusuf, Shaira Tabassum

    Traffic simulations, essential for planning urban transit infrastructure interventions, require vehicle-category-specific origin-destination (OD) data. Existing data sources are imperfect: sparse tollbooth sensors provide accurate vehicle counts by category, while extensive mobility data from cellular network activity captures aggregated crowd movement, but

  27. Xiaolin Wen, Changlin Li, Manusha Karunathilaka, Can Liu

    Many expressive visualizations are shared online only as bitmap images, making them difficult to redesign or adapt to new data. Reusing such image-based visualizations requires substantial expertise and is often time-consuming, even for experienced visualization practitioners. Existing work on reproducing visualizations often relies on structured SVG or spec

  28. Wai Man Si, Mingjie Li, Michael Backes, Yang Zhang

    Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe behaviors inherited from pre-training. Current alignment methods like SFT and RLHF primarily encourage models to generate preferred responses, but do not explicitly remove the unsafe subnetworks that trigger har

  29. Nan Yang, Bahman Javadi, Rodrigo Neves Calheiros, David Boland

    Low Earth Orbit (LEO) mega-constellations extend the cloud-to-edge continuum into space, enabling satellite edge computing. However, Federated Learning (FL) in this environment is fundamentally energy-constrained due to dynamic inter-satellite connectivity, heterogeneous onboard computing hardware, and strict power budgets. We propose CroSatFL, a sustainable

  30. Kotaro Miyake, Yasuhiro Yamaguchi

    The nature of the $X(3872)$ and other exotic hadrons has been a subject of extensive investigation. While various theoretical models have been proposed, experimental evidence suggests that the $X(3872)$ may be a mixture state of a hadronic molecule and a $c\bar{c}$ core. In this work, we perform a systematic study of hidden-charm tetraquark candidates $X(386

  31. Zhiling Yan, Sicheng Chen, Tianyi Zhang, Nan Ying

    Segmentation is a critical task in computational pathology, as it identifies areas affected by disease or abnormal growth and is essential for diagnosis and treatment. However, acquiring high-quality pixel-level supervised segmentation data requires significant workload demands from experienced pathologists, limiting the application of deep learning. To over

  32. Pritesh Jha

    We present PIIBench, a unified benchmark corpus for Personally Identifiable Information (PII) detection in natural language text. Existing resources for PII detection are fragmented across domain-specific corpora with mutually incompatible annotation schemes, preventing systematic comparison of detection systems. We consolidate ten publicly available dataset

  33. Abhishek Sawaika, Durga Pritam Suggisetti, Udaya Parampalli, Rajkumar Buyya

    Learning with large-scale datasets and information-critical applications, such as in High Energy Physics (HEP), demands highly complex, large-scale models that are both robust and accurate. To tackle this issue and cater to the learning requirements, we envision using a federated learning framework with a quantum-enhanced model. Specifically, we design a hyb

  34. Taiyo Narita, Hideyuki Miyahara

    We introduce a novel characterization of phase transitions based on hypothesis testing. In our formulation, a phase transition is defined as the breakdown of statistical indistinguishability under vanishing parameter perturbations in the thermodynamic limit. This perspective provides a general, order-parameter-free framework that does not rely on model-speci

  35. Zhenggang Tang, Yuehao Wang, Yuchen Fan, Jun-Kun Chen

    Recent text-to-scene generation approaches largely reduced the manual efforts required to create 3D scenes. However, their focus is either to generate a scene layout or to generate objects, and few generate both. The generated scene layout is often simple even with LLM's help. Moreover, the generated scene is often inconsistent with the text input that conta

  36. Hürkan Şahin, Van Huyen Dang, Erdi Sayar, Alper Yegenoglu

    Reinforcement learning (RL) often struggles in real-world tasks with high-dimensional state spaces and long horizons, where sparse or fixed rewards severely slow down exploration and cause agents to get trapped in local optima. This paper presents a fuzzy logic based reward shaping method that integrates human intuition into RL reward design. By encoding exp

  37. Junjie Wen, Junlin He, Fei Ma, Jinqiang Cui

    Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existing representations struggle to jointly satisfy these requirements, and densely propagating pixel-wise semantics to 3D often results in subst

  38. Dongxin Guo, Jikun Wu, Siu Ming Yiu

    Spiking transformers achieve competitive accuracy with conventional transformers while offering $38$-$57\times$ energy efficiency on neuromorphic hardware, yet no theoretical framework guides their design. This paper establishes the first comprehensive expressivity theory for spiking self-attention. We prove that spiking attention with Leaky Integrate-and-Fi

  39. Zhixiang Yin, Changjun Gao, Yun-Long Zhang

    We investigate black holes with non-minimal couplings between the electromagnetic field and spacetime curvature, focusing on their event horizons, shadows, and photon rings. Such couplings can naturally arise from both classical effective field theories of gravity and quantum effects in curved spacetime. Starting from a general action with three independent

  40. Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao

    AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schr\"odinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy a

  41. Advait Sarkar

    Software adoption has traditionally been understood through instrumental lenses, such as usability, cost, security, and interoperability. We argue that a new, ideological dimension is reshaping adoption decisions: one we term digital patriotism, the individual counterpart to the state ideology of digital sovereignty. Through two studies, we trace this phenom

  42. Hongjia Chen, Chun-Hua Zhang, Zhongming Teng, Lei Du

    This paper considers the numerical solution of generalized Sylvester matrix equations, which arise in many scientific and engineering applications but remain challenging to solve efficiently, particularly when the coefficient matrices are general and the spectral radius of the associated operator is large but not greater than $1$. We propose a new iterative

  43. Timo Ziegler

    Classical simulation of quantum operations is essential for algorithm design, noise characterisation, and benchmarking of quantum hardware. The most general physically realisable operation can be described by a positive linear map acting on a hermitian operator, representing either a density matrix or an observable. Established simulators vectorise the densi

  44. Dongxin Guo, Jikun Wu, Siu Ming Yiu

    Early-exit neural networks enable adaptive computation by allowing confident predictions to exit at intermediate layers, achieving 2-8$\times$ inference speedup. Despite widespread deployment, their generalization properties lack theoretical understanding -- a gap explicitly identified in recent surveys. This paper establishes a unified PAC-Bayesian framewor

  45. Jorge Sánchez, Guadalupe García-Isla, Sandra Perez-Herrero, Beatriz Trenor

    Designing optimizers that remain effective under tight evaluation budgets is critical in expensive black-box settings such as cardiac digital twinning. We propose Frenetic Cat-inspired Particle Optimization (FCPO), a hybrid swarm method that couples particle swarm optimization-like dynamics with an explicit-state Markov switching controller to schedule explo

  46. Ankit Maloo

    We introduce the first version of KWBench (Knowledge Work Bench), a benchmark for unprompted problem recognition in large language models: can an LLM identify a professional scenario before attempting to solve it. Existing frontier benchmarks have saturated, and most knowledge-work evaluations to date reduce to extraction or task completion against a specifi

  47. Enrico Calloni, Annalisa Allocca, Antonino Chiummo, Rosario De Rosa

    We present a new differential mechanical gradiometer for the detection of low-frequency Gravitational Waves. The frequency range is 0.05 to 1 Hz, a frequency gap not covered either by future space-based detectors such as LISA or by ground-based observatories such as Einstein Telescope or Cosmic Explorer. The proposed detection principle is similar to antenna

  48. Peter Vamplew, Cameron Foale

    This research note identifies a previously overlooked distinction between multi-objective reinforcement learning (MORL), and more conventional single-objective reinforcement learning (RL). It has previously been noted that the optimal policy for an MORL agent with a non-linear utility function is required to be conditioned on both the current environmental s

  49. Jinlun Ye, Jiang Liao, Runhe Lai, Xinhua Lu

    Vision-language models (VLMs) such as CLIP exhibit strong Out-of-distribution (OOD) detection capabilities by aligning visual and textual representations. Recent CLIP-based test-time adaptation methods further improve detection performance by incorporating external OOD labels. However, such labels are finite and fixed, while the real OOD semantic space is in

  50. Lam T. Nguyen, Khanh Kieu

    We demonstrate the use of two-photon excitation for observing the ground state optically detected magnetic resonance (ODMR) of nitrogen-vacancy centers in diamonds at room temperature. An ultrafast femtosecond laser at 1040 nm was used for excitation, while fluorescence signal read out was achieved through a combination of a PMT and a lock-in amplifier. The

  51. Satoshi Suguira, Kam-Fung Cheung, Michael G. H. Bell, Hitomi Nakanishi

    Transit network design plays an important role in public transport. With the simplicity of spanning tree, this paper adopts the concept of spanning tree to help (re-)design a public transit network that addresses passenger utility by minimizing the total passenger-kilometers, which can be formulated as a mixed-integer optimization model. However, searching f

  52. Jingke Chen, Jingrui Zhong, Tazneen Hossain Tani, Zidong Su

    Despite the high accuracy of 'black box' deep learning models, drug discovery still relies on protein-ligand interaction principles and heuristics. To improve interpretability of protein-small molecule binding predictions, we developed the PWRules framework, which applies binding affinity data to identify privileged small molecule fragments and subsequently

  53. Joe Bond, Jacob Pake, Cristina David, Andrew McNutt

    \emph{Literate programming}, introduced by Knurth, interleaves code and prose so that a program can be read as both executable and explanatory text. We propose \emph{literate execution}, which inverts this relationship: rather than embedding code within a static narrative, we treat documentation -- and other expository elements such as visualisations -- as f

  54. Stein Andreas Bethuelsen, Frank Namugera

    We consider a general class of contact processes on $\mathbb{Z}^d$ with potentially long-range interactions. By adapting well established renormalization arguments to the long-range setting we extend by now classical results for finite-range processes to this more general setting. Particularly, we provide general conditions on the decay of the interactions u

  55. Yi-Lin Ge, Bing-Shu Hu, Ling-Yun Deng, Xiao-Ming Lu

    The geometry of quantum states has profound implications in quantum multiparameter estimation. While the Riemannian structure of quantum state space is well understood, the full understanding of the curvature structure of mixed quantum states is still an open problem. Inspired by the Yang-Mills action in non-Abelian gauge theory, we propose a scalar quantify

  56. David L. Condrey

    We introduce PoSME (Proof of Sequential Memory Execution), a cryptographic primitive that enforces sustained sequential computation via latency-bound pointer chasing over a mutable arena. Each step reads data-dependent addresses, writes a block whose value and causal hash are mutually dependent (symbiotic binding), and chains the result into a global transcr

  57. Xiang Xia, Wuyang Zhang, Jiazheng Liu, Cheng Yan

    Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive language generation due to their potential for parallel decoding and global refinement of the entire sequence. To unlock this potential, DLM inference must carefully balance generation quality and decoding speed. Recent block-wise DLM decoding methods improve this trad

  58. Yi Bing, Zheng Ran, Fu Jinyang, Liu Long

    Efficient and stable solution of partial differential equations (PDEs) is central to scientific and engineering applications, yet existing numerical solvers rely heavily on matrix based discretizations, while learning based methods require costly training and often suffer from limited generalization. In this work, we proposes a PDE energy driven framework th

  59. Sing-Hong Chan, Pochung Chen

    We present a general framework for extracting conformal data from critical two-dimensional classical lattice models using finite-size tensor-network flow. The central idea is to identify, from transfer-matrix spectra, a self-consistent finite-size window together with a crossover scale that separates the finite-size-scaling regime from the finite-entanglemen

  60. Qianchen Gong, Yingpeng Liu, Yan Zhang, Muhua Zheng

    Experimental evidence indicates that intracellular chloride concentration regulates the excitation and inhibition (EI) balance, yet the mechanisms by which activity-dependent chloride dynamics drive seizure evolution and stage transitions remain unclear. We present a conductance-based neuronal network in which EI balance emerges from chloride homeostasis via

  61. Qianshi Wang, Xilong Qu, Wenbin Pei, Nan Li

    Influence maximization (IM) is a fundamental problem in complex network analysis, with a wide range of real-world applications. To date, existing approaches to influential node identification in IM have predominantly relied on standard graphs, failing to capture higher-order intrinsic interactions embedded in many real-world systems. Hypergraphs can be emplo

  62. J. Manuel Cabrera, A. G. Andarcia Caballero, J. M. Paulin Fuentes

    Non-commutative electrodynamics obtained through the Seiberg-Witten map ceases to have equivalent action-level and equation-level realizations once fixed external currents are introduced, and in the action-level construction associated with the Banerjee current map the canonical location of this source-induced obstruction has remained unclear. Working in the

  63. Yongyun Chen, Qiusheng Gu, Junhui Fan, Dingrong Xiong

    A long-term variability study spanning a range of black hole mass systems, from microquasars hosting stellar-mass black holes to active galactic nuclei (AGNs) harboring supermassive black holes, provides new insights into the physics of relativistic jets. In this work, we investigate the optical variability of both jetted and nonjetted AGNs. We apply a stoch

  64. Hidetoshi Kawase, Toshihiro Ota

    In finite-width deep neural networks, the empirical kernel $G$ evolves stochastically across layers. We develop a collective kernel effective field theory (EFT) for pre-activation ResNets based on a $G$-only closure hierarchy and diagnose its finite validity window. Exploiting the exact conditional Gaussianity of residual increments, we derive an exact stoch

  65. Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Anh Tuan Luu

    Uncertainty estimation is a promising approach to detect hallucinations in large language models (LLMs). Recent approaches commonly depend on model internal states to estimate uncertainty. However, they suffer from strict assumptions on how hidden states should evolve across layers, and from information loss by solely focusing on last or mean tokens. To addr

  66. Oleg Solozobov

    Machine learning systems in fraud detection, credit scoring, and clinical risk assessment operate under delayed ground truth: outcome labels arrive days to months after the decision they evaluate. During this blind period, governance evidence degrades through mechanisms that neither drift detection methods nor governance frameworks adequately address. This p

  67. Yusheng Huang, Shuang Yang, Zhaojie Liu, Han Li

    Generative recommendation (GR) has emerged as a widely adopted paradigm in industrial sequential recommendation. Current GR systems follow a similar pipeline: tokenization for item indexing, next-token prediction as the training objective and auto-regressive decoding for next-item generation. However, existing GR research mainly focuses on architecture desig

  68. Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma

    Chromatic Correlation Clustering (CCC) extends Correlation Clustering by assigning semantic colors to edges and requiring each cluster to receive a single color label. Unlike standard CC, whose LP relaxation has integrality gap 2 on complete graphs and admits a 2.06-approximation, the analogous LP for CCC has a strict lower bound of 2.11, and the best known

  69. Dongguk Kim, Dongjo Kim, Jeongsu Bok, Beomkyu Kim

    Jet quenching in heavy-ion collisions probes parton energy loss in the quark--gluon plasma (QGP), but the extracted transport properties may not be universally constrained across centrality, beam energy, and observable class. In this work, we perform an analysis of the compatibility and predictive transferability of Bayesian constraints obtained from a six-p

  70. Yichen Xu, Yuanhang Liu, Chuhan Wang, Zihan Zhao

    While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first large-scale benchmark for evaluating MLLMs as automatic sports referees. Spanning 11 sports with 925 curated videos and 6,

  71. Siyuan Wang, Hanchen Gao, Guangming Zhu, Jiang Lu

    Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex structural contours, they lack color and texture, which text effectively provides despite omitting spatial contours. Motivated by the complementary nature of these modalities, we pr

  72. En Chen, Xi Chen

    Bubble N68 in the G35 complex shows clear cloud-cloud collision (CCC) signatures. Its semi-ring-like morphology harbors many significant massive star formation tracers: 6 HII regions, 4 6.7 GHz masers, 5 Midcourse Space Experiment sources, 9 radio peaks, and nearly 10 O/B-type stars. We also identified 163 young stellar objects (45 Class I, 5 Flat, 113 Class

  73. Lachlan Drake, Lawrence Ong, Duy T. Ngo

    Low earth orbit (LEO) satellites are a key technology to enable connectivity for rural and remote users. Communication satellites in LEO can provide coverage to much larger areas than terrestrial or aerial systems, while offering improved data rates when compared with geostationary systems. However, a major challenge with LEO satellite communications is the

  74. Takeshi Yoshimura, Valentijn Dymphnus van de Beek, Tatsuhiro Chiba

    Distributed LLM serving systems optimize per-request latency and throughput. However, under long-context workloads, inference accuracy becomes more variable. When incorrect responses trigger retries, accuracy directly translates into cumulative user-visible delay that is not captured by single-shot latency metrics. In this work, we argue that under long-cont

  75. Manaswini Piduguralla, Souvik Sarkar, Arunmoezhi Ramachandran, Sathya Peri

    Blockchain technology enhances transparency by maintaining a distributed ledger among mutually untrusting parties. Despite its advantages, scalability and availability remain critical bottlenecks that hinder widespread adoption. The increasing complexity of blockchain nodes further necessitates robust fault tolerance and high throughput to ensure seamless op

  76. Mohammad Mogharen Askarin, Jiankun Hu, Min Wang, Xuefei Yin

    Three-dimensional (3D) fingerprint recognition and identification offer several advantages over traditional two-dimensional (2D) recognition systems. The contactless nature of 3D fingerprints enhances hygiene and security, reducing the risk of contamination and spoofing. In addition to surface ridge and valley patterns, 3D fingerprints capture depth, curvatu

  77. Sicheng Chen, Chad Wong, Tianyi Zhang, Enhui Chai

    Whole Slide Image (WSI) analysis is pivotal in computational pathology, enabling cancer diagnosis by integrating morphological and architectural cues across magnifications. Multiple Instance Learning (MIL) serves as the standard framework for WSI analysis. Recently, Mamba has become a promising backbone for MIL, overtaking Transformers due to its efficiency

  78. Xidong Wu, Yukuan Zhang, Yuqiong Ji, Reza Shirkavand

    Large language model (LLM) routing has emerged as a critical strategy to balance model performance and cost-efficiency by dynamically selecting services from various model providers. However, LLM routing adds an intermediate layer between users and LLMs, creating new privacy risks to user data. These privacy risks have not been systematically studied. Althou

  79. Sankalp Gilda, Shlok Gilda

    Large language models exhibit systematic limitations in structured logical reasoning: they conflate hypothesis generation with verification, cannot distinguish conjecture from validated knowledge, and allow weak reasoning steps to propagate unchecked through inference chains. We present a symbolic reasoning scaffold that operationalizes Peirce's tripartite i

  80. Wenshuo Wang

    This position paper argues that large language model (LLM) reasoning should be studied as latent-state trajectory formation rather than as faithful surface chain-of-thought (CoT). This matters because claims about faithfulness, interpretability, reasoning benchmarks, and inference-time intervention all depend on what the field takes the primary object of rea

  81. Zehao Wang, Lanjun Wang

    Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling their deployment in high-stakes domains such as healthcare and education. While prior jailbreak attack studies have focused on the safety of final answers, little attention has been given to the safety of the reaso

  82. F. Zaidi, S. Akther, M. Sajjad Athar, S. K. Singh

    The charged current $\nu_\mu(\bar{\nu}_\mu)$-induced deep inelastic scattering (DIS) from an $^{40}\mathrm{Ar}$ target is studied using a microscopic framework that incorporates nuclear medium effects due to Fermi motion, binding energy, nucleon correlations, mesonic ($\pi$ and $\rho$) contributions, and nuclear shadowing and antishadowing across the relevan

  83. Mathumetha Palani, Kavya Puthumana, Ayantika Das, Ganapathy Krishnamurthi

    The advent of handheld fundus imaging devices has made ophthalmologic diagnosis and disease screening more accessible, efficient, and cost-effective. However, images captured from these setups often suffer from artifacts such as flash reflections, exposure variations, and motion-induced blur, which degrade image quality and hinder downstream analysis. While

  84. Sangbum Cho, Yuya Koda, Jung Hoon Lee

    For the genus-$4$ Heegaard surface in the $3$-sphere, we present a sufficient condition for a non-separating weak reducing pair to be separated by a reducing sphere for the surface. As a consequence, we reduce the connectivity problem in the reducing sphere complex for the surface to the problem of showing that any two vertices, whose representative reducing

  85. Huayu Dai, Xi-Hao Fang, Mitsutoshi Fujita, Song He

    In this paper, we analyze the double Wick rotation of a rotating BTZ black hole and the entanglement entropy. We derive the transition matrix dual to the double Wick-rotated BTZ black hole, which has the usual shape at an imaginary chemical potential. In the dual gravity side, the double Wick rotated BTZ black hole, which is obtained as a quotient, is equal

  86. Chuyang Wei, Maohang Gao, Zhixin Han, Kefei Chen

    Many high-stakes decisions depend on forecasts made before outcomes are known. In this future prediction setting, the central challenge is that public evidence evolves over time, while the main supervision signal arrives only after resolution: the realized outcome mainly assesses final correctness, offering only coarse guidance on what to track, what to veri

  87. Junguang Yao, Wenye Liu, Stjepan Picek, Yue Zheng

    Visual speaker recognition based on lip motion offers a silent, hands-free, and behavior-driven biometric solution that remains effective even when acoustic cues are unavailable. Compared to traditional methods that rely heavily on appearance-dependent representations, lip motion encodes subject-specific behavioral dynamics driven by consistent articulation

  88. Ki Sen Hung, Xi Yang, Chang Liu, Haoran Li

    A central goal of LLM alignment is to balance helpfulness with harmlessness, yet these objectives conflict when the same knowledge serves both legitimate and malicious purposes. This tension is amplified by context-sensitive alignment: we observe that domain-specific contexts (e.g., chemistry) selectively relax defenses for domain-relevant harmful knowledge,

  89. Chathranee Jayathilaka, Mark B. Flegg

    Biochemical signalling cascades transduce extracellular stimuli into cellular responses through sequences of discrete, node-to-node activations. While signal fidelity depends critically on local interaction kinetics, the mechanisms governing information propagation in realistic, highly variable kinetic contexts remain poorly understood. In this paper, we dev

  90. Jize Wang, Xuanxuan Liu, Yining Li, Songyang Zhang

    The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use benchmarks remain misaligned with real-world requirements, relying on AI-generated queries, dummy tools, and limited system-level coordination. To address this, we propose GTA-2, a h

  91. Tsao-Hsien Chen, Lingfei Yi

    We establish versions of Matsuki duality for loop groups. The main result is a bijection between symmetric loop group orbits and real polynomial loop group orbits on the affine Grassmannians or affine flag varieties. Along the way we obtain orbit parametrizations and make connections with vector bundles on real and twistor-$\mathbb P^1$ and Kottwitz sets .

  92. Enhui Chai, Sicheng Chen, Tianyi Zhang, Xingyu Li

    Pathological diagnosis is highly reliant on image analysis, where Regions of Interest (ROIs) serve as the primary basis for diagnostic evidence, while whole-slide image (WSI)-level tasks primarily capture aggregated patterns. To extract these critical morphological features, ROI-level Foundation Models (FMs) based on Vision Transformers (ViTs) and large-scal

  93. Tianle Liang, Yifu Chen, Shengpeng Ji, Yijun Chen

    Recent end-to-end spoken dialogue models enable natural interaction. However, as user demands become increasingly complex, models that rely solely on conversational abilities often struggle to cope. Incorporating agentic capabilities is therefore essential: by enabling tool use, these models can extend their knowledge boundaries and better solve real-world t

  94. Chenyi Huang, Haoting Zhang, Jingxu Xu, Zeyu Zheng

    Agent \texttt{skills} are structured collections of instructions, tools, and supporting resources that help large language model (LLM) agents perform particular classes of tasks. Empirical evidence shows that the design of \texttt{skills} can materially affect agent task performance, yet systematically optimizing \texttt{skills} remains challenging. Since a

  95. Geunyoung Jung, Soohong Kim, Inseok Kong, Jiyoung Jung

    The advent of deep neural networks has led to remarkable progress in 3D point cloud recognition, but they remain vulnerable to adversarial attacks. Although various defense methods have been studied, they suffer from a trade-off between robustness and transferability. We propose Adversarial Point Counterattack (APC) to achieve both simultaneously. APC is a l

  96. Ruxin Ding, Jianfeng Ren, Heng Yu, Jiawei Li

    Spatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle this, STLBP features are often extracted on three orthogonal planes, which sacrifice inter-plane correlation. In this work, we propose a Locality-Preserving Pixel-Difference Hashing (LP$^{2}$DH) framework that

  97. Zijun Wang, Haoqin Tu, Weidong Zhou, Yiyang Zhou

    Everyday tasks come with a target, and pretraining models around this target is what turns them into experts. In this paper, we study target-oriented language model (LM) pretraining by introducing Neuron-Activated Graph Ranking (NAG-based Ranking), a training-free and interpretable framework for target pretraining data selection. Rather than using black-box

  98. Xiaoyu Yang, En Yu, Wei Duan, Jie Lu

    Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with complex human values and domain-specific requirements. Nevertheless, current research primarily focuses on mitigating exogenous distribution shifts arising from data-centric factors, the non-stationarity inherent i

  99. Haojie Li, Junwei Du, Guanfeng Liu, Feng Jiang

    Disentanglement techniques used in collaborative filtering uncover interaction intents between nodes, improving the interpretability of node representations and enhancing recommendation performance. However, existing disentanglement methods still face two problems. First, they focus on local structural features derived from direct node interactions and overl

  100. Catherine Ning, Yu Ma, Cindy Beini Wang, Sean McMahon

    Left ventricular ejection fraction (LVEF) assessment depends on echocardiography, limiting access in primary care and resource-constrained settings. We developed a multimodal machine-learning framework that combines engineered 12-lead ECG timeseries features with structured EHR variables to classify LVEF into four clinically used strata: normal (>50%), mildl