Skip to content

November 2025 arXiv papers — page 104

Showing 10,30110,400 of 22,271 papers

  1. Lorenzo Valentini, Diego Forlivesi, Andrea Talarico, Marco Chiani

    Hardware-friendly quantum low-density parity-check (QLDPC) decoders are commonly built upon belief propagation (BP) processing. Yet, quantum degeneracy often prevents BP from achieving reliable convergence. To overcome this fundamental limitation, we propose the restart belief (RB) decoder, an iterative BP-based algorithm inspired by branch-and-bound optimiz

  2. Samuel F. Lewin, Alexis K. Kaminski, Arun Balakrishna, Miles M. P. Couchman

    The behaviour of internal waves propagating in a background shear flow is studied in the case where the direction of shear is orthogonal to gravity. Ray-tracing theory is used to predict properties of the wave state at locations where instability occurs. Local wave energy growth is found to result from two distinct mechanisms: an increase in wave steepness d

  3. Tobias Kaltenmark, Chris Nill, Christian Groß, Igor Lesanovsky

    Lattice spin models featuring kinetic constraints constitute a paradigmatic setting for the investigation of glassiness and localization phenomena. The intricate dynamical behavior of these systems is a result of the dramatically reduced connectivity between many-body configurations. This truncation of transition pathways often leads to a fragmentation of th

  4. Zihan Li, Tengfei Wang, Wentian Gan, Hao Zhan

    Lightweight building surface models are crucial for digital city, navigation, and fast geospatial analytics, yet conventional multi-view geometry pipelines remain cumbersome and quality-sensitive due to their reliance on dense reconstruction, meshing, and subsequent simplification. This work presents SF-Recon, a method that directly reconstructs lightweight

  5. Noam Tsfaty, Avishai Weizman, Liav Cohen, Moshe Tshuva

    We address the challenge of detecting rare and diverse anomalies in surveillance videos using only video-level supervision. Our dual-backbone framework combines convolutional and transformer representations through top-k pooling, achieving 90.7% area under the curve (AUC) on the UCF-Crime dataset.

  6. Tetsuya Kaji, Elena Manresa

    We revisit the saving behavior of elderly singles using an adversarial structural estimation framework by Kaji, Manresa and Pouliot (2023). The method bridges the simulated method of moments (SMM) and maximum-likelihood estimation by embedding a flexible discriminator, implemented as a neural network, that adaptively selects the most informative features of

  7. Taras Sereda, Tom St. John, Burak Bartan, Natalie Serrino

    GPU kernels are critical for ML performance but difficult to optimize across diverse accelerators. We present KForge, a platform-agnostic framework built on two collaborative LLM-based agents: a generation agent that produces and iteratively refines programs through compilation and correctness feedback, and a performance analysis agent that interprets profil

  8. Zhe Sun, Yujun Cai, Jiayu Yao, Yiwei Wang

    Large Audio-Language Models (LALMs) have recently shown impressive progress in speech recognition, audio captioning, and auditory question answering. Yet, whether these models can perceive spatial dynamics, particularly the motion of sound sources, remains unclear. In this work, we uncover a systematic motion perception deficit in current ALLMs. To investiga

  9. Zeyang Sun, Xidong Mu, Shuai Han, Sai Xu

    This paper investigates a pinching-antenna (PA)-enabled cognitive radio network, where both the primary transmitter (PT) and secondary transmitter (ST) are equipped with a single waveguide and multiple PAs to facilitate simultaneous spectrum sharing. Under a general Ricean fading channel model, a closed-form analytical expression for the average spectral eff

  10. Manish Patel, Subhajit Paul, Debasish Chaudhuri

    We study the dynamics of inertial active particles in a one-dimensional chain with harmonic nearest-neighbor interactions, highlighting the interplay of persistence, interaction, and inertial timescales. Using a Green's function approach, we derive the mean-squared displacement (MSD) and mean-squared change in velocity (MSCV), revealing multiple crossovers b

  11. Lingfeng Zhang, Yuchen Zhang, Hongsheng Li, Haoxiang Fu

    Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks. However, the spatial intelligence capabilities of existing VLMs in UAV scenarios remain largely unexplored, raising concerns about their effectiveness in navigating and interpreting dynamic en

  12. Sukanya Randhawa, Guntaj Randhawa, Clemens Langer, Francis Andorful

    Resilient road infrastructure is a cornerstone of the UN Sustainable Development Goals. Yet a primary indicator of network functionality and resilience is critically lacking: a comprehensive global baseline of road surface information. Here, we overcome this gap by applying a deep learning framework to a global mosaic of Planetscope satellite imagery from 20

  13. F. Yang, L. Q. Chen

    Starting from the purely microscopic model, we go beyond conventional mean-field theory and develop a self-consistent microscopic thermodynamic framework for disordered 2D superconductors. It incorporates the fermionic Bogoliubov quasiparticles, bosonic Nambu-Goldstone (NG) quantum and thermal phase fluctuations in the presence of long-range Coulomb interact

  14. Dancheng Lu, Zexin Wang, Guangjun Zhu

    The homological shift algebra and the projective dimension function of complementary edge ideals are investigated. Let $G$ be a connected graph, and let $I$ be its complementary edge ideal. For bipartite graphs $G$, we show that the projective dimension of $I^s$ increases strictly with $s$ until reaching its maximum value. For trees and cycles, explicit expr

  15. José J. Gil

    We show that the antisymmetric Mueller generator provides a universal algebraic kernel for geometric phase in classical polarization optics and in quantum two-level systems. For any ideal retarder, the antisymmetric 3x3 block of its Mueller matrix (the antisymmetric generator of the adjoint SU(2) action on the Stokes vector) encodes the angular-velocity vect

  16. Yu-Hao Zhang, Liang-Duan Liu, Ze-Xin Du, Guang-Lei Wu

    We present TransFit-CSM, a fast and physically consistent framework for modeling interaction-powered transients. The method self-consistently couples the ejecta circumstellar medium (CSM) shock dynamics to radiative diffusion from a moving heating boundary tied to the shocks, so that both the photon escape path and the effective diffusion time evolve with ra

  17. Keshav Gupta, Akshat Sanghvi, Shreyas Reddy Palley, Astitva Srivastava

    3D Gaussian Splatting has emerged as a transformative technique in novel view synthesis, primarily due to its high rendering speed and photorealistic fidelity. However, its memory footprint scales rapidly with scene complexity, often reaching several gigabytes. Existing methods address this issue by introducing compression strategies that exploit primitive-l

  18. Yoonhak Nam, Kazuyuki Sekizawa

    Old, thermally bright neutron stars imply internal heating at late times. Among candidate mechanisms, vortex creep heating (VCH) provides a robust link between spin-down and frictional dissipation in the pinned inner-crust superfluid, yet its interplay with fast DUrca cooling in massive stars remains insufficiently explored. We (i) implement VCH in our cooli

  19. Jack B. Coughlin, Archis Joglekar, Jonathan Brodrick, Alexander Lavin

    This work presents a case study of a heterogeneous multiphysics solver from the nuclear fusion domain. At the macroscopic scale, an auto-differentiable ODE solver in JAX computes the evolution of the pulsed power circuit and bulk plasma parameters for a compressing Z Pinch. The ODE solver requires a closure for the impedance of the plasma load obtained via r

  20. Junlong Li, Huaiyuan Xu, Sijie Cheng, Kejun Wu

    Driven by recent advances in vision-language models (VLMs) and egocentric perception research, the emerging topic of an egocentric procedural AI assistant (EgoProceAssist) is introduced to step-by-step support daily procedural tasks in a first-person view. In this paper, we start by identifying three core tasks in EgoProceAssist: egocentric procedural error

  21. Amit Shivam, Kiran Kumari, Fernando A. C. C. Fontes

    This paper proposes a hybrid-gain finite-time sliding-mode control (HG-FTSMC) strategy for a class of perturbed nonlinear systems. The controller combines a finite-time reaching law that drives the sliding variable to a predefined boundary layer with an inner mixed-power or exponential law that guarantees rapid convergence within the layer while maintaining

  22. Yushuo Zheng, Jiangyong Ying, Huiyu Duan, Chunyi Li

    Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, \textit{etc}. To bridge this gap, we introduce \textbf{G

  23. V. S. Ryumshin, Yu. D. Panov, V. A. Ulitko, A. S. Moskvin

    The results of numerical simulation using a modified Monte Carlo method with a thermostat algorithm for a pseudospin model of orthonickelates are presented. Temperature phase diagrams are constructed for various degrees of filling and for various parameters of the model, and the effect of local correlations on the critical temperatures of the model orthonick

  24. Filippo Cenacchi, Longbing Cao, Mitchell McEwan, Deborah Richards

    We target passive dementia screening from short camera-facing talking head video, developing a facial temporal micro dynamics analysis for language free detection of early neuro cognitive change. This enables unscripted, in the wild video analysis at scale to capture natural facial behaviors, transferrable across devices, topics, and cultures without active

  25. Spyros Tserkis, Muhammad Umer, Dimitris G. Angelakis

    The increasing depth of quantum circuits presents a major limitation for the execution of quantum algorithms, as the limited coherence time of physical qubits leads to noise that manifests as errors during computation. In this work, we focus on CNOT ladder circuits, which find applications in several quantum computing tasks, including the preparation of GHZ

  26. Shalini Maiti, Amar Budhiraja, Bhavul Gauri, Gaurav Chaurasia

    Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but their training remains resource- and time-intensive, requiring massive compute power and careful orchestration of training procedures. Model souping-the practice of averaging weights from multiple models of the same architecture-has emerged as a promising pre-

  27. Mordechai Guri

    This paper introduces the Pico-Cloud, a micro-edge cloud architecture built on ultra-minimal hardware platforms such as the Raspberry Pi Zero and comparable single-board computers. The Pico-Cloud delivers container-based virtualization, service discovery, and lightweight orchestration directly at the device layer, enabling local operation with low latency an

  28. Huanxin Chen, Chun Xia, Hechao Chen

    Solar prominences usually have a horizontally elongated body with many feet extending to the solar surface, resembling a multi-arch bridge with many bridge piers. The basic mechanism by which solar prominences acquire these common structures during their evolution, however, remains an unresolved question. For the first time, our three-dimensional magneto-fri

  29. Thanh Nguyen

    This paper develops and empirically evaluates a Sharpe-driven stock selection and liquidity-constrained portfolio optimization framework designed for the Chinese equity market. The proposed methodology integrates three sequential stages: Sharpe-ratio-based universe selection, liquidity-adjusted mean-variance optimization, and multi-layered risk management im

  30. Aleksandar Stanković, Dejan Lisica

    We present reproducible, edge-aware baselines for ogbn-proteins in PyTorch Geometric (PyG). We study two system choices that dominate practice: (i) how 8-dimensional edge evidence is aggregated into node inputs, and (ii) how edges are used inside message passing. Our strongest baseline is GraphSAGE with sum-based edge-to-node features. We compare LayerNorm (

  31. Yu Wen, Shuyong Gao, Shuping Zhang, Miao Huang

    Referring camouflaged object detection (Ref-COD) aims to identify hidden objects by incorporating reference information such as images and text descriptions. Previous research has transformed reference images with salient objects into one-dimensional prompts, yielding significant results. We explore ways to enhance performance through multi-context fusion of

  32. Fuyao Zhang, Jiaming Zhang, Che Wang, Xiongtao Sun

    The reliance of mobile GUI agents on Multimodal Large Language Models (MLLMs) introduces a severe privacy vulnerability: screenshots containing Personally Identifiable Information (PII) are often sent to untrusted, third-party routers. These routers can exploit their own MLLMs to mine this data, violating user privacy. Existing privacy perturbations fail the

  33. Zhuo Chen, Zhongqun Zhang, Yihua Cheng, Ales Leonardis

    Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, specifically the contact force, leading to reduced stability of

  34. Qin Guo, Haonan Tong, Sihua Wang, Peiyuan Si

    This study proposes a novel approach to ensure the security of textual data transmission in a semantic communication system. In the proposed system, a sender transmits textual information to a receiver, while a potential eavesdropper attempts to intercept the information. At the sender side, the text is initially preprocessed, where each sentence is annotate

  35. Matt Luckcuck, Maike Schwammberger, Mengwei Xu

    This EPTCS volume contains the papers from the Seventh International Workshop on Formal Methods for Autonomous Systems (FMAS 2025), which was held between the 17th and 19th of November 2025. The goal of the FMAS workshop series is to bring together leading researchers who are using formal methods to tackle the unique challenges that autonomous systems presen

  36. Nadav Bojan Sellam, Meital Bojan, Paul Schanda, Alex Bronstein

    Accurate protein structures are essential for understanding biological function, yet incorporating experimental data into protein generative models remains a major challenge. Most predictors of experimental observables are non-differentiable, making them incompatible with gradient-based conditional sampling. This is especially limiting in nuclear magnetic re

  37. Xiaoqi Han, Ru Li, Ran Yi, Hongye Tan

    Multimodal Model Editing (MMED) aims to correct erroneous knowledge in multimodal models. Existing evaluation methods, adapted from textual model editing, overstate success by relying on low-similarity or random inputs, obscure overfitting. We propose a comprehensive locality evaluation framework, covering three key dimensions: random-image locality, no-imag

  38. Junjie Wu, Guohong Fu

    Multimodal misinformation floods on various social media, and continues to evolve in the era of AI-generated content (AIGC). The emerged misinformation with low creation cost and high deception poses significant threats to society. While recent studies leverage general-purpose multimodal large language models (MLLMs) to achieve remarkable results in detectio

  39. Nadine du Toit, Kristian K. Muller-Nedebock, Giuseppe Pellicane

    This paper extends a field-theoretical dynamical networking formalism for mesoscopic polymer dynamics to explicitly include dedicated cross-linker particles. Cross-linkers are represented within a Martin-Siggia-Rose generating functional and reversibly coupled to polymers through Gaussian networking fields, enabling an approximation scheme that reduces their

  40. Arka Pal, Teo Kitanovski, Arthur Liang, Akilesh Potti

    Large language models (LLMs) are increasingly deployed in agentic and multi-turn workflows where they are tasked to perform actions of significant consequence. In order to deploy them reliably and manage risky outcomes in these settings, it is helpful to access model uncertainty estimates. However, confidence elicitation methods for LLMs are typically not ev

  41. Thanh Nguyen

    Cryptocurrency trading has attracted tremendous attention from both retail and institutional investors. However, most traders fail to scale their assets under management due to fragile strategies that collapse during adverse markets. The primary causes are oversized leverage, speculative position sizing, and the absence of robust risk management or hedging m

  42. Patrick Parschan, Charlott Jakob

    This article presents the first systematic review of unsupervised and semi-supervised computational text-based ideal point estimation (CT-IPE) algorithms, methods designed to infer latent political positions from textual data. These algorithms are widely used in political science, communication, computational social science, and computer science to estimate

  43. Alan G. Paredes Cetina, Kaouther Benguessoum, Raoni Lourenço, Sylvain Kubler

    Recent advances in deep learning have improved multivariate time series (MTS) classification and regression by capturing complex patterns, but their lack of transparency hinders decision-making. Explainable AI (XAI) methods offer partial insights, yet often fall short of conveying the full decision space. Counterfactual Explanations (CE) provide a promising

  44. Huatian Hu, Xin Shu, Zhiwei Hu, Di Zheng

    Pushing nanoscale optical confinement to its ultimate limits defines the regime of nano-cavity quantum electrodynamics (nano-cQED), where light--matter interactions approach the fundamental quantum limits of individual atoms, e.g., picocavities. However, realizing such extreme confinement in a stable and controllable manner remains a key challenge. Here, we

  45. Vaidehi Nattoja, Tobias Toll

    Ultra-peripheral heavy-ion collisions (UPCs) provide a distinct environment for high-energy QCD research, focusing on the production of vector mesons. This proceeding details recent advancements in the Sar$t$re Monte Carlo event generator, a dipole model-based tool, to better describe UPC data. We present the incorporation of the full photon flux, accounting

  46. Boris Kriuk

    Traditional gradient boosting algorithms employ static tree structures with fixed splitting criteria that remain unchanged throughout training, limiting their ability to adapt to evolving gradient distributions and problem-specific characteristics across different learning stages. This work introduces MorphBoost, a new gradient boosting framework featuring s

  47. Jun Sashihara, Yukihisa Fujita, Kota Nakamura, Masahiro Kuwahara

    Data marketplaces, which mediate the purchase and exchange of data from third parties, have attracted growing attention for reducing the cost and effort of data collection while enabling the trading of diverse datasets. However, a systematic understanding of the interactions between market participants, data, and regulations remains limited. To address this

  48. Malek Al Abed, Sebiha Demir, Anne Groteklaes, Elodie Germani

    Portable ultra-low-field MRI (uLF-MRI, 0.064 T) offers accessible neuroimaging for neonatal care but suffers from low signal-to-noise ratio and poor diagnostic quality compared to high-field (HF) MRI. We propose MRIQT, a 3D conditional diffusion framework for image quality transfer (IQT) from uLF to HF MRI. MRIQT combines realistic K-space degradation for ph

  49. Rion Shimazu, Suguru Endo, Shigeo Hakkaku, Shinobu Saito

    Quantum error mitigation (QEM) has been proposed as a class of hardware-friendly error suppression techniques. While QEM has been primarily studied for mitigating errors in the estimation of expectation values of observables, recent works have explored its application to estimating noiseless probability distributions. In this work, we propose two protocols t

  50. Petar Orlić

    Let $N$ be a positive integer. For every $d\mid N$ such that $(d,N/d)=1$ there exists an Atkin-Lehner involution $w_d$ of the modular curve $X_0(N)$. Let $B(N)$ be the group of all such involutions. In this paper we determine all $\mathbb C$ and $\mathbb Q$-tetragonal quotient curves $X_0(N)/W_N$, where $W_N\subseteq B(N)$ such that $4\leq|W_N|\leq 2^{\omega

  51. Mary Chriselda Antony Oliver, Michael Roberts, Carola-Bibiane Schönlieb, Matthew Thorpe

    The manifold hypothesis posits that high-dimensional data typically resides on low-dimensional sub spaces. In this paper, we assume manifold hypothesis to investigate graph-based semi-supervised learning methods. In particular, we examine Laplace Learning in the Wasserstein space, extending the classical notion of graph-based semi-supervised learning algorit

  52. Xavier R. Advincula, Kara D. Fong, Yongkang Wang, Christoph Schran

    Hydrophobic solid-water interfaces underpin processes in nanofluidics, electrochemistry, and energy technologies. Microscopic insights into these systems are often inferred from our understanding of the air-water interface, which is assumed to exhibit similar behavior. Here, we challenge this paradigm by combining heterodyne-detected vibrational sum-frequenc

  53. I. Asiáin

    Throughout this thesis, we investigate how effective field theories, combined with unitarization techniques, can be used to explore physics beyond the Standard Model, with particular emphasis on the dynamical origin of electroweak symmetry breaking. Since effective theories often produce amplitudes that violate unitarity at high energies, restoring unitarity

  54. Michele Persiani, Thomas Hellstrom

    When a robot is asked to verbalize its plan it can do it in many ways. For example, a seemingly natural strategy is incremental, where the robot verbalizes its planned actions in plan order. However, an important aspect of this type of strategy is that it misses considerations on what is effectively informative to communicate, because not considering what th

  55. Tyler Loakman, Joseph James, Chenghua Lin

    With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of vision and language. In this work, we benchmark the extent to which VLMs are able to act as highly-trained phoneticians, interpreting spectrograms and waveforms of speech. To do t

  56. Devina Mohan, Anna M. M. Scaife

    Bayesian neural networks (BNNs) are most commonly optimised with first-order optimisers such as stochastic gradient descent. However, when optimising for parameters of probabilistic models, incorporating second order information during optimisation can lead to a more direct path in the distribution space and faster convergence. In this work we examine whethe

  57. Yuxiang Zhang, Zhengxu Yu, Weihang Pan, Zhongming Jin

    Emerging reasoning LLMs such as OpenAI-o1 and DeepSeek-R1 have achieved strong performance on complex reasoning tasks by generating long chain-of-thought (CoT) traces. However, these long CoTs result in increased token usage, leading to higher inference latency and memory consumption. As a result, balancing accuracy and reasoning efficiency has become essent

  58. Qida Tan, Hongyu Yang, Wenchao Du

    Appearance-based gaze estimation, aiming to predict accurate 3D gaze direction from a single facial image, has made promising progress in recent years. However, most methods suffer significant performance degradation in cross-domain evaluation due to interference from gaze-irrelevant factors, such as expressions, wearables, and image quality. To alleviate th

  59. Mohamed Salem, Inyoung Kim

    The transformer architecture has demonstrated strong performance in classification tasks involving structured and high-dimensional data. However, its success often hinges on large- scale training data and careful regularization to prevent overfitting. In this paper, we intro- duce a novel likelihood-guided variational Ising-based regularization framework for

  60. Zachary J. Hoelscher, Kelly Holley-Bockelmann, Akaxia Cruz, N. Nicole Sanchez

    Though the nature of dark matter remains elusive, two models have come to prominence with testable predictions: cold dark matter (CDM) and self-interacting dark matter (SIDM). While CDM remains the widely accepted model, SIDM was introduced to potentially help resolve the discrepancies between the predictions of the CDM model and observational data, in parti

  61. Satvik Dixit, Koichi Saito, Zhi Zhong, Yuki Mitsufuji

    Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires generating audio that is both semantically aligned with visible events and temporally aligned with their timing. Yet, there is a

  62. Xiaodong He, Xiao Wang, Jianda Wu

    We present a hybrid lattice Hamiltonian truncation method that integrates the numerical renormalization group (NRG) with a truncated lattice integrable spectrum. The technique is tailored for generic deformations of integrable lattice models, where the NRG enables a controlled incorporation of high-energy states. The method extends the basis set more effecti

  63. Yuchen Bi, Jintian Zhu

    In the spin case, we can establish a mass-capacity inequality for generalized asymptotically flat manifolds $(M,g,E)$ with nonnegative scalar curvature, where the equality implies that $(M,g)$ is harmonically conformal to $\mathbb R^n\setminus S$ for a closed bounded subset $S$ of $\mathbb R^n$ with Hausdorff dimension no greater than $\frac{n-2}{2}$.

  64. Guillaume Infantes, Stéphanie Roussel, Antoine Jacquet, Emmanuel Benazera

    The Resource-Constrained Project Scheduling Problem (RCPSP) is a classical scheduling problem that has received significant attention due to of its numerous applications in industry. However, in practice, task durations are subject to uncertainty that must be considered in order to propose resilient scheduling. In this paper, we address the RCPSP variant wit

  65. Muriel Zoë Stiefel, Paolo Massa, Alessia Guidetti, Marina Battaglia

    Solar hard X-ray observations provide diagnostics of the hottest plasmas and of nonthermal electron populations present during solar flares and coronal mass ejections. HXR images of specific energy ranges often contain overlapping contributions of these components, complicating their interpretation. This is even more challenging as HXR imagers generally use

  66. Bao-Duy Le, Dinh-Thi Nguyen

    We study vortex patterns of a two-dimensional Bose-Einstein condensate rotating close to the centrifugal limit, treating the two signs of the contact interaction with the method each requires: for repulsion, a GPU-accelerated variational minimization with exact projection onto the Lowest Landau Level (LLL); for attraction, imaginary-time evolution of the ful

  67. Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang

    Despite recent advancements in 3D-text cross-modal alignment, existing state-of-the-art methods still struggle to align fine-grained textual semantics with detailed geometric structures, and their alignment performance degrades significantly when scaling to large-scale 3D databases. To overcome this limitation, we introduce 3DAlign-DAER, a unified framework

  68. Tushar Jogi

    Modeling microstructural evolution at large strains requires mechanical formulations that remain thermodynamically consistent while capturing significant lattice rotations and transformation-induced stresses. However, most existing finite-strain microelasticity and phase-field approaches apply macroscopic boundary conditions heuristically, preventing proper

  69. Yonghui Yu, Jiahang Cai, Xun Wang, Wenwu Yang

    Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single person pose estimation. This design relies on heuristic operations such as detection, RoI cropping, and non-maximum suppression (NMS), limiting both accuracy and efficiency. In this paper, we

  70. Yihuai Zhang, Huan Yu

    Modeling and congestion mitigation of mixed-autonomy traffic systems consisting of human-driven vehicles (HVs) and autonomous vehicles (AVs) have become increasingly critical with the rapid development of autonomous driving technology. This paper develops an event-triggered control (ETC) framework for mitigating congestion in such systems, which are modeled

  71. Pavel Arkhipov, Vladimir Kolmogorov

    Greedy minimum weight spanning tree packings have proven to be useful in connectivity-related problems. We study the process of greedy minimum weight base packings in general matroids and explore its applications. For general matroids, we observe two characterizations of the limit of the base packings (``the vector of ideal loads''). Specialized to graphic m

  72. Junhee Lee, ChaeBeen Bang, MyoungChul Kim, MyeongAh Cho

    Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by treating all abnormal events as a single category, overlooking the diverse semantic and temporal characteristics intrinsic to

  73. Piyush Paliwal, Aftab Alam

    Lattice thermal conductivity ($\kappa_L$) is a key physical property governing heat transport in solids, with direct relevance to thermoelectrics, thermal barrier coatings, and heat management applications. However, while experimental determination of $\kappa_L$ is challenging, its theoretical calculation via ab initio methods particularly using density func

  74. Hao Hu, Yifan Feng, Ruoxue Li, Rundong Xue

    Retrieval-Augmented Generation (RAG) enhances the response quality and domain-specific performance of large language models (LLMs) by incorporating external knowledge to combat hallucinations. In recent research, graph structures have been integrated into RAG to enhance the capture of semantic relations between entities. However, it primarily focuses on low-

  75. Mingzhi Xiao, Yuki Takayama

    This study examines how congestion pricing shapes housing market outcomes and spatial equity in New York City. Using high-frequency sales and rental data and a combination of propensity score matching difference-in-differences, geographic regression discontinuity, and event study designs, the analysis identifies distinct short-run adjustment patterns trigger

  76. Natalie Neumeyer, Jan Rabe, Mathias Trabs

    In a multivariate nonparametric regression setting we construct explicit asymptotic uniform confidence bands for centered purely random forests. Since the most popular example in this class of random forests, namely the uniformly centered purely random forests, is well known to suffer from suboptimal rates, we propose a new type of purely random forests, cal

  77. Zhixin Ou, Peng Liang, Jianchen Han, Baihui Liu

    Dynamic sequences with varying lengths have been widely used in the training of Transformer-based large language models (LLMs). However, current training frameworks adopt a pre-defined static parallel strategy for these sequences, causing neither communication-parallelization cancellation on short sequences nor out-of-memory on long sequences. To mitigate th

  78. Alberto Gomez, Jorge Oliveira, Ramon Casero, Agis Chartsias

    Ultrasound (US) machines display images on a built-in monitor, but routine transfer to hospital systems relies on DICOM. We propose a fully automatic method to generate labeled data that can be used to train a screen detector model, and a pipeline to use that model to extract and rectify the US image from a photograph of the monitor, without any need for hum

  79. Hung-Ming Huang, Yu-Hsin Yang, Fu-Chieh Chang, Yun-Chia Hsu

    As IC design grows more complex, automating comprehension and documentation of RTL code has become increasingly important. Engineers currently should manually interpret existing RTL code and write specifications, a slow and error-prone process. Although LLMs have been studied for generating RTL from specifications, automated specification generation remains

  80. Robert Turnbull

    Application of phylogenetic methods to textual traditions has traditionally treated all changes as equivalent even though it is widely recognized that certain types of variants were more likely to be introduced than others. While it is possible to give weights to certain changes using a maximum parsimony evaluation criterion, it is difficult to state a prior

  81. Vincent Guillemet, Michael Unser

    The sampling of functions of bounded variation (BV) is a long-standing problem in op- timization. The ability to sample such functions has relevance in the field of variational inverse problems, where the standard theory fails to guarantee the mere existence of solutions when the loss functional involves samples of BV functions. In this paper, we prove the c

  82. Soyul Lee, Seungmin Baek, Dongbo Min

    Monocular 3D object detection is a cost-effective solution for applications like autonomous driving and robotics, but remains fundamentally ill-posed due to inherently ambiguous depth cues. Recent DETR-based methods attempt to mitigate this through global attention and auxiliary depth prediction, yet they still struggle with inaccurate depth estimates. Moreo

  83. Jiangwei Long, Zihui Liu, Yizhi Li, Jianxin Zhong

    We present a systematic numerical construction of a universal quantum gate set for topological quantum computation based on the non-semisimple Ising anyons model. By employing a Genetic Algorithm-enhanced Solovay-Kitaev Algorithm (GA-enhanced SKA), we achieve high-fidelity approximations of standard single-qubit gates (Hadamard H-gate and phase T-gate) with

  84. Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang

    Multi-agent systems (MAS) built on large language models (LLMs) often suffer from inefficient "free-for-all" communication, leading to exponential token costs and low signal-to-noise ratios that hinder their practical deployment. We challenge the notion that more communication is always beneficial, hypothesizing instead that the core issue is the absence of

  85. Yantong Liu, Junjie Wu, Lingling Lao

    Color codes present distinct advantages for fault-tolerant quantum computing, such as high encoding rates and the transversal implementation of Clifford gates. However, existing matching-based decoders for the color codes such as the restricted decoder (Kubica and Delfosse, 2023), suffer from limited decoding performance. Inspired by the global decoding insi

  86. Ying Jiang, Jiayin Lu, Yunuo Chen, Yumeng He

    Painting embodies a unique form of visual storytelling, where the creation process is as significant as the final artwork. Although recent advances in generative models have enabled visually compelling painting synthesis, most existing methods focus solely on final image generation or patch-based process simulation, lacking explicit stroke structure and fail

  87. Haoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang

    Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning-the ability to comprehend object locations, orientations, and inter-object relationships in dynamic 3D scenes-remains a key unsolved challenge. Existing approaches primarily rel

  88. Diego Ortego, Marlon Rodríguez, Mario Almagro, Kunal Dahiya

    Foundation models have revolutionized artificial intelligence across numerous domains, yet their transformative potential remains largely untapped in Extreme Multi-label Classification (XMC). Queries in XMC are associated with relevant labels from extremely large label spaces, where it is critical to strike a balance between efficiency and performance. There

  89. Osama Al Sheikh Ali, Sotiris Koutsoftas, Ze Zhang, Knut Akesson

    This paper presents an integrated navigation framework for Autonomous Mobile Robots (AMRs) that unifies environment representation, trajectory generation, and Model Predictive Control (MPC). The proposed approach incorporates a quadtree-based method to generate structured, axis-aligned collision-free regions from occupancy maps. These regions serve as both a

  90. Logan Jackson, Victor Boyer, Tanner Rima, Edward Cazalas

    The use of electrical motors and other remote systems are important tools in radiation environments. Certain harsh radiation environments, such as particle accelerators, require the use of remote systems during operation. Stepper motors are one motor, in particular, that have acquired interest in the nuclear field for use in these remote systems. The stepper

  91. Akash Karthikeyan, Yash Vardhan Pant

    Self-play reinforcement learning has demonstrated significant success in learning complex strategic and interactive behaviors in competitive multi-agent games. However, achieving such behaviors in continuous decision spaces remains challenging. Ensuring adaptability and generalization in self-play settings is critical for achieving competitive performance in

  92. Aishwarya Venkataramanan, Sai Karthikeya Vemuri, Adithya Ashok Chalain Valapil, Joachim Denzler

    Coherent anti-Stokes Raman scattering (CARS) spectroscopy is a powerful and rapid technique widely used in medicine, material science, and chemical analyses. However, its effectiveness is hindered by the presence of a non-resonant background that interferes with and distorts the true Raman signal. Deep learning methods have been employed to reconstruct the t

  93. Fu Zhang, Yuming Zhao

    This study introduces a hybrid quantum-classical dispatching framework designed for power systems with high renewable penetration. The proposed method integrates a variational quantum algorithm with classical optimization to provide noise-resilient performance under realistic hardware constraints. Extensive numerical tests and a real-world case study demonst

  94. Jonas Länzlinger, Katharina Müller, Burkhard Stiller, Bruno Rodrigues

    Depression affects over millions people worldwide, yet diagnosis still relies on subjective self-reports and interviews that may not capture authentic behavior. We present IHearYou, an approach to automated depression detection focused on speech acoustics. Using passive sensing in household environments, IHearYou extracts voice features and links them to DSM

  95. Min Li, Lailai Zhu

    Spinning ice discs in nature have been reported for more than a century, yet laboratory experiments have yielded diverse observations and contradictory explanations, leaving the mechanism behind the disc motion elusive. Here we combine numerical simulations and scaling analysis to investigate a freely moving ice disc in a lab-scale water tank. We observe the

  96. Chiel van der Laan, Alessandro Corbetta

    Pedestrians in crowds frequently move as part of small groups, constituting up to 70% of individuals. Dyads (groups of two) are the most frequent. Understanding quantitatively the dynamics of dyads walking in crowds is therefore an essential building block towards a fundamental comprehension of crowd behavior as a whole, and mandatory for accurate crowd dyna

  97. Ronit D. Gross, Yanir Harel, Ido Kanter

    The translation of written language has been known since the 3rd century BC; however, its necessity has become increasingly common in the information age. Today, many translators exist, based on encoder-decoder deep architectures, nevertheless, no quantitative objective methods are available to assess their performance, likely because the entropy of even a s

  98. Mansi Mishra

    If $A$ is in the $p$-Schatten class on $\mathbb{R}^n$, $1\leq p \leq \frac{4n}{2n-1}$, then the quantum translates of $A$ are linearly independent. Moreover, there exists a non-zero operator in the $p$-Schatten class on $\mathbb{R}^n$, $p>\frac{4n}{2n-1}$ whose quantum translates are linearly dependent.

  99. Mingxuan Tian, Haochen Mu, Donghong Ding, Mengjiao Li

    With the development of digital twins and smart manufacturing systems, there is an urgent need for real-time distortion field prediction to control defects in metal Additive Manufacturing (AM). However, numerical simulation methods suffer from high computational cost, long run-times that prevent real-time use, while conventional Machine learning (ML) models

  100. Lefan Dolg, Moritz Scharfstädt, Andrea Bergschneider, Dante M. Kennes

    We investigate exciton confinement to a quantum wire in monolayer $\text{MoSe}_2$ where the confinement is achieved by a p-i-n junction. We employ an effective-mass exciton model and solve the problem numerically, reflecting device geometries found in experimental state-of-the-art set up. Our method allows us to investigate the entire spectrum of confined st