Skip to content

October 2025 arXiv papers — page 134

Showing 13,30113,400 of 25,213 papers

  1. Jucai Yang, Liang Li, Yiwei Gu, Haiqin Wu

    The integration of blockchain technology into healthcare presents a paradigm shift for secure data management, enabling decentralized and tamper-proof storage and sharing of sensitive Electronic Health Records (EHRs). However, existing blockchain-based healthcare systems, while providing robust access control, commonly overlook the high latency in user-side

  2. Simon Kiefhaber, Stefan Roth, Simone Schaub-Meyer

    Cost volumes are used in every modern optical flow estimator, but due to their computational and space complexity, they are often a limiting factor regarding both processing speed and the resolution of input frames. Motivated by our empirical observation that cost volumes lose their importance once all other network parts of, e.g., a RAFT-based pipeline have

  3. Fitim Abdullahu, Helmut Grabner

    Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on large-scale visual and textual data has demonstrated impressive capabilities. We explore these models' potential to understand to what exten

  4. Eun Woo Im, Muhammad Kashif Ali, Vivek Gupta

    Large Vision-Language Models (LVLMs) have demonstrated remarkable multimodal capabilities, but they inherit the tendency to hallucinate from their underlying language models. While visual contrastive decoding has been proposed to mitigate this issue, existing methods often apply generic visual augmentations that disregard the specific context provided by the

  5. Mathias Federolf, Alexander Steinhoff, Monika Emmerling, Matthias Florian

    Hybrid interlayer excitons in bilayer MoS2 are a promising platform for nonlinear optics due to their intrinsic dipolar character, which combines in-plane and out-ofplane dipole moments. In this work, we directly probe the nonlinear exciton-exciton interactions of hybrid interlayer excitons. By applying an external out-of-plane electric field, we polarize th

  6. Arune Makareviciute, Qun Yang, Tomoki Fujita, Oliver Gross

    A secondary $\beta$-relaxation process is often the dominant source of atomic dynamics below $T_\mathrm{g}$ in many glass forming systems. Recent studies reported the presence of $\beta$-relaxations in amorphous phase-change materials (PCMs) and showed that suppressing the $\beta$-relaxation via annealing in Ge$_{15}$Sb$_{85}$ can effectively slow down its c

  7. Yanshan Xiao, Kaihong Wu, Bo Liu

    Label distribution learning (LDL) is a paradigm that each sample is associated with a label distribution. At present, the existing approaches are proposed for the single-view LDL problem with labeled data, while the multi-view LDL problem with labeled and unlabeled data has not been considered. In this paper, we put forward the multi-view semi-supervised lab

  8. Simon Lupart, Mohammad Aliannejadi, Evangelos Kanoulas

    We present ChatR1, a reasoning framework based on reinforcement learning (RL) for conversational question answering (CQA). Reasoning plays an important role in CQA, where user intent evolves across dialogue turns, and utterances are often underspecified, requiring contextual interpretation, query reformulation, and dynamic coordination between retrieval and

  9. Yang Cao, Sikun Yang, Hao Tian, Kai He

    Anomaly detection is a critical task in data mining and management with applications spanning fraud detection, network security, and log monitoring. Despite extensive research, existing unsupervised anomaly detection methods still face fundamental challenges including conflicting distributional assumptions, computational inefficiency, and difficulty handling

  10. Chenxin Yu, Hao Ma, Xu Li, Xiao-Lei Zhang

    Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings underexploited. We introduce a query-based spatial audio source extraction framework for recovering dry target signals from fi

  11. Yang Li, Aming Wu, Zihao Zhang, Yahong Han

    In this paper, we focus on Novel Class Discovery for Point Cloud Segmentation (3D-NCD), aiming to learn a model that can segment unlabeled (novel) 3D classes using only the supervision from labeled (base) 3D classes. The key to this task is to setup the exact correlations between the point representations and their base class labels, as well as the represent

  12. Jannick Borowitz, Ernestine Großmann, Mattthias Schimek

    Finding maximum-weight independent sets in graphs is an important NP-hard optimization problem. Given a vertex-weighted graph $G$, the task is to find a subset of pairwise non-adjacent vertices of $G$ with maximum weight. Most recently published practical exact algorithms and heuristics for this problem use a variety of data-reduction rules to compute (near-

  13. Arnaud Chéritat, Pascale Roesch

    Parabolic renormalization associates to a holomorphic map f with a parabolic fixed point another holomorphic map with a parabolic fixed point. This procedure is essential for understanding the phenomenon of parabolic enrichment, which occurs when one perturbs f appropriately. Shishikura defined in [Shi98] (see also [LY14]) a class of maps that is stable unde

  14. The Tracker Group of the CMS Collaboration

    The high-luminosity upgrade of the CERN LHC requires the replacement of the CMS tracking detector to cope with the increased radiation fluence while maintaining its excellent performance. An extensive R\&D program, aiming at using 3D pixel silicon sensors in the innermost barrel layer of the detector, has been carried out by CMS in collaboration with the FBK

  15. Aya Kaysan Bahjat

    An automatic document classification system is presented that detects textual content in images and classifies documents into four predefined categories (Invoice, Report, Letter, and Form). The system supports both offline images (e.g., files on flash drives, HDDs, microSD) and real-time capture via connected cameras, and is designed to mitigate practical ch

  16. Yuanhao Li, Keyuan Lai, Tianqi Wang, Qihao Liu

    Accurate property data for chemical elements is crucial for materials design and manufacturing, but many of them are difficult to measure directly due to equipment constraints. While traditional methods use the properties of other elements or related properties for prediction via numerical analyses, they often fail to model complex relationships. After all,

  17. Pablo Miralles-González, Javier Huertas-Tato, Alejandro Martín, David Camacho

    Computational stylometry studies writing style through quantitative textual patterns, enabling applications such as authorship attribution, identity linking, and plagiarism detection. Despite the relevance of language modeling to these tasks, the pre-training of modern large language models (LLMs) has been underutilized in authorship attribution and verifica

  18. Ziheng Ni, Congcong Liu, Cai Shang, Yiming Sun

    The ranking stage serves as the central optimization and allocation hub in advertising systems, governing economic value distribution through eCPM and orchestrating the user-centric blending of organic and advertising content. Prevailing ranking models often rely on fragmented modules and hand-crafted features, limiting their ability to interpret complex use

  19. Alessandro Brusaferri, Andrea Ballarino

    Dynamical downscaling is crucial for deriving high-resolution meteorological fields from coarse-scale simulations, enabling detailed analysis for critical applications such as weather forecasting and renewable energy modeling. Generative Diffusion models (DMs) have recently emerged as powerful data-driven tools for this task, offering reconstruction fidelity

  20. Ilaria Rinaldi, Giorgio Cartechini, Angelo Schiavi, Jan Gajewski

    At the Maastro Proton Therapy Center in Maastricht, patient-specific quality assurance (PSQA) using an independent GPU-accelerated Monte Carlo (MC) calculation has fully replaced conventional measurements, which are time-consuming and have limited sensitivity to clinically relevant errors. A fully automated and robust pipeline was developed, integrating two

  21. Vincel Hoang Ngoc Minh

    To factorize and to decompose the graphs of representative functions on the free monoid X * (generated by the alphabet X ) with values in the ring A containing Q, we examine various products of series (as concatenation, shuffle and its $\phi$ -deformations) and co-products, which are such that their associated non graded bialgebras are isomorphic, for A is a

  22. Rui Xu, Xingyuan Chen, Wenxing Huang, Minxuan Huang

    Conformal Prediction (CP) provides distribution-free uncertainty quantification by constructing prediction sets that guarantee coverage of the true labels. This reliability makes CP valuable for high-stakes federated learning scenarios such as multi-center healthcare. However, standard CP assumes i.i.d. data, which is violated in federated settings where cli

  23. Jakub Wójcik, Wojciech Bruzda, Ignacy Stachura, Remigiusz Augusiak

    Whether every pure genuinely multipartite entangled (GME) state necessarily exhibits genuine multipartite nonlocality (GMNL) remains an open question. By combining a recently proposed Bell inequality [I. Stachura \textit{et al.}, \href{https://iopscience.iop.org/article/10.1088/1367-2630/ad7753}{New J. Phys. \textbf{26}, 093029 (2024)}] with Hardy's paradox

  24. Vincel Hoang Ngoc Minh

    Two confluent rewriting systems in noncommutatives polynomials are constructed using the equations allowing the identification of the local coordinates (of second kind) of the graphs of the $\zeta$ polymorphism as being (shuffle or quasi-shuffle) characters and bridging two algebraic structures of polyzetas. In each system, the left side of each rewriting ru

  25. J. M. Kempf, H. Latter

    The buoyancy stability properties of the ICM are modified because of the anisotropic transport of heat along the magnetic field lines. This feature gives rise to the MTI when the temperature gradient is aligned with the gravity, which occurs in the outskirts of galaxy clusters. Most previous linear analyses of the MTI adopted a local, Boussinesq approach. Ho

  26. Jean Gillibert, Emmanuel Hallouin, Aaron Levin

    Combining $2$-descent techniques with Riemann-Roch and B\'ezout's theorems, we give an upper bound on the number of rational points of bounded height on elliptic and hyperelliptic curves over function fields of characteristic $\neq 2$. We deduce an upper bound on the number of $S$-integral points, where $S$ is a finite set of places. As a primary application

  27. Xuxin Cheng, Ke Zeng, Zhiquan Cao, Linyi Dai

    Enhancing customer experience is essential for business success, particularly as service demands grow in scale and complexity. Generative artificial intelligence and Large Language Models (LLMs) have empowered intelligent interaction systems to deliver efficient, personalized, and 24/7 support. In practice, intelligent interaction systems encounter several c

  28. Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra

    We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the

  29. Vindhyawasini Prasad

    Dark matter (DM) plays a crucial role in explaining the observed astrophysical anomalies, galaxy rotation curves, and other fundamental characteristics of the Universe. Many extensions of the Standard Model (SM), such as the dark hidden-sector model, provide an attractive framework in which DM can couple to SM particles via various portals. These portals giv

  30. Anton F. Seleznev, Maxim V. Kulesh

    We performed a statistical study of two old open clusters NGC 188 and M 67 using Gaia DR3 data. No tidal tails of the clusters were detected, which most likely had been destroyed when the cluster passed through the Galactic plane. The size estimates of the clusters depend on the range of astrometric parameters and stellar magnitudes of the stars used for sta

  31. Nishant Chandna, Akshat Kaushal

    LiDAR Simultaneous Localization and Mapping (SLAM) systems are essential for enabling precise navigation and environmental reconstruction across various applications. Although current point-to-plane ICP algorithms perform effec- tively in structured, feature-rich environments, they struggle in scenarios with sparse features, repetitive geometric structures,

  32. Tzung-En Hsieh, Michael S. Moritz, Andreas Mölkner, Christoph Wichmann

    We studied the surface properties of Ga-Cu based liquid metal alloys, a promising material system for supported catalytically active liquid metal solutions (SCALMS). The impact of Cu dilution in the (liquid) Ga matrix is in-detail investigated by X-ray and UV photoelectron spectroscopy (XPS/UPS) and Machine-Learned-Force Field (ML-FF) calculations. With decr

  33. Arthur Vogels, Benjamin Wong, Yann Choho, Annabelle Blangero

    Activation steering methods control large language model (LLM) behavior by modifying internal activations at inference time. However, most existing activation steering methods rely on a fixed steering strength, leading to either insufficient control or unadapted intervention that degrades text plausibility and coherence. We introduce In-Distribution Steering

  34. Haoyang Wu, Siheng Wu, William X. Liu, Fangui Zeng

    ALOHA2 is an enhanced version of the dual-arm teleoperated robot ALOHA, featuring higher performance and robustness compared to the original design, while also being more ergonomic. Like ALOHA, ALOHA2 consists of two grippers and two ViperX 6-DoF arms, as well as two smaller WidowX arms. Users control the follower mechanical arms by operating the leader mech

  35. Stefania Gatti, Erica Ipocoana, Alain Miranville

    In this paper, we propose a new non-isothermal Allen-Cahn (Ginzburg-Landau) model for tumor growth. After deriving it using a microforces approach, we study its well-posedness. In particular, we are able to prove the existence and uniqueness of a local and global-in-time solution to our PDE system.

  36. JiaKui Hu, Zhengjian Yao, Lujia Jin, Yinghao Chen

    This study introduces a Masked Degradation Classification Pre-Training method (MaskDCPT), designed to facilitate the classification of degradation types in input images, leading to comprehensive image restoration pre-training. Unlike conventional pre-training methods, MaskDCPT uses the degradation type of the image as an extremely weak supervision, while sim

  37. Sungnyun Kim, Kangwook Jang, Sungwoo Cho, Joon Son Chung

    This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directly in the language space. Our framework, DualHyp, empowers a large language model (LLM) to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and

  38. D Huilier

    The University of Strasbourg, fundamentally humanistic since its creation, has a complicated history, being sometimes German, sometimes French through the ages. If one focus, from 1871, on the area of mathematics, we can identify two periods. Political goals led to send high-level scientists, sometimes rather young, to develop theoretical, later applied math

  39. Fuma Omori, Atsushi Yano, Takuya Azumi

    Autonomous driving systems, critical for safety, require real-time guarantees and can be modeled as DAGs. Their acceleration features, such as caches and pipelining, often result in execution times below the worst-case. Thus, a probabilistic approach ensuring constraint satisfaction within a probability threshold is more suitable than worst-case guarantees f

  40. Zehao Sha

    In this paper, we study constant scalar curvature K\"ahler (cscK) metrics on complete non-compact K\"ahler--Einstein manifolds. We give sufficient conditions under which a cscK perturbation of a K\"ahler--Einstein metric must remain K\"ahler--Einstein. As a model case, we prove that the Bergman metric on a bounded strictly pseudoconvex domain is K\"ahler--Ei

  41. Maxim Kirsebom, Philipp Kunde, Tomas Persson, Mike Todd

    We consider the minimal distance between orbits of measure preserving dynamical systems. In the spirit of dynamical shrinking target problems we identify distance rates for which almost sure asymptotic closeness properties can be ensured. More precisely, we consider the set $E_n$ of pairs of points whose orbits up to time $n$ have minimal distance to each ot

  42. Keyan Zhou, Zecheng Tang, Lingfeng Ming, Guanghao Zhou

    The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on

  43. Jean-François Babadjian, Blanche Buet, Michael Goldman

    This paper studies the effect of anisotropy on sharp or diffuse interfaces models. When the surface tension is a convex function of the normal to the interface, the anisotropy is said to be weak. This usually ensures the lower semicontinuity of the associated energy. If, however, the surface tension depends on the normal in a nonconvex way, this so-called st

  44. BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson

    Using $e^+e^-$ collision data at 19 center-of-mass energies ranging from $4.396$ to $4.951~\mathrm{GeV}$ corresponding to a total integrated luminosity of $8.86~{\rm fb}^{-1}$ collected by the BESIII detector, the process $e^+e^-\to K^{0}K^-\pi^+ J/\psi+c.c.$ is observed for the first time, with a statistical significance of $9.4\sigma$ summing up all the da

  45. Hanyin Cheng, Ruitong Zhang, Yuning Lu, Peng Chen

    While Time Series Foundation Models (TSFMs) have demonstrated remarkable success in Multivariate Time Series Anomaly Detection (MTSAD), however, in real-world industrial scenarios, many time series comprise not only numerical variables such as temperature and flow, but also numerous discrete state variables that describe the system status, such as valve on/o

  46. Ine Gevers, Walter Daelemans

    Large language models (LLMs) have achieved striking successes on many benchmarks, yet recent studies continue to expose fundamental weaknesses. In this paper, we introduce Concept, a simple word-guessing board game, as a benchmark for probing abductive reasoning. Our results show that this game, easily solved by humans (with a success rate of over 90\%), is

  47. Ivan Lee, Taylor Berg-Kirkpatrick

    Recent studies suggest that very small language models (SLMs) can generate surprisingly coherent text when trained on simplified, child-directed corpora such as TinyStories. These findings have been interpreted as evidence that readability -- characterized by accessible vocabulary, familiar narrative structure, and simple syntax -- plays a key role in enabli

  48. Stephan Kleinbölting, Nigel Goldenfeld, Johannes Berg

    Phylogenetic trees capture evolutionary relationships among species and reflect the forces that shaped them. While many studies rely on branch length information, the topology of phylogenetic trees (particularly their degree of imbalance) offers a robust framework for inferring evolutionary dynamics when timing data is uncertain. Classical metrics, such as t

  49. Lingwei Ma, Qi Xiong, Zhenqiu Zhang

    This paper presents the nonlinear potential theory for mixed local and nonlocal $p$-Laplace type equations with coefficients and measure data, involving both superquadratic and subquadratic cases. We prove a class of universal pointwise estimates for the solution and its gradient via Riesz and Wolff potentials. These are achieved by imposing various low regu

  50. Malte Fliedner, Julian Golak, Yağmur Gül, Simone Neumann

    Growing demand for sustainable logistics and higher space utilization, driven by e-commerce and urbanization, increases the need for storage systems that are both energy- and space-efficient. Compact storage systems aim to maximize space utilization in limited storage areas and are therefore particularly suited in densely-populated urban areas where space is

  51. Emanuele Artioli, Farzad Tashtarian, Christian Timmerer

    As the popularity of video streaming entertainment continues to grow, understanding how users engage with the content and react to its changes becomes a critical success factor for every stakeholder. User engagement, i.e., the percentage of video the user watches before quitting, is central to customer loyalty, content personalization, ad relevance, and A/B

  52. Alejandro Guerra-Manzanares, Omar El-Herraoui, Michail Maniatakos, Farah E. Shamout

    One of the key challenges of collaborative machine learning, without data sharing, is multimodal data heterogeneity in real-world settings. While Federated Learning (FL) enables model training across multiple clients, existing frameworks, such as horizontal and vertical FL, are only effective in `ideal' settings that meet specific assumptions. Hence, they st

  53. Cyril Letrouit

    The stability of optimal transport maps with respect to perturbations of the marginals is a question of interest for several reasons, ranging from the justification of the linearized optimal transport framework to numerical analysis and statistics. Under various assumptions on the source measure, it is known that optimal transport maps are stable with respec

  54. Jun Ming Hou, Long Chen, Xuan Zheng, Jia Wei Wu

    Generative models such as AlphaFold and MatterGen can directly generate novel material structures with desired properties, accelerating the new materials discovery and revolutionizing the material design paradigm from traditional trial-and-error approach to intelligent on-demand generation. AlphaFold is focused on protein prediction with specific aperiodic s

  55. Joshua J. Collier, David R. Gozzard, John S. Wallis, Benjamin P. Dix-Matthews

    The maximum baseline, and therefore resolution, of optical astronomical interferometers is limited by attenuation and phase noise within the optical path between the apertures and beam combiner, as well as the practical challenges of constructing optical delay lines more than a few hundred meters in length. We implement off-band phase stabilization on two fi

  56. Weiqi Guo, Guanjun Liu, Ziyuan Zhou

    Multi-Agent Deep Reinforcement Learning (MADRL) has shown potential for cooperative and competitive tasks such as autonomous driving and strategic gaming. However, models trained by MADRL are vulnerable to adversarial perturbations on states and actions. Therefore, it is essential to investigate the robustness of MADRL models from an attack perspective. Exis

  57. Björn Filter, Ralf Möller, Özgür Lütfü Özçep

    Collaborative machine learning enables multiple data owners to jointly train models for improved predictive performance. However, ensuring incentive compatibility and fair contribution-based rewards remains a critical challenge. Prior work by Sim and colleagues (Rachel Hwee Ling Sim et al: Collaborative machine learning with incentive-aware model rewards. In

  58. Richard Medina Rodriguez

    In this article we study the well-posedness of the Boltzmann equation near its hydrodynamic limit on a bounded domain. We consider two types of domains, namely $C^2$ domains with Maxwell boundary conditions where the accommodation coefficient is a continuous space dependent function $\iota \in [\iota_0,1]$ for any $\iota_0 \in (0,1]$, or cylindrical domains

  59. Daniil Ignatev, Denis Paperno, Massimo Poesio

    The task of perspective-aware classification introduces a bottleneck in terms of parametric efficiency that did not get enough recognition in existing studies. In this article, we aim to address this issue by applying an existing architecture, the hypernetwork+adapters combination, to perspectivist classification. Ultimately, we arrive at a solution that can

  60. Jahidul Arafat, Sanjaya Poudel

    Nanopore sequencing enables real-time long-read DNA sequencing with reads exceeding 10 kilobases, but inherent error rates of 12-15 percent present significant computational challenges for read alignment. The critical seed chaining step must connect exact k-mer matches between reads and reference genomes while filtering spurious matches, yet state-of-the-art

  61. Quan Yuan, Qi Fang, Shishuo Fu, Haijun Li

    Hetyei introduced in 2019 the homogenized Linial arrangement and showed that its regions are counted by the median Genocchi numbers. In the course of devising a different proof of Hetyei's result, Lazar and Wachs considered another hyperplane arrangement that is associated with certain bipartite graph called Ferrers graph. We bijectively label the regions of

  62. Jiarui Li, Yuhan Chai, Lei Du, Chenyun Duan

    Rule-based network intrusion detection systems play a crucial role in the real-time detection of Web attacks. However, most existing works primarily focus on automatically generating detection rules for new attacks, often overlooking the relationships between new attacks and existing rules, which leads to significant redundancy within the ever-expanding rule

  63. Albert Suceava, Sankalpa Hazra, Aiden Ross, Ian Reed Philippi

    The search for thin film electro-optic (EO) materials that can retain superior performance under cryogenic conditions has become critical for quantum computing. Barium titanate thin films show large linear EO coefficients in the tetragonal phase at room temperature, which is severely degraded down to ~200 pm V$^{-1}$ in the rhombohedral phase at cryogenic te

  64. Jingmin An, Yilong Song, Ruolin Yang, Nai Ding

    Large Language Models (LLMs) demonstrate human-level or even superior language abilities, effectively modeling syntactic structures, yet the specific computational modules responsible remain unclear. A key question is whether LLM behavioral capabilities stem from mechanisms akin to those in the human brain. To address these questions, we introduce the Hierar

  65. Haoyu Zhang, Yuxuan Cheng, Wenqi Fan, Yulong Chen

    Graph neural networks (GNNs) have achieved remarkable success in various domains, yet they often struggle with domain adaptation due to significant structural distribution shifts and insufficient exploration of transferable patterns. One of the main reasons behind this is that traditional approaches do not treat global and local patterns discriminatingly so

  66. Chunhao Lu, Qiang Lu, Meichen Dong, Jake Luo

    Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of various modalities. To unify multi-modal processing, we propose a novel architecture called MDM (Multi-modal Diffusion Mamba). MDM utilizes a Mamba-based multi-step selection diffu

  67. Nguyen Gia Hien, Michael Reiter, Duong Ngoc Son

    We classify CR maps from the hyperquadric of signature $l>0$ in $\mathbb{C}^n$, $n\geq 3$, to the local model for the tube over the null cone of a symmetric form in $\mathbb{C}^{n+1}$, up to CR automorphisms of the source and target. In contrast to the setting of the Heisenberg hypersurface in $\mathbb{C}^3$ (i.e., the case $l=0$), studied earlier in Reiter-

  68. Dimitrios Voulanas, Eduardo Gildin

    Accurate and robust surrogate modeling is essential for the real-time control and optimization of large-scale subsurface systems, such as geological CO2 storage and waterflood management. This study investigates the limits of classical Dynamic Mode Decomposition with control (DMDc) and introduces CCKM, as a robust alter-native, in enforcing control in pressu

  69. Minji Kim, Taekyung Kim, Bohyung Han

    Video Large Language Models (VideoLLMs) extend the capabilities of vision-language models to spatiotemporal inputs, enabling tasks such as video question answering (VideoQA). Despite recent advances in VideoLLMs, their internal mechanisms on where and how they extract and propagate video and textual information remain less explored. In this study, we investi

  70. Zhiyuan Zhao, Yubin Wen, Siyu Yang, Lichen Ning

    Crowd counting is a task of estimating the number of the crowd through images, which is extremely valuable in the fields of intelligent security, urban planning, public safety management, and so on. However, the existing counting methods have some problems in practical application on embedded systems for these fields, such as excessive model parameters, abun

  71. Wei Meng, Yuan You, Shuang-Nan Zhang, Jia-Ying Cao

    For accreting black holes (BHs), the lamp-post scenario is a simple and popular model: a hot and point-like corona is located above the black hole, irradiating the accretion disk with hard X-ray radiation, which is believed to be generated by inverse Compton scattering in the corona. Although the lamp-post model successfully explains the disk reflection comp

  72. Yunze Wei, Kaiwen Wei, Shibo Du, Jianyu Wang

    Network protocol testing is fundamental for modern network infrastructure. However, traditional network protocol testing methods are labor-intensive and error-prone, requiring manual interpretation of specifications, test case design, and translation into executable artifacts, typically demanding one person-day of effort per test case. Existing model-based a

  73. Emily C. Adlam, Kelvin J. McQueen, Mordecai Waegell

    What are the physical requirements for agency? We investigate whether a purely quantum system (one evolving unitarily in a coherent regime without decoherence or collapse) can satisfy three minimal conditions for agency: an agent must be able to create a world-model, use it to evaluate the likely consequences of alternative actions, and reliably perform the

  74. Li Liang, Bo Miao, Xinyu Wang, Naveed Akhtar

    Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large-scale benchmark for generating 3D outdoor sema

  75. Xuanchen Wang, Heng Wang, Weidong Cai

    Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their ability to capture rhythmic and structural cues that drive movement. We propose MotionBeat, a framework for motion-aligned music representation

  76. Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh

    The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimod

  77. Ashutosh Dixit, Hichem Hajaiej, Tuhina Mukherjee

    This study investigates the existence, uniqueness, and multiplicity of positive solutions for a system of fractional differential equations given by: \begin{equation*} (-\Delta)^{s_i} u_{i}+\lambda_{i} u_{i}=\sum_{j=1}^{n} \alpha_{i j}\left|u_{j}\right|^{q_{i j}}\left|u_{i}\right|^{p_{i j}-2} u_{i} , u_i\in {\mathscr{D}^{s_i,2}\left(\mathbb{R}^{N}\right)}, i

  78. Jian-Wen Zhou

    We performed N-body simulations of both individual cluster evolution and subcluster coalescence, demonstrating that cluster evolution and its outcomes strongly depend on the cluster formation process through comparisons of different gas expulsion modes and formation channels. The evolution of star clusters is significantly shaped by the gas expulsion mode, w

  79. Ikki Mitsuhashi, Katherine A. Suess, Joel Leja, Pratika Dayal

    We report the discovery of two z ~ 12 galaxy candidates with unusually red UV slopes (betaUV ~> -1.5), and probe the origin of such colors at cosmic dawn. From Prospector fits to the UNCOVER/MegaScience dataset -- deep JWST/NIRCam imaging of Abell 2744 in 20 broad- and medium-bands -- we identify several new z > 10 galaxies. Medium-band data improve redshift

  80. Tuoxin Li, Juncheng Wei, Haidong Yang

    In this paper, we consider the following prescribed scalar curvature problem: \begin{equation*} -\Delta u = K(x) u^{\frac{n+2}{n-2}}, \quad u>0\quad\hbox{in}\quad \mathbb{R}^n, \quad u \in D^{1,2}(\mathbb{R}^n), \end{equation*} where $K(x)$ is a volcano-like positive function such that $$ K(x)= K(r_0)- c_0 | |x|- r_0|^m + O( | |x|- r_0|^{m+\theta}),\quad r_0

  81. Toshio Mikami

    We introduce a stochastic optimal transport for the Langevin dynamics with positive mass and study its zero--mass limit. The new aspect of this paper is that we only fix the initial and terminal probability distributions of the positions of particles under consideration, but not those of their velocities with Heisenberg's uncertainty principle in mind. In th

  82. Haochuan Xu, Yun Sing Koh, Shuhuai Huang, Zirun Zhou

    Vision-Language-Action (VLA) models have achieved revolutionary progress in robot learning, enabling robots to execute complex physical robot tasks from natural language instructions. Despite this progress, their adversarial robustness remains underexplored. In this work, we propose both adversarial patch attack and corresponding defense strategies for VLA m

  83. Sebastian mateos Nicolajsen

    I here conduct an exploration of programming language extensibility, making an argument for an often overlooked component of conventional language design. Now, this is not a technical detailing of these components, rather, I attempt to provide an overview as I myself have lacked during my time investigating programming languages. Thus, read this as an introd

  84. Yukuan Zhang, Jiarui Zhao, Shangqing Nie, Jin Kuang

    Multimodal semantic cues, such as textual descriptions, have shown strong potential in enhancing target perception for tracking. However, existing methods rely on static textual descriptions from large language models, which lack adaptability to real-time target state changes and prone to hallucinations. To address these challenges, we propose a unified mult

  85. Yinglong Yan, Jun Yue, Shaobo Xia, Hanmeng Sun

    Vector extraction retrieves structured vector geometry from raster images, offering high-fidelity representation and broad applicability. Existing methods, however, are usually tailored to a single vector type (e.g., polygons, polylines, line segments), requiring separate models for different structures. This stems from treating instance attributes (category

  86. Arghya Mukherjee, Arnab Hazra, Dootika Vats

    Spatial generalized linear mixed-effects models are popularly used to analyze spatially indexed univariate responses. However, with modern technology, it is common to observe vector-valued mixed-type responses, e.g., a combination of binary, count, or continuous types, at each location. Methods for jointly modeling such mixed-type multivariate spatial respon

  87. Inha Kang, Youngsun Lim, Seonho Lee, Jiho Choi

    State-of-the-art vision-language models (VLMs) suffer from a critical failure in understanding negation, often referred to as affirmative bias. This limitation is particularly severe in described object detection (DOD) tasks. To address this, we propose two primary contributions: (1) a new dataset pipeline and (2) a novel, lightweight adaptation recipe. Firs

  88. Mihir Gupta

    In this paper, we examine the asymptotic behavior of the longest increasing subsequence (LIS) in a uniformly random permutation of $n$ elements. We rely on the Robinson--Schensted--Knuth correspondence, Young tableaux, and key classical results -- including the Erd\H{o}s--Szekeres theorem and the Hook Length Formula -- to demonstrate that the expected LIS le

  89. Jalal Khan, Manzoor Khan, Sherzod Turaev, Sumbal Malik

    The driving environment perception has a vital role for autonomous driving and nowadays has been actively explored for its realization. The research community and relevant stakeholders necessitate the development of Deep Learning (DL) models and AI-enabled solutions to enhance autonomous vehicles (AVs) for smart mobility. There is a need to develop a model t

  90. Yi Zhang, Lili Xie, Ruihong Qiu, Jiajun Liu

    Recommender systems (RecSys) have become critical tools for enhancing user engagement by delivering personalized content across diverse digital platforms. Recent advancements in large language models (LLMs) demonstrate significant potential for improving RecSys, primarily due to their exceptional generalization capabilities and sophisticated contextual under

  91. Yan Tan, Chenhao Ye, Qinghai Zhang, Shubo Zhao

    The secant method, as an important approach for solving nonlinear equations, is introduced in nearly all numerical analysis textbooks. However, most textbooks only briefly address the Q-order of convergence of this method, with few providing rigorous mathematical proofs. This paper establishes a rigorous proof for the Q-order of convergence of the secant met

  92. Divyanshu Singh, Ashman Mehra, Kavya Makwana, Snehanshu Saha

    Urban mobility systems face persistent challenges of congestion, underutilized vehicles, and rising emissions driven by private point-to-point commuting. Although ride-sharing platforms exist, their profit-driven incentive structures often fail to align individual participation with broader community benefit. We introduce Altruistic Ride Sharing (ARS), a dec

  93. Hang-Cheng Dong, Yibo Jiao, Fupeng Wei, Guodong Liu

    Industrial surface defect inspection for sample-wise quality control (QC) must simultaneously decide whether a given sample contains defects and localize those defects spatially. In real production lines, extreme foreground-background imbalance, defect sparsity with a long-tailed scale distribution, and low contrast are common. As a result, pixel-centric tra

  94. Andrei Khrennikov

    This paper presents a comprehensive review of the Social Laser Theory (SLT) as a natural extension of the broader framework of Quantum-Like Modeling (QLM). While QLM applies the mathematical formalism of quantum theory . such as Hilbert space representations, interference, and non-commutative observables - to model context-dependent and non-classical phenome

  95. Tony J. Puthenpurakal

    Let $(A,\mathfrak{m})$ be an excellent Gorenstein local ring of dimension $d \geq 2$ which is an isolated singularity. Let $\widehat{A}$ denote the completion of $A$. If $G(A)$ is the Grothendieck group of $A$ then by $G(A)_\mathbb{Q}$ we denote $G(A)\otimes_\mathbb{Z} \mathbb{Q}$. We prove that the natural map $G(A)_\mathbb{Q} \rightarrow G(\widehat{A})_\ma

  96. Y. Yang, C. A. Morales

    We introduce the concept of topological expansive flow. We prove that this concept is invariant by topological conjugacy and reduces to expansivity in the compact case. We characterize tiopological expansive flows as rescaling expansive flows for which the singularities are isolated points of the space. Finally, we prove that the growth rate of the periodic

  97. Yiyuan He, Minxian Xu, Jingfeng Wu, Jianmin Hu

    Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, which separates prompt prefill from auto-regressive decode, has emerged as a promising architecture by isolating their heterogeneous compute and memory demands. However, current disa

  98. Samuel R. Grant, Geraint H. Jones

    During October - November 2025, interstellar comet 3I/ATLAS, will pass upstream of the Europa Clipper and Hera spacecraft. Here, we identify two potential opportunities for in-situ observations of 3I's ion tail by immersion, facilitated by the close alignment between the comet's hyperbolic trajectory with the ecliptic plane. During the period 30 October - 6

  99. Philipp Grundhuber, Mhd Modar Halimeh, Emanuël A. P. Habets

    This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech recordings while preserving content and speaker identity. We build upon previous work using the EnCodec architecture, a

  100. Yufei He, Juncheng Liu, Yue Liu, Yibo Li

    A fundamental limitation of current AI agents is their inability to learn complex skills on the fly at test time, often behaving like "clever but clueless interns" in novel environments. This severely limits their practical utility. To systematically measure and drive progress on this challenge, we first introduce the Jericho Test-Time Learning (J-TTL) bench