Skip to content

November 2025 arXiv papers — page 53

Showing 5,2015,300 of 22,271 papers

  1. Jingqian Zhao, Bingbing Wang, Geng Tu, Yice Zhang

    Data contamination poses a significant challenge to the fairness of LLM evaluations in natural language processing tasks by inadvertently exposing models to test data during training. Current studies attempt to mitigate this issue by modifying existing datasets or generating new ones from freshly collected information. However, these methods fall short of en

  2. Qian Jiang, Qianqian Wang, Xin Jin, Michal Wozniak

    Remote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an important issue is to obtain a color image with high spatial resol

  3. Guangyuan Li, Bo Li, Jinwei Chen, Xiaobin Hu

    Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with multiple interacting subjects, where dynamic subjects fail to follow realistic motion patterns during scene evolution. Secon

  4. Sudipta Ghosh, Mike Miller Eismeier

    We determine the framed instanton homology with coefficients in $\mathbb F = \mathbb Z/2$ for Dehn surgeries on a knot in the $3$-sphere. The dimension of these groups is seen to have a close relationship with a homology cobordism invariant due to Froyshov. As an application, we show that $r$-surgery on a non-trivial knot cannot be nondegenerate $SU(2)$-abel

  5. Jihun Park, Junyong Shin, Jinsung Park, Yo-Seb Jeon

    This paper proposes robust nonlinear transform coding (Robust-NTC), a generalizable digital joint source-channel coding (JSCC) framework that couples variational latent modeling with channel-adaptive transmission. Unlike learning-based JSCC methods that implicitly absorb channel variations, Robust-NTC explicitly models element-wise latent distributions via a

  6. Richard Golnik, Thomas Gatter, Peter F. Stadler, Nicola Vassena

    Autocatalysis is an important feature of metabolic networks, contributing crucially to the self-maintenance of organisms. Autocatalytic subsystems of chemical reaction networks (CRNs) are characterized in terms of algebraic conditions on submatrices of the stoichiometric matrix. Here, we derive sufficient conditions for subgraphs supporting irreducible autoc

  7. Martín-Olalla, José María

    The relationship between the vanishing of the heat capacities as $T\to0^+$ and the thermal stability is examined. The heat capacities vanish as fast as or faster than $T$ as $T\to0^+$ for states at the phase space boundary ($T=0$) to sustain the standard thermal stability criterion $U_{ss}>0$. Conversely, weakly vanishing heat capacities, which signify a los

  8. Ayca Duran, Christoph Waibel, Bernd Bickel, Iro Armeni

    Building integrated photovoltaic (BIPV) facades represent a promising pathway towards urban decarbonization, especially where roof areas are insufficient and ground-mounted arrays are infeasible. Although machine learning-based approaches to support photovoltaic (PV) planning on rooftops are well researched, automated approaches for facades still remain scar

  9. Ye Xu

    It is postulated that heavy dark matter $\phi$ with a mass on the order of TeV, once captured by the Earth, can decay into relativistic milli-charged particles (MCPs). These MCPs are potentially detectable at the IceCube neutrino telescope. In this study, MCPs are modeled within the massless hidden photon framework, where they interact with nuclei via a runn

  10. Christoph Brause, Dieter Rautenbach, Laurin Schwartze

    Kamyczura introduced the notion of a majority additive $k$-coloring of a graph $G$ as a function $c: V(G) \to \{1,2,\ldots,k\}$ such that $$\left|\left\{u \in N_G(v):\sum_{w \in N_G(u)} c(w) = s \right\}\right|\leq \max\left\{1,\frac{d_G(v)}{2}\right\}$$ for every vertex $v$ of $G$ and every positive integer $s$. We show that every graph $G$ of maximum degre

  11. Bowen Duan, Ge Zhang

    When liquids are cooled rapidly, they bypass crystallization and instead enter a supercooled state and then a glass state. Previous studies have shown that the static structure factors of high-temperature liquids, supercooled liquids, and glasses exhibit only subtle differences, leading to the conclusion that the glass transition cannot be predicted solely f

  12. Suzie Kim, Hye-Bin Shin, Hyo-Jeong Jang

    In this work, we investigate how implicit neural feed back can accelerate reinforcement learning in complex robotic manipulation settings. While prior electroencephalogram (EEG) guided reinforcement learning studies have primarily focused on navigation or low-dimensional locomotion tasks, we aim to understand whether such neural evaluative signals can improv

  13. Geunho Noh, Panayotis Benetatos

    We study two cross-linked polymer systems in the strong stretching regime. The first consists of two polymers sharing one endpoint, with the other two endpoints coupled by a harmonic potential. Within the weakly bending approximation, we analyze the tensile elastic response for freely jointed or wormlike chains; for the latter, the approximation applies eith

  14. Colin Faverjon, Marina Poulet

    Mahler equations arise in a wide range of contexts including the study of finite automata, regular sequences, algebraic series over Fp(z), and periods of Drinfeld modules. Introduced a century ago by K. Mahler to study the transcendence of certain complex numbers, they have recently been the subject of several works establishing a deep connection between suc

  15. Lilian Say, Christophe Denis, Rafael Pinot

    The increasing use of machine learning in sensitive applications demands algorithms that simultaneously preserve data privacy and ensure fairness across potentially sensitive sub-populations. While privacy and fairness have each been extensively studied, their joint treatment remains poorly understood. Existing research often frames them as conflicting objec

  16. Wengyi Zhan, Mingbao Lin, Zhihang Lin, Rongrong Ji

    Multimodal large language models (MLLMs) deliver impressive vision-language reasoning but suffer steep inference latency because self-attention scales quadratically with sequence length and thousands of visual tokens contributed by high-resolution images. Naively pruning less-informative visual tokens reduces this burden, yet indiscriminate removal can strip

  17. Yuzhi Chen, Yuanchang Xie, Lei Zhao, Pan Liu

    Multimodal trajectory prediction generates multiple plausible future trajectories to address vehicle motion uncertainty from intention ambiguity and execution variability. However, HD map-dependent models suffer from costly data acquisition, delayed updates, and vulnerability to corrupted inputs, causing prediction failures. Map-free approaches lack global c

  18. Yiming Wang, Shaofei Wang, Marko Mihajlovic, Siyu Tang

    3D Gaussian Splatting (3DGS) has emerged as a leading approach for high-quality novel view synthesis, with numerous variants extending its applicability to a broad spectrum of 3D and 4D scene reconstruction tasks. Despite its success, the representational capacity of 3DGS remains limited by the use of 3D Gaussian kernels to model local variations. Recent wor

  19. Hector Bouton, Laurent Desvillettes, Helge Dietert

    In this work, we adapt our recent article [BDD25] to the setting of Dirichlet boundary conditions. A key part is the study of the parabolic equation $a\partial_t w - \Delta w = f$ with a rough coefficient $a$, homogeneous Dirichlet boundary conditions, and the special assumption $\partial_tw \ge 0$. We then apply it to prove existence of global strong soluti

  20. Bing Wu, Chang Zou, Changlin Li, Duojun Huang

    We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architec

  21. Shuyang Liu, Yuan Jin, Rui Lin, Shizhe Chen

    Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source multi-scale representations module to obtain complementary segment- and track-level features, (2) a hierarchical augmentation stra

  22. Dezhi Ran, Shuxiao Xie, Mingfang Ji, Anmin Liu

    High-performance GPU kernels are critical for efficient LLM serving, yet their optimization remains a bottleneck requiring deep system expertise. While code LLMs show promise in generating functionally correct code, kernel optimization is intrinsically a search problem over a vast optimization space. The fundamental mismatch prevents existing LLM agents from

  23. Liutong Han, Chu Kang, Mingjie Xing, Yanjun Wu

    Intrinsic functions are specialized functions provided by the compiler that efficiently operate on architecture-specific hardware, allowing programmers to write optimized code in a high-level language that fully exploits hardware features. Using intrinsics to vectorize core code blocks is a standard optimization method in high-performance libraries, often re

  24. Lei Ming, Himanshu Chaudhary, Shi-Dong Liang, Hong-Hao Zhang

    We consider an $f(Q, T)$ gravity theory with a Schr\"{o}dinger type vectorial non-metricity. In the presence of such a non-metricity, the length of vectors is preserved under autoparallel transport. We obtain the field equations assuming a vanishing total scalar curvature, implemented by a Lagrange multiplier, and investigate their cosmological implications.

  25. Yu Zhang, Haoan Ping, Yuchen Li, Zhenshan Bing

    Recent salient object detection (SOD) methods aim to improve performance in four key directions: semantic enhancement, boundary refinement, auxiliary task supervision, and multi-modal fusion. In pursuit of continuous gains, these approaches have evolved toward increasingly sophisticated architectures with multi-stage pipelines, specialized fusion modules, ed

  26. Yang Xiang, Yixin Ji, Juntao Li, Min Zhang

    Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning benchmarks. However, their long chain-of-thought reasoning processes incur significant inference overhead. Pruning has emerged as a promising approach to reducing computational costs. However, existing efforts have primarily focused on large language models (LLMs), wh

  27. Wai-Kit Lam, Arnab Sen

    We study correlation decay for the maximum weight matching problem on sparse graphs with i.i.d. edge weights. We show exponential decay of correlations when the underlying graphs are locally tree-like with uniformly bounded degree and the edge weights are exponential. We also prove a polynomial rate of decay of correlations for any finite graph with maximum

  28. Xingyu Huang, Fei Jiang, Jianli Xiao

    With the rapid development of large language models (LLMs), the applications of LLMs have grown substantially. In the education domain, LLMs demonstrate significant potential, particularly in automatic text generation, which enables the creation of intelligent and adaptive learning content. This paper proposes a new LLMs framework, which is named as Reading

  29. Vidi Team, Chia-Wen Kuo, Chuang Huang, Dawei Du

    Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video unde

  30. Bo Jiang, Weijun Zhao, Beibei Wang, Xiao Wang

    Recently, fine-tuning large-scale pre-trained GNNs has yielded remarkable attention in adapting pre-trained GNN models for downstream graph learning tasks. One representative fine-tuning method is to exploit adapter (termed AdapterGNN) which aims to 'augment' the pre-trained model by inserting a lightweight module to make the 'augmented' model better adapt t

  31. Xiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou

    Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies induce biased model representations and corrupt statistical esti

  32. Changsheng Luo, Yushi Wang, Wenhan Cai, Mingguo Zhao

    Accurate proprioceptive odometry is fundamental for legged robot navigation in GPS-denied and visually degraded environments where conventional visual odometry systems fail. Current approaches face critical limitations: analytical filtering methods suffer from modeling uncertainties and cumulative drift, hybrid learning-filtering approaches remain constraine

  33. Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang, Kaixin Wang

    Scaling vision-language-action (VLA) model pre-training requires large volumes of diverse, high-quality manipulation trajectories. Most current data is obtained via human teleoperation, which is expensive and difficult to scale. Reinforcement learning (RL) methods learn useful skills through autonomous exploration, making them a viable approach for generatin

  34. Sana Alamgeer, Mylene Farias, Marcelo Carvalho

    The main goal of the project is to design a new model that predicts regions of interest in 360$^{\circ}$ videos. The region of interest (ROI) plays an important role in 360$^{\circ}$ video streaming. For example, ROIs are used to predict view-ports, intelligently cut the videos for live streaming, etc so that less bandwidth is used. Detecting view-ports in a

  35. Rui Li, Ronglong Dou, Ting Gao, Qinan Li

    The electronic structure of the ICl+ molecular ion is investigated by using high-level multireference configuration interaction (MRCI) method. To improve computational accuracy, Davidson corrections, spin-orbit coupling (SOC), and core-valence electron correlations effects are incorporated into the calculations. The potential energy curves (PECs) of 21 elect

  36. Yujing Wang, Weize Hong

    We present a novel framework that integrates Large Language Models (LLMs) into the Git bisect process for semantic fault localization. Traditional bisect assumes deterministic predicates and binary failure states assumptions often violated in modern software development due to flaky tests, nonmonotonic regressions, and semantic divergence from upstream repos

  37. Sophie Béraud-Dufour, Ilona Legroux, Thierry Coppola, Patricia Lebrun

    Stroke is a leading cause of disability and death worldwide, with ischemic strokes accounting for nearly 80% of cases. Fewer than 5% of patients receive the sole validated pharmacotherapy, intravenous thrombolysis, highlighting the urgent need for novel therapies. Within this landscape, the exploration of natural molecules emerges as a promising avenue, part

  38. Masoomali Fatehkia, Enes Altinisik, Husrev Taha Sencar

    Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of ove

  39. Yilin Wen, Kechuan Dong, Yusuke Sugano

    Online test-time adaptation addresses the train-test domain gap by adapting the model on unlabeled streaming test inputs before making the final prediction. However, online adaptation for 3D human pose estimation suffers from error accumulation when relying on self-supervision with imperfect predictions, leading to degraded performance over time. To mitigate

  40. Mohammad Nour Al Awad, Sergey Ivanov, Olga Tikhonova

    Large Language Models (LLMs) are increasingly integrated into code editors to provide AI-powered code suggestions. Yet many of these suggestions are ignored, resulting in wasted computation, increased latency, and unnecessary interruptions. We introduce a lightweight pre-filtering model that predicts the likelihood of suggestion acceptance before invoking th

  41. Václav Tran, Jakub Šmíd, Ladislav Lenc, Jean-Pierre Salmon

    Text summarization is the task of automatically condensing longer texts into shorter, coherent summaries while preserving the original meaning and key information. Although this task has been extensively studied in English and other high-resource languages, Czech summarization, particularly in the context of historical documents, remains underexplored. This

  42. Ishmam Tashdeed, Md. Atiqur Rahman, Sabrina Islam, Md. Azam Hossain

    Personalized federated learning (PFL) possesses the unique capability of preserving data confidentiality among clients while tackling the data heterogeneity problem of non-independent and identically distributed (Non-IID) data. Its advantages have led to widespread adoption in domains such as medical image segmentation. However, the existing approaches mostl

  43. Yubo Wang, Hui He, Chaoxi Niu, Zhendong Niu

    Due to the inherent complexity, temporal patterns in real-world time series often evolve across multiple intertwined scales, including long-term periodicity, short-term fluctuations, and abrupt regime shifts. While existing literature has designed many sophisticated decomposition approaches based on the time or frequency domain to partition trend-seasonality

  44. Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu

    Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such metho

  45. Iona Ann Sebastian, S. M. Sunoj

    Fractional cumulative residual inaccuracy (FCRI) measure allows to determine regions of discrepancy between systems, depending on their respective fractional and chaotic map parameters. Most of the theoretical results and applications related to the FCRI of the lifetime random variable are based on the distribution function approach. However, there are situa

  46. Heger Arfaoui, Mohammed Iheb Hergli, Beya Benzina, Slimane BenMiled

    Focus group discussions generate rich qualitative data but their analysis traditionally relies on labor-intensive manual coding that limits scalability and reproducibility. We present a systematic framework for applying BERTopic to focus group transcripts using data from ten focus groups exploring HPV vaccine perceptions in Tunisia (1,075 utterances). We con

  47. Mohammad Nour Al Awad, Sergey Ivanov, Olga Tikhonova

    Large Language Models (LLMs) have transformed code auto-completion by generating context-aware suggestions. Yet, deciding when to present these suggestions remains underexplored, often leading to interruptions or wasted inference calls. We propose an adaptive timing mechanism that dynamically adjusts the delay before offering a suggestion based on real-time

  48. Mincheol Jeon, Euinam Huh

    Personalized Federated Learning (PFL) faces persistent challenges, including domain heterogeneity from diverse client data, data imbalance due to skewed participation, and strict communication constraints. Traditional federated learning often lacks personalization, as a single global model cannot capture client-specific characteristics, leading to biased pre

  49. Hongyu Lyu, Thomas Monninger, Julie Stephany Berrio Perez, Mao Shan

    Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local maps from on-board sensors. However, existing methods typically rely on costly 3D map annotations for training, which limits their generalization and sca

  50. Binglin Liu, Yucheng Wang, Zheyuan Zhang, Jiyuan Lu

    The adaptation of teaching slides to instructors' situated teaching needs, including pedagogical styles and their students' context, is a critical yet time-consuming task for educators. Through a series of educator interviews, we first identify and systematically categorize the key friction points that impede this adaptation process. Grounded in these findin

  51. Yasiru Laksara, Uthayasanker Thayasivam

    The utility of deep learning models, such as CheXNet, in high stakes clinical settings is fundamentally constrained by their purely deterministic nature, failing to provide reliable measures of predictive confidence. This project addresses this critical gap by integrating robust Uncertainty Quantification (UQ) into a high performance diagnostic platform for

  52. Xiaofan Li, Chenming Wu, Yanpeng Sun, Jiaming Zhou

    Visual autoregressive models achieve remarkable generation quality through next-scale predictions across multi-scale token pyramids. However, the conventional method uses uniform scale downsampling to build these pyramids, leading to aliasing artifacts that compromise fine details and introduce unwanted jaggies and moir\'e patterns. To tackle this issue, we

  53. Wenxin He, Bin Xu

    In this paper, we study the complex structures of complete hyperk\"ahler four-manifolds of infinite topological type arising from the Gibbons-Hawking ansatz. We show that for almost all complex structures in the hyperk\"ahler family, the manifold is biholomorphic to a hypersurface in $\mathbb{C}^3$ defined by an explicit entire function. For the remaining co

  54. Fang Wang, Lance Kosca, Adrienne Kosca, Marko Gacesa

    This paper introduces HGNN(O), an AutoML GNN hypermodel framework for outcome prediction on event-sequence data. Building on our earlier work on graph convolutional network hypermodels, HGNN(O) extends four architectures-One Level, Two Level, Two Level Pseudo Embedding, and Two Level Embedding-across six canonical GNN operators. A self-tuning mechanism based

  55. Lei Ke, Hubery Yin, Gongye Liu, Zhengyao Lv

    With the success of flow matching in visual generation, sampling efficiency remains a critical bottleneck for its practical application. Among flow models' accelerating methods, ReFlow has been somehow overlooked although it has theoretical consistency with flow matching. This is primarily due to its suboptimal performance in practical scenarios compared to

  56. Jonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye

    We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes to support goal-directed reasoning. While this form of reasoning is fundamental to human cognition, current vision-lang

  57. Huadai Liu, Kaicheng Luo, Wen Wang, Qian Chen

    Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and lack human preference alignment. We introduce PrismAudio, th

  58. Shivam Pal, Sakshi Varshney, Piyush Rai

    Deep neural networks are prone to learning shortcuts, spurious correlations present in the training data that undermine out-of-distribution (OOD) generalization. Most prior work mitigates shortcut learning through input-space reweighting, either relying on explicit shortcut labels or inferring shortcut structure from heuristics such as per-sample loss. Moreo

  59. Kaize Shi, Xueyao Sun, Xiaohui Tao, Lin Li

    Large Language Models (LLMs) face information overload when handling long contexts, particularly in Retrieval-Augmented Generation (RAG) where extensive supporting documents often introduce redundant content. This issue not only weakens reasoning accuracy but also increases computational overhead. We propose an unsupervised context compression framework that

  60. Shaobo Wang, Tianle Niu, Runkang Yang, Deshan Liu

    The scalability of video understanding models is increasingly limited by the prohibitive storage and computational costs of large-scale video datasets. While data synthesis has improved data efficiency in the image domain, its extension to video remains challenging due to pervasive temporal redundancy and complex spatiotemporal dynamics. In this work, we unc

  61. Fang Wang, Paolo Ceravolo, Ernesto Damiani

    Existing deep learning models for Predictive Process Monitoring (PPM) struggle with temporal irregularities, particularly stochastic event durations and overlapping timestamps, limiting their adaptability across heterogeneous datasets. We propose a dual input neural network strategy that separates event and sequence attributes, using a duration-aware pseudo-

  62. Kanav Arora, Girish Narayanswamy, Shwetak Patel, Richard Li

    Heart rate estimation from photoplethysmography (PPG) signals generated by wearable devices such as smartwatches and fitness trackers has significant implications for the health and well-being of individuals. Although prior work has demonstrated deep learning models with strong performance in the heart rate estimation task, in order to deploy these models on

  63. Boyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan

    By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. T

  64. Edgar Dobriban

    Over the last few months, AI models including large language models have improved greatly. There are now several documented examples where they have helped professional mathematical scientists prove new results, sometimes even helping resolve known open problems. In this short note, we add another example to the list, by documenting how we were able to solve

  65. Mohammadreza Amiri, Monireh Hosseini

    Despite being among the most common psychological disorders, anxiety-related conditions are still primarily identified through subjective assessments, such as clinical interviews and self-evaluation questionnaires. These conventional methods often require significant time and may vary depending on the evaluator. However, the emergence of advanced artificial

  66. Aakash Gore, Anoushka Dey, Aryan Mishra

    Knowledge distillation has emerged as a powerful technique for model compression, enabling the transfer of knowledge from large teacher networks to compact student models. However, traditional knowledge distillation methods treat all teacher predictions equally, regardless of the teacher's confidence in those predictions. This paper proposes an uncertainty-a

  67. Xiele Wu, Zicheng Zhang, Mingtao Chen, Yixian Liu

    Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and algorithms have been proposed, existing approaches are constrained by two limitations: the absence of systematic definitions for evaluation dimensions, and the isolated treatment of

  68. Alvin Wei Ming Tan, Jane Yang, Tarun Sepuri, Khai Loong Aw

    Figuring out which objects or concepts words refer to is a central language learning challenge for young children. Most models of this process posit that children learn early object labels from co-occurrences of words and their referents that occur when someone around them talks about an object in the immediate physical environment. But how aligned in time a

  69. Fufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao

    We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient events in long videos. VideoPerceiver adopts a two-stage training framework. During supervised fine-tuning (SFT), we co

  70. Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang

    Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from potential information loss and non-end-to-end training. In contrast, existing pixel space models bypass VAEs but are computationally prohibitive for high-resolution synthesis. To res

  71. Georgi Bebrov

    Here we concerned with quantum key distribution - a way to establish common cryptographic key between several parties. The work proposes a combination between quantum key distribution and systematic polar coding (an error correction algorithm) frameworks - quantum key distribution based on systematic polar coding. This results in obtaining key rates greater

  72. Said Laaroua

    We develop an effective mapping for low-redshift photon propagation that captures the leading path-dependent deviations from the standard FLRW redshift. Instead of relying on exact integrations of the Sachs optical equations, we introduce a minimal deformation of the redshift relation z_eff(z) = z minus alpha times f(z), where alpha is a small amplitude and

  73. Siyuan Wei, Chunjie Wang, Xiao Liu, Xiaosheng Yan

    3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leave two key ambiguities unresolved: viewpoint ambiguity, where spatial language presumes unknown camera poses, and object referring ambiguity,

  74. Nimeshika Udayangani, Sarah Erfani, Christopher Leckie

    Out-of-Distribution (OOD) detection in semantic segmentation aims to localize anomalous regions at the pixel level, advancing beyond traditional image-level OOD techniques to better suit real-world applications such as autonomous driving. Recent literature has successfully explored the adaptation of commonly used image-level OOD methods--primarily based on c

  75. Onat Gungor, Roshan Sood, Jiasheng Zhou, Tajana Rosing

    Large Language Models (LLMs) are highly effective for cybersecurity question answering (QA) but are difficult to deploy on edge devices due to their size. Quantization reduces memory and compute requirements but often degrades accuracy and increases vulnerability to adversarial attacks. We present EAGER, an edge-aligned defense framework that integrates para

  76. Yoichi Izunaga, Kota Kurihara, Hokuto Nagano, Daiki Uchida

    We analyze the axiomatic properties of a class of probability estimators derived from Distributionally Robust Optimization (DRO) with $q$-norm ambiguity sets ($q$-DRO), a principled approach to the zero-frequency problem. While classical estimators such as Laplace smoothing are characterized by strong linearity axioms like Ratio Preservation, we show that $q

  77. Jiawei Hou, Shenghao Zhang, Can Wang, Zheng Gu

    Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on complex multi-stage pipelines that are prone to error propagation

  78. Congren Dai, Yue Yang, Krinos Li, Huichi Zhou

    Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains insufficiently examined. We introduce Musical Score Understanding Benchmark (MSU-Bench), a human-curated benchmark for score-

  79. Sing-Yuan Yeh, Chun-Hao Yang

    Persistent homology (PH) is a crucial concept in computational topology, providing a multiscale topological description of a space. It is particularly significant in topological data analysis, which aims to make statistical inference from a topological perspective. In this work, we introduce a new topological summary for Bayesian neural networks, termed the

  80. Yu Jia, Chengyu Wu, Hao Wu, Jiaqing Yang

    In this paper, we investigate the inverse Stokes problem of determining a discontinuous viscosity coefficient $\mu$ in a bounded domain $\Omega\subset\mathbb{R}^3$. By analyzing the singularity of the Dirichlet Green's functions in $H^1$-norm and constructing a specifically coupled Stokes-Brinkman system in a localized domain, we prove a global uniqueness th

  81. Yuqiu Jiang, Xiaozhen Qiao, Yifan Chen, Ye Zheng

    Human-Object Interaction (HOI) detection is a fundamental task in computer vision, empowering machines to comprehend human-object relationships in diverse real-world scenarios. Recent advances in VLMs have significantly improved HOI detection by leveraging rich cross-modal representations. However, most existing VLM-based approaches rely heavily on additiona

  82. Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang

    Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extending them to multi-skill settings remains challenging: directly merging VLA experts trained on different tasks results in near-zero success r

  83. Yezheng Gao

    In this paper, we establish several inequalities comparing formal slopes with p-adic slopes of solvable differential modules over the punctured open unit disc. Our approach is based on a delicate analysis of Newton polygons and the log-convexity of generic radius functions.

  84. Linxiao Cao, Ruitao Wang, Jindong Li, Zhipeng Zhou

    Retrieval-augmented generation (RAG) enables large language models (LLMs) to access external knowledge, helping mitigate hallucinations and enhance domain-specific expertise. Graph-based RAG enhances structural reasoning by introducing explicit relational organization that enables information propagation across semantically connected text units. However, the

  85. Chun-Meng Tang, Chun-Gui Duan, Liang Tang, Cong-Feng Qiao

    In this work, we present a systematic calculation of the mass spectrum for tetraquark hybrid states, focusing on the $8_{[c\bar{c}]}\otimes 8_{[G]}\otimes 8_{[c\bar{c}]}$ color configuration, within the framework of QCD sum rules. As an extension of our previous work on $0^{++}$ and $0^{-+}$ states, we now construct 18 distinct interpolating currents with $J

  86. A. S. Yurkov

    It makes sense to consider a helical waveguide with a fine pitch approximately, replacing the turns with anisotropic conductivity: infinite along the turns and zero across them. This approach has been known for a long time, but calculation formulas within it have only been obtained for the case where the winding does not contain a dielectric core. This paper

  87. Qinglei Cao, Ziyao Tang, Xiaoqin Tang

    X-ray imaging, based on penetration, enables detailed visualization of internal structures. Building on this capability, existing implicit 3D reconstruction methods have adapted the NeRF model and its variants for internal CT reconstruction. However, these approaches often neglect the significance of objects' anatomical priors for implicit learning, limiting

  88. Sreesritha Sai, Sai Venkata Suma Sreeja, Sai Sri Deepthi, Nikhil

    Accurate assessment of post-disaster damage is essential for prioritizing emergency response, yet current practices rely heavily on manual interpretation of satellite imagery.This approach is time-consuming, subjective, and difficult to scale during large-area disasters. Although recent deep-learning models for semantic segmentation and change detection have

  89. Yi Xu, Chaofan Fan, Jinxin Hu, Yu Zhang

    Ranking models have become an important part of modern personalized recommendation systems. However, significant challenges persist in handling high-cardinality, heterogeneous, and sparse feature spaces, particularly regarding model scalability and efficiency. We identify two key bottlenecks: (i) Representation Bottleneck: Driven by the high cardinality and

  90. Takayuki Sakuma

    We study a \emph{QDisCoCirc}-inspired, chunked diagram-to-circuit quantum natural language processing (QNLP) model for three-class sentiment classification of financial texts. In our classical simulations, we keep the Hilbert-space dimension manageable by decomposing each sentence into short contiguous chunks. Each chunk is mapped to a shallow quantum circui

  91. Farshad Darabi, Juan Ruben Gomez-Solano

    We investigate experimentally the single-particle motion in water of silica colloidal beads half-coated with carbon under the action of a converging laser beam. The beads are self-propelled in this medium by means of self-thermophoresis resulting from local heating as a result of light absorption by their carbon cap. Within a certain laser power range, we fi

  92. Zhaoyuan Meng, Leyu Chen, Jin-Peng Liu, Guowei He

    We propose an end-to-end quantum algorithm to simulate rapidly distorted turbulence via linear combination of Hamiltonian (LCHS). The algorithm comprises three primary stages: the efficient preparation of an initial turbulent state with a prescribed energy spectrum, its subsequent time evolution via LCHS, and the direct measurement of key turbulence statisti

  93. Matthew Hampsey, Pieter van Goor, Ravi Banavar, Robert Mahony

    Mechanical control systems such as aerial, marine, space, and terrestrial robots often naturally admit a state-space that has the structure of a Lie group. The kinetic energy of such systems is commonly invariant to the induced action by the Lie group, and the system dynamics can be written as a coupled ordinary differential equation on the group and the dua

  94. Chengyu Wu, Yushan Xue, Jiaqing Yang

    In this paper, we present the first well-posedness result for elastic scattering by locally rough interfaces in both two and three dimensions. Inspired by the Helmholtz decomposition, we discover a fundamental identity for the stress vector, revealing an intrinsic relationship among the generalized stress vector, the Lame constants and certain tangential dif

  95. Dinesh Kumar

    We derive a simple sufficient condition for the local asymptotic stability of spatially discrete, continuous-time reaction-diffusion systems of networked dynamical systems at a homogeneous equilibrium point. The framework explicitly accommodates \emph{heterogeneous} local dynamics -- patches at different nodes governed by structurally distinct functional for

  96. Jessalyn N. Sebastian, Volodymyr M. Minin

    Many quantities characterizing infectious disease outbreaks - like the effective reproduction number ($R_t$), defined as the average number of secondary infections a newly infected individual will cause over the course of their infection - need to be modeled as time-varying parameters. It is common practice to use Gaussian random walks as priors for estimati

  97. Arindam Chakraborty

    The article presents various Witt type vector field realizations of 2-D Cayley-Klein algebras with non-vanishing curvatures. The expressions of the vector fields involve Jacobi elliptic functions whose moduli are directly related to the parameters that appear in the corresponding matrix representation obtained from a bi-orthogonal set of vectors. First, the

  98. Roger Guzman, Ján Rusz, Ang Li, Juan Carlos Idrobo

    X-ray linear dichroism has been pivotal for probing electronic anisotropies, but its inherent limited spatial resolution precludes atomic-scale investigations of orbital polarization. Here we introduce a versatile electron linear dichroism methodology in scanning transmission electron microscopy that overcomes these constraints. By exploiting momentum-transf

  99. Zi-Dong Zhang, Zhen-Hui Qin, Yi-Han He, Yun-Fei Cheng

    High overtone bulk acoustic resonators are essential components in microwave signal processing and emerging quantum technologies; however, conventional designs suffer from limited impedance matching, spurious mode interference, and restricted scalability. Here we introduce a laterally excited high overtone thickness shear bulk acoustic resonator, abbreviated

  100. Yejing Wang, Shengyu Zhou, Jinyu Lu, Ziwei Liu

    Generative Recommendation (GR), powered by Large Language Models (LLMs), represents a promising new paradigm for industrial recommender systems. However, their practical application is severely hindered by high inference latency, which makes them infeasible for high-throughput, real-time services and limits their overall business impact. While Speculative De