November 2025 arXiv papers — page 53
Showing 5,201–5,300 of 22,271 papers
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
cs.CLJingqian Zhao, Bingbing Wang, Geng Tu, Yice Zhang
Data contamination poses a significant challenge to the fairness of LLM evaluations in natural language processing tasks by inadvertently exposing models to test data during training. Current studies attempt to mitigate this issue by modifying existing datasets or generating new ones from freshly collected information. However, these methods fall short of en
MFmamba: A Multi-function Network for Panchromatic Image Resolution Restoration Based on State-Space Model
cs.CVQian Jiang, Qianqian Wang, Xin Jin, Michal Wozniak
Remote sensing images are becoming increasingly widespread in military, earth resource exploration. Because of the limitation of a single sensor, we can obtain high spatial resolution grayscale panchromatic (PAN) images and low spatial resolution color multispectral (MS) images. Therefore, an important issue is to obtain a color image with high spatial resol
Guangyuan Li, Bo Li, Jinwei Chen, Xiaobin Hu
Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with multiple interacting subjects, where dynamic subjects fail to follow realistic motion patterns during scene evolution. Secon
Sudipta Ghosh, Mike Miller Eismeier
We determine the framed instanton homology with coefficients in $\mathbb F = \mathbb Z/2$ for Dehn surgeries on a knot in the $3$-sphere. The dimension of these groups is seen to have a close relationship with a homology cobordism invariant due to Froyshov. As an application, we show that $r$-surgery on a non-trivial knot cannot be nondegenerate $SU(2)$-abel
Jihun Park, Junyong Shin, Jinsung Park, Yo-Seb Jeon
This paper proposes robust nonlinear transform coding (Robust-NTC), a generalizable digital joint source-channel coding (JSCC) framework that couples variational latent modeling with channel-adaptive transmission. Unlike learning-based JSCC methods that implicitly absorb channel variations, Robust-NTC explicitly models element-wise latent distributions via a
Richard Golnik, Thomas Gatter, Peter F. Stadler, Nicola Vassena
Autocatalysis is an important feature of metabolic networks, contributing crucially to the self-maintenance of organisms. Autocatalytic subsystems of chemical reaction networks (CRNs) are characterized in terms of algebraic conditions on submatrices of the stoichiometric matrix. Here, we derive sufficient conditions for subgraphs supporting irreducible autoc
Thermal stability originates the vanishing of the specific heats at the absolute zero
cond-mat.stat-mechMartín-Olalla, José María
The relationship between the vanishing of the heat capacities as $T\to0^+$ and the thermal stability is examined. The heat capacities vanish as fast as or faster than $T$ as $T\to0^+$ for states at the phase space boundary ($T=0$) to sustain the standard thermal stability criterion $U_{ss}>0$. Conversely, weakly vanishing heat capacities, which signify a los
Ayca Duran, Christoph Waibel, Bernd Bickel, Iro Armeni
Building integrated photovoltaic (BIPV) facades represent a promising pathway towards urban decarbonization, especially where roof areas are insufficient and ground-mounted arrays are infeasible. Although machine learning-based approaches to support photovoltaic (PV) planning on rooftops are well researched, automated approaches for facades still remain scar
Measurement of Milli-Charged Particles with a running electromagnetic coupling constant at IceCube
astro-ph.HEYe Xu
It is postulated that heavy dark matter $\phi$ with a mass on the order of TeV, once captured by the Earth, can decay into relativistic milli-charged particles (MCPs). These MCPs are potentially detectable at the IceCube neutrino telescope. In this study, MCPs are modeled within the massless hidden photon framework, where they interact with nuclei via a runn
Christoph Brause, Dieter Rautenbach, Laurin Schwartze
Kamyczura introduced the notion of a majority additive $k$-coloring of a graph $G$ as a function $c: V(G) \to \{1,2,\ldots,k\}$ such that $$\left|\left\{u \in N_G(v):\sum_{w \in N_G(u)} c(w) = s \right\}\right|\leq \max\left\{1,\frac{d_G(v)}{2}\right\}$$ for every vertex $v$ of $G$ and every positive integer $s$. We show that every graph $G$ of maximum degre
Bowen Duan, Ge Zhang
When liquids are cooled rapidly, they bypass crystallization and instead enter a supercooled state and then a glass state. Previous studies have shown that the static structure factors of high-temperature liquids, supercooled liquids, and glasses exhibit only subtle differences, leading to the conclusion that the glass transition cannot be predicted solely f
Suzie Kim, Hye-Bin Shin, Hyo-Jeong Jang
In this work, we investigate how implicit neural feed back can accelerate reinforcement learning in complex robotic manipulation settings. While prior electroencephalogram (EEG) guided reinforcement learning studies have primarily focused on navigation or low-dimensional locomotion tasks, we aim to understand whether such neural evaluative signals can improv
Geunho Noh, Panayotis Benetatos
We study two cross-linked polymer systems in the strong stretching regime. The first consists of two polymers sharing one endpoint, with the other two endpoints coupled by a harmonic potential. Within the weakly bending approximation, we analyze the tensile elastic response for freely jointed or wormlike chains; for the latter, the approximation applies eith
Colin Faverjon, Marina Poulet
Mahler equations arise in a wide range of contexts including the study of finite automata, regular sequences, algebraic series over Fp(z), and periods of Drinfeld modules. Introduced a century ago by K. Mahler to study the transcendence of certain complex numbers, they have recently been the subject of several works establishing a deep connection between suc
Fairness Meets Privacy: Integrating Differential Privacy and Demographic Parity in Multi-class Classification
stat.MLLilian Say, Christophe Denis, Rafael Pinot
The increasing use of machine learning in sensitive applications demands algorithms that simultaneously preserve data privacy and ensure fairness across potentially sensitive sub-populations. While privacy and fairness have each been extensively studied, their joint treatment remains poorly understood. Existing research often frames them as conflicting objec
Wengyi Zhan, Mingbao Lin, Zhihang Lin, Rongrong Ji
Multimodal large language models (MLLMs) deliver impressive vision-language reasoning but suffer steep inference latency because self-attention scales quadratically with sequence length and thousands of visual tokens contributed by high-resolution images. Naively pruning less-informative visual tokens reduces this burden, yet indiscriminate removal can strip
GContextFormer: A global context-aware hybrid multi-head attention approach with scaled additive aggregation for multimodal trajectory prediction
cs.AIYuzhi Chen, Yuanchang Xie, Lei Zhao, Pan Liu
Multimodal trajectory prediction generates multiple plausible future trajectories to address vehicle motion uncertainty from intention ambiguity and execution variability. However, HD map-dependent models suffer from costly data acquisition, delayed updates, and vulnerability to corrupted inputs, causing prediction failures. Map-free approaches lack global c
Neural Texture Splatting: Expressive 3D Gaussian Splatting for View Synthesis, Geometry, and Dynamic Reconstruction
cs.CVYiming Wang, Shaofei Wang, Marko Mihajlovic, Siyu Tang
3D Gaussian Splatting (3DGS) has emerged as a leading approach for high-quality novel view synthesis, with numerous variants extending its applicability to a broad spectrum of 3D and 4D scene reconstruction tasks. Despite its success, the representational capacity of 3DGS remains limited by the use of 3D Gaussian kernels to model local variations. Recent wor
H{\"o}lder regularity of parabolic equations with Dirichlet boundary conditions and application to reaction-diffusion and reaction-cross-diffusion systems
math.APHector Bouton, Laurent Desvillettes, Helge Dietert
In this work, we adapt our recent article [BDD25] to the setting of Dirichlet boundary conditions. A key part is the study of the parabolic equation $a\partial_t w - \Delta w = f$ with a rough coefficient $a$, homogeneous Dirichlet boundary conditions, and the special assumption $\partial_tw \ge 0$. We then apply it to prove existence of global strong soluti
Bing Wu, Chang Zou, Changlin Li, Duojun Huang
We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architec
Shuyang Liu, Yuan Jin, Rui Lin, Shizhe Chen
Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source multi-scale representations module to obtain complementary segment- and track-level features, (2) a hierarchical augmentation stra
Dezhi Ran, Shuxiao Xie, Mingfang Ji, Anmin Liu
High-performance GPU kernels are critical for efficient LLM serving, yet their optimization remains a bottleneck requiring deep system expertise. While code LLMs show promise in generating functionally correct code, kernel optimization is intrinsically a search problem over a vast optimization space. The fundamental mismatch prevents existing LLM agents from
Liutong Han, Chu Kang, Mingjie Xing, Yanjun Wu
Intrinsic functions are specialized functions provided by the compiler that efficiently operate on architecture-specific hardware, allowing programmers to write optimized code in a high-level language that fully exploits hardware features. Using intrinsics to vectorize core code blocks is a standard optimization method in high-performance libraries, often re
Schr\"{o}dinger-type $f(Q,T)$ gravity-nonmetricity driven cosmological evolution from inflation to the late Universe
gr-qcLei Ming, Himanshu Chaudhary, Shi-Dong Liang, Hong-Hao Zhang
We consider an $f(Q, T)$ gravity theory with a Schr\"{o}dinger type vectorial non-metricity. In the presence of such a non-metricity, the length of vectors is preserved under autoparallel transport. We obtain the field equations assuming a vanishing total scalar curvature, implemented by a Lagrange multiplier, and investigate their cosmological implications.
Yu Zhang, Haoan Ping, Yuchen Li, Zhenshan Bing
Recent salient object detection (SOD) methods aim to improve performance in four key directions: semantic enhancement, boundary refinement, auxiliary task supervision, and multi-modal fusion. In pursuit of continuous gains, these approaches have evolved toward increasingly sophisticated architectures with multi-stage pipelines, specialized fusion modules, ed
Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
cs.CLYang Xiang, Yixin Ji, Juntao Li, Min Zhang
Large Reasoning Models (LRMs) have demonstrated remarkable performance on complex reasoning benchmarks. However, their long chain-of-thought reasoning processes incur significant inference overhead. Pruning has emerged as a promising approach to reducing computational costs. However, existing efforts have primarily focused on large language models (LLMs), wh
Wai-Kit Lam, Arnab Sen
We study correlation decay for the maximum weight matching problem on sparse graphs with i.i.d. edge weights. We show exponential decay of correlations when the underlying graphs are locally tree-like with uniformly bounded degree and the edge weights are exponential. We also prove a polynomial rate of decay of correlations for any finite graph with maximum
Generating Reading Comprehension Exercises with Large Language Models for Educational Applications
cs.CLXingyu Huang, Fei Jiang, Jianli Xiao
With the rapid development of large language models (LLMs), the applications of LLMs have grown substantially. In the education domain, LLMs demonstrate significant potential, particularly in automatic text generation, which enables the creation of intelligent and adaptive learning content. This paper proposes a new LLMs framework, which is named as Reading
Vidi Team, Chia-Wen Kuo, Chuang Huang, Dawei Du
Video has emerged as the primary medium for communication and creativity on the Internet, driving strong demand for scalable, high-quality video production. Vidi models continue to evolve toward next-generation video creation and have achieved state-of-the-art performance in multimodal temporal retrieval (TR). In its second release, Vidi2 advances video unde
Bo Jiang, Weijun Zhao, Beibei Wang, Xiao Wang
Recently, fine-tuning large-scale pre-trained GNNs has yielded remarkable attention in adapting pre-trained GNN models for downstream graph learning tasks. One representative fine-tuning method is to exploit adapter (termed AdapterGNN) which aims to 'augment' the pre-trained model by inserting a lightweight module to make the 'augmented' model better adapt t
Rethinking Long-tailed Dataset Distillation: A Uni-Level Framework with Unbiased Recovery and Relabeling
cs.CVXiao Cui, Yulei Qin, Xinyue Li, Wengang Zhou
Dataset distillation creates a small distilled set that enables efficient training by capturing key information from the full dataset. While existing dataset distillation methods perform well on balanced datasets, they struggle under long-tailed distributions, where imbalanced class frequencies induce biased model representations and corrupt statistical esti
Changsheng Luo, Yushi Wang, Wenhan Cai, Mingguo Zhao
Accurate proprioceptive odometry is fundamental for legged robot navigation in GPS-denied and visually degraded environments where conventional visual odometry systems fail. Current approaches face critical limitations: analytical filtering methods suffer from modeling uncertainties and cumulative drift, hybrid learning-filtering approaches remain constraine
Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories
cs.RORushuai Yang, Zhiyuan Feng, Tianxiang Zhang, Kaixin Wang
Scaling vision-language-action (VLA) model pre-training requires large volumes of diverse, high-quality manipulation trajectories. Most current data is obtained via human teleoperation, which is expensive and difficult to scale. Reinforcement learning (RL) methods learn useful skills through autonomous exploration, making them a viable approach for generatin
Sana Alamgeer, Mylene Farias, Marcelo Carvalho
The main goal of the project is to design a new model that predicts regions of interest in 360$^{\circ}$ videos. The region of interest (ROI) plays an important role in 360$^{\circ}$ video streaming. For example, ROIs are used to predict view-ports, intelligently cut the videos for live streaming, etc so that less bandwidth is used. Detecting view-ports in a
Theoretical study on excited states of ICl+ molecular ion considering spin-orbit coupling
physics.atom-phRui Li, Ronglong Dou, Ting Gao, Qinan Li
The electronic structure of the ICl+ molecular ion is investigated by using high-level multireference configuration interaction (MRCI) method. To improve computational accuracy, Davidson corrections, spin-orbit coupling (SOC), and core-valence electron correlations effects are incorporated into the calculations. The potential energy curves (PECs) of 21 elect
Yujing Wang, Weize Hong
We present a novel framework that integrates Large Language Models (LLMs) into the Git bisect process for semantic fault localization. Traditional bisect assumes deterministic predicates and binary failure states assumptions often violated in modern software development due to flaky tests, nonmonotonic regressions, and semantic divergence from upstream repos
Sophie Béraud-Dufour, Ilona Legroux, Thierry Coppola, Patricia Lebrun
Stroke is a leading cause of disability and death worldwide, with ischemic strokes accounting for nearly 80% of cases. Fewer than 5% of patients receive the sole validated pharmacotherapy, intravenous thrombolysis, highlighting the urgent need for novel therapies. Within this landscape, the exploration of natural molecules emerges as a promising avenue, part
Masoomali Fatehkia, Enes Altinisik, Husrev Taha Sencar
Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of ove
Robust Long-term Test-Time Adaptation for 3D Human Pose Estimation through Motion Discretization
cs.CVYilin Wen, Kechuan Dong, Yusuke Sugano
Online test-time adaptation addresses the train-test domain gap by adapting the model on unlabeled streaming test inputs before making the final prediction. However, online adaptation for 3D human pose estimation suffers from error accumulation when relying on self-supervision with imperfect predictions, leading to degraded performance over time. To mitigate
Pre-Filtering Code Suggestions using Developer Behavioral Telemetry to Optimize LLM-Assisted Programming
cs.SEMohammad Nour Al Awad, Sergey Ivanov, Olga Tikhonova
Large Language Models (LLMs) are increasingly integrated into code editors to provide AI-powered code suggestions. Yet many of these suggestions are ignored, resulting in wasted computation, increased latency, and unnecessary interruptions. We introduce a lightweight pre-filtering model that predicts the likelihood of suggestion acceptance before invoking th
Václav Tran, Jakub Šmíd, Ladislav Lenc, Jean-Pierre Salmon
Text summarization is the task of automatically condensing longer texts into shorter, coherent summaries while preserving the original meaning and key information. Although this task has been extensively studied in English and other high-resource languages, Czech summarization, particularly in the context of historical documents, remains underexplored. This
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
cs.CVIshmam Tashdeed, Md. Atiqur Rahman, Sabrina Islam, Md. Azam Hossain
Personalized federated learning (PFL) possesses the unique capability of preserving data confidentiality among clients while tackling the data heterogeneity problem of non-independent and identically distributed (Non-IID) data. Its advantages have led to widespread adoption in domains such as medical image segmentation. However, the existing approaches mostl
Yubo Wang, Hui He, Chaoxi Niu, Zhendong Niu
Due to the inherent complexity, temporal patterns in real-world time series often evolve across multiple intertwined scales, including long-term periodicity, short-term fluctuations, and abrupt regime shifts. While existing literature has designed many sophisticated decomposition approaches based on the time or frequency domain to partition trend-seasonality
Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu
Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such metho
Iona Ann Sebastian, S. M. Sunoj
Fractional cumulative residual inaccuracy (FCRI) measure allows to determine regions of discrepancy between systems, depending on their respective fractional and chaotic map parameters. Most of the theoretical results and applications related to the FCRI of the lifetime random variable are based on the distribution function approach. However, there are situa
Heger Arfaoui, Mohammed Iheb Hergli, Beya Benzina, Slimane BenMiled
Focus group discussions generate rich qualitative data but their analysis traditionally relies on labor-intensive manual coding that limits scalability and reproducibility. We present a systematic framework for applying BERTopic to focus group transcripts using data from ten focus groups exploring HPV vaccine perceptions in Tunisia (1,075 utterances). We con
Mohammad Nour Al Awad, Sergey Ivanov, Olga Tikhonova
Large Language Models (LLMs) have transformed code auto-completion by generating context-aware suggestions. Yet, deciding when to present these suggestions remains underexplored, often leading to interruptions or wasted inference calls. We propose an adaptive timing mechanism that dynamically adjusts the delay before offering a suggestion based on real-time
Mincheol Jeon, Euinam Huh
Personalized Federated Learning (PFL) faces persistent challenges, including domain heterogeneity from diverse client data, data imbalance due to skewed participation, and strict communication constraints. Traditional federated learning often lacks personalization, as a single global model cannot capture client-specific characteristics, leading to biased pre
Hongyu Lyu, Thomas Monninger, Julie Stephany Berrio Perez, Mao Shan
Autonomous driving systems benefit from high-definition (HD) maps that provide critical information about road infrastructure. The online construction of HD maps offers a scalable approach to generate local maps from on-board sensors. However, existing methods typically rely on costly 3D map annotations for training, which limits their generalization and sca
Binglin Liu, Yucheng Wang, Zheyuan Zhang, Jiyuan Lu
The adaptation of teaching slides to instructors' situated teaching needs, including pedagogical styles and their students' context, is a critical yet time-consuming task for educators. Through a series of educator interviews, we first identify and systematically categorize the key friction points that impede this adaptation process. Grounded in these findin
Enhancing Multi-Label Thoracic Disease Diagnosis with Deep Ensemble-Based Uncertainty Quantification
cs.CVYasiru Laksara, Uthayasanker Thayasivam
The utility of deep learning models, such as CheXNet, in high stakes clinical settings is fundamentally constrained by their purely deterministic nature, failing to provide reliable measures of predictive confidence. This project addresses this critical gap by integrating robust Uncertainty Quantification (UQ) into a high performance diagnostic platform for
Xiaofan Li, Chenming Wu, Yanpeng Sun, Jiaming Zhou
Visual autoregressive models achieve remarkable generation quality through next-scale predictions across multi-scale token pyramids. However, the conventional method uses uniform scale downsampling to build these pyramids, leading to aliasing artifacts that compromise fine details and introduce unwanted jaggies and moir\'e patterns. To tackle this issue, we
Wenxin He, Bin Xu
In this paper, we study the complex structures of complete hyperk\"ahler four-manifolds of infinite topological type arising from the Gibbons-Hawking ansatz. We show that for almost all complex structures in the hyperk\"ahler family, the manifold is biholomorphic to a hypersurface in $\mathbb{C}^3$ defined by an explicit entire function. For the remaining co
Fang Wang, Lance Kosca, Adrienne Kosca, Marko Gacesa
This paper introduces HGNN(O), an AutoML GNN hypermodel framework for outcome prediction on event-sequence data. Building on our earlier work on graph convolutional network hypermodels, HGNN(O) extends four architectures-One Level, Two Level, Two Level Pseudo Embedding, and Two Level Embedding-across six canonical GNN operators. A self-tuning mechanism based
Lei Ke, Hubery Yin, Gongye Liu, Zhengyao Lv
With the success of flow matching in visual generation, sampling efficiency remains a critical bottleneck for its practical application. Among flow models' accelerating methods, ReFlow has been somehow overlooked although it has theoretical consistency with flow matching. This is primarily due to its suboptimal performance in practical scenarios compared to
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
cs.CVJonathan Lee, Xingrui Wang, Jiawei Peng, Luoxin Ye
We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes to support goal-directed reasoning. While this form of reasoning is fundamental to human cognition, current vision-lang
PrismAudio: Decomposed Chain-of-Thoughts and Multi-dimensional Rewards for Video-to-Audio Generation
cs.SDHuadai Liu, Kaicheng Luo, Wen Wang, Qian Chen
Video-to-Audio (V2A) generation requires balancing four critical perceptual dimensions: semantic consistency, audio-visual temporal synchrony, aesthetic quality, and spatial accuracy; yet existing methods suffer from objective entanglement that conflates competing goals in single loss functions and lack human preference alignment. We introduce PrismAudio, th
Shivam Pal, Sakshi Varshney, Piyush Rai
Deep neural networks are prone to learning shortcuts, spurious correlations present in the training data that undermine out-of-distribution (OOD) generalization. Most prior work mitigates shortcut learning through input-space reweighting, either relying on explicit shortcut labels or inferring shortcut structure from heuristics such as per-sample loss. Moreo
Kaize Shi, Xueyao Sun, Xiaohui Tao, Lin Li
Large Language Models (LLMs) face information overload when handling long contexts, particularly in Retrieval-Augmented Generation (RAG) where extensive supporting documents often introduce redundant content. This issue not only weakens reasoning accuracy but also increases computational overhead. We propose an unsupervised context compression framework that
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
cs.CVShaobo Wang, Tianle Niu, Runkang Yang, Deshan Liu
The scalability of video understanding models is increasingly limited by the prohibitive storage and computational costs of large-scale video datasets. While data synthesis has improved data efficiency in the image domain, its extension to video remains challenging due to pervasive temporal redundancy and complex spatiotemporal dynamics. In this work, we unc
Leveraging Duration Pseudo-Embeddings in Multilevel LSTM and GCN Hypermodels for Outcome-Oriented PPM
cs.LGFang Wang, Paolo Ceravolo, Ernesto Damiani
Existing deep learning models for Predictive Process Monitoring (PPM) struggle with temporal irregularities, particularly stochastic event durations and overlapping timestamps, limiting their adaptability across heterogeneous datasets. We propose a dual input neural network strategy that separates event and sequence attributes, using a duration-aware pseudo-
Kanav Arora, Girish Narayanswamy, Shwetak Patel, Richard Li
Heart rate estimation from photoplethysmography (PPG) signals generated by wearable devices such as smartwatches and fitness trackers has significant implications for the health and well-being of individuals. Although prior work has demonstrated deep learning models with strong performance in the heart rate estimation task, in order to deploy these models on
VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
cs.CVBoyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan
By leveraging tool-augmented Multimodal Large Language Models (MLLMs), multi-agent frameworks are driving progress in video understanding. However, most of them adopt static and non-learnable tool invocation mechanisms, which limit the discovery of diverse clues essential for robust perception and reasoning regarding temporally or spatially complex videos. T
Edgar Dobriban
Over the last few months, AI models including large language models have improved greatly. There are now several documented examples where they have helped professional mathematical scientists prove new results, sometimes even helping resolve known open problems. In this short note, we add another example to the list, by documenting how we were able to solve
Leveraging Metaheuristic Approaches to Improve Deep Learning Systems for Anxiety Disorder Detection
cs.CVMohammadreza Amiri, Monireh Hosseini
Despite being among the most common psychological disorders, anxiety-related conditions are still primarily identified through subjective assessments, such as clinical interviews and self-evaluation questionnaires. These conventional methods often require significant time and may vary depending on the evaluator. However, the emergence of advanced artificial
Aakash Gore, Anoushka Dey, Aryan Mishra
Knowledge distillation has emerged as a powerful technique for model compression, enabling the transfer of knowledge from large teacher networks to compact student models. However, traditional knowledge distillation methods treat all teacher predictions equally, regardless of the teacher's confidence in those predictions. This paper proposes an uncertainty-a
Xiele Wu, Zicheng Zhang, Mingtao Chen, Yixian Liu
Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and algorithms have been proposed, existing approaches are constrained by two limitations: the absence of systematic definitions for evaluation dimensions, and the isolated treatment of
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
cs.CVAlvin Wei Ming Tan, Jane Yang, Tarun Sepuri, Khai Loong Aw
Figuring out which objects or concepts words refer to is a central language learning challenge for young children. Most models of this process posit that children learn early object labels from co-occurrences of words and their referents that occur when someone around them talks about an object in the immediate physical environment. But how aligned in time a
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
cs.CVFufangchen Zhao, Liao Zhang, Daiqi Shi, Yuanjun Gao
We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient events in long videos. VideoPerceiver adopts a two-stage training framework. During supervised fine-tuning (SFT), we co
Zhennan Chen, Junwei Zhu, Xu Chen, Jiangning Zhang
Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from potential information loss and non-end-to-end training. In contrast, existing pixel space models bypass VAEs but are computationally prohibitive for high-resolution synthesis. To res
Georgi Bebrov
Here we concerned with quantum key distribution - a way to establish common cryptographic key between several parties. The work proposes a combination between quantum key distribution and systematic polar coding (an error correction algorithm) frameworks - quantum key distribution based on systematic polar coding. This results in obtaining key rates greater
Said Laaroua
We develop an effective mapping for low-redshift photon propagation that captures the leading path-dependent deviations from the standard FLRW redshift. Instead of relying on exact integrations of the Sachs optical equations, we introduce a minimal deformation of the redshift relation z_eff(z) = z minus alpha times f(z), where alpha is a small amplitude and
Siyuan Wei, Chunjie Wang, Xiao Liu, Xiaosheng Yan
3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leave two key ambiguities unresolved: viewpoint ambiguity, where spatial language presumes unknown camera poses, and object referring ambiguity,
Nimeshika Udayangani, Sarah Erfani, Christopher Leckie
Out-of-Distribution (OOD) detection in semantic segmentation aims to localize anomalous regions at the pixel level, advancing beyond traditional image-level OOD techniques to better suit real-world applications such as autonomous driving. Recent literature has successfully explored the adaptation of commonly used image-level OOD methods--primarily based on c
EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
cs.CROnat Gungor, Roshan Sood, Jiasheng Zhou, Tajana Rosing
Large Language Models (LLMs) are highly effective for cybersecurity question answering (QA) but are difficult to deploy on edge devices due to their size. Quantization reduces memory and compute requirements but often degrades accuracy and increases vulnerability to adversarial attacks. We present EAGER, an edge-aligned defense framework that integrates para
An Axiomatic Analysis of Distributionally Robust Optimization with $q$-Norm Ambiguity Sets for Probability Smoothing
math.OCYoichi Izunaga, Kota Kurihara, Hokuto Nagano, Daiki Uchida
We analyze the axiomatic properties of a class of probability estimators derived from Distributionally Robust Optimization (DRO) with $q$-norm ambiguity sets ($q$-DRO), a principled approach to the zero-frequency problem. While classical estimators such as Laplace smoothing are characterized by strong linearity axioms like Ratio Preservation, we show that $q
Jiawei Hou, Shenghao Zhang, Can Wang, Zheng Gu
Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on complex multi-stage pipelines that are prone to error propagation
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
cs.SDCongren Dai, Yue Yang, Krinos Li, Huichi Zhou
Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains insufficiently examined. We introduce Musical Score Understanding Benchmark (MSU-Bench), a human-curated benchmark for score-
Sing-Yuan Yeh, Chun-Hao Yang
Persistent homology (PH) is a crucial concept in computational topology, providing a multiscale topological description of a space. It is particularly significant in topological data analysis, which aims to make statistical inference from a topological perspective. In this work, we introduce a new topological summary for Bayesian neural networks, termed the
Recovering discontinuous viscosity coefficients for inverse Stokes problems by boundary measurements
math.APYu Jia, Chengyu Wu, Hao Wu, Jiaqing Yang
In this paper, we investigate the inverse Stokes problem of determining a discontinuous viscosity coefficient $\mu$ in a bounded domain $\Omega\subset\mathbb{R}^3$. By analyzing the singularity of the Dirichlet Green's functions in $H^1$-norm and constructing a specifically coupled Stokes-Brinkman system in a localized domain, we prove a global uniqueness th
Yuqiu Jiang, Xiaozhen Qiao, Yifan Chen, Ye Zheng
Human-Object Interaction (HOI) detection is a fundamental task in computer vision, empowering machines to comprehend human-object relationships in diverse real-world scenarios. Recent advances in VLMs have significantly improved HOI detection by leveraging rich cross-modal representations. However, most existing VLM-based approaches rely heavily on additiona
Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang
Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extending them to multi-skill settings remains challenging: directly merging VLA experts trained on different tasks results in near-zero success r
Yezheng Gao
In this paper, we establish several inequalities comparing formal slopes with p-adic slopes of solvable differential modules over the punctured open unit disc. Our approach is based on a delicate analysis of Newton polygons and the log-convexity of generic radius functions.
Linxiao Cao, Ruitao Wang, Jindong Li, Zhipeng Zhou
Retrieval-augmented generation (RAG) enables large language models (LLMs) to access external knowledge, helping mitigate hallucinations and enhance domain-specific expertise. Graph-based RAG enhances structural reasoning by introducing explicit relational organization that enables information propagation across semantically connected text units. However, the
QCD sum rule predictions on gluonic tetraquark states with $J^{PC}=0^{+-},0^{--}$ and $1^{\pm \pm}$
hep-phChun-Meng Tang, Chun-Gui Duan, Liang Tang, Cong-Feng Qiao
In this work, we present a systematic calculation of the mass spectrum for tetraquark hybrid states, focusing on the $8_{[c\bar{c}]}\otimes 8_{[G]}\otimes 8_{[c\bar{c}]}$ color configuration, within the framework of QCD sum rules. As an extension of our previous work on $0^{++}$ and $0^{-+}$ states, we now construct 18 distinct interpolating currents with $J
A. S. Yurkov
It makes sense to consider a helical waveguide with a fine pitch approximately, replacing the turns with anisotropic conductivity: infinite along the turns and zero across them. This approach has been known for a long time, but calculation formulas within it have only been obtained for the case where the winding does not contain a dielectric core. This paper
Qinglei Cao, Ziyao Tang, Xiaoqin Tang
X-ray imaging, based on penetration, enables detailed visualization of internal structures. Building on this capability, existing implicit 3D reconstruction methods have adapted the NeRF model and its variants for internal CT reconstruction. However, these approaches often neglect the significance of objects' anatomical priors for implicit learning, limiting
Sreesritha Sai, Sai Venkata Suma Sreeja, Sai Sri Deepthi, Nikhil
Accurate assessment of post-disaster damage is essential for prioritizing emergency response, yet current practices rely heavily on manual interpretation of satellite imagery.This approach is time-consuming, subjective, and difficult to scale during large-area disasters. Although recent deep-learning models for semantic segmentation and change detection have
STORE: Semantic Tokenization, Orthogonal Rotation and Efficient Attention for Scaling Up Ranking Models
cs.IRYi Xu, Chaofan Fan, Jinxin Hu, Yu Zhang
Ranking models have become an important part of modern personalized recommendation systems. However, significant challenges persist in handling high-cardinality, heterogeneous, and sparse feature spaces, particularly regarding model scalability and efficiency. We identify two key bottlenecks: (i) Representation Bottleneck: Driven by the high cardinality and
Takayuki Sakuma
We study a \emph{QDisCoCirc}-inspired, chunked diagram-to-circuit quantum natural language processing (QNLP) model for three-class sentiment classification of financial texts. In our classical simulations, we keep the Hilbert-space dimension manageable by decomposing each sentence into short contiguous chunks. Each chunk is mapped to a shallow quantum circui
Farshad Darabi, Juan Ruben Gomez-Solano
We investigate experimentally the single-particle motion in water of silica colloidal beads half-coated with carbon under the action of a converging laser beam. The beads are self-propelled in this medium by means of self-thermophoresis resulting from local heating as a result of light absorption by their carbon cap. Within a certain laser power range, we fi
Zhaoyuan Meng, Leyu Chen, Jin-Peng Liu, Guowei He
We propose an end-to-end quantum algorithm to simulate rapidly distorted turbulence via linear combination of Hamiltonian (LCHS). The algorithm comprises three primary stages: the efficient preparation of an initial turbulent state with a prescribed energy spectrum, its subsequent time evolution via LCHS, and the direct measurement of key turbulence statisti
Matthew Hampsey, Pieter van Goor, Ravi Banavar, Robert Mahony
Mechanical control systems such as aerial, marine, space, and terrestrial robots often naturally admit a state-space that has the structure of a Lie group. The kinetic energy of such systems is commonly invariant to the induced action by the Lie group, and the system dynamics can be written as a coupled ordinary differential equation on the group and the dua
Chengyu Wu, Yushan Xue, Jiaqing Yang
In this paper, we present the first well-posedness result for elastic scattering by locally rough interfaces in both two and three dimensions. Inspired by the Helmholtz decomposition, we discover a fundamental identity for the stress vector, revealing an intrinsic relationship among the generalized stress vector, the Lame constants and certain tangential dif
Dinesh Kumar
We derive a simple sufficient condition for the local asymptotic stability of spatially discrete, continuous-time reaction-diffusion systems of networked dynamical systems at a homogeneous equilibrium point. The framework explicitly accommodates \emph{heterogeneous} local dynamics -- patches at different nodes governed by structurally distinct functional for
Jessalyn N. Sebastian, Volodymyr M. Minin
Many quantities characterizing infectious disease outbreaks - like the effective reproduction number ($R_t$), defined as the average number of secondary infections a newly infected individual will cause over the course of their infection - need to be modeled as time-varying parameters. It is common practice to use Gaussian random walks as priors for estimati
Arindam Chakraborty
The article presents various Witt type vector field realizations of 2-D Cayley-Klein algebras with non-vanishing curvatures. The expressions of the vector fields involve Jacobi elliptic functions whose moduli are directly related to the parameters that appear in the corresponding matrix representation obtained from a bi-orthogonal set of vectors. First, the
Roger Guzman, Ján Rusz, Ang Li, Juan Carlos Idrobo
X-ray linear dichroism has been pivotal for probing electronic anisotropies, but its inherent limited spatial resolution precludes atomic-scale investigations of orbital polarization. Here we introduce a versatile electron linear dichroism methodology in scanning transmission electron microscopy that overcomes these constraints. By exploiting momentum-transf
Impedance-matched High-Overtone Thickness-Shear Bulk Acoustic Resonators with Scalable Mode Volume
cond-mat.mes-hallZi-Dong Zhang, Zhen-Hui Qin, Yi-Han He, Yun-Fei Cheng
High overtone bulk acoustic resonators are essential components in microwave signal processing and emerging quantum technologies; however, conventional designs suffer from limited impedance matching, spurious mode interference, and restricted scalability. Here we introduce a laterally excited high overtone thickness shear bulk acoustic resonator, abbreviated
Yejing Wang, Shengyu Zhou, Jinyu Lu, Ziwei Liu
Generative Recommendation (GR), powered by Large Language Models (LLMs), represents a promising new paradigm for industrial recommender systems. However, their practical application is severely hindered by high inference latency, which makes them infeasible for high-throughput, real-time services and limits their overall business impact. While Speculative De