Skip to content

May 2025 arXiv papers — page 35

Showing 3,4013,500 of 24,552 papers

  1. Mattias Linde, Daniel Lindmark, Sandra Ålstig, Martin Servin

    We present a simulation framework for lunar construction work involving multiple autonomous machines. The framework supports modelling of construction scenarios and autonomy solutions, execution of the scenarios in simulation, and analysis of work time and energy consumption throughout the construction project. The simulations are based on physics-based mode

  2. Tristan S. W. Stevens, Oisín Nolan, Oudom Somphone, Jean-Luc Robert

    Three-dimensional ultrasound enables real-time volumetric visualization of anatomical structures. Unlike traditional 2D ultrasound, 3D imaging reduces reliance on precise probe orientation, potentially making ultrasound more accessible to clinicians with varying levels of experience and improving automated measurements and post-exam analysis. However, achiev

  3. San Jiang, Kan You, Ruqin Zhou, Xing Zhang

    Feature matching dominates the time costs in structure from motion (SfM). The primary contribution of this study is a GPU data schedule algorithm for efficient feature matching of Unmanned aerial vehicle (UAV) images. The core idea is to divide the whole dataset into blocks based on matrix band reduction (MBR) and achieve efficient feature matching via GPU-a

  4. Sam O'Connor Russell, Naomi Harte

    Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% i

  5. Xiaoqing Cheng, Ruizhe Chen, Hongying Zan, Yuxiang Jia

    Mitigating social bias in large language models (LLMs) has become an increasingly important research objective. However, existing debiasing methods often incur high human and computational costs, exhibit limited effectiveness, and struggle to scale to larger models and open-ended generation tasks. To address these limitations, this paper proposes BiasFilter,

  6. Ruxiao Chen, Dezheng Han, Wenjie Han, Shuaishuai Guo

    Assistive systems for visually impaired individuals must deliver rapid, interpretable, and adaptive feedback to facilitate real-time navigation. Current approaches face a trade-off between latency and semantic richness: natural language-based systems provide detailed guidance but are too slow for dynamic scenarios, while emergent communication frameworks off

  7. Runkai Li, Jia Xiong, Xi Wang

    High-Level Synthesis (HLS) serves as an agile hardware development tool that streamlines the circuit design by abstracting the register transfer level into behavioral descriptions, while allowing designers to customize the generated microarchitectures through optimization directives. However, the combinatorial explosion of possible directive configurations y

  8. Arnulf Jentzen, Julian Kranz, Adrian Riekert

    Averaging techniques such as Ruppert--Polyak averaging and exponential movering averaging (EMA) are powerful approaches to accelerate optimization procedures of stochastic gradient descent (SGD) optimization methods such as the popular ADAM optimizer. However, depending on the specific optimization problem under consideration, the type and the parameters for

  9. Elisa Ancarani, Julie Tores, Lucile Sassatelli, Hui-Yin Wu

    Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification and introduce a new AI task to the ML community: characterize and quantify complex multimodal (visual, speech, audio) tem

  10. H. L. Dao

    In this work, we introduce the first type of non-Euclidean neural quantum state (NQS) ansatz, in the form of the hyperbolic GRU (a variant of recurrent neural networks (RNNs)), to be used in the Variational Monte Carlo method of approximating the ground state energy for quantum many-body systems. In particular, we examine the performances of NQS ansatzes con

  11. Maret Einasto

    The richest and largest structures in the cosmic web are galaxy superclusters, their complexes (associations of several almost connected very rich superclusters), and planes. Superclusters represent a special environment where the evolution of galaxies and galaxy groups and clusters differs from the evolution of these systems in a low-density environment. Th

  12. Shun Sato, Issei Sato

    Mathematical expressions play a central role in scientific discovery. Symbolic regression aims to automatically discover such expressions from given numerical data. Recently, Neural symbolic regression (NSR) methods that involve Transformers pre-trained on synthetic datasets have gained attention for their fast inference, but they often perform poorly, espec

  13. Tomoo Kikuchi, Lien Pham

    We develop a model where currency issuers provide liquidity, while users in a trade network choose currency usage for trade settlement. We identify a feedback mechanism where a user's currency preference spillovers to others and increases the issuer's commitment to liquidity provision, which in turn increases the adoption of the currency. Our findings highli

  14. Hanbin Ko, Chang-Min Park

    The development of large-scale image-text pair datasets has significantly advanced self-supervised learning in Vision-Language Processing (VLP). However, directly applying general-domain architectures such as CLIP to medical data presents challenges, particularly in handling negations and addressing the inherent data imbalance of medical datasets. To address

  15. Pauline Vidal, Emily Bourne, Virginie Grandgirard, Michel Mehrenberger

    We present a semi-Lagrangian method for the numerical resolution of Vlasov-type equations on multi-patch meshes. Following N. Crouseilles et al. [A parallel Vlasov solver based on local cubic spline interpolation on patches. Journal of Computational Physics (2009)], we employ a local cubic spline interpolation with Hermite boundary conditions between the pat

  16. Indujan Sivanesarajaha, Leon Abelmann, Uwe Hartmann

    Amorphous sputtered Co-based thin films are widely used as soft magnetic materials in applications such as sensors, inductors and magnetic flux concentrators. The magnetic properties of these films can be controlled by deposition parameters like film thickness, argon pressure, deposition rate and others. In this study, we present a detailed investigation of

  17. Maja Stahl, Timon Ziegenbein, Joonsuk Park, Henning Wachsmuth

    Training large language models (LLMs) to follow instructions has significantly enhanced their ability to tackle unseen tasks. However, despite their strong generalization capabilities, instruction-following LLMs encounter difficulties when dealing with tasks that require domain knowledge. This work introduces a specialized instruction fine-tuning for the dom

  18. Xiaoxing Ren, Alessio Moreschini, Zhongda Chu, Yulong Gao

    In this paper, we develop a two-stage data-driven approach to address the adjustable robust optimization problem, where the uncertainty set is adjustable to manage infeasibility caused by significant or poorly quantified uncertainties. In the first stage, we synthesize an uncertainty set to ensure the feasibility of the problem as much as possible using the

  19. Coşku Can Horuz, Geoffrey Kasenbacher, Saya Higuchi, Sebastian Kairat

    Modeling sophisticated activation functions within deep learning architectures has evolved into a distinct research direction. Functions such as GELU, SELU, and SiLU offer smooth gradients and improved convergence properties, making them popular choices in state-of-the-art models. Despite this trend, the classical ReLU remains appealing due to its simplicity

  20. Megan Li, Wendy Bickersteth, Ningjing Tang, Jason Hong

    Due to its general-purpose nature, Generative AI is applied in an ever-growing set of domains and tasks, leading to an expanding set of risks of harm impacting people, communities, society, and the environment. These risks may arise due to failures during the design and development of the technology, as well as during its release, deployment, or downstream u

  21. Shujie HU, Xurong Xie, Mengzhe Geng, Jiajun Deng

    This paper proposes a novel MoE-based speaker adaptation framework for foundation models based dysarthric speech recognition. This approach enables zero-shot adaptation and real-time processing while incorporating domain knowledge. Speech impairment severity and gender conditioned adapter experts are dynamically combined using on-the-fly predicted speaker-de

  22. Hendra I. Nurdin

    An effective approach to modeling non-Markovian quantum systems is to embed a principal (quantum) system of interest into a larger quantum system. A widely employed embedding is one that uses another quantum system, referred to as the auxiliary system, which is coupled to the principal system, and both the principal and auxiliary can be coupled to quantum wh

  23. Longhao Li, Yangze Li, Hongfei Xue, Jie Liu

    CTC-based streaming ASR has gained significant attention in real-world applications but faces two main challenges: accuracy degradation in small chunks and token emission latency. To mitigate these challenges, we propose Delayed-KD, which applies delayed knowledge distillation on CTC posterior probabilities from a non-streaming to a streaming model. Specific

  24. Camilla Quaresmini, Giacomo Zanotti

    Automatic Gender Recognition (AGR) systems are an increasingly widespread application in the Machine Learning (ML) landscape. While these systems are typically understood as detecting gender, they often classify datapoints based on observable features correlated at best with either male or female sex. In addition to questionable binary assumptions, from an e

  25. Ran Li, Shimin Di, Yuchen Liu, Chen Jing

    Previous study suggest that powerful Large Language Models (LLMs) trained with Reinforcement Learning with Verifiable Rewards (RLVR) only refines reasoning path without improving the reasoning capacity in math tasks while supervised-finetuning(SFT) with distillation can. We study this from the view of Scientific information extraction (SciIE) where LLMs and

  26. Xinyu Xia, Xingjun Ma, Yunfeng Hu, Ting Qu

    Ensuring robust and generalizable autonomous driving requires not only broad scenario coverage but also efficient repair of failure cases, particularly those related to challenging and safety-critical scenarios. However, existing scenario generation and selection methods often lack adaptivity and semantic relevance, limiting their impact on performance impro

  27. Hanyu Cheng, Eleonora Di Valentino, Luca Visinelli

    Cosmic strings, topological defects predicted by high-energy theories, may contribute to the late-time expansion of the Universe, effectively mimicking dynamical dark energy. We investigate four phenomenological extensions of the $\Lambda$CDM model involving a residual string network: (i) a non-relativistic component with positive energy density (Model~1), (

  28. Mikko Impiö, Philipp M. Rehsen, Tiina Laamanen, Arne J. Beermann

    This paper presents the AquaMonitor dataset, the first large computer vision dataset of aquatic invertebrates collected during routine environmental monitoring. While several large species identification datasets exist, they are rarely collected using standardized collection protocols, and none focus on aquatic invertebrates, which are particularly laborious

  29. Zhicheng Feng, Gunter Malle, Jiping Zhang

    This paper is motivated by the study of Alperin's weight conjecture in the representation theory of finite groups. We generalize the notion of $e$-cuspidality in the $e$-Harish-Chandra theory of finite reductive groups, and define generic weights in non-defining characteristic. We show that the generic weights play an analogous role as the weights defined by

  30. Lei Yu, Yechao Zhang, Ziqi Zhou, Yang Wu

    With the rapid development of the Vision-Language Model (VLM), significant progress has been made in Visual Question Answering (VQA) tasks. However, existing VLM often generate inaccurate answers due to a lack of up-to-date knowledge. To address this issue, recent research has introduced Retrieval-Augmented Generation (RAG) techniques, commonly used in Large

  31. Mingchen Shao, Xinfa Zhu, Chengyou Wang, Bingshen Mu

    Despite remarkable achievements, automatic speech recognition (ASR) in low-resource scenarios still faces two challenges: high-quality data scarcity and high computational demands. This paper proposes EThai-ASR, the first to apply large language models (LLMs) to Thai ASR and create an efficient LLM-based ASR system. EThai-ASR comprises a speech encoder, a co

  32. Jiahao Hu, Ruiyang Zhang, Zhiliang Chen, Yunpeng Lu

    The Circular Electron-Positron Collider (CEPC), a proposed next-generation $e^+e^-$ collider to enable high-precision studies of the Higgs boson and potential new physics, imposes rigorous demands on detector technologies, particularly the vertex detector. JadePix-3 is a prototype Monolithic Active Pixel Sensor (MAPS) designed for the CEPC vertex detector. T

  33. Yujin Choi, Youngjoo Park, Junyoung Byun, Jaewook Lee

    Retrieval-augmented generation (RAG) mitigates the hallucination problem in large language models (LLMs) and has proven effective for personalized usages. However, delivering private retrieved documents directly to LLMs introduces vulnerability to membership inference attacks (MIAs), which try to determine whether the target data point exists in the private

  34. Lukas Schmidbauer, Wolfgang Mauerer

    In the foreseeable future, toolchains for quantum computing should offer automatic means of transforming a high level problem formulation down to a hardware executable form. Thereby, it is crucial to find (multiple) transformation paths that are optimised for (hardware specific) metrics. We zoom into this pictured tree of transformations by focussing on k-SA

  35. Emmanuel Kowalski, Théo Untrau

    The Wasserstein distance between probability measures on compact spaces provides a natural invariant quantitative measure of equidistribution, which is partly similar to the classical discrepancy appearing in Erd\"os-Tur\'an type inequalities in the case of tori, but is a more intrinsic quantity. We recall the basic properties of Wasserstein distances and pr

  36. Dezhi Song, Fuyang Hang, Gang Yao, Jun Zhang

    The intrinsic antiferromagnetic topological insulators in the Mn-Bi-Te family, composed of superlattice-like MnBi2Te4/(Bi2Te3)n (n = 0, 1, 2, 3...) layered structure, present intriguing states of matter such as quantum anomalous Hall effect and the axion insulator. However, the surface state gap, which is the prerequisite for the observation of these states,

  37. Yansen Zhang, Xiaokun Zhang, Ziqiang Cui, Chen Ma

    Recommender systems often suffer from noisy interactions like accidental clicks or popularity bias. Existing denoising methods typically identify users' intent in their interactions, and filter out noisy interactions that deviate from the assumed intent. However, they ignore that interactions deemed noisy could still aid model training, while some ``clean''

  38. Nayara Carral-Sainz, Toraya Fernández-Ruiz, Jorge Íñiguez, Javier Junquera

    We present a systematic, quasi-automated methodology for generating electronic models in the framework of second-principles density functional theory (SPDFT). This approach enables the construction of accurate and computationally efficient models by deriving all necessary parameters from first-principles calculations on a carefully designed training set. A k

  39. Jozefien D'haeseleer, Vladislav Taranchuk

    In this paper we study the chromatic number of the Grassmann graphs $J_q(n, m)$. We show that $\binom{n-m+1}{1}_q \leq \chi(J_q(n, m)) \leq \binom{n}{1}_q$, which is analogous to the best-known bounds for the chromatic number of the Johnson graphs $J(n, m)$. When $m = 2$, determining $\chi(J_q(n, 2))$ is equivalent to determining the smallest number of parti

  40. Samuel Stucki, Jan Deriu, Mark Cieliebak

    This work investigates the performance of Voice Adaptation models for Swiss German dialects, i.e., translating Standard German text to Swiss German dialect speech. For this, we preprocess a large dataset of Swiss podcasts, which we automatically transcribe and annotate with dialect classes, yielding approximately 5000 hours of weakly labeled training materia

  41. Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang

    Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images), owing to the scarcity of high-quality paired datasets and the lack of robust multi-task learning frameworks. Recently, mult

  42. Keno Hassler, Philipp Görz, Stephan Lipp

    Over 70% of security vulnerabilities in critical software systems today result from memory safety violations. To address this challenge, fuzzing and static analysis are widely used automated methods to discover such vulnerabilities. Fuzzing generates random program inputs to identify faults at runtime, while static analysis reasons about the code to detect p

  43. Pengjie Shen, Xueliang Zhang, Zhong-Qiu Wang

    We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regressive connections, where the estimated target speech at previous frames is leveraged as extra input features to help the DNN estimate the tar

  44. Di Wu, Jiaxin Fan, Junzhe Zang, Guanbo Wang

    Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle with the temporal reasoning, spatial understanding, and commonsense grounding needed for planning in interactive environments. In th

  45. Laetitia Chapel, Romain Tavenard, Samuel Vaiter

    Optimal Transport (OT) has attracted significant interest in the machine learning community, not only for its ability to define meaningful distances between probability distributions -- such as the Wasserstein distance -- but also for its formulation of OT plans. Its computational complexity remains a bottleneck, though, and slicing techniques have been deve

  46. Haihan Zhang, Weicheng Lin, Yuanshi Liu, Cong Fang

    This paper considers a canonical problem in kernel regression: how good are the model performances when it is trained by the popular online first-order algorithms, compared to the offline ones, such as ridge and ridgeless regression? In this paper, we analyze the foundational single-pass Stochastic Gradient Descent (SGD) in kernel regression under source con

  47. Claus Metzner, Achim Schilling, Andreas Maier, Patrick Krauss

    Previous work has shown that the dynamical regime of Recurrent Neural Networks (RNNs) - ranging from oscillatory to chaotic and fixpoint behavior - can be controlled by the global distribution of weights in connection matrices with statistically independent elements. However, it remains unclear how network dynamics respond to organizational regularities in t

  48. Haipeng Zhou, Sicheng Yang, Sihan Yang, Jing Qin

    Survival prediction aims to evaluate the risk level of cancer patients. Existing methods primarily rely on pathology and genomics data, either individually or in combination. From the perspective of cancer pathogenesis, epigenetic changes, such as methylation data, could also be crucial for this task. Furthermore, no previous endeavors have utilized textual

  49. Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun, Farid Boussaid

    Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human movements, leading to unnatural deformations. To tackle this issue, we present LatentMove, a DiT-based framework specificall

  50. Le Xu, Chenxing Li, Yong Ren, Yujie Chen

    Current vision-guided audio captioning systems frequently fail to address audiovisual misalignment in real-world scenarios, such as dubbed content or off-screen sounds. To bridge this critical gap, we present an entropy-aware gated fusion framework that dynamically modulates visual information flow through cross-modal uncertainty quantification. Our novel ap

  51. L. L. Kurchaninov, E. A. Ladygin, V. P. Ladygin, A. A. Semak

    The analog front-end electronics based on the constant fraction discrimination method is designed and optimized for the Multigap Resistive Plate Chamber (MRPC) timing measurements. The total time resolution of 40 ps has been obtained for 10 and 12 gaps MRPCs using cosmic setup and a muon beam at the IHEP U-70 accelerator in Protvino, which complies with the

  52. Timofei Snegirev

    Superconformal extensions of the perfect fluid equations, which realize $N=1,2$ Schrodinger superalgebra, are constructed within the Hamiltonian formalism. They are built by introducing real (for $N=1$) or complex (for $N=2$) anticommuting field variables as superpartners for the density and velocity of a fluid. The full set of conserved charges associated w

  53. Hao Yang, Haoxuan Li, Mengyue Yang, Xu Chen

    The order of training samples plays a crucial role in large language models (LLMs), significantly impacting both their external performance and internal learning dynamics. Traditional methods for investigating this effect generally require retraining the model with various sample orders, which is computationally infeasible for LLMs. In this work, we improve

  54. Michael Grohs, Adrian Rebmann, Jana-Rebecca Rehse

    Conformance checking techniques detect undesired process behavior by comparing process executions that are recorded in event logs to desired behavior that is captured in a dedicated process model. If such models are not available, conformance checking techniques are not applicable, but organizations might still be interested in detecting undesired behavior i

  55. Nachuan Xiao, Xiaoyin Hu, Xin Liu, Kim-Chuan Toh

    In this paper, we focus on the nonconvex-nonconvex bilevel optimization problem (BLO), where both upper-level and lower-level objectives are nonconvex, with the upper-level problem potentially being nonsmooth. We develop a two-timescale momentum-accelerated subgradient method (TMG) that employs two-timescale stepsizes, and establish its local convergence whe

  56. Kaiyuan Li, Xiaoyue Chen, Chen Gao, Yong Li

    Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image tokens results in significant computational overhead, and the use of dynamic high-resolution inputs further increases this burden. Previous approaches have attempted to reduce the numb

  57. Jingyu Zhang, Ahmed Elgohary, Xiawei Wang, A S M Iftekhar

    Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a novel benchmark construction framework that "distills" jailbreak attacks into high-quality and easily-updatable safety benchmarks. JBDistill utilizes a small set of development model

  58. Tetsushi Ito, Daichi Takeuchi, Takahiro Tsushima

    The van der Geer--van der Vlugt curves form a class of Artin--Schreier coverings of the projective line over finite fields. We provide an explicit formula for their $L$-polynomials in characteristic $2$, expressed in terms of characters of maximal abelian subgroups of associated Heisenberg groups. For this purpose, we develop new methods specific to characte

  59. Nuolin Sun, Linyuan Wang, Dongyang Li, Bin Yan

    Adversarial attacks have received increasing attention and it has been widely recognized that classical DNNs have weak adversarial robustness. The most commonly used adversarial defense method, adversarial training, improves the adversarial accuracy of DNNs by generating adversarial examples and retraining the model. However, adversarial training requires a

  60. Xinyi Chen, Chenxiang Ma, Yujie Wu, Kay Chen Tan

    Temporal processing is vital for extracting meaningful information from time-varying signals. Recent advancements in Spiking Neural Networks (SNNs) have shown immense promise in efficiently processing these signals. However, progress in this field has been impeded by the lack of effective and standardized benchmarks, which complicates the consistent measurem

  61. Oskar Høgberg Simensen, Dennis Christensen, Nils Lid Hjort

    We propose a new method of histogram construction, providing a fully Bayesian approach to irregular histograms. Our procedure applies Bayesian model selection to a piecewise constant model of the underlying distribution, resulting in a method that selects both the number of bins as well as their location based on the data in a fully automatic fashion. We sho

  62. Zeming Zhuang, Kun Meng, Hongsheng Zhang

    We study the phase transition and critical phenomenon of charged black holes in Einstein-Maxwell-scalar (EMs) theory. Through comprehensive analysis of thermodynamic behaviors manifested in P-V diagrams, G(T,P) surfaces, and C_P curves, we establish that these black holes exhibit van der Waals-type phase transition behavior. The derived critical exponents go

  63. Tawfiq Ammari, Anna Gutowska, Jacob Ziff, Casey Randazzo

    As the COVID-19 pandemic evolved, the Centers for Disease Control and Prevention (CDC) used Twitter to disseminate safety guidance and updates, reaching millions of users. This study analyzes two years of tweets from, to, and about the CDC using a mixed methods approach to examine discourse characteristics, credibility, and user engagement. We found that the

  64. Hasan Yucedag, Adam Jatowt

    This paper introduces Guess the Age of Photos, a web platform engaging users in estimating the years of historical photographs through two gamified modes: Guess the Year (predicting a single image's year) and Timeline Challenge (comparing two images to identify the older). Built with Python, Flask, Bootstrap, and PostgreSQL, it uses a 10,150-image subset of

  65. Marco Limongi, Lorenzo Roberti, Agnese Falla, Alessandro Chieffi

    In Limongi et al. (2024) we presented and discussed the main evolutionary properties and final fate of stars in the mass range 7-15 Msun. The evolutions of those models were computed by means of a medium size nuclear network that guaranteed a proper calculation of the nuclear energy generation and hence a good modeling of the physical evolution of these star

  66. Jinming Zhang, Xuanru Zhou, Jiachen Lian, Shuhe Li

    Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these

  67. Zi-Hao Zhou, Jun-Jie Wang, Tong Wei, Min-Ling Zhang

    Contrastive learning has achieved remarkable success in learning effective representations, with supervised contrastive learning often outperforming self-supervised approaches. However, in real-world scenarios, data annotations are often ambiguous or inaccurate, meaning that class labels may not reliably indicate whether two examples belong to the same class

  68. Miika Toikkanen, June-Woo Kim

    Respiratory sound datasets are limited in size and quality, making high performance difficult to achieve. Ensemble models help but inevitably increase compute cost at inference time. Soft label training distills knowledge efficiently with extra cost only at training. In this study, we explore soft labels for respiratory sound classification as an architectur

  69. Zheng Wei

    We establish a form of 2-adjunction (tentatively termed the *fundamental 2-adjunction*), building on the fundamental adjunction proposed by Olivia Caramello and Riccardo Zanfa, which provides a constructive method for the associated stack functor. Additionally, we investigate 2-local homeomorphisms through the lens of indexed fibrations.

  70. M. Gorgone, F. Oliveri, A. Ricciardello, P. Rogolino

    In this paper, after reviewing the form of the constitutive equations for a third grade Korteweg fluid, recently derived by means of an extended Liu procedure, an equilibrium problem is investigated. By considering a two--dimensional setting, it is derived a single nonlinear elliptic equation such that the equilibrium conditions are identically satisfied. Su

  71. Manchao Bao, Shengjiang Fang, Tao Yue, Xuemei Hu

    Long-distance depth imaging holds great promise for applications such as autonomous driving and robotics. Direct time-of-flight (dToF) imaging offers high-precision, long-distance depth sensing, yet demands ultra-short pulse light sources and high-resolution time-to-digital converters. In contrast, indirect time-of-flight (iToF) imaging often suffers from ph

  72. Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham, Van Nguyen

    Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a novel L2S system that generates intelligible and expressive speech from silent talking face videos. Leveraging source-fil

  73. Bangde Du, Ziyi Ye, Zhijing Wu, Jankowska Monika

    As Large Language Models (LLMs) continue to exhibit increasingly human-like capabilities, aligning them with human values has become critically important. Contemporary advanced techniques, such as prompt learning and reinforcement learning, are being deployed to better align LLMs with human values. However, while these approaches address broad ethical consid

  74. Ritwik Murali, Akash Ravi

    Software systems have grown as an indispensable commodity used across various industries, and almost all essential services depend on them for effective operation. The software is no longer an independent or stand-alone piece of code written by a developer but rather a collection of packages designed by multiple developers across the globe. Ensuring the reli

  75. Farjana Siddiqua, Catalin Trenchea

    We analyze an advection-diffusion-reaction problem with non-homogeneous boundary conditions that models the chromatography process.We prove stability and error estimates for both constant and affine adsorption, using the symplectic one-step implicit midpoint method for time discretization and finite elements for spatial discretization. In addition, we perfor

  76. Yuanjian Xu, Jianing Hao, Kunsheng Tang, Jingnan Chen

    Financial markets exhibit complex dynamics where localized events trigger ripple effects across entities. Previous event studies, constrained by static single-company analyses and simplistic assumptions, fail to capture these ripple effects. While large language models (LLMs) offer emergent reasoning capabilities, their direct application falters due to stru

  77. Zhihong Tang

    Document Image Enhancement (DIE) serves as a critical component in Document AI systems, where its performance substantially determines the effectiveness of downstream tasks. To address the limitations of existing methods confined to single-degradation restoration or grayscale image processing, we present Global with Local Parametric Generation Enhancement Ne

  78. Jing Du, Haley Stone, Yang Yang, Ashna Desai

    Accurate forecasting of Avian Influenza Virus (AIV) outbreaks within wild bird populations necessitates models that account for complex, multi-scale transmission patterns driven by diverse factors. While conventional spatiotemporal epidemic models are robust for human-centric diseases, they rely on spatial homophily and diffusive transmission between geograp

  79. Won Joon Sohn, Jeffrey Lim, Po T. Wang, Susan J. Shaw

    Bi-directional brain computer interfaces (BD-BCIs) may restore brain-controlled walking and artificial leg sensation after spinal cord injury. Current BD-BCIs provide only simplistic "tingling" feedback, which lacks proprioceptive information to perceive critical gait events (leg swing, double support). This information must also be perceived adequately fast

  80. Yidian Wu, Rui Liu, Runbin Luo, Wensi Wang

    Mass drainage is frequently observed in solar filaments. During filament eruptions, falling material most likely flows along magnetic field lines, which may provide important clues for the magnetic structures of filaments. Here we study three filament eruptions exhibiting significant mass draining, often manifested as falling threads at a constant speed rang

  81. Qiuchen Wang, Ruixue Ding, Yu Zeng, Zehui Chen

    Effectively retrieving, reasoning and understanding visually rich information remains a challenge for RAG methods. Traditional text-based methods cannot handle visual-related information. On the other hand, current vision-based RAG approaches are often limited by fixed pipelines and frequently struggle to reason effectively due to the insufficient activation

  82. Ruicheng Yin, Xuan Gao, Changze Lv, Xiaohua Wang

    Continual pre-training has demonstrated significant potential in enhancing model performance, particularly in domain-specific scenarios. The most common approach for packing data before continual pre-training involves concatenating input texts and splitting them into fixed-length sequences. While straightforward and efficient, this method often leads to exce

  83. Siqi Fan, Bowen Qin, Peng Han, Shuo Shang

    Recent thinking models trained with reinforcement learning and backward-checking CoT often suffer from overthinking: they produce excessively long outputs even on simple problems, wasting computation. Existing evaluations, based on token efficiency, give an incomplete view as they neglect problem difficulty and intermediate computation costs. We formalize re

  84. Carl Corea, Timotheus Kampik, Nico Potyka

    We investigate a new form of (privacy-preserving) inconsistency measurement for multi-party communication. Intuitively, for two knowledge bases K_A, K_B (of two agents A, B), our results allow to quantitatively assess the degree of inconsistency for K_A U K_B without having to reveal the actual contents of the knowledge bases. Using secure multi-party comput

  85. Yifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu

    Panoramic video generation enables immersive 360{\deg} content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited

  86. Weiming Li, Zeng Li, Siyu Wang, Yanqing Yin

    We study distributed principal component analysis (PCA) in high-dimensional settings under the spiked model. In such regimes, sample eigenvectors can deviate significantly from population ones, introducing a persistent bias. Existing distributed PCA methods are sensitive to this bias, particularly when the number of machines is small. Their consistency typic

  87. Jörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev

    The successful training of deep neural networks requires addressing challenges such as overfitting, numerical instabilities leading to divergence, and increasing variance in the residual stream. A common solution is to apply regularization and normalization techniques that usually require tuning additional hyperparameters. An alternative is to force all para

  88. Shangkun Huang, Yuxuan Du, Jingwen Yang, Dejun Zhang

    This paper presents the system developed to address the MISP 2025 Challenge. For the diarization system, we proposed a hybrid approach combining a WavLM end-to-end segmentation method with a traditional multi-module clustering technique to adaptively select the appropriate model for handling varying degrees of overlapping speech. For the automatic speech rec

  89. Martin Huang, Samuel Muller, Garth Tarr

    Stability selection has gained popularity as a method for enhancing the performance of variable selection algorithms while controlling false discovery rates. However, achieving these desirable properties depends on correctly specifying the stable threshold parameter, which can be challenging. An arbitrary choice of this parameter can substantially alter the

  90. Menghui Zhang, Jing Zhang, Lin Chen, Li Zhuo

    Livestreaming often involves interactions between streamers and objects, which is critical for understanding and regulating web content. While human-object interaction (HOI) detection has made some progress in general-purpose video downstream tasks, when applied to recognize the interaction behaviors between a streamer and different objects in livestreaming,

  91. Nasir Hussain, Haohan Chen, Chanh Tran, Philip Huang

    Recognizing vulnerabilities in stripped binary files presents a significant challenge in software security. Although some progress has been made in generating human-readable information from decompiled binary files with Large Language Models (LLMs), effectively and scalably detecting vulnerabilities within these binary files is still an open problem. This pa

  92. Akihiko Fukui

    TOI-2285 b is a sub-Neptune-sized planet orbiting a nearby M dwarf, discovered through the TESS photometric survey and ground-based follow-up observations. The planet was initially reported to have an orbital period of 27.27 d, making it one of the lowest temperature sub-Neptunes transiting a bright M dwarf. However, additional TESS data reveal that its true

  93. Jing-An Sun, Hang Fan, Junchao Gong, Ben Fei

    Data assimilation (DA) aims to estimate the full state of a dynamical system by combining partial and noisy observations with a prior model forecast, commonly referred to as the background. In atmospheric applications, this problem is fundamentally ill-posed due to the sparsity of observations relative to the high-dimensional state space. Traditional methods

  94. Tianmai M. Zhang, Neil F. Abernethy

    Recent advancements in large language models have sparked interest in utilizing them to aid the peer review process of scientific publication amid the peer review crisis. However, having AI models generate full reviews in the same way as human reviewers risks exacerbating the irresponsible use of LLM-generated reviews and instigating intentional manipulation

  95. Wataru Ikeda, Masashi Hatano, Ryosei Hara, Mariko Isogawa

    Estimating human pose using a front-facing egocentric camera is essential for applications such as sports motion analysis, VR/AR, and AI for wearable devices. However, many existing methods rely on RGB cameras and do not account for low-light environments or motion blur. Event-based cameras have the potential to address these challenges. In this work, we int

  96. Changze Qiao, Mingming Lu

    With large language models (LLMs) demonstrating remarkable capabilities, there has been a surge in research on leveraging LLMs to build general-purpose multi-modal agents. However, existing approaches either rely on computationally expensive end-to-end training using large-scale multi-modal data or adopt tool-use methods that lack the ability to continuously

  97. Shangkun Huang, Jing Deng, Jintao Kang, Rong Zheng

    The performance bottleneck of Automatic Speech Recognition (ASR) in stuttering speech scenarios has limited its applicability in domains such as speech rehabilitation. This paper proposed an LLM-driven ASR-SED multi-task learning framework that jointly optimized the ASR and Stuttering Event Detection (SED) tasks. We proposed a dynamic interaction mechanism w

  98. Eric Hoffbeck, Johan Leray, Bruno Vallette

    In this paper, we settle the homotopy properties of the infinity-morphisms of homotopy (bial)-gebras over properads, i.e. algebraic structures made up of operations with several inputs and outputs. We start by providing the literature with characterizations for the various types of infinity-morphisms, the most seminal one being the equivalence between infini

  99. Simone Bendazzoli, Sanna Persson, Mehdi Astaraki, Sebastian Pettersson

    The integration of Artificial Intelligence (AI) into clinical workflows requires robust collaborative platforms that are able to bridge the gap between technical innovation and practical healthcare applications. This paper introduces MAIA (Medical Artificial Intelligence Assistant), an open-source platform designed to facilitate interdisciplinary collaborati

  100. Jatin Gupta, Akhil Sharma, Saransh Singhania, Ali Imam Abidi

    In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full advantage of their legal rights due to limited access and awareness of apposite legal information. This paper thus introduces Legal Assist AI, a highly efficient framework designed to provide legal assistance in the