Skip to content

March 2025 arXiv papers — page 179

Showing 17,80117,900 of 23,633 papers

  1. Gaurav Patel, Qiang Qiu

    Machine Unlearning has recently garnered significant attention, aiming to selectively remove knowledge associated with specific data while preserving the model's performance on the remaining data. A fundamental challenge in this process is balancing effective unlearning with knowledge retention, as naive optimization of these competing objectives can lead to

  2. Amin Akhavan

    We investigate the role of external constraints in quantum field theory using the path integral formalism. We begin by reviewing the quantization of constrained systems and extend the analysis to cases where constraints are added to the action via auxiliary fields. These constraints involve both the degrees of freedom and their time derivatives. Using the re

  3. Mohit Pandey, Gopeshh Subbaraj, Artem Cherkasov, Martin Ester

    Generative Flow Networks (GFlowNets) have recently emerged as a suitable framework for generating diverse and high-quality molecular structures by learning from rewards treated as unnormalized distributions. Previous works in this framework often restrict exploration by using predefined molecular fragments as building blocks, limiting the chemical space that

  4. Jnana Ranjan Das, Santanu Sinha, Alex Hansen, Sitangshu B. Santra

    We present a percolation model that is inspired by recent works on immiscible two-phase flow in a mixed-wet porous medium made of a mixture of grains with two different wettability properties. The percolation model is constructed on a dual lattice where the sites on the primal lattice represent the grains of the porous medium, and the bonds on the dual latti

  5. Alex Calderwood, John Joon Young Chung, Yuqian Sun, Melissa Roemmele

    According to the recently introduced theory of artistic support tools, creativity support tools exert normative influences over artistic production, instantiating a normative ground that shapes both the process and product of artistic expression. We argue that the normative ground of most existing automated writing tools is misaligned with writerly values an

  6. Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen

    Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities. Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL). The exploration of MCL is in its early

  7. Rashik Shrestha, Madhav Rijal, Trevor Smith, Yu Gu

    This study presents Flower Pose Estimation (FloPE), a real-time flower pose estimation framework for computationally constrained robotic pollination systems. Robotic pollination has been proposed to supplement natural pollination to ensure global food security due to the decreased population of natural pollinators. However, flower pose estimation for pollina

  8. Kenneth Stephenson

    There exists an extensive and fairly comprehensive discrete analytic function theory which is based on circle packing. This paper introduces a faithful discrete analogue of the classical Schwarzian derivative to this theory and develops its basic properties. The motivation comes from the current lack of circle packing algorithms in spherical geometry, and th

  9. Panagiotis Kourtesis, Andrea Lizarraga, Sarah E. MacPherson

    Objective: Immersive virtual reality (VR) enhances ecologically validity and facilitates intuitive and ergonomic hand interactions for performing neuropsychological assessments. However, its comparability to traditional computerized methods remains unclear. This study investigates the convergent validity, user experience, and usability of VR-based versus PC-

  10. Hristo N. Djidjev

    Node embedding is a key technique for representing graph nodes as vectors while preserving structural and relational properties, which enables machine learning tasks like feature extraction, clustering, and classification. While classical methods such as DeepWalk, node2vec, and graph convolutional networks learn node embeddings by capturing structural and re

  11. Nikolay Mikhaylovskiy

    It is known for some time that autocorrelations of words in human-written texts decay according to a power law. Recent works have also shown that the autocorrelations decay in texts generated by LLMs is qualitatively different from the literary texts. Solid state physics tie the autocorrelations decay laws to the states of matter. In this work, we empiricall

  12. M. H. Shahzamanian

    In this paper, we introduce and study a class of monoids, called Layered Catalan Monoids (\( {LC}_n \)), which satisfy the structural conditions for $\ll$-smoothness as defined in~\cite{Sha-Det2}. These monoids are defined by specific identities inspired by Catalan monoids. We establish their canonical forms and compute their determinant, proving that it is

  13. Shlok Nahar, Devashish Tupkary, Norbert Lütkenhaus

    Security analyses in quantum key distribution (QKD) and other adversarial quantum tasks often assume perfect device models. However, real-world implementations often deviate from these models. Thus, it is important to develop security proofs that account for such deviations from ideality. In this work, we extend the idea of squashing maps to develop a genera

  14. Altaf Allah Abbassi, Leuson Da Silva, Amin Nikanjam, Foutse Khomh

    Large Language Models (LLMs) are widely adopted for automated code generation with promising results. Although prior research has assessed LLM-generated code and identified various quality issues -- such as redundancy, poor maintainability, and sub-optimal performance a systematic understanding and categorization of these inefficiencies remain unexplored. Wi

  15. Evgeny Mukhin, Alexander Varchenko

    In [J. Lond. Math. Soc. 109 (2024), e12884, 22 pages, arXiv:2208.09721], the difference qKZ equations were considered modulo a prime number $p$ and a family of polynomial solutions of the qKZ equations modulo $p$ was constructed by an elementary procedure as suitable $p$-approximations of the hypergeometric integrals. In this paper, we study in detail the fi

  16. Vijayamanikandan Vijayarangan, Harshavardhana A. Uranakara, Francisco E. Hernández-Pérez, Hong G. Im

    Using the information theory, this study provides insights into how the construction of latent space of autoencoder (AE) using deep neural network (DNN) training finds a smooth low-dimensional manifold in the stiff dynamical system. Our recent study [1] reported that an autoencoder (AE) combined with neural ODE (NODE) as a surrogate reduced order model (ROM)

  17. Kazuya Izumi, Shuhey Koyama, Yoichi Ochiai

    Avatars on displays lack the ability to engage with the physical environment through gaze. To address this limitation, we propose a gaze synthesis method that enables animated avatars to establish gaze communication with the physical environment using a camera-behind-the-display system. The system uses a display that rapidly alternates between visible and tr

  18. RB Yadav, Arpan Sharma

    In this article, we give a characterisation of crossed homomorphisms on Lie superalgebras as a Maurer-Cartan element of a graded Lie algebra. Using this characterisation we study cohomology of these crossed homomorphisms. As an application of this cohomology we study formal deformation of crossed homomorphisms. We show that linear deformations of these homom

  19. Jack Foxabbott, Rohan Subramani, Francis Rhys Ward

    Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs) but have a more compact and informative structure. However, MAIDs cannot, in general, represent settings of incomplete information -- wherein agents have different beliefs about t

  20. Jieyang Chen, Qian Gong, Yanliang Li, Xin Liang

    The rapid growth of scientific data is surpassing advancements in computing, creating challenges in storage, transfer, and analysis, particularly at the exascale. While data reduction techniques such as lossless and lossy compression help mitigate these issues, their computational overhead introduces new bottlenecks. GPU-accelerated approaches improve perfor

  21. Md Ohiduzzaman Ovi, Maliha Sanjana, Fahad Fahad, Mahjabin Runa

    Pediatric dental segmentation is critical in dental diagnostics, presenting unique challenges due to variations in dental structures and the lower number of pediatric X-ray images. This study proposes a custom SegUNet model with a VGG19 backbone, designed explicitly for pediatric dental segmentation and applied to the Children's Dental Panoramic Radiographs

  22. Zongren Zou, Zhicheng Wang, George Em Karniadakis

    We explore the capability of physics-informed neural networks (PINNs) to discover multiple solutions. Many real-world phenomena governed by nonlinear differential equations (DEs), such as fluid flow, exhibit multiple solutions under the same conditions, yet capturing this solution multiplicity remains a significant challenge. A key difficulty is giving appro

  23. Ritika Nagpal, S. K. J. Pacif, Farruh Atamurotov, Rasmikanta Pati

    In this study, we explore the impact of the interacting parameter on dark matter in a model resulting from a parametrization of dark energy density. To ensure a model-independent approach, we treat \( r_d \) as a free parameter, avoiding assumptions about the physics of the early Universe or specific recombination models. This approach allows late-time cosmo

  24. William Webb, Arturas Medeisis, Leo Fulvio Minervini

    This article discusses the key principles of radio spectrum management with a focus on spectrum allocation and access. We show the current regime's inherent rigidity and constrained possibilities for introducing new radiocommunication services and applications. The article proposes how governments and spectrum users could cooperate in taking spectrum managem

  25. Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu

    Object detection in videos plays a crucial role in advancing applications such as public safety and anomaly detection. Existing methods have explored different techniques, including CNN, deep learning, and Transformers, for object detection and video classification. However, detecting tiny objects, e.g., guns, in videos remains challenging due to their small

  26. Tieqiao Wang, Sinisa Todorovic

    Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- e

  27. Luiz A. Ferreira, Aliaksei Mikhaliuk, Yakov Shnir

    We present and study new non-topological soliton solutions in the $U(1)$ gauged non-linear $O(3)$ sigma model with a symmetry breaking potential in 3+1 dimensional flat space-time. The configurations are endowed with an electric and magnetic field and also carry a nonvanishing angular momentum density. We discuss properties of these solitons and investigate

  28. Ming-Hua Chang, Steffen Backes, Donghui Lu, Nicolas Gauthier

    Understanding how renormalized quasiparticles emerge in strongly correlated electron materials provides a challenge for both experiment and theory. It has been predicted that distinctive spin and orbital screening mechanisms drive this process in multiorbital materials with strong Coulomb and Hund's interactions. Here, we provide the experimental evidence of

  29. Chandan Kumar Sah, Ankit Kumar Shaw, Xiaoli Lian, Arsalan Shahid Baig

    Autonomous vehicles (AVs) require reliable traffic sign recognition and robust lane detection capabilities to ensure safe navigation in complex and dynamic environments. This paper introduces an integrated approach combining advanced deep learning techniques and Multimodal Large Language Models (MLLMs) for comprehensive road perception. For traffic sign reco

  30. Zhitong Xiong, Yi Wang, Weikang Yu, Adam J Stewart

    Earth observation (EO) spans a broad spectrum of modalities, including optical, radar, multispectral, and hyperspectral data, each capturing distinct environmental signals. However, current vision-language models in EO, particularly CLIP-based variants, remain confined to individual modalities, limiting generalization and scalability across diverse tasks. We

  31. Yuxuan Li, Sheng Jinag, Bizhu Wang

    With technology advancing and the pursuit of new audiovisual experiences strengthening, the metaverse has gained surging enthusiasm. However, it faces practical hurdles as substantial data like high-resolution virtual scenes must be transmitted between cloud platforms and VR devices. Specifically, the VR device's wireless transmission hampered by insufficien

  32. Sizhen Bian, Vitor Fortes Rey, Siyu Yuan, Paul Lukowicz

    While human body capacitance ($HBC$) has been explored as a novel wearable motion sensing modality, its competence has never been quantitatively demonstrated compared to that of the dominant inertial measurement unit ($IMU$) in practical scenarios. This work is thus motivated to evaluate the contribution of $HBC$ in wearable motion sensing. A real-life case

  33. Marco Iannotta, Johannes A. Stork, Erik Schaffernicht, Todor Stoyanov

    With the rising demand for flexible manufacturing, robots are increasingly expected to operate in dynamic environments where local -- such as slight offsets or size differences in workpieces -- are common. We propose to address the problem of adapting robot behaviors to these task variations with a sample-efficient hierarchical reinforcement learning approac

  34. Georg Hahn, Sebastian Schneeweiss, Shirley Wang

    Computable phenotypes are used to characterize patients and identify outcomes in studies conducted using healthcare claims and electronic health record data. Chart review studies establish reference labels against which computable phenotypes are compared to understand their measurement characteristics, the quantity of interest, for instance the positive pred

  35. Qizhen Lan, Qing Tian

    Dense visual prediction tasks, such as detection and segmentation, are crucial for time-critical applications (e.g., autonomous driving and video surveillance). While deep models achieve strong performance, their efficiency remains a challenge. Knowledge distillation (KD) is an effective model compression technique, but existing feature-based KD methods rely

  36. Arnaldo J. Vargas

    This work presents a model for testing Lorentz and CPT symmetry using rovibrational transitions within the electronic ground state of the molecular hydrogen ion (H$^+_2$). The model is based on the Standard-Model Extension (SME) and incorporates minimal and nonminimal effects. Our analysis concludes that sidereal variation studies of these transitions could

  37. Anna N. Morozovska, Eugene. A. Eliseev, Oleksiy V. Bereznikov, Mykola Ye. Yelisieiev

    The contribution of flexoelectric coupling to the long-range order parameter fluctuations in ferroics can be critically important to the ferron dispersion and related polar, pyroelectric and electrocaloric properties. Here we calculate analytically the dispersion relations of soft optic and acoustic flexocoupling-induced phonons and ferrons by incorporating

  38. Faaiq Waqar, Jungyoun Kwak, Junmo Lee, Minji Shon

    The Last Level Cache (LLC) is the processor's critical bridge between on-chip and off-chip memory levels - optimized for high density, high bandwidth, and low operation energy. To date, high-density (HD) SRAM has been the conventional device of choice; however, with the slowing of transistor scaling, as reflected in the industry's almost identical HD SRAM ce

  39. Georg Hahn, Moulinath Banerjee, Bodhisattva Sen

    The estimation of regression parameters in one dimensional broken stick models is a research area of statistics with an extensive literature. We are interested in extending such models by aiming to recover two or more intersecting (hyper)planes in multiple dimensions. In contrast to approaches aiming to recover a given number of piecewise linear components u

  40. Zifan Zhang, Minghong Fang, Dianwei Chen, Xianfeng Yang

    Digital network twins (DNTs) are virtual representations of physical networks, designed to enable real-time monitoring, simulation, and optimization of network performance. When integrated with machine learning (ML) techniques, particularly federated learning (FL) and reinforcement learning (RL), DNTs emerge as powerful solutions for managing the complexitie

  41. C. O. Edet, K. Słowik, N. Ali, M. Asjad

    Controlling heat flow at the quantum level is a key challenge for next-generation quantum technologies, including thermal management and quantum information processing. Here, we investigate quantum heat transport in an asymmetrically driven hybrid magnon-photon system in contact with two thermal baths at different temperatures. We demonstrate that external d

  42. Jeongmin Lee, Sunkyung Park, Minji Lee, Dongjun Lee

    This paper presents a framework designed to tackle a range of planning problems arise in manipulation, which typically involve complex geometric-physical reasoning related to contact and dynamic constraints. We introduce the Contact Factor Graph (CFG) to graphically model these diverse factors, enabling us to perform inference on the graphs to approximate th

  43. Yueh-Chun Wu, Matthew DeCapua, ZhongChen Xu, Takashi Taniguchi

    Moir\'e superlattices created by stacking atomic layers of transition metal dichalcogenide semiconductors have emerged as a class of fascinating artificial photonic and electronic materials. An appealing attribute of these structures is the inheritance of the valley degree of freedom from the constituent monolayers. Recent studies show evidence that the vall

  44. Tuoc Phan, Dario A. Valdebenito

    We study an inviscid limit problem for a class of Navier-Stokes equations with vanishing measurable viscous coefficients in 3-dimensional spatial domains whose boundaries are oscillatory, depending on a small parameter, and become flat when the parameter converges to zero. Under some sufficient conditions on the anisotropic vanishing rates of the eigenvalues

  45. Clara Sanchez-Perez, Paula Rivas-Lazaro, Elisa García-Tabarés, Iván García Vara

    A comprehensive evaluation of the effect and limitations of variable current density and electrolyte composition on layer porosity and microstructure changes of porous silicon (pSi) multilayer stacks is reported. Following these results, the development and optimization of a four-layer stack architecture is reported through addition of super-low porosity lay

  46. Vinay Kumar Verma, Shreyas Sunil Kulkarni, Happy Mittal, Deepak Gupta

    Question Answering (QA) and Visual Question Answering (VQA) are well-studied problems in the language and vision domain. One challenging scenario involves multiple sources of information, each of a different modality, where the answer to the question may exist in one or more sources. This scenario contains richer information but is highly complex to handle.

  47. Jobir Adashev, Xursanoy Berdalova, Feruza Toshtemirova

    In this paper we investigate classifications of all (transposed) Poisson algebras of the associated associative null-filiform algebra

  48. Dibakar Roychowdhury

    We explore various field theory aspects of integrable $ \eta $-deformed geometry in type IIB supergravity by employing several holographic probes. These include the computation of holographic timelike entanglement entropy and estimation of various other field theory observables for example, the flow central charge and the quantum complexity. We also discuss

  49. Jirui Guo, Mauricio Romo, Lucy Smith

    We study the properties of B-branes in a class of nonabelian GLSMs realizing the canonical line bundle $K_{Gr(2,N)}$ in their geometric phase. By analysing the hemisphere partition function, i.e. B-brane central charge, we propose a grade restriction rule and the corresponding window categories for a specific class of paths between phases. We find very strik

  50. Qitan Lv, Tianyu Liu, Hong Wang

    Large language models (LLMs) have been widely adopted in mathematical optimization in scientific scenarios for their extensive knowledge and advanced reasoning capabilities. Existing methods mainly focus on utilizing LLMs to solve optimization problems in a prompt-based manner, which takes observational feedback as additional textual descriptions. However, d

  51. Aidan Pellow-Jarman, Shane McFarthing, Doo Hyung Kang, Pilsun Yoo

    A novel hybrid quantum-classical approach has been developed to efficiently address the multireference quantum chemistry problem. The Handover Iterative Variational Quantum Eigensolver (HiVQE) is designed to accurately estimate ground-state wavefunctions by leveraging both quantum and classical computing resources. In this framework, noisy intermediate-scale

  52. Haryo Akbarianto Wibowo, Haiyue Song, Hideki Tanaka, Masao Utiyama

    Large Language Models (LLMs) have grown increasingly expensive to deploy, driving the need for effective model compression techniques. While block pruning offers a straightforward approach to reducing model size, existing methods often struggle to maintain performance or require substantial computational resources for recovery. We present IteRABRe, a simple

  53. Bojan Lukić, Thorben Knust, Andreas Rausch

    Due to the increasing complexity and interconnectedness of different components in modern automotive software systems there is a great number of interactions between these system components and their environment. These interactions result in unique temporal behaviors we call underlying scenarios. The signal data from all system components, which is recorded

  54. Mario I. Molina

    We study a nonlinear magnetic metamaterial modeled as a split-ring resonator array, where the standard discrete laplacian is replaced by its fractional form. We find a closed-form expression for the dispersion relation as a function of the fractional exponent s and the gain/loss parameter {\gamma} and examine the conditions under which stable magneto-inducti

  55. Xiangyu Yin, Jiaxu Liu, Zhen Chen, Jinwei Hu

    Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailb

  56. Hao Yan, Marzi Heidari, Yuhong Guo

    Domain Generalization (DG) aims to train models that can generalize to unseen testing domains by leveraging data from multiple training domains. However, traditional DG methods rely on the availability of multiple diverse training domains, limiting their applicability in data-constrained scenarios. Single Domain Generalization (SDG) addresses the more realis

  57. Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang

    Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual grounding, though they inevitably require fine-tuning and additional model components to explicitly generate bounding boxes or se

  58. Alessandro T. Gifford, Radoslaw M. Cichy, Thomas Naselaris, Kendrick Kay

    Now published in Nature Communications DOI: https://doi.org/10.1038/s41467-026-69345-9 Large-scale visual neural datasets such as the Natural Scenes Dataset (NSD) are boosting computational neuroscience research by enabling models of the brain with performances beyond what was possible just a decade ago. However, because the stimuli of these datasets typical

  59. Gourav Kumar, V. Vetrivel

    In this article, we provide a modification to the Bregman Golden Ratio Algorithm (B-GRAAL). We analyze the B-GRAAL algorithm with a new step size rule, where the step size increases after a certain number of iterations and does not require prior knowledge of the global Lipschitz constant of the cost operator. Under suitable assumptions, we establish the glob

  60. Shabnam Ghasemirad, Si Liu, Christoph Sprenger, Luca Multazzu

    Isolation bugs, stemming especially from design-level defects, have been repeatedly found in carefully designed and extensively tested production databases over decades. In parallel, various frameworks for modeling database transactions and reasoning about their isolation guarantees have been developed. What is missing however is a mathematically rigorous an

  61. Deepti Jain, Hee Taek Yi, Xiong Yao, Alessandro R. Mazza

    The intrinsic magnetic topological insulator (IMTI) family $[\mathrm{MnTe}][\mathrm{Bi}_{2}\mathrm{Te}_{3}]_{\mathrm{n}}$ has demonstrated magneto-topological properties dependent on $n$, making it a promising platform for advanced electronics and spintronics. However, due to technical barriers in sample synthesis, their properties in the large $n$ limit rem

  62. Shuangzhi Li, Junlong Shen, Lei Ma, Xingyu Li

    LiDAR-based 3D object detection models often struggle to generalize to real-world environments due to limited object diversity in existing datasets. To tackle it, we introduce the first generalized cross-domain few-shot (GCFS) task in 3D object detection, aiming to adapt a source-pretrained model to both common and novel classes in a new domain with only few

  63. Ramanath Cowsik, Dawson Huth

    The leaky-box model and the attendant concept of path-length distribution of cosmic rays were invented in the mid-1960's. Even though versatile computational packages such as GALPROP and DRAGON with the diffusion approach are now available for analyzing cosmic ray data, the concepts of the leaky-box and path-length distribution continue to be adopted extensi

  64. Xiongfei Zhao, Hou-Wan Long, Zhengzhe Li, Jiangchuan Liu

    The rapid growth of Blockchain and Decentralized Finance (DeFi) has introduced new challenges and vulnerabilities that threaten the integrity and efficiency of the ecosystem. This study identifies critical issues such as Transaction Order Dependence (TOD), Blockchain Extractable Value (BEV), and Transaction Importance Diversity (TID), which collectively unde

  65. Bojan Lukić

    This paper gives an overview on how to develop a dense and deep neural network for making a time series prediction. First, the history and cornerstones in Artificial Intelligence and Machine Learning will be presented. After a short introduction to the theory of Artificial Intelligence and Machine Learning, the paper will go deeper into the techniques for co

  66. Siyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin

    Multimodal image-tabular learning is gaining attention, yet it faces challenges due to limited labeled data. While earlier work has applied self-supervised learning (SSL) to unlabeled data, its task-agnostic nature often results in learning suboptimal features for downstream tasks. Semi-supervised learning (SemiSL), which combines labeled and unlabeled data,

  67. Songping Wang, Xinquan Yue, Yueming Lyu, Caifeng Shan

    Kolmogorov-Arnold Networks (KANs) have emerged as a transformative model paradigm, significantly impacting various fields. However, their adversarial robustness remains less underexplored, especially across different KAN architectures. To explore this critical safety issue, we conduct an analysis and find that due to overfitting to the specific basis functio

  68. A. Zhadyranova, M. Koussour, Zh. Kanibekova, V. Zhumabekova

    We investigate the divergence-free parametric form of the deceleration parameter within the simplest non-minimal matter-geometry coupling in $f(R,T)$ gravity, where $R$ is the Ricci scalar and $T$ is the trace of the energy-momentum tensor. Specifically, we consider the linear model $f(R,T) = R + 2\lambda T$, where $\lambda$ governs the interaction between m

  69. Elena Agliari, Andrea Alessandrelli, Paulo Duarte Mourao, Alberto Fachechi

    We consider $L$-directional associative memories, composed of $L$ Hopfield networks, displaying imitative Hebbian intra-network interactions and anti-imitative Hebbian inter-network interactions, where couplings are built over a set of hidden binary patterns. We evaluate the model's performance in reconstructing the whole set of hidden binary patterns when p

  70. Jeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis

    We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which learns language-agnostic speech representations by predictin

  71. Gubio Gomes de Lima, Gustavo Miranda, Tiago de Souza Farias

    As t\'ecnicas de aprendizado de m\'aquina emergiram no contexto cient\'ifico e se desenvolveram como ferramentas poderosas para enfrentar uma ampla gama de desafios na sociedade. A integra\c{c}\~ao dessas t\'ecnicas com a f\'isica tem conduzido a abordagens inovadoras na compreens\~ao, controle e simula\c{c}\~ao de fen\^omenos f\'isicos. Este artigo visa pro

  72. Anh Thai, Songyou Peng, Kyle Genova, Leonidas Guibas

    Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D VQA tasks, progress in the 3D domain has been significantly s

  73. Sizhen Bian, Gerald Pirkl, Jingyuan Cheng, Paul Lukowicz

    Using oscillating magnetic fields for indoor positioning is a robust way to resist dynamic environments. This work presents the hard- and software-related optimizations of an induced magnetic field positioning system. We describe a new coil architecture for both the transmitter and receiver, reducing inter-axes cross-talk. A new analog circuit design on the

  74. Shaobin Zhuang, Zhipeng Huang, Binxin Yang, Ying Zhang

    Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure natural instance/scene interactions. We formalize this overlooked yet critical editing paradigm as "Get-In-Video Editing", w

  75. Higinio Serrano, Bernardo Uribe, Miguel A. Xicoténcatl

    We present the fundamental properties of the K-theory groups of complex vector bundles endowed with actions of magnetic groups. In this work we show that the magnetic equivariant K-theory groups define an equivariant cohomology theory, we determine its coefficients, we show Bott's, Thom's and the degree shift isomorphism, we present the Atiyah-Hirzeburh spec

  76. Surender Baswana, Abhyuday Pandey

    Let $G=(V,E)$ be an undirected unweighted multi-graph and $S\subseteq V$ be a subset of vertices. A set of edges with the least cardinality whose removal disconnects $S$, that is, there is no path between at least one pair of vertices from $S$, is called a Steiner mincut for $S$ or simply an $S$-mincut. Connectivity Carcass is a compact data structure storin

  77. Yael Kapon, Dror Merhav, Gal Finkelstein-Zuta, Omer Blumen

    Protein aggregation into insoluble amyloid-like fibrils is implicated in a wide range of diseases and understanding its nucleation process is a key for mechanistic insights and advancing therapeutics. The electronic charge of the amyloidogenic monomers significantly influences their self-assembly process. However, the impact of electron spin interactions bet

  78. M. M. McKinnon

    A number of polarization estimators have been developed for a variety of astrophysical applications to compensate measurements of linear polarization for a bias contributed by the instrumental noise. Most derivations of the estimators assume that the amplitude and orientation of the polarization vector are constant. This assumption generally is not valid for

  79. Benjamin Jensen, Ian Reynolds, Yasir Atalan, Michael Garcia

    As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instr

  80. Michela Varagnolo, Eric Vasserot

    We compare the integral category O of shifted affine quantum groups of symmetric and non symmetric types. To do so we compute the K-theoretic analog of the Coulomb branches with symmetrizers introduced by Nakajima and Weekes. This yields an equivalence of the category O with a module category over a new type of quiver Hecke algebras. At the decategorified le

  81. Wei-En Tai, Yu-Lin Shih, Cheng Sun, Yu-Chiang Frank Wang

    Amodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications including autonomous driving, robotic manipulation, and scene understanding. While existing methods require training both front-end detectors and mask decoders jointly, this approach lacks flexibil

  82. Muzhi Dai, Jiashuo Sun, Zhiyuan Zhao, Shixuan Liu

    Aligning large vision-language models (LVLMs) with human preferences is challenging due to the scarcity of fine-grained, high-quality, and multimodal preference data without human annotations. Existing methods relying on direct distillation often struggle with low-confidence data, leading to suboptimal performance. To address this, we propose CAREVL, a novel

  83. Xiao-Yu Zhang, Pan-Pan Shi, Feng-Kun Guo

    The absence of observed charmonium-like states with the exotic quantum numbers $J^{PC}=1^{-+}$ has prompted us to investigate the production rates of the $1^{-+}$ $D\bar D_1(2420)$ and $D^*\bar D_1(2420)$ hadronic molecules, which we refer to as $\eta_{c1}$ and $\eta_{c1}^{\prime}$, respectively, in electron-positron collisions. Assuming a hadronic molecular

  84. Jan Poštulka, Petr Slavíček, Johannes Kästner, Germán Molpeceres

    Context. Radical chemical reactions on cosmic dust grains play a crucial role in forming various chemical species. Among different radicals, the hydroxyl (OH) is one of the most important ones, with a rather specific chemistry. Aims. The goal of this work is to simulate the recombination dynamics of hydroxyl radicals and the subsequent formation of hydrogen

  85. X. Wang, D. Stroobandt

    Packing is a crucial step of FPGA design, directly impacting interconnect complexity, routing congestion, and overall performance. This paper presents a post-packing interconnect-aware analysis, illustrating how dense (sparse) packing changes the interconnection structure. We introduce a new metric, RDensity, to define post-packing density and investigate it

  86. Seth Hardy

    For $f$ a Steinhaus random multiplicative function, we prove convergence in distribution of the appropriately normalised partial sums \[ \frac{{(\log \log x)}^{1/4}}{\sqrt{x}} \sum_{\substack{n \leq x \\ P(n) > \sqrt{x}}} f(n), \] where $P(n)$ denotes the largest prime factor of $n$. We find that the limiting distribution is given by the square root of an in

  87. Lídia M. André, Jonathan A. Tawn

    Fully describing the entire data set is essential in multivariate risk assessment, since moderate levels of one variable can influence another, potentially leading it to be extreme. Additionally, modelling both non-extreme and extreme events within a single framework avoids the need to select a threshold vector used to determine an extremal region, or the re

  88. Yinuo Liu, Zenghui Yuan, Guiyao Tie, Jiawen Shi

    Multimodal retrieval-augmented generation (RAG) enhances the visual reasoning capability of vision-language models (VLMs) by dynamically accessing information from external knowledge bases. In this work, we introduce \textit{Poisoned-MRAG}, the first knowledge poisoning attack on multimodal RAG systems. Poisoned-MRAG injects a few carefully crafted image-tex

  89. Stefan Schoepf, Muhammad Zaid Hameed, Ambrish Rawat, Kieran Fraser

    With LLM usage rapidly increasing, their vulnerability to jailbreaks that create harmful outputs are a major security risk. As new jailbreaking strategies emerge and models are changed by fine-tuning, continuous testing for security vulnerabilities is necessary. Existing Red Teaming methods fall short in cost efficiency, attack success rate, attack diversity

  90. Kun Xiang, Zhili Liu, Zihao Jiang, Yunshuang Nie

    In this paper, we address the challenging task of multimodal mathematical reasoning by incorporating the ability of "slow thinking" into multimodal large language models (MLLMs). Our core idea is that different levels of reasoning abilities can be combined dynamically to tackle questions with different complexity. To this end, we propose a paradigm of Self-s

  91. Rishabh Gupta, Shivam Gupta, Jaskirat Singh, Sabre Kais

    Short-term patterns in financial time series form the cornerstone of many algorithmic trading strategies, yet extracting these patterns reliably from noisy market data remains a formidable challenge. In this paper, we propose an entropy-assisted framework for identifying high-quality, non-overlapping patterns that exhibit consistent behavior over time. We gr

  92. Kenta Yoshimura, Kazuyuki Sekizawa

    Phase transitions of matter under changes of external environment such as temperature and magnetic field have attracted great interests to various quantum many-body systems. Several phase transitions must have occurred in neutron stars as well such as transitions from normal to superfluid/superconducting phases and crust formation. In this work, we extend th

  93. Ulrich Haisch

    We present a two-loop analysis of the contributions to Higgs production via gluon-gluon fusion arising from the triple-gluon operator in the Standard Model effective field theory (SMEFT). Our discussion covers all aspects of renormalization group (RG) improved perturbation theory, including matching and running within the SMEFT. This study can therefore be s

  94. Leidy M. L. Abril, André A. Moreira, José S. Andrade, Hans J. Herrmann

    Extending the Schramm--Loewner Evolution (SLE) to model branching structures while preserving conformal invariance and other stochastic properties remains a formidable research challenge. Unlike simple paths, branching structures, or trees, must be associated with discontinuous driving functions. Moreover, the driving function of a particular tree is not uni

  95. Minghao Fu, Danning Li, Aryan Gadhiya, Benjamin Lambright

    This paper addresses a major challenge in acoustic event detection, in particular infant cry detection in the presence of other sounds and background noises: the lack of precise annotated data. We present two contributions for supervised and unsupervised infant cry detection. The first is an annotated dataset for cry segmentation, which enables supervised mo

  96. Andrés Fernando Barón Sandoval, Milena Radenkovic

    Access to educational materials in remote Amazonian communities is challenged by limited communication infrastructure. This paper proposes a novel delay-tolerant network (DTN) approach for content distribution and compares the Epidemic, MaxProp, and PRoPHETv2 routing protocols using the ONE simulator under dynamically changing educational file sizes. Results

  97. Debayan Jana, Astik Haldar, Abhik Basu

    We present a hydrodynamic theory of anisotropic and inversion-asymmetric moving active permeable fluid membranes. These are described by an anisotropic Kardar-Parisi-Zhang equation. Depending upon the anisotropy parameters, the membrane is either effectively isotropic and algebraically rough with translational short, but orientational long range order, or un

  98. Aarushi Kalra

    Social media algorithms are thought to amplify variation in user beliefs, thus contributing to radicalization. However, quantitative evidence on how algorithms and user preferences jointly shape harmful online engagement is limited. I conduct an individually randomized experiment with 8 million users of an Indian TikTok-like platform, replacing algorithmic r

  99. J. W. P. Hirschfeld, J. A. Thas

    Arcs and caps are fundamental structures in finite projective spaces. They can be generalised. Here, a survey is given of some important results on these objects, in particular on generalised ovals and generalised ovoids. The paper also contains recent results and several open problems.

  100. Łukasz Struski, Michał B. Bednarczyk, Igor T. Podolak, Jacek Tabor

    We present a novel technique for constructing differentiable order-type operations, including soft ranking, soft top-k selection, and soft permutations. Our approach leverages an efficient closed-form formula for the inverse of the function LapSum, defined as the sum of Laplace distributions. This formulation ensures low computational and memory complexity i