March 2025 arXiv papers — page 179
Showing 17,801–17,900 of 23,633 papers
Gaurav Patel, Qiang Qiu
Machine Unlearning has recently garnered significant attention, aiming to selectively remove knowledge associated with specific data while preserving the model's performance on the remaining data. A fundamental challenge in this process is balancing effective unlearning with knowledge retention, as naive optimization of these competing objectives can lead to
Amin Akhavan
We investigate the role of external constraints in quantum field theory using the path integral formalism. We begin by reviewing the quantization of constrained systems and extend the analysis to cases where constraints are added to the action via auxiliary fields. These constraints involve both the degrees of freedom and their time derivatives. Using the re
Mohit Pandey, Gopeshh Subbaraj, Artem Cherkasov, Martin Ester
Generative Flow Networks (GFlowNets) have recently emerged as a suitable framework for generating diverse and high-quality molecular structures by learning from rewards treated as unnormalized distributions. Previous works in this framework often restrict exploration by using predefined molecular fragments as building blocks, limiting the chemical space that
Jnana Ranjan Das, Santanu Sinha, Alex Hansen, Sitangshu B. Santra
We present a percolation model that is inspired by recent works on immiscible two-phase flow in a mixed-wet porous medium made of a mixture of grains with two different wettability properties. The percolation model is constructed on a dual lattice where the sites on the primal lattice represent the grains of the porous medium, and the bonds on the dual latti
Alex Calderwood, John Joon Young Chung, Yuqian Sun, Melissa Roemmele
According to the recently introduced theory of artistic support tools, creativity support tools exert normative influences over artistic production, instantiating a normative ground that shapes both the process and product of artistic expression. We argue that the normative ground of most existing automated writing tools is misaligned with writerly values an
Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs
cs.LGDingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen
Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities. Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL). The exploration of MCL is in its early
Rashik Shrestha, Madhav Rijal, Trevor Smith, Yu Gu
This study presents Flower Pose Estimation (FloPE), a real-time flower pose estimation framework for computationally constrained robotic pollination systems. Robotic pollination has been proposed to supplement natural pollination to ensure global food security due to the decreased population of natural pollinators. However, flower pose estimation for pollina
Kenneth Stephenson
There exists an extensive and fairly comprehensive discrete analytic function theory which is based on circle packing. This paper introduces a faithful discrete analogue of the classical Schwarzian derivative to this theory and develops its basic properties. The motivation comes from the current lack of circle packing algorithms in spherical geometry, and th
Immersive Virtual Reality Assessments of Working Memory and Psychomotor Skills: A Comparison between Immersive and Non-Immersive Assessments
cs.HCPanagiotis Kourtesis, Andrea Lizarraga, Sarah E. MacPherson
Objective: Immersive virtual reality (VR) enhances ecologically validity and facilitates intuitive and ergonomic hand interactions for performing neuropsychological assessments. However, its comparability to traditional computerized methods remains unclear. This study investigates the convergent validity, user experience, and usability of VR-based versus PC-
Hristo N. Djidjev
Node embedding is a key technique for representing graph nodes as vectors while preserving structural and relational properties, which enables machine learning tasks like feature extraction, clustering, and classification. While classical methods such as DeepWalk, node2vec, and graph convolutional networks learn node embeddings by capturing structural and re
Nikolay Mikhaylovskiy
It is known for some time that autocorrelations of words in human-written texts decay according to a power law. Recent works have also shown that the autocorrelations decay in texts generated by LLMs is qualitatively different from the literary texts. Solid state physics tie the autocorrelations decay laws to the states of matter. In this work, we empiricall
M. H. Shahzamanian
In this paper, we introduce and study a class of monoids, called Layered Catalan Monoids (\( {LC}_n \)), which satisfy the structural conditions for $\ll$-smoothness as defined in~\cite{Sha-Det2}. These monoids are defined by specific identities inspired by Catalan monoids. We establish their canonical forms and compute their determinant, proving that it is
Shlok Nahar, Devashish Tupkary, Norbert Lütkenhaus
Security analyses in quantum key distribution (QKD) and other adversarial quantum tasks often assume perfect device models. However, real-world implementations often deviate from these models. Thus, it is important to develop security proofs that account for such deviations from ideality. In this work, we extend the idea of squashing maps to develop a genera
Altaf Allah Abbassi, Leuson Da Silva, Amin Nikanjam, Foutse Khomh
Large Language Models (LLMs) are widely adopted for automated code generation with promising results. Although prior research has assessed LLM-generated code and identified various quality issues -- such as redundancy, poor maintainability, and sub-optimal performance a systematic understanding and categorization of these inefficiencies remain unexplored. Wi
Evgeny Mukhin, Alexander Varchenko
In [J. Lond. Math. Soc. 109 (2024), e12884, 22 pages, arXiv:2208.09721], the difference qKZ equations were considered modulo a prime number $p$ and a family of polynomial solutions of the qKZ equations modulo $p$ was constructed by an elementary procedure as suitable $p$-approximations of the hypergeometric integrals. In this paper, we study in detail the fi
Vijayamanikandan Vijayarangan, Harshavardhana A. Uranakara, Francisco E. Hernández-Pérez, Hong G. Im
Using the information theory, this study provides insights into how the construction of latent space of autoencoder (AE) using deep neural network (DNN) training finds a smooth low-dimensional manifold in the stiff dynamical system. Our recent study [1] reported that an autoencoder (AE) combined with neural ODE (NODE) as a surrogate reduced order model (ROM)
AnimeGaze: Real-Time Mutual Gaze Synthesis for Anime-Style Avatars in Physical Environments via Behind-Display Camera
cs.HCKazuya Izumi, Shuhey Koyama, Yoichi Ochiai
Avatars on displays lack the ability to engage with the physical environment through gaze. To address this limitation, we propose a gaze synthesis method that enables animated avatars to establish gaze communication with the physical environment using a camera-behind-the-display system. The system uses a display that rapidly alternates between visible and tr
RB Yadav, Arpan Sharma
In this article, we give a characterisation of crossed homomorphisms on Lie superalgebras as a Maurer-Cartan element of a graded Lie algebra. Using this characterisation we study cohomology of these crossed homomorphisms. As an application of this cohomology we study formal deformation of crossed homomorphisms. We show that linear deformations of these homom
Jack Foxabbott, Rohan Subramani, Francis Rhys Ward
Multi-agent influence diagrams (MAIDs) are probabilistic graphical models which represent strategic interactions between agents. MAIDs are equivalent to extensive form games (EFGs) but have a more compact and informative structure. However, MAIDs cannot, in general, represent settings of incomplete information -- wherein agents have different beliefs about t
Jieyang Chen, Qian Gong, Yanliang Li, Xin Liang
The rapid growth of scientific data is surpassing advancements in computing, creating challenges in storage, transfer, and analysis, particularly at the exascale. While data reduction techniques such as lossless and lossy compression help mitigate these issues, their computational overhead introduces new bottlenecks. GPU-accelerated approaches improve perfor
Enhanced Pediatric Dental Segmentation Using a Custom SegUNet with VGG19 Backbone on Panoramic Radiographs
eess.IVMd Ohiduzzaman Ovi, Maliha Sanjana, Fahad Fahad, Mahjabin Runa
Pediatric dental segmentation is critical in dental diagnostics, presenting unique challenges due to variations in dental structures and the lower number of pediatric X-ray images. This study proposes a custom SegUNet model with a VGG19 backbone, designed explicitly for pediatric dental segmentation and applied to the Children's Dental Panoramic Radiographs
Learning and discovering multiple solutions using physics-informed neural networks with random initialization and deep ensemble
cs.LGZongren Zou, Zhicheng Wang, George Em Karniadakis
We explore the capability of physics-informed neural networks (PINNs) to discover multiple solutions. Many real-world phenomena governed by nonlinear differential equations (DEs), such as fluid flow, exhibit multiple solutions under the same conditions, yet capturing this solution multiplicity remains a significant challenge. A key difficulty is giving appro
Ritika Nagpal, S. K. J. Pacif, Farruh Atamurotov, Rasmikanta Pati
In this study, we explore the impact of the interacting parameter on dark matter in a model resulting from a parametrization of dark energy density. To ensure a model-independent approach, we treat \( r_d \) as a free parameter, avoiding assumptions about the physics of the early Universe or specific recombination models. This approach allows late-time cosmo
William Webb, Arturas Medeisis, Leo Fulvio Minervini
This article discusses the key principles of radio spectrum management with a focus on spectrum allocation and access. We show the current regime's inherent rigidity and constrained possibilities for introducing new radiocommunication services and applications. The article proposes how governments and spectrum users could cooperate in taking spectrum managem
Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu
Object detection in videos plays a crucial role in advancing applications such as public safety and anomaly detection. Existing methods have explored different techniques, including CNN, deep learning, and Transformers, for object detection and video classification. However, detecting tiny objects, e.g., guns, in videos remains challenging due to their small
Tieqiao Wang, Sinisa Todorovic
Most recent work on action segmentation relies on pre-computed frame features from models trained on other tasks and typically focuses on framewise encoding and labeling without explicitly modeling action segments. To overcome these limitations, we introduce the End-to-End Action Segmentation Transformer (EAST), which processes raw video frames directly -- e
Luiz A. Ferreira, Aliaksei Mikhaliuk, Yakov Shnir
We present and study new non-topological soliton solutions in the $U(1)$ gauged non-linear $O(3)$ sigma model with a symmetry breaking potential in 3+1 dimensional flat space-time. The configurations are endowed with an electric and magnetic field and also carry a nonvanishing angular momentum density. We discuss properties of these solitons and investigate
Ming-Hua Chang, Steffen Backes, Donghui Lu, Nicolas Gauthier
Understanding how renormalized quasiparticles emerge in strongly correlated electron materials provides a challenge for both experiment and theory. It has been predicted that distinctive spin and orbital screening mechanisms drive this process in multiorbital materials with strong Coulomb and Hund's interactions. Here, we provide the experimental evidence of
Advancing Autonomous Vehicle Intelligence: Deep Learning and Multimodal LLM for Traffic Sign Recognition and Robust Lane Detection
cs.CVChandan Kumar Sah, Ankit Kumar Shaw, Xiaoli Lian, Arsalan Shahid Baig
Autonomous vehicles (AVs) require reliable traffic sign recognition and robust lane detection capabilities to ensure safe navigation in complex and dynamic environments. This paper introduces an integrated approach combining advanced deep learning techniques and Multimodal Large Language Models (MLLMs) for comprehensive road perception. For traffic sign reco
Zhitong Xiong, Yi Wang, Weikang Yu, Adam J Stewart
Earth observation (EO) spans a broad spectrum of modalities, including optical, radar, multispectral, and hyperspectral data, each capturing distinct environmental signals. However, current vision-language models in EO, particularly CLIP-based variants, remain confined to individual modalities, limiting generalization and scalability across diverse tasks. We
Yuxuan Li, Sheng Jinag, Bizhu Wang
With technology advancing and the pursuit of new audiovisual experiences strengthening, the metaverse has gained surging enthusiasm. However, it faces practical hurdles as substantial data like high-resolution virtual scenes must be transmitted between cloud platforms and VR devices. Specifically, the VR device's wireless transmission hampered by insufficien
Hybrid CNN-Dilated Self-attention Model Using Inertial and Body-Area Electrostatic Sensing for Gym Workout Recognition, Counting, and User Authentification
eess.SPSizhen Bian, Vitor Fortes Rey, Siyu Yuan, Paul Lukowicz
While human body capacitance ($HBC$) has been explored as a novel wearable motion sensing modality, its competence has never been quantitatively demonstrated compared to that of the dominant inertial measurement unit ($IMU$) in practical scenarios. This work is thus motivated to evaluate the contribution of $HBC$ in wearable motion sensing. A real-life case
Marco Iannotta, Johannes A. Stork, Erik Schaffernicht, Todor Stoyanov
With the rising demand for flexible manufacturing, robots are increasingly expected to operate in dynamic environments where local -- such as slight offsets or size differences in workpieces -- are common. We propose to address the problem of adapting robot behaviors to these task variations with a sample-efficient hierarchical reinforcement learning approac
Georg Hahn, Sebastian Schneeweiss, Shirley Wang
Computable phenotypes are used to characterize patients and identify outcomes in studies conducted using healthcare claims and electronic health record data. Chart review studies establish reference labels against which computable phenotypes are compared to understand their measurement characteristics, the quantity of interest, for instance the positive pred
Qizhen Lan, Qing Tian
Dense visual prediction tasks, such as detection and segmentation, are crucial for time-critical applications (e.g., autonomous driving and video surveillance). While deep models achieve strong performance, their efficiency remains a challenge. Knowledge distillation (KD) is an effective model compression technique, but existing feature-based KD methods rely
Arnaldo J. Vargas
This work presents a model for testing Lorentz and CPT symmetry using rovibrational transitions within the electronic ground state of the molecular hydrogen ion (H$^+_2$). The model is based on the Standard-Model Extension (SME) and incorporates minimal and nonminimal effects. Our analysis concludes that sidereal variation studies of these transitions could
Anna N. Morozovska, Eugene. A. Eliseev, Oleksiy V. Bereznikov, Mykola Ye. Yelisieiev
The contribution of flexoelectric coupling to the long-range order parameter fluctuations in ferroics can be critically important to the ferron dispersion and related polar, pyroelectric and electrocaloric properties. Here we calculate analytically the dispersion relations of soft optic and acoustic flexocoupling-induced phonons and ferrons by incorporating
Optimization and Benchmarking of Monolithically Stackable Gain Cell Memory for Last-Level Cache
cs.ETFaaiq Waqar, Jungyoun Kwak, Junmo Lee, Minji Shon
The Last Level Cache (LLC) is the processor's critical bridge between on-chip and off-chip memory levels - optimized for high density, high bandwidth, and low operation energy. To date, high-density (HD) SRAM has been the conventional device of choice; however, with the slowing of transistor scaling, as reflected in the industry's almost identical HD SRAM ce
Georg Hahn, Moulinath Banerjee, Bodhisattva Sen
The estimation of regression parameters in one dimensional broken stick models is a research area of statistics with an extensive literature. We are interested in extending such models by aiming to recover two or more intersecting (hyper)planes in multiple dimensions. In contrast to approaches aiming to recover a given number of piecewise linear components u
Synergizing AI and Digital Twins for Next-Generation Network Optimization, Forecasting, and Security
cs.NIZifan Zhang, Minghong Fang, Dianwei Chen, Xianfeng Yang
Digital network twins (DNTs) are virtual representations of physical networks, designed to enable real-time monitoring, simulation, and optimization of network performance. When integrated with machine learning (ML) techniques, particularly federated learning (FL) and reinforcement learning (RL), DNTs emerge as powerful solutions for managing the complexitie
C. O. Edet, K. Słowik, N. Ali, M. Asjad
Controlling heat flow at the quantum level is a key challenge for next-generation quantum technologies, including thermal management and quantum information processing. Here, we investigate quantum heat transport in an asymmetrically driven hybrid magnon-photon system in contact with two thermal baths at different temperatures. We demonstrate that external d
Jeongmin Lee, Sunkyung Park, Minji Lee, Dongjun Lee
This paper presents a framework designed to tackle a range of planning problems arise in manipulation, which typically involve complex geometric-physical reasoning related to contact and dynamic constraints. We introduce the Contact Factor Graph (CFG) to graphically model these diverse factors, enabling us to perform inference on the graphs to approximate th
Highly tunable valley polarization of potential-trapped moir\'e excitons in WSe2/WS2 heterojunctions
cond-mat.mes-hallYueh-Chun Wu, Matthew DeCapua, ZhongChen Xu, Takashi Taniguchi
Moir\'e superlattices created by stacking atomic layers of transition metal dichalcogenide semiconductors have emerged as a class of fascinating artificial photonic and electronic materials. An appealing attribute of these structures is the inheritance of the valley degree of freedom from the constituent monolayers. Recent studies show evidence that the vall
An inviscid limit problem for Navier-Stokes equations in 3D domains with oscillatory boundaries
math.APTuoc Phan, Dario A. Valdebenito
We study an inviscid limit problem for a class of Navier-Stokes equations with vanishing measurable viscous coefficients in 3-dimensional spatial domains whose boundaries are oscillatory, depending on a small parameter, and become flat when the parameter converges to zero. Under some sufficient conditions on the anisotropic vanishing rates of the eigenvalues
Ultra-flexible silicon foils with seamless detachability: the effect of porous multilayered structures prepared through modulated electrolyte composition
physics.app-phClara Sanchez-Perez, Paula Rivas-Lazaro, Elisa García-Tabarés, Iván García Vara
A comprehensive evaluation of the effect and limitations of variable current density and electrolyte composition on layer porosity and microstructure changes of porous silicon (pSi) multilayer stacks is reported. Following these results, the development and optimization of a four-layer stack architecture is reported through addition of super-low porosity lay
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering
cs.CLVinay Kumar Verma, Shreyas Sunil Kulkarni, Happy Mittal, Deepak Gupta
Question Answering (QA) and Visual Question Answering (VQA) are well-studied problems in the language and vision domain. One challenging scenario involves multiple sources of information, each of a different modality, where the answer to the question may exist in one or more sources. This scenario contains richer information but is highly complex to handle.
Jobir Adashev, Xursanoy Berdalova, Feruza Toshtemirova
In this paper we investigate classifications of all (transposed) Poisson algebras of the associated associative null-filiform algebra
Dibakar Roychowdhury
We explore various field theory aspects of integrable $ \eta $-deformed geometry in type IIB supergravity by employing several holographic probes. These include the computation of holographic timelike entanglement entropy and estimation of various other field theory observables for example, the flow central charge and the quantum complexity. We also discuss
Jirui Guo, Mauricio Romo, Lucy Smith
We study the properties of B-branes in a class of nonabelian GLSMs realizing the canonical line bundle $K_{Gr(2,N)}$ in their geometric phase. By analysing the hemisphere partition function, i.e. B-brane central charge, we propose a grade restriction rule and the corresponding window categories for a specific class of paths between phases. We find very strik
Qitan Lv, Tianyu Liu, Hong Wang
Large language models (LLMs) have been widely adopted in mathematical optimization in scientific scenarios for their extensive knowledge and advanced reasoning capabilities. Existing methods mainly focus on utilizing LLMs to solve optimization problems in a prompt-based manner, which takes observational feedback as additional textual descriptions. However, d
HIVQE: Handover Iterative Variational Quantum Eigensolver for Efficient Quantum Chemistry Calculations
quant-phAidan Pellow-Jarman, Shane McFarthing, Doo Hyung Kang, Pilsun Yoo
A novel hybrid quantum-classical approach has been developed to efficiently address the multireference quantum chemistry problem. The Handover Iterative Variational Quantum Eigensolver (HiVQE) is designed to accurately estimate ground-state wavefunctions by leveraging both quantum and classical computing resources. In this framework, noisy intermediate-scale
Haryo Akbarianto Wibowo, Haiyue Song, Hideki Tanaka, Masao Utiyama
Large Language Models (LLMs) have grown increasingly expensive to deploy, driving the need for effective model compression techniques. While block pruning offers a straightforward approach to reducing model size, existing methods often struggle to maintain performance or require substantial computational resources for recovery. We present IteRABRe, a simple
Analysis of Patterns in Recorded Signals of Software Systems With a Variance Based Segmentation Algorithm
stat.APBojan Lukić, Thorben Knust, Andreas Rausch
Due to the increasing complexity and interconnectedness of different components in modern automotive software systems there is a great number of interactions between these system components and their environment. These interactions result in unique temporal behaviors we call underlying scenarios. The signal data from all system components, which is recorded
Mario I. Molina
We study a nonlinear magnetic metamaterial modeled as a split-ring resonator array, where the standard discrete laplacian is replaced by its fractional form. We find a closed-form expression for the dispersion relation as a function of the fractional exponent s and the gain/loss parameter {\gamma} and examine the conditions under which stable magneto-inducti
Xiangyu Yin, Jiaxu Liu, Zhen Chen, Jinwei Hu
Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailb
Hao Yan, Marzi Heidari, Yuhong Guo
Domain Generalization (DG) aims to train models that can generalize to unseen testing domains by leveraging data from multiple training domains. However, traditional DG methods rely on the availability of multiple diverse training domains, limiting their applicability in data-constrained scenarios. Single Domain Generalization (SDG) addresses the more realis
Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang
Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual grounding, though they inevitably require fine-tuning and additional model components to explicitly generate bounding boxes or se
Alessandro T. Gifford, Radoslaw M. Cichy, Thomas Naselaris, Kendrick Kay
Now published in Nature Communications DOI: https://doi.org/10.1038/s41467-026-69345-9 Large-scale visual neural datasets such as the Natural Scenes Dataset (NSD) are boosting computational neuroscience research by enabling models of the brain with performances beyond what was possible just a decade ago. However, because the stimuli of these datasets typical
Gourav Kumar, V. Vetrivel
In this article, we provide a modification to the Bregman Golden Ratio Algorithm (B-GRAAL). We analyze the B-GRAAL algorithm with a new step size rule, where the step size increases after a certain number of iterations and does not require prior knowledge of the global Lipschitz constant of the cost operator. Under suitable assumptions, we establish the glob
Shabnam Ghasemirad, Si Liu, Christoph Sprenger, Luca Multazzu
Isolation bugs, stemming especially from design-level defects, have been repeatedly found in carefully designed and extensively tested production databases over decades. In parallel, various frameworks for modeling database transactions and reasoning about their isolation guarantees have been developed. What is missing however is a mathematically rigorous an
Single-layer magnet phase in intrinsic magnetic topological insulators, $[\mathrm{MnTe}][\mathrm{Bi}_{2}\mathrm{Te}_{3}]_{\mathrm{n}}$, far beyond the thermodynamic limit
cond-mat.mes-hallDeepti Jain, Hee Taek Yi, Xiong Yao, Alessandro R. Mazza
The intrinsic magnetic topological insulator (IMTI) family $[\mathrm{MnTe}][\mathrm{Bi}_{2}\mathrm{Te}_{3}]_{\mathrm{n}}$ has demonstrated magneto-topological properties dependent on $n$, making it a promising platform for advanced electronics and spintronics. However, due to technical barriers in sample synthesis, their properties in the large $n$ limit rem
From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot Learning
cs.CVShuangzhi Li, Junlong Shen, Lei Ma, Xingyu Li
LiDAR-based 3D object detection models often struggle to generalize to real-world environments due to limited object diversity in existing datasets. To tackle it, we introduce the first generalized cross-domain few-shot (GCFS) task in 3D object detection, aiming to adapt a source-pretrained model to both common and novel classes in a new domain with only few
Ramanath Cowsik, Dawson Huth
The leaky-box model and the attendant concept of path-length distribution of cosmic rays were invented in the mid-1960's. Even though versatile computational packages such as GALPROP and DRAGON with the diffusion approach are now available for analyzing cosmic ray data, the concepts of the leaky-box and path-length distribution continue to be adopted extensi
Mitigating Blockchain extractable value (BEV) threats by Distributed Transaction Sequencing in Blockchains
cs.CRXiongfei Zhao, Hou-Wan Long, Zhengzhe Li, Jiangchuan Liu
The rapid growth of Blockchain and Decentralized Finance (DeFi) has introduced new challenges and vulnerabilities that threaten the integrity and efficiency of the ecosystem. This study identifies critical issues such as Transaction Order Dependence (TOD), Blockchain Extractable Value (BEV), and Transaction Importance Diversity (TID), which collectively unde
Applied Machine Learning Methods with Long-Short Term Memory Based Recurrent Neural Networks for Multivariate Temperature Prediction
cs.LGBojan Lukić
This paper gives an overview on how to develop a dense and deep neural network for making a time series prediction. First, the history and cornerstones in Artificial Intelligence and Machine Learning will be presented. After a short introduction to the theory of Artificial Intelligence and Machine Learning, the paper will go deeper into the techniques for co
STiL: Semi-supervised Tabular-Image Learning for Comprehensive Task-Relevant Information Exploration in Multimodal Classification
cs.CVSiyi Du, Xinzhe Luo, Declan P. O'Regan, Chen Qin
Multimodal image-tabular learning is gaining attention, yet it faces challenges due to limited labeled data. While earlier work has applied self-supervised learning (SSL) to unlabeled data, its task-agnostic nature often results in learning suboptimal features for downstream tasks. Semi-supervised learning (SemiSL), which combines labeled and unlabeled data,
Songping Wang, Xinquan Yue, Yueming Lyu, Caifeng Shan
Kolmogorov-Arnold Networks (KANs) have emerged as a transformative model paradigm, significantly impacting various fields. However, their adversarial robustness remains less underexplored, especially across different KAN architectures. To explore this critical safety issue, we conduct an analysis and find that due to overfitting to the specific basis functio
A. Zhadyranova, M. Koussour, Zh. Kanibekova, V. Zhumabekova
We investigate the divergence-free parametric form of the deceleration parameter within the simplest non-minimal matter-geometry coupling in $f(R,T)$ gravity, where $R$ is the Ricci scalar and $T$ is the trace of the energy-momentum tensor. Specifically, we consider the linear model $f(R,T) = R + 2\lambda T$, where $\lambda$ governs the interaction between m
Elena Agliari, Andrea Alessandrelli, Paulo Duarte Mourao, Alberto Fachechi
We consider $L$-directional associative memories, composed of $L$ Hopfield networks, displaying imitative Hebbian intra-network interactions and anti-imitative Hebbian inter-network interactions, where couplings are built over a set of hidden binary patterns. We evaluate the model's performance in reconstructing the whole set of hidden binary patterns when p
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
cs.CVJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis
We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which learns language-agnostic speech representations by predictin
Gubio Gomes de Lima, Gustavo Miranda, Tiago de Souza Farias
As t\'ecnicas de aprendizado de m\'aquina emergiram no contexto cient\'ifico e se desenvolveram como ferramentas poderosas para enfrentar uma ampla gama de desafios na sociedade. A integra\c{c}\~ao dessas t\'ecnicas com a f\'isica tem conduzido a abordagens inovadoras na compreens\~ao, controle e simula\c{c}\~ao de fen\^omenos f\'isicos. Este artigo visa pro
Anh Thai, Songyou Peng, Kyle Genova, Leonidas Guibas
Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D vision-language models (VLMs) have achieved remarkable success in 2D VQA tasks, progress in the 3D domain has been significantly s
Sizhen Bian, Gerald Pirkl, Jingyuan Cheng, Paul Lukowicz
Using oscillating magnetic fields for indoor positioning is a robust way to resist dynamic environments. This work presents the hard- and software-related optimizations of an induced magnetic field positioning system. We describe a new coil architecture for both the transmitter and receiver, reducing inter-axes cross-talk. A new analog circuit design on the
Shaobin Zhuang, Zhipeng Huang, Binxin Yang, Ying Zhang
Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure natural instance/scene interactions. We formalize this overlooked yet critical editing paradigm as "Get-In-Video Editing", w
Higinio Serrano, Bernardo Uribe, Miguel A. Xicoténcatl
We present the fundamental properties of the K-theory groups of complex vector bundles endowed with actions of magnetic groups. In this work we show that the magnetic equivariant K-theory groups define an equivariant cohomology theory, we determine its coefficients, we show Bott's, Thom's and the degree shift isomorphism, we present the Atiyah-Hirzeburh spec
Surender Baswana, Abhyuday Pandey
Let $G=(V,E)$ be an undirected unweighted multi-graph and $S\subseteq V$ be a subset of vertices. A set of edges with the least cardinality whose removal disconnects $S$, that is, there is no path between at least one pair of vertices from $S$, is called a Steiner mincut for $S$ or simply an $S$-mincut. Connectivity Carcass is a compact data structure storin
Yael Kapon, Dror Merhav, Gal Finkelstein-Zuta, Omer Blumen
Protein aggregation into insoluble amyloid-like fibrils is implicated in a wide range of diseases and understanding its nucleation process is a key for mechanistic insights and advancing therapeutics. The electronic charge of the amyloidogenic monomers significantly influences their self-assembly process. However, the impact of electron spin interactions bet
M. M. McKinnon
A number of polarization estimators have been developed for a variety of astrophysical applications to compensate measurements of linear polarization for a bias contributed by the instrumental noise. Most derivations of the estimators assume that the amplitude and orientation of the polarization vector are constant. This assumption generally is not valid for
Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
cs.CYBenjamin Jensen, Ian Reynolds, Yasir Atalan, Michael Garcia
As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of large language models (LLMs) is crucial. This study presents a novel benchmark designed to evaluate the biases and preferences of seven prominent foundation models-Llama 3.1 8B Instr
Michela Varagnolo, Eric Vasserot
We compare the integral category O of shifted affine quantum groups of symmetric and non symmetric types. To do so we compute the K-theoretic analog of the Coulomb branches with symmetrizers introduced by Nakajima and Weekes. This yields an equivalence of the category O with a module category over a new type of quiver Hecke algebras. At the decategorified le
Wei-En Tai, Yu-Lin Shih, Cheng Sun, Yu-Chiang Frank Wang
Amodal instance segmentation, which aims to detect and segment both visible and invisible parts of objects in images, plays a crucial role in various applications including autonomous driving, robotic manipulation, and scene understanding. While existing methods require training both front-end detectors and mask decoders jointly, this approach lacks flexibil
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
cs.CVMuzhi Dai, Jiashuo Sun, Zhiyuan Zhao, Shixuan Liu
Aligning large vision-language models (LVLMs) with human preferences is challenging due to the scarcity of fine-grained, high-quality, and multimodal preference data without human annotations. Existing methods relying on direct distillation often struggle with low-confidence data, leading to suboptimal performance. To address this, we propose CAREVL, a novel
Xiao-Yu Zhang, Pan-Pan Shi, Feng-Kun Guo
The absence of observed charmonium-like states with the exotic quantum numbers $J^{PC}=1^{-+}$ has prompted us to investigate the production rates of the $1^{-+}$ $D\bar D_1(2420)$ and $D^*\bar D_1(2420)$ hadronic molecules, which we refer to as $\eta_{c1}$ and $\eta_{c1}^{\prime}$, respectively, in electron-positron collisions. Assuming a hadronic molecular
Diffusive vs. non-diffusive paths to interstellar hydrogen peroxide. A machine learning-based molecular dynamics study
astro-ph.GAJan Poštulka, Petr Slavíček, Johannes Kästner, Germán Molpeceres
Context. Radical chemical reactions on cosmic dust grains play a crucial role in forming various chemical species. Among different radicals, the hydroxyl (OH) is one of the most important ones, with a rather specific chemistry. Aims. The goal of this work is to simulate the recombination dynamics of hydroxyl radicals and the subsequent formation of hydrogen
X. Wang, D. Stroobandt
Packing is a crucial step of FPGA design, directly impacting interconnect complexity, routing congestion, and overall performance. This paper presents a post-packing interconnect-aware analysis, illustrating how dense (sparse) packing changes the interconnection structure. We introduce a new metric, RDensity, to define post-packing density and investigate it
The distribution of partial sums of random multiplicative functions with a large prime factor
math.NTSeth Hardy
For $f$ a Steinhaus random multiplicative function, we prove convergence in distribution of the appropriately normalised partial sums \[ \frac{{(\log \log x)}^{1/4}}{\sqrt{x}} \sum_{\substack{n \leq x \\ P(n) > \sqrt{x}}} f(n), \] where $P(n)$ denotes the largest prime factor of $n$. We find that the limiting distribution is given by the square root of an in
Gaussian mixture copulas for flexible dependence modelling in the body and tails of joint distributions
stat.MELídia M. André, Jonathan A. Tawn
Fully describing the entire data set is essential in multivariate risk assessment, since moderate levels of one variable can influence another, potentially leading it to be extreme. Additionally, modelling both non-extreme and extreme events within a single framework avoids the need to select a threshold vector used to determine an extremal region, or the re
Yinuo Liu, Zenghui Yuan, Guiyao Tie, Jiawen Shi
Multimodal retrieval-augmented generation (RAG) enhances the visual reasoning capability of vision-language models (VLMs) by dynamically accessing information from external knowledge bases. In this work, we introduce \textit{Poisoned-MRAG}, the first knowledge poisoning attack on multimodal RAG systems. Poisoned-MRAG injects a few carefully crafted image-tex
Stefan Schoepf, Muhammad Zaid Hameed, Ambrish Rawat, Kieran Fraser
With LLM usage rapidly increasing, their vulnerability to jailbreaks that create harmful outputs are a major security risk. As new jailbreaking strategies emerge and models are changed by fine-tuning, continuous testing for security vulnerabilities is necessary. Existing Red Teaming methods fall short in cost efficiency, attack success rate, attack diversity
Can Atomic Step Decomposition Enhance the Self-structured Reasoning of Multimodal Large Models?
cs.CVKun Xiang, Zhili Liu, Zihao Jiang, Yunshuang Nie
In this paper, we address the challenging task of multimodal mathematical reasoning by incorporating the ability of "slow thinking" into multimodal large language models (MLLMs). Our core idea is that different levels of reasoning abilities can be combined dynamically to tackle questions with different complexity. To this end, we propose a paradigm of Self-s
Rishabh Gupta, Shivam Gupta, Jaskirat Singh, Sabre Kais
Short-term patterns in financial time series form the cornerstone of many algorithmic trading strategies, yet extracting these patterns reliably from noisy market data remains a formidable challenge. In this paper, we propose an entropy-assisted framework for identifying high-quality, non-overlapping patterns that exhibit consistent behavior over time. We gr
Phase transitions in the inner crust of neutron stars within the superfluid band theory: Competition between $^1\text{S}_0$ pairing and spin polarization under finite temperature and magnetic field
nucl-thKenta Yoshimura, Kazuyuki Sekizawa
Phase transitions of matter under changes of external environment such as temperature and magnetic field have attracted great interests to various quantum many-body systems. Several phase transitions must have occurred in neutron stars as well such as transitions from normal to superfluid/superconducting phases and crust formation. In this work, we extend th
Ulrich Haisch
We present a two-loop analysis of the contributions to Higgs production via gluon-gluon fusion arising from the triple-gluon operator in the Standard Model effective field theory (SMEFT). Our discussion covers all aspects of renormalization group (RG) improved perturbation theory, including matching and running within the SMEFT. This study can therefore be s
Leidy M. L. Abril, André A. Moreira, José S. Andrade, Hans J. Herrmann
Extending the Schramm--Loewner Evolution (SLE) to model branching structures while preserving conformal invariance and other stochastic properties remains a formidable research challenge. Unlike simple paths, branching structures, or trees, must be associated with discontinuous driving functions. Moreover, the driving function of a particular tree is not uni
Minghao Fu, Danning Li, Aryan Gadhiya, Benjamin Lambright
This paper addresses a major challenge in acoustic event detection, in particular infant cry detection in the presence of other sounds and background noises: the lack of precise annotated data. We present two contributions for supervised and unsupervised infant cry detection. The first is an annotated dataset for cry segmentation, which enables supervised mo
AmazonNetLink: Enabling Education Access in Remote Amazonian Regions through Delay-Tolerant Networks
cs.NIAndrés Fernando Barón Sandoval, Milena Radenkovic
Access to educational materials in remote Amazonian communities is challenged by limited communication infrastructure. This paper proposes a novel delay-tolerant network (DTN) approach for content distribution and compares the Epidemic, MaxProp, and PRoPHETv2 routing protocols using the ONE simulator under dynamically changing educational file sizes. Results
Debayan Jana, Astik Haldar, Abhik Basu
We present a hydrodynamic theory of anisotropic and inversion-asymmetric moving active permeable fluid membranes. These are described by an anisotropic Kardar-Parisi-Zhang equation. Depending upon the anisotropy parameters, the membrane is either effectively isotropic and algebraically rough with translational short, but orientational long range order, or un
Aarushi Kalra
Social media algorithms are thought to amplify variation in user beliefs, thus contributing to radicalization. However, quantitative evidence on how algorithms and user preferences jointly shape harmful online engagement is limited. I conduct an individually randomized experiment with 8 million users of an Indian TikTok-like platform, replacing algorithmic r
J. W. P. Hirschfeld, J. A. Thas
Arcs and caps are fundamental structures in finite projective spaces. They can be generalised. Here, a survey is given of some important results on these objects, in particular on generalised ovals and generalised ovoids. The paper also contains recent results and several open problems.
Łukasz Struski, Michał B. Bednarczyk, Igor T. Podolak, Jacek Tabor
We present a novel technique for constructing differentiable order-type operations, including soft ranking, soft top-k selection, and soft permutations. Our approach leverages an efficient closed-form formula for the inverse of the function LapSum, defined as the sum of Laplace distributions. This formulation ensures low computational and memory complexity i