May 2025 arXiv papers — page 89
Showing 8,801–8,900 of 24,552 papers
Rembert Daems, Manfred Opper, Guillaume Crevecoeur, Tolga Birdal
We present a hierarchical, control theory inspired method for variational inference (VI) for neural stochastic differential equations (SDEs). While VI for neural SDEs is a promising avenue for uncertainty-aware reasoning in time-series, it is computationally challenging due to the iterative nature of maximizing the ELBO. In this work, we propose to decompose
Alexander Boeschoten, Giacomo Sorelli, Manuel Gessner, Claude Fabre
We investigate the problem of estimating simultaneously multiple parameters encoded in the shape of the modes on which the light is expanded. For this, we generalize the mode-encoded parameter estimation theory as introduced in Ref.[1] to a multi-parameter scenario. We derive the general expression for the Quantum Fisher information matrix and establish the
Ranjith Merugu, Mohammad Sameer Suhail, Akshay P Sarashetti, Venkata Bharath Reddy Reddem
Recent advancements in video restoration have focused on recovering high-quality video frames from low-quality inputs. Compared with static images, the performance of video restoration significantly depends on efficient exploitation of temporal correlations among successive video frames. The numerous techniques make use of temporal information via flow-based
Yasuyuki Hatsuda, Takaki Matsumoto, Kazumi Okuyama
It is well-known that the partition function of the Jackiw-Teitelboim (JT) gravity is obtained by an integral transformation of volumes of moduli spaces for Riemann surfaces, also known as the Weil-Petersson volumes. This fact enables us to compute the perturbative genus expansion of the partition function by solving a KdV-type non-linear partial differentia
High-pressure high-temperature solution growth, structural, and superconducting properties of Fe-substituted MgB2 single crystals
cond-mat.supr-conN. D. Zhigadlo, R. Puzniak
Clarifying the impact of Fe doping on the structural and superconducting properties of MgB2 is crucial, considering that iron is commonly used as a sheath material for the fabrication of metal-clad MgB2 wires and tapes. To date the effects of Fe doping have only been investigated in polycrystalline samples, but the obtained results are controversial. Here, w
Samuel Humeau, Damien Pous
A tuple (s1,t1,s2,t2) of vertices in a simple undirected graph is 2-linked when there are two vertex-disjoint paths respectively from s1 to t1 and s2 to t2. A graph is 2-linked when all such tuples are 2-linked. We give a new and simple proof of the ``two paths theorem'', a characterisation of edge-maximal graphs which are not 2-linked as webs: particular ne
Martin Goodfellow, Robbie Booth, Andrew Fagan, Alasdair Lambert
Students often do not fully understand the code they have written. This sometimes does not become evident until later in their education, which can mean it is harder to fix their incorrect knowledge or misunderstandings. In addition, being able to fully understand code is increasingly important in a world where students have access to generative artificial i
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
cs.CLSong Jin, Juntian Zhang, Yuhan Liu, Xun Zhang
Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interactions. While agent-based simulation is promising, existing platforms often lack a mechanism for user actions to dynamically reshape the environment. To bridge this gap, we introduce Re
Luca Nils Philipp, Eva Münzel, Julian Lüttig, Roland Mitrić
Forming new hybrid quasiparticles by strong light-matter coupling is a promising tool for tailoring photophysics and photochemistry of molecules. Thus, the ultrafast dynamics of polaritons formed upon strong light-matter coupling has been extensively studied by pump-probe spectroscopy. Although it was predicted that the partial photonic character of polarito
Dayanand Mishra
In this talk, I will present the calculation of LCSR predictions for the $B \to K$ Hadronic Matrix Elements (HME) at low $q^2$ using light meson distribution amplitudes. I will discuss the results obtained.
Ying-Ying Jin, Ye-Qing Sheng, Yi-Ting Wang, Li-Hong Xie
We present a characterization of paratopological gyrogroups that can be topologically embedded as subgyrogroups into a product of first-countable $T_{i}$ paratopological gyrogroups for $i = 0, 1, 2$. Specifically, we demonstrate that a strongly paratopological gyrogroup $G$ is topologically isomorphic to a subgyrogroup of a topological product of first-count
Jing Bi, Pinxin Liu, Ali Vosoughi, Jiarui Wu
The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey complex physical actions and spatial relationships. We address this limitation by proposing a language-driven framework that translates procedural text into coherent visual instructions.
Web Element Relocalization in Evolving Web Applications: A Comparative Analysis and Extension Study
cs.SEAnton Kluge, Andrea Stocco
Fragile web tests, primarily caused by locator breakages, are a persistent challenge in web development. Hence, researchers have proposed techniques for web-element re-identification in which algorithms utilize a range of element properties to relocate elements on updated versions of websites based on similarity scoring. In this paper, we replicate the origi
Enrico Da Ronche
In this paper we generalize the notion of logarithmic vector-valued modular form in order to give a general definition of matrix-valued Hilbert modular forms. We prove that they admit unique polynomial Fourier expansions and we build examples in some particular cases.
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
cs.CVXiaoran Yin, Xu Luo, Hao Wu, Lianli Gao
The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental information available at each step, primarily through visual observations. As a result, current approaches, which typically rely on reactive pol
Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang
While reinforcement learning (RL) has demonstrated remarkable success in enhancing large language models (LLMs), it has primarily focused on single-turn tasks such as solving math problems. Training effective web agents for multi-turn interactions remains challenging due to the complexity of long-horizon decision-making across dynamic web interfaces. In this
Meta-Calibration of the Cosmic Magnification Coefficient: Toward Unbiased Weak Lensing Reconstruction by Counting Galaxies
astro-ph.COJian Qin, Pengjie Zhang, Zhu Chen, Liping Fu
Weak lensing alters galaxy sizes and fluxes, influencing the clustering patterns of galaxies through cosmic magnification. This effect enables the reconstruction of weak lensing convergence $\hat{\kappa}$ maps for DES and DECaLS by linearly combining galaxy overdensities across magnitude bins in the $g$, $r$, and $z$ photometry bands \citep{Qin+,Qin2+}. In t
Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment
cs.CVSoh Takahashi, Masaru Sasaki, Ken Takeda, Masafumi Oizumi
The learning mechanisms by which humans acquire internal representations of objects are not fully understood. Deep neural networks (DNNs) have emerged as a useful tool for investigating this question, as they have internal representations similar to those of humans as a byproduct of optimizing their objective functions. While previous studies have shown that
Yoichi Aoki, Soichiro Murakami, Ukyo Honda, Akihiko Kato
In natural language generation for advertising, creating diverse and engaging ad texts is crucial for capturing a broad audience and avoiding advertising fatigue. Regardless of the importance of diversity, the impact of the diversity-enhancing methods in ad text generation -- mainly tested on tasks such as summarization and machine translation -- has not bee
Iker de las Heras, Benjamin Klopsch, Anitha Thillaisundaram
We establish that finitely generated non-abelian direct products $G$ of free pro-$p$ groups have full Hausdorff spectrum with respect to the lower $p$-series $\mathcal{L}$. This complements similar results with respect to other standard filtration series and a recent theorem showing that the Hausdorff spectrum $\text{hspec}^\mathcal{L}(G)$ of a $p$-adic anal
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
cs.CLRuizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao
Retrieval-Augmented Generation (RAG) leverages large language models (LLMs) combined with external contexts to enhance the accuracy and reliability of generated responses. However, reliably attributing generated content to specific context segments, context attribution, remains challenging due to the computationally intensive nature of current methods, which
Qin Chen, Yuanyi Ren, Xiaojun Ma, Yuyang Shi
Predictive analysis is a cornerstone of modern decision-making, with applications in various domains. Large Language Models (LLMs) have emerged as powerful tools in enabling nuanced, knowledge-intensive conversations, thus aiding in complex decision-making tasks. With the burgeoning expectation to harness LLMs for predictive analysis, there is an urgent need
Critical mean field equations for equilibrium turbulence with sign-changing prescribed functions
math.APLinlin Sun, Xiaobao Zhu
Let $(M,g)$ be a compact Riemann surface with unit area. We investigate the mean field equation for equilibrium turbulence: \begin{align} \begin{cases} -\Delta u = \rho_1\left(\frac{h_1e^{u}}{\int_Mh_1e^udv_g}-1\right) - \rho_2\left(\frac{h_2e^{-u}}{\int_Mh_2e^{-u}dv_g}-1\right), \\ \int_Mudv_g=0, \end{cases} \end{align} where $\rho_1=8\pi$ and $\rho_2\in(0,
Nikolay Stanishev, Yuhang Lu, Touradj Ebrahimi
Pose-invariant face recognition has become a challenging problem for modern AI-based face recognition systems. It aims at matching a profile face captured in the wild with a frontal face registered in a database. Existing methods perform face frontalization via either generative models or learning a pose robust feature representation. In this paper, a new me
Sreetama Sarkar, Yue Che, Alex Gavin, Peter A. Beerel
Despite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucinations", generating texts misaligned with the visual context. Existing methods aimed at reducing hallucinations through inference time intervention incur a significant increase in latency. To mitigate this, we present SPIN, a t
Guanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin
Recently, large language models (LLMs) have shown remarkable reasoning capabilities via large-scale reinforcement learning (RL). However, leveraging the RL algorithm to empower effective multi-tool collaborative reasoning in LLMs remains an open challenge. In this paper, we introduce Tool-Star, an RL-based framework designed to empower LLMs to autonomously i
Chaeeun Kim, Seungone Kim
Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in multi-step reasoning and calling search engines at appropriate steps. However, existing retrieval-augmented reasoning approaches rely on separate retrieval models, limiting the LRM's role in retrieval to deciding when to retrieve and how to query. This separation not only increases ha
Muhammad Farid Adilazuarda, Chen Cecilia Liu, Iryna Gurevych, Alham Fikri Aji
Adapting cultural values in Large Language Models (LLMs) presents significant challenges, particularly due to biases and limited training data. Prior work primarily aligns LLMs with different cultural values using World Values Survey (WVS) data. However, it remains unclear whether this approach effectively captures cultural nuances or produces distinct cultu
Robust Longitudinal-lateral Look-ahead Pursuit Path-Following Control: Fast Finite-Time Stability Guarantee
eess.SYZimao Sheng, Hong'an Yang, Shuxiang Yang, Zirui Yu
This paper addresses the challenging problem of robust path-following for fixed-wing unmanned aerial vehicles (UAVs) in complex environments with bounded external disturbances and non-smooth predefined paths. Due to the unique aerodynamic characteristics and flight constraints of fixed-wing UAVs, achieving accurate and fast stable path following remains diff
Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi
As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution method
Xinxin Chen, Yong Han, Yanqi Qiu, Zipeng Wang
We give a complete solution to the Mandelbrot-Kahane problem for the microcanonical cascade measures by determing their exact Fourier dimensions. We also discuss the Frostman regularity as well as the bi-H\"older continuity of the Dubins-Freedman random homeomorphisms.
Kishan Gupta, Srikanth Korse, Andreas Brendel, Nicola Pia
In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, engendered bitrate and computational complexity. Most conventional and neural speech codecs operate on wideband (WB) speech signals to achiev
Huazi Pan, Yanjun Zhang, Leo Yu Zhang, Scott Adams
Manipulation of local training data and local updates, i.e., the poisoning attack, is the main threat arising from the collaborative nature of the federated learning (FL) paradigm. Most existing poisoning attacks aim to manipulate local data/models in a way that causes denial-of-service (DoS) issues. In this paper, we introduce a novel attack method, named F
AdvReal: Physical Adversarial Patch Generation Framework for Security Evaluation of Object Detection Systems
cs.CVYuanhao Huang, Yilong Ren, Jinlei Wang, Lujia Huo
Autonomous vehicles are typical complex intelligent systems with artificial intelligence at their core. However, perception methods based on deep learning are extremely vulnerable to adversarial samples, resulting in security accidents. How to generate effective adversarial examples in the physical world and evaluate object detection systems is a huge challe
Xiaoqing Zhang, Huabin Zheng, Ang Lv, Yuhan Liu
Large language models (LLMs) have been observed to suddenly exhibit advanced reasoning abilities during reinforcement learning (RL), resembling an ``aha moment'' triggered by simple outcome-based rewards. While RL has proven effective in eliciting such breakthroughs in tasks involving mathematics, coding, and vision, it faces significant challenges in multi-
Yang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee
Despite recent progress in large-scale reinforcement learning (RL) for reasoning, the training recipe for building high-performing reasoning models remains elusive. Key implementation details of frontier models, such as DeepSeek-R1, including data curation strategies and RL training recipe, are often omitted. Moreover, recent research indicates distillation
Qian Deng, Le Hui, Jin Xie, Jian Yang
Bounding box supervision has gained considerable attention in weakly supervised 3D instance segmentation. While this approach alleviates the need for extensive point-level annotations, obtaining accurate bounding boxes in practical applications remains challenging. To this end, we explore the inaccurate bounding box, named sketchy bounding box, which is imit
Consistent and Compatible Modelling of Cyber Intrusions and Incident Response Demonstrated in the Context of Malware Attacks on Critical Infrastructure
cs.CRPeter Maynard, Yulia Cherdantseva, Avi Shaked, Pete Burnap
Cyber Security Incident Response (IR) Playbooks are used to capture the steps required to recover from a cyber intrusion. Individual IR playbooks should focus on a specific type of incident and be aligned with the architecture of a system under attack. Intrusion modelling focuses on a specific potential cyber intrusion and is used to identify where and what
Koki Nagakura, Tatsuki Fushimi, Ayaka Tsutsui, Yoichi Ochiai
This paper presents a method for generating dynamic caustic patterns by utilising dual-optimised holographic fields with Phased Array Transducer (PAT). Building on previous research in static caustic optimisation and ultrasonic manipulation, this approach employs computational techniques to dynamically shape fluid surfaces, thereby creating controllable and
Trajectory-Independent Flexibility Envelopes of Energy-Constrained Systems with State-Dependent Losses
eess.SYJulie Rousseau, Carlo Tajoli, Hanmin Cai, Philipp Heer
As non-dispatchable renewable power units become prominent in electric power grids, demand-side flexibility appears as a key element of future power systems' operation. Power and energy bounds are intuitive metrics to describe the flexibility of energy-constrained loads. However, to be used in operation, any power consumption trajectory fulfilling the power
Yuxin Kang, Xin Zeng, Wuji Zhang, Chunfang Sun
The generation of magnon entanglement and squeezing plays a crucial role in quantum information processing. In this study, we propose a scheme based on a chiral cavity-magnon system, which consists of a torus-shaped cavity and two yttrium iron garnet spheres. The magnon mode of each yttrium iron garnet sphere is selectively coupled to one of the two degenera
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)
cs.ROZhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang
Reinforcement Learning (RL) can mitigate the causal confusion and distribution shift inherent to imitation learning (IL). However, applying RL to end-to-end autonomous driving (E2E-AD) remains an open problem for its training difficulty, and IL is still the mainstream paradigm in both academia and industry. Recently Model-based Reinforcement Learning (MBRL)
LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice
cs.SETristan Karch, Jakhongir Saydaliev, Isabella Di Lenardo, Frédéric Kaplan
Cadastral data reveal key information about the historical organization of cities but are often non-standardized due to diverse formats and human annotations, complicating large-scale analysis. We explore as a case study Venice's urban history during the critical period from 1740 to 1808, capturing the transition following the fall of the ancient Republic an
Mateja Hrast, Georgios M. Koutentakis, Mikhail Maslov, Mikhail Lemeshko
Helical dichroism (HD) is a proposed method for the resolution of molecular chirality, employing the orbital angular momentum (OAM) of light. Going beyond the conventional assumptions about HD, this work proposes a rigid theoretical framework for the analysis of the HD, based on molecular symmetries and rotational eigenstates. We derive the rotational select
Benjamin Vendeville, Liana Ermakova, Pierre De Loor
The general public often encounters complex texts but does not have the time or expertise to fully understand them, leading to the spread of misinformation. Automatic Text Simplification (ATS) helps make information more accessible, but its evaluation methods have not kept up with advances in text generation, especially with Large Language Models (LLMs). In
Quantum-Driven Multihead Inland Waterbody Detection With Transformer-Encoded CYGNSS Delay-Doppler Map Data
eess.IVChia-Hsiang Lin, Jhao-Ting Lin, Po-Ying Chiu, Shih-Ping Chen
Inland waterbody detection (IWD) is critical for water resources management and agricultural planning. However, the development of high-fidelity IWD mapping technology remains unresolved. We aim to propose a practical solution based on the easily accessible data, i.e., the delay-Doppler map (DDM) provided by NASA's Cyclone Global Navigation Satellite System
Deyu Song, Xiangyin Zhang, Zipei Yu, Kaiyu Qin
Multi-view Synthetic Aperture Radar (SAR) imaging can effectively enhance the performance of tasks such as automatic target recognition and image information fusion. Unmanned aerial vehicles (UAVs) have the advantages of flexible deployment and cost reduction. A swarm of UAVs equipped with synthetic aperture radar imaging equipment is well suited to meet the
Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge
eess.ASMing Cheng, Fei Su, Cancan Li, Juan Liu
This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework to generate initial predictions using single-channel audio. Then, we extend the original S2SND framework to create a new version, Multi-Chan
Omni TM-AE: A Scalable and Interpretable Embedding Model Using the Full Tsetlin Machine State Space
cs.LGAhmed K. Kadhim, Lei Jiao, Rishad Shafik, Ole-Christoffer Granmo
The increasing complexity of large-scale language models has amplified concerns regarding their interpretability and reusability. While traditional embedding models like Word2Vec and GloVe offer scalability, they lack transparency and often behave as black boxes. Conversely, interpretable models such as the Tsetlin Machine (TM) have shown promise in construc
Kaiyu He, Tong Zhou, Yubo Chen, Delai Qiu
Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify the cross-lingual ability of LLMs accurately, we propose a Word-Level Cross-Lingual Translation Task. To find how LLMs learn cross-lingual ability, we trace the outputs of LLMs' int
Haoming Huang, Musen Zhang, Jianxin Yang, Zhen Li
Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interaction (HRI). However, existing gaze estimation methods merely predict either the gaze direction or the Point-of-Gaze (PoG) on the screen, failing to provide sufficient information for a comprehensive six Degree-of-
Filling in the Blanks? A Systematic Review and Theoretical Conceptualisation for Measuring WikiData Content Gaps
cs.SIMarisa Ripoll, Neal Reeves, Anelia Kurteva, Elena Simperl
Wikidata is a collaborative knowledge graph which provides machine-readable structured data for Wikimedia projects including Wikipedia. Managed by a community of volunteers, it has grown to become the most edited Wikimedia project. However, it features a long-tail of items with limited data and a number of systematic gaps within the available content. In thi
Priyanka Mishra, Nevill Gonzalez Szwacki
We present a comprehensive first-principles investigation of the structural, electronic, and vibrational properties of four layered boron nitride (BN) polymorphs--AA-stacked ($e$-BN), AA$^\prime$-stacked ($h$-BN), ABC-stacked ($r$-BN), and AB-stacked ($b$-BN). Using density functional theory and density functional perturbation theory with and without van der
Songlin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan
The attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains such as language. Rotary position encoding (RoPE) has emerged as the de facto standard approach for position encoding and is part of many mod
On the Identification of Exotic Compact Binaries with Gravitational Waves: a Phenomenological approach
gr-qcShrobana Ghosh, Mark Hannam
Gravitational wave (GW) astronomy has been hailed as a gateway to discovering unexpected phenomena in the universe. Over the last decade there have been close to one hundred GW observations of compact-binary mergers. While these signals are largely consistent with mergers of binary black holes, binary neutron stars, or black hole-neutron star systems, some e
Zhixun Li, Bin Cao, Rui Jiao, Liang Wang
Materials are the foundation of modern society, underpinning advancements in energy, electronics, healthcare, transportation, and infrastructure. The ability to discover and design new materials with tailored properties is critical to solving some of the most pressing global challenges. In recent years, the growing availability of high-quality materials data
VLM-SAFE: Vision-Language Model-Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving
cs.ROYansong Qu, Zilin Huang, Zihao Sheng, Jiancong Chen
Autonomous driving policy learning with reinforcement learning (RL) is fundamentally limited by low sample efficiency, weak generalization, and a dependence on unsafe online trial-and-error interactions. Although safe RL introduces explicit constraints or costs, existing methods often fail to capture the semantic meaning of safety in real driving scenes, lea
Zijia Lu, A S M Iftekhar, Gaurav Mittal, Tianjian Meng
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of proc
Carles Broto, Ran Levi, Bob Oliver
We compare four different types of realizability for saturated fusion systems over discrete $p$-toral groups. For example, when $G$ is a locally finite group all of whose $p$-subgroups are artinian (hence discrete $p$-toral), we show that it has ``weakly Sylow'' $p$-subgroups and give explicit constructions of saturated fusion systems and associated
Julie Rousseau, Philipp Heer, Kristina Orehounig, Gabriela Hug
Loads represent a promising flexibility source to support the integration of renewable energy sources, as they may shift their energy consumption over time. By computing the aggregated flexibility of power and energy-constrained loads, aggregators can communicate the group's flexibility without sharing individual private information. However, this computatio
PCMamba: Physics-Informed Cross-Modal State Space Model for Dual-Camera Compressive Hyperspectral Imaging
eess.IVGe Meng, Zhongnan Cai, Jingyan Tu, Yingying Wang
Panchromatic (PAN) -assisted Dual-Camera Compressive Hyperspectral Imaging (DCCHI) is a key technology in snapshot hyperspectral imaging. Existing research primarily focuses on exploring spectral information from 2D compressive measurements and spatial information from PAN images in an explicit manner, leading to a bottleneck in HSI reconstruction. Various p
Feng Liu, Bingyu Nan, Xuezhong Qian, Xiaolan Fu
When emotions are repressed, an individual's true feelings may be revealed through micro-expressions. Consequently, micro-expressions are regarded as a genuine source of insight into an individual's authentic emotions. However, the transient and highly localised nature of micro-expressions poses a significant challenge to their accurate recognition, with the
Privacy-Aware Cyberterrorism Network Analysis using Graph Neural Networks and Federated Learning
cs.CRAnas Ali, Mubashar Husain, Peter Hans
Cyberterrorism poses a formidable threat to digital infrastructures, with increasing reliance on encrypted, decentralized platforms that obscure threat actor activity. To address the challenge of analyzing such adversarial networks while preserving the privacy of distributed intelligence data, we propose a Privacy-Aware Federated Graph Neural Network (PA-FGN
Maximilian Krause, Nicola Simon, Claudius Klein, Jens Gibmeier
Diffraction-based stress analysis of textured materials depends on understanding their elastic heterogeneity and its influence on microscopic strain distributions, which is generally done by using simplifying assumptions for crystallite interactions to calculate tensorial stress factors or in the case of very strong textures, by considering the material phas
Junbo Zhang, Heinrich Dinkel, Yadong Niu, Chenyu Liu
We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse domains. By encompassing tasks spanning speech, environmental sounds, and music, X-ARES provides two evaluation approaches for evaluating audio representations: linear fine-tuning
Huanyu Liu, Ge Li, Jia Li, Hao Zhu
How to design reinforcement learning (RL) tasks that effectively unleash the reasoning capability of large language models (LLMs) remains an open question. Existing RL tasks (e.g., math, programming, and constructing reasoning tasks) suffer from three key limitations: (1) Scalability. They rely heavily on human annotation or expensive LLM synthesis to genera
Weiyang Guo, Jing Li, Wenya Wang, YU LI
The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the \textbf{M}ulti-\textbf{T}urn \textbf{S}afety \textbf{A}li
Hongru Song, Yu-an Liu, Ruqing Zhang, Jiafeng Guo
Retrieval-augmented generation (RAG) systems can effectively mitigate the hallucination problem of large language models (LLMs),but they also possess inherent vulnerabilities. Identifying these weaknesses before the large-scale real-world deployment of RAG systems is of great importance, as it lays the foundation for building more secure and robust RAG syste
Guoqiang Chen, Huiqi Sun, Daguang Liu, Zhiqi Wang
Binary analysis plays a pivotal role in security domains such as malware detection and vulnerability discovery, yet it remains labor-intensive and heavily reliant on expert knowledge. General-purpose large language models (LLMs) perform well in programming analysis on source code, while binaryspecific LLMs are underexplored. In this work, we present ReCopilo
A collaborative constrained graph diffusion model for the generation of realistic synthetic molecules
cs.LGManuel Ruiz-Botella, Marta Sales-Pardo, Roger Guimerà
Developing new molecular compounds is crucial to address pressing challenges, from health to environmental sustainability. However, exploring the molecular space to discover new molecules is difficult due to the vastness of the space. Here we introduce CoCoGraph, a collaborative and constrained graph diffusion model capable of generating molecules that are g
Single-shot 3D characterization the spatiotemporal optical vortex via a spatiotemporal wavefront sensor (STWFS)
physics.opticsXiuyu Yao, Ping Zhu, Youjian Yi, Zezhao Gong
The advent of spatiotemporal wave packets (STWPs), represented by spatiotemporal optical vortices (STOVs), has paved the way for the exploration in optics and photonics. To date, despite considerable efforts, a comprehensive and efficient practical means to characterizing wave packets with such complex structures is still lacking. In this study, we introduce
Huishuai Zhang, Bohan Wang, Luoxin Chen
We introduce AdamS, a simple yet effective alternative to Adam for large language model (LLM) pretraining and post-training. By leveraging a novel denominator, i.e., the root of weighted sum of squares of the momentum and the current gradient, AdamS eliminates the need for second-moment estimates. Hence, AdamS is efficient, matching the memory and compute fo
Neuromorphic-based metaheuristics: A new generation of low power, low latency and small footprint optimization algorithms
cs.NEEl-ghazali Talbi
Neuromorphic computing (NC) introduces a novel algorithmic paradigm representing a major shift from traditional digital computing of Von Neumann architectures. NC emulates or simulates the neural dynamics of brains in the form of Spiking Neural Networks (SNNs). Much of the research in NC has concentrated on machine learning applications and neuroscience simu
Yusuke Makita, Keisuke Izumi, Daisuke Yoshida, Keiya Uemichi
We analytically construct static regular solutions describing wormholes that connect multiple asymptotic regions, supported by a phantom scalar field. The solutions are static and axially symmetric, and are constructed using the gravitational soliton formalism, in which the equations of motion reduce to the Laplace equations on a two-dimensional sheet. Howev
Estelle Chigot, Dennis G. Wilson, Meriem Ghrib, Thomas Oberlin
Semantic segmentation models trained on synthetic data often perform poorly on real-world images due to domain gaps, particularly in adverse conditions where labeled data is scarce. Yet, recent foundation models enable to generate realistic images without any training. This paper proposes to leverage such diffusion models to improve the performance of vision
Asit karan, Anil Kumar, Monika Sinha, Ritam Mallick
In this study, we investigate the impact of dark matter on the structure and deformation of magnetars. We assume a perturbative approach for the magnetic field deformation and that the dark matter only interacts gravitationally with hadronic matter. Assuming that dark matter is significantly softer than hadronic matter, we find that the magnetic field can af
Gur Keinan, Omer Ben-Porat
We introduce a game-theoretic framework examining strategic interactions between a platform and its content creators in the presence of AI-generated content. Our model's main novelty is in capturing creators' dual strategic decisions: The investment in content quality and their (possible) consent to share their content with the platform's GenAI, both of whic
Lina Gerlach, Tobias Winkler, Erika Ábrahám, Borzoo Bonakdarpour
Markov decision processes model systems subject to nondeterministic and probabilistic uncertainty. A plethora of verification techniques addresses variations of reachability properties, such as: Is there a scheduler resolving the nondeterminism such that the probability to reach an error state is above a threshold? We consider an understudied extension that
Statistical properties of non-linear observables of fractal Gaussian fields with a focus on spatial-averaging observables and on composite operators
cond-mat.stat-mechCecile Monthus
The statistical properties of non-linear observables of the fractal Gaussian field $\phi(\vec x)$ of negative Hurst exponent $H<0$ in dimension $d$ are revisited with a focus on spatial-averaging observables and on the properties of the finite parts $\phi_n(\vec x)$ of the ill-defined composite operators $\phi^n(\vec x) $. For the special case $n=2$ of quadr
Yu Xin, Zu-dong Zhao, Suo Tang
We investigate the nonlinear Compton photon source for upcoming laser-particle experiments in the collision scenario of high-energy electron beams and relativistic laser pulses. The stronger laser field could not only improve the scattering probability but also induce broader photon beam divergence. To maximize the photon flux in a realistic narrow angular a
Classical solutions to a mixed-type PDE with a Keldysh-type degeneracy and accelerating transonic solutions to the Euler-Poisson system
math.APMyoungjean Bae, Ben Duan, Chunjing Xie
In this paper, we first prove the existence of classical solutions to a class of Keldysh-type equations. Next, we apply this existence result to prove the structural stability of one-dimensional smooth transonic solutions to the steady Euler-Poisson system. Most importantly, the solutions constructed in this paper are classical solutions to the Euler-Poisson
Admission Control of Quasi-Reversible Queueing Systems: Optimization and Reinforcement Learning
cs.LGCéline Comte, Pascal Moyal
In this paper, we introduce a versatile scheme for optimizing the arrival rates of quasi-reversible queueing systems. We first propose an alternative definition of quasi-reversibility that encompasses reversibility and highlights the importance of the definition of customer classes. Then we introduce balanced arrival control policies, which generalize the no
Mudassir Ibrahim Awan, Seokhee Jeon
Accurate prediction of perceptual attributes of haptic textures is essential for advancing VR and AR applications and enhancing robotic interaction with physical surfaces. This paper presents a deep learning-based multi-modal framework, incorporating visual and tactile data, to predict perceptual texture ratings by leveraging multi-feature inputs. To achieve
Chenxu Guo, Jiachen Lian, Xuanru Zhou, Jinming Zhang
Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This w
Jingli Li, Yiyan Ma, Bo Ai, Weijie Yuan
With the rapid growth of the low-altitude economy, the demand for cellular-enabled low-altitude wireless networks (LAWN) is rising significantly. The three-dimensional mobility of drones will lead to frequent handovers (HOs) in cellular networks, while traditional reference signal received power (RSRP)-based criteria may fail to capture the dynamic environme
Pierre Achkar, Tim Gollub, Martin Potthast
The exponential growth of scientific publications has made it increasingly difficult for researchers to stay updated and synthesize knowledge effectively. This paper presents XSum, a modular pipeline for multi-document summarization (MDS) in the scientific domain using Retrieval-Augmented Generation (RAG). The pipeline includes two core components: a questio
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
cs.CLTaeyoon Kwon, Dongwook Choi, Hyojun Kim, Sunghwan Kim
LLM-powered embodied agents have shown success on conventional object-rearrangement tasks, but providing personalized assistance that leverages user-specific knowledge from past interactions presents new challenges. We investigate these challenges through the lens of agents' memory utilization along two critical dimensions: object semantics (identifying obje
Javad Mirzaei, Jeebak Mitra, Gwenael Poitau
With increased 5G deployments, network densification is higher than ever to support the exponentially high throughput requirements. However, this has meant a significant increase in energy consumption, leading to higher operational expenditure (OpEx) for network operators creating an acute need for improvements in network energy savings (NES). A key determin
Marian Verhelst, Luca Benini, Naveen Verma
The rapidly growing importance of Machine Learning (ML) applications, coupled with their ever-increasing model size and inference energy footprint, has created a strong need for specialized ML hardware architectures. Numerous ML accelerators have been explored and implemented, primarily to increase task-level throughput per unit area and reduce task-level en
Javad Haghighat, Tolga M. Duman
DNA storage systems face significant challenges, including insertion, deletion, and substitution (IDS) errors. Therefore, designing effective synchronization codes, i.e., codes capable of correcting IDS errors, is essential for DNA storage systems. Marker codes are a favorable choice for this purpose. In this paper, we extend the notion of marker codes by ma
Daniele Avitabile, Francesca Cavallini, Svetlana Dubinkina, Gabriel J. Lord
We study neural field equations, which are prototypical models of large-scale cortical activity, subject to random data. We view this spatially-extended, nonlocal evolution equation as a Cauchy problem on abstract Banach spaces, with randomness in the synaptic kernel, firing rate function, external stimuli, and initial conditions. We determine conditions on
Nima Rasekh
The unstraightening construction due to Lurie establishes an equivalence between presheaves and fibrations, using one prominent model of $(\infty,1)$-categories, namely quasi-categories. In this work we generalize this result by proving that for all $\infty$-cosmoi of $(\infty,1)$-categories in the sense of Riehl and Verity, which includes quasi-categories b
Yaxin Hou, Yuheng Jia
This paper studies the long-tailed semi-supervised learning (LTSSL) with distribution mismatch, where the class distribution of the labeled training data follows a long-tailed distribution and mismatches with that of the unlabeled training data. Most existing methods introduce auxiliary classifiers (experts) to model various unlabeled data distributions and
Yunhui Jang, Jaehyung Kim, Sungsoo Ahn
Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation. However, current LLMs struggle to interpret SMILES, even failing to carr
Fannar Steinn Aðalsteinsson, Björn Borgar Magnússon, Mislav Milicevic, Adam Nirving Davidsson
Code reviews are a critical yet time-consuming aspect of modern software development, increasingly challenged by growing system complexity and the demand for faster delivery. This paper presents a study conducted at WirelessCar Sweden AB, combining an exploratory field study of current code review practices with a field experiment involving two variations of
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification
cs.CVAmirreza Mahbod, Rupert Ecker, Ramona Woitek
Accurate classification of skin lesions from dermatoscopic images is essential for diagnosis and treatment of skin cancer. In this study, we investigate the utility of a dermatology-specific foundation model, PanDerm, in comparison with two Vision Transformer (ViT) architectures (ViT base and Swin Transformer V2 base) for the task of skin lesion classificati
Yitao Yang, Erjian Liu, Bin Jia, Ed Manley
Mobility is a fundamental feature of human life, and through it our interactions with the world and people around us generate complex and consequential social phenomena. Social segregation, one such process, is increasingly acknowledged as a product of one's entire lived experience rather than mere residential location. Increasingly granular sources of data
Lin Li
Using an intangible intensity factor that is orthogonal to the Fama--French factors, we compare the role of intangible investment in predicting stock returns over the periods 1963--1992 and 1993--2022. For 1963--1992, intangible investment is weak in predicting stock returns, but for 1993--2022, the predictive power of intangible investment becomes very stro
Proof-of-Principle Experiment on a Displacement-Noise-Free Neutron Interferometer for Gravitational Wave Detection
physics.ins-detShoki Iwaguchi, Takuhiro Fujiie, Taro Nambu, Masaaki Kitaguchi
The displacement-noise-free interferometer (DFI) is designed to eliminate all displacement-induced noise while retaining sensitivity to gravitational wave (GW) signals. Ground-based DFIs suffer from physical arm-length limitations, resulting in poor sensitivity at frequencies below 1 kHz. To address this, previous research introduced a neutron-based DFI, whi
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
cs.CVRenjie Wei, Songqiang Xu, Qingyu Guo, Meng Li
Visual autoregressive (VAR) modeling has marked a paradigm shift in image generation from next-token prediction to next-scale prediction. VAR predicts a set of tokens at each step from coarse to fine scale, leading to better image quality and faster inference speed compared to existing diffusion models. However, the large parameter size and computation cost