Skip to content

May 2025 arXiv papers — page 89

Showing 8,8018,900 of 24,552 papers

  1. Rembert Daems, Manfred Opper, Guillaume Crevecoeur, Tolga Birdal

    We present a hierarchical, control theory inspired method for variational inference (VI) for neural stochastic differential equations (SDEs). While VI for neural SDEs is a promising avenue for uncertainty-aware reasoning in time-series, it is computationally challenging due to the iterative nature of maximizing the ELBO. In this work, we propose to decompose

  2. Alexander Boeschoten, Giacomo Sorelli, Manuel Gessner, Claude Fabre

    We investigate the problem of estimating simultaneously multiple parameters encoded in the shape of the modes on which the light is expanded. For this, we generalize the mode-encoded parameter estimation theory as introduced in Ref.[1] to a multi-parameter scenario. We derive the general expression for the Quantum Fisher information matrix and establish the

  3. Ranjith Merugu, Mohammad Sameer Suhail, Akshay P Sarashetti, Venkata Bharath Reddy Reddem

    Recent advancements in video restoration have focused on recovering high-quality video frames from low-quality inputs. Compared with static images, the performance of video restoration significantly depends on efficient exploitation of temporal correlations among successive video frames. The numerous techniques make use of temporal information via flow-based

  4. Yasuyuki Hatsuda, Takaki Matsumoto, Kazumi Okuyama

    It is well-known that the partition function of the Jackiw-Teitelboim (JT) gravity is obtained by an integral transformation of volumes of moduli spaces for Riemann surfaces, also known as the Weil-Petersson volumes. This fact enables us to compute the perturbative genus expansion of the partition function by solving a KdV-type non-linear partial differentia

  5. N. D. Zhigadlo, R. Puzniak

    Clarifying the impact of Fe doping on the structural and superconducting properties of MgB2 is crucial, considering that iron is commonly used as a sheath material for the fabrication of metal-clad MgB2 wires and tapes. To date the effects of Fe doping have only been investigated in polycrystalline samples, but the obtained results are controversial. Here, w

  6. Samuel Humeau, Damien Pous

    A tuple (s1,t1,s2,t2) of vertices in a simple undirected graph is 2-linked when there are two vertex-disjoint paths respectively from s1 to t1 and s2 to t2. A graph is 2-linked when all such tuples are 2-linked. We give a new and simple proof of the ``two paths theorem'', a characterisation of edge-maximal graphs which are not 2-linked as webs: particular ne

  7. Martin Goodfellow, Robbie Booth, Andrew Fagan, Alasdair Lambert

    Students often do not fully understand the code they have written. This sometimes does not become evident until later in their education, which can mean it is harder to fix their incorrect knowledge or misunderstandings. In addition, being able to fully understand code is increasingly important in a world where students have access to generative artificial i

  8. Song Jin, Juntian Zhang, Yuhan Liu, Xun Zhang

    Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interactions. While agent-based simulation is promising, existing platforms often lack a mechanism for user actions to dynamically reshape the environment. To bridge this gap, we introduce Re

  9. Luca Nils Philipp, Eva Münzel, Julian Lüttig, Roland Mitrić

    Forming new hybrid quasiparticles by strong light-matter coupling is a promising tool for tailoring photophysics and photochemistry of molecules. Thus, the ultrafast dynamics of polaritons formed upon strong light-matter coupling has been extensively studied by pump-probe spectroscopy. Although it was predicted that the partial photonic character of polarito

  10. Dayanand Mishra

    In this talk, I will present the calculation of LCSR predictions for the $B \to K$ Hadronic Matrix Elements (HME) at low $q^2$ using light meson distribution amplitudes. I will discuss the results obtained.

  11. Ying-Ying Jin, Ye-Qing Sheng, Yi-Ting Wang, Li-Hong Xie

    We present a characterization of paratopological gyrogroups that can be topologically embedded as subgyrogroups into a product of first-countable $T_{i}$ paratopological gyrogroups for $i = 0, 1, 2$. Specifically, we demonstrate that a strongly paratopological gyrogroup $G$ is topologically isomorphic to a subgyrogroup of a topological product of first-count

  12. Jing Bi, Pinxin Liu, Ali Vosoughi, Jiarui Wu

    The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey complex physical actions and spatial relationships. We address this limitation by proposing a language-driven framework that translates procedural text into coherent visual instructions.

  13. Anton Kluge, Andrea Stocco

    Fragile web tests, primarily caused by locator breakages, are a persistent challenge in web development. Hence, researchers have proposed techniques for web-element re-identification in which algorithms utilize a range of element properties to relocate elements on updated versions of websites based on similarity scoring. In this paper, we replicate the origi

  14. Enrico Da Ronche

    In this paper we generalize the notion of logarithmic vector-valued modular form in order to give a general definition of matrix-valued Hilbert modular forms. We prove that they admit unique polynomial Fourier expansions and we build examples in some particular cases.

  15. Xiaoran Yin, Xu Luo, Hao Wu, Lianli Gao

    The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental information available at each step, primarily through visual observations. As a result, current approaches, which typically rely on reactive pol

  16. Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang

    While reinforcement learning (RL) has demonstrated remarkable success in enhancing large language models (LLMs), it has primarily focused on single-turn tasks such as solving math problems. Training effective web agents for multi-turn interactions remains challenging due to the complexity of long-horizon decision-making across dynamic web interfaces. In this

  17. Jian Qin, Pengjie Zhang, Zhu Chen, Liping Fu

    Weak lensing alters galaxy sizes and fluxes, influencing the clustering patterns of galaxies through cosmic magnification. This effect enables the reconstruction of weak lensing convergence $\hat{\kappa}$ maps for DES and DECaLS by linearly combining galaxy overdensities across magnitude bins in the $g$, $r$, and $z$ photometry bands \citep{Qin+,Qin2+}. In t

  18. Soh Takahashi, Masaru Sasaki, Ken Takeda, Masafumi Oizumi

    The learning mechanisms by which humans acquire internal representations of objects are not fully understood. Deep neural networks (DNNs) have emerged as a useful tool for investigating this question, as they have internal representations similar to those of humans as a byproduct of optimizing their objective functions. While previous studies have shown that

  19. Yoichi Aoki, Soichiro Murakami, Ukyo Honda, Akihiko Kato

    In natural language generation for advertising, creating diverse and engaging ad texts is crucial for capturing a broad audience and avoiding advertising fatigue. Regardless of the importance of diversity, the impact of the diversity-enhancing methods in ad text generation -- mainly tested on tasks such as summarization and machine translation -- has not bee

  20. Iker de las Heras, Benjamin Klopsch, Anitha Thillaisundaram

    We establish that finitely generated non-abelian direct products $G$ of free pro-$p$ groups have full Hausdorff spectrum with respect to the lower $p$-series $\mathcal{L}$. This complements similar results with respect to other standard filtration series and a recent theorem showing that the Hausdorff spectrum $\text{hspec}^\mathcal{L}(G)$ of a $p$-adic anal

  21. Ruizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao

    Retrieval-Augmented Generation (RAG) leverages large language models (LLMs) combined with external contexts to enhance the accuracy and reliability of generated responses. However, reliably attributing generated content to specific context segments, context attribution, remains challenging due to the computationally intensive nature of current methods, which

  22. Qin Chen, Yuanyi Ren, Xiaojun Ma, Yuyang Shi

    Predictive analysis is a cornerstone of modern decision-making, with applications in various domains. Large Language Models (LLMs) have emerged as powerful tools in enabling nuanced, knowledge-intensive conversations, thus aiding in complex decision-making tasks. With the burgeoning expectation to harness LLMs for predictive analysis, there is an urgent need

  23. Linlin Sun, Xiaobao Zhu

    Let $(M,g)$ be a compact Riemann surface with unit area. We investigate the mean field equation for equilibrium turbulence: \begin{align} \begin{cases} -\Delta u = \rho_1\left(\frac{h_1e^{u}}{\int_Mh_1e^udv_g}-1\right) - \rho_2\left(\frac{h_2e^{-u}}{\int_Mh_2e^{-u}dv_g}-1\right), \\ \int_Mudv_g=0, \end{cases} \end{align} where $\rho_1=8\pi$ and $\rho_2\in(0,

  24. Nikolay Stanishev, Yuhang Lu, Touradj Ebrahimi

    Pose-invariant face recognition has become a challenging problem for modern AI-based face recognition systems. It aims at matching a profile face captured in the wild with a frontal face registered in a database. Existing methods perform face frontalization via either generative models or learning a pose robust feature representation. In this paper, a new me

  25. Sreetama Sarkar, Yue Che, Alex Gavin, Peter A. Beerel

    Despite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucinations", generating texts misaligned with the visual context. Existing methods aimed at reducing hallucinations through inference time intervention incur a significant increase in latency. To mitigate this, we present SPIN, a t

  26. Guanting Dong, Yifei Chen, Xiaoxi Li, Jiajie Jin

    Recently, large language models (LLMs) have shown remarkable reasoning capabilities via large-scale reinforcement learning (RL). However, leveraging the RL algorithm to empower effective multi-tool collaborative reasoning in LLMs remains an open challenge. In this paper, we introduce Tool-Star, an RL-based framework designed to empower LLMs to autonomously i

  27. Chaeeun Kim, Seungone Kim

    Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in multi-step reasoning and calling search engines at appropriate steps. However, existing retrieval-augmented reasoning approaches rely on separate retrieval models, limiting the LRM's role in retrieval to deciding when to retrieve and how to query. This separation not only increases ha

  28. Muhammad Farid Adilazuarda, Chen Cecilia Liu, Iryna Gurevych, Alham Fikri Aji

    Adapting cultural values in Large Language Models (LLMs) presents significant challenges, particularly due to biases and limited training data. Prior work primarily aligns LLMs with different cultural values using World Values Survey (WVS) data. However, it remains unclear whether this approach effectively captures cultural nuances or produces distinct cultu

  29. Zimao Sheng, Hong'an Yang, Shuxiang Yang, Zirui Yu

    This paper addresses the challenging problem of robust path-following for fixed-wing unmanned aerial vehicles (UAVs) in complex environments with bounded external disturbances and non-smooth predefined paths. Due to the unique aerodynamic characteristics and flight constraints of fixed-wing UAVs, achieving accurate and fast stable path following remains diff

  30. Gaofei Shen, Hosein Mohebbi, Arianna Bisazza, Afra Alishahi

    As the capabilities of large-scale pre-trained models evolve, understanding the determinants of their outputs becomes more important. Feature attribution aims to reveal which parts of the input elements contribute the most to model outputs. In speech processing, the unique characteristics of the input signal make the application of feature attribution method

  31. Xinxin Chen, Yong Han, Yanqi Qiu, Zipeng Wang

    We give a complete solution to the Mandelbrot-Kahane problem for the microcanonical cascade measures by determing their exact Fourier dimensions. We also discuss the Frostman regularity as well as the bi-H\"older continuity of the Dubins-Freedman random homeomorphisms.

  32. Kishan Gupta, Srikanth Korse, Andreas Brendel, Nicola Pia

    In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, engendered bitrate and computational complexity. Most conventional and neural speech codecs operate on wideband (WB) speech signals to achiev

  33. Huazi Pan, Yanjun Zhang, Leo Yu Zhang, Scott Adams

    Manipulation of local training data and local updates, i.e., the poisoning attack, is the main threat arising from the collaborative nature of the federated learning (FL) paradigm. Most existing poisoning attacks aim to manipulate local data/models in a way that causes denial-of-service (DoS) issues. In this paper, we introduce a novel attack method, named F

  34. Yuanhao Huang, Yilong Ren, Jinlei Wang, Lujia Huo

    Autonomous vehicles are typical complex intelligent systems with artificial intelligence at their core. However, perception methods based on deep learning are extremely vulnerable to adversarial samples, resulting in security accidents. How to generate effective adversarial examples in the physical world and evaluate object detection systems is a huge challe

  35. Xiaoqing Zhang, Huabin Zheng, Ang Lv, Yuhan Liu

    Large language models (LLMs) have been observed to suddenly exhibit advanced reasoning abilities during reinforcement learning (RL), resembling an ``aha moment'' triggered by simple outcome-based rewards. While RL has proven effective in eliciting such breakthroughs in tasks involving mathematics, coding, and vision, it faces significant challenges in multi-

  36. Yang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee

    Despite recent progress in large-scale reinforcement learning (RL) for reasoning, the training recipe for building high-performing reasoning models remains elusive. Key implementation details of frontier models, such as DeepSeek-R1, including data curation strategies and RL training recipe, are often omitted. Moreover, recent research indicates distillation

  37. Qian Deng, Le Hui, Jin Xie, Jian Yang

    Bounding box supervision has gained considerable attention in weakly supervised 3D instance segmentation. While this approach alleviates the need for extensive point-level annotations, obtaining accurate bounding boxes in practical applications remains challenging. To this end, we explore the inaccurate bounding box, named sketchy bounding box, which is imit

  38. Peter Maynard, Yulia Cherdantseva, Avi Shaked, Pete Burnap

    Cyber Security Incident Response (IR) Playbooks are used to capture the steps required to recover from a cyber intrusion. Individual IR playbooks should focus on a specific type of incident and be aligned with the architecture of a system under attack. Intrusion modelling focuses on a specific potential cyber intrusion and is used to identify where and what

  39. Koki Nagakura, Tatsuki Fushimi, Ayaka Tsutsui, Yoichi Ochiai

    This paper presents a method for generating dynamic caustic patterns by utilising dual-optimised holographic fields with Phased Array Transducer (PAT). Building on previous research in static caustic optimisation and ultrasonic manipulation, this approach employs computational techniques to dynamically shape fluid surfaces, thereby creating controllable and

  40. Julie Rousseau, Carlo Tajoli, Hanmin Cai, Philipp Heer

    As non-dispatchable renewable power units become prominent in electric power grids, demand-side flexibility appears as a key element of future power systems' operation. Power and energy bounds are intuitive metrics to describe the flexibility of energy-constrained loads. However, to be used in operation, any power consumption trajectory fulfilling the power

  41. Yuxin Kang, Xin Zeng, Wuji Zhang, Chunfang Sun

    The generation of magnon entanglement and squeezing plays a crucial role in quantum information processing. In this study, we propose a scheme based on a chiral cavity-magnon system, which consists of a torus-shaped cavity and two yttrium iron garnet spheres. The magnon mode of each yttrium iron garnet sphere is selectively coupled to one of the two degenera

  42. Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang

    Reinforcement Learning (RL) can mitigate the causal confusion and distribution shift inherent to imitation learning (IL). However, applying RL to end-to-end autonomous driving (E2E-AD) remains an open problem for its training difficulty, and IL is still the mainstream paradigm in both academia and industry. Recently Model-based Reinforcement Learning (MBRL)

  43. Tristan Karch, Jakhongir Saydaliev, Isabella Di Lenardo, Frédéric Kaplan

    Cadastral data reveal key information about the historical organization of cities but are often non-standardized due to diverse formats and human annotations, complicating large-scale analysis. We explore as a case study Venice's urban history during the critical period from 1740 to 1808, capturing the transition following the fall of the ancient Republic an

  44. Mateja Hrast, Georgios M. Koutentakis, Mikhail Maslov, Mikhail Lemeshko

    Helical dichroism (HD) is a proposed method for the resolution of molecular chirality, employing the orbital angular momentum (OAM) of light. Going beyond the conventional assumptions about HD, this work proposes a rigid theoretical framework for the analysis of the HD, based on molecular symmetries and rotational eigenstates. We derive the rotational select

  45. Benjamin Vendeville, Liana Ermakova, Pierre De Loor

    The general public often encounters complex texts but does not have the time or expertise to fully understand them, leading to the spread of misinformation. Automatic Text Simplification (ATS) helps make information more accessible, but its evaluation methods have not kept up with advances in text generation, especially with Large Language Models (LLMs). In

  46. Chia-Hsiang Lin, Jhao-Ting Lin, Po-Ying Chiu, Shih-Ping Chen

    Inland waterbody detection (IWD) is critical for water resources management and agricultural planning. However, the development of high-fidelity IWD mapping technology remains unresolved. We aim to propose a practical solution based on the easily accessible data, i.e., the delay-Doppler map (DDM) provided by NASA's Cyclone Global Navigation Satellite System

  47. Deyu Song, Xiangyin Zhang, Zipei Yu, Kaiyu Qin

    Multi-view Synthetic Aperture Radar (SAR) imaging can effectively enhance the performance of tasks such as automatic target recognition and image information fusion. Unmanned aerial vehicles (UAVs) have the advantages of flexible deployment and cost reduction. A swarm of UAVs equipped with synthetic aperture radar imaging equipment is well suited to meet the

  48. Ming Cheng, Fei Su, Cancan Li, Juan Liu

    This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework to generate initial predictions using single-channel audio. Then, we extend the original S2SND framework to create a new version, Multi-Chan

  49. Ahmed K. Kadhim, Lei Jiao, Rishad Shafik, Ole-Christoffer Granmo

    The increasing complexity of large-scale language models has amplified concerns regarding their interpretability and reusability. While traditional embedding models like Word2Vec and GloVe offer scalability, they lack transparency and often behave as black boxes. Conversely, interpretable models such as the Tsetlin Machine (TM) have shown promise in construc

  50. Kaiyu He, Tong Zhou, Yubo Chen, Delai Qiu

    Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify the cross-lingual ability of LLMs accurately, we propose a Word-Level Cross-Lingual Translation Task. To find how LLMs learn cross-lingual ability, we trace the outputs of LLMs' int

  51. Haoming Huang, Musen Zhang, Jianxin Yang, Zhen Li

    Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interaction (HRI). However, existing gaze estimation methods merely predict either the gaze direction or the Point-of-Gaze (PoG) on the screen, failing to provide sufficient information for a comprehensive six Degree-of-

  52. Marisa Ripoll, Neal Reeves, Anelia Kurteva, Elena Simperl

    Wikidata is a collaborative knowledge graph which provides machine-readable structured data for Wikimedia projects including Wikipedia. Managed by a community of volunteers, it has grown to become the most edited Wikimedia project. However, it features a long-tail of items with limited data and a number of systematic gaps within the available content. In thi

  53. Priyanka Mishra, Nevill Gonzalez Szwacki

    We present a comprehensive first-principles investigation of the structural, electronic, and vibrational properties of four layered boron nitride (BN) polymorphs--AA-stacked ($e$-BN), AA$^\prime$-stacked ($h$-BN), ABC-stacked ($r$-BN), and AB-stacked ($b$-BN). Using density functional theory and density functional perturbation theory with and without van der

  54. Songlin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan

    The attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains such as language. Rotary position encoding (RoPE) has emerged as the de facto standard approach for position encoding and is part of many mod

  55. Shrobana Ghosh, Mark Hannam

    Gravitational wave (GW) astronomy has been hailed as a gateway to discovering unexpected phenomena in the universe. Over the last decade there have been close to one hundred GW observations of compact-binary mergers. While these signals are largely consistent with mergers of binary black holes, binary neutron stars, or black hole-neutron star systems, some e

  56. Zhixun Li, Bin Cao, Rui Jiao, Liang Wang

    Materials are the foundation of modern society, underpinning advancements in energy, electronics, healthcare, transportation, and infrastructure. The ability to discover and design new materials with tailored properties is critical to solving some of the most pressing global challenges. In recent years, the growing availability of high-quality materials data

  57. Yansong Qu, Zilin Huang, Zihao Sheng, Jiancong Chen

    Autonomous driving policy learning with reinforcement learning (RL) is fundamentally limited by low sample efficiency, weak generalization, and a dependence on unsafe online trial-and-error interactions. Although safe RL introduces explicit constraints or costs, existing methods often fail to capture the semantic meaning of safety in real driving scenes, lea

  58. Zijia Lu, A S M Iftekhar, Gaurav Mittal, Tianjian Meng

    Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of proc

  59. Carles Broto, Ran Levi, Bob Oliver

    We compare four different types of realizability for saturated fusion systems over discrete $p$-toral groups. For example, when $G$ is a locally finite group all of whose $p$-subgroups are artinian (hence discrete $p$-toral), we show that it has ``weakly Sylow'' $p$-subgroups and give explicit constructions of saturated fusion systems and associated

  60. Julie Rousseau, Philipp Heer, Kristina Orehounig, Gabriela Hug

    Loads represent a promising flexibility source to support the integration of renewable energy sources, as they may shift their energy consumption over time. By computing the aggregated flexibility of power and energy-constrained loads, aggregators can communicate the group's flexibility without sharing individual private information. However, this computatio

  61. Ge Meng, Zhongnan Cai, Jingyan Tu, Yingying Wang

    Panchromatic (PAN) -assisted Dual-Camera Compressive Hyperspectral Imaging (DCCHI) is a key technology in snapshot hyperspectral imaging. Existing research primarily focuses on exploring spectral information from 2D compressive measurements and spatial information from PAN images in an explicit manner, leading to a bottleneck in HSI reconstruction. Various p

  62. Feng Liu, Bingyu Nan, Xuezhong Qian, Xiaolan Fu

    When emotions are repressed, an individual's true feelings may be revealed through micro-expressions. Consequently, micro-expressions are regarded as a genuine source of insight into an individual's authentic emotions. However, the transient and highly localised nature of micro-expressions poses a significant challenge to their accurate recognition, with the

  63. Anas Ali, Mubashar Husain, Peter Hans

    Cyberterrorism poses a formidable threat to digital infrastructures, with increasing reliance on encrypted, decentralized platforms that obscure threat actor activity. To address the challenge of analyzing such adversarial networks while preserving the privacy of distributed intelligence data, we propose a Privacy-Aware Federated Graph Neural Network (PA-FGN

  64. Maximilian Krause, Nicola Simon, Claudius Klein, Jens Gibmeier

    Diffraction-based stress analysis of textured materials depends on understanding their elastic heterogeneity and its influence on microscopic strain distributions, which is generally done by using simplifying assumptions for crystallite interactions to calculate tensorial stress factors or in the case of very strong textures, by considering the material phas

  65. Junbo Zhang, Heinrich Dinkel, Yadong Niu, Chenyu Liu

    We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse domains. By encompassing tasks spanning speech, environmental sounds, and music, X-ARES provides two evaluation approaches for evaluating audio representations: linear fine-tuning

  66. Huanyu Liu, Ge Li, Jia Li, Hao Zhu

    How to design reinforcement learning (RL) tasks that effectively unleash the reasoning capability of large language models (LLMs) remains an open question. Existing RL tasks (e.g., math, programming, and constructing reasoning tasks) suffer from three key limitations: (1) Scalability. They rely heavily on human annotation or expensive LLM synthesis to genera

  67. Weiyang Guo, Jing Li, Wenya Wang, YU LI

    The proliferation of jailbreak attacks against large language models (LLMs) highlights the need for robust security measures. However, in multi-round dialogues, malicious intentions may be hidden in interactions, leading LLMs to be more prone to produce harmful responses. In this paper, we propose the \textbf{M}ulti-\textbf{T}urn \textbf{S}afety \textbf{A}li

  68. Hongru Song, Yu-an Liu, Ruqing Zhang, Jiafeng Guo

    Retrieval-augmented generation (RAG) systems can effectively mitigate the hallucination problem of large language models (LLMs),but they also possess inherent vulnerabilities. Identifying these weaknesses before the large-scale real-world deployment of RAG systems is of great importance, as it lays the foundation for building more secure and robust RAG syste

  69. Guoqiang Chen, Huiqi Sun, Daguang Liu, Zhiqi Wang

    Binary analysis plays a pivotal role in security domains such as malware detection and vulnerability discovery, yet it remains labor-intensive and heavily reliant on expert knowledge. General-purpose large language models (LLMs) perform well in programming analysis on source code, while binaryspecific LLMs are underexplored. In this work, we present ReCopilo

  70. Manuel Ruiz-Botella, Marta Sales-Pardo, Roger Guimerà

    Developing new molecular compounds is crucial to address pressing challenges, from health to environmental sustainability. However, exploring the molecular space to discover new molecules is difficult due to the vastness of the space. Here we introduce CoCoGraph, a collaborative and constrained graph diffusion model capable of generating molecules that are g

  71. Xiuyu Yao, Ping Zhu, Youjian Yi, Zezhao Gong

    The advent of spatiotemporal wave packets (STWPs), represented by spatiotemporal optical vortices (STOVs), has paved the way for the exploration in optics and photonics. To date, despite considerable efforts, a comprehensive and efficient practical means to characterizing wave packets with such complex structures is still lacking. In this study, we introduce

  72. Huishuai Zhang, Bohan Wang, Luoxin Chen

    We introduce AdamS, a simple yet effective alternative to Adam for large language model (LLM) pretraining and post-training. By leveraging a novel denominator, i.e., the root of weighted sum of squares of the momentum and the current gradient, AdamS eliminates the need for second-moment estimates. Hence, AdamS is efficient, matching the memory and compute fo

  73. El-ghazali Talbi

    Neuromorphic computing (NC) introduces a novel algorithmic paradigm representing a major shift from traditional digital computing of Von Neumann architectures. NC emulates or simulates the neural dynamics of brains in the form of Spiking Neural Networks (SNNs). Much of the research in NC has concentrated on machine learning applications and neuroscience simu

  74. Yusuke Makita, Keisuke Izumi, Daisuke Yoshida, Keiya Uemichi

    We analytically construct static regular solutions describing wormholes that connect multiple asymptotic regions, supported by a phantom scalar field. The solutions are static and axially symmetric, and are constructed using the gravitational soliton formalism, in which the equations of motion reduce to the Laplace equations on a two-dimensional sheet. Howev

  75. Estelle Chigot, Dennis G. Wilson, Meriem Ghrib, Thomas Oberlin

    Semantic segmentation models trained on synthetic data often perform poorly on real-world images due to domain gaps, particularly in adverse conditions where labeled data is scarce. Yet, recent foundation models enable to generate realistic images without any training. This paper proposes to leverage such diffusion models to improve the performance of vision

  76. Asit karan, Anil Kumar, Monika Sinha, Ritam Mallick

    In this study, we investigate the impact of dark matter on the structure and deformation of magnetars. We assume a perturbative approach for the magnetic field deformation and that the dark matter only interacts gravitationally with hadronic matter. Assuming that dark matter is significantly softer than hadronic matter, we find that the magnetic field can af

  77. Gur Keinan, Omer Ben-Porat

    We introduce a game-theoretic framework examining strategic interactions between a platform and its content creators in the presence of AI-generated content. Our model's main novelty is in capturing creators' dual strategic decisions: The investment in content quality and their (possible) consent to share their content with the platform's GenAI, both of whic

  78. Lina Gerlach, Tobias Winkler, Erika Ábrahám, Borzoo Bonakdarpour

    Markov decision processes model systems subject to nondeterministic and probabilistic uncertainty. A plethora of verification techniques addresses variations of reachability properties, such as: Is there a scheduler resolving the nondeterminism such that the probability to reach an error state is above a threshold? We consider an understudied extension that

  79. Cecile Monthus

    The statistical properties of non-linear observables of the fractal Gaussian field $\phi(\vec x)$ of negative Hurst exponent $H<0$ in dimension $d$ are revisited with a focus on spatial-averaging observables and on the properties of the finite parts $\phi_n(\vec x)$ of the ill-defined composite operators $\phi^n(\vec x) $. For the special case $n=2$ of quadr

  80. Yu Xin, Zu-dong Zhao, Suo Tang

    We investigate the nonlinear Compton photon source for upcoming laser-particle experiments in the collision scenario of high-energy electron beams and relativistic laser pulses. The stronger laser field could not only improve the scattering probability but also induce broader photon beam divergence. To maximize the photon flux in a realistic narrow angular a

  81. Myoungjean Bae, Ben Duan, Chunjing Xie

    In this paper, we first prove the existence of classical solutions to a class of Keldysh-type equations. Next, we apply this existence result to prove the structural stability of one-dimensional smooth transonic solutions to the steady Euler-Poisson system. Most importantly, the solutions constructed in this paper are classical solutions to the Euler-Poisson

  82. Céline Comte, Pascal Moyal

    In this paper, we introduce a versatile scheme for optimizing the arrival rates of quasi-reversible queueing systems. We first propose an alternative definition of quasi-reversibility that encompasses reversibility and highlights the importance of the definition of customer classes. Then we introduce balanced arrival control policies, which generalize the no

  83. Mudassir Ibrahim Awan, Seokhee Jeon

    Accurate prediction of perceptual attributes of haptic textures is essential for advancing VR and AR applications and enhancing robotic interaction with physical surfaces. This paper presents a deep learning-based multi-modal framework, incorporating visual and tactile data, to predict perceptual texture ratings by leveraging multi-feature inputs. To achieve

  84. Chenxu Guo, Jiachen Lian, Xuanru Zhou, Jinming Zhang

    Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This w

  85. Jingli Li, Yiyan Ma, Bo Ai, Weijie Yuan

    With the rapid growth of the low-altitude economy, the demand for cellular-enabled low-altitude wireless networks (LAWN) is rising significantly. The three-dimensional mobility of drones will lead to frequent handovers (HOs) in cellular networks, while traditional reference signal received power (RSRP)-based criteria may fail to capture the dynamic environme

  86. Pierre Achkar, Tim Gollub, Martin Potthast

    The exponential growth of scientific publications has made it increasingly difficult for researchers to stay updated and synthesize knowledge effectively. This paper presents XSum, a modular pipeline for multi-document summarization (MDS) in the scientific domain using Retrieval-Augmented Generation (RAG). The pipeline includes two core components: a questio

  87. Taeyoon Kwon, Dongwook Choi, Hyojun Kim, Sunghwan Kim

    LLM-powered embodied agents have shown success on conventional object-rearrangement tasks, but providing personalized assistance that leverages user-specific knowledge from past interactions presents new challenges. We investigate these challenges through the lens of agents' memory utilization along two critical dimensions: object semantics (identifying obje

  88. Javad Mirzaei, Jeebak Mitra, Gwenael Poitau

    With increased 5G deployments, network densification is higher than ever to support the exponentially high throughput requirements. However, this has meant a significant increase in energy consumption, leading to higher operational expenditure (OpEx) for network operators creating an acute need for improvements in network energy savings (NES). A key determin

  89. Marian Verhelst, Luca Benini, Naveen Verma

    The rapidly growing importance of Machine Learning (ML) applications, coupled with their ever-increasing model size and inference energy footprint, has created a strong need for specialized ML hardware architectures. Numerous ML accelerators have been explored and implemented, primarily to increase task-level throughput per unit area and reduce task-level en

  90. Javad Haghighat, Tolga M. Duman

    DNA storage systems face significant challenges, including insertion, deletion, and substitution (IDS) errors. Therefore, designing effective synchronization codes, i.e., codes capable of correcting IDS errors, is essential for DNA storage systems. Marker codes are a favorable choice for this purpose. In this paper, we extend the notion of marker codes by ma

  91. Daniele Avitabile, Francesca Cavallini, Svetlana Dubinkina, Gabriel J. Lord

    We study neural field equations, which are prototypical models of large-scale cortical activity, subject to random data. We view this spatially-extended, nonlocal evolution equation as a Cauchy problem on abstract Banach spaces, with randomness in the synaptic kernel, firing rate function, external stimuli, and initial conditions. We determine conditions on

  92. Nima Rasekh

    The unstraightening construction due to Lurie establishes an equivalence between presheaves and fibrations, using one prominent model of $(\infty,1)$-categories, namely quasi-categories. In this work we generalize this result by proving that for all $\infty$-cosmoi of $(\infty,1)$-categories in the sense of Riehl and Verity, which includes quasi-categories b

  93. Yaxin Hou, Yuheng Jia

    This paper studies the long-tailed semi-supervised learning (LTSSL) with distribution mismatch, where the class distribution of the labeled training data follows a long-tailed distribution and mismatches with that of the unlabeled training data. Most existing methods introduce auxiliary classifiers (experts) to model various unlabeled data distributions and

  94. Yunhui Jang, Jaehyung Kim, Sungsoo Ahn

    Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation. However, current LLMs struggle to interpret SMILES, even failing to carr

  95. Fannar Steinn Aðalsteinsson, Björn Borgar Magnússon, Mislav Milicevic, Adam Nirving Davidsson

    Code reviews are a critical yet time-consuming aspect of modern software development, increasingly challenged by growing system complexity and the demand for faster delivery. This paper presents a study conducted at WirelessCar Sweden AB, combining an exploratory field study of current code review practices with a field experiment involving two variations of

  96. Amirreza Mahbod, Rupert Ecker, Ramona Woitek

    Accurate classification of skin lesions from dermatoscopic images is essential for diagnosis and treatment of skin cancer. In this study, we investigate the utility of a dermatology-specific foundation model, PanDerm, in comparison with two Vision Transformer (ViT) architectures (ViT base and Swin Transformer V2 base) for the task of skin lesion classificati

  97. Yitao Yang, Erjian Liu, Bin Jia, Ed Manley

    Mobility is a fundamental feature of human life, and through it our interactions with the world and people around us generate complex and consequential social phenomena. Social segregation, one such process, is increasingly acknowledged as a product of one's entire lived experience rather than mere residential location. Increasingly granular sources of data

  98. Lin Li

    Using an intangible intensity factor that is orthogonal to the Fama--French factors, we compare the role of intangible investment in predicting stock returns over the periods 1963--1992 and 1993--2022. For 1963--1992, intangible investment is weak in predicting stock returns, but for 1993--2022, the predictive power of intangible investment becomes very stro

  99. Shoki Iwaguchi, Takuhiro Fujiie, Taro Nambu, Masaaki Kitaguchi

    The displacement-noise-free interferometer (DFI) is designed to eliminate all displacement-induced noise while retaining sensitivity to gravitational wave (GW) signals. Ground-based DFIs suffer from physical arm-length limitations, resulting in poor sensitivity at frequencies below 1 kHz. To address this, previous research introduced a neutron-based DFI, whi

  100. Renjie Wei, Songqiang Xu, Qingyu Guo, Meng Li

    Visual autoregressive (VAR) modeling has marked a paradigm shift in image generation from next-token prediction to next-scale prediction. VAR predicts a set of tokens at each step from coarse to fine scale, leading to better image quality and faster inference speed compared to existing diffusion models. However, the large parameter size and computation cost