May 2025 arXiv papers — page 78
Showing 7,701–7,800 of 24,552 papers
Ibuki Terashima, Tetsuo Hyodo
We study the internal structure of exotic hadrons, especially focusing on the relation between the compositeness and physical observables. Defined as the probability of finding hadronic molecular components in the wave function, compositeness serves as a quantitative measure of the internal structure of exotic hadrons. We utilize the coupled-channel potentia
Hexiang Tan, Fei Sun, Sha Liu, Du Su
As large language models (LLMs) often generate plausible but incorrect content, error detection has become increasingly critical to ensure truthfulness. However, existing detection methods often overlook a critical problem we term as self-consistent error, where LLMs repeatedly generate the same incorrect response across multiple stochastic samples. This wor
Soumya Dutta, Avni Jain, Sriram Ganapathy
Given a pair of source and reference speech recordings, speech-to-speech (S2S) emotion style transfer involves the generation of an output speech that mimics the emotion characteristics of the reference while preserving the content and speaker attributes of the source. In this paper, we propose a speech-to-speech zero-shot emotion style transfer framework, t
Shixian Luo, Zezhou Zhu, Yu Yuan, Yuncheng Yang
Geometric spatial reasoning forms the foundation of many applications in artificial intelligence, yet the ability of large language models (LLMs) to operate over geometric spatial information expressed in procedural code remains underexplored. In this paper, we address this gap by formalizing the Program-to-Geometry task, which challenges models to translate
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
cs.LGDeyang Kong, Qi Guo, Xiangyu Xi, Wei Wang
Reinforcement learning exhibits potential in enhancing the reasoning abilities of large language models, yet it is hard to scale for the low sample efficiency during the rollout phase. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of p
A Framework for Spontaneous Brillouin Noise: Unveiling Fundamental Limits in Brillouin Metrology
physics.opticsSimeng Jin, Shuai Yao, Zhisheng Yang, Zixuan Du
Spontaneous Brillouin scattering (SpBS) provides a non-contact tool for probing the mechanical and thermodynamic properties of materials, enabling important applications such as distributed optical fiber sensing and high-resolution Brillouin microscopy. Achieving metrological precision in these systems relies critically on identifying fundamental noise sourc
Chengda Lu, Xiaoyu Fan, Yu Huang, Rongwu Xu
Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the question: Does CoT reasoning really reduce harmfulness from jailbr
Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He
Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenge
Mostafa Jamshidian, Adam Wittek, Saeideh Sekhavat, Farah Alkhatib
This article presents a dataset used in the article "Kinematics of Abdominal Aortic Aneurysms", published in the Journal of Biomechanics. The dataset is publicly available for download from the Zenodo data repository (https://doi.org/10.5281/zenodo.15477710). The dataset includes time-resolved 3D computed tomography angiography (4D-CTA) images of abdominal a
Huanran Chen, Yinpeng Dong, Zeming Wei, Yao Huang
We discover the emergence of \textit{basins} in the loss landscape of large language models. As model scale increases, LLMs become progressively more resilient to random perturbations in the parameter space, giving rise to expansive stability regions where models exhibit nearly identical performance, but outside of which their capabilities collapse. We obser
Chuhao Zhou, Jianfei Yang
Embodied agents operating in smart homes must understand human behavior through diverse sensory inputs and communicate via natural language. While Vision-Language Models (VLMs) have enabled impressive language-grounded perception, their reliance on visual data limits robustness in real-world scenarios with occlusions, poor lighting, or privacy constraints. I
Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport
eess.IVTaoran Zheng, Yan Yang, Xing Li, Xiang Gu
Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i.e., retrospective reconstruction. However, training on simulated pairs commonly leads to
Bridging Electronic Health Records and Clinical Texts: Contrastive Learning for Enhanced Clinical Tasks
cs.CLSara Ketabi, Dhanesh Ramachandram
Conventional machine learning models, particularly tree-based approaches, have demonstrated promising performance across various clinical prediction tasks using electronic health record (EHR) data. Despite their strengths, these models struggle with tasks that require deeper contextual understanding, such as predicting 30-day hospital readmission. This can b
Alessandra Teresa Cignarella, Anastasia Giachanou, Els Lefever
Stereotypes influence social perceptions and can escalate into discrimination and violence. While NLP research has extensively addressed gender bias and hate speech, stereotype detection remains an emerging field with significant societal implications. In this work is presented a survey of existing research, analyzing definitions from psychology, sociology,
Hanze Zhang, Ke Cheng, Rong Chen, Xingda Wei
This paper reveals that locking can significantly degrade the performance of applications on disaggregated memory (DM), sometimes by several orders of magnitude, due to contention on the NICs of memory nodes (MN-NICs). To address this issue, we present DecLock, a locking mechanism for DM that employs decentralized coordination for ownership transfer across c
Ivana Kesić, Carolina Fortuna, Mihael Mohorčič, Blaž Bertalanič
Time series segmentation (TSS) is one of the time series (TS) analysis techniques, that has received considerably less attention compared to other TS related tasks. In recent years, deep learning architectures have been introduced for TSS, however their reliance on sliding windows limits segmentation granularity due to fixed window sizes and strides. To over
Tony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc Mézard
Diffusion models have achieved remarkable success across a wide range of generative tasks. A key challenge is understanding the mechanisms that prevent their memorization of training data and allow generalization. In this work, we investigate the role of the training dynamics in the transition from generalization to memorization. Through extensive experiment
Yuting Huang, Ziquan Fang, Zhihao Zeng, Lu Chen
Spatio-temporal prediction plays a crucial role in intelligent transportation, weather forecasting, and urban planning. While integrating multi-modal data has shown potential for enhancing prediction accuracy, key challenges persist: (i) inadequate fusion of multi-modal information, (ii) confounding factors that obscure causal relations, and (iii) high compu
Hessa Alawwad
This paper aims to address a gap in major Islamic topics by developing an ontology for the Book of Purification in Islam. Many authoritative Islamic texts begin with the Book of Purification, as it is essential for performing prayer (the second pillar of Islam after Shahadah, the profession of faith) and other religious duties such as Umrah and Hajj. The ont
Jonathan Bennion, Shaona Ghosh, Mantek Singh, Nouha Dziri
Various AI safety datasets have been developed to measure LLMs against evolving interpretations of harm. Our evaluation of five recently published open-source safety benchmarks reveals distinct semantic clusters using UMAP dimensionality reduction and kmeans clustering (silhouette score: 0.470). We identify six primary harm categories with varying benchmark
Sharad Duwal, Mir Nafis Sharear Shopnil, Abhishek Tyagi, Adiba Mahbub Proma
Multimodal out-of-context (OOC) misinformation is misinformation that repurposes real images with unrelated or misleading captions. Detecting such misinformation is challenging because it requires resolving the context of the claim before checking for misinformation. Many current methods, including LLMs and LVLMs, do not perform this contextualization step.
Giuseppe Soda, Alessandro Iorio, Leonardo Rizzo
This study brings a network perspective to papal elections by mapping the relational architecture of the College of Cardinals. Using publicly available data sources, such as official Vatican directories and episcopal consecration records, we assemble a multiplex network that captures cardinals' co-membership in various collegial bodies of the Vatican and the
Xiaohan Chen, Weiqin Zou, Lianyi Zhi, Qianshuang Meng
Word embedding (WE) techniques are advanced textual semantic representation models oriented from the natural language processing (NLP) area. Inspired by their effectiveness in facilitating various NLP tasks, more and more researchers attempt to adopt these WE models for their software engineering (SE) tasks, of which semantic representation of software artif
Gael Meigniez, Hiraku Nozawa
We prove that if the leaves of a minimal Lie foliation are locally isometric to a symmetric space of non-compact type without a Poincare disk factor, then the foliation is smoothly conjugate to a homogeneous Lie foliation up to finite covering. This result generalizes and strengthens Zimmer's theorem, which characterizes minimal Lie foliations with leaves is
Mohammad Kasra Habib, Daniel Graziotin, Stefan Wagner
Requirements elicitation and specification remains a labor-intensive, manual process prone to inconsistencies and gaps, presenting a significant challenge in modern software engineering. Emerging studies underscore the potential of employing large language models (LLMs) for automated requirements generation to support requirements elicitation and specificati
Jiahui Gong, Jingtao Ding, Fanjin Meng, Chen Yang
In recent years, foundational models have revolutionized the fields of language and vision, demonstrating remarkable abilities in understanding and generating complex data; however, similar advances in user behavior modeling have been limited, largely due to the complexity of behavioral data and the challenges involved in capturing intricate temporal and con
Smitha Kumar, Michael A. Lones, Manuel Maarek, Hind Zantout
The rapid advancement of Large Language Models (LLMs) has opened new avenues in education. This study examines the use of LLMs in supporting learning in machine learning education; in particular, it focuses on the ability of LLMs to identify common errors of practice (pitfalls) in machine learning code, and their ability to provide feedback that can guide le
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
cs.HCYuheng Lu, Qian Yu, Hongru Wang, Zeming Liu
Graphical User Interface (GUI) agents, which autonomously operate on digital interfaces through natural language instructions, hold transformative potential for accessibility, automation, and user experience. A critical aspect of their functionality is grounding - the ability to map linguistic intents to visual and structural interface elements. However, exi
Yuxin Wang, Qi Liu, Yueyue Feng, Jinyu Xia
Recently, the von Neumann-Jordan type constants C(X) has defined by Takahashi. A new skew generalized constant Cp({\lambda},\mu,X) based on C(X) constant is given in this paper. First, we will obtain some basic properties of this new constant. Moreover, some relations between this new constant and other constants are investigated. Specially, with the Banach-
Geeta Chandra Raju Bethala, Hao Huang, Niraj Pudasaini, Abdullah Mohamed Ali
We present a hierarchical policy-learning framework that enables a legged humanoid to cooperatively carry extended loads with a human partner using only haptic cues for intent inference. At the upper tier, a lightweight behavior-cloning network consumes six-axis force/torque streams from dual wrist-mounted sensors and outputs whole-body planar velocity comma
Guilherme Korol, Antonio Carlos Schneider Beck, Jeronimo Castrillon
Dynamic DNN optimization techniques such as layer-skipping offer increased adaptability and efficiency gains but can lead to i) a larger memory footprint as in decision gates, ii) increased training complexity (e.g., with non-differentiable operations), and iii) less control over performance-quality trade-offs due to its inherent input-dependent execution. T
Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports
cs.CLHayato Aida, Kosuke Takahashi, Takahiro Omi
With recent advancements in Large Language Models (LLMs) and growing interest in retrieval-augmented generation (RAG), the ability to understand table structures has become increasingly important. This is especially critical in financial domains such as securities reports, where highly accurate question answering (QA) over tables is required. However, tables
Early stakeholder engagement for a possible new multipurpose research reactor for Canada
physics.soc-phZ. Yamani, L. Walters, A. Siddiqui, K. Huynh
Canada has a rich history of nuclear technology development. Since the 1940s, nuclear research infrastructure and facilities, such as National Research Universal (NRU) reactor at Canadian Nuclear Laboratories in Chalk River, Ontario, have played a key role for R&D and for building Canadian expertise and competency in nuclear technology. The NRU reactor retir
Rizoi Bakhromzod
The paper reassesses the largely neglected contribution of Avicenna (Ibn Sina, 980-1037) to medieval Tajik-Persian astronomy. Drawing on published primary and secondary sources, it reconstructs the main directions of his scientific activity - the construction of an observatory at Isfahan, the design of a high-precision angular instrument that anticipates the
Carlotta Pacifici, Simone A. Padoan, Jaroslav Mysiak
Climate extremes such as floods, storms, and heatwaves have caused severe economic and human losses across Europe in recent decades. To support the European Union's climate resilience efforts, we propose a statistical framework for short-to-medium-term prediction of tail risks related to extreme economic losses and fatalities. Our approach builds on Extreme
Jingtong Gao, Ling Pan, Yejing Wang, Rui Zhong
Reinforcement Learning (RL) has become a key approach for enhancing the reasoning capabilities of large language models. However, prevalent RL approaches like proximal policy optimization and group relative policy optimization suffer from sparse, outcome-based rewards and weak exploration incentives, limiting their effectiveness. Specifically, sparse rewards
Andrew Fowlie
Sampling from multi-modal distributions and estimating marginal likelihoods, also known as evidences and normalizing constants, are well-known challenges in statistical computation. They can be overcome by nested sampling, which evolves a set of live points through a sequence of distributions upwards in likelihood. We introduce PolyStan -- a nested sampling
Bo Wang, De-Xing Huang, Xiao-Hu Zhou, Mei-Jiang Gui
Synthetic X-ray angiographies generated by modern generative models hold great potential to reduce the use of contrast agents in vascular interventional procedures. However, low-quality synthetic angiographies can significantly increase procedural risk, underscoring the need for reliable image quality assessment (IQA) methods. Existing IQA models, however, f
Haoran He, Jiajun Liang, Xintao Wang, Pengfei Wan
As the marginal cost of scaling computation (data and parameters) during model pre-training continues to increase substantially, test-time scaling (TTS) has emerged as a promising direction for improving generative model performance by allocating additional computation at inference time. While TTS has demonstrated significant success across multiple language
Tzu-Miao Chou
This work investigates the boundary and defect effects on the modular data in SU$(N)_k$ Chern-Simons theories, focusing on how different boundary conditions and symmetry defects modify the fusion rules and braiding statistics of anyons. Using the framework of modular tensor categories (MTCs) and Frobenius algebra objects, explicit expressions for the modifie
Shuhang Xu, Fangwei Zhong
Metaphors are a crucial way for humans to express complex or subtle ideas by comparing one concept to another, often from a different domain. However, many large language models (LLMs) struggle to interpret and apply metaphors in multi-agent language games, hindering their ability to engage in covert communication and semantic evasion, which are crucial for
Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
cs.CLQingyu Lu, Liang Ding, Siyi Cao, Xuebo Liu
Agents powered by large language models (LLMs) have demonstrated strong planning and decision-making capabilities in complex embodied environments. However, such agents often suffer from inefficiencies in multi-turn interactions, frequently trapped in repetitive loops or issuing ineffective commands, leading to redundant computational overhead. Instead of re
Large language model as user daily behavior data generator: balancing population diversity and individual personality
cs.LGHaoxin Li, Jingtao Ding, Jiahui Gong, Yong Li
Predicting human daily behavior is challenging due to the complexity of routine patterns and short-term fluctuations. While data-driven models have improved behavior prediction by leveraging empirical data from various platforms and devices, the reliance on sensitive, large-scale user data raises privacy concerns and limits data availability. Synthetic data
PathoSCOPE: Few-Shot Pathology Detection via Self-Supervised Contrastive Learning and Pathology-Informed Synthetic Embeddings
cs.CVSinchee Chin, Yinuo Ma, Xiaochen Yang, Jing-Hao Xue
Unsupervised pathology detection trains models on non-pathological data to flag deviations as pathologies, offering strong generalizability for identifying novel diseases and avoiding costly annotations. However, building reliable normality models requires vast healthy datasets, as hospitals' data is inherently biased toward symptomatic populations, while pr
Jihan Yao, Yushi Hu, Yujie Yi, Bin Han
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we present MMMG, a comprehensive and human-aligned benchmark for multimodal generation across 4 modality combinations (ima
Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho
Large language models (LLMs) excel at complex reasoning tasks but remain computationally expensive, limiting their practical deployment. To address this, recent works have focused on distilling reasoning capabilities into smaller language models (sLMs) using chain-of-thought (CoT) traces from teacher LLMs. However, this approach struggles in scenarios requir
Tzu-Miao Chou
This work studies anyonic excitations in warped and curved AdS$_3$ backgrounds via Chern-Simons theory. By incorporating geometric deformations such as conical defects, it is shown that curvature modifies the fusion and braiding properties through corrections to modular data in SU($N$)$_k$ models. Analytical models and numerical simulations reveal how these
Till Freihaut, Luca Viano, Volkan Cevher, Matthieu Geist
This paper provides the first expert sample complexity characterization for learning a Nash equilibrium from expert data in Markov Games. We show that a new quantity named the single policy deviation concentrability coefficient is unavoidable in the non-interactive imitation learning setting, and we provide an upper bound for behavioral cloning (BC) featurin
Zixian Guo, Ming Liu, Qilong Wang, Zhilong Ji
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end training to achieve multi-modal understanding in a unified process. Effective alignment needs high-quality pre-training data and a carefully designed training process. Current LVLMs f
Mingzhe Xie, Siqi Yang, Wenhao Ma, Minghui Liu
We presen a novel determination of the down-to-up composition ratio using the forward-backward asymmetry observed in the proton-proton collisions at the LHC. This method offers unique insights into the flavor-specific difference between down and up quarks, which are difficult to isolate in traditional cross-section measurements due to the inherent mixing of
Bálint Takács
In this paper, nonstandard multistep methods are considered. It is shown that under some (sufficient and necessary) conditions, these methods attain the same order as their standard counterparts - to prove this statement, a nonstandard version of Taylor's series is constructed. The preservation of some qualitative properties (boundedness, the linear combinat
Multi-shot readout error benchmark of the nitrogen-vacancy center's electronic qubit
cond-mat.mes-hallPéter Boross, Domonkos Svastits, Győző Egri, András Pályi
The ground-state electronic spin of a negatively charged nitrogen-vacancy center in diamond can be used for room-temperature experiments showing coherent qubit functionality. At room temperature, photoluminescence-based qubit readout has a low single-shot fidelity; however, the populations of the qubit's two basis states can be inferred using multi-shot read
Alessio Devoto, Jary Pomponi, Mattia Merluzzi, Paolo Di Lorenzo
This paper presents an adaptive framework for edge inference based on a dynamically configurable transformer-powered deep joint source channel coding (DJSCC) architecture. Motivated by a practical scenario where a resource constrained edge device engages in goal oriented semantic communication, such as selectively transmitting essential features for object d
Continuing Isaacson's Legacy: A general metric theory perspective on gravitational memory and the non-linearity of gravity
gr-qcJann Zosso
The challenge of defining a physical notion of gravitational waves, together with the associated dynamical degrees of freedom of a gravity theory, is a long-standing problem that famously lead to the discovery the Bondi-Metzner-Sachs (BMS) spacetime symmetry at null infinity and its connection to gravitational memory. Here, we show that the second major cont
Attention-ResUNet and EfficientSASM-UNet: UNet based frameworks for Lung and Nodule segmentation
eess.IVMuhammad Abdullah, Furqan Shaukat
Lung cancer has been one of the major threats across the world with the highest mortalities. Computer-aided detection (CAD) can help in early detection and thus can help increase the survival rate. Accurate lung parenchyma segmentation (to include the juxta-pleural nodules) and lung nodule segmentation, the primary symptom of lung cancer, play a crucial role
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
cs.CLJiawei Kong, Hao Fang, Xiaochen Yang, Kuofeng Gao
Recent studies have widely investigated backdoor attacks on Large Language Models (LLMs) by inserting harmful question-answer (QA) pairs into their training data. However, we revisit existing attacks and identify two critical limitations: (1) directly embedding harmful content into the training data compromises safety alignment, resulting in attack efficacy
Yuxin Wang, Qi Liu, Haoyu Zhou, Jinyu Xia
In this paper, we build upon the TX constant that was introduced by Alonso and Llorens-Fuster in 2008. Through the incorporation of suitable parameters, we have successfully generalized the aforementioned constant into two novel forms of geometric constants, which are denoted as T1({\lambda},\mu,X ) and T2(\k{appa},{\tau},X ). First, we obtained some basic p
Yusheng Zhao, Qixin Zhang, Xiao Luo, Weizhi Zhang
Large language models (LLMs) have been used in many zero-shot learning problems, with their strong generalization ability. Recently, adopting LLMs in text-attributed graphs (TAGs) has drawn increasing attention. However, the adoption of LLMs faces two major challenges: limited information on graph structure and unreliable responses. LLMs struggle with text a
Linbao Li, Yannan Liu, Daojing He, Yu Li
Safety alignment in large language models (LLMs) is increasingly compromised by jailbreak attacks, which can manipulate these models to generate harmful or unintended content. Investigating these attacks is crucial for uncovering model vulnerabilities. However, many existing jailbreak strategies fail to keep pace with the rapid development of defense mechani
On a Relation between Euler characteristics of \MakeLowercase{de} Rham cohomology and Koszul cohomology of graded local cohomology modules
math.ACTony J. Puthenpurakal, Rakesh B. T. Reddy
Let $K$ be a field of characteristic zero. Let $R = K[X_0, X_1,\ldots,X_n]$ be standard graded. Let $A_{n+1}(K)$ be the $(n + 1)^{th}$ Weyl algebra over $K$. Let $I$ be a homogeneous ideal of $R$ and let $M = H^i_I(R)$ for some $i \geq 0$. By a result of Lyubeznik, $M$ is a graded holonomic $A_{n +1}(K)$-module for each $i \geq 0$. Let $\chi^c(\mathbf{\parti
Soumya Dutta, Smruthi Balaji, Varada R, Viveka Salinamakki
Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Challenge which focused on these complexities, we present Abhinaya, a system integrating speech-based, text-based, and speech-text models. Our ap
Majid Moradi, Mostafa Annabestani
A new estimation scheme based on the split-step quantum walk (SSQW) revealed that by just setting a single parameter, SSQW can potentially achieve quantum Crame\'r-Rao bound in multiparameter estimation. This parameter even does not involve the parameterization but the initial state and unlike ordinary Quantum walk (OQW) there is no necessity for an entangle
Giampaolo Liuzzi, Stefano Lucidi
In this work, we are concerned with the worst case complexity analysis of "a posteriori" methods for unconstrained multi-objective optimization problems where objective function values can only be obtained by querying a black box. We present two main algorithms, namely DFMOnew and DFMOlight which are based on a linesearch expansion technique. In particular,
Manuel Valle Torre, Thom van der Velden, Marcus Specht, Catharine Oertel
Generative AI offers potential for educational support, but often lacks pedagogical grounding and awareness of the student's learning context. Furthermore, researching student interactions with these tools within authentic learning environments remains challenging. To address this, we present JELAI, an open-source platform architecture designed to integrate
AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model
astro-ph.IMTijmen de Haan, Yuan-Sen Ting, Tirthankar Ghosal, Tuan Dung Nguyen
General-purpose large language models (LLMs), despite their broad capabilities, often struggle with specialized domain knowledge. This gap hinders their deployment as reliable research agents in demanding fields such as astronomy. Building on our prior work with AstroSage-Llama-3.1-8B, this study introduces AstroSage-Llama-3.1-70B, a 70-billion parameter dom
MinkUNeXt-SI: Improving point cloud-based place recognition including spherical coordinates and LiDAR intensity
cs.LGJudith Vilella-Cantos, Juan José Cabrera, Luis Payá, Mónica Ballesta
In autonomous navigation systems, the solution of the place recognition problem is crucial for their safe functioning. But this is not a trivial solution, since it must be accurate regardless of any changes in the scene, such as seasonal changes and different weather conditions, and it must be generalizable to other environments. This paper presents our meth
Florian Barthel, Wieland Morgenstern, Paul Hinzer, Anna Hilsmann
Recently, 3D GANs based on 3D Gaussian splatting have been proposed for high quality synthesis of human heads. However, existing methods stabilize training and enhance rendering quality from steep viewpoints by conditioning the random latent vector on the current camera position. This compromises 3D consistency, as we observe significant identity changes whe
Zhihao Du, Changfeng Gao, Yuxuan Wang, Fan Yu
In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversi
Laurent Chupin, Thierry Dubois
This article is devoted to questions concerning the existence of solutions for partial differential equation problems modeling granular flows. The models studied take into account the complex threshold rheology of these flows, as well as the dilatance effects. It is the coupling of these two physical phenomena that ensures stability and the existence of diss
Andrew Conrad, Roderick Cochran, Daniel Sanchez-Rosales, Samantha Isaac
Quantum key distribution is a point-to-point communication protocol that leverages quantum mechanics to enable secure information exchange. Commonly, the transmitter and receiver stations are at fixed locations, and the single-photon quantum states are transmitted over fiber or free space. Here, we describe a modular, platform-agnostic, quantum key distribut
Kushal Khatiwada, Jayden Hopper, Joseph Cheatham, Ayan Joshi
The Internet of Things (IoT) and Large Language Models (LLMs) have been two major emerging players in the information technology era. Although there has been significant coverage of their individual capabilities, our literature survey sheds some light on the integration and interaction of LLMs and IoT devices - a mutualistic relationship in which both partie
Tianqi Zheng, Yi Li, Yu Xiang, Qiongyi He
Certifying maximal quantum randomness without assumptions about system dimension remains a pivotal challenge for secure communication and foundational studies. Here, we introduce a generalized framework to directly certify maximal randomness from observed probability distributions across systems with arbitrary user numbers, without relying on the Bell-inequa
Carlos Franzreb, Arnab Das, Tim Polzehl, Sebastian Möller
Speaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech. The achieved privacy is commonly evaluated with a speaker recognition model trained on anonymized speech. Although this represents a strong attack, it is unclear which aspects of speech are exploited to identify the speakers. Our research sets out to unv
Exploring gravastar-like structures with strongly interacting quark matter shell in the framework of $f(Q)$ gravity under conformal symmetry
gr-qcDebadri Bhattacharjee, Pradip Kumar Chattopadhyay
In this work, we investigate gravastar-like structures in static and spherically symmetric space-time within the framework of $f(Q)$ gravity coupled with conformal symmetry. We have modified the conventional gravastar model by introducing a strongly interacting quark matter shell which maintains the apex of causal limit through the EoS, $p=\rho-2B_{g}$, wher
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
eess.IVMasataka Kobayashi, Shintaro Shiba, Quan Kong, Norimasa Kobori
With the growing adoption of autonomous driving, the advancement of sensor technology is crucial for ensuring safety and reliable operation. Sensor fusion techniques that combine multiple sensors such as LiDAR, radar, and cameras have proven effective, but the integration of multiple devices increases both hardware complexity and cost. Therefore, developing
Hainuo Wang, Qiming Hu, Xiaojie Guo
Restoring images degraded by adverse weather remains a significant challenge due to the highly non-uniform and spatially heterogeneous nature of weather-induced artifacts, e.g., fine-grained rain streaks versus widespread haze. Accurately estimating the underlying degradation can intuitively provide restoration models with more targeted and effective guidanc
Ze Zhang, Qian Dong
Accurate indoor node localization is critical for practical Wireless Sensor Network (WSN) applications, as Global Positioning System (GPS) fails to provide reliable Line-of-Sight (LoS) conditions in most indoor environments. Real-world localization scenarios often involve unknown obstacles with unpredictable shapes, sizes, quantities, and layouts. These obst
Predictive posterior sampling from non-stationnary Gaussian process priors via Diffusion models with application to climate data
stat.MLGabriel V Cardoso, Mike Pereira
Bayesian models based on Gaussian processes (GPs) offer a flexible framework to predict spatially distributed variables with uncertainty. But the use of nonstationary priors, often necessary for capturing complex spatial patterns, makes sampling from the predictive posterior distribution (PPD) computationally intractable. In this paper, we propose a two-step
Ownership Verification of DNN Models Using White-Box Adversarial Attacks with Specified Probability Manipulation
cs.LGTeruki Sano, Minoru Kuribayashi, Masao Sakai, Shuji Isobe
In this paper, we propose a novel framework for ownership verification of deep neural network (DNN) models for image classification tasks. It allows verification of model identity by both the rightful owner and third party without presenting the original model. We assume a gray-box scenario where an unauthorized user owns a model that is illegally copied fro
Frédéric Mangolte
This review is an elaboration of a presentation given at the Real algebraic geometry and singularities conference in honor of Wojciech Kucharz's 70th birthday in Krakow in 2022.
Doncey Albin, Miles Mena, Annika Thomas, Harel Biggie
Multi-robot systems (MRSs) are valuable for tasks such as search and rescue due to their ability to coordinate over shared observations. A central challenge in these systems is aligning independently collected perception data across space and time, i.e., multi-robot data association. While recent advances in collaborative SLAM (C-SLAM), map merging, and inte
Multiphysics Bench: Benchmarking and Investigating Scientific Machine Learning for Multiphysics PDEs
cs.LGChangfan Yang, Lichen Bai, Yinpeng Wang, Shufei Zhang
Solving partial differential equations (PDEs) with machine learning has recently attracted great attention, as PDEs are fundamental tools for modeling real-world systems that range from fundamental physical science to advanced engineering disciplines. Most real-world physical systems across various disciplines are actually involved in multiple coupled physic
Peggy Cellier, Mireille Ducassé, Sébastien Ferré, Olivier Ridoux
This chapter illustrates the basic concepts of fault localization using a data mining technique. It utilizes the Trityp program to illustrate the general method. Formal concept analysis and association rule are two well-known methods for symbolic data mining. In their original inception, they both consider data in the form of an object-attribute table. In th
Xueji Fang, Liyuan Ma, Zhiyang Chen, Mingyuan Zhou
Recent advances in text-to-video generation, particularly with autoregressive models, have enabled the synthesis of high-quality videos depicting individual scenes. However, extending these models to generate long, cross-scene videos remains a significant challenge. As the context length grows during autoregressive decoding, computational costs rise sharply,
Lukas Froschauer, Jonatan Langlet, Andreas Kassler
Real-time traffic monitoring is critical for network operators to ensure performance, security, and visibility, especially as encryption becomes the norm. AI and ML have emerged as powerful tools to create deeper insights from network traffic, but collecting the fine-grained features needed at terabit speeds remains a major bottleneck. We introduce Direct Fe
Siqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen
Large language models (LLMs) have shown emerging potential in spatiotemporal reasoning, making them promising candidates for building urban agents that support diverse urban downstream applications. Despite these benefits, existing studies primarily focus on evaluating urban LLM agent on outcome-level metrics (e.g., prediction accuracy, traffic efficiency),
Reasoning Meets Personalization: Unleashing the Potential of Large Reasoning Model for Personalized Generation
cs.CLSichun Luo, Guanzhi Deng, Jian Xu, Zerui Yang
Personalization is a critical task in modern intelligent systems, with applications spanning diverse domains, including interactions with large language models (LLMs). Recent advances in reasoning capabilities have significantly enhanced LLMs, enabling unprecedented performance in tasks such as mathematics and coding. However, their potential for personaliza
Ján Karabáš, Edita Máčajová, Roman Nedela, Martin Škoviera
We study two measures of uncolourability of cubic graphs, their colouring defect and perfect matching index. The colouring defect of a cubic graph $G$ is the smallest number of edges left uncovered by three perfect matchings; the perfect matching index of $G$ is the smallest number of perfect matchings that together cover all edges of $G$. We provide a compl
Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li
Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks that bypass safety alignment. However, there remains a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare jailbreak attacks a
Denisa Qosja, Kilian Barth, Simon Wagner
In radar systems, high resolution in the Doppler dimension is important for detecting slow-moving targets as it allows for more distinct separation between these targets and clutter, or stationary objects. However, achieving sufficient resolution is constrained by hardware capabilities and physical factors, leading to the development of processing techniques
Geometric Interpretations and Applications of the Berger-Ebin and York $L^2$-Orthogonal Decompositions
math.DGSergey Stepanov, Irina Tsyganok
The Berger-Ebin and York $L^2$-orthogonal decompositions of the vector space of symmetric bilinear differential two-forms are fundamental tools in global Riemannian geometry. In this paper, we investigate the structure of Ricci tensors on compact Riemannian manifolds, with a particular focus on compact Ricci almost solitons, utilizing both the Berger-Ebin an
p2-TQA: A Process-based Preference Learning Framework for Self-Improving Table Question Answering Models
cs.CLWei Zhou, Mohsen Mesgar, Heike Adel, Annemarie Friedrich
Table question answering (TQA) focuses on answering questions based on tabular data. Developing TQA systems targets effective interaction with tabular data for tasks such as cell retrieval and data analysis. While recent work has leveraged fine-tuning to improve TQA systems, existing approaches often under-utilize available data and neglect the potential of
Using low-cost sensors to improve NO2 concentration maps derived from physico-chemical models
stat.APEmma Thulliez, Camille Coron
Urban air quality is a major concern today. Concentrations of pollutants, such as nitrogen dioxide, must be monitored to ensure that they do not exceed hazardous thresholds. For this reason, scarse reference stations, which are generally managed by air quality monitoring associations, are located in major cities. Two recent approaches enable fine-scale mappi
Chun-Ju Lai, Daniel K. Nakano, Arik Wilbert
The authors define a Category $\mathcal{O}$ for any quasi-reductive Lie superalgebra $\mathfrak{g}$ with respect to a triangular decomposition. This much needed approach unifies many important constructions in the existing literature in a rigorous fashion. Our Category $\mathcal{O}$ encompasses all highest weight categories for Lie (super)algebras as well as
Élie Bretin, Eliott Kacedan, Laurent Seppecher
This article aims to present a general analysis of a class of inverse problems that consists in recovering the elliptic parameter maps in systems of PDEs, such as the linear elastic system, from the knowledge of some of their solutions. This identification problem is reformulated as a first-order linear system of the form $\nabla\bm{\mu} + \bm{B} \cdot \bm{\
Model Already Knows the Best Noise: Bayesian Active Noise Selection via Attention in Video Diffusion Model
cs.CVKwanyoung Kim, Sanghyun Kim
The choice of initial noise strongly affects quality and prompt alignment in video diffusion; different seeds for the same prompt can yield drastically different results. While recent methods use externally designed priors (e.g., frequency filtering or inter-frame smoothing), they often overlook internal model signals that indicate inherently preferable seed
Shahin Hakemi, Naveed Akhtar, Ghulam Mubashar Hassan, Ajmal Mian
Despite the remarkable performance of generative Diffusion Models (DMs), their internal working is still not well understood, which is potentially problematic. This paper focuses on exploring the important notion of bias-variance tradeoff in diffusion models. Providing a systematic foundation for this exploration, it establishes that at one extreme, the diff
Zhufeng Yao
We prove for a $\Theta-$positive representation from a discrete subgroup $\Gamma\subset \mathsf{PSL}(2,\mathbb{R})$, the critical exponent for any $\alpha\in \Theta$ is not greater than one. When $\Gamma$ is geometrically finite, the equality holds if and only if $\Gamma$ is a lattice.
Shrey Pandit, Ashwin Vinod, Liu Leqi, Ying Ding
Aligning large language models (LLMs) to accurately detect hallucinations remains a significant challenge due to the sophisticated nature of hallucinated text. Recognizing that hallucinated samples typically exhibit higher deceptive quality than traditional negative samples, we use these carefully engineered hallucinations as negative examples in the DPO ali
Novobo: Supporting Teachers' Peer Learning of Instructional Gestures by Teaching a Mentee AI-Agent Together
cs.HCJiaqi Jiang, Kexin Huang, Roberto Martinez-Maldonado, Huan Zeng
Instructional gestures are essential for teaching, as they enhance communication and support student comprehension. However, existing training methods for developing these embodied skills can be time-consuming, isolating, or overly prescriptive. Research suggests that developing these tacit, experiential skills requires teachers' peer learning, where they le