May 2025 arXiv papers — page 47
Showing 4,601–4,700 of 24,552 papers
Xinbo Wu, Abhishek Umrawal, Lav R. Varshney
As large language models (LLMs) grow more capable, concerns about their safe deployment have also grown. Although alignment mechanisms have been introduced to deter misuse, they remain vulnerable to carefully designed adversarial prompts. In this work, we present a scalable attack strategy: intent-hiding adversarial prompting, which conceals malicious intent
Daehyeon Baek, Jieun Choi, Jimyoung Son, Kyungmin Bin
As large language models become increasingly prevalent, memory bandwidth constraints significantly limit inference throughput, motivating post-training quantization (PTQ). In this paper, we propose FireQ, a co-designed PTQ framework and an INT4-FP8 matrix multiplication kernel that accelerates LLM inference across all linear layers. Specifically, FireQ quant
Thomas Hiemstra, David Hasler, Domenico Paone, Fabian Reichert
The satellite mission EAGLE-1 represents an important step towards a future pan-European secure quantum key distribution (QKD) network. The public-private partnership behind the mission consists of a consortium of universities, research institutes, and companies partially funded by ESA, the European Union, and supported by national delegations. This unique c
Observation of charge density wave excitonic order parameter in topological insulator monolayer WTe2
cond-mat.str-elLiam Watson, Joan Ripoll, Zhengjue Tong, Amit Kumar
Strong electron-hole interactions in a semimetal or narrow-gap semiconductor may drive a ground state of condensed excitons. Monolayer WTe2 has been proposed as a host material for such an exciton condensate, but the order parameter - the key signature of a macroscopic quantum-coherent condensate - has not been observed. Here we use Fourier-transform scannin
Hexiong Yang, Mingrui Chen, Huaibo Huang, Junxian Duan
Inspired by the great success of Masked Language Modeling (MLM) in the natural language domain, the paradigm of self-supervised pre-training and fine-tuning has also achieved remarkable progress in the field of DNA sequence modeling. However, previous methods often relied on massive pre-training data or large-scale base models with huge parameters, imposing
Di Yu, Changze Lv, Xin Du, Linshan Jiang
Most edge-cloud collaboration frameworks rely on the substantial computational and storage capabilities of cloud-based artificial neural networks (ANNs). However, this reliance results in significant communication overhead between edge devices and the cloud and high computational energy consumption, especially when applied to resource-constrained edge device
Jingjun Yang, Liangwei Fan, Jinpu Zhang, Xiangkai Lian
The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, asynchronous information from event streams, failing to leverage
Competition between heating and cooling effects in an optomechanical oscillator using a squeezed field
quant-phVinh N. T. Pham, Chu Manh Hoang, Nguyen Duy Vy
Squeezed light is a useful phenomenon that can be exploited to improve the sensitivity of specific classes of detectors based on optomechanical effects. Recently, there has been significant interest in the potential application of a squeezed field in the cooling of an optomechanical oscillator. It has been shown that this field could cool an oscillator below
Jiawei Tang, Yuheng Jia
Label distribution learning (LDL) is an effective method to predict the relative label description degree (a.k.a. label distribution) of a sample. However, the label distribution is not a complete representation of an instance because it overlooks the absolute intensity of each label. Specifically, it's impossible to obtain the total description degree of hi
Piotr T. Grochowski, Radim Filip
Quantum metrology enables sensitivity to approach the limits set by fundamental physical laws. Even a single continuous mode offers enhanced precision, with the improvement scaling with its occupation number. Due to their high information capacity, continuous modes allow for the engineering of quantum non-Gaussian states, which not only improve metrological
Sebastian Schertler, Oliver Lang, Jonas Lindenberger, Stefan Schuster
In many industrial applications, signals with short periodic pulses, caused by repeated steps in the manufacturing process, are present, and their fundamental frequency or period may be of interest. Fundamental frequency estimation is in many cases performed by describing the periodic signal as a multiharmonic signal and employing the corresponding maximum l
Linli Ma, Suzhen Lin, Jianchao Zeng, Zanxia Jin
Image fusion aims to combine complementary information from multiple source images to generate more comprehensive scene representations. Existing methods primarily rely on the stacking and design of network architectures to enhance the fusion performance, often ignoring the impact of dataset scene bias on model training. This oversight leads the model to lea
Peiyuan Zhi, Peiyang Li, Jianqin Yin, Baoxiong Jia
Robotic loco-manipulation tasks often involve contact-rich interactions with the environment, requiring the joint modeling of contact force and robot position. However, recent visuomotor policies often focus solely on learning position or force control, overlooking their co-learning. In this work, we propose the first unified policy for legged robots that jo
Dawei Feng, Di Mei, Huiri Tan, Lei Ren
Large Language Models (LLMs) have shown remarkable proficiency in natural language understanding (NLU), opening doors for innovative applications. We introduce StreamLink - an LLM-driven distributed data system designed to improve the efficiency and accessibility of data engineering tasks. We build StreamLink on top of distributed frameworks such as Apache S
Lanxiang Zheng, Ruidong Mei, Mingxin Wei, Hao Ren
Object search in large-scale, unstructured environments remains a fundamental challenge in robotics, particularly in dynamic or expansive settings such as outdoor autonomous exploration. This task requires robust spatial reasoning and the ability to leverage prior experiences. While Large Language Models (LLMs) offer strong semantic capabilities, their appli
Guangcong Zheng, Jianlong Yuan, Bo Wang, Haoyang Huang
Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because their step-by-step process naturally leads to a serious error accumulation (drift). Also, many existing ways to make long v
Yuli Chen, Bo Cheng, Jiale Han, Yingying Zhang
Pruning has recently been widely adopted to reduce the parameter scale and improve the inference efficiency of Large Language Models (LLMs). Mainstream pruning techniques often rely on uniform layerwise pruning strategies, which can lead to severe performance degradation at high sparsity levels. Recognizing the varying contributions of different layers in LL
AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset
cs.CLSoichiro Murakami, Peinan Zhang, Hidetaka Kamigaito, Hiroya Takamura
Identifying factors that make ad text attractive is essential for advertising success. This study proposes AdParaphrase v2.0, a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with v1.0, this dataset is 2
Yuhao Wang, Ruiyang Ren, Yucheng Wang, Wayne Xin Zhao
Long-form question answering (LFQA) requires open-ended long-form responses that synthesize coherent, factually grounded content from multi-source evidence. This makes reinforcement learning (RL) reward design critical. The reward must be verifiable for faithful grounding and stable optimization. However, many standard rewards assume a unique target with an
Kai Chen, Taihang Zhen, Hewei Wang, Kailai Liu
As large language models (LLMs) are increasingly deployed in healthcare, ensuring their safety, particularly within collaborative multi-agent configurations, is paramount. In this paper we introduce MedSentry, a benchmark comprising 5 000 adversarial medical prompts spanning 25 threat categories with 100 subthemes. Coupled with this dataset, we develop an en
Quantum Dynamics Predicts Coherent Oscillatory Behavior in the Early-times of a Photoisomerization Reaction
physics.chem-phMohammad Aarabi, Emanuele Marsili, Massimo Olivucci, David Lauvergnat
In this work, we study the quantum dynamics of a photoisomerization reaction employing a two-electronic-state three-vibrational-mode model of the 2-cis-penta-2,4-dieniminium cation (cis-PSB3). In particular, we address two main issues: the challenges encountered in properly converging quantum dynamics calculations, even when a reduced-dimensionality molecula
Larger cities, more commuters, more crime? The role of inter-city commuting in the scaling of urban crime
physics.soc-phSimon Puttock, Umberto Barros, Diego Pinheiro, Marcos Oliveira
Cities attract a daily influx of non-resident commuters, reflecting their roles within wider urban networks -- not as isolated places. However, it remains unclear how this interconnectivity shapes the way crime scales with population, given that larger cities tend to receive more commuters and experience more crime. In this work, we investigate how inter-cit
A candidate for True Type-2 AGN without hidden central BLRs Identified by central Tidal Disruption Event
astro-ph.GAGu Ying, Zheng Qi, Cheng Peizheng, Li Xiao
In this manuscript, through applications of TDE (tidal disruption event) expected variability properties, a potential candidate for True type-2 AGN without hidden central broad line regions (=TT2AGN) is reported in the SDSS J233454.07+145712.9 (=SDSS J2334). Through analyzing the 20-years optical light curves of SDSS J2334 from different Sky Survey projects,
Zibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li
Zero-shot object navigation (ZSON) allows robots to find target objects in unfamiliar environments using natural language instructions, without relying on pre-built maps or task-specific training. Recent general-purpose models, such as large language models (LLMs) and vision-language models (VLMs), equip agents with semantic reasoning abilities to estimate t
Hyomin Kim, Yunhui Jang, Sungsoo Ahn
Large language models (LLMs) have large potential for molecular optimization, as they can gather external chemistry tools and enable collaborative interactions to iteratively refine molecular candidates. However, this potential remains underexplored, particularly in the context of structured reasoning, interpretability, and comprehensive tool-grounded molecu
Dang Nguyen, Jiping Li, Jinghao Zheng, Baharan Mirzasoleiman
Synthetically augmenting training datasets with diffusion models has become an effective strategy for improving the generalization of image classifiers. However, existing approaches typically increase dataset size by 10-30x and struggle to ensure generation diversity, leading to substantial computational overhead. In this work, we introduce TADA (TArgeted Di
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer
Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, w
Domain Decomposition Subspace Neural Network Method for Solving Linear and Nonlinear Partial Differential Equations
math.NAZhenxing Fu, Hongliang Liu, Zhiqiang Sheng, Baixue Xing
This paper proposes a domain decomposition subspace neural network method for efficiently solving linear and nonlinear partial differential equations. By combining the principles of domain decomposition and subspace neural networks, the method constructs basis functions using neural networks to approximate PDE solutions. It imposes $C^k$ continuity condition
Convergence of Clipped-SGD for Convex $(L_0,L_1)$-Smooth Optimization with Heavy-Tailed Noise
math.OCSavelii Chezhegov, Aleksandr Beznosikov, Samuel Horváth, Eduard Gorbunov
Gradient clipping is a widely used technique in Machine Learning and Deep Learning (DL), known for its effectiveness in mitigating the impact of heavy-tailed noise, which frequently arises in the training of large language models. Additionally, first-order methods with clipping, such as Clip-SGD, exhibit stronger convergence guarantees than SGD under the $(L
Krishna Singh Rajput, Tejas Anvekar, Chitta Baral, Vivek Gupta
Recent advances in multimodal question answering have primarily focused on combining heterogeneous modalities or fine-tuning multimodal large language models. While these approaches have shown strong performance, they often rely on a single, generalized reasoning strategy, overlooking the unique characteristics of each modality ultimately limiting both accur
Shiqi Yang, Ziyi Huang, Wengran Xiao, Xinyu Shen
This study focuses on the problem of credit default prediction, builds a modeling framework based on machine learning, and conducts comparative experiments on a variety of mainstream classification algorithms. Through preprocessing, feature engineering, and model training of the Home Credit dataset, the performance of multiple models including logistic regre
Yiqi Huang, Travis Davies, Jiahuan Yan, Jiankai Sun
Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their reliance on raw RGB inputs and handcrafted features often leads to overfitting and poor 3D reasoning under varied lighting, occ
Junsik Kim, Jinwook Park, Kangil Kim
In knowledge graph embedding, leveraging relation specific entity transformation has markedly enhanced performance. However, the consistency of embedding differences before and after transformation remains unaddressed, risking the loss of valuable inductive bias inherent in the embeddings. This inconsistency stems from two problems. First, transformation rep
Parametrizing the reconstruction performance of Super-Kamiokande in the Sub-GeV to TeV neutrino energy range
hep-exC. Jesús-Valls
Super-Kamiokande is a paramount detector for studying atmospheric, astrophysical and accelerator neutrino physics. This work extracts and characterizes the neutrino reconstruction performance of Super-Kamiokande using the public data release from its latest atmospheric neutrino analysis. Energy and zenith angle reconstruction performances are derived and mod
Linear-Time Computation of the Frobenius Normal Form for Symmetric Toeplitz Matrices via Graph-Theoretic Decomposition
math.COHojin Chu, Homoon Ryu
We introduce a linear-time algorithm for computing the Frobenius normal form (FNF) of symmetric Toeplitz matrices by utilizing their inherent structural properties through a graph-theoretic approach. Previous results of the authors established that the FNF of a symmetric Toeplitz matrix is explicitly represented as a direct sum of symmetric irreducible Toepl
The Role of AI in Early Detection of Life-Threatening Diseases: A Retinal Imaging Perspective
eess.IVTariq M Khan, Toufique Ahmed Soomro, Imran Razzak
Retinal imaging has emerged as a powerful, non-invasive modality for detecting and quantifying biomarkers of systemic diseases-ranging from diabetes and hypertension to Alzheimer's disease and cardiovascular disorders but current insights remain dispersed across platforms and specialties. Recent technological advances in optical coherence tomography (OCT/OCT
Sungwon Kim, Namkyeong Lee, Yunyoung Doh, Seungmin Shin
Mesh-based 3D static analysis methods have recently emerged as efficient alternatives to traditional computational numerical solvers, significantly reducing computational costs and runtime for various physics-based analyses. However, these methods primarily focus on surface topology and geometry, often overlooking the inherent thickness of real-world 3D obje
Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning
Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce
Bo-Kai Ruan, Zi-Xiang Ni, Bo-Lun Huang, Teng-Fang Hsiao
Diffusion models achieve impressive performance in high-fidelity image generation but often struggle with rare concepts that appear infrequently in the training distribution. Prior work attempts to address this issue by prompt switching, where generation begins with a frequent proxy prompt and later transitions to the original rare prompt. However, such desi
Yurui Lai, Taiyan Zhang, Renchi Yang
Despite plentiful successes achieved by graph representation learning in various domains, the training of graph neural networks (GNNs) still remains tenaciously challenging due to the tremendous computational overhead needed for sizable graphs in practice. Recently, graph data distillation (GDD), which seeks to distill large graphs into compact and informati
Yao Lu, Tengfei Ma, Zeyu Wang, Zhuangzhi Chen
With the rapid development of wireless communications and the growing complexity of digital modulation schemes, traditional manual modulation recognition methods struggle to extract reliable signal features and meet real-time requirements in modern scenarios. Recently, deep learning based Automatic Modulation Recognition (AMR) approaches have greatly improve
Spatial organisation of multiple species of active particles interacting with an interface
cond-mat.softLove Grover, Rajeev Kapri, Abhishek Chaudhuri
We investigate the steady-state organisation of active particles residing on an interface. Particle activity induces interface deformations, while the local shape of the interface guides particle movement. We consider multiple species of particles which can locally pull on the interface or push it. This coupled system exhibits a wide variety of behaviours, i
Yida Zhang, Qiuyan Liu, Hongtao Luo, Yuqi Xia
To address the limited wave domain signal processing capabilities of traditional single-polarized stacked intelligent metasurfaces (SIMs) in holographic multiple-input multiple-output (HMIMO) systems, which stems from limited integration space, this paper proposes a dual-polarized SIM (DPSIM) architecture. By stacking dual-polarized reconfigurable intelligen
Antonio Tudisco, Deborah Volpe, Giovanna Turvani
Effective and accurate diagnosis of diseases such as cancer, diabetes, and heart failure is crucial for timely medical intervention and improving patient survival rates. Machine learning has revolutionized diagnostic methods in recent years by developing classification models that detect diseases based on selected features. However, these classification task
Enrico Sabatini
Let $RQ$ be the path algebra of a Dynkin quiver $Q$ over a commutative noetherian ring $R$. We show that any homotopically smashing t-structure in the derived category of $RQ$ is compactly generated. We also give a complete description of the compactly generated t-structures in terms of poset homomorphisms from the prime spectrum of the ring $\mathrm{Spec}(R
Hemanth Saratchandran, Damien Teney, Simon Lucey
Transformers have reshaped machine learning by utilizing attention mechanisms to capture complex patterns in large datasets, leading to significant improvements in performance. This success has contributed to the belief that "bigger means better", leading to ever-increasing model sizes. This paper challenge this ideology by showing that many existing transfo
Giulia Cavagnari, Giuseppe Savaré, Giacomo Enrico Sodini
We study the convergence of stochastic time-discretization schemes for evolution equations driven by random velocity fields, including examples like stochastic gradient descent and interacting particle systems. Using a unified framework based on Multivalued Probability Vector Fields, we analyze these dynamics at the level of probability measures in the Wasse
Mozib Bin Awal, Prabwal Phukon
We investigate the thermodynamic phase structure of four-dimensional \textit{R}-charged black holes--characterized by four independent $U(1)$ charges--through the lens of Lyapunov exponents associated with unstable circular orbits of both massless and massive particles. Considering three distinct charge configurations (equal, partially unequal, and fully une
Guozheng Dai, Yiyun He, Ke Wang, Yizhe Zhu
We establish sparse Hanson-Wright inequalities for quadratic forms of sparse $\alpha$-sub-exponential random vectors with exponent parameter $\alpha\in(0, 2]$. In the regime $0< \alpha\le 1$ we derive a refined inequality that is optimal in several canonical models. These results extend the classical Hanson-Wright bound to the sparse setting. Illustrative ap
Yuka Yamaguchi
Any three basic hypergeometric series ${}_{2}\phi_{1}$ whose respective parameters $a, b, c$ and a variable $x$ are shifted by integer powers of $q$ are linearly related with coefficients that are rational functions of $a, b, c, q$, and $x$. This relation is called a three-term relation for ${}_{2}\phi_{1}$. In this paper, we prove that the coefficients of t
Antonio Tudisco, Deborah Volpe, Giovanna Turvani
Accurate and reliable diagnosis of diseases is crucial in enabling timely medical treatment and enhancing patient survival rates. In recent years, Machine Learning has revolutionized diagnostic practices by creating classification models capable of identifying diseases. However, these classification problems often suffer from significant class imbalances, wh
Describe Me Something You Do Not Remember - Challenges and Risks of Exposure Design Using Generative Artificial Intelligence for Therapy of Complex Post-traumatic Disorder
cs.HCAnnalisa Degenhard, Stefan Tschöke, Michael Rietzler, Enrico Rukzio
Post-traumatic stress disorder (PTSD) is associated with sudden, uncontrollable, and intense flashbacks of traumatic memories. Trauma exposure psychotherapy has proven effective in reducing the severity of trauma-related symptoms. It involves controlled recall of traumatic memories to train coping mechanisms for flashbacks and enable autobiographical integra
Joon-Seung Choi, Dong-Min Byun, Hyung-Seok Oh, Seong-Whan Lee
Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propos
Juan A. Rodriguez, Haotian Zhang, Abhay Puri, Aarash Feizi
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both globa
Horst Lewitschnig, Marcus Mayrhofer, Peter Filzmoser
Mission profiles cover the conditions that a component, e.g., an electronic component of a vehicle, is exposed to during its lifecycle. Currently, these profiles typically provide descriptive summaries, such as histograms, of single stress parameters like temperature, humidity, or voltage. This is highly aggregated information. New requirements for electric
Reconfigurable Defect States in Non-Hermitian Topolectrical Chains with Gain and Loss
cond-mat.mes-hallS. M. Rafi-Ul-Islam, Zhuo Bin Siu, Md Saddam Hossain Razo, Mansoor B. A. Jalil
We investigate the interplay between the non-Hermitian skin effect (NHSE), parity-time (PT) symmetry, and topological defect states in a finite non-Hermitian Su-Schrieffer-Heeger (SSH) chain. In the conventional NHSE regime, non-reciprocal hopping leads to an asymmetric localization of all eigenstates at one edge of the system, including the bulk and topolog
Anja Himmerlich, Núria Castelló-Mor, Esteban Currás-Rivera, Yana Gurimskaya
Boron-doped silicon detectors used in high radiation environments like the future HL-LHC show a degradation in device performance due to the radiation induced deactivation of the active boron dopant. This effect, known as the so-called Acceptor Removal Effect (ARE), depends on particle type, particle energy and radiation dose and is usually explained by the
Integrating Intermediate Layer Optimization and Projected Gradient Descent for Solving Inverse Problems with Diffusion Models
cs.CVYang Zheng, Wen Li, Zhaoqiang Liu
Inverse problems (IPs) involve reconstructing signals from noisy observations. Recently, diffusion models (DMs) have emerged as a powerful framework for solving IPs, achieving remarkable reconstruction performance. However, existing DM-based methods frequently encounter issues such as heavy computational demands and suboptimal convergence. In this work, buil
Enhancing Wearable Tap Water Audio Detection through Subclass Annotation in the HD-Epic Dataset
cs.HCRobin Burchard, Kristof Van Laerhoven
Wearable human activity recognition has been shown to benefit from the inclusion of acoustic data, as the sounds around a person often contain valuable context. However, due to privacy concerns, it is usually not ethically feasible to record and save microphone data from the device, since the audio could, for instance, also contain private conversations. Rat
AmirEmad Ghassami, James M. Robins, Andrea Rotnitzky
In various statistical settings, the goal is to estimate a function which is restricted by the statistical model only through a conditional moment restriction. Prominent examples include the nonparametric instrumental variable framework for estimating the structural function of the outcome variable, and the proximal causal inference framework for estimating
On the construction of solutions of the Davey--Stewartson I equation using an open Toda chain
nlin.SII. T. Habibullin, A. R. Khakimova
An effective method for constructing explicit solutions to the Davey--Stewartson type integrable equations is discussed based on the use of a dressing chain. The application of the method is exemplified by the equation DS I, for which a new class of explicit solutions is constructed, containing freedom in two arbitrary functions. In this case the generalized
Tamar Bar-On, Ido Efrat
Let $p$ be a prime number. For a field $F$ containing a root of unity of order $p$, let $H^\bullet(F)=H^\bullet(F,\mathbb{F}_p)$ be the mod-$p$ Galois cohomology graded $\mathbb{F}_p$-algebra of $F$. By the Norm Residue Theorem, $H^\bullet(F)$ is a purely quadratic graded-commutative algebra, and is therefore determined by the cup product $\cup\colon H^1(F)\
Daniël Paulusma, Johannes Rauch, Erik Jan van Leeuwen
The NP-complete problems Colouring and k-Colouring $(k\geq 3$) are well studied on $H$-free graphs, i.e., graphs that do not contain some fixed graph $H$ as an induced subgraph. We research to what extent the known polynomial-time algorithms for $H$-free graphs can be generalized if we only know some of the edges of the input graph. We do this by considering
Dalit Ken-Dror Feldman, Daniel Benoliel
Artificial Knowledge (AK) systems are transforming decision-making across critical domains such as healthcare, finance, and criminal justice. However, their growing opacity presents governance challenges that current regulatory approaches, focused predominantly on explainability, fail to address adequately. This article argues for a shift toward validation a
Kazuharu Kidera, Takuma Miyaguchi, Hideyoshi Yanagisawa
We constructed a computational model of the driver's brain for steering tasks using the active inference framework, grounded in the free energy principle - a theory from computational neuroscience. This model enables quantitative estimation of how accurately the brain learns vehicle dynamics and performs appropriate steering, using a measure called variation
Jiaping Xiao, Cheng Wen Tsao, Yuhang Zhang, Mir Feroskhan
Path planning is a critical component in autonomous drone operations, enabling safe and efficient navigation through complex environments. Recent advances in foundation models, particularly large language models (LLMs) and vision-language models (VLMs), have opened new opportunities for enhanced perception and intelligent decision-making in robotics. However
Taïga Gonçalves, Tomo Miyazaki, Shinichiro Omachi
Multi-targeted adversarial attacks aim to mislead classifiers toward specific target classes using a single perturbation generator with a conditional input specifying the desired target class. Existing methods face two key limitations: (1) a single generator supports only a limited number of predefined target classes, and (2) it requires access to the victim
Noy Sternlicht, Tom Hope
A hallmark of human innovation is recombination -- the creation of novel ideas by integrating elements from existing concepts and mechanisms. In this work, we introduce CHIMERA, the first large-scale Knowledge Base (KB) of recombination examples automatically mined from the scientific literature. CHIMERA enables empirical analysis of how scientists recombine
Yaroslav V. Bazaikin, Yury D. Efremenko, Anton S. Galaev
Let $P$ be a pseudogroup of local diffeomorphisms of an $n$-dimensional smooth manifold $M$. Following Losik we consider characteristic classes of the quotient $M/P$ as elements of the de~Rham cohomology of the second order frame bundles over $M/P$ coming from the generators of the Gelfand-Fuchs cohomology. We provide explicit expressions for the classes tha
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
cs.CVZhehan Kan, Yanlin Liu, Kun Yin, Xinghua Jiang
DeepSeek R1 has significantly advanced complex reasoning for large language models (LLMs). While recent methods have attempted to replicate R1's reasoning capabilities in multimodal settings, they face limitations, including inconsistencies between reasoning and final answers, model instability and crashes during long-chain exploration, and low data learning
Jungyoub Cha, Hyunjong Kim, Sungzoon Cho
Speculative decoding is a widely used technique for accelerating inference in large language models (LLMs), but its performance degrades as input length grows, with significant drops even at moderate lengths. Yet, this early degradation has remained largely underexplored. We introduce SpecExtend, a drop-in enhancement that improves speculative decoding on lo
Non-invasive maturity assessment of iPSC-CMs based on optical maturity characteristics using interpretable AI
cs.LGFabian Scheurer, Alexander Hammer, Mario Schubert, Robert-Patrick Steiner
Human induced pluripotent stem cell-derived cardiomyocytes (iPSC-CMs) are an important resource for the identification of new therapeutic targets and cardioprotective drugs. After differentiation iPSC-CMs show an immature, fetal-like phenotype. Cultivation of iPSC-CMs in lipid-supplemented maturation medium (MM) strongly enhances their structural, metabolic
TimePro: Efficient Multivariate Long-term Time Series Forecasting with Variable- and Time-Aware Hyper-state
cs.LGXiaowen Ma, Zhenliang Ni, Shuai Xiao, Xinghao Chen
In long-term time series forecasting, different variables often influence the target variable over distinct time intervals, a challenge known as the multi-delay issue. Traditional models typically process all variables or time points uniformly, which limits their ability to capture complex variable relationships and obtain non-trivial time representations. T
Adaptive Candidate Retrieval with Dynamic Knowledge Graph Construction for Cold-Start Recommendation
cs.IRWooseong Yang, Weizhi Zhang, Yuqing Liu, Yuwei Han
The cold-start problem remains a critical challenge in real-world recommender systems, as new items with limited interaction data or insufficient information are frequently introduced. Despite recent advances leveraging external knowledge such as knowledge graphs (KGs) and large language models (LLMs), recommender systems still face challenges in practical e
Hongjia Liu, Rongzhen Zhao, Haohan Chen, Joni Pajarinen
Learning object-level, structured representations is widely regarded as a key to better generalization in vision and underpins the design of next-generation Pre-trained Vision Models (PVMs). Mainstream Object-Centric Learning (OCL) methods adopt Slot Attention or its variants to iteratively aggregate objects' super-pixels into a fixed set of query feature ve
Zhucong Li, Powei Chang, Jin Xiao, Zhijian Zhou
Although LLM-based agents are proven to master tool orchestration in scientific fields, particularly chemistry, their single-task performance remains limited by underlying tool constraints. To this end, we propose tool amplification, a novel paradigm that enhances the collective capabilities of specialized tools through optimized, dynamic coordination within
Heng Tang, Feng Liu, Xinbo Chen, Jiawei Chen
Recent years have witnessed extensive exploration of Large Language Models (LLMs) on the field of Recommender Systems (RS). There are currently two commonly used strategies to enable LLMs to have recommendation capabilities: 1) The "Guidance-Only" strategy uses in-context learning to exploit and amplify the inherent semantic understanding and item recommenda
Seungheon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Wei-Hsiang Liao
In music production, manipulating audio effects (Fx) parameters through natural language has the potential to reduce technical barriers for non-experts. We present LLM2Fx, a framework leveraging Large Language Models (LLMs) to predict Fx parameters directly from textual descriptions without requiring task-specific training or fine-tuning. Our approach addres
Physics-Informed Neural Network for Cross-Domain Predictive Control of Tapered Amplifier Thermal Stabilization
eess.SYYanpei Shi, Bo Feng, Yuxin Zhong, Haochen Guo
Thermally induced laser noise poses a critical limitation to the sensitivity of quantum sensor arrays employing ultra-stable amplified lasers, primarily stemming from nonlinear gain-temperature coupling effects in tapered amplifiers (TAs). To address this challenge, we present a robust intelligent control strategy that synergistically integrates an encoder-d
Huaian Diao, Kaixin Lu, Ruixiang Tang, Weisheng Zhou
In our earlier work [13], we introduced a novel quasi-Minnaert resonance for three-dimensional elastic wave scattering in the sub-wavelength regime. Therein, we provided a rigorous analysis of the boundary localization and surface resonance phenomena for both the total and scattered waves, achieved through carefully selected incident waves and tailored physi
CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models
cs.CLXiaqiang Tang, Jian Li, Keyu Hu, Du Nan
Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on "factual statements" that rephrase source materials while overlooking "cognitive statements" that involve making inferences from the given context. Consequently, evaluating
Junbin Li, Xi-Ping Zhu
In this paper, we study the instability of naked singularities arising in the Einstein equations coupled with isothermal perfect fluid. We show that the spherically symmetric self-similar naked singularities of this system, are unstable to trapped surface formation, under $C^{1,\alpha}$ perturbations of an external massless scalar field. We viewed this as a
Robust and Explainable Detector of Time Series Anomaly via Augmenting Multiclass Pseudo-Anomalies
cs.LGKohei Obata, Yasuko Matsubara, Yasushi Sakurai
Unsupervised anomaly detection in time series has been a pivotal research area for decades. Current mainstream approaches focus on learning normality, on the assumption that all or most of the samples in the training set are normal. However, anomalies in the training set (i.e., anomaly contamination) can be misleading. Recent studies employ data augmentation
Eric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou
Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framewor
Dislocations in a multi-layered elastic solid with applications to fault and interface identifications
math.APHuaian Diao, Hongyu Liu, Qingle Meng
This paper investigates an elastic dislocation problem within a bounded and multi-layered solid governed by the Lam\'e system. We address the simultaneous reconstruction of the faults, the jumps in displacement and traction fields across the faults, and the interfaces of layers using a single passive boundary measurement. This inverse problem is particularly
The Extreme Universe Observatory on a Super-Pressure Balloon II: Mission, Payload, and Flight
astro-ph.HEJames. H. Adams, Denis Allard, Phillip Alldredge, Luis Anchordoqui
The Extreme Universe Space Observatory on a Super Pressure Balloon 2 (EUSO-SPB2) is a pathfinder mission toward a space-based observatory such as the Probe of Extreme Multi-Messenger Astrophysics (POEMMA). The aim of POEMMA is the observation of Ultra High Energy COsmic Rays (UHECRs) in order to elucidate their nature and origins and to discover $\gtrsim$ 20
Ryota Ushio, Takashi Ishida, Masashi Sugiyama
While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid to the fundamental question: to what extent can we improve our models? This paper provides a means of answering this question in the setting of binary classification, which is practical and theoretically support
Jingze Ding, Zijian Zhou, Xiaodan Shao, Bingli Jiao
Polarforming emerges as a promising technique for manipulating the polarization of electromagnetic (EM) waves by shaping the polarization of an antenna into a desired state. By dynamically adjusting antenna polarization, polarforming enables real-time polarization matching or mismatching with received EM waves, thereby leveraging polarization degrees of free
Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan
Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning-yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward task. In this work, we introduce PARTONOMY, an LMM benchmark designed for pixel-level part grounding. We construct PART
Ground state solutions of a class of (2,q)-Laplacian Schr\"odinger equations with inhomogeneous nonlinearity
math.APYing Huang, Tingjian Luo, Youde Wang
In this paper, we systematically investigate the ground state solutions of a class of (2,q)-Laplacian Schr\"odinger equations with inhomogeneous nonlinearity. By analyzing global and local constrained variational problems, we establish the existence, non-existence, and asymptotic behavior of ground states, addressing the mass-subcritical,mass-critical, and m
Performance of prior event rate ratio method in the presence of differential mortality or dropout
stat.APYin Bun Cheung, Xiangmei Ma
Purpose: Prior event rate ratio (PERR) method was proposed to control for measured or unmeasured confounders in real-world evaluation of effectiveness and safety of medical treatments using electronic medical records data. A widely cited simulation study showed that PERR estimate of treatment effect was biased in the presence of differential morality/dropout
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
eess.ASIshan D. Biyani, Nirmesh J. Shah, Ashishkumar P. Gudmalwar, Pankaj Wasnik
Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. I
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
cs.SDHaiyun Li, Zhiyong Wu, Xiaofeng Xie, Jingran Xie
Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail in zero-shot VC scenarios, where models synthesize audio from an audio prompt without training. To address this, we propose VoiceMark, the f
Yifei Wang, Weimin Bai, Colin Zhang, Debing Zhang
In this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, $f$-distill, etc, inside a theory-driven framework which we name the \textbf{\emph{Uni-Instruct}}. Uni-Instruct is motivated by our proposed diffusion expansion theory of the $f$-divergence family. Then we introduce key theories tha
Zonghao Chen, Toni Karvonen, Heishiro Kanagawa, François-Xavier Briol
Approximation of a target probability distribution using a finite set of points is a problem of fundamental importance in numerical integration. Several authors have proposed to select points by minimising a maximum mean discrepancy (MMD), but the non-convexity of this objective typically precludes global minimisation. Instead, we consider the concept of \em
Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng
Large Multimodal Models (LMMs) have recently demonstrated remarkable visual understanding performance on both vision-language and vision-centric tasks. However, they often fall short in integrating advanced, task-specific capabilities for compositional reasoning, which hinders their progress toward truly competent general vision models. To address this, we p
Lipei Du, Ulrich Heinz
We present a novel multimessenger approach to extract the effective radial flow of the quark-gluon plasma (QGP) by jointly analyzing thermal photon and dilepton spectra in heavy-ion collisions. A key feature of this method is that it circumvents the need for a directly unmeasurable reference -- the photon temperature in the absence of flow -- by establishing
Interactive OT Gym: A Reinforcement Learning-Based Interactive Optical tweezer (OT)-Driven Microrobotics Simulation Platform
cs.ROZongcai Tan, Dandan Zhang
Optical tweezers (OT) offer unparalleled capabilities for micromanipulation with submicron precision in biomedical applications. However, controlling conventional multi-trap OT to achieve cooperative manipulation of multiple complex-shaped microrobots in dynamic environments poses a significant challenge. To address this, we introduce Interactive OT Gym, a r
Yuan Zhang, Wenxuan Xu, Mohamed Darouach, Tyrone Fernando
Target output controllers aim at regulating a system's target outputs by placing poles of a suitable subsystem using partial state feedback, where full state controllability is not required. This paper establishes existence conditions for such controllers using input and partial state data, where the system dynamics are unknown. The approach bypasses traditi
Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen
LLM-based agent systems are emerging as a new software paradigm and have been widely adopted across diverse domains such as medicine, robotics, and programming. However, maintaining these systems requires substantial effort, as they are inevitably prone to bugs and continually evolve to meet changing external requirements. Therefore, automatically resolving