Skip to content

May 2025 arXiv papers — page 47

Showing 4,6014,700 of 24,552 papers

  1. Xinbo Wu, Abhishek Umrawal, Lav R. Varshney

    As large language models (LLMs) grow more capable, concerns about their safe deployment have also grown. Although alignment mechanisms have been introduced to deter misuse, they remain vulnerable to carefully designed adversarial prompts. In this work, we present a scalable attack strategy: intent-hiding adversarial prompting, which conceals malicious intent

  2. Daehyeon Baek, Jieun Choi, Jimyoung Son, Kyungmin Bin

    As large language models become increasingly prevalent, memory bandwidth constraints significantly limit inference throughput, motivating post-training quantization (PTQ). In this paper, we propose FireQ, a co-designed PTQ framework and an INT4-FP8 matrix multiplication kernel that accelerates LLM inference across all linear layers. Specifically, FireQ quant

  3. Thomas Hiemstra, David Hasler, Domenico Paone, Fabian Reichert

    The satellite mission EAGLE-1 represents an important step towards a future pan-European secure quantum key distribution (QKD) network. The public-private partnership behind the mission consists of a consortium of universities, research institutes, and companies partially funded by ESA, the European Union, and supported by national delegations. This unique c

  4. Liam Watson, Joan Ripoll, Zhengjue Tong, Amit Kumar

    Strong electron-hole interactions in a semimetal or narrow-gap semiconductor may drive a ground state of condensed excitons. Monolayer WTe2 has been proposed as a host material for such an exciton condensate, but the order parameter - the key signature of a macroscopic quantum-coherent condensate - has not been observed. Here we use Fourier-transform scannin

  5. Hexiong Yang, Mingrui Chen, Huaibo Huang, Junxian Duan

    Inspired by the great success of Masked Language Modeling (MLM) in the natural language domain, the paradigm of self-supervised pre-training and fine-tuning has also achieved remarkable progress in the field of DNA sequence modeling. However, previous methods often relied on massive pre-training data or large-scale base models with huge parameters, imposing

  6. Di Yu, Changze Lv, Xin Du, Linshan Jiang

    Most edge-cloud collaboration frameworks rely on the substantial computational and storage capabilities of cloud-based artificial neural networks (ANNs). However, this reliance results in significant communication overhead between edge devices and the cloud and high computational energy consumption, especially when applied to resource-constrained edge device

  7. Jingjun Yang, Liangwei Fan, Jinpu Zhang, Xiangkai Lian

    The integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, asynchronous information from event streams, failing to leverage

  8. Vinh N. T. Pham, Chu Manh Hoang, Nguyen Duy Vy

    Squeezed light is a useful phenomenon that can be exploited to improve the sensitivity of specific classes of detectors based on optomechanical effects. Recently, there has been significant interest in the potential application of a squeezed field in the cooling of an optomechanical oscillator. It has been shown that this field could cool an oscillator below

  9. Jiawei Tang, Yuheng Jia

    Label distribution learning (LDL) is an effective method to predict the relative label description degree (a.k.a. label distribution) of a sample. However, the label distribution is not a complete representation of an instance because it overlooks the absolute intensity of each label. Specifically, it's impossible to obtain the total description degree of hi

  10. Piotr T. Grochowski, Radim Filip

    Quantum metrology enables sensitivity to approach the limits set by fundamental physical laws. Even a single continuous mode offers enhanced precision, with the improvement scaling with its occupation number. Due to their high information capacity, continuous modes allow for the engineering of quantum non-Gaussian states, which not only improve metrological

  11. Sebastian Schertler, Oliver Lang, Jonas Lindenberger, Stefan Schuster

    In many industrial applications, signals with short periodic pulses, caused by repeated steps in the manufacturing process, are present, and their fundamental frequency or period may be of interest. Fundamental frequency estimation is in many cases performed by describing the periodic signal as a multiharmonic signal and employing the corresponding maximum l

  12. Linli Ma, Suzhen Lin, Jianchao Zeng, Zanxia Jin

    Image fusion aims to combine complementary information from multiple source images to generate more comprehensive scene representations. Existing methods primarily rely on the stacking and design of network architectures to enhance the fusion performance, often ignoring the impact of dataset scene bias on model training. This oversight leads the model to lea

  13. Peiyuan Zhi, Peiyang Li, Jianqin Yin, Baoxiong Jia

    Robotic loco-manipulation tasks often involve contact-rich interactions with the environment, requiring the joint modeling of contact force and robot position. However, recent visuomotor policies often focus solely on learning position or force control, overlooking their co-learning. In this work, we propose the first unified policy for legged robots that jo

  14. Dawei Feng, Di Mei, Huiri Tan, Lei Ren

    Large Language Models (LLMs) have shown remarkable proficiency in natural language understanding (NLU), opening doors for innovative applications. We introduce StreamLink - an LLM-driven distributed data system designed to improve the efficiency and accessibility of data engineering tasks. We build StreamLink on top of distributed frameworks such as Apache S

  15. Lanxiang Zheng, Ruidong Mei, Mingxin Wei, Hao Ren

    Object search in large-scale, unstructured environments remains a fundamental challenge in robotics, particularly in dynamic or expansive settings such as outdoor autonomous exploration. This task requires robust spatial reasoning and the ability to leverage prior experiences. While Large Language Models (LLMs) offer strong semantic capabilities, their appli

  16. Guangcong Zheng, Jianlong Yuan, Bo Wang, Haoyang Huang

    Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because their step-by-step process naturally leads to a serious error accumulation (drift). Also, many existing ways to make long v

  17. Yuli Chen, Bo Cheng, Jiale Han, Yingying Zhang

    Pruning has recently been widely adopted to reduce the parameter scale and improve the inference efficiency of Large Language Models (LLMs). Mainstream pruning techniques often rely on uniform layerwise pruning strategies, which can lead to severe performance degradation at high sparsity levels. Recognizing the varying contributions of different layers in LL

  18. Soichiro Murakami, Peinan Zhang, Hidetaka Kamigaito, Hiroya Takamura

    Identifying factors that make ad text attractive is essential for advertising success. This study proposes AdParaphrase v2.0, a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with v1.0, this dataset is 2

  19. Yuhao Wang, Ruiyang Ren, Yucheng Wang, Wayne Xin Zhao

    Long-form question answering (LFQA) requires open-ended long-form responses that synthesize coherent, factually grounded content from multi-source evidence. This makes reinforcement learning (RL) reward design critical. The reward must be verifiable for faithful grounding and stable optimization. However, many standard rewards assume a unique target with an

  20. Kai Chen, Taihang Zhen, Hewei Wang, Kailai Liu

    As large language models (LLMs) are increasingly deployed in healthcare, ensuring their safety, particularly within collaborative multi-agent configurations, is paramount. In this paper we introduce MedSentry, a benchmark comprising 5 000 adversarial medical prompts spanning 25 threat categories with 100 subthemes. Coupled with this dataset, we develop an en

  21. Mohammad Aarabi, Emanuele Marsili, Massimo Olivucci, David Lauvergnat

    In this work, we study the quantum dynamics of a photoisomerization reaction employing a two-electronic-state three-vibrational-mode model of the 2-cis-penta-2,4-dieniminium cation (cis-PSB3). In particular, we address two main issues: the challenges encountered in properly converging quantum dynamics calculations, even when a reduced-dimensionality molecula

  22. Simon Puttock, Umberto Barros, Diego Pinheiro, Marcos Oliveira

    Cities attract a daily influx of non-resident commuters, reflecting their roles within wider urban networks -- not as isolated places. However, it remains unclear how this interconnectivity shapes the way crime scales with population, given that larger cities tend to receive more commuters and experience more crime. In this work, we investigate how inter-cit

  23. Gu Ying, Zheng Qi, Cheng Peizheng, Li Xiao

    In this manuscript, through applications of TDE (tidal disruption event) expected variability properties, a potential candidate for True type-2 AGN without hidden central broad line regions (=TT2AGN) is reported in the SDSS J233454.07+145712.9 (=SDSS J2334). Through analyzing the 20-years optical light curves of SDSS J2334 from different Sky Survey projects,

  24. Zibo Zhou, Yue Hu, Lingkai Zhang, Zonglin Li

    Zero-shot object navigation (ZSON) allows robots to find target objects in unfamiliar environments using natural language instructions, without relying on pre-built maps or task-specific training. Recent general-purpose models, such as large language models (LLMs) and vision-language models (VLMs), equip agents with semantic reasoning abilities to estimate t

  25. Hyomin Kim, Yunhui Jang, Sungsoo Ahn

    Large language models (LLMs) have large potential for molecular optimization, as they can gather external chemistry tools and enable collaborative interactions to iteratively refine molecular candidates. However, this potential remains underexplored, particularly in the context of structured reasoning, interpretability, and comprehensive tool-grounded molecu

  26. Dang Nguyen, Jiping Li, Jinghao Zheng, Baharan Mirzasoleiman

    Synthetically augmenting training datasets with diffusion models has become an effective strategy for improving the generalization of image classifiers. However, existing approaches typically increase dataset size by 10-30x and struggle to ensure generation diversity, leading to substantial computational overhead. In this work, we introduce TADA (TArgeted Di

  27. Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer

    Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, w

  28. Zhenxing Fu, Hongliang Liu, Zhiqiang Sheng, Baixue Xing

    This paper proposes a domain decomposition subspace neural network method for efficiently solving linear and nonlinear partial differential equations. By combining the principles of domain decomposition and subspace neural networks, the method constructs basis functions using neural networks to approximate PDE solutions. It imposes $C^k$ continuity condition

  29. Savelii Chezhegov, Aleksandr Beznosikov, Samuel Horváth, Eduard Gorbunov

    Gradient clipping is a widely used technique in Machine Learning and Deep Learning (DL), known for its effectiveness in mitigating the impact of heavy-tailed noise, which frequently arises in the training of large language models. Additionally, first-order methods with clipping, such as Clip-SGD, exhibit stronger convergence guarantees than SGD under the $(L

  30. Krishna Singh Rajput, Tejas Anvekar, Chitta Baral, Vivek Gupta

    Recent advances in multimodal question answering have primarily focused on combining heterogeneous modalities or fine-tuning multimodal large language models. While these approaches have shown strong performance, they often rely on a single, generalized reasoning strategy, overlooking the unique characteristics of each modality ultimately limiting both accur

  31. Shiqi Yang, Ziyi Huang, Wengran Xiao, Xinyu Shen

    This study focuses on the problem of credit default prediction, builds a modeling framework based on machine learning, and conducts comparative experiments on a variety of mainstream classification algorithms. Through preprocessing, feature engineering, and model training of the Home Credit dataset, the performance of multiple models including logistic regre

  32. Yiqi Huang, Travis Davies, Jiahuan Yan, Jiankai Sun

    Achieving generalizable and precise robotic manipulation across diverse environments remains a critical challenge, largely due to limitations in spatial perception. While prior imitation-learning approaches have made progress, their reliance on raw RGB inputs and handcrafted features often leads to overfitting and poor 3D reasoning under varied lighting, occ

  33. Junsik Kim, Jinwook Park, Kangil Kim

    In knowledge graph embedding, leveraging relation specific entity transformation has markedly enhanced performance. However, the consistency of embedding differences before and after transformation remains unaddressed, risking the loss of valuable inductive bias inherent in the embeddings. This inconsistency stems from two problems. First, transformation rep

  34. C. Jesús-Valls

    Super-Kamiokande is a paramount detector for studying atmospheric, astrophysical and accelerator neutrino physics. This work extracts and characterizes the neutrino reconstruction performance of Super-Kamiokande using the public data release from its latest atmospheric neutrino analysis. Energy and zenith angle reconstruction performances are derived and mod

  35. Hojin Chu, Homoon Ryu

    We introduce a linear-time algorithm for computing the Frobenius normal form (FNF) of symmetric Toeplitz matrices by utilizing their inherent structural properties through a graph-theoretic approach. Previous results of the authors established that the FNF of a symmetric Toeplitz matrix is explicitly represented as a direct sum of symmetric irreducible Toepl

  36. Tariq M Khan, Toufique Ahmed Soomro, Imran Razzak

    Retinal imaging has emerged as a powerful, non-invasive modality for detecting and quantifying biomarkers of systemic diseases-ranging from diabetes and hypertension to Alzheimer's disease and cardiovascular disorders but current insights remain dispersed across platforms and specialties. Recent technological advances in optical coherence tomography (OCT/OCT

  37. Sungwon Kim, Namkyeong Lee, Yunyoung Doh, Seungmin Shin

    Mesh-based 3D static analysis methods have recently emerged as efficient alternatives to traditional computational numerical solvers, significantly reducing computational costs and runtime for various physics-based analyses. However, these methods primarily focus on surface topology and geometry, often overlooking the inherent thickness of real-world 3D obje

  38. Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning

    Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce

  39. Bo-Kai Ruan, Zi-Xiang Ni, Bo-Lun Huang, Teng-Fang Hsiao

    Diffusion models achieve impressive performance in high-fidelity image generation but often struggle with rare concepts that appear infrequently in the training distribution. Prior work attempts to address this issue by prompt switching, where generation begins with a frequent proxy prompt and later transitions to the original rare prompt. However, such desi

  40. Yurui Lai, Taiyan Zhang, Renchi Yang

    Despite plentiful successes achieved by graph representation learning in various domains, the training of graph neural networks (GNNs) still remains tenaciously challenging due to the tremendous computational overhead needed for sizable graphs in practice. Recently, graph data distillation (GDD), which seeks to distill large graphs into compact and informati

  41. Yao Lu, Tengfei Ma, Zeyu Wang, Zhuangzhi Chen

    With the rapid development of wireless communications and the growing complexity of digital modulation schemes, traditional manual modulation recognition methods struggle to extract reliable signal features and meet real-time requirements in modern scenarios. Recently, deep learning based Automatic Modulation Recognition (AMR) approaches have greatly improve

  42. Love Grover, Rajeev Kapri, Abhishek Chaudhuri

    We investigate the steady-state organisation of active particles residing on an interface. Particle activity induces interface deformations, while the local shape of the interface guides particle movement. We consider multiple species of particles which can locally pull on the interface or push it. This coupled system exhibits a wide variety of behaviours, i

  43. Yida Zhang, Qiuyan Liu, Hongtao Luo, Yuqi Xia

    To address the limited wave domain signal processing capabilities of traditional single-polarized stacked intelligent metasurfaces (SIMs) in holographic multiple-input multiple-output (HMIMO) systems, which stems from limited integration space, this paper proposes a dual-polarized SIM (DPSIM) architecture. By stacking dual-polarized reconfigurable intelligen

  44. Antonio Tudisco, Deborah Volpe, Giovanna Turvani

    Effective and accurate diagnosis of diseases such as cancer, diabetes, and heart failure is crucial for timely medical intervention and improving patient survival rates. Machine learning has revolutionized diagnostic methods in recent years by developing classification models that detect diseases based on selected features. However, these classification task

  45. Enrico Sabatini

    Let $RQ$ be the path algebra of a Dynkin quiver $Q$ over a commutative noetherian ring $R$. We show that any homotopically smashing t-structure in the derived category of $RQ$ is compactly generated. We also give a complete description of the compactly generated t-structures in terms of poset homomorphisms from the prime spectrum of the ring $\mathrm{Spec}(R

  46. Hemanth Saratchandran, Damien Teney, Simon Lucey

    Transformers have reshaped machine learning by utilizing attention mechanisms to capture complex patterns in large datasets, leading to significant improvements in performance. This success has contributed to the belief that "bigger means better", leading to ever-increasing model sizes. This paper challenge this ideology by showing that many existing transfo

  47. Giulia Cavagnari, Giuseppe Savaré, Giacomo Enrico Sodini

    We study the convergence of stochastic time-discretization schemes for evolution equations driven by random velocity fields, including examples like stochastic gradient descent and interacting particle systems. Using a unified framework based on Multivalued Probability Vector Fields, we analyze these dynamics at the level of probability measures in the Wasse

  48. Mozib Bin Awal, Prabwal Phukon

    We investigate the thermodynamic phase structure of four-dimensional \textit{R}-charged black holes--characterized by four independent $U(1)$ charges--through the lens of Lyapunov exponents associated with unstable circular orbits of both massless and massive particles. Considering three distinct charge configurations (equal, partially unequal, and fully une

  49. Guozheng Dai, Yiyun He, Ke Wang, Yizhe Zhu

    We establish sparse Hanson-Wright inequalities for quadratic forms of sparse $\alpha$-sub-exponential random vectors with exponent parameter $\alpha\in(0, 2]$. In the regime $0< \alpha\le 1$ we derive a refined inequality that is optimal in several canonical models. These results extend the classical Hanson-Wright bound to the sparse setting. Illustrative ap

  50. Yuka Yamaguchi

    Any three basic hypergeometric series ${}_{2}\phi_{1}$ whose respective parameters $a, b, c$ and a variable $x$ are shifted by integer powers of $q$ are linearly related with coefficients that are rational functions of $a, b, c, q$, and $x$. This relation is called a three-term relation for ${}_{2}\phi_{1}$. In this paper, we prove that the coefficients of t

  51. Antonio Tudisco, Deborah Volpe, Giovanna Turvani

    Accurate and reliable diagnosis of diseases is crucial in enabling timely medical treatment and enhancing patient survival rates. In recent years, Machine Learning has revolutionized diagnostic practices by creating classification models capable of identifying diseases. However, these classification problems often suffer from significant class imbalances, wh

  52. Annalisa Degenhard, Stefan Tschöke, Michael Rietzler, Enrico Rukzio

    Post-traumatic stress disorder (PTSD) is associated with sudden, uncontrollable, and intense flashbacks of traumatic memories. Trauma exposure psychotherapy has proven effective in reducing the severity of trauma-related symptoms. It involves controlled recall of traumatic memories to train coping mechanisms for flashbacks and enable autobiographical integra

  53. Joon-Seung Choi, Dong-Min Byun, Hyung-Seok Oh, Seong-Whan Lee

    Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propos

  54. Juan A. Rodriguez, Haotian Zhang, Abhay Puri, Aarash Feizi

    Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both globa

  55. Horst Lewitschnig, Marcus Mayrhofer, Peter Filzmoser

    Mission profiles cover the conditions that a component, e.g., an electronic component of a vehicle, is exposed to during its lifecycle. Currently, these profiles typically provide descriptive summaries, such as histograms, of single stress parameters like temperature, humidity, or voltage. This is highly aggregated information. New requirements for electric

  56. S. M. Rafi-Ul-Islam, Zhuo Bin Siu, Md Saddam Hossain Razo, Mansoor B. A. Jalil

    We investigate the interplay between the non-Hermitian skin effect (NHSE), parity-time (PT) symmetry, and topological defect states in a finite non-Hermitian Su-Schrieffer-Heeger (SSH) chain. In the conventional NHSE regime, non-reciprocal hopping leads to an asymmetric localization of all eigenstates at one edge of the system, including the bulk and topolog

  57. Anja Himmerlich, Núria Castelló-Mor, Esteban Currás-Rivera, Yana Gurimskaya

    Boron-doped silicon detectors used in high radiation environments like the future HL-LHC show a degradation in device performance due to the radiation induced deactivation of the active boron dopant. This effect, known as the so-called Acceptor Removal Effect (ARE), depends on particle type, particle energy and radiation dose and is usually explained by the

  58. Yang Zheng, Wen Li, Zhaoqiang Liu

    Inverse problems (IPs) involve reconstructing signals from noisy observations. Recently, diffusion models (DMs) have emerged as a powerful framework for solving IPs, achieving remarkable reconstruction performance. However, existing DM-based methods frequently encounter issues such as heavy computational demands and suboptimal convergence. In this work, buil

  59. Robin Burchard, Kristof Van Laerhoven

    Wearable human activity recognition has been shown to benefit from the inclusion of acoustic data, as the sounds around a person often contain valuable context. However, due to privacy concerns, it is usually not ethically feasible to record and save microphone data from the device, since the audio could, for instance, also contain private conversations. Rat

  60. AmirEmad Ghassami, James M. Robins, Andrea Rotnitzky

    In various statistical settings, the goal is to estimate a function which is restricted by the statistical model only through a conditional moment restriction. Prominent examples include the nonparametric instrumental variable framework for estimating the structural function of the outcome variable, and the proximal causal inference framework for estimating

  61. I. T. Habibullin, A. R. Khakimova

    An effective method for constructing explicit solutions to the Davey--Stewartson type integrable equations is discussed based on the use of a dressing chain. The application of the method is exemplified by the equation DS I, for which a new class of explicit solutions is constructed, containing freedom in two arbitrary functions. In this case the generalized

  62. Tamar Bar-On, Ido Efrat

    Let $p$ be a prime number. For a field $F$ containing a root of unity of order $p$, let $H^\bullet(F)=H^\bullet(F,\mathbb{F}_p)$ be the mod-$p$ Galois cohomology graded $\mathbb{F}_p$-algebra of $F$. By the Norm Residue Theorem, $H^\bullet(F)$ is a purely quadratic graded-commutative algebra, and is therefore determined by the cup product $\cup\colon H^1(F)\

  63. Daniël Paulusma, Johannes Rauch, Erik Jan van Leeuwen

    The NP-complete problems Colouring and k-Colouring $(k\geq 3$) are well studied on $H$-free graphs, i.e., graphs that do not contain some fixed graph $H$ as an induced subgraph. We research to what extent the known polynomial-time algorithms for $H$-free graphs can be generalized if we only know some of the edges of the input graph. We do this by considering

  64. Dalit Ken-Dror Feldman, Daniel Benoliel

    Artificial Knowledge (AK) systems are transforming decision-making across critical domains such as healthcare, finance, and criminal justice. However, their growing opacity presents governance challenges that current regulatory approaches, focused predominantly on explainability, fail to address adequately. This article argues for a shift toward validation a

  65. Kazuharu Kidera, Takuma Miyaguchi, Hideyoshi Yanagisawa

    We constructed a computational model of the driver's brain for steering tasks using the active inference framework, grounded in the free energy principle - a theory from computational neuroscience. This model enables quantitative estimation of how accurately the brain learns vehicle dynamics and performs appropriate steering, using a measure called variation

  66. Jiaping Xiao, Cheng Wen Tsao, Yuhang Zhang, Mir Feroskhan

    Path planning is a critical component in autonomous drone operations, enabling safe and efficient navigation through complex environments. Recent advances in foundation models, particularly large language models (LLMs) and vision-language models (VLMs), have opened new opportunities for enhanced perception and intelligent decision-making in robotics. However

  67. Taïga Gonçalves, Tomo Miyazaki, Shinichiro Omachi

    Multi-targeted adversarial attacks aim to mislead classifiers toward specific target classes using a single perturbation generator with a conditional input specifying the desired target class. Existing methods face two key limitations: (1) a single generator supports only a limited number of predefined target classes, and (2) it requires access to the victim

  68. Noy Sternlicht, Tom Hope

    A hallmark of human innovation is recombination -- the creation of novel ideas by integrating elements from existing concepts and mechanisms. In this work, we introduce CHIMERA, the first large-scale Knowledge Base (KB) of recombination examples automatically mined from the scientific literature. CHIMERA enables empirical analysis of how scientists recombine

  69. Yaroslav V. Bazaikin, Yury D. Efremenko, Anton S. Galaev

    Let $P$ be a pseudogroup of local diffeomorphisms of an $n$-dimensional smooth manifold $M$. Following Losik we consider characteristic classes of the quotient $M/P$ as elements of the de~Rham cohomology of the second order frame bundles over $M/P$ coming from the generators of the Gelfand-Fuchs cohomology. We provide explicit expressions for the classes tha

  70. Zhehan Kan, Yanlin Liu, Kun Yin, Xinghua Jiang

    DeepSeek R1 has significantly advanced complex reasoning for large language models (LLMs). While recent methods have attempted to replicate R1's reasoning capabilities in multimodal settings, they face limitations, including inconsistencies between reasoning and final answers, model instability and crashes during long-chain exploration, and low data learning

  71. Jungyoub Cha, Hyunjong Kim, Sungzoon Cho

    Speculative decoding is a widely used technique for accelerating inference in large language models (LLMs), but its performance degrades as input length grows, with significant drops even at moderate lengths. Yet, this early degradation has remained largely underexplored. We introduce SpecExtend, a drop-in enhancement that improves speculative decoding on lo

  72. Fabian Scheurer, Alexander Hammer, Mario Schubert, Robert-Patrick Steiner

    Human induced pluripotent stem cell-derived cardiomyocytes (iPSC-CMs) are an important resource for the identification of new therapeutic targets and cardioprotective drugs. After differentiation iPSC-CMs show an immature, fetal-like phenotype. Cultivation of iPSC-CMs in lipid-supplemented maturation medium (MM) strongly enhances their structural, metabolic

  73. Xiaowen Ma, Zhenliang Ni, Shuai Xiao, Xinghao Chen

    In long-term time series forecasting, different variables often influence the target variable over distinct time intervals, a challenge known as the multi-delay issue. Traditional models typically process all variables or time points uniformly, which limits their ability to capture complex variable relationships and obtain non-trivial time representations. T

  74. Wooseong Yang, Weizhi Zhang, Yuqing Liu, Yuwei Han

    The cold-start problem remains a critical challenge in real-world recommender systems, as new items with limited interaction data or insufficient information are frequently introduced. Despite recent advances leveraging external knowledge such as knowledge graphs (KGs) and large language models (LLMs), recommender systems still face challenges in practical e

  75. Hongjia Liu, Rongzhen Zhao, Haohan Chen, Joni Pajarinen

    Learning object-level, structured representations is widely regarded as a key to better generalization in vision and underpins the design of next-generation Pre-trained Vision Models (PVMs). Mainstream Object-Centric Learning (OCL) methods adopt Slot Attention or its variants to iteratively aggregate objects' super-pixels into a fixed set of query feature ve

  76. Zhucong Li, Powei Chang, Jin Xiao, Zhijian Zhou

    Although LLM-based agents are proven to master tool orchestration in scientific fields, particularly chemistry, their single-task performance remains limited by underlying tool constraints. To this end, we propose tool amplification, a novel paradigm that enhances the collective capabilities of specialized tools through optimized, dynamic coordination within

  77. Heng Tang, Feng Liu, Xinbo Chen, Jiawei Chen

    Recent years have witnessed extensive exploration of Large Language Models (LLMs) on the field of Recommender Systems (RS). There are currently two commonly used strategies to enable LLMs to have recommendation capabilities: 1) The "Guidance-Only" strategy uses in-context learning to exploit and amplify the inherent semantic understanding and item recommenda

  78. Seungheon Doh, Junghyun Koo, Marco A. Martínez-Ramírez, Wei-Hsiang Liao

    In music production, manipulating audio effects (Fx) parameters through natural language has the potential to reduce technical barriers for non-experts. We present LLM2Fx, a framework leveraging Large Language Models (LLMs) to predict Fx parameters directly from textual descriptions without requiring task-specific training or fine-tuning. Our approach addres

  79. Yanpei Shi, Bo Feng, Yuxin Zhong, Haochen Guo

    Thermally induced laser noise poses a critical limitation to the sensitivity of quantum sensor arrays employing ultra-stable amplified lasers, primarily stemming from nonlinear gain-temperature coupling effects in tapered amplifiers (TAs). To address this challenge, we present a robust intelligent control strategy that synergistically integrates an encoder-d

  80. Huaian Diao, Kaixin Lu, Ruixiang Tang, Weisheng Zhou

    In our earlier work [13], we introduced a novel quasi-Minnaert resonance for three-dimensional elastic wave scattering in the sub-wavelength regime. Therein, we provided a rigorous analysis of the boundary localization and surface resonance phenomena for both the total and scattered waves, achieved through carefully selected incident waves and tailored physi

  81. Xiaqiang Tang, Jian Li, Keyu Hu, Du Nan

    Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on "factual statements" that rephrase source materials while overlooking "cognitive statements" that involve making inferences from the given context. Consequently, evaluating

  82. Junbin Li, Xi-Ping Zhu

    In this paper, we study the instability of naked singularities arising in the Einstein equations coupled with isothermal perfect fluid. We show that the spherically symmetric self-similar naked singularities of this system, are unstable to trapped surface formation, under $C^{1,\alpha}$ perturbations of an external massless scalar field. We viewed this as a

  83. Kohei Obata, Yasuko Matsubara, Yasushi Sakurai

    Unsupervised anomaly detection in time series has been a pivotal research area for decades. Current mainstream approaches focus on learning normality, on the assumption that all or most of the samples in the training set are normal. However, anomalies in the training set (i.e., anomaly contamination) can be misleading. Recent studies employ data augmentation

  84. Eric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou

    Composed image retrieval (CIR) is the task of retrieving a target image specified by a query image and a relative text that describes a semantic modification to the query image. Existing methods in CIR struggle to accurately represent the image and the text modification, resulting in subpar performance. To address this limitation, we introduce a CIR framewor

  85. Huaian Diao, Hongyu Liu, Qingle Meng

    This paper investigates an elastic dislocation problem within a bounded and multi-layered solid governed by the Lam\'e system. We address the simultaneous reconstruction of the faults, the jumps in displacement and traction fields across the faults, and the interfaces of layers using a single passive boundary measurement. This inverse problem is particularly

  86. James. H. Adams, Denis Allard, Phillip Alldredge, Luis Anchordoqui

    The Extreme Universe Space Observatory on a Super Pressure Balloon 2 (EUSO-SPB2) is a pathfinder mission toward a space-based observatory such as the Probe of Extreme Multi-Messenger Astrophysics (POEMMA). The aim of POEMMA is the observation of Ultra High Energy COsmic Rays (UHECRs) in order to elucidate their nature and origins and to discover $\gtrsim$ 20

  87. Ryota Ushio, Takashi Ishida, Masashi Sugiyama

    While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid to the fundamental question: to what extent can we improve our models? This paper provides a means of answering this question in the setting of binary classification, which is practical and theoretically support

  88. Jingze Ding, Zijian Zhou, Xiaodan Shao, Bingli Jiao

    Polarforming emerges as a promising technique for manipulating the polarization of electromagnetic (EM) waves by shaping the polarization of an antenna into a desired state. By dynamically adjusting antenna polarization, polarforming enables real-time polarization matching or mismatching with received EM waves, thereby leveraging polarization degrees of free

  89. Ansel Blume, Jeonghwan Kim, Hyeonjeong Ha, Elen Chatikyan

    Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning-yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward task. In this work, we introduce PARTONOMY, an LMM benchmark designed for pixel-level part grounding. We construct PART

  90. Ying Huang, Tingjian Luo, Youde Wang

    In this paper, we systematically investigate the ground state solutions of a class of (2,q)-Laplacian Schr\"odinger equations with inhomogeneous nonlinearity. By analyzing global and local constrained variational problems, we establish the existence, non-existence, and asymptotic behavior of ground states, addressing the mass-subcritical,mass-critical, and m

  91. Yin Bun Cheung, Xiangmei Ma

    Purpose: Prior event rate ratio (PERR) method was proposed to control for measured or unmeasured confounders in real-world evaluation of effectiveness and safety of medical treatments using electronic medical records data. A widely cited simulation study showed that PERR estimate of treatment effect was biased in the presence of differential morality/dropout

  92. Ishan D. Biyani, Nirmesh J. Shah, Ashishkumar P. Gudmalwar, Pankaj Wasnik

    Speech time reversal refers to the process of reversing the entire speech signal in time, causing it to play backward. Such signals are completely unintelligible since the fundamental structures of phonemes and syllables are destroyed. However, they still retain tonal patterns that enable perceptual speaker identification despite losing linguistic content. I

  93. Haiyun Li, Zhiyong Wu, Xiaofeng Xie, Jingran Xie

    Voice cloning (VC)-resistant watermarking is an emerging technique for tracing and preventing unauthorized cloning. Existing methods effectively trace traditional VC models by training them on watermarked audio but fail in zero-shot VC scenarios, where models synthesize audio from an audio prompt without training. To address this, we propose VoiceMark, the f

  94. Yifei Wang, Weimin Bai, Colin Zhang, Debing Zhang

    In this paper, we unify more than 10 existing one-step diffusion distillation approaches, such as Diff-Instruct, DMD, SIM, SiD, $f$-distill, etc, inside a theory-driven framework which we name the \textbf{\emph{Uni-Instruct}}. Uni-Instruct is motivated by our proposed diffusion expansion theory of the $f$-divergence family. Then we introduce key theories tha

  95. Zonghao Chen, Toni Karvonen, Heishiro Kanagawa, François-Xavier Briol

    Approximation of a target probability distribution using a finite set of points is a problem of fundamental importance in numerical integration. Several authors have proposed to select points by minimising a maximum mean discrepancy (MMD), but the non-convexity of this objective typically precludes global minimisation. Instead, we consider the concept of \em

  96. Yufei Zhan, Hongyin Zhao, Yousong Zhu, Shurong Zheng

    Large Multimodal Models (LMMs) have recently demonstrated remarkable visual understanding performance on both vision-language and vision-centric tasks. However, they often fall short in integrating advanced, task-specific capabilities for compositional reasoning, which hinders their progress toward truly competent general vision models. To address this, we p

  97. Lipei Du, Ulrich Heinz

    We present a novel multimessenger approach to extract the effective radial flow of the quark-gluon plasma (QGP) by jointly analyzing thermal photon and dilepton spectra in heavy-ion collisions. A key feature of this method is that it circumvents the need for a directly unmeasurable reference -- the photon temperature in the absence of flow -- by establishing

  98. Zongcai Tan, Dandan Zhang

    Optical tweezers (OT) offer unparalleled capabilities for micromanipulation with submicron precision in biomedical applications. However, controlling conventional multi-trap OT to achieve cooperative manipulation of multiple complex-shaped microrobots in dynamic environments poses a significant challenge. To address this, we introduce Interactive OT Gym, a r

  99. Yuan Zhang, Wenxuan Xu, Mohamed Darouach, Tyrone Fernando

    Target output controllers aim at regulating a system's target outputs by placing poles of a suitable subsystem using partial state feedback, where full state controllability is not required. This paper establishes existence conditions for such controllers using input and partial state data, where the system dynamics are unknown. The approach bypasses traditi

  100. Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen

    LLM-based agent systems are emerging as a new software paradigm and have been widely adopted across diverse domains such as medicine, robotics, and programming. However, maintaining these systems requires substantial effort, as they are inevitably prone to bugs and continually evolve to meet changing external requirements. Therefore, automatically resolving