Skip to content

May 2025 arXiv papers — page 48

Showing 4,7014,800 of 24,552 papers

  1. Suguman Bansal, Ramneet Singh

    This paper presents a novel symbolic algorithm for the Maximal End Component (MEC) decomposition of a Markov Decision Process (MDP). The key idea behind our algorithm INTERLEAVE is to interleave the computation of Strongly Connected Components (SCCs) with eager elimination of redundant state-action pairs, rather than performing these computations sequentiall

  2. Feng Jiang, Zihao Zheng, Xiuping Cui, Maoliang Li

    With the development of Embodied Artificial intelligence, the end-to-end control policy such as Vision-Language-Action (VLA) model has become the mainstream. Existing VLA models faces expensive computing/storage cost, which need to be optimized. Quantization is considered as the most effective method which can not only reduce the memory cost but also achieve

  3. Yu Xu, Biqiang Mu, Tianshi Chen

    There have been increasing interests on the Volterra series identification with the kernel-based regularization method. The major difficulties are on the kernel design and efficiency of the corresponding implementation. In this paper, we first assume that the underlying system to be identified is the Wiener-Hammerstein (WH) system with polynomial nonlinearit

  4. Nikola Andrejic, Milica Spasic, Igor Mihajlovic, Petra Milosavljevic

    This work introduces Ui2i, a novel model for unpaired image-to-image translation, trained on content-wise unpaired datasets to enable style transfer across domains while preserving content. Building on CycleGAN, Ui2i incorporates key modifications to better disentangle content and style features, and preserve content integrity. Specifically, Ui2i employs U-N

  5. Jingping Nie, Dung T. Tran, Karan Thakkar, Vasudha Kowtha

    Auscultation, particularly heart sound, is a non-invasive technique that provides essential vital sign information. Recently, self-supervised acoustic representation foundation models (FMs) have been proposed to offer insights into acoustics-based vital signs. However, there has been little exploration of the extent to which auscultation is encoded in these

  6. Hao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang

    Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a novel self-supervised framework that enhances interpretability by tokenizing inertial measurement unit signals into semant

  7. Yanran Tang, Ruihong Qiu, Zi Huang

    Legal case retrieval plays a pivotal role in the legal domain by facilitating the efficient identification of relevant cases, supporting legal professionals and researchers to propose legal arguments and make informed decision-making. To improve retrieval accuracy, the Competition on Legal Information Extraction and Entailment (COLIEE) is held annually, offe

  8. Sunwoo Kim, Soo Yong Lee, Jaemin Yoo, Kijung Shin

    While graph neural networks (GNNs) have shown remarkable performance across diverse graph-related tasks, their high-dimensional hidden representations render them black boxes. In this work, we propose Graph Lingual Network (GLN), a GNN built on large language models (LLMs), with hidden representations in the form of human-readable text. Through careful promp

  9. Jiatong Shi, Hye-Jin Shim, Shinji Watanabe

    Subjective listening tests remain the golden standard for speech quality assessment, but are costly, variable, and difficult to scale. In contrast, existing objective metrics, such as PESQ, F0 correlation, and DNSMOS, typically capture only specific aspects of speech quality. To address these limitations, we introduce Uni-VERSA, a unified network that simult

  10. Xiangyu Zhao, Wanghan Xu, Bo Liu, Yuhao Zhou

    The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at the graduate level-remains underexplored due to a lack of benchmarks reflecting the depth and complexity of geoscientific reasoning. Existing datasets often rely on synthetic data o

  11. Kunpeng Zhao, Asahi Miyazaki, Tsuyoshi Okita

    Human Activity Recognition (HAR) has recently witnessed advancements with Transformer-based models. Especially, ActionFormer shows us a new perspectives for HAR in the sense that this approach gives us additional outputs which detect the border of the activities as well as the activity labels. ActionFormer was originally proposed with its input as image/vide

  12. Peiwen Yuan, Yiwei Li, Shaoxiong Feng, Xinglin Wang

    LLM-as-Benchmark-Generator methods have been widely studied as a supplement to human annotators for scalable evaluation, while the potential biases within this paradigm remain underexplored. In this work, we systematically define and validate the phenomenon of inflated performance in models evaluated on their self-generated benchmarks, referred to as self-bi

  13. Zilong Wang, Jingfeng Yang, Sreyashi Nag, Samarth Varshney

    Large language models (LLMs) have exhibited extraordinary performance in a variety of tasks while it remains challenging for them to solve complex multi-step tasks as agents. In practice, agents sensitive to the outcome of certain key steps which makes them likely to fail the task because of a subtle mistake in the planning trajectory. Recent approaches reso

  14. Hui-Min Wang, Kai Liao, Shao-Wen Wei

    We study the orbital dynamics and relativistic precession effects in the spacetime of rotating braneworld black holes within the Randall-Sundrum framework. For test particles on spherical orbits, we analyze three conserved quantities-energy, angular momentum, and Carter constant-and examine how the innermost stable spherical orbit depends on the tidal charge

  15. Jianfeng Yu, Yanyong Hong

    In this paper, we introduce the definition of extended $\mathcal{O}$-operators on a Novikov algebra $(A,\circ)$ associated to an $A$-bimodule Novikov algebra which is a generalization of the definition of $\mathcal{O}$-operators and show that there are new Novikov algebra structures on the $A$-bimodule Novikov algebra obtained from extended $\mathcal{O}$-ope

  16. Zhuoyu Cheng, Kohei Hatano, Eiji Takimoto

    We consider a bandit optimization problem for nonconvex and non-smooth functions, where in each trial the loss function is the sum of a linear function and a small but arbitrary perturbation chosen after observing the player's choice. We give both expected and high probability regret bounds for the problem. Our result also implies an improved high-probabilit

  17. Cheonsu Jeong, Seongmin Sim, Hyoyoung Cho, Sungsu Kim

    This paper presents an intelligent work automation approach in the context of contemporary digital transformation by integrating generative AI and Intelligent Document Processing (IDP) technologies with an Automation Agent to realize End-to-End (E2E) automation of corporate financial expense processing tasks. While traditional Robotic Process Automation (RPA

  18. Hanlin Wang, Chak Tou Leong, Jiashuo Wang, Jian Wang

    Reinforcement learning (RL) holds significant promise for training LLM agents to handle complex, goal-oriented tasks that require multi-step interactions with external environments. However, a critical challenge when applying RL to these agentic tasks arises from delayed rewards: feedback signals are typically available only after the entire task is complete

  19. Linshanshan Wang, Mengyan Li, Zongqi Xia, Molei Liu

    Electronic Health Records (EHR) offer rich real-world data for personalized medicine, providing insights into disease progression, treatment responses, and patient outcomes. However, their sparsity, heterogeneity, and high dimensionality make them difficult to model, while the lack of standardized ground truth further complicates predictive modeling. To addr

  20. Shahrooz Pouryousef, Ali Montazeralghaem

    Collaborative information from user-item interactions is a fundamental source of signal in successful recommender systems. Recently, researchers have attempted to incorporate this knowledge into large language model-based recommender approaches (LLMRec) to enhance their performance. However, there has been little fundamental analysis of whether LLMs can effe

  21. Xiangyu Sun, Runnan Chen, Mingming Gong, Dong Xu

    Sparse-view scene reconstruction often faces significant challenges due to the constraints imposed by limited observational data. These limitations result in incomplete information, leading to suboptimal reconstructions using existing methodologies. To address this, we present Intern-GS, a novel approach that effectively leverages rich prior knowledge from v

  22. Zesen Lyu, Dandan Zhang, Wei Ye, Fangdi Li

    Spatial reasoning is a core component of human cognition, enabling individuals to perceive, comprehend, and interact with the physical world. It relies on a nuanced understanding of spatial structures and inter-object relationships, serving as the foundation for complex reasoning and decision-making. To investigate whether current vision-language models (VLM

  23. Jingchao Fang, Mina Lee

    Have you ever read a blog or social media post and suspected that it was written--at least in part--by artificial intelligence (AI)? While transparently acknowledging contributors to writing is generally valued, why some writers choose to disclose or withhold AI involvement remains unclear. In this work, we ask what factors shape writers' decisions to disclo

  24. Liu Dai, Haina Wang, Weikang Wan, Hao Su

    Building embodied agents capable of accomplishing arbitrary tasks is a core objective towards achieving embodied artificial general intelligence (E-AGI). While recent work has advanced such general robot policies, their training and evaluation are often limited to tasks within specific scenes, involving restricted instructions and scenarios. Existing benchma

  25. Alberto Pliego Marugán, Jesús M. Pinar-Pérez, Fausto Pedro García Márquez

    Efficient maintenance has always been essential for the successful application of engineering systems. However, the challenges to be overcome in the implementation of Industry 4.0 necessitate new paradigms of maintenance optimization. Machine learning techniques are becoming increasingly used in engineering and maintenance, with reinforcement learning being

  26. Sihyeon Lee, Hyunjoo Song, Jong-chan Lee, Yoon Jin Lee

    Deploying large language models (LLMs) in clinical settings faces critical trade-offs: cloud LLMs, with their extensive parameters and superior performance, pose risks to sensitive clinical data privacy, while local LLMs preserve privacy but often fail at complex clinical interpretation tasks. We propose MedOrchestra, a hybrid framework where a cloud LLM dec

  27. Il Gyeong Choi, Dong-han Yeom

    We consider the possibility that the weak cosmic censorship conjecture can be violated using the Oppenheimer-Snyder collapse model with a perfect fluid star interior. Metric models with a naked singularity can be used, and the Oppenheimer-Snyder collapse might be possible; the null energy condition is also satisfied in these models. However, this is just a t

  28. Pascal Zwick, Nils Friederich, Maximilian Beichter, Lennart Hilbert

    Enhancing the efficiency of high-quality image generation using Diffusion Models (DMs) is a significant challenge due to the iterative nature of the process. Flow Matching (FM) is emerging as a powerful generative modeling paradigm based on a simulation-free training objective instead of a score-based one used in DMs. Typical FM approaches rely on a Gaussian

  29. H. A. Adarsha, Chandrachur Chakraborty, Sudip Bhattacharyya

    We introduce a novel mechanism -- Magnetically Arrested Transmutation (MAT) -- which could be a viable model to account for the observed over-representation of magnetic white dwarfs (WDs) near the Galactic centre (GC), and the presence of a magnetar as opposed to the absence of ordinary pulsars in the same region. In this scenario, compact stars accumulate a

  30. Gao Huayu, Huang Tengjiu, Ye Xiaolong, Tsuyoshi Okita

    AI-based motion capture is an emerging technology that offers a cost-effective alternative to traditional motion capture systems. However, current AI motion capture methods rely entirely on observed video sequences, similar to conventional motion capture. This means that all human actions must be predefined, and movements outside the observed sequences are n

  31. Zaijun Ye, Chen-Song Zhang, Wansheng Wang

    Neural operators have emerged as powerful tools for learning solution operators of partial differential equations. However, in time-dependent problems, standard training strategies such as teacher forcing introduce a mismatch between training and inference, leading to compounding errors in long-term autoregressive predictions. To address this issue, we propo

  32. Haicheng Liao, Zhenning Li, Guohui Zhang, Keqiang Li

    Predicting the trajectories of vehicles is crucial for the development of autonomous driving (AD) systems, particularly in complex and dynamic traffic environments. In this study, we introduce HiT (Human-like Trajectory Prediction), a novel model designed to enhance trajectory prediction by incorporating behavior-aware modules and dynamic centrality measures

  33. Mehdi Neshat, Nataliia Y. Sergiienko, Leandro S. P. da Silva, Seyedali Mirjalili

    Floating hybrid wind-wave systems combine offshore wind platforms with wave energy converters (WECs) to create cost-effective and reliable energy solutions. Adequately designed and tuned WECs are essential to avoid unwanted loads disrupting turbine motion while efficiently harvesting wave energy. These systems diversify energy sources, enhancing energy secur

  34. Kengo Suzuki, Takeshi Iwashita

    Low-precision computing is essential for efficiently utilizing memory bandwidth and computing cores. While many mixed-precision algorithms have been developed for iterative sparse linear solvers, effectively leveraging half-precision (fp16) arithmetic remains challenging. This study introduces a novel nested Krylov approach that integrates the flexible GMRES

  35. Kui Wu, Shuhang Xu, Hao Chen, Churan Wang

    We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tracking systems in recovering from tracking failure. Our approach combines the off-the-shelf active tracking methods with VLMs' reasoning capabilities, deploying a fast visual polic

  36. Sobirjon Shoyimardonov

    In this paper, we investigate a discrete-time phytoplankton-zooplankton model that incorporates a linear predator functional response alongside a Holling-type toxin distribution. Both Holling type II and type III cases are considered, and we derive conditions on the model parameters that guarantee the existence of positive fixed points. We classify all fixed

  37. Reza Nematirad, Anil Pahwa, Balasubramaniam Natarajan

    Time series forecasting plays a crucial role in many real-world applications, and numerous complex forecasting models have been proposed in recent years. Despite their architectural innovations, most state-of-the-art models report only marginal improvements -- typically just a few thousandths in standard error metrics. These models often incorporate complex

  38. Fuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang

    Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained temporal reasoning. While reinforcement learning (RL) has been explored to address this issue recently, existing RL approaches remain limit

  39. Zechen Li, Lanqing Yang, Yiheng Bian, Hao Pan

    Indoor environments typically contain diverse RF signals distributed across multiple frequency bands, including NB-IoT, Wi-Fi, and millimeter-wave. Consequently, wideband RF modeling is essential for practical applications such as joint deployment of heterogeneous RF systems, cross-band communication, and distributed RF sensing. Although 3D Gaussian Splattin

  40. Shihan Zhao, Stefanos Nikolaidis

    Quality-Diversity (QD) optimization is an emerging field that focuses on finding a set of behaviorally diverse and high-quality solutions. While the quality is typically defined w.r.t. a single objective function, recent work on Multi-Objective Quality-Diversity (MOQD) extends QD optimization to simultaneously optimize multiple objective functions. This open

  41. Ding Xia, Xinyue Gui, Fan Gao, Dongyuan Li

    The absence of explicit communication channels between automated vehicles (AVs) and other road users requires the use of external Human-Machine Interfaces (eHMIs) to convey messages effectively in uncertain scenarios. Currently, most eHMI studies employ predefined text messages and manually designed actions to perform these messages, which limits the real-wo

  42. Kui Wu, Hao Chen, Churan Wang, Fakhri Karray

    User-Centric Embodied Visual Tracking (UC-EVT) presents a novel challenge for reinforcement learning-based models due to the substantial gap between high-level user instructions and low-level agent actions. While recent advancements in language models (e.g., LLMs, VLMs, VLAs) have improved instruction comprehension, these models face critical limitations in

  43. Chen Lu, Mingjin Li, Jianren Long

    On the one hand, the fractional order derivative characterization of the Besov-Morrey type space $B_{p}^{K}(s)$ is established by $K$-Carleson measures, and it was also shown that $f \in B_{p}^{K}(s_1) \Leftrightarrow f^{\left(\frac{s_2 - s_1}{p}\right)} \in B_{p}^{K}(s_2)$, which extended the results of Sun et al. on the fractional derivative of Morrey type

  44. Ignacio Esponda, Demian Pouzo

    We study learning in complete-information games, allowing the players' models of their environment to be misspecified. We introduce Berk--Nash rationalizability: the largest self-justified set of actions -- meaning each action in the set is optimal under some belief that is a best fit to outcomes generated by joint play within the set. We show that, in a mod

  45. Nicy Scaria, Silvester John Joseph Kennedy, Krishna Agarwal, Diksha Seth

    Small Language Models (SLMs) offer privacy and efficiency for educational deployment, yet their utility depends on reliable multistep reasoning. Existing benchmarks often prioritize final answer accuracy, obscuring 'right answer, wrong procedure' failures that can reinforce student misconceptions. This work investigates SLM physics reasoning reliability, sta

  46. Shaohua Zhang, Yuchong Luo, Shangchun Xie, Chao Gao

    We present, for the first time, a systematic study of quasar-associated 2175 \AA\ dust absorbers using spectroscopic data from the Sloan Digital Sky Survey (SDSS) Data Release 16 (DR16). By analyzing the optical spectra and multi-band magnitudes of 557,674 quasars in the redshift range of $0.7 \le z \le 2.4$, we identify 843 absorbers that share the same red

  47. Yang Wang, Wenxuan Zhu, Xuehui Quan, Heyi Wang

    This paper addresses the challenges of fault prediction and delayed response in distributed systems by proposing an intelligent prediction method based on temporal feature learning. The method takes multi-dimensional performance metric sequences as input. We use a Gated Recurrent Unit (GRU) to model the evolution of system states over time. An attention mech

  48. Zixuan Hu, Yichun Hu, Xiaotong Li, Shixiang Tang

    Wild Test-Time Adaptation (WTTA) is proposed to adapt a source model to unseen domains under extreme data scarcity and multiple shifts. Previous approaches mainly focused on sample selection strategies, while overlooking the fundamental problem on underlying optimization. Initially, we critically analyze the widely-adopted entropy minimization framework in W

  49. Jiong Li, Qing-Hu Chen

    We investigate the spectral properties and quantum criticality of the two-photon Rabi-Stark model. Using the exact solution of this model, we rigorously derive a condition for complete spectral collapse, where all bound states vanish. In this case, the energy gap closes at a critical coupling, signaling a continuous quantum phase transition. The correspondin

  50. Riku Matsubayashi, Shiki Ogata, Taishi Ihara, Hiroyasu Matsudaira

    The magnetic field dependence of the spin part of Knight shift, which is proportional to the superconducting-state spin susceptibility, was investigated at two nuclear sites, $^{31}$P and $^{139}$La in a conventional s-wave superconductor LaRu$_4$P$_{12}$. After the analyses, we confirmed that the superconducting-state spin susceptibility is proportional to

  51. Ryosuke Kohita, Akira Kasuga

    Cloud architecture design is a complex process requiring both technical expertise and architectural knowledge to develop solutions from frequently ambiguous requirements. We present CloudArchitectBuddy, a system-driven cloud architecture design support application with two key mechanisms: (1) structured state management that enhances design understanding thr

  52. Koki Matsuishi, Tsuyoshi Okita

    In deep multi-instance learning, the number of applicable instances depends on the data set. In histopathology images, deep learning multi-instance learners usually assume there are hundreds to thousands instances in a bag. However, when the number of instances in a bag increases to 256 in brain hematoma CT, learning becomes extremely difficult. In this pape

  53. Yong Wu, Weihang Pan, Ke Li, Chen Binhui

    Large language models (LLMs) have shown remarkable reasoning capabilities, yet aligning such abilities to small language models (SLMs) remains a challenge due to distributional mismatches and limited model capacity. Existing reasoning datasets, typically designed for powerful LLMs, often lead to degraded performance when directly applied to weaker models. In

  54. Isabella Novik, Hailun Zheng

    For a $(d-1)$-dimensional simplicial complex $\Delta$ and $1\leq i\leq d$, let $f_{i-1}$ be the number of $(i-1)$-faces of $\Delta$ and $m_i$ be the number of missing $i$-faces of $\Delta$. In the nineties, Kalai asked for a characterization of the $m$-numbers of simplicial polytopes and spheres -- a problem that remains wide open to this day. Here, we study

  55. Woomin Song, Jihoon Tack, Sangwoo Mo, Seunghyuk Oh

    State-space models (SSMs) offer a promising architecture for sequence modeling, providing an alternative to Transformers by replacing expensive self-attention with linear recurrences. In this paper, we propose a simple yet effective trick to enhance SSMs within given computational budgets by sparsifying them. Our intuition is that tokens in SSMs are highly r

  56. Zachary C. Brown, David Carlson

    The field of hypothesis generation promises to reduce costs in neuroscience by narrowing the range of interventional studies needed to study various phenomena. Existing machine learning methods can generate scientific hypotheses from complex datasets, but many approaches assume causal relationships are static over time, limiting their applicability to system

  57. Marc A. Tunnell, David F. Gleich

    Despite hundreds of papers on preconditioned linear systems of equations, there remains a significant lack of comprehensive performance benchmarks comparing various preconditioners for solving symmetric positive definite (SPD) systems. In this paper, we present a comparative study of 79 matrices using a broad range of preconditioners. Specifically, we evalua

  58. Sharmistha Dey, Nahid Chaudhary, Ulrich Kentsch, Rajendra Singh

    The alterations in the magnetic properties and electronic structure of chemical vapor deposition (CVD) grown nano-dimensional molybdenum disulfide (MoS2) after low energy ion irradiation are thoroughly investigated. The formation of pure hexagonal 2-H phase has been identified by Raman spectroscopy and X-ray diffraction (XRD). The pristine samples are irradi

  59. Haobo Li, Eunseo Jung, Zixin Chen, Zhaowei Wang

    Multimodal time series forecasting is foundational in various fields, such as utilizing satellite imagery and numerical data for predicting typhoons in climate science. However, existing multimodal approaches primarily focus on utilizing text data to help time series forecasting, leaving the visual data in existing time series datasets untouched. Furthermore

  60. Xulin Gu, Xinhao Zhong, Zhixing Wei, Yimin Zhou

    Dataset distillation (DD) has emerged as a powerful paradigm for dataset compression, enabling the synthesis of compact surrogate datasets that approximate the training utility of large-scale ones. While significant progress has been achieved in distilling image datasets, extending DD to the video domain remains challenging due to the high dimensionality and

  61. Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha, Mitesh M. Khapra

    What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing. We compare: (i) training from scratch, (ii) fine-tuning English F5 on Indian data, and (iii) fine-tuning on both Indian and English data t

  62. Saharsh Barve, Andy Mao, Jiayue Melissa Shi, Prerna Juneja

    Recent advances in generative AI have enabled visual content creation through text-to-image (T2I) generation. However, despite their creative potential, T2I models often replicate and amplify societal stereotypes -- particularly those related to gender, race, and culture -- raising important ethical concerns. This paper proposes a theory-driven bias detectio

  63. Shenkai Zhao, Xinao Zhang, Lipeng Pan, Xiaobin Xu

    Semi-supervised classification based on active learning has made significant progress, but the existing methods often ignore the uncertainty estimation (or reliability) of the prediction results during the learning process, which makes it questionable whether the selected samples can effectively update the model. Hence, this paper proposes an evidential deep

  64. S. A. Avdonin, V. S. Mikhaylov

    We study the boundary control problems for the wave, heat, and Schr\"odinger equations on a finite graph. We suppose that the graph is a tree (i.e., it does not contain cycles), and on each edge an equation is defined. The control is acting through the Dirichlet condition applied to all or all but one boundary vertices. The exact controllability in $L_2$-cla

  65. A. S. Mikhaylov, V. S. Mikhaylov

    We consider the inverse dynamic problem for a dynamical system with discrete time associated with a semi-infinite complex Jacobi matrix. We propose two approaches of recovering coefficients from dynamic response operator and answer a question on the characterization of dynamic inverse data.

  66. Taehyo Kim, Qiran Jia, Mony J. de Leon, Hai Shu

    False discovery rate (FDR) control methods are essential for voxel-wise multiple testing in neuroimaging data analysis, where hundreds of thousands or even millions of tests are conducted to detect brain regions associated with disease-related changes. Classical FDR control methods (e.g., BH, q-value, and LocalFDR) assume independence among tests and often l

  67. Mingxuan Sun, Juntao Jiang, Zhiqiang Yang, Shenao Kong

    Microalgae, vital for ecological balance and economic sectors, present challenges in detection due to their diverse sizes and conditions. This paper summarizes the second "Vision Meets Algae" (VisAlgae 2023) Challenge, aiming to enhance high-throughput microalgae cell detection. The challenge, which attracted 369 participating teams, includes a dataset of 10

  68. Kianté Brantley, Mingyu Chen, Zhaolin Gao, Jason D. Lee

    Reinforcement learning (RL) has emerged as a powerful tool for fine-tuning large language models (LLMs) to improve complex reasoning abilities. However, state-of-the-art policy optimization methods often suffer from high computational overhead and memory consumption, primarily due to the need for multiple generations per prompt and the reliance on critic net

  69. Rosen Ting-Ying Yu, Cyril Picard, Faez Ahmed

    Bayesian optimization (BO) struggles in high dimensions, where Gaussian-process surrogates demand heavy retraining and brittle assumptions, slowing progress on real engineering and design problems. We introduce GIT-BO, a Gradient-Informed BO framework that couples TabPFN v2, a tabular foundation model that performs zero-shot Bayesian inference in context, wi

  70. Pengyuan Li, Boris Glavic, Dieter Gawlick, Vasudha Krishnaswamy

    Provenance-based data skipping compactly over-approximates the provenance of a query using so-called provenance sketches and utilizes such sketches to speed-up the execution of subsequent queries by skipping irrelevant data. However, a sketch captured at some time in the past may become stale if the data has been updated subsequently. Thus, there is a need t

  71. Wei Tao, Ya-Hui Zhao, Li-Ting Wang, Qin Chang

    We present the Non-Relativistic QCD (NRQCD) calculations at the next-to-leading order (NLO) of $\alpha_s$ for $B_c^*\to \eta_c$ vector, axial-vector, tensor and axial-tensor form factors, and obtain complete analytical expressions for the form factors, along with their asymptotic forms in the hierarchical heavy quark limit. Our results show that the NLO corr

  72. Enrique Ernesto Álvarez, Maximiliano Luis Riddick

    Hereby we propose a Bayesian method of estimation for the semiparametric Additive Hazards Model (AHM) from Survival Analysis under right-censoring. With this aim, we review the AHM revisiting the likelihood function, so as to comment on the challenges posed by Bayesian estimation from the full likelihood. Through an algorithmic reformulation of that likeliho

  73. Haodong Lu, Xinyu Zhang, Kristen Moore, Jason Xue

    Continual learning (CL) enables deep networks to acquire new knowledge while avoiding catastrophic forgetting. The powerful generalization ability of pre-trained models (PTMs), such as the Contrastive Language-Image Pre-training (CLIP) model, has inspired a range of CL methods targeting new and specialized tasks, providing rich multi-modal embeddings that su

  74. Danush Khanna, Pratinav Seth, Sidhaarth Sredharan Murali, Aditya Kumar Guru

    Mental manipulation is a subtle yet pervasive form of abuse in interpersonal communication, making its detection critical for safeguarding potential victims. However, due to manipulation's nuanced and context-specific nature, identifying manipulative language in complex, multi-turn, and multi-person conversations remains a significant challenge for large lan

  75. Tianhua Qi, Shiyan Wang, Cheng Lu, Tengfei Song

    Controllable emotional voice conversion (EVC) aims to manipulate emotional expressions to increase the diversity of synthesized speech. Existing methods typically rely on predefined labels, reference audios, or prespecified factor values, often overlooking individual differences in emotion perception and expression. In this paper, we introduce PromptEVC that

  76. Weiping Yao, Yifu Cai, Zehua Tian

    We investigate the holographic entanglement entropy (HEE) with dark matter in a higher-dimensional AdS black hole spacetime including full back reaction, revealing its role as a diagnostic tool for critical phenomena in strongly coupled systems. By analyzing the HEE, we uncover distinct signatures of the metal/superconductor phase transition, demonstrating t

  77. Sadaf Safa, Ali Abedi, Shehroz S. Khan

    Student engagement plays a crucial role in the successful delivery of educational programs. Automated engagement measurement helps instructors monitor student participation, identify disengagement, and adapt their teaching strategies to enhance learning outcomes effectively. This paper identifies two key challenges in this problem: class imbalance and incorp

  78. Lingyu Qiu, Ke Jiang, Xiaoyang Tan

    In this paper, we propose a new cross-domain face forgery detection method that is insensitive to different and possibly unseen forgery methods while ensuring an acceptable low false positive rate. Although existing face forgery detection methods are applicable to multiple domains to some degree, they often come with a high false positive rate, which can gre

  79. Boyi Zeng, Shixiang Song, Siyuan Huang, Yixuan Wang

    Humans ponder before articulating complex sentence elements, enabling deeper cognitive processing through focused effort. In this work, we introduce this pondering process into language models by repeatedly invoking the forward process within a single token generation step. During pondering, instead of generating an actual token sampled from the prediction d

  80. Yuxiang Zhang, Jianhua Zhang, Xidong Hu, Jiwei Zhang

    Accurate radar cross section (RCS) modeling is crucial for characterizing target scattering and improving the precision of Integrated Sensing and Communication (ISAC) channel modeling. Existing RCS models are typically designed for specific target types, leading to increased complexity and lack of generalization. This makes it difficult to standardize RCS mo

  81. Woochang Sim, Hyunseok Ryu, Kyungmin Choi, Sungwon Han

    The Abstraction and Reasoning Corpus (ARC) poses a stringent test of general AI capabilities, requiring solvers to infer abstract patterns from only a handful of examples. Despite substantial progress in deep learning, state-of-the-art models still achieve accuracy rates of merely 40-55% on 2024 ARC Competition, indicative of a significant gap between their

  82. Heng Tan, Hua Yan, Yu Yang

    While reinforcement learning (RL) has achieved notable success in various domains, training effective policies for complex tasks remains challenging. Agents often converge to local optima and fail to maximize long-term rewards. Existing approaches to mitigate training bottlenecks typically fall into two categories: (i) Automated policy refinement, which iden

  83. Zikang Guo, Benfeng Xu, Xiaorui Wang, Zhendong Mao

    Complex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting erroneous trajectories in agentic workflows. However, existing approaches only exploit such capability in the post-acti

  84. Hyein Yoon, O. Ivy Wong, Aeree Chung, Shan Huang

    We investigate the star formation and neutral atomic hydrogen (HI) gas properties of galaxies along three large-scale filaments and two galaxy groups in the wide field around the Virgo cluster. Our goal is to understand how galaxies are processed in low-density environments before falling into high-density regions. Combining the spatial distribution of galax

  85. Seongmin Kim, Kwangmin Lee, Sewon Park, Jaeyong Lee

    In multivariate statistics, estimating the covariance matrix is essential for understanding the interdependence among variables. In high-dimensional settings, where the number of covariates increases with the sample size, it is well known that the eigenstructure of the sample covariance matrix is inconsistent. The inverse-Wishart prior, a standard choice for

  86. Kai Yang, Hui Ma, Shaoyu Dou

    Anomalies are common in network system monitoring. When manifested as network threats to be mitigated, service outages to be prevented, and security risks to be ameliorated, detecting such anomalous network behaviors becomes of great importance. However, the growing scale and complexity of the mobile communication networks, as well as the ever-increasing amo

  87. Weiyin Gong, Kai Zhang, Yanghai Zhang, Qi Liu

    Multimodal intent recognition (MIR) seeks to accurately interpret user intentions by integrating verbal and non-verbal information across video, audio and text modalities. While existing approaches prioritize text analysis, they often overlook the rich semantic content embedded in non-verbal cues. This paper presents a novel Wavelet-Driven Multimodal Intent

  88. Tianwa Chen, Barbara Weber, Graeme Shanks, Gianluca Demartini

    A range of integrated modeling approaches have been developed to enable a holistic representation of business process logic together with all relevant business rules. These approaches address inherent problems with separate documentation of business process models and business rules. In this study, we explore how expert process workers make sense of the info

  89. Haoyang Feng, Yanjun Dai, Yuan Gao

    For personalized marketing, a new challenge of how to effectively algorithm the A/B testing to maximize user response is urgently to be overcome. In this paper, we present a new approach, the RL-LLM-AB test framework, for using reinforcement learning strategy optimization combined with LLM to automate and personalize A/B tests. The RL-LLM-AB test is built up

  90. Yukun Zhang, Xueqing Zhou

    We propose a novel framework, Continuous_Time Attention, which infuses partial differential equations (PDEs) into the Transformer's attention mechanism to address the challenges of extremely long input sequences. Instead of relying solely on a static attention matrix, we allow attention weights to evolve over a pseudo_time dimension via diffusion, wave, or r

  91. Hanxi Guo, Siyuan Cheng, Kaiyuan Zhang, Guangyu Shen

    Large language models (LLMs) have become integral to modern software development, producing vast amounts of AI-generated source code. While these models boost programming productivity, their misuse introduces critical risks, including code plagiarism, license violations, and the propagation of insecure programs. As a result, robust detection of AI-generated

  92. Muxi Diao, Lele Yang, Hongbo Yin, Zhexu Wang

    Effective autonomous driving hinges on robust reasoning across perception, prediction, planning, and behavior. However, conventional end-to-end models fail to generalize in complex scenarios due to the lack of structured reasoning. While recent vision-language models (VLMs) have been applied to driving tasks, they typically rely on isolated modules and stati

  93. Yang He, Xiao Ding, Bibo Cai, Yufei Zhang

    While reasoning-augmented large language models (RLLMs) significantly enhance complex task performance through extended reasoning chains, they inevitably introduce substantial unnecessary token consumption, particularly for simpler problems where Short Chain-of-Thought (Short CoT) suffices. This overthinking phenomenon leads to inefficient resource usage wit

  94. Xu Kang, Siqi Jiang, Kangwei Xu, Jiahao Li

    Terpenoids are a crucial class of natural products that have been studied for over 150 years, but their interdisciplinary nature (spanning chemistry, pharmacology, and biology) complicates knowledge integration. To address this, the authors developed TeroSeek, a curated knowledge base (KB) built from two decades of terpenoid literature, coupled with an AI-po

  95. Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi

    Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods often renders reproduction a labor-intensive endeavor, necessitating profound domain expertise. To address this, we introduce the paper lineage, which systematically mines implicit knowledge from the cited literatur

  96. Mulyanto, Hadyan Luthfan Prihadi, Emir Syahreza Fadhilla, Ardian Nata Atmaja

    We investigate the holographic renormalization of scalar-torsion gravity in a four-dimensional bulk spacetime with non-minimal derivative coupling. The asymptotic behavior of the static equations leads to an anti-de Sitter geometry for negative cosmological constants, allowing for a holographic interpretation via the AdS/CFT correspondence. The existence of

  97. Qinzhuo Wu, Pengzhi Gao, Wei Liu, Jian Luan

    Graphical User Interface (GUI) agents have gained substantial attention due to their impressive capabilities to complete tasks through multiple interactions within GUI environments. However, existing agents primarily focus on enhancing the accuracy of individual actions and often lack effective mechanisms for detecting and recovering from errors. To address

  98. Nathan Monette, Alistair Letcher, Michael Beukman, Matthew T. Jackson

    For reinforcement learning agents to be deployed in high-risk settings, they must achieve a high level of robustness to unfamiliar scenarios. One method for improving robustness is unsupervised environment design (UED), a suite of methods aiming to maximise an agent's generalisability across configurations of an environment. In this work, we study UED from a

  99. Yue Fang, Zhi Jin, Jie An, Hongshen Chen

    Temporal Logic (TL), especially Signal Temporal Logic (STL), enables precise formal specification, making it widely used in cyber-physical systems such as autonomous driving and robotics. Automatically transforming NL into STL is an attractive approach to overcome the limitations of manual transformation, which is time-consuming and error-prone. However, due

  100. M. Maus, M. White, N. Sailer, A. Baleato Lizancos

    The spectroscopic data from DESI Data Release 1 (DR1) galaxies enables the analysis of 3D clustering by fitting galaxy power spectra and reconstructed correlation functions in redshift space. Given low measurements of the amplitude of structure from cosmic shear at $z\sim1$, redshift space distortions (RSD) + Baryon Acoustic Oscillation (BAO) signals from DE