Skip to content

May 2025 arXiv papers — page 79

Showing 7,8017,900 of 24,552 papers

  1. Nikolaos Anastasiou, Spyros Kondylatos, Ioannis Papoutsis

    Accurate prediction of wildfire spread is crucial for effective risk management, emergency response, and strategic resource allocation. In this study, we present a deep learning (DL)-based framework for forecasting the final extent of burned areas, using data available at the time of ignition. We leverage a spatio-temporal dataset that covers the Mediterrane

  2. Yuchen He, Jianbing Lv, Liqi Cheng, Lingyu Meng

    Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficult

  3. Abolhassan Mohammadi, Bao-Fei Li, Tao Zhu

    We investigate the cosmological impacts of the alternative effective mass function of the modified Mukhanov-Sasaki equation in loop quantum cosmology, which is obtained from the polymerization of the classical mass function derived in the comoving gauge. This alternative effective mass function is distinct from those in the dressed metric and the hybrid appr

  4. Jinyuan Feng, Chaopeng Wei, Tenghai Qiu, Tianyi Hu

    In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off between model capacity and computation overhead. However, current MoE variants fall short on heterogeneous datasets, ignoring

  5. Zijie Qiu, Jiaqi Wei, Xiang Zhang, Sheng Xu

    De novo peptide sequencing is a critical task in proteomics. However, the performance of current deep learning-based methods is limited by the inherent complexity of mass spectrometry data and the heterogeneous distribution of noise signals, leading to data-specific biases. We present RankNovo, the first deep reranking framework that enhances de novo peptide

  6. Qiyu Chen, Huiyuan Luo, Haiming Yao, Wei Luo

    Anomaly detection plays a vital role in the inspection of industrial images. Most existing methods require separate models for each category, resulting in multiplied deployment costs. This highlights the challenge of developing a unified model for multi-class anomaly detection. However, the significant increase in inter-class interference leads to severe mis

  7. Xiaoyu Ye, Songjie Cheng, Yongtao Wang, Yajiao Xiong

    Recent advances in text-to-video (T2V) diffusion models have significantly enhanced the quality of generated videos. However, their capability to produce explicit or harmful content introduces new challenges related to misuse and potential rights violations. To address this newly emerging threat, we propose unlearning-based concept erasing as a solution. Fir

  8. Zuowu Zheng, Ze Wang, Fan Yang, Jiangke Fan

    Traditional online industrial advertising systems suffer from the limitations of multi-stage cascaded architectures, which often discard high-potential candidates prematurely and distribute decision logic across disconnected modules. While recent generative recommendation approaches provide end-to-end solutions, they fail to address critical advertising requ

  9. Ding Tang, Jiecheng Zhou, Jiakai Hu, Shengwei Li

    Recent advancements in large language models (LLMs) necessitate extensive computational resources, prompting the use of diverse hardware accelerators from multiple vendors. However, traditional distributed training frameworks struggle to efficiently utilize hyper-heterogeneous clusters comprising thousands of chips due to significant disparities in software

  10. Chuhan Lei, Xiaoqin Zhan

    This paper investigates $2$-$(v,5,\lambda)$ designs $\mathcal{D}$ admitting a block-transitive automorphism group $G$. We first prove that if $G$ is point-imprimitive, then $v$ must be one of 16, 21, or 81. We further provide a complete classification of all such designs for $v=16$ and $v=21$. Secondly, we demonstrate that if $G$ is point-primitive, then it

  11. Rafna Rafeek, Debasish Mondal

    An information engine harnesses energy from a single heat bath, utilising the gathered information. This study explores the best control strategy of a Brownian information engine (BIE), confined in a potential energy surface (PES) of arbitrary shape, and experiencing a measurement outcome-based feedback cycle. The feedback site corresponds to an instantaneou

  12. Jean-Marie Malherbe

    This paper is based on a dataset of many strongly polarized solar lines belonging to the ''second solar spectrum'', i.e. the spectrum near the limb in linear scattering polarization. The observations were done at the Pic du Midi Turret Dome in 2006. The solar spectra were recorded at high spectral resolution (R = 400000) with the spectrograph slit orthogonal

  13. Ruiqi Xing

    Medical image segmentation faces persistent challenges due to severe class imbalance and the frequency-specific distribution of anatomical structures. Most conventional CNN-based methods operate in the spatial domain and struggle to capture minority class signals, often affected by frequency aliasing and limited spectral selectivity. Transformer-based models

  14. Kaixing Yang, Xulong Tang, Ziqiao Peng, Yuxuan Hu

    Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance movement from audio signals. However, traditional methods underutilize genre conditioning, often treating it as auxiliary mo

  15. Bardh Prenkaj, Efstratios Zaradoukas, Gjergji Kasneci

    Counterfactual explainability seeks to uncover model decisions by identifying minimal changes to the input that alter the predicted outcome. This task becomes particularly challenging for graph data due to preserving structural integrity and semantic meaning. Unlike prior approaches that rely on forward perturbation mechanisms, we introduce Graph Inverse Sty

  16. Rabiou Issa, Kokou Mawulonmi Robert Afansounoudji, Komi Sodoga, David Lauvergnat

    In this study, we provide a novel wave packet propagation method that generalizes the Hagedorn approach by introducing alternative primitive basis sets that are better suited to describe different physical processes. More precisely, in our propagation scheme, we can mix basis sets with time-dependent parameters (the Hagedorn basis set) and time-independent o

  17. Mingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang

    Despite recent progress in text-to-image (T2I) generation, existing models often struggle to faithfully capture user intentions from short and under-specified prompts. While prior work has attempted to enhance prompts using large language models (LLMs), these methods frequently generate stylistic or unrealistic content due to insufficient grounding in visual

  18. Arindam Panda, Roland G. Winkler, Sunil P. Singh

    The rheological properties of tangentially propelled flexible polymers under linear shear flow are studied by computer simulations and are compared with analytical calculations. We find a significant impact of the coupled nonequilibrium active and shear forces on the polymer characteristics. The polar activity enhances shear-induced stretching along the flow

  19. Leonora Vesterbacka, Faton Rekathati, Robin Kurtz, Justyna Sikora

    This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual training datasets, substantial improvements in performance can be achieved by fine-tuning existing multilingual models, as sho

  20. Shiyu Ni, Keping Bi, Jiafeng Guo, Xueqi Cheng

    Large language models (LLMs) often produce incorrect answers with high confidence, yet the factors associated with such overconfidence remain insufficiently understood. We study this problem through the lens of knowledge popularity. Using entity-centric factual QA derived from Wikidata triplets, we characterize popularity through question entity popularity,

  21. Kent K. Chang, Mackenzie Hanh Cramer, Anna Ho, Ti Ti Nguyen

    While multimodal large language models (LLMs) excel at dialogue, whether they can adequately parse the structure of conversation -- conversational roles and threading -- remains underexplored. In this work, we introduce a suite of tasks and release TV-MMPC, a new annotated dataset, for multimodal conversation structure understanding. Our evaluation reveals t

  22. Denise Aregba-Driollet, Thomas Bellotti

    The concept of equilibrium is a general tool to fill the gap between macroscopic and mesoscopic information, both within kinetic systems and kinetic schemes. This work explores the use of equilibria to devise numerical boundary conditions for multi-dimensional vectorial lattice Boltzmann schemes tackling systems of hyperbolic conservation laws. In the scalar

  23. Jingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang

    This paper presents a pioneering exploration of reinforcement learning (RL) via group relative policy optimization for unified multimodal large language models (ULMs), aimed at simultaneously reinforcing generation and understanding capabilities. Through systematic pilot studies, we uncover the significant potential of ULMs to enable the synergistic co-evolu

  24. Pavan Ravishankar, Rushabh Shah, Daniel B. Neill

    We propose a fair machine learning algorithm to model interpretable differences between observed and desired human decision-making, with the latter aimed at reducing disparity in a downstream outcome impacted by the human decision. Prior work learns fair representations without considering the outcome in the decision-making process. We model the outcome disp

  25. Bin Wang, Heming Yang, Jinfang Sheng

    Recent studies have shown that by introducing prior knowledge, multi-scale analysis of complex and non-stationary time series in real environments can achieve good results in the field of long-term forecasting. However, affected by channel-independent methods, models based on multi-scale analysis may produce suboptimal prediction results due to the autocorre

  26. Sascha Kurz

    To each nodal hypersurface one can associate a binary linear code. Here we show that the binary linear code associated to sextics in $\mathbb{P}^3$ with the maximum number of $65$ nodes, as e.g. the Barth sextic, is unique. We also state possible candidates for codes that might be associated with a hypothetical septic attaining the currently best known upper

  27. Vendi Ardianto Nugroho, Byung Moo Lee

    Millimeter-wave (mmWave) communication enables high data rates for cellular-connected Unmanned Aerial Vehicles (UAVs). However, a robust beam management remains challenging due to significant path loss and the dynamic mobility of UAVs, which can destabilize the UAV-base station (BS) link. This research presents a GPS-aided deep learning (DL) model that simul

  28. Yeongjae Cho, Keonwoo Kim, Taebaek Hwang, Sungzoon Cho

    Recent advancements in Large Vision-Language Models (LVLMs) have significantly expanded their utility in tasks like image captioning and visual question answering. However, they still struggle with object hallucination, where models generate descriptions that inaccurately reflect the visual content by including nonexistent objects or misrepresenting existing

  29. Hai Jiang, Chushan Zheng, Jiawei Pan, Yuanpin Zhou

    Background: Accurate assessment of metastatic burden in axillary lymph nodes is crucial for guiding breast cancer treatment decisions, yet conventional imaging modalities struggle to differentiate metastatic burden levels and capture comprehensive lymph node characteristics. This study leverages dual-energy computed tomography (DECT) to exploit spectral-spat

  30. Kwangsu Lee

    Distributed broadcast encryption (DBE) is a variant of broadcast encryption (BE) that can efficiently transmit a message to a subset of users, in which users independently generate user private keys and user public keys instead of a central trusted authority generating user keys. In this paper, we propose a DBE scheme with constant size ciphertexts, constant

  31. M. Yaser Yagan, Ali E. Pusane, Ali Gorcin, Ibrahim Hokelek

    The majority of spatial signal processing techniques focus on increasing the total system capacity and providing high data rates for intended user(s). Unlike the existing studies, this paper introduces a novel interference modulation method that exploits the correlation between wireless channels to enable low-data-rate transmission towards additional users w

  32. Juliett Suárez Ferreira, Marija Slavkovik, Jorge Casillas

    Algorithmic decision-making systems sometimes produce errors or skewed predictions toward a particular group, leading to unfair results. Debiasing practices, applied at different stages of the development of such systems, occasionally introduce new forms of unfairness or exacerbate existing inequalities. We focus on post-processing techniques that modify alg

  33. Ye Du, Chen Yang, Nanxi Yu, Wanyu Lin

    De novo peptide sequencing is a fundamental computational technique for ascertaining amino acid sequences of peptides directly from tandem mass spectrometry data, eliminating the need for reference databases. Cutting-edge models usually encode the observed mass spectra into latent representations from which peptides are predicted autoregressively. However, t

  34. Fred Diamond, Payman L Kassaei

    Let $p$ be a prime, $F$ a totally real field in which $p$ is unramified, and $X/\overline{\mathbb{F}}_p$ a Shimura variety associated to ${\rm Res}_{F/\mathbb{Q}} {\rm GL}_2$ (or a PEL Hilbert modular variety). A mod $p$ Hilbert modular form of weight $\kappa$ can be defined as a section of an automorphic line bundle $\mathcal{L}_\kappa$ on $X$. We consider

  35. L. Yildiz, D. Kayki, M. F. Ciappina

    We demonstrate that dark matter interactions can profoundly influence stellar nucleosynthesis in the early universe by altering thermodynamic gradients and modifying nuclear reaction rates within primordial stars. Incorporating a dark matter-modified Fermi-Dirac distribution and accounting for localized energy injection from annihilation heating, our model p

  36. E. Kankare, T. Kangas, M. Fraser, S. Mattila

    We study a sample of narrow-line transients that share characteristics with the Type IIn classified supernova (SN) 1994W, a prototypical member of this class of events, via investigation of their explosion sites and spectrophotometric data. The normalised cumulative rank (NCR) method was used to compare the explosion sites of 10 events to the star-formation

  37. Salahuddin Alawadhi, Noorhan Abbas

    Integrating Retrieval Augmented Generation (RAG) with Large Language Models (LLMs) has shown the potential to provide precise, contextually relevant responses in knowledge intensive domains. This study investigates the ap-plication of RAG for ABB circuit breakers, focusing on accuracy, reliability, and contextual relevance in high-stakes engineering environm

  38. Mingxuan Gao, Jingjing Chen, Yun Long, Xiaomeng Xu

    Background: Silence is a common phenomenon in classrooms, yet its implicit nature limits a clear understanding of students' underlying learning statuses. Aim: This study proposed a nuanced framework to classify classroom silence based on class events and student status, and examined neurophysiological markers to reveal similarities and differences in silent

  39. Wenhan Chang, Tianqing Zhu, Yu Zhao, Shuangyong Song

    In the era of rapid generative AI development, interactions with large language models (LLMs) pose increasing risks of misuse. Prior research has primarily focused on attacks using template-based prompts and optimization-oriented methods, while overlooking the fact that LLMs possess strong unconstrained deceptive capabilities to attack other LLMs. This paper

  40. Tuomas Kangas, Panos Charalampopoulos, Takashi Nagao, Lin Yan

    We present our observations and analysis of SN 2023gpw, a hydrogen-rich superluminous supernova (SLSN II) with broad emission lines in its post-peak spectra. Unlike previously observed SLSNe II, its light curve suggests an abrupt drop during a solar conjunction between ~80 and ~180 d after the light-curve peak, possibly analogous to a normal hydrogen-rich su

  41. Rafał Karczewski, Markus Heinonen, Alison Pouplin, Søren Hauberg

    We present a novel geometric perspective on the latent space of diffusion models. We first show that the standard pullback approach, utilizing the deterministic probability flow ODE decoder, is fundamentally flawed. It provably forces geodesics to decode as straight segments in data space, effectively ignoring any intrinsic data geometry beyond the ambient E

  42. P. A. Russkikh, G. Sh. Boltachev

    The possibility of significant increase of generated pulsed magnetic fields by the inductor system of a single-turn solenoid and magnetic flux concentrator without initialization of low-cycle fatigue mechanism is theoretically studied by varying the size of the inductor system, the material of the concentrator and the parameters of the discharge circuit. The

  43. Suraina Gupta, Santu Prasad Jana, Pawan Kumar Gupta, Anjan K. Gupta

    We report a super-insulating behavior, in a device having granular Pb film on back-gated few-layer $\mathrm{MoS_2}$, below an onset temperature same as the critical temperature $T_{\rm C}\approx7$ K of bulk Pb. Below $T_{\rm C}$, the current-voltage characteristics exhibit a threshold voltage marking a crossover between the low-bias insulating and the high-b

  44. Daniel Frolovsky, Sergei V. Ketov

    The data release from the Atacama Cosmology Telescope (ACT) imposes stronger constraints on primordial black holes (PBHs) formation in single-field inflation models versus the Planck data. In particular, the updated Cosmic Microwave Background (CMB) radiation measurements favour a {\it higher} scalar spectral index $n_s$ and its {\it positive} running $\alph

  45. Binh Nguyen, Shuji Shi, Ryan Ofman, Thai Le

    Recent advances in text-to-speech technologies have enabled realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio anti-spoofing systems are critical for detecting such threats, prior work has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexpl

  46. Shuhang Xu, Weijian Deng, Yixuan Zhou, Fangwei Zhong

    Concepts serve as fundamental abstractions that support human reasoning and categorization. However, it remains unclear whether large language models truly capture such conceptual structures or primarily rely on surface-level pattern memorization. Existing benchmarks are largely static and fact oriented, which limits their ability to probe fine-grained seman

  47. Aditya Gautam

    The rapid proliferation of misinformation in digital media demands solutions that go beyond isolated Large Language Model(LLM) or AI Agent based detection methods. This paper introduces a novel multi-agent framework that covers the complete misinformation lifecycle: classification, detection, correction, and source verification to deliver more transparent an

  48. Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason

    Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the models at each label generation step. We find that the initial p

  49. Shiji Zhao, Qihui Zhu, Shukun Xiong, Shouwei Ruan

    Large pre-trained Vision Language Models (VLMs) demonstrate excellent generalization capabilities but remain highly susceptible to adversarial examples, posing potential security risks. To improve the robustness of VLMs against adversarial examples, adversarial prompt tuning methods are proposed to align the text feature with the adversarial image feature wi

  50. Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Yang Yuan

    Policy gradient algorithms have been successfully applied to enhance the reasoning capabilities of large language models (LLMs). KL regularization is ubiquitous, yet the design surface, choice of KL direction (forward vs. reverse), normalization (normalized vs. unnormalized), and estimator ($k_1/k_2/k_3$), is scattered across the literature and often intertw

  51. Qiaosheng Chen, Kaijia Huang, Xiao Zhou, Weiqing Luo

    The rapid growth of open source machine learning (ML) resources, such as models and datasets, has accelerated IR research. However, existing platforms like Hugging Face do not explicitly utilize structured representations, limiting advanced queries and analyses such as tracing model evolution and recommending relevant datasets. To fill the gap, we construct

  52. Seokmin Ko, Ambuj Tewari, Kihyuk Hong

    We study offline constrained reinforcement learning with general function approximation in discounted constrained Markov decision processes. Prior methods either require full data coverage for evaluating intermediate policies, lack oracle efficiency, or requires the knowledge of data-generating distribution for policy extraction. We propose PDOCRL, an oracle

  53. Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang

    Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained in both contextual coverage and inference efficiency due to its inherently sequential process. To overcome these challenges, we propose leap multi-token prediction~(L-MTP), an inno

  54. Jun Li, Lingsheng Meng

    In this research, to solve the large indefinite least squares problem, we firstly transform its normal equation into a sparse block three-by-three linear systems, then use GMRES method with an accelerated preconditioner to solve it. The construction idea of the preconditioner comes from the thought of Luo et.al [Luo, WH., Gu, XM., Carpentieri, B., BIT 62, 19

  55. Minsoo Khang, Sangjun Park, Teakgyu Hong, Dawoon Jung

    Large Language Models (LLMs) have made substantial progress in recent years, yet evaluating their capabilities in practical Retrieval-Augmented Generation (RAG) scenarios remains challenging. In practical applications, LLMs must demonstrate complex reasoning, refuse to answer appropriately, provide precise citations, and effectively understand document layou

  56. Konstantinos Gkouliaras, Vasileios Theos, True Miller, Brian Jowers

    Quantum key distribution (QKD), one of the latest cryptographic techniques, founded on the laws of quantum mechanics rather than mathematical complexity, promises for the first time unconditional secure remote communications. Integrating this technology into the next generation nuclear systems - designed for universal data collection and real-time sharing as

  57. Yuehan Jin, Xiaoqing Liu, Yiyuan Yang, Zhiwen Yu

    Multimodal emotion recognition analyzes emotions by combining data from multiple sources. However, real-world noise or sensor failures often cause missing or corrupted data, creating the Incomplete Multimodal Emotion Recognition (IMER) challenge. In this paper, we propose Robust Hybrid Diffusion Recovery (RoHyDR), a novel framework that performs missing-moda

  58. Vladimir Baulin, Austin Cook, Daniel Friedman, Janna Lumiruusu

    The prevailing model for disseminating scientific knowledge relies on individual publications dispersed across numerous journals and archives. This legacy system is ill suited to the recent exponential proliferation of publications, contributing to insurmountable information overload, issues surrounding reproducibility and retractions. We introduce the Disco

  59. Tianxiang Dai, Yixuan Shao, Chenkai Mao, Yu Wu

    Nanophotonic freeform design has the potential to push the performance of optical components to new limits, but there remains a challenge to effectively perform optimization while reliably enforcing design and manufacturing constraints. We present Neuroshaper, a framework for freeform geometric parameterization in which nanophotonic device layouts are define

  60. Marco Brandizi, Carlos Bobed, Luca Garulli, Arné de Klerk

    Linked Data and labelled property graphs (LPG) are two data management approaches with complementary strengths and weaknesses, making their integration beneficial for sharing datasets and supporting software ecosystems. In this paper, we introduce rdf2pg, an extensible framework for mapping RDF data to semantically equivalent LPG formats and data-bases. Util

  61. D. R. Scott, T. Dial, A. Bera, A. T. Deller

    We present microsecond-resolution, coherently-dedispersed, polarimetric measurements of 35 fast radio bursts (FRBs) detected during the Commensal Real-time ASKAP Fast Transients (CRAFT) incoherent sum (ICS) survey with the Australian Square Kilometre Array Pathfinder (ASKAP). We find a wide diversity of time-frequency morphology and polarisation properties b

  62. Chi-Yuan Hsiao, Ke-Han Lu, Kai-Wei Chang, Chih-Kai Yang

    End-to-end training of Spoken Language Models (SLMs) commonly involves adapting pre-trained text-based Large Language Models (LLMs) to the speech modality through multi-stage training on diverse tasks such as ASR, TTS and spoken question answering (SQA). Although this multi-stage continual learning equips LLMs with both speech understanding and generation ca

  63. Junyan Zhang, Yiming Huang, Shuliang Liu, Yubo Gao

    The rapid adoption of LLMs has overshadowed the potential advantages of traditional BERT-like models in text classification. This study challenges the prevailing "LLM-centric" trend by systematically comparing three category methods, i.e., BERT-like models fine-tuning, LLM internal state utilization, and zero-shot inference across six high-difficulty dataset

  64. Landon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas

    Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features. To identify these interactions, most existing approaches require enumerating all possible combinations of features up to a given order, causing them to scale poorly with the number of inputs $n$. Recently, Kang et al. (2025) proposed SP

  65. Haruka Matsumoto, Hiroto Isomura, Keita Kojima, Ryutaro Okuma

    We report thermoelectric properties of sintered samples of undoped, W-doped, and Sb-doped ReSTe crystallized in a cubic MoSBr-type structure. All samples exhibited p-type thermoelectric properties. ReSTe and Re0.993W0.007STe exhibited the largest dimensionless figure of merit ZT, reaching 0.4 at 660 K. This high performance is attributed to large power facto

  66. Jingde Huang, Zhangyu Huang, Chenyu Li, Jiantong Liu

    The motor control board has various defects such as inconsistent color differences, incorrect plug-in positions, solder short circuits, and more. These defects directly affect the performance and stability of the motor control board, thereby having a negative impact on product quality. Therefore, studying the defect detection technology of the motor control

  67. Reza Marzban, Hamed Abiri, Raphael Pestourie, Ali Adibi

    HiLAB (Hybrid inverse-design with Latent-space learning, Adjoint-based partial optimizations, and Bayesian optimization) is a new paradigm for inverse design of nanophotonic structures. Combining early-terminated topological optimization (TO) with a Vision Transformer-based variational autoencoder (VAE) and a Bayesian search, HiLAB addresses multi-functional

  68. Haotian Liu, Yuchuang Tong, Zhengtao Zhang

    In physical Human-Robot Collaboration (pHRC), accurate human intent estimation and rational human-robot role allocation are crucial for safe and efficient assistance. Existing methods that rely on short-term motion data for intention estimation lack multi-step prediction capabilities, hindering their ability to sense intent changes and adjust human-robot ass

  69. Amin Pishehvar, Zixin Yan, Zhaoyou Wang, Yu Jiang

    Floquet engineering has been recently recognized as an important tool for manipulating the coherent magnon-photon interaction in cavity electromagnonics systems at microwave frequencies. In spite of the novel hybrid magnonic functionalities that have been demonstrated, the effect of the Floquet drive has been relatively weak due to the limited driving effici

  70. Haoran Li, Muhao Guo, Yang Weng, Marija Ilic

    Non-stationary power system dynamics, influenced by renewable energy variability, evolving demand patterns, and climate change, are becoming increasingly complex. Accurately capturing these dynamics requires a model capable of adapting to environmental factors. Traditional models, including Recurrent Neural Networks (RNNs), lack efficient mechanisms to encod

  71. Yue Xiao, Yi He, Yaqing Zhang, Xin Lin

    Under extreme conditions, autonomous drifting enables vehicles to follow predefined paths at large slip angles, significantly enhancing the control system's capability to handle hazardous scenarios. Four-wheel-drive and four-wheel-steering (4WD-4WS) vehicles, which have been extensively studied, offer superior path-following precision and enhanced maneuverab

  72. Hirotaka Tashiro

    Based on the analogies of arithmetic topology, we show a topological analogue of Hilbert's Satz 90 for idele groups and utilize our previously established Hasse norm principle to present a proof of an Iyanaga--Tamagawa type genus formula for finite abelian branched covers over integral homology 3-spheres in a very parallel manner to the case of number theory

  73. Saketh Reddy Vemula, Parameswari Krishnamurthy

    Identification of hallucination spans in black-box language model generated text is essential for applications in the real world. A recent attempt at this direction is SemEval-2025 Task 3, Mu-SHROOM-a Multilingual Shared Task on Hallucinations and Related Observable Over-generation Errors. In this work, we present our solution to this problem, which capitali

  74. Hai Jiang, Qiongting Liu, Yuanpin Zhou, Jiawei Pan

    Placenta Accreta Spectrum Disorders (PAS) pose significant risks during pregnancy, frequently leading to postpartum hemorrhage during cesarean deliveries and other severe clinical complications, with bleeding severity correlating to the degree of placental invasion. Consequently, accurate prenatal diagnosis of PAS and its subtypes-placenta accreta (PA), plac

  75. Yiqing Guo, Nagur Cherukuru, Eric Lehmann, S. L. Kesav Unnithan

    Nitrate ($\text{NO}_3^-$) is a form of dissolved inorganic nitrogen derived primarily from anthropogenic sources. The recent increase in river-discharged nitrate poses a major risk for coral bleaching in the Great Barrier Reef (GBR) lagoon. Although nitrate is an optically inactive (i.e., colourless) constituent, previous studies have demonstrated there is a

  76. Chao Lei, Nir Lipovetzky, Krista A. Ehinger, Yanchuan Chang

    Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus

  77. Yusheng Zhao, Xiao Luo, Weizhi Zhang, Wei Ju

    The ability to reason is one of the most fundamental capabilities of large language models (LLMs), enabling a wide range of downstream tasks through sophisticated problem-solving. A critical aspect of this is code reasoning, which involves logical reasoning with formal languages (i.e., programming code). In this paper, we enhance this capability of LLMs by e

  78. Faruk Alpay

    In this second installment of the Alpay Algebra framework, I formally define identity as a fixed point that emerges through categorical recursion. Building upon the transfinite operator $\varphi^\infty$, I characterize identity as the universal solution to a self-referential functorial equation over a small cartesian closed category. I prove the existence an

  79. Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu

    Test-time compute has empowered multimodal large language models to generate extended reasoning chains, yielding strong performance on tasks such as multimodal math reasoning. However, this improved reasoning ability often comes with increased hallucination: as generations become longer, models tend to drift away from image-grounded content and rely more hea

  80. Olivier Toubia, George Z. Gui, Tianyi Peng, Daniel J. Merlau

    LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-qua

  81. Yuning Shen, Lihao Wang, Huizhuo Yuan, Yan Wang

    Understanding protein dynamics is critical for elucidating their biological functions. The increasing availability of molecular dynamics (MD) data enables the training of deep generative models to efficiently explore the conformational space of proteins. However, existing approaches either fail to explicitly capture the temporal dependencies between conforma

  82. Victor OK Li, Yang Han, Jacqueline CK Lam, Lawrence YL Cheung

    This study introduces Reverse-Speech-Finder (RSF), a groundbreaking neural network backtracking architecture designed to enhance Alzheimer's Disease (AD) diagnosis through speech analysis. Leveraging the power of pre-trained large language models, RSF identifies and utilizes the most probable AD-specific speech markers, addressing both the scarcity of real A

  83. Yuchen Zhang, Yaxiong Wang, Yujiao Wu, Lianwei Wu

    The detection and grounding of multimedia manipulation has emerged as a critical challenge in combating AI-generated disinformation. While existing methods have made progress in recent years, we identify two fundamental limitations in current approaches: (1) Underestimation of MLLM-driven deception risk: prevailing techniques primarily address rule-based tex

  84. Uyoung Jeong, Jonathan Freer, Seungryul Baek, Hyung Jin Chang

    We study multi-dataset training (MDT) for pose estimation, where skeletal heterogeneity presents a unique challenge that existing methods have yet to address. In traditional domains, \eg regression and classification, MDT typically relies on dataset merging or multi-head supervision. However, the diversity of skeleton types and limited cross-dataset supervis

  85. Cam McLeman, Christopher Rasmussen

    Heavenly abelian varieties are abelian varieties defined over number fields that exhibit constrained $\ell$-adic Galois representations for some rational prime $\ell$. At the ICMS Workshop held in November 2024, we presented evidence for two finiteness conjectures around the distribution of heavenly elliptic curves over quadratic number fields, one in terms

  86. Jiangning Zhu, Yuxing Zhou, Zheng Wang, Juntao Yao

    Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies in their inaccurate visual grounding of infographic elements, including charts and human-recognizable objects (HROs) such

  87. Jiangjie Wu, Lixuan Chen, Zhenghao Li, Xin Li

    High-quality 3D fetal brain MRI reconstruction from motion-corrupted 2D slices is crucial for clinical diagnosis. Reliable slice-to-volume registration (SVR)-based motion correction and super-resolution reconstruction (SRR) methods are essential. Deep learning (DL) has demonstrated potential in enhancing SVR and SRR when compared to conventional methods. How

  88. Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao

    Retrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of k

  89. Xiang Liu, Zhaoxiang Liu, Peng Wang, Kohou Wang

    When using supervised fine-tuning (SFT) to adapt large language models (LLMs) to specific domains, a significant challenge arises: should we use the entire SFT dataset for fine-tuning? Common practice often involves fine-tuning directly on the entire dataset due to limited information on the LLM's past training data. However, if the SFT dataset largely overl

  90. Miruna Oprescu, Brian M Cho, Nathan Kallus

    We study the problem of estimating the average treatment effect (ATE) in adaptive experiments where treatment can only be encouraged -- rather than directly assigned -- via a binary instrumental variable. Building on semiparametric efficiency theory, we derive the efficiency bound for ATE estimation under arbitrary, history-dependent instrument-assignment po

  91. Chengrui Zhou, Yuandeng Shen, Chun Xia, Hao Liang

    Magnetic flux emergence is traditionally considered a key trigger of solar filament eruptions; however, its role in suppressing filament eruptions remains less understood. Using multi-wavelength observations from the Solar Dynamics Observatory, this study investigates a unique case of flux emergence below a quiescent filament from January 3 to 5, 2016, where

  92. Junru Lin, Mingzhe Liu, Songze Li, Xuechao Wang

    Recent years have witnessed a rapid development of platform economy, as it effectively addresses the trust dilemma between untrusted online buyers and merchants. However, malicious platforms can misuse users' funds and information, causing severe security concerns. Previous research efforts aimed at enhancing security in platform payment systems often sacrif

  93. Roelien C Timmer, Yufang Hou, Stephen Wan

    An important task in machine learning (ML) research is comparing prior work, which is often performed via ML leaderboards: a tabular overview of experiments with comparable conditions (e.g., same task, dataset, and metric). However, the growing volume of literature creates challenges in creating and maintaining these leaderboards. To ease this burden, resear

  94. Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu

    Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating external knowledge. Current hybrid RAG system retrieves evidence from both knowledge graphs (KGs) and text documents to support LLM reasoning. However, it faces challenges like handling multi-hop reasoning, multi-entity questions, multi-source verification, and effect

  95. Jiaming Liang, Renato D. C. Monteiro, Honghao Zhang

    This paper considers the stochastic convex composite optimization problem and presents multi-cut stochastic approximation (SA) methods for solving it, whose models in expectation overestimate its objective function. The multi-cut model obtained by taking the maximum of a finite number of linearizations of the stochastic objective function provides a biased e

  96. Davrbek Oltiboev

    In this note we consider a tent-like family with a cusp at the singular point and show that the linear response holds for certain perturbations of this family. This contrasts the tent-like maps with finite derivatives at the singularity. Our results extend the results of Bahsoun and Galatolo to the larger class of singularities and we obtain the linear respo

  97. Kazuki Hayashi, Shintaro Ozaki, Yusuke Sakai, Hidetaka Kamigaito

    Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. As capabilities improve, applications in navigation, education, and accessibility are becoming practical. These settings require accommodation of perceptual variation rather than assuming a uniform visual experience. Color perception illustrate

  98. Xinran Zheng, Xingzhi Qian, Huichi Zhou, Shuo Yang

    Language models (LMs) show promise for vulnerability detection but struggle with long, real-world code due to sparse and uncertain vulnerability locations. These issues, exacerbated by token limits, often cause models to miss vulnerability-related signals, thereby impairing effective learning. A key intuition is to enhance LMs with concise, information-rich

  99. Jingwen Cheng, Ruikun Li, Huandong Wang, Yong Li

    Predicting the behavior of complex systems is critical in many scientific and engineering domains, and hinges on the model's ability to capture their underlying dynamics. Existing methods encode the intrinsic dynamics of high-dimensional observations through latent representations and predict autoregressively. However, these latent representations lose the i

  100. Guiquan Sun, Xikun Zhang, Jingchao Ni, Dongjin Song

    Heterogeneous graph neural networks have seen rapid progress in web applications such as social networks, knowledge graphs, and recommendation systems, driven by the inherent heterogeneity of web data. However, existing methods typically assume static graphs, while real-world graphs are continuously evolving. This dynamic nature requires models to adapt to n