Skip to content

May 2025 arXiv papers — page 68

Showing 6,7016,800 of 24,552 papers

  1. Guillaume Ballif, Laurent Pfeiffer, Jakob Ruess

    In this work, we present a general method to establish properties of multi-dimensional continuous-time Markov chains representing stochastic reaction networks. This method consists of grouping states together (via a partition of the state space), then constructing two one-dimensional birth and death processes that lower and upper bound the initial process un

  2. Eunjin Roh, Yigitcan Kaya, Christopher Kruegel, Giovanni Vigna

    We present MADCAT, a self-supervised approach designed to address the concept drift problem in malware detection. MADCAT employs an encoder-decoder architecture and works by test-time training of the encoder on a small, balanced subset of the test-time data using a self-supervised objective. During test-time training, the model learns features that are usefu

  3. N. A. Moraga, F. Castillo, D. D. Ofengeim, A. Reisenegger

    The high quiescent X-ray luminosity observed in some magnetars is widely attributed to the decay and evolution of their ultra-strong magnetic fields. Several dissipation mechanisms have been proposed, each operating with different efficiencies depending on the region of the star. In this context, ambipolar diffusion, i.e., the relative motion of charged part

  4. Jiaming Hu, Jiawei Wang, Henrik I Christensen

    Efficient tabletop rearrangement planning seeks to find high-quality solutions while minimizing total cost. However, the task is challenging due to object dependencies and limited buffer space for temporary placements. The complexity increases for mobile robots, which must navigate around the table with restricted access. A*-based methods yield high-quality

  5. Wei Shen, Xiaonan He, Chuheng Zhang, Xuyun Zhang

    Reward-driven proactive dialogue agents require precise estimation of user satisfaction as an intrinsic reward signal to determine optimal interaction strategies. Specifically, this framework triggers clarification questions when detecting potential user dissatisfaction during interactions in the industrial dialogue system. Traditional works typically rely o

  6. Wenchao Zhang, Jiahe Tian, Runze He, Jizhong Han

    Recent text-to-image (T2I) generation models have advanced significantly, enabling the creation of high-fidelity images from textual prompts. However, existing evaluation benchmarks primarily focus on the explicit alignment between generated images and prompts, neglecting the alignment with real-world knowledge beyond prompts. To address this gap, we introdu

  7. Jiajun Hu, Jian Xiao

    For $(n-2)$ free divisor classes on a smooth projective variety of dimension $n$, the product of these free divisor classes induces a Lefschetz type operator acting on the N\'{e}ron-Severi space or the cohomology group of $(1,1)$ classes. We give a characterization of this kernel space, when the collection of these free divisor classes is supercritical. This

  8. Xiaohe Li, Pengfei Li, Zide Fan, Ying Geng

    Multi-view multi-object tracking (MVMOT) has found widespread applications in intelligent transportation, surveillance systems, and urban management. However, existing studies rarely address genuinely free-viewpoint MVMOT systems, which could significantly enhance the flexibility and scalability of cooperative tracking systems. To bridge this gap, we first c

  9. Ryohei Miyadera, Enchong Li, Akito Tsujii

    We define a variant of the two-dimensional Silver Dollar game. Two coins are placed on a chessboard of unbounded size, and two players take turns choosing one of the coins and moving it. Coins are to be moved to the left or upward vertically as far as desired. If a coin is dropped off the board, players cannot use this coin. Jumping a coin over another coin

  10. Mahmudul Hasan

    Breast cancer is the most commonly occurring cancer worldwide. This cancer caused 670,000 deaths globally in 2022, as reported by the WHO. Yet since health officials began routine mammography screening in age groups deemed at risk in the 1980s, breast cancer mortality has decreased by 40% in high-income nations. Every day, a greater and greater number of peo

  11. Junyu Chen, Junzhuo Li, Zhen Peng, Wenjie Wang

    Quantization and fine-tuning are crucial for deploying large language models (LLMs) on resource-constrained edge devices. However, fine-tuning quantized models presents significant challenges, primarily stemming from: First, the mismatch in data types between the low-precision quantized weights (e.g., 4-bit) and the high-precision adaptation weights (e.g., 1

  12. Terry Yi Zhong, Esther Janse, Cristian Tejedor-Garcia, Louis ten Bosch

    Speech-based Parkinson's disease (PD) detection has gained attention for its automated, cost-effective, and non-intrusive nature. As research studies usually rely on data from diagnostic-oriented speech tasks, this work explores the feasibility of diagnosing PD on the basis of speech data not originally intended for diagnostic purposes, using the Turn-Taking

  13. Harshit Raj, Sanjeev Dhurandhar, Massimo Tinto

    We quantify the advantages of a recently proposed data processing technique to search for continuous gravitational wave (GW) signals from isolated rotating asymmetric neutron stars in data measured by ground-based GW interferometers. This technique relies on the symmetry of the motion around the Sun of an Earth-bound gravitational wave interferometer. By mul

  14. Meng Li, Guangda Huzhang, Haibo Zhang, Xiting Wang

    Direct Preference Optimization (DPO) has emerged as a promising framework for aligning Large Language Models (LLMs) with human preferences by directly optimizing the log-likelihood difference between chosen and rejected responses. However, existing methods assign equal importance to all tokens in the response, while humans focus on more meaningful parts. Thi

  15. Guanxing Lu, Wenkai Guo, Chubin Zhang, Yuheng Zhou

    Recent high-capacity vision-language-action (VLA) models have demonstrated impressive performance on a range of robotic manipulation tasks by imitating human demonstrations. However, exploiting offline data with limited visited states will cause execution failure in out-of-distribution scenarios. Intuitively, an exploration-based method that improves on onli

  16. Botao Amber Hu, Helena Rong

    In Artificial Life (ALife) research, replicating Open-Ended Evolution (OEE)-the continuous emergence of novelty observed in biological life-has usually been pursued within isolated, closed system simulations, such as Tierra and Avida, which have typically plateaued after an initial burst of novelty, failing to achieve sustained OEE. Scholars suggest that OEE

  17. Hassan Nagib

    Following recent evidence that even ZPG boundary layers do not exhibit a purely logarithmic extended overlap region, reconsideration of recently advanced logarithmic plus linear extended overlap region in wall-bounded flows leads to a revision of the model for the extended overlap region. The significant difference between the two representations is a separa

  18. Ke Huang, Yue Zhou, Xi He, Weibo Chen

    Cybroc is a series of kinetic art installations exploring the recent proliferating populist longevity activism through the satirical cyborgization of broccoli. The artwork augments the symbol of health food-broccoli-with prosthetic limbs to perform so-called longevity-enhancing exercises such as cold plunges, treadmill running, brachiation (arm-swinging), sl

  19. Hui-Min Yang, Xuan Luo, Hua-Xing Chen, Wei Chen

    We investigate charmed hybrid baryons using the QCD sum rule method within the framework of heavy quark effective theory. We construct twenty-eight interpolating currents for charmed hybrid baryons, seven of which are employed in QCD sum rule analyses of nineteen states with quark-gluon configurations $qqcg$, $qscg$, and $sscg$ ($q = u/d$). The masses of the

  20. Igor Chagas Santos

    In this paper, we classify the generic singularities of 2-parameter plane congruences in $\mathbb{R^4}$ and the generic singularities of affine normal plane congruences. We also study the generic singularities of the family of affine distance functions.

  21. Quentin Changeat, Deborah Bardet, Katy Chubb, Achrene Dyrek

    Context: Before JWST, telescope observations were not sensitive enough to constrain the nature of clouds in exo-atmospheres. Recent observations, however, have inferred cloud signatures as well as haze-enhanced scattering slopes motivating the need for modern inversion techniques and a deeper understanding of the JWST information content. Aims: We aim to inv

  22. Hongyu Cao, Junjie Lu, Xuewei Zhang, Yulin Hui

    Off-road navigation remains challenging for autonomous robots due to the harsh terrain and clustered obstacles. In this letter, we extend the YOPO (You Only Plan Once) end-to-end navigation framework to off-road environments, explicitly focusing on forest terrains, consisting of a high-performance, multi-sensor supported off-road simulator YOPO-Sim, a zero-s

  23. Ziming Wang, Nan Xue, Rebecka Jörnsten

    The goal of point cloud assembly is to reconstruct a complete 3D shape by aligning multiple point cloud pieces. This work presents a novel equivariant solver for assembly tasks based on flow matching models. We first theoretically show that the key to learning equivariant distributions via flow matching is to learn related vector fields. Based on this result

  24. Guodong Du, Zitao Fang, Jing Li, Junlin Li

    Foundation models and their checkpoints have significantly advanced deep learning, boosting performance across various applications. However, fine-tuned models often struggle outside their specific domains and exhibit considerable redundancy. Recent studies suggest that combining a pruned fine-tuned model with the original pre-trained model can mitigate forg

  25. Zihan Weng, Lucas Gomez, Taylor Whittington Webb, Pouya Bashivan

    Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years. Yet, they still lag behind human capabilities in specific visual tasks such as counting or relational reasoning. To understand the underlying limitations, we adopt methodologies from cognitive science, analyzing VLM performance along core cognitive axes: Per

  26. Martin Čech, Lucile Devin, Daniel Fiorilli, Kaisa Matomäki

    We study the one-level density of low-lying zeros in the family of Maass form $L$-functions of prime level $N$ tending to infinity. Generalizing the influential work of Iwaniec, Luo and Sarnak to this context, Alpoge et al. have proven the Katz-Sarnak prediction for test functions whose Fourier transform is supported in $(-\frac32,\frac32)$. In this paper, w

  27. Yukun Zhang, Qi Dong, Mengkang Li

    Understanding how latent representations evolve during generation is a central open problem in large language model interpretability. We introduce \textbf{Dynamical Manifold Evolution Theory} (DMET), a phenomenological framework that models LLM generation as a controlled dynamical system evolving along a trajectory on a low-dimensional semantic manifold. DME

  28. Shi Jin, Chundan Zhang

    In this paper we study quantum simulation algorithms on the elastic wave equations using the Schr\"odingerisation method. The Schr\"odingerisation method transforms any linear PDEs into a system of Schr\"odinger-type PDEs -with unitary evolution-using the warped phase transformation that maps the equations in one higher dimension. This makes them suitable fo

  29. Yi Jiang, Sendong Zhao, Jianbo Li, Haochun Wang

    The Retrieval-Augmented Generation (RAG) framework introduces a retrieval module to dynamically inject retrieved information into the input context of large language models (LLMs), and has demonstrated significant success in various NLP tasks. However, the current study points out that there is a preference gap between retrievers and LLMs in the RAG framewor

  30. Sourav Kumar Das, Md. Julkar Naeen, MD. Jahidul Islam, Md. Anisul Haque Sajeeb

    Bangla or Bengali is the national language of Bangladesh, people from different regions don't talk in proper Bangla. Every division of Bangladesh has its own local language like Sylheti, Chittagong etc. In recent years some papers were published on Bangla language like sentiment analysis, fake news detection and classifications, but a few of them were on Ban

  31. Oluwaseyi Giwa

    Dynamic causal discovery in wireless networks is essential due to evolving interference, fading, and mobility, which complicate traditional static causal models. This paper addresses causal inference challenges in dynamic fading wireless environments by proposing a sequential regression-based algorithm with a novel application of the NOTEARS acyclicity const

  32. Xu Zhang, Kun Zhang, Wenxin Ma, Rongsheng Wang

    ICD Coding aims to assign a wide range of medical codes to a medical text document, which is a popular and challenging task in the healthcare domain. To alleviate the problems of long-tail distribution and the lack of annotations of code-specific evidence, many previous works have proposed incorporating code knowledge to improve coding performance. However,

  33. Ali Khalifa, Michael Breuer

    In this study, agglomerate breakage in homogeneous isotropic turbulence is investigated using particle-resolved direct numerical simulations. Single agglomerates composed of 500 monodisperse spherical particles are considered, and their interaction with the turbulent flow is resolved through an immersed boundary method coupled with a soft-sphere discrete ele

  34. Viacheslav Sinii, Alexey Gorbatovski, Artem Cherepanov, Boris Shaposhnikov

    We show that training a single $d$-dimensional steering vector per layer with reinforcement learning, while freezing all base weights, matches the accuracy of fully RL-tuned reasoning models on mathematical-reasoning tasks. On an 8 billion-parameter model this adds only $\approx 0.0016\%$ additional parameters and reproduces performance across a range of bas

  35. Jiabin Tang, Lianghao Xia, Zhonghang Li, Chao Huang

    The powerful reasoning capabilities of Large Language Models (LLMs) in mathematics and coding, combined with their ability to automate complex tasks through agentic frameworks, present unprecedented opportunities for accelerating scientific innovation. In this paper, we introduce AI-Researcher, a fully autonomous research system that transforms how AI-driven

  36. Anton Lipin

    Suppose $X$ and $Y$ are topological spaces, $|X| = \Delta(X)$ and $|Y| = \Delta(Y)$. We investigate resolvability of the product $X \times Y$. We prove that: I. If $|X| = |Y| = \omega$ and $X,Y$ are Hausdorff, then $X \times Y$ is maximally resolvable; II. If $2^\kappa = \kappa^+$, $\{|X|, \mathrm{cf}|X|\} \cap \{\kappa, \kappa^+\} \ne \emptyset$ and $\mathr

  37. Gaurav Negi, Dhairya Dalal, Omnia Zayed, Paul Buitelaar

    This paper introduces the Unified Opinion Concepts (UOC) ontology to integrate opinions within their semantic context. The UOC ontology bridges the gap between the semantic representation of opinion across different formulations. It is a unified conceptualisation based on the facets of opinions studied extensively in NLP and semantic structures described thr

  38. Xiangcun Meng

    Eccentric millisecond pulsar + helium white dwarf (MSP + He WD) systems have attracted increasing attention, with the rotationally delayed accretion-induced collapse (RD-AIC) scenario proposed as a possible formation channel. Given the similarity between the formation channels of He WDs and subdwarf B (sdB) stars, eccentric MSP + sdB binaries could also exis

  39. Antoni Gomila, Vincent C. Müller

    The declared goal of this paper is to fill this gap: "... cognitive systems research needs questions or challenges that define progress. The challenges are not (yet more) predictions of the future, but a guideline to what are the aims and what would constitute progress." -- the quotation being from the project description of EUCogII, the project for the Euro

  40. Sebastian Gherghe, Iván Moyano, Israel Michael Sigal

    In this paper, we consider the time-dependent Born-Oppenheimer approximation (BOA) of a classical quantum molecule involving a possibly large number of nuclei and electrons, described by a Schr\"odinger equation. In the spirit of Born and Oppenheimer's original idea we study quantitatively the approximation of the molecular evolution. We obtain an iterable a

  41. Chun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan

    Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requiring the extraction of multigranular visual cues from images and their integration with external world knowledge for systematic reasoning. Current approaches to geo-localization tasks

  42. Peixuan Zhang, Shuchen Weng, Chengxuan Zhu, Binghao Tang

    In daily life, images as common affective stimuli have widespread applications. Despite significant progress in text-driven image editing, there is limited work focusing on understanding users' emotional requests. In this paper, we introduce AIEdiT for Affective Image Editing using Text descriptions, which evokes specific emotions by adaptively shaping multi

  43. Can Yaras, Alec S. Xu, Pierre Abillama, Changwoo Lee

    Transformers have achieved state-of-the-art performance across various tasks, but suffer from a notable quadratic complexity in sequence length due to the attention mechanism. In this work, we propose MonarchAttention -- a novel approach to sub-quadratic attention approximation via Monarch matrices, an expressive class of structured matrices. Based on the va

  44. Ziyang Cheng, Zhixun Li, Yuhan Li, Yixin Song

    Nowadays, real-world data, including graph-structure data, often arrives in a streaming manner, which means that learning systems need to continuously acquire new knowledge without forgetting previously learned information. Although substantial existing works attempt to address catastrophic forgetting in graph machine learning, they are all based on training

  45. V. K. Suman, T. K. Sengupta

    The role of round-off errors on the receptivity and instability of fluid flows are conclusively established for the first time using high accuracy simulations of the benchmark two-dimensional (2D) Taylor-Green vortex problem using double and quadruple precisions. Employing the fourth order Runge-Kutta (RK4) method for temporal discretization and Fourier pseu

  46. Yu Han, Aaron Ceross, Jeroen H. M. Bergmann

    Regulatory affairs, which sits at the intersection of medicine and law, can benefit significantly from AI-enabled automation. Classification task is the initial step in which manufacturers position their products to regulatory authorities, and it plays a critical role in determining market access, regulatory scrutiny, and ultimately, patient safety. In this

  47. Cayo Viegas, Rohit Gheyi, Márcio Ribeiro

    Recent advancements in Large Language Models (LLMs) have significantly expanded the capabilities of artificial intelligence in natural language processing tasks. Despite this progress, their performance in specialized domains such as computer science remains relatively unexplored. Understanding the proficiency of LLMs in these domains is critical for evaluat

  48. Rafiu Adekoya Badekale, Adewale Akinfaderin

    Climate policy scenario generation and evaluation have traditionally relied on integrated assessment models (IAMs) and expert-driven qualitative analysis. These methods enable stakeholders, such as policymakers and researchers, to anticipate impacts, plan governance strategies, and develop mitigation measures. However, traditional methods are often time-inte

  49. Ihtesham Ibn Malek, Hafiz Imtiaz, Samia Subrina

    Perovskite solar cells (PSCs) without a hole transport layer (HTL) offer a cost-effective and stable alternative to conventional architectures, utilizing only an absorber layer and an electron transport layer (ETL). This study presents a machine learning (ML)-driven framework to optimize the efficiency and stability of HTL-free PSCs by integrating experiment

  50. Alexander Flamant, Bram Mesland, Adam Rennie

    We compare the constructions of Levi-Civita connections for noncommutative algebras developed in arXiv:1505.07330, arXiv:1809.06721, arXiv:2403.13735. The assumptions in these various constructions differ, but when they are all defined, we provide direct translations between them. An essential assumption is that the (indefinite) Hermitian inner product on di

  51. Zhenyu Wei, Zhijiang Shao, Lorenz T. Biegler

    Multiple parafoil landing is an enabling technology for massive supply delivery missions. However, it is still an open question to design a collision-free, computation-efficient guidance and control method for unpowered parafoils. To address this issue, this paper proposes a coordinated guidance and control method for multiple parafoil landing. First, the mu

  52. Guoxiu He, Xin Song, Futing Wang, Aixin Sun

    Knowledge editing aims to update the embedded knowledge within Large Language Models (LLMs). However, existing approaches, whether through parameter modification or external memory integration, often suffer from inconsistent evaluation objectives and experimental setups. To address this gap, we conduct a comprehensive benchmarking study. In addition to fact-

  53. Moldir Seidaliyeva, Victor Denisov, Irene Denisova

    Within the framework of the parameterized post-Maxwellian vacuum electrodynamics, the propagation of an X-ray or gamma-ray pulse through the electromagnetic field of a relativistically rotating pulsar is studied. Expressions are obtained for the trajectory of this pulse and the law The effect of electromagnetic radiation birefringence in the field of a relat

  54. Aleksandr Tsymbalov, Mikhail Khovrichev

    Machine learning models for text classification are trained to predict a class for a given text. To do this, training and validation samples must be prepared: a set of texts is collected, and each text is assigned a class. These classes are usually assigned by human annotators with different expertise levels, depending on the specific classification task. Co

  55. Xin Wang, Han-Xiao Tao, Re-Bing Wu

    Quantum machine learning models incorporating data re-uploading circuits have garnered significant attention due to their exceptional expressivity and trainability. However, their ability to generate accurate predictions on unseen data, referred to as the predictive performance, remains insufficiently investigated. This study reveals a fundamental limitation

  56. Yang Liu, Silin Cheng, Xinwei He, Sebastien Ourselin

    Weakly supervised referring expression comprehension(WREC) and segmentation(WRES) aim to learn object grounding based on a given expression using weak supervision signals like image-text pairs. While these tasks have traditionally been modeled separately, we argue that they can benefit from joint learning in a multi-task framework. To this end, we propose We

  57. Zhihao Zhang, Yiran Zhang, Xiyue Zhou, Liting Huang

    Infodemics and health misinformation have significant negative impact on individuals and society, exacerbating confusion and increasing hesitancy in adopting recommended health measures. Recent advancements in generative AI, capable of producing realistic, human like text and images, have significantly accelerated the spread and expanded the reach of health

  58. Zhixing Wang, Le Zheng, Shi Yan, Ruud J. G. van Sloun

    Extended object tracking methods based on random matrices, founded on Bayesian filters, have been able to achieve efficient recursive processes while jointly estimating the kinematic states and extension of the targets. Existing random matrix approaches typically assume that the evolution of state and extension follows a first-order Markov process, where the

  59. Raphaël Merx, Hanna Suominen, Lois Hong, Nick Thieberger

    Machine translation (MT) systems that support low-resource languages often struggle on specialized domains. While researchers have proposed various techniques for domain adaptation, these approaches typically require model fine-tuning, making them impractical for non-technical users and small organizations. To address this gap, we propose Tulun, a versatile

  60. Anastasios Apsemidis, Karin Weyermair, Hans Peter Stüger, Sabrina Kuchling

    Wastewater data can be very useful for epidemic control during a disease outbreak and proper synthesis of different sources of information can be integrated towards an alerting system, that can be used for decision support. Wastewater data are considered to be of high quality, since they do not depend on testing and can take into account asymptomatic cases.

  61. Shashank Raj, Kalyanmoy Deb

    We present EvoSort, a general-purpose adaptive parallel parallel sorting framework accessible at the Python level. EvoSort employs a Genetic Algorithm (GA) to automatically discover and refine critical parameters, including insertion sort thresholds and algorithm selection (e.g., versus LSD radix sort). By adapting continuously to input data and system archi

  62. Yuanhe Zhang, Xinyue Wang, Haoran Gao, Zhenhong Zhou

    Large Language Models (LLMs), due to substantial computational requirements, are vulnerable to resource consumption attacks, which can severely degrade server performance or even cause crashes, as demonstrated by denial-of-service (DoS) attacks designed for LLMs. However, existing works lack mitigation strategies against such threats, resulting in unresolved

  63. Bin Ren, Yawei Li, Xu Zheng, Yuqian Fu

    Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across different degradation types. Existing approaches either sacrifice efficiency for versatility or fail to capture the distinct representational requirements of various degradations. We

  64. Kazuya Nakayam, Takanari Yasui

    UVA to NIR with multi-directional photo responses have been found on metal (Au)/n-Si device. A reasonable explanation has not been found in various physical models of Si-devices for the phenomena. We approached a zero-gap at X (reciprocal point) in two conduction bands of Si to analysis the optical response with the inter-band phonon scatterings. The calcula

  65. Eric Chamoun, Nedjma Ousidhoum, Michael Schlichtkrull, Andreas Vlachos

    Clarifying the research framing of NLP artefacts (e.g., models, datasets, etc.) is crucial to aligning research with practical applications. Recent studies manually analyzed NLP research across domains, showing that few papers explicitly identify key stakeholders, intended uses, or appropriate contexts. In this work, we propose to automate this analysis, dev

  66. Zhengyu Chen, Yudong Wang, Teng Xiao, Ruochen Zhou

    Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We inv

  67. Achini Jayawardane, Rajitha Senanayake, Erfan Khordad, Jamie Evans

    Cell-free wireless networks have attracted significant interest for their ability to eliminate cell-edge effects and deliver uniformly high service quality through macro-diversity. In this paper, we develop an algorithm to jointly optimize uplink transmit powers and dynamic user-centric access point (AP) clusters in a centralized cell-free network. This appr

  68. Sicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong

    Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic. To bridge this gap, we introduce ReasonMap, a novel benchmark specifically designed to evaluate these capabilities. Reaso

  69. Peng Xiao, Hongbo Zhao, Yijun Wang, Jianxin Lin

    Restoring real-world degraded images, such as old photographs or low-resolution images, presents a significant challenge due to the complex, mixed degradations they exhibit, such as scratches, color fading, and noise. Recent data-driven approaches have struggled with two main challenges: achieving high-fidelity restoration and providing object-level control

  70. Zixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen

    Large Language Models (LLMs) have achieved remarkable success in Natural Language Processing (NLP), yet their cross-lingual performance consistency remains a significant challenge. This paper introduces a novel methodology for efficiently identifying inherent cross-lingual weaknesses in LLMs. Our approach leverages beam search and LLM-based simulation to gen

  71. Yu Zhang, Wanli Jiang, Zhengyu Yang

    The multi-objective alignment of Large Language Models (LLMs) is essential for ensuring foundational models conform to diverse human preferences. Current research in this field typically involves either multiple policies or multiple reward models customized for various preferences, or the need to train a preference-specific supervised fine-tuning (SFT) model

  72. Giacomo Turri, Luigi Bonati, Kai Zhu, Massimiliano Pontil

    We introduce an encoder-only approach to learn the evolution operators of large-scale non-linear dynamical systems, such as those describing complex natural phenomena. Evolution operators are particularly well-suited for analyzing systems that exhibit complex spatio-temporal patterns and have become a key analytical tool across various scientific communities

  73. Chonghua Han, Yuan Yuan, Jingtao Ding, Jie Feng

    The success of foundation models in language has inspired a new wave of general-purpose models for human mobility. However, existing approaches struggle to scale effectively due to two fundamental limitations: a failure to use meaningful basic units to represent movement, and an inability to capture the vast diversity of patterns found in large-scale data. I

  74. Zishun Yu, Shangzhe Li, Xinhua Zhang

    Large language models have led to significant progress across many NLP tasks, although their massive sizes often incur substantial computational costs. Distillation has become a common practice to compress these large and highly capable models into smaller, more efficient ones. Many existing language model distillation methods can be viewed as behavior cloni

  75. Christoffer Tarmet

    This paper investigates the concept of an optimal ratio for regular polytopes in $n$-dimensional space within the framework of the Generalized Chaos Game. The optimal ratio, $r_{\text{opt}}$, is defined as the value at which the self-similar regions of the resulting fractal touch but do not overlap. Using a series of Python simulations, we explore how the op

  76. Zhen Li, Duan Li, Yukai Guo, Xinyuan Guo

    Infographic charts are a powerful medium for communicating abstract data by combining visual elements (e.g., charts, images) with textual information. However, their visual and structural richness poses challenges for large vision-language models (LVLMs), which are typically trained on plain charts. To bridge this gap, we introduce ChartGalaxy, a million-sca

  77. Elise M. Sänger

    We performed tests of General Relativity on gravitational wave signal GW230529_181500, which comes from what is most likely a neutrons star merging with a black hole in the lower mass gap. We used two different frameworks to perform parameterized inspiral tests. We find that the signal is consistent with General Relativity for all deviation parameters and we

  78. Joysankar Majumdar, Sakshi Maurya, Raj Prince

    In October 2024, the object BL Lacertae experienced the brightest flaring event in gamma-ray ($>$100 MeV) with a historically bright $\gamma$-ray flux of $\sim$2.59 $\times 10^{-5}$ erg cm$^{-2}$ s$^{-1}$ with a detection of a 175.7 GeV photon with Fermi-LAT. This event was also followed by very high-energy $\gamma$-ray detection with LHAASO, VERITAS, and MA

  79. Aaron Beyen, Christian Maes, Ji-Hui Pei

    We consider a slow elastic string with Klein-Gordon dynamics coupled to a bath of run-and-tumble particles. We derive and solve the induced Langevin-Klein-Gordon string dynamics with explicit expressions for the streaming term, friction coefficient, and noise variance. These parameters are computed exactly in a weak coupling expansion. The induced friction i

  80. Evgeny Ugolkov, Xupeng He, Hyung Kwak, Hussein Hoteit

    We present a memory-efficient algorithm for significantly enhancing the quality of segmented 3D micro-Computed Tomography (micro-CT) images of rocks using a generative model. The proposed model achieves a 16x increase in resolution and corrects inaccuracies in segmentation caused by the overlapping X-ray attenuation in micro-CT measurements across different

  81. Zhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu

    Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While post-training quantization (PTQ) presents a promising approach to accelerate Video DiT models, existing methods suffer from two critical limitations: (1) dependence on computation-

  82. Bohan Ouyang, Maurizio Grasselli, Hao Wu

    We study a thermodynamically consistent diffuse interface model that describes the motion of a two-phase flow of two viscous incompressible Newtonian fluids with unmatched densities and a soluble surfactant in a bounded domain of two or three dimensions. The resulting hydrodynamic system consists of a nonhomogeneous Navier-Stokes system for the (volume avera

  83. Santiago Berrezueta-Guzman, María Dolón-Poza, Stefan Wagner

    This study evaluates the integration of AI-powered robots in early childhood education, focusing on their impact on emotional self-regulation, engagement, and collaborative skills. A ten-week experimental design involving two groups of children assessed the robot's effectiveness through progress assessments, parental surveys, and teacher feedback. Results de

  84. Sangwoo Park, Matteo Zecchin, Osvaldo Simeone

    Selecting artificial intelligence (AI) models, such as large language models (LLMs), from multiple candidates requires accurate performance estimation. This is ideally achieved through empirical evaluations involving abundant real-world data. However, such evaluations are costly and impractical at scale. To address this challenge, autoevaluation methods leve

  85. Pankaj Kumar, Subhankar Mishra

    Large Language Models (LLMs) have emerged as a promising cornerstone for the development of natural language processing (NLP) and artificial intelligence (AI). However, ensuring the robustness of LLMs remains a critical challenge. To address these challenges and advance the field, this survey provides a comprehensive overview of current studies in this area.

  86. Xu Zheng, Chenfei Liao, Yuqian Fu, Kaiyu Lei

    Recent advances in Multimodal Large Language Models (MLLMs) have shown promising results in integrating diverse modalities such as texts and images. MLLMs are heavily influenced by modality bias, often relying on language while under-utilizing other modalities like visual inputs. This position paper argues that MLLMs are deeply affected by modality bias. Fir

  87. Dev Gurung, Shiva Raj Pokhrel

    Inspired by the power of large language models (LLMs), our research adapts them to quantum federated learning (QFL) to boost efficiency and performance. We propose a federated fine-tuning method that distills an LLM within QFL, allowing each client to locally adapt the model to its own data while preserving privacy and reducing unnecessary global updates. Th

  88. Alberto Enciso, Antonio J. Fernández, David Meyer

    We show how to regularize vortex sheets by means of smooth, compactly supported vorticities that asymptotically evolve according to the Birkhoff-Rott vortex sheet dynamics. More precisely, consider a vortex sheet initial datum $\omega^0_{\mathrm{sing}}$, which is a signed Radon measure supported on a closed curve. We construct a family of initial vorticities

  89. Ruidong Han, Bin Yin, Shangyu Chen, He Jiang

    Scaling law has been extensively validated in many domains such as natural language processing and computer vision. In the recommendation system, recent work has adopted generative recommendations to achieve scalability, but their generative approaches require abandoning the carefully constructed cross features of traditional recommendation models. We found

  90. Murathan Kurfalı, Shorouq Zahra, Joakim Nivre, Gabriele Messori

    Climate-Eval is a comprehensive benchmark designed to evaluate natural language processing models across a broad range of tasks related to climate change. Climate-Eval aggregates existing datasets along with a newly developed news classification dataset, created specifically for this release. This results in a benchmark of 25 tasks based on 13 datasets, cove

  91. Yicheng Lin, Yunlong Jiang, Xujia Jiao, Bin Han

    Robust long-term visual localization in complex industrial environments is critical for mobile robotic systems. Existing approaches face limitations: handcrafted features are illumination-sensitive, learned features are computationally intensive, and semantic- or marker-based methods are environmentally constrained. Handcrafted and learned features share sim

  92. Daniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu Wyart

    Models such as Word2Vec and GloVe construct word embeddings based on the co-occurrence probability $P(i,j)$ of words $i$ and $j$ in text corpora. The resulting vectors $W_i$ not only group semantically similar words but also exhibit a striking linear analogy structure -- for example, $W_{\text{king}} - W_{\text{man}} + W_{\text{woman}} \approx W_{\text{queen

  93. Xiaodong Wang, Peixi Peng

    Real-world driving requires people to observe the current environment, anticipate the future, and make appropriate driving decisions. This requirement is aligned well with the capabilities of world models, which understand the environment and predict the future. However, recent world models in autonomous driving are built explicitly, where they could predict

  94. Shiyun Xie, Zhiru Wang, Yinghao Zhu, Xu Wang

    Recently, 3D Gaussian Splatting (3DGS) has excelled in novel view synthesis (NVS) with its real-time rendering capabilities and superior quality. However, it encounters challenges for high-resolution novel view synthesis (HRNVS) due to the coarse nature of primitives derived from low-resolution input views. To address this issue, we propose SuperGS, an expan

  95. Sadegh Keshavarzi, Gregory Chockler, Alexey Gotsman

    Recent advances in secure hardware technologies, such as Intel SGX or ARM TrustZone, offer an opportunity to substantially reduce the costs of Byzantine fault-tolerance by placing the program code and state within a secure enclave known as a Trusted Execution Environment (TEE). However, the protection offered by a TEE only applies during program execution. O

  96. Siwei Liu, Jinyuan Fang, Han Zhou, Yingxu Wang

    Large Language Models (LLMs) have demonstrated effectiveness in code generation tasks. To enable LLMs to address more complex coding challenges, existing research has focused on crafting multi-agent systems with agentic workflows, where complex coding tasks are decomposed into sub-tasks, assigned to specialized agents. Despite their effectiveness, current ap

  97. Haleema Bibi, Sadia Saleem, Zakia Jalil, Muhammad Nasir

    Flooding is the most devastating phenomenon occurring globally, particularly in mountainous regions, risk dramatically increases due to complex terrains and extreme climate changes. These situations are damaging livelihoods, agriculture, infrastructure, and human lives. This study uses the Kabul River between Pakistan and Afghanistan as a case study to refle

  98. Jingran Xie, Xiang Li, Hui Wang, Yue Yu

    Large language models (LLMs) have shown remarkable generalization across tasks, leading to increased interest in integrating speech with LLMs. These speech LLMs (SLLMs) typically use supervised fine-tuning to align speech with text-based LLMs. However, the lack of annotated speech data across a wide range of tasks hinders alignment efficiency, resulting in p

  99. Steven Ndung'u, Trienko Grobler, Stefan J. Wijnholds, George Azzopardi

    Detecting anomalies in radio astronomy is challenging due to the vast amounts of data and the rarity of labeled anomalous examples. Addressing this challenge requires efficient methods capable of identifying unusual radio galaxy morphologies without relying on extensive supervision. This work introduces an innovative approach to anomaly detection based on mo

  100. Xiao Chen, Sihang Zhou, Ke Liang, Xiaoyu Sun

    Chain-of-thought (CoT) distillation allows a large language model (LLM) to guide a small language model (SLM) in reasoning tasks. Existing methods train the SLM to learn the long rationale in one iteration, resulting in two issues: 1) Long rationales lead to a large token-level batch size during training, making gradients of core reasoning tokens (i.e., the