Skip to content

May 2025 arXiv papers — page 63

Showing 6,2016,300 of 24,552 papers

  1. Shiyue Wang, Haozheng Xu, Yuhan Zhang, Jingran Lin

    Multi-Agent Path Finding (MAPF) is a fundamental problem in artificial intelligence and robotics, requiring the computation of collision-free paths for multiple agents navigating from their start locations to designated goals. As autonomous systems become increasingly prevalent in warehouses, urban transportation, and other complex environments, MAPF has evo

  2. Lyle Regenwetter, Yazan Abu Obaideh, Fabien Chiotti, Ioanna Lykourentzou

    We introduce BikeBench, an engineering design benchmark for evaluating generative models on problems with multiple real-world objectives and constraints. As generative AI's reach continues to grow, evaluating its capability to understand physical laws, human guidelines, and hard constraints grows increasingly important. Engineering product design lies at the

  3. Jingwei Wu, Zhewei Huang, Chang Liu

    In the past decade, image foundation models (IFMs) have achieved unprecedented progress. However, the potential of directly using IFMs for video self-supervised representation learning has largely been overlooked. In this study, we propose an advancing video self-supervised learning (AdViSe) approach, aimed at significantly reducing the training overhead of

  4. Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding

    Recent large language models (LLMs) exhibit impressive reasoning but often over-think, generating excessively long responses that hinder efficiency. We introduce DIET ( DIfficulty-AwarE Training), a framework that systematically cuts these "token calories" by integrating on-the-fly problem difficulty into the reinforcement learning (RL) process. DIET dynamic

  5. Idit Keidar, Andrew Lewis-Pye, Ehud Shapiro, Nimrod Talmon

    Permissionless-consensus-based Decentralised Autonomous Organisations (DAOs) are the prevailing paradigm for participant-governed digital organisations. As participants have verified resources but no trusted identities, this ecosystem is necessarily plutocratic (one coin -- one vote). Here we offer, for the first time, a democratic (one person -- one vote) p

  6. Abhay Negi, Omey M. Manyar, Dhanush Kumar Varma Penmetsa, Satyandra K. Gupta

    Contact-rich assembly of complex, non-convex parts with tight tolerances remains a formidable challenge. Purely model-based methods struggle with discontinuous contact dynamics, while model-free methods require vast data and often lack precision. In this work, we introduce a hybrid framework that uses only contact-state information between a complex peg and

  7. Zifan Wang, Teli Ma, Yufei Jia, Xun Yang

    Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-based perception approaches often struggle with sensor noise, lighting variability, computational overhead from intermediate representations (e.g., elevation maps), and difficulties

  8. Shaohao Rui, Kaitao Chen, Weijie Ma, Xiaosong Wang

    Recent advances in reinforcement learning with verifiable, rule-based rewards have greatly enhanced the reasoning capabilities and out-of-distribution generalization of VLMs/LLMs, obviating the need for manually crafted reasoning chains. Despite these promising developments in the general domain, their translation to medical imaging remains limited. Current

  9. Abdelaziz Salama, Mohammed M. H. Qazzaz, Syed Danial Ali Shah, Maryam Hafeez

    This work proposes an integrated approach for optimising Federated Learning (FL) communication in dynamic and heterogeneous network environments. Leveraging the modular flexibility of the Open Radio Access Network (ORAN) architecture and multiple Radio Access Technologies (RATs), we aim to enhance data transmission efficiency and mitigate client-server commu

  10. Xiang Li, Rongrong Wang, Qing Qu

    Classifier-free guidance (CFG) is a core technique powering state-of-the-art image generation systems, yet its underlying mechanisms remain poorly understood. In this work, we begin by analyzing CFG in a simplified linear diffusion model, where we show its behavior closely resembles that observed in the nonlinear case. Our analysis reveals that linear CFG im

  11. Zonglin Yang, Wanhao Liu, Ben Gao, Yujie Liu

    Large language models (LLMs) have shown promise in automating scientific hypothesis generation, yet existing approaches primarily yield coarse-grained hypotheses lacking critical methodological and experimental details. We introduce and formally define the new task of fine-grained scientific hypothesis discovery, which entails generating detailed, experiment

  12. Tyler Ward, Aaron Moseley, Abdullah-Al-Zubaer Imran

    Segmentation is one of the most important tasks in the medical imaging pipeline as it influences a number of image-based decisions. To be effective, fully supervised segmentation approaches require large amounts of manually annotated training data. However, the pixel-level annotation process is expensive, time-consuming, and error-prone, hindering progress a

  13. Ariel Smooha, Jitender Kumar, Dan Yudilevich, John W. Rosenberg

    Single-molecule magnets (SMMs) are molecules that can function as nanoscale magnets with potential use as magnetic memory bits. While SMMs can retain magnetization at low temperatures, characterizing them on surface and at room temperature remains challenging and requires specialized nanoscale techniques. Here, we use single nitrogen-vacancy (NV) centers in

  14. Richard He Bai, Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly

    The latency bottleneck of traditional text-to-speech (TTS) systems fundamentally hinders the potential of streaming large language models (LLMs) in conversational AI. These TTS systems, typically trained and inferenced on complete utterances, introduce unacceptable delays, even with optimized inference speeds, when coupled with streaming LLM outputs. This is

  15. Meher Bhaskar Madiraju, Meher Sai Preetam Madiraju

    Hyperparameter optimization (HPO) is a critical yet challenging aspect of machine learning model development, significantly impacting model performance and generalization. Traditional HPO methods often struggle with high dimensionality, complex interdependencies, and computational expense. This paper introduces OptiMindTune, a novel multi-agent framework des

  16. Federico Becca, Alberto Parola

    We extend the previously defined many-body marker for two-dimensional $\mathbb{Z}_2$ topological insulators [I. Gilardoni {\it et al.}, Phys. Rev. B {\bf 106}, L161106 (2022)] to distinguish trivial, weak-, and strong-topological insulators in three dimensions, in presence of the inversion symmetry. The marker is written in term of ground-state expectation v

  17. Han Yin, Yang Xiao, Rohan Kumar Das, Jisheng Bai

    Audio generation systems now create very realistic soundscapes that can enhance media production, but also pose potential risks. Several studies have examined deepfakes in speech or singing voice. However, environmental sounds have different characteristics, which may make methods for detecting speech and singing deepfakes less effective for real-world sound

  18. Morteza Nattagh Najafi, Fatemeh Foroughirad

    We investigate the space time fractional nonlinear Schrodinger equation (FNLSE) incorporating the modified Riemann Liouville derivative introduced by Jumari. The equation is characterized by two parameters: the fractional derivative parameter (alpha, which captures the memory effects) and the non linearity parameter a. We present analytical solutions via thr

  19. Yunhai Hu, Tianhua Xia, Zining Liu, Rahul Raman

    Speculative decoding (SD) has emerged as a powerful method for accelerating autoregressive generation in large language models (LLMs), yet its integration into vision-language models (VLMs) remains underexplored. We introduce DREAM, a novel speculative decoding framework tailored for VLMs that combines three key innovations: (1) a cross-attention-based mecha

  20. I. Fernández de Fuentes, E. Raymenants, B. Undseth, O. Pietx-Casas

    The simplicity of encoding a qubit in the state of a single electron spin and the potential for their integration into industry-standard microchips continue to drive the field of semiconductor-based quantum computing. However, after decades of progress, validating universal logic in these platforms has advanced little beyond first-principles demonstrations o

  21. David Schimel, Andres Baresch, Adam Chlus, Phil Townsend

    Plant functional trait variation in tropical forests is central to predicting ecosystem responses to change. Informaiton on traits is limited relative to the diversity of climate, landforms, disturbance regimes and species present. These traits are central to modeled predictions of ecosystem change. We used a new spaceborne imagining spectrometer from the It

  22. Abuzer Gündüz, Osama A. Naji, Mehmet Özen

    This article studies the notion of $S-r-$ideals in commutative ring $H$, where $S$ is a multiplicatively closed subset of $H$. Some basic properties of $S-r-$ideals are given. Various characterizations of $S-r-$ideals are presented. Also, $S-uz-$ring is defined and it is proved that $H$ is an $S-uz-$ring if and only if every maximal ideal disjoint from $S$ i

  23. Chanyeol Choi, Alejandro Lopez-Lira, Yongjae Lee, Jihoon Kwon

    Extracting structured and quantitative insights from unstructured financial filings is essential in investment research, yet remains time-consuming and resource-intensive. Conventional approaches in practice rely heavily on labor-intensive manual processes, limiting scalability and delaying the research workflow. In this paper, we propose an efficient and sc

  24. Xinyao Liao, Wei Wei, Xiaoye Qu, Qiyuan He

    Recent advances in text-to-image (T2I) diffusion model fine-tuning leverage reinforcement learning (RL) to align generated images with learnable reward functions. The existing approaches reformulate denoising as a Markov decision process for RL-driven optimization. However, they suffer from reward sparsity, receiving only a single delayed reward per generate

  25. Shaohao Rui, Haoyang Su, Jinyi Xiang, Lian-Ming Wu

    Accurate prediction of major adverse cardiovascular events recurrence risk in acute myocardial infarction patients based on postoperative cardiac MRI and associated clinical notes is crucial for precision treatment and personalized intervention. Existing methods primarily focus on risk stratification capability while overlooking the need for intermediate rob

  26. Peiran Sun

    Adversarial attack reveals the vulnerability of deep learning models. It is assumed that high curvature may give rise to rough decision boundary and thus result in less robust models. However, the most commonly used \textit{curvature} is the curvature of loss function, scores or other parameters from within the model as opposed to decision boundary curvature

  27. Maya Bechler-Speicher, Andrea Zerio, Maor Huri, Marie Vibeke Vestergaard

    Real-world temporal data often consists of multiple signal types recorded at irregular, asynchronous intervals. For instance, in the medical domain, different types of blood tests can be measured at different times and frequencies, resulting in fragmented and unevenly scattered temporal data. Similar issues of irregular sampling occur in other domains, such

  28. Bastiaan Cnossen, Tobias Lenz, Sil Linskens

    Given an $\infty$-category $C$ equipped with suitable wide subcategories $I, P \subset E\subset C$, we show that the $(\infty,2)$-category $\text{S}{\scriptstyle\text{PAN}}_2(C,E)_{P,I}$ of higher (or iterated) spans defined by Haugseng has the universal property that 2-functors $\text{S}{\scriptstyle\text{PAN}}_2(C,E)_{P,I} \to \mathbb D$ correspond precise

  29. Nursulu Sagimbayeva, Ruveyda Betül Bahçeci, Ingmar Weber

    Inconsistent political statements represent a form of misinformation. They erode public trust and pose challenges to accountability, when left unnoticed. Detecting inconsistencies automatically could support journalists in asking clarification questions, thereby helping to keep politicians accountable. We propose the Inconsistency detection task and develop

  30. Jiayi Xin, Sukwon Yun, Jie Peng, Inyoung Choi

    Modality fusion is a cornerstone of multimodal learning, enabling information integration from diverse data sources. However, vanilla fusion methods are limited by (1) inability to account for heterogeneous interactions between modalities and (2) lack of interpretability in uncovering the multimodal interactions inherent in the data. To this end, we propose

  31. Yaoyang Liu, Junlin Li, Yinjun Wu, Zhen Chen

    Although Multi-Vector Retrieval (MVR) has achieved the state of the art on many information retrieval (IR) tasks, its performance highly depends on how to decompose queries into smaller pieces, say phrases or tokens. However, optimizing query decomposition for MVR performance is not end-to-end differentiable. Even worse, jointly solving this problem and trai

  32. Hongxu Pan, Shuxian Hu, Mo Zhou, Zhibin Wang

    Researchers have proposed various methods of incorporating more structured information into the design of Graph Neural Networks (GNNs) to enhance their expressiveness. However, these methods are either computationally expensive or lacking in provable expressiveness. In this paper, we observe that the chords increase the complexity of the graph structure whil

  33. Yang Xiao, Jiashuo Wang, Ruifeng Yuan, Chunpu Xu

    Large language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progre

  34. Rushiraj Gadhvi, Priyansh Desai, Siddharth

    Automated pose correction remains a significant challenge in AI-driven fitness systems, despite extensive research in activity recognition. This work presents PosePilot, a novel system that integrates pose recognition with real-time personalized corrective feedback, overcoming the limitations of traditional fitness solutions. Using Yoga, a discipline requiri

  35. Qiong Deng, Minghui Du, Peng Xu, Liang Huang

    The $\mu$Hz gravitational wave band holds crucial insights into coalescing supermassive black hole binaries and stochastic backgrounds but remains inaccessible due to technical challenges. We demonstrate that geocentric space-based GW detectors (e.g., TianQin, gLISA, GADFLI) can bridge this gap by considering orbital resonance effects, circumventing the need

  36. Pradyumna Shyama Prasad, Minh Nhat Nguyen

    Can LLMs accurately adjust their confidence when facing opposition? Building on previous studies measuring calibration on static fact-based question-answering tasks, we evaluate Large Language Models (LLMs) in a dynamic, adversarial debate setting, uniquely combining two realistic factors: (a) a multi-turn format requiring models to update beliefs as new inf

  37. A. Jung

    This book offers a hands-on introduction to building and understanding federated learning (FL) systems. FL enables multiple devices -- such as smartphones, sensors, or local computers -- to collaboratively train machine learning (ML) models, while keeping their data private and local. It is a powerful solution when data cannot or should not be centralized du

  38. Kefan Wang, Hao Wang, Wei Guo, Yong Liu

    Click-through rate (CTR) prediction is a critical task in online advertising and recommender systems, relying on effective modeling of feature interactions. Explicit interactions capture predefined relationships, such as inner products, but often suffer from data sparsity, while implicit interactions excel at learning complex patterns through non-linear tran

  39. Yaoting Gui, Yuqiao Li, Jun Sun

    This paper extends the results of [GLS24], where the existence of a constant harmonic mean curvature foliation was established in the setting of a 3-dimensional asymptotically Schwarzschild manifold. Here, we generalize this construction to higher dimensions, proving the existence of foliations by constant harmonic mean curvature hypersurfaces in an asymptot

  40. Lukas Exl, Sebastian Schaffer

    We present an extension of the tensor grid method for stray field computation on rectangular domains that incorporates higher-order basis functions. Both the magnetization and the resulting magnetic field are represented using higher-order B-spline bases, which allow for increased accuracy and smoothness. The method employs a super-potential formulation, whi

  41. Xun Gong, Anqi Lv, Zhiming Wang, Huijia Zhu

    While speech large language models (SpeechLLMs) have advanced standard automatic speech recognition (ASR), contextual biasing for named entities and rare words remains challenging, especially at scale. To address this, we propose BR-ASR: a Bias Retrieval framework for large-scale contextual biasing (up to 200k entries) via two innovations: (1) speech-and-bia

  42. Akhila Yaragoppa, Siddharth

    Understanding the emotional impact of videos is crucial for applications in content creation, advertising, and Human-Computer Interaction (HCI). Traditional affective computing methods rely on self-reported emotions, facial expression analysis, and biosensing data, yet they often overlook the role of visual saliency -- the naturally attention-grabbing region

  43. Zhuo Liu, Moxin Li, Xun Deng, Qifan Wang

    LLM-as-a-Judge employs large language models (LLMs), such as GPT-4, to evaluate the quality of LLM-generated responses, gaining popularity for its cost-effectiveness and strong alignment with human evaluations. However, training proxy judge models using evaluation data generated by powerful teacher models introduces a critical yet previously overlooked issue

  44. Jan Held, Renaud Vandeghen, Adrien Deliege, Abdullah Hamdi

    The field of computer graphics was revolutionized by models such as Neural Radiance Fields and 3D Gaussian Splatting, displacing triangles as the dominant representation for photogrammetry. In this paper, we argue for a triangle comeback. We develop a differentiable renderer that directly optimizes triangles via end-to-end gradients. We achieve this by rende

  45. Wei Zhang, Ju Xing, Xiaoqi Li

    Penetration testing refers to the process of simulating hacker attacks to evaluate the security of information systems . This study aims not only to clarify the theoretical foundations of penetration testing but also to explain and demonstrate the complete testing process, including how network system administrators may simulate attacks using various penetra

  46. Debdeep Sanyal, Agniva Maiti, Umakanta Maharana, Dhruv Kumar

    Effective teaching requires adapting instructional strategies to accommodate the diverse cognitive and behavioral profiles of students, a persistent challenge in education and teacher training. While Large Language Models (LLMs) offer promise as tools to simulate such complex pedagogical environments, current simulation frameworks are limited in two key resp

  47. Shiri Artstein-Avidan, Arnon Chor

    We study the c-affine surface area $\Omega^c$, recently introduced by Sch\"utt, Werner and Yalikun. We show that on the class of ball-bodies, $\Omega^c$ is maximized by a ball of radius $\frac{n}{n+1}$, and that a Santal\'o-type inequality holds: $\Omega^c(K) \Omega^c(K^c) \leq \Omega^c(\frac{1}{2} B_2^n)^2$. We also produce some more intricate inequalities

  48. Atahan Karagoz

    We identify a conserved quantity in continuous-time optimization dynamics, termed computational inertia. Defined as the sum of kinetic energy (parameter velocity) and potential energy (loss), this scalar remains invariant under idealized, frictionless training. We formalize this conservation law, derive its analytic decay under damping and stochastic perturb

  49. Steven Samuels, William Campbell, Michael E. Tobar, Maxim Goryachev

    A low-noise cryogenic microwave spectroscopy experiment was performed on a high-purity lithium fluoride (LiF) crystal. The spectroscopy data revealed avoided level crossing interactions in whispering gallery modes, indicative of electron spin resonance (ESR) coupling with paramagnetic impurities. Analysis of the interaction spectra identified distinct spin s

  50. Vincenzo Mallardo, Christian Dunser, Gernot Beer

    This paper is concerned with the Boundary Element simulation of elastic domains that contain thin inclusions that have elastic material properties, which are different to the domain. With thin inclusions we mean inclusions with extreme aspect ratios, i.e. where one dimension is much smaller than the other ones. Examples of this are reinforcements in civil/me

  51. Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa

    Reconstructing 3D hand mesh is challenging but an important task for human-computer interaction and AR/VR applications. In particular, RGB and/or depth cameras have been widely used in this task. However, methods using these conventional cameras face challenges in low-light environments and during motion blur. Thus, to address these limitations, event camera

  52. Swee Hong Chan, Alex Kontorovich, Igor Pak

    For a simple graph $G=(V,E)$ and edge $e\in E$, the effective resistance is defined as a ratio $\frac{\tau(G/e)}{\tau(G)}$, where $\tau(G)$ denotes the number of spanning trees in $G$. We resolve the inverse problem for the effective resistance for planar graphs. Namely, we determine (up to a constant) the smallest size of a simple planar graph with a given

  53. Thomas P. Kehler, Scott E. Page, Alex Pentland, Martin Reeves

    We propose a general framework for human-AI collaboration that amplifies the distinct capabilities of both types of intelligence. We refer to this as Generative Collective Intelligence (GCI). GCI employs AI in dual roles: as interactive agents and as technology that accumulates, organizes, and leverages knowledge. In this second role, AI creates a cognitive

  54. Eric Tillmann Bill, Enis Simsar, Thomas Hofmann

    We introduce JEDI, a test-time adaptation method that enhances subject separation and compositional alignment in diffusion models without requiring retraining or external supervision. JEDI operates by minimizing semantic entanglement in attention maps using a novel Jensen-Shannon divergence based objective. To improve efficiency, we leverage adversarial opti

  55. Debdeep Sanyal, Umakanta Maharana, Yash Sinha, Hong Ming Tan

    Role-based access control (RBAC) and hierarchical structures are foundational to how information flows and decisions are made within virtually all organizations. As the potential of Large Language Models (LLMs) to serve as unified knowledge repositories and intelligent assistants in enterprise settings becomes increasingly apparent, a critical, yet under exp

  56. Ashirbad Mishra, Jinyu Zhao, Soumik Dey, Hansi Wu

    In the domain of sponsored search advertising, the focus of Keyphrase recommendation has largely been on exact match types, which pose issues such as high management expenses, limited targeting scope, and evolving search query patterns. Alternatives like Broad match types can alleviate certain drawbacks of exact matches but present challenges like poor targe

  57. Firoj Alam, Md Arid Hasan, Shammur Absar Chowdhury

    Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce SpokenNativQA, the first multilingual and culturally aligned spoken question-answering (SQA) dataset designed to evaluate

  58. Shun Xu, Jianzhi Han

    Let $V$ be a vertex operator algebra and $g$ an automorphism of $V$ of finite order $T$. For any $m, n \in(1/T) \mathbb N$, an $A_{g,n}(V)\!-\!A_{g,m}(V)$ bimodule $A_{g,n, m}(V)=V/O_{g,n,m}(V)$ was defined by Dong and Jiang, where $O_{g,n,m}(V)$ is the sum of three certain subspaces $O_{g,n, m}^{\prime}(V), O_{g,n, m}^{\prime \prime}(V)$ and $O_{g,n, m}^{\p

  59. Jialun Pei, Diandian Guo, Donghui Yang, Zhixi Li

    In endoscopic surgery, a clear and high-quality visual field is critical for surgeons to make accurate intraoperative decisions. However, persistent visual degradation, including smoke generated by energy devices, lens fogging from thermal gradients, and lens contamination due to blood or tissue fluid splashes during surgical procedures, severely impairs vis

  60. Raju Biswas

    Let $\mathcal{H}$ be the class of harmonic functions $f=h+\overline{g}$ in the unit disk $\mathbb{D}:=\{z\in\mathbb{C}:|z|<1\}$, where $h$ and $g$ are analytic in $\mathbb{D}$ with the normalization $h(0)=g(0)=h'(0)-1=0$. Let $\mathcal{D}_{\mathcal{H}}^0(\alpha, M)$ denote the class of functions $f=h+ \overline{g}\in\mathcal{H}$ satisfying the conditions $\l

  61. Yuze Wang, Mariana Belgiu, Haiyang Wu, Dandan Zhong

    Satellite Image Time Series (SITS) is crucial for agricultural semantic segmentation. However, Cloud contamination introduces time gaps in SITS, disrupting temporal dependencies and causing feature shifts, leading to degraded performance of models trained on complete SITS. Existing methods typically address this by reconstructing the entire SITS before predi

  62. Matthias Blaschke, Fabian Pauly

    Measurements of the thermal conductance of single-molecule junctions have recently been reported for the first time. It is presently unclear, how much the heat transport can be controlled through molecule-internal effects. The search for molecules with lowest and highest thermal conductance is complicated by the gigantic chemical space. Here we describe a sy

  63. Marius Causemann, Miroslav Kuchta

    This paper presents a scalable and robust solver for a cell-by-cell poroelasticity model, describing the mechanical interactions between brain cells embedded in extracellular space. Explicitly representing the complex cellular shapes, the proposed approach models both intracellular and extracellular spaces as distinct poroelastic media, separated by a permea

  64. Prasanth Shyamsundar

    The paper titled "An implementation of neural simulation-based inference for parameter estimation in ATLAS" by the ATLAS collaboration (arXiv:2412.01600v1 [hep-ex]) describes the implementation of neural simulation-based inference for a measurement analysis performed by ATLAS. The uncertainties in the analysis arising from the finiteness of the simulated dat

  65. Q. G. Duan, Benyun Zhao, Mingqiao Han Yijun Huang, Ben M. Chen

    Scene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic color representation of gaussian primitives and the isotropic requirements of semantic features, leading to insufficient

  66. Jingcheng Dong

    Let $\C$ be a self-dual fusion category of rank $4$ which has a nontrivial proper fusion subcategory. We identify three new families of Grothendieck rings for $\C$: one of them is completely determined, the other two are parameterized by several non-negative integers.

  67. Yajie Ji, Yanlai Chen, Shawn Koohy

    We propose S$^2$GPT-PINN, a sparse and small model for solving parametric partial differential equations (PDEs). Similar to Small Language Models (SLMs), S$^2$GPT-PINN is tailored to domain-specific (families of) PDEs and characterized by its compact architecture and minimal computational power. Leveraging a small amount of extremely high quality data via a

  68. Zhenyu Li, Özlem Tuğfe Demir, Emil Björnson, Cicek Cavdar

    This paper investigates the application of reconfigurable intelligent surfaces (RISs) to improve fronthaul link survivability in cell-free massive MIMO (CF mMIMO) systems. To enhance the fronthaul survivability, two complementary mechanisms are considered. Firstly, RIS is set to provide reliable line-of-sight (LOS) connectivity and enhance the mmWave backup

  69. Shenggan Cheng, Yuanxin Wei, Lansong Diao, Yong Liu

    Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing tasks. Despite these advances, diffusion-based video generation remains computationally intensive, especially for high-resolution, long-duration videos. Prior work accelerates its i

  70. Byungki Ryu, Ji Hui Son, Sungjin Park, Jaywan Chung

    This study presents a curated thermoelectric material database, teMatDb, constructed by digitizing literature-reported data. It includes temperature-dependent thermoelectric properties (TEPs), Seebeck coefficient, electrical resistivity, thermal conductivity, and figure of merit (ZT), along with metadata on materials and their corresponding publications. A s

  71. Shuyu Wang, Weiqi Li, Qian Wang, Shijie Zhao

    Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these advances, existing image editing methods still face challenges in achieving high precision and semantic accuracy in complex scenarios. Recent studies address this issue by incorpor

  72. Shengdong Han, Shangdong Yang, Xin Zhang, Yuxuan Li

    Resolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small target detection, its application to closely-spaced infrared smal

  73. Xuyang Liu, Zichen Wen, Shaobo Wang, Junjie Chen

    The advancement of large language models (LLMs) and multi-modal LLMs (MLLMs) has historically relied on scaling model parameters. However, as hardware limits constrain further model growth, the primary computational bottleneck has shifted to the quadratic cost of self-attention over increasingly long sequences by ultra-long text contexts, high-resolution ima

  74. Bowen Li, Zekun Chen, Xuefei Chen, Luhao Zhang

    A wireless wearable Electrical Impedance Tomography (EIT) system has been developed utilizing the AD5933 chip to achieve real-time imaging of lung respiration. The system employs a voltage excitation method tailored to human impedance characteristics, injecting current by applying a known voltage and measuring the resulting current through the body. Addition

  75. Weijie Su

    Large language models (LLMs) represent a new paradigm for processing unstructured data, with applications across an unprecedented range of domains. In this paper, we address, through two arguments, whether the development and application of LLMs would genuinely benefit from foundational contributions from the statistics discipline. First, we argue affirmativ

  76. Yuxuan Nie, Yutong Song, Jinjie Yang, Yupeng Song

    Drug combinations are essential in cancer therapy, leveraging synergistic drug-drug interactions (DDI) to enhance efficacy and combat resistance. However, the vast combinatorial space makes experimental screening impractical, and existing computational models struggle to capture the complex, bidirectional nature of DDIs, often relying on independent drug enc

  77. Wenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang

    Recent advances have demonstrated the effectiveness of Reinforcement Learning (RL) in improving the reasoning capabilities of Large Language Models (LLMs). However, existing works inevitably rely on high-quality instructions and verifiable rewards for effective training, both of which are often difficult to obtain in specialized domains. In this paper, we pr

  78. Tengfei Bai, Pengfei Guo, Jingshi Xu

    Let $X$ be a Banach space such that there exists a Banach space $^\ast X$ and $ ( ^\ast X )^ \ast = X $. In this paper, we introduce $X$-valued Bourgain-Morrey spaces. We show that $^\ast X$-valued block spaces are the predual of $X$-valued Bourgain-Morrey spaces. We obtain the completeness, denseness and Fatou property of $^\ast X$-valued block spaces. We g

  79. Satya N. Majumdar, Alberto Rosso

    We study a continuous time branching process where an individual splits into two daughters with rate b and dies with rate a, starting from a single individual at t=0. We show that the model can be mapped exactly to a random walk problem where the population size N(t) performs a random walk on a positive semi-infinite lattice. The hopping rate of this random

  80. Ruiwen Dong, Doron Shafrir

    Let $T$ be a positive integer, and $\mathcal{M}$ be a finitely presented module over the Laurent polynomial ring $\mathbb{Z}_{/T}[X_1^{\pm}, \ldots, X_N^{\pm}]$. We consider S-unit equations over $\mathcal{M}$: these are equations of the form $x_1 m_1 + \cdots + x_K m_K = m_0$, where the variables $x_1, \ldots, x_K$ range over the set of monomials (with coef

  81. Marion Pillas

    This study evaluates ejecta properties from multi-messenger observations to understand the absence of detectable KN associated to the four NSBH candidates from May 2023 to July 2024: we use GW public information and joint observations taken from 05.2023 to 07.2024 (LVK, ATLAS, DECam, GECKO, GOTO, GRANDMA, SAGUARO, TESS, WINTER, ZTF) in the followup of S23051

  82. Feiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng

    Our research reveals a new privacy risk associated with the vision-language model (VLM) agentic framework: the ability to infer sensitive attributes (e.g., age and health information) and even abstract ones (e.g., personality and social traits) from a set of personal images, which we term "image private attribute profiling." This threat is particularly sever

  83. Myeongseok Nam, Wongi Park, Minsol Kim, Hyejin Hur

    Recently, 3D Gaussian Splatting (3D-GS) based on Thermal Infrared (TIR) imaging has gained attention in novel-view synthesis, showing real-time rendering. However, novel-view synthesis with thermal infrared images suffers from transmission effects, emissivity, and low resolution, leading to floaters and blur effects in rendered images. To address these probl

  84. Jiahe Qin, Junpeng Li, Changchun Hua, Yana Yang

    Label Proportion Learning (LLP) addresses the classification problem where multiple instances are grouped into bags and each bag contains information about the proportion of each class. However, in practical applications, obtaining precise supervisory information regarding the proportion of instances in a specific class is challenging. To better align with r

  85. Lakshya Joshi, Arya Deshmukh, Atharv Chhabra, Chetan Gupta

    In this paper, we present algorithms to solve matrix multiplication problems in the MPC model. In particular, we consider the problem under various processor/memory constraints in the MPC model and prove the following results. 1. Multiplication of two rectangular matrices of size $d \times n$ and $n \times d$ ( where $d \leq n$) respectively can be done in,

  86. Frank Shih, Zhenghao Jiang, Faming Liang

    Uncertainty quantification (UQ) in scientific machine learning is increasingly critical as neural networks are widely adopted to tackle complex problems across diverse scientific disciplines. For physics-informed neural networks (PINNs), a prominent model in scientific machine learning, uncertainty is typically quantified using Bayesian or dropout methods. H

  87. Tengfei Bai, Pengfei Guo, Jingshi Xu

    Let $(X,\mu)$ be a space of homogeneous type satisfying $\mu(X) =\infty$, the doubling property and the reverse doubling condition. Let $L$ be a nonnegative self-adjoint operator on $L^2(X)$ whose heat kernel enjoys a Gaussian upper bound. We introduce the weighted homogeneous Bourgain-Morrey-Besov type spaces and Triebel-Lizorkin type spaces associated with

  88. Mingyu Huang, Shasha Zhou, Yuxuan Chen, Ke Li

    We are living in an era of "big literature", where the volume of digital scientific publications is growing exponentially. While offering new opportunities, this also poses challenges for understanding literature landscapes, as traditional manual reviewing is no longer feasible. Recent large language models (LLMs) have shown strong capabilities for literatur

  89. Wang Yu-Hang, Liu ying, Fang liang, Wang Xuelin

    Adversarial Training (AT) is a cornerstone defense, but many variants overlook foundational feature representations by primarily focusing on stronger attack generation. We introduce Adversarial Evolution Training (AET), a simple yet powerful framework that strategically prepends an Empirical Risk Minimization (ERM) phase to conventional AT. We hypothesize th

  90. Shang Liu, Zhongze Cai, Hanzhao Wang, Zhongyao Ma

    Human-annotated data plays a vital role in training large language models (LLMs), such as supervised fine-tuning and human preference alignment. However, it is not guaranteed that paid human annotators produce high-quality data. In this paper, we study how to incentivize human annotators to do so. We start from a principal-agent model to model the dynamics b

  91. Yan Xia, Hao Feng, Hongwei Sun, Junjie Wang

    Low-rank representation learning has emerged as a powerful tool for recovering missing values in power load data due to its ability to exploit the inherent low-dimensional structures of spatiotemporal measurements. Among various techniques, low-rank factorization models are favoured for their efficiency and interpretability. However, their performance is hig

  92. Andrei Moroianu, Mihaela Pilca

    We study conformal product structures on compact reducible Riemannian manifolds, and show that under a suitable technical assumption, the underlying Riemannian mani\-folds are either conformally flat, or triple products, \emph{i.e.} locally isometric to Riemannian manifolds of the form $(M,g)$ with $M=M_1\times M_2\times M_3$ and $g=e^{2f}g_1+g_2+g_3$, where

  93. Lea Bold, Lukas Lanza, Karl Worthmann

    We design a two-component controller to achieve reference tracking with output constraints - exemplified on systems of relative degree two. One component is a data-driven or learning-based predictive controller, which uses data samples to learn a model and predict the future behavior of the system. We exemplify this component concisely by data-enabled predic

  94. Tengfei Bai, Pengfei Guo, Jingshi Xu

    We introduce Bourgain-Morrey-Lorentz spaces and give a description of the predual of Bourgain-Morrey-Lorentz spaces via the block spaces. As an application of duality, we obtain the boundedness of Hardy-Littlewood maximal operator, sharp maximal operator, Calder\'on-Zygmund operator, fractional integral operator, commutator on Bourgain-Morrey-Lorentz spaces.

  95. Yuchao He, Mengda Wu, Yonghui Xia, Meirong Zhang

    This paper develops a methodological framework for addressing a novel and application-oriented inverse nodal problem in Sturm-Liouville operators, having significant applications in seismic wave analysis and submarine underwater radar (sonar) detection. By utilizing a given finite set of nodal data, we propose an optimization framework to find the potential

  96. Jin Zhang, Fan Gao, Linyu Li, Yongbin Yu

    The rise of large language models has led to significant performance breakthroughs in named entity recognition (NER) for high-resource languages, yet there remains substantial room for improvement in low- and medium-resource languages. Existing multilingual NER methods face severe language interference during the multi-language adaptation process, manifested

  97. Takumi Tagaki, Seiya Nishikawa, Shuji Ishihara

    We investigate the behavior of self-propelled particles under cyclic stretching, inspired by the characteristic pattern dynamics observed in microtubule (MT) motility assays subjected to uniaxial cyclic substrate stretching. We develop a self-propelled particle model that incorporates the elastic energy acting on the filaments due to substrate deformation, s

  98. Wenyang Luo, Wayne Xin Zhao, Jing Sha, Shijin Wang

    The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely focused on simpler tasks like MGSM. To address this gap, we introduce MMATH, a benchmark for multilingual complex reasoni

  99. Yuqi Liu, Qin Jin, Tianyuan Qu, Xuan Liu

    Understanding accurate atomic temporal event is essential for video comprehension. However, current video-language benchmarks often fall short to evaluate Large Multi-modal Models' (LMMs) temporal event understanding capabilities, as they can be effectively addressed using image-language models. In this paper, we introduce RTime-QA, a novel benchmark specifi

  100. Xingrui Liu, Jieming Ke, Yanlong Zhao

    This paper investigates the optimality analysis of the recursive least-squares (RLS) algorithm for autoregressive systems with exogenous inputs (ARX systems). A key challenge in analyzing is managing the potential unboundedness of the parameter estimates, which may diverge to infinity. Previous approaches addressed this issue by assuming that both the true p