Skip to content

May 2025 arXiv papers — page 124

Showing 12,30112,400 of 24,552 papers

  1. Nishant Suresh Aswani, Saif Eddin Jabari

    This paper explores a simple question: can we model the internal transformations of a neural network using dynamical systems theory? We introduce Koopman autoencoders to capture how neural representations evolve through network layers, treating these representations as states in a dynamical system. Our approach learns a surrogate model that predicts how neur

  2. Yanbin Yin, Kun Zhou, Zhen Wang, Xiangdong Zhang

    The recent explosion of large language models (LLMs), each with its own general or specialized strengths, makes scalable, reliable benchmarking more urgent than ever. Standard practices nowadays face fundamental trade-offs: closed-ended question-based benchmarks (eg MMLU) struggle with saturation as newer models emerge, while crowd-sourced leaderboards (eg C

  3. Filippo Dell'Oro

    We provide a growth bound for the operator norm of $C_0$-semigroups on Hilbert spaces under a corresponding growth bound on the resolvent of the semigroup generator. For some super-linear resolvent growths, our estimate is sharper than the ones currently available in the literature.

  4. Cameron Cornell, Lewis Mitchell, Matthew Roughan

    Causal networks offer an intuitive framework to understand influence structures within time series systems. However, the presence of cycles can obscure dynamic relationships and hinder hierarchical analysis. These networks are typically identified through multivariate predictive modelling, but enforcing acyclic constraints significantly increases computation

  5. Seanie Lee, Sangwoo Park, Dong Bok Lee, Dominik Wagner

    Low-Rank Adaptation (LoRA), which introduces a product of two trainable low-rank matrices into frozen pre-trained weights, is widely used for efficient fine-tuning of language models in federated learning (FL). However, when combined with differentially private stochastic gradient descent (DP-SGD), LoRA faces substantial noise amplification: DP-SGD perturbs

  6. Shimon Garti

    We force over a model of AD to obtain the consistency of the Galvin number having countable cofinality.

  7. Jiawen Xu, Odej Kao, Margret Keuper

    Open set recognition (OSR) is devised to address the problem of detecting novel classes during model inference. Even in recent vision models, this remains an open issue which is receiving increasing attention. Thereby, a crucial challenge is to learn features that are relevant for unseen categories from given data, for which these features might not be discr

  8. Konstantina Lelova, Gregory F. Cooper, Sofia Triantafillou

    Transporting causal information across populations is a critical challenge in clinical decision-making. Causal modeling provides criteria for identifiability and transportability, but these require knowledge of the causal graph, which rarely holds in practice. We propose a Bayesian method that combines observational data from the target domain with experimen

  9. Hieu-Nghia Huynh-Nguyen, Ngoc Son Nguyen, Huynh Nguyen Dang, Thieu Vo

    Text-to-speech (TTS) systems have seen significant advancements in recent years, driven by improvements in deep learning and neural network architectures. Viewing the output speech as a data distribution, previous approaches often employ traditional speech representations, such as waveforms or spectrograms, within the Flow Matching framework. However, these

  10. Marios H. Michael, Gunda Kipp, Alexander M. Potts, Matthew W. Day

    Two-dimensional materials and van der Waals (vdW) heterostructures host many strongly correlated and topological quantum phases on the $\sim$ meV energy scale. Direct electrodynamical signatures of such states are thus expected to appear in the terahertz (THz) frequency range (1 THz $\sim$ 4 meV). Because the typical size of vdW heterostructures ($\sim$10 $\

  11. Jingchuan Wang, Maoqi Liu, Liwang Lu, Alan Pak Tao Lau

    We comprehensively analyze the fiber nonlinearity crosstalks between DAS and communication channels through numerical results and 40 x 800-Gb/s 90-km experimental demonstration. Our findings indicate that conventional pulse-based DAS is unsuitable for in-band DWDM coexistence system, whereas pulse-compression DAS shows negligible penalties with legacy cohere

  12. Shuji Shinohara, Daiki Morita, Hayato Hirai, Ryosuke Kuribayashi

    This study proposes the novel Bayesian and inverse Bayesian (BIB) inference framework that incorporates symmetry bias into the Bayesian updating process to perform both conventional and inverse Bayesian updates concurrently. Conventional Bayesian inference is constrained by a fundamental trade-off between adaptability to abrupt environmental changes and accu

  13. Won Sang Chung, L. M. Nieto, Soroush Zare, Hassan Hassanabadi

    In this work, we explore both the ordinary $q$-Gaussian distribution and a new one defined here, determining both their mean and variance, and we use them to construct solutions of the $q$-deformed diffusion differential equation. This approach allows us to realize that the standard deviation of the distribution must be a function of time. In one case, we de

  14. Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang

    Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-Judge'' evaluators, though promising, remain constrained to specific tasks and aspects. In this paper, we argue that, on one hand, based on the interconnected nature of aspects, lea

  15. Yaguang Li

    Mass loss on the red giant branch (RGB) influences stellar evolution, properties of stellar populations, and Galactic chemical enrichment, yet remains poorly constrained observationally. Current models provide limited insight into how stellar properties, particularly how metallicity and mass, affect RGB mass loss. Here, I introduce a new observational approa

  16. Lin Zuo, Hao Yang, Bingwei Long

    We construct the three-nucleon contact potentials in the Jacobi partial-wave basis. The potentials are built in the separable form as the products of the antisymmetrized three-nucleon states in which the nucleons are arbitrarily close to each other. We compile the three-nucleon contact potentials up to $\mathcal{O}(Q^2)$. These contact potentials are by cons

  17. Wenhao Zhu, Yuhang Xie, Guojie Song, Xin Zhang

    The rapid evolution of large language models (LLMs) has revolutionized various fields, including the identification and discovery of human values within text data. While traditional NLP models, such as BERT, have been employed for this task, their ability to represent textual data is significantly outperformed by emerging LLMs like GPTs. However, the perform

  18. Yiling Tao, Shuyi Wang, Jiaxi Yang, Guido Zuccon

    This paper reports on findings from a comparative study on the effectiveness and efficiency of federated unlearning strategies within Federated Online Learning to Rank (FOLTR), with specific attention to systematically analysing the unlearning capabilities of methods in a verifiable manner. Federated approaches to ranking of search results have recently garn

  19. Yiru Jiao, Simeon C. Calvert, Sander van Cranenburgh, Hans van Lint

    Accurately and proactively alerting drivers or automated systems to emerging collisions is crucial for road safety, particularly in highly interactive and complex urban environments. Existing methods either require labour-intensive annotation of sparse risk, struggle to consider varying contextual factors, or are tailored to limited scenarios. Here we presen

  20. Biagio Ricceri

    In this note, we establish a multiplicity theorem for a nonlocal discrete problem of the type $$\cases{-\left(a\sum_{m=1}^{n+1}|x_m-x_{m-1}|^2+b\right)(x_{k+1}-2x_k+x_{k-1})=h_k(x_k)\hskip 10pt k=1,...,n, \cr & \cr x_0=x_{n+1}=0\cr}$$ assuming $a>0$ and (for the first time) $b<0$.

  21. Hemanth Saratchandran, Simon Lucey

    Transformers have transformed modern machine learning, driving breakthroughs in computer vision, natural language processing, and robotics. At the core of their success lies the attention mechanism, which enables the modeling of global dependencies among input tokens. However, we reveal that the attention block in transformers suffers from inherent ill-condi

  22. K. Kudaybergenov, A. Arziev, P. Orinbaev

    In this paper, a spectral theorem is proved for self-adjoint cyclically compact partial integral operators in the space of functions with mixed norm, which is a Kaplansky--Hilbert module. The decomposition through eigenfunctions, integral representation using orthogonal projectors, and functional calculus are established. The results generalize Mercer theore

  23. Zhongni Hou, Miao Su, Xiaolong Jin, Zixuan Li

    Temporal Knowledge Graphs (TKGs), which utilize quadruples in the form of (subject, predicate, object, timestamp) to describe temporal facts, have attracted extensive attention. N-tuple TKGs (N-TKGs) further extend traditional TKGs by utilizing n-tuples to incorporate auxiliary elements alongside core elements (i.e., subject, predicate, and object) of facts,

  24. Teemu Pennanen, Ari-Pekka Perkkiö

    This paper studies stochastic optimization problems and associated Bellman equations in formats that allow for reduced dimensionality of the cost-to-go functions. In particular, we study stochastic control problems in the ``decision-hazard-decision'' form where at each stage, the system state is controlled both by predictable as well as adapted controls. Suc

  25. Minrui Xu, Jiani Fan, Xinyu Huang, Conghao Zhou

    With the continuous evolution of Large Language Models (LLMs), LLM-based agents have advanced beyond passive chatbots to become autonomous cyber entities capable of performing complex tasks, including web browsing, malicious code and deceptive content generation, and decision-making. By significantly reducing the time, expertise, and resources, AI-assisted c

  26. Ki-Young Choi, Yu Seon Jeong, Sung Hyun Kim, Yeong Gyun Kim

    In this study, as an extension of our previous work, we estimate the sensitivity of the Search for Hidden Particles (SHiP) experiment to the 3+1 model using the charged-current deep inelastic scattering event spectrum. We employ the Feldman-Cousins method with a parametric bootstrap to account for nuisance parameters and systematic uncertainties. In the prev

  27. Ratko Darda, Takehiko Yasuda

    Let $F$ be a global field of characteristic $p > 0$ and $G$ a finite abelian $p$-group. In this paper we treat the question of counting $G$-torsors over $F$ for certain heights developed in [DY25].

  28. Mikhail A. Mikheenko

    It is a known fact that any unimodular equation over an abelian group has a solution in that group itself. It is also known that for metabelian groups this does not hold; moreover, there is a unimodular equation over some metabelian group which has no solutions in any larger metabelian group. Here we present the proof of an analagous fact for solvable groups

  29. Kai Zhang, Xingyu Chen, Xiaofeng Zhang

    Large Multimodal Models (LMMs) have become a pivotal research focus in deep learning, demonstrating remarkable capabilities in 3D scene understanding. However, current 3D LMMs employing thousands of spatial tokens for multimodal reasoning suffer from critical inefficiencies: excessive computational overhead and redundant information flows. Unlike 2D VLMs pro

  30. Jitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao

    Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key challenges: (1) information loss from hard pruning, (2) inefficient alignment of representations, and (3) underutilization of informative activations, particularly from Feed-Forward

  31. Soohwan Lee, Seoyeong Hwang, Kyungho Lee

    Recent advancements in HCI and AI have predominantly centered on individual user experiences, often neglecting the emergent dynamics of group interactions. This provocation introduces Group Experience(GX) to capture the collective perceptual, emotional, and cognitive dimensions that arise when individuals interact in cohesive groups. We challenge the convent

  32. Andreas Maurischat

    Anderson t-modules are analogs of abelian varieties in positive characteristic. Associated to such a t-module, there are its t-motive and its dual t-motive. When dealing with these objects, several questions occur which one would like to solve algorithmically. For example, for a given t-module one would like to decide whether its t-motive is indeed finitely

  33. Xueqiang Ouyang, Jia Wei

    Infertility, a pressing global health concern, affects a substantial proportion of individuals worldwide. While advancements in assisted reproductive technology (ART) have offered effective interventions, conventional in vitro fertilization-embryo transfer (IVF-ET) procedures still encounter significant hurdles in enhancing pregnancy success rates. Key chall

  34. Z. H. Cho, J. H. Han, D. H. Suk, H. J. Jeung

    We propose a novel MRI (Magnetic Resonance Imaging) technique based quantum bit (qubit) generation with water proton NMR (1H-NMR), distinct from previously proposed NMR chemical shift or spectroscopic techniques based qubit generation. We briefly review prior NMR-based techniques in the context of quantum computing, focusing on MRI-related methods. The propo

  35. Shang Li

    For a reductive group $G$ over a discretely valued Henselian field $k$, using valuations of root datum and concave functions, the Bruhat--Tits theory defines an important class of open bounded subgroups of $G(k)$ which are essential objects in representation theory and arithmetic geometry. Moreover, these subgroups are uniquely determined by smooth affine gr

  36. Kai Liang

    This paper discusses the enumeration of independent sets in king graphs of size $m \times n$, based on the tensor network contractions algorithm given in reference~\cite{tilEnum}. We transform the problem into Wang tiling enumeration within an $(m+1) \times (n+1)$ rectangle and compute the results for all cases where $m + n \leq 79$ using tensor network cont

  37. Miroslav Kolar, Daniel Sevcovic

    We investigate the motion of a family of closed curves evolving according to the geometric evolution law on a given two dimensional manifold which is embedded or immersed in the three-dimensional Euclidean space. We derive a system of nonlinear parabolic equations describing the motion of curves belonging to a given two-dimensional manifold. Using the abstra

  38. Zichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Mian

    Human motion synthesis in complex scenes presents a fundamental challenge, extending beyond conventional Text-to-Motion tasks by requiring the integration of diverse modalities such as static environments, movable objects, natural language prompts, and spatial waypoints. Existing language-conditioned motion models often struggle with scene-aware motion gener

  39. Andrey A. Shavrin

    The soft-wall holographic composite Higgs model assumes first-order phase transition from the dynamical inner symmetry breaking. This research focuses on the implications of the semi-analytical perturbative solution of the dual 5-dimensional theory as an effective description of the strongly coupled composite Higgs sector. We clarify the thermodynamical desc

  40. Junyi Hu, Tian Bai, Fengyi Wu, Zhenming Peng

    Feature fusion plays a pivotal role in achieving high performance in vision models, yet existing attention-based fusion techniques often suffer from substantial computational overhead and implementation complexity, particularly in resource-constrained settings. To address these limitations, we introduce the Plug-and-Play Hierarchical C2F Transformer (P$^2$HC

  41. Tenglong Li, Jindong Li, Guobin Shen, Dongcheng Zhao

    Spiking transformers are emerging as a promising architecture that combines the energy efficiency of Spiking Neural Networks (SNNs) with the powerful attention mechanisms of transformers. However, existing hardware accelerators lack support for spiking attention, exhibit limited throughput in exploiting fine-grained sparsity, and struggle with scalable paral

  42. Behzad Eslam Panah, Narges Heidari, Mana Soleimani, Maryam Kaveh

    This paper is motivated by the application of the inverse isoperimetric inequality to establish constraints on the parameters of gravity's rainbow. We investigate the thermodynamic (in)stability conditions for $d-$dimensional energy-dependent black holes, which are recognized as $d-$ dimensional black holes within the framework of gravity's rainbow. To achie

  43. Zhanglin Wu, Daimeng Wei, Xiaoyu Chen, Hengchao Shang

    Large language model (LLM) shows promising performances in a variety of downstream tasks, such as machine translation (MT). However, using LLMs for translation suffers from high computational costs and significant latency. Based on our evaluation, in most cases, translations using LLMs are comparable to that generated by neural machine translation (NMT) syst

  44. Chengcheng Xiang, Li Zhong, Eric Mugnier, Nathaniel Nguyen

    Access-control misconfigurations are among the main causes of today's data breaches in web applications. However, few techniques are available to support automatic and systematic testing for access-control changes and detecting risky changes to prevent severe consequences. As a result, those critical security configurations often lack testing, or are tested

  45. Yaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang

    In Text-to-SQL, execution feedback is essential for guiding large language models (LLMs) to reason accurately and generate reliable SQL queries. However, existing methods treat execution feedback solely as a post-hoc signal for correction or selection, failing to integrate it into the generation process. This limitation hinders their ability to address reaso

  46. Danqing Chen, Tobias Ladner, Ahmed Rayen Mhadhbi, Matthias Althoff

    As large language models become integral to high-stakes applications, ensuring their robustness and fairness is critical. Despite their success, large language models remain vulnerable to adversarial attacks, where small perturbations, such as synonym substitutions, can alter model predictions, posing risks in fairness-critical areas, such as gender bias mit

  47. Haibin He, Maoyuan Ye, Jing Zhang, Xiantao Cai

    Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmarks emphasize evaluating LMMs' abilities of relatively simple visual question answering, visual-text parsing, etc. However, the extent to which LMMs can deal with complex logical re

  48. Eric Palmerduca, Hong Qin

    We show that for massless helicity $h$ particles, the angular momentum eigenstates are given in an appropriate coordinate system by the spin-weighted spherical harmonics ${_{-h}Y_{jm}}$ of spin-weight $-h$. In particular, these are simultaneous eigenstates of the Hamiltonian, helicity, $J^2$, and $J_z$. The appearance of the spin-weighted spherical harmonics

  49. Maximilian Balthasar Mansky, Tobias Rohe, Gerhard Stenzel, Alejandro Bravo de la Serna

    Many computational problems are unchanged under some symmetry operation. In classical machine learning, this can be reflected with the layer structure of the neural network. In quantum machine learning, the ansatz can be tuned to correspond to the specific symmetry of the problem. We investigate this adaption of the quantum circuit to the problem symmetry on

  50. Sunghwan Kim, Dongjin Kang, Taeyoon Kwon, Hyungjoo Chae

    Reward models (RMs) play a crucial role in reinforcement learning from human feedback (RLHF), aligning model behavior with human preferences. However, existing benchmarks for reward models show a weak correlation with the performance of optimized policies, suggesting that they fail to accurately assess the true capabilities of RMs. To bridge this gap, we exp

  51. Chenlin Ming, Chendi Qu, Mengzhang Cai, Qizhi Pei

    Large Language Models (LLMs) have achieved impressive performance through Supervised Fine-tuning (SFT) on diverse instructional datasets. When training on multiple capabilities simultaneously, the mixture training dataset, governed by volumes of data from different domains, is a critical factor that directly impacts the final model's performance. Unlike many

  52. Donghwa Shin, Edwin Zhang

    Transformers have recently gained popularity in time series forecasting due to their ability to capture long-term dependencies. However, many existing models focus only on capturing temporal dependencies while omitting intricate relationships between variables. Recent models have tried tackling this by explicitly modeling both cross-time and cross-variate de

  53. Zipeng Wang, Kenan Zhang

    In this note, we obtianed hypercontractive inequalities between different weighted Bergman spaces. In addition, we establish Nikol'ski\u{\i}-type inequalities for weighted Bergman spaces with optimal constants.

  54. Sahil Tomar, Rajeshwar Tripathi, Sandeep Kumar

    Bone fractures are a leading cause of morbidity and disability worldwide, imposing significant clinical and economic burdens on healthcare systems. Traditional X ray interpretation is time consuming and error prone, while existing machine learning and deep learning solutions often demand extensive feature engineering, large, annotated datasets, and high comp

  55. Juntao Zhao, Jiuru Li, Chuan Wu

    CPUs are critical for LLM serving due to their availability, cost efficiency, and edge applicability. However, efficient CPU serving is hindered by conflicting prefill/decode resource demands under non-disaggregated deployment constraints--existing solutions fail to avoid cross-phase interference, ignore sub-NUMA hardware structures, and deliver suboptimal d

  56. Haochen Yuan, Minting Pan, Yunbo Wang, Siyu Gao

    Reinforcement learning (RL) has shown significant promise for sequential portfolio optimization tasks, such as stock trading, where the objective is to maximize cumulative returns while minimizing risks using historical data. However, traditional RL approaches often produce policies that merely memorize the optimal yet impractical buying and selling behavior

  57. Matias Quintana, Youlong Gu, Xiucheng Liang, Yujun Hou

    Understanding people's preferences is crucial for urban planning, yet current approaches often combine responses from multi-cultural populations, obscuring demographic differences and risking amplifying biases. We conducted a largescale urban visual perception survey of streetscapes worldwide using street view imagery, examining how demographics -- including

  58. Maki Nagata, Fumi Egusa, Fumiya Maeda, Kazuki Tokuda

    High-velocity clouds (HVCs), which are gas clouds moving at high velocity relative to the galactic disk, may play a critical role in galaxy evolution, potentially supplying gas to the disk and triggering star formation. In this study, we focus on the nearby face-on barred spiral galaxy M83, where high spatial resolution, high-sensitivity CO (1-0) data are av

  59. Yunsong Wei

    We study the invariant algebraic D-modules on an affine variety under the action of an algebraic group.For linear algebraic groups with the multiplication action by themselves, such D-modules correspond to representations of their Lie algebra. For unipotent algebraic groups, we show that two invariant D-modules are isomorphic if and only if they lie in the s

  60. Wenya Guo, Zhengkun Zhang, Xumeng Liu, Ying Zhang

    Instruction data selection aims to identify a high-quality subset from the training set that matches or exceeds the performance of the full dataset on target tasks. Existing methods focus on the instruction-to-response mapping, but neglect the human preference for diverse responses. In this paper, we propose Preference-oriented Data Selection method (ProDS)

  61. Martha Teiko Teye, Ori Maoz, Matthias Rottmann

    Multi-object tracking from LiDAR point clouds presents unique challenges due to the sparse and irregular nature of the data, compounded by the need for temporal coherence across frames. Traditional tracking systems often rely on hand-crafted features and motion models, which can struggle to maintain consistent object identities in crowded or fast-moving scen

  62. Zhibiao Wang, Yunlong Zhou, Ziwei Zhang, Mengmei Zhang

    Graph Transformers, leveraging the global attention to capture long-range dependencies in graph structures, have significantly advanced graph machine learning, but face prohibitive computational complexity. Tokenized Graph Learning Models (TGLMs) address this issue by converting graphs into ordered token lists for scalable processing. Besides, TGLMs also emp

  63. Daigo Nakajima, Kanji Tanaka, Daiki Iwata, Kouki Terashima

    This paper proposes MOON (Multi-Objective Optimization-driven Object-goal Navigation), a novel framework designed for efficient navigation in large-scale, complex indoor environments. While existing methods often rely on local heuristics, they frequently fail to address the strategic trade-offs between competing objectives in vast areas. To overcome this, we

  64. Filippo Leveni

    Object detection and identification is surely a fundamental topic in the computer vision field; it plays a crucial role in many applications such as object tracking, industrial robots control, image retrieval, etc. We propose a feature-based approach for detecting and identifying distorted occurrences of a given template in a scene image by incremental group

  65. Filippo Leveni

    Anomaly detection is a fundamental problem in domains such as healthcare, manufacturing, and cybersecurity. This thesis proposes new unsupervised methods for anomaly detection in both structured and streaming data settings. In the first part, we focus on structure-based anomaly detection, where normal data follows low-dimensional manifolds while anomalies de

  66. Filippo Leveni, Matteo Mistura, Francesco Iubatti, Carmine Giangregorio

    Malware are malicious programs that are grouped into families based on their penetration technique, source code, and other characteristics. Classifying malware programs into their respective families is essential for building effective defenses against cyber threats. Machine learning models have a huge potential in malware detection on mobile devices, as mal

  67. Yunsong Wei

    The various types of compactifications of symmetric spaces and locally symmetric spaces are well-studied. Among them, the De Concini-Procesi compactification, also known as the wonderful compactification, of symmetric varieties has been found to have many applications. Intuitively, this compactification provides information at infinity. The diagonal action a

  68. Hangyu Li, Qin Zhao, Haoran Xu, Xinyu Jiang

    Teleoperation is a cornerstone of embodied-robot learning, and bimanual dexterous teleoperation in particular provides rich demonstrations that are difficult to obtain with fully autonomous systems. While recent studies have proposed diverse hardware pipelines-ranging from inertial motion-capture gloves to exoskeletons and vision-based interfaces-there is st

  69. Zepeng Liu, Tianyu Liu, Hongmei Ma, Chun-Hua Yuan

    Temporal optics has attracted much attention due to its ability for lossless stretching of ultrafast temporal pulses. At the same time, spatial SU(1,1) interferometers have been widely used because of their high sensitivity to phase changes. On this basis, we studied a temporal SU(1,1) interferometer based on a temporal Fourier transform system and injected

  70. Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura, Takato Horii

    Recent studies have revealed that human emotions exhibit a high-dimensional, complex structure. A full capturing of this complexity requires new approaches, as conventional models that disregard high dimensionality risk overlooking key nuances of human emotions. Here, we examined the extent to which the latest generation of rapidly evolving Multimodal Large

  71. Dong Kyu Cho, Inwoo Hwang, Sanghack Lee

    Data augmentation is a popular tool for single source domain generalization, which expands the source domain by generating simulated ones, improving generalization on unseen target domains. In this work, we show that the performance of such augmentation-based methods in the target domains universally fluctuates during training, posing challenges in model sel

  72. Weiliang Tang, Dong Jing, Jia-Hui Pan, Zhiwu Lu

    Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. Our key insight is that these emerging abilities can naturally extend to robotic manipulation by enabling LMMs to directly infer the next goal in language via reasoning, rather than

  73. Yeseul Jeon, Rajarshi Guhaniyogi, Aaron Scheffler

    This article addresses the challenge of modeling the amplitude of spatially indexed low frequency fluctuations (ALFF) in resting state functional MRI as a function of cortical structural features and a multi-task coactivation network in the Adolescent Brain Cognitive Development (ABCD) Study. It proposes a generative model that integrates effects of spatiall

  74. Jinhua Zhang, Wei Long, Minghao Han, Weiyi You

    Essential to visual generation is efficient modeling of visual data priors. Conventional next-token prediction methods define the process as learning the conditional probability distribution of successive tokens. Recently, next-scale prediction methods redefine the process to learn the distribution over multi-scale representations, significantly reducing gen

  75. Cheng Yuan, Yufei Jiang, Xu Zhu

    We propose a multi-reference and adaptive nonlinear transform source-channel coding (MA-NTSCC) system for wireless image semantic transmission to improve rate-distortion (RD) performance by introducing multi-dimensional contexts into the entropy model of the state-of-the-art (SOTA) NTSCC system. Improvements in RD performance of the proposed MA-NTSCC system

  76. Tuan A. Hoang, Thanh V. Pham, Chuyen T. Nguyen

    This paper studies the performance of physical layer security (PLS) in a multi-user hybrid heterogeneous visible light communication (VLC) and radio frequency (RF) wireless communication system with simultaneous lightwave information and power transfer (SLIPT). In the considered system, VLC is used for downlink (DL) while RF is employed for uplink (UL) trans

  77. Chenghua Gong, Rui Sun, Yuhao Zheng, Juyuan Zhang

    Advanced epidemic forecasting is critical for enabling precision containment strategies, highlighting its strategic importance for public health security. While recent advances in Large Language Models (LLMs) have demonstrated effectiveness as foundation models for domain-specific tasks, their potential for epidemic forecasting remains largely unexplored. In

  78. Hongjoon Ahn, Heewoong Choi, Jisu Han, Taesup Moon

    Offline goal-conditioned reinforcement learning (GCRL) offers a practical learning paradigm in which goal-reaching policies are trained from abundant state-action trajectory datasets without additional environment interaction. However, offline GCRL still struggles with long-horizon tasks, even with recent advances that employ hierarchical policy structures,

  79. Zeyi Ren, Jingreng Lei, Yichen Jin, Ermo Hua

    The development of edge computing places critical demands on energy-efficient model deployment for multiple-input multiple-output (MIMO) detection tasks. Deploying deep unfolding models such as PGD-Nets and ADMM-Nets into resource-constrained edge devices using quantization methods is challenging. Existing quantization methods based on quantization aware tra

  80. Andrés Ahumada Gómez, Mauricio Che, Manuel Cuerno

    We present theoretical properties of the space of metric pairs equipped with the Gromov--Hausdorff distance. First, we establish the classical metric separability and the geometric geodesicity of this space. Second, we prove an Arzel\`a--Ascoli-type theorem for metric pairs. Third, extending a result by Cassorla, we show that the set of pairs consisting of a

  81. Junbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang

    Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to reconstruct real-world landscapes conditioned on environmental soundscapes. To address this gap, we present Geo-contextual Soundscape-to-Landscape (GeoS2L) generation, a novel and pra

  82. Xin-Wei Yi, Wei Li, Jing-Yang You, Bo Gu

    Recent strain-stabilized superconductivity at ambient pressure in La$_3$Ni$_2$O$_{7}$ films opens new avenues for nickelates research, in parallel with its pressure-induced counterpart. Using density functional theory calculations, we elucidate the critical factors bridging strain- and pressure-driven superconductivity in La$_3$Ni$_2$O$_{7}$ by comprehensive

  83. Chensen Lin, Ruian Tie, Shihong Yi, Xiaohui Zhong

    High-resolution wind information is essential for wind energy planning and power forecasting, particularly in regions with complex terrain. However, most AI-based weather forecasting models operate at kilometer-scale resolution, constrained by the reanalysis datasets they are trained on. Here we introduce FuXi-CFD, an AI-based downscaling framework designed

  84. Jie Ou, Jinyu Guo, Shuaihong Jiang, Zhaokun Wang

    Retrieval-augmented generation (RAG) has emerged as a pivotal method for expanding the knowledge of large language models. To handle complex queries more effectively, researchers developed Adaptive-RAG (A-RAG) to enhance the generated quality through multiple interactions with external knowledge bases. Despite its effectiveness, A-RAG exacerbates the pre-exi

  85. Zeyu Huang, Markus Rupp, Stefan Schwarz

    In this work, we present a recent investigation on leveraging large reconfigurable intelligent surfaces (RIS) as anchors for positioning in wireless communication systems. Unlike existing approaches, we explicitly address the uncertainty arising from the substantial physical size of the RIS, particularly relevant when a user equipment resides in the near fie

  86. Zhaoyang Li, Qianqian Yang, Zehui Xiong, Zhiguo Shi

    Accurate channel prediction is essential in massive multiple-input multiple-output (m-MIMO) systems to improve precoding effectiveness and reduce the overhead of channel state information (CSI) feedback. However, existing methods often suffer from accumulated prediction errors and poor generalization to dynamic wireless environments. Large language models (L

  87. Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou

    Large language models and large multimodal models (LLMs and LMMs) deliver strong generative performance but suffer from slow decoding, a problem that becomes more severe when handling visual inputs, whose sequences typically contain many more tokens with lower information density than text. Speculative decoding accelerates LLM inference by letting a compact

  88. Han Meng, Yancan Chen, Yunan Li, Yitian Yang

    Mental-health stigma remains a pervasive social problem that hampers treatment-seeking and recovery. Existing resources for training neural models to finely classify such stigma are limited, relying primarily on social-media or synthetic data without theoretical underpinnings. To remedy this gap, we present an expert-annotated, theory-informed corpus of huma

  89. Xiaoyuan Wang, Fredric S. Cohen, Shixin Xu, Yongqiang Cai

    Cholesterol is known to modulate the structure and function of biological membranes. In this study, we use self-consistent field theory (SCFT) to investigate phospholipid/cholesterol bilayer membranes modeled with two types of diblock copolymers. These copolymer-based bilayers serve as biomimetic platforms with applications in areas such as drug delivery. Ou

  90. Yangyang Xu, Hongyu Zhao, Chengzhong Zhang, Chenglin Liao

    This letter proposes a fractional-order battery model based on the Caputo definition. A closed-form time-domain solution is derived, enabling a simple recursive expression for discrete-time implementation. The model requires only the current and previous time-step states in each iteration, significantly reducing memory usage compared to the conventional Gr\"

  91. Hyunseong Kim, Gyunghyun Jang, Seungwon Jin, Dongbin Shin

    The Josephson junction is the fundamental nonlinear building block of superconducting quantum technologies. Its macroscopic quantum tunneling physics underpins superconducting quantum computing, sensing, and communication, but scaling these platforms to utility-scale architectures places increasingly stringent demands on junction materials, interfaces, and f

  92. Jia Xu Wei

    Modern comparison sorts like quicksort suffer from performance inconsistencies due to suboptimal pivot selection, leading to $(O(N^2))$ worst-case complexity, while in-place merge sort variants face challenges with data movement overhead. We introduce Wave Sort, a novel in-place sorting algorithm that addresses these limitations through a dynamic pivot selec

  93. Haoyuan Wu, Rui Ming, Jilong Gao, Hangyu Zhao

    Large language models (LLMs) achieve remarkable performance in code generation tasks. However, a significant performance disparity persists between popular programming languages (e.g., Python, C++) and others. To address this capability gap, we leverage the code translation task to train LLMs, thereby facilitating the transfer of coding proficiency across di

  94. Louis H Kauffman

    This paper introduces a new algebra, the crossing algebra, that is applied to count the number of components for arborescent knots, links, tangles or states (of a state polynomial expansion such as the Kauffman bracket). This algebra is foundational, and it is related to generalisations of boolean logic and to aspects of foundations based in diagrams and net

  95. V. I. Danilov

    In 1962, Gale and Shapley \cite{GS} introduced the concept of stable marriages and proved their existence. Since then, the statement of the stability problem has been highly generalized. And a lot of proofs has emerged for the existence in these more general statements. It's time to review them and identify the similarities and differences. First, we will br

  96. Yulin Gong, Rachel Bean

    We apply two machine learning methods, a CNN deep-leaning model and a gradient-boosting decision tree, to estimate individual cluster optical depths from observed properties derived from multiple complementary datasets. The models are trained and tested with simulated N-body derived halo catalogs and synthetic full-sky CMB maps designed to mirror data from t

  97. Chenxuan Zhang, Qingwen Wu, Xiao Fan, Luis C. Ho

    James Webb Space Telescope (JWST) has revealed a new class of high-redshift, very red, compact broad-line sources, termed as "little red dots" (LRDs). The physical mechanism driving these properties remains elusive. We construct spectral energy distributions (SEDs) with spectroscopic redshift for 28 LRDs and find they exhibit V-shaped SEDs with a common brea

  98. Jingyang Peng, Wenyuan Shen, Jiarui Rao, Jionghao Lin

    Recent advances in Generative Artificial Intelligence (GenAI) have transformed educational content creation, particularly in developing tutor training materials. However, biases embedded in AI-generated content--such as gender, racial, or national stereotypes--raise significant ethical and educational concerns. Despite the growing use of GenAI, systematic me

  99. Haoyuan Wu, Xueyi Chen, Rui Ming, Jilong Gao

    Large language models (LLMs) demonstrate significant reasoning capabilities, particularly through long chain-of-thought (CoT) processes, which can be elicited by reinforcement learning (RL). However, prolonged CoT reasoning presents limitations, primarily verbose outputs due to excessive introspection. The reasoning process in these LLMs often appears to fol

  100. Taiqiang Wu, Runming Yang, Jiayi Li, Pengfei Hu

    Large language models (LLMs) consistently benefit from further fine-tuning on various tasks. However, we observe that directly tuning the Instruct (i.e., instruction-tuned) models often leads to marginal improvements and even performance degeneration. Notably, paired Base models, the foundation for these Instruct variants, contain highly similar weight value