Skip to content

May 2025 arXiv papers — page 114

Showing 11,30111,400 of 24,552 papers

  1. Yanheng He, Jiahe Jin, Pengfei Liu

    Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesi

  2. Ajitesh Bankula, Praney Bankula

    Cross-lingual transfer has become a crucial aspect of multilingual NLP, as it allows for models trained on resource-rich languages to be applied to low-resource languages more effectively. Recently massively multilingual pre-trained language models (e.g., mBERT, XLM-R) demonstrate strong zero-shot transfer capabilities[14] [13]. This paper investigates cross

  3. Junyu Luo, Yusheng Zhao, Xiao Luo, Zhiping Xiao

    Unsupervised efficient domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, while maintaining low storage cost and high retrieval efficiency. However, existing methods typically fail to address potential noise in the target domain, and directly align high-level features across domains, thus resultin

  4. Soyabul Islam Lincoln, Mirza Mohd Shahriar Maswood

    A common neurodegenerative disease, Alzheimer's disease requires a precise diagnosis and efficient treatment, particularly in light of escalating healthcare expenses and the expanding use of artificial intelligence in medical diagnostics. Many recent studies shows that the combination of brain Magnetic Resonance Imaging (MRI) and deep neural networks have ac

  5. Ruihan Liu, Xiaoyi Wu, Xijun Chen, Liang Hu

    A comprehensive understanding of 3D scenes is essential for autonomous vehicles (AVs), and among various perception tasks, occupancy estimation plays a central role by providing a general representation of drivable and occupied space. However, most existing occupancy estimation methods rely on LiDAR or cameras, which perform poorly in degraded environments s

  6. Ruixiao Li, Fahao Chen, Peng Li

    Speculative decoding accelerates Large Language Model (LLM) inference by employing a small speculative model (SSM) to generate multiple candidate tokens and verify them using the LLM in parallel. This technique has been widely integrated into LLM inference serving systems. However, inference requests typically exhibit uncertain execution time, which poses a

  7. Fu Luo, Xi Lin, Mengyuan Zhong, Fei Liu

    Neural Combinatorial Optimisation (NCO) is a promising learning-based approach for solving Vehicle Routing Problems (VRPs) without extensive manual design. While existing constructive NCO methods typically follow an appending-based paradigm that sequentially adds unvisited nodes to partial solutions, this rigid approach often leads to suboptimal results. To

  8. Chengyu Shen, Zhen Hao Wong, Runming He, Hao Liang

    Large Language Models (LLMs) have recently achieved remarkable progress in mathematical reasoning. To enable such capabilities, many existing works distill strong reasoning models into long chains of thought or design algorithms to construct high-quality math question-answer (QA) data for training. However, these efforts primarily focus on generating correct

  9. Naoki Hayashi, Takuro Kutsuna, Sawa Takamuku

    In statistical learning, models are classified as regular or singular depending on whether the mapping from parameters to probability distributions is injective. Most models with hierarchical structures or latent variables are singular, for which conventional criteria such as the Akaike Information Criterion and the Bayesian Information Criterion are inappli

  10. S. Alipour, A. T. Rezakhani, Alireza Tavanfar, K. Mölmer

    Quantum computing employs controllable interactions to perform sequences of logical gates and entire algorithms on quantum registers. This paradigm has been widely explored, e.g., for simulating dynamics of manybody systems by decomposing their Hamiltonian evolution in a series of quantum gates. Here, we introduce a method for quantum simulation in which the

  11. Zhanpeng Zhou, Yongyi Yang, Mahito Sugiyama, Junchi Yan

    Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distinct phase transition, yet our understanding of this transition is still incomplete. In this paper, we introduce an interval-wise perspective that compares network states across a ti

  12. Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo, Emily Wenger

    Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties and metrics on which these models align, surprisingly little work has explored causes of this similarity. To advance this line of inquiry, this work explores how two factors - datase

  13. Róbert Csordás, Christopher D. Manning, Christopher Potts

    Modern LLMs are increasingly deep, and depth correlates with performance, albeit with diminishing returns. However, do these models use their depth efficiently? Do they compose more features to create higher-order computations that are impossible in shallow models, or do they merely spread the same kinds of computation out over more layers? To address these

  14. Ramon van den Akker, Bas J. M. Werker, Bo Zhou

    We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized experiments across various fields. While algorithms for maximizing rewards or, equivalently, minimizing regret have received considerable attention, our focus centers on statistica

  15. Yingwei Zhang, Ke Bu, Zhuoran Zhuang, Tao Xie

    The past decades witness the significant advancements in time series forecasting (TSF) across various real-world domains, including e-commerce and disease spread prediction. However, TSF is usually constrained by the uncertainty dilemma of predicting future data with limited past observations. To settle this question, we explore the use of Cross-Future Behav

  16. Yuning Jiang, Feiyang Shang, Freedy Tan Wei You, Huilin Wang

    The dynamic landscape of cybersecurity demands precise and scalable solutions for vulnerability management in heterogeneous systems, where configuration-specific vulnerabilities are often misidentified due to inconsistent data in databases like the National Vulnerability Database (NVD). Inaccurate Common Platform Enumeration (CPE) data in NVD further leads t

  17. Chenxi Liu, Tianyi Xiong, Yanshuo Chen, Ruibo Chen

    The task adaptation and alignment of Large Multimodal Models (LMMs) have been significantly advanced by instruction tuning and further strengthened by recent preference optimization. Yet, most LMMs still suffer from severe modality imbalance during reasoning, i.e., outweighing language prior biases over visual inputs, which bottlenecks their generalization t

  18. Jiangxia Cao, Pengbo Xu, Yin Cheng, Kaiwei Guo

    In this paper, we provide our milestone ensemble sort work and the first-hand practical experience, Pantheon, which transforms ensemble sorting from a "human-curated art" to a "machine-optimized science". Compared with formulation-based ensemble sort, our Pantheon has the following advantages: (1) Personalized Joint Training: our Pantheon is jointly trained

  19. Victor Valencia Torres

    The quark-gluon plasma (QGP) produced in ultrarelativistic heavy-ion collisions has exhibited properties of a mostly perfect fluid. These properties can be observed through the hydrodynamic expansion of the QGP. Experimentally, this was established by measuring azimuthal anisotropies in the final state, known as elliptic flow ($v_2$) or higher order harmonic

  20. Jin Li, Meijin Li, Nan Yang, Li Wang

    As one of the primary detection targets for contemporary gravitational wave (GW) observatories, the stochastic gravitational wave background (SGWB) holds significant potential for enhancing our understanding of the early universe's formation and evolution. Studies indicate that the SGWB spectrum from cosmic strings can span an extraordinarily broad frequency

  21. Zhen Xiong, Yujun Cai, Zhecheng Li, Yiwei Wang

    Recent advances in test-time scaling have enabled Large Language Models (LLMs) to display sophisticated reasoning abilities via extended Chain-of-Thought (CoT) generation. Despite their potential, these Reasoning LLMs (RLMs) often demonstrate counterintuitive and unstable behaviors, such as performance degradation under few-shot prompting, that challenge our

  22. Yiting Zhang, Shichen Li

    Manipulating deformable linear objects (DLOs) is challenging due to their complex dynamics and the need for safe interaction in contact-rich environments. Most existing models focus on shape prediction alone and fail to account for contact and tension constraints, which can lead to damage to both the DLO and the robot. In this work, we propose a certifiably

  23. Ji Zhang, Shihan Wu, Xu Luo, Hao Wu

    Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correlate task-irrelevant visual features with actions, limiting t

  24. Junyang Wang, Haiyang Xu, Xi Zhang, Ming Yan

    The exponential rise in mobile device usage necessitates streamlined automation for effective task management, yet many AI frameworks fall short due to inadequate operational expertise. While manually written knowledge can bridge this gap, it is often burdensome and inefficient. We introduce Mobile-Agent-V, an innovative framework that utilizes video as a gu

  25. Jingqi Tong, Jixin Tang, Hangcheng Li, Yurong Mou

    Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherently provide rich visual elements and mechanics that are easy t

  26. Huiliang Zhang, Di Wu, Arnaud Zinflou, Stephane Dellacherie

    Multivariate Time Series (MTS) forecasting has a wide range of applications in both industry and academia. Recent advances in Spatial-Temporal Graph Neural Network (STGNN) have achieved great progress in modelling spatial-temporal correlations. Limited by computational complexity, most STGNNs for MTS forecasting focus primarily on short-term and local spatia

  27. Ming-Yang Li, Chen-Xun Weng, Wen-Bo Liu, Mengya Zhu

    Accurate and tamper-resistant timestamps are essential for applications demanding verifiable chronological ordering, such as legal documentation and digital intellectual property protection. Classical timestamp protocols rely on computational assumptions for security, rendering them vulnerable to quantum attacks, which is a critical limitation given the rapi

  28. Shaocong Pei, A-Man Zhang, Chang Liu, Tianyuan Zhang

    We investigate the influence of ambient temperature on the dynamics of spark-generated cavitation bubbles over a broad temperature range of 23 to 90$^\circ \text{C}$. Increasing temperature, the attenuation of collapse intensity of a bubble in a free field is quantitatively characterised through the Rayleigh factor, minimum bubble volume, and maximum collaps

  29. Benjamin Spetzler, Elizaveta Spetzler, Saba Zamankhani, Dilara Abdel

    Modeling hysteretic switching dynamics in memristive devices is computationally demanding due to coupled ionic and electronic transport processes. This challenge is particularly relevant for emerging two-dimensional (2D) devices, which feature high-dimensional design spaces that remain largely unexplored. We introduce a physics-guided modeling framework that

  30. Maysam Rabbani, Zahra Akbari

    The literature documents the effects of the pandemic on birthrate, birthweight, and pregnancy complications. This study contributes to this growing body of research by examining multiple facets of the phenomenon. Using the 2012-2022 hospital inpatient discharge data of New York, we implemented fixed-effects regression models and reported three key findings.

  31. Jiahao Yu, Haozhuang Liu, Yeqiu Yang, Lu Chen

    Regression models are crucial in recommender systems. However, retransformation bias problem has been conspicuously neglected within the community. While many works in other fields have devised effective bias correction methods, all of them are post-hoc cures externally to the model, facing practical challenges when applied to real-world recommender systems.

  32. Ziqian Wang, Xianjun Xia, Xinfa Zhu, Lei Xie

    The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often

  33. Eric T. Johnson, Michael Zingale

    We investigate the properties of mixed H/He flames in X-ray bursts using 2D hydrodynamic simulations. We find that as the initial hydrogen abundance of the atmosphere increases, the flame is less energetic and propagates slower. The simulation outcome, whether a flame forms and whether there's runaway burning at the base of the atmosphere, is very sensitive

  34. Yi-Fei Xia, Zi-Xiang Xu, Yu-Ting Yan, An Chen

    Non-Hermitian systems have recently shown new possibilities to manipulate wave scattering by exploiting loss, yet coherent perfect absorption at an exceptional point (CPA EP) remains elusive in acoustics. Here we demonstrate it based on a two-channel waveguide with compact lossy resonators. We realize imbalanced losses crucial for CPA EP by using active comp

  35. Ming-Chen Sun, Shi-Han Zhao, Rui-Xuan Gao, He-Sheng Liu

    Atmospheric neutrinos (ATNs) offer a paradigm for understanding neutrino properties, while it is critical to quantify uncertainties in flux modeling. Since ATNs are produced simultaneously with cosmic ray muons, precision measurements of cosmic ray muons, including arrival direction, energy spectra, and spin polarization, will help reduce ATN production unce

  36. Yu-Sik Kim, Jae-Young Kim

    Recent studies of individual track-like TeV-PeV IceCube neutrino events suggest that strongly jetted AGNs, blazars, can be plausible sources of extragalactic high-energy neutrinos. Although the broadband emission and neutrinos from such blazars can be modeled by hadronic jets with inverse Compton processes, various models show degeneracies. One of the reason

  37. Lanlan Kang, Jian Wang, Jian QIn, Yiqin Liang

    The ThinPrep Cytologic Test (TCT) is the most widely used method for cervical cancer screening, and the sample quality directly impacts the accuracy of the diagnosis. Traditional manual evaluation methods rely on the observation of pathologist under microscopes. These methods exhibit high subjectivity, high cost, long duration, and low reliability. With the

  38. Peisong Niu, Ziqing Ma, Tian Zhou, Weiqi Chen

    Weather forecasting has long posed a significant challenge for humanity. While recent AI-based models have surpassed traditional numerical weather prediction (NWP) methods in global forecasting tasks, overfitting remains a critical issue due to the limited availability of real-world weather data spanning only a few decades. Unlike fields like computer vision

  39. Jingzheng Li, Tiancheng Wang, Xingyu Peng, Jiacheng Chen

    Autonomous Driving (AD) systems demand the high levels of safety assurance. Despite significant advancements in AD demonstrated on open-source benchmarks like Longest6 and Bench2Drive, existing datasets still lack regulatory-compliant scenario libraries for closed-loop testing to comprehensively evaluate the functional safety of AD. Meanwhile, real-world AD

  40. John Harding, Remi Salinas Schmeis

    We provide two results. The first gives a finite graph constructed from consideration of mutually unbiased bases that occurs as a subgraph of the orthogonality space of $\mathbb{C}^3$ but not of that of $\mathbb{R}^3$. The second is a companion result to the result of Tau and Tserunyan \cite{Tau} that every countable graph occurs as an induced subgraph of th

  41. Pei-Yao Chen, Chen Wang, Fang Yan, Chao-Yang Zhang

    Understanding the propagation and attenuation patterns of ground vibrations is critical for evaluating the impact of environmental disturbances on large-scale scientific facilities. However, complex site conditions often result in intricate vibration behaviors, limiting the accuracy of traditional predictive methods. This study proposes a hybrid iterative fi

  42. Chengda Song, Jing He, Guanghui Yuan

    Numerous vector angular spectrum methods have been presented to model the vectorial nature of diffractive electromagnetic field, facilitating optical field engineering in polarization-related and high numerical aperture systems. However, balancing accuracy and efficiency in state-of-the-art vector methods is challenging, especially with not well-defined inci

  43. Maysam Rabbani, Elijah Gervais

    Recent literature reports mixed evidence on whether birthweight has decreased during the pandemic. In this paper, we use New York's hospital inpatient discharge data and contribute to this ongoing debate in multiple ways. First, we corroborate that birthweight has declined during the pandemic by 7g (grams). Second, we provide the first empirical evidence tha

  44. Yi Zhang, Wenfu Xu, Zhiqiang Tan

    For sensitivity analysis against unmeasured confounding, we build on the marginal sensitivity model (MSM) and propose a new model, deMSM, by incorporating a second constraint on the shift of potential outcome distributions caused by unmeasured confounders in addition to the constraint on the shift of treatment probabilities. We show that deMSM leads to inter

  45. Haoyu Wang, Zhi Sun, Shuangfeng Han, Xiaoyun Wang

    Frequency-domain channel extrapolation is effective in reducing pilot overhead for massive multiple-input multiple-output (MIMO) systems. Recently, Deep learning (DL) based channel extrapolator has become a promising candidate for modeling complex frequency-domain dependency. Nevertheless, current DL extrapolators fail to operate in unseen environments under

  46. Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim

    Recent reasoning-focused language models achieve high accuracy by generating lengthy intermediate reasoning paths before producing final answers. While this approach is effective in solving problems that require logical thinking, long reasoning paths significantly increase memory usage and reduce throughput of token generation, limiting the practical deploym

  47. Xue Dong, Xuexing Lu, Yu Ye

    An upward planar order on an acyclic directed graph $G$ is a special linear extension of the edge poset of $G$ that satisfies the nesting condition. This order was introduced to combinatorially characterize upward plane graphs and progressive plane graphs (commonly known as plane string diagrams). In this paper, motivated by the theory of graphical calculus

  48. Sevvandi Kandanaarachchi, Cheng Soon Ong

    Social networks have a small number of large hubs, and a large number of small dense communities. We propose a generative model that captures both hub and dense structures. Based on recent results about graphons on line graphs, our model is a graphon mixture, enabling us to generate sequences of graphs where each graph is a combination of sparse and dense gr

  49. Z. H. Zhang, L. G. Wang

    The distance signless Laplacian matrix of a graph $G$ is define as $Q(G)=$Tr$(G)+D(G)$, where Tr$(G)$ and $D(G)$ are the diagonal matrix of vertex transmissions and the distance matrix of $G$, respectively. Denote by $E_G(v)$ the set of all edges incident to a vertex $v$ in $G$. A fractional matching of a graph $G$ is a function $f:E(G) \rightarrow [0,1]$ su

  50. Guobin Shen, Dongcheng Zhao, Linghao Feng, Xiang He

    Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit harmful outputs. Despite growing efforts in LLM safety research, existing evaluations are often fragmented, focused on isolated attack or defense techniques, and lack systematic, rep

  51. Musharraf N. Alruwaill, Saraju P. Mohanty, Elias Kougianos

    The growing utilization of Internet of Medical Things (IoMT) devices, including smartwatches and wearable medical devices, has facilitated real-time health monitoring and data analysis to enhance healthcare outcomes. These gadgets necessitate improved security measures to safeguard sensitive health data while tackling scalability issues in real-time settings

  52. Tiancheng Jiang, Henry Wang, Md Sirajus Salekin, Parmida Atighehchian

    Vision Language Models (VLMs) have demonstrated strong performance in multi-modal tasks by effectively aligning visual and textual representations. However, most video understanding VLM research has been domain-agnostic, leaving the understanding of their transfer learning capability to specialized domains under-explored. In this work, we address this by exp

  53. Junchao Hong, Jiayan Yang, Yi Bi, Bo Yang

    Coronal bright points are typical small-scale coronal brightenings that consist of a bundle of miniature coronal loops. Using the ultra-high-resolution coronal images from the Extreme Ultraviolet Image onboard Solar Obiter, we report the first observational evidence of oscillatory magnetic reconnection at a coronal bright point (CBP). The reconnection is cha

  54. Gonzalo E. Constante-Flores, Hao Chen, Can Li

    Deep learning models are increasingly deployed in safety-critical tasks where predictions must satisfy hard constraints, such as physical laws, fairness requirements, or safety limits. However, standard architectures lack built-in mechanisms to enforce such constraints, and existing approaches based on regularization or projection are often limited to simple

  55. Tian Sun, Yuqi Chen, Baihua Zheng, Weiwei Sun

    In real-world applications, GPS trajectories often suffer from low sampling rates, with large and irregular intervals between consecutive GPS points. This sparse characteristic presents challenges for their direct use in GPS-based systems. This paper addresses the task of map-constrained trajectory recovery, aiming to enhance trajectory sampling rates of GPS

  56. Ruqin Zhou, Chenguang Dai, Wanshou Jiang, Yongsheng Zhang

    Vectorized HD map is essential for autonomous driving. Remarkable work has been achieved in recent years, but there are still major issues: (1) in the generation of the BEV features, single modality-based methods are of limited perception capability, while direct concatenation-based multi-modal methods fail to capture synergies and disparities between differ

  57. Arihant Tripathi, Liam Dugan, Charis Gao, Maggie Huan

    As state-of-the-art language models continue to improve, the need for robust detection of machine-generated text becomes increasingly critical. However, current state-of-the-art machine text detectors struggle to adapt to new unseen domains and generative models. In this paper we present DoGEN (Domain Gating Ensemble Networks), a technique that allows detect

  58. Yu Zheng

    The proliferation of artificial intelligence has enabled a diversity of applications that bridge the gap between digital and physical worlds. As physical environments are too complex to model through a single information acquisition approach, it is crucial to fuse multimodal data generated by different sources, such as sensors, devices, systems, and people,

  59. Ruihao Zheng, Jingda Deng, Zhenkun Wang

    The weak Pareto boundary ($WPB$) refers to a boundary in the objective space of a multi-objective optimization problem, characterized by weak Pareto optimality rather than Pareto optimality. The $WPB$ brings severe challenges to multi-objective evolutionary algorithms (MOEAs), as it may mislead the algorithms into finding dominance-resistant solutions (DRSs)

  60. Rutwig Campoamor-Stursberg, Francisco J. Herranz, Javier de Lucas

    The Lie-Hamilton approach for $t$-dependent Hamiltonians is extended to cover the so-called nonlinear Lie-Hamilton systems, which are no longer related to a linear $t$-dependent combination of a basis of a finite-dimensional Lie algebra of functions $\mathcal{W}$, but an arbitrary $t$-dependent function on $\mathcal{W}$. This novel formalism is accomplished

  61. Yusheng Zhao, Chi Zhang, Yuxuan Du

    Characterizing the ground state properties of quantum systems is fundamental to capturing their behavior but computationally challenging. Recent advances in AI have introduced novel approaches, with diverse machine learning (ML) and deep learning (DL) models proposed for this purpose. However, the necessity and specific role of DL models in these tasks remai

  62. Sahil Shah, Harsh Goel, Sai Shankar Narasimhan, Minkyu Choi

    Modern video understanding systems excel at tasks such as scene classification, object detection, and short video retrieval. However, as video analysis becomes increasingly central to real-world applications, there is a growing need for proactive video agents for the systems that not only interpret video streams but also reason about events and take informed

  63. Guo-Ping Li, Meng-Qi Wu, Ke-Jian He, Qing-Quan Jiang

    In this paper, based on the action of a complex scalar field minimally coupled to a gravitational field, we numerically obtain a series of massive boson star solutions in a spherically symmetric background with a quartic-order self-interaction potential. Then, considering a thin accretion flow with a certain four-velocity, we further investigate the observab

  64. Paratat Bejrakarbum, Paolo Bertozzini, Supaporn Theesoongnern

    We investigate the notion of involutive weak cubical $\omega$-categories via Penon's approach: as algebras for the monad induced by the free involutive strict $\omega$-category functor on cubical $\omega$-sets. A few examples of involutive weak cubical $\omega$-categories are provided.

  65. Melissa Lee, Anthony Pisani

    Given a permutation group $G \le \mathrm{Sym}(\Omega)$, a subset $B$ of $\Omega$ is said to be a base if its pointwise stabiliser in $G$ is trivial, and the base size $b(G)$ is the minimum size of a base. In the notable case $b(G) = 2$, Burness and Giudici define the Saxl graph of $G$ to be the graph on $\Omega$ with bases of size 2 as edges. Later work of F

  66. Amal Raj, Vivek Balachandran

    Quantum computing leverages quantum mechanics to achieve computational advantages over classical hardware, but the use of third-party quantum compilers in the Noisy Intermediate-Scale Quantum (NISQ) era introduces risks of intellectual property (IP) exposure. We address this by proposing a novel obfuscation technique that protects proprietary quantum circuit

  67. Tianle Yang, Chengzhe Sun, Siwei Lyu, Phil Rose

    This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected to be more difficult for deepfake models to replicate. The results demonstrate that certain segmental features commonly

  68. Satoru Shinoda, Hideaki Kawaguchi

    In clinical biomarker studies, the Dynamic Network Biomarker (DNB) is sometimes used. DNB is a composite variable derived from the variance and the Pearson correlation coefficient of biological signals. When applying DNB to clinical data, it is important to account for confounding bias. However, little attention has been paid to statistical causal inference

  69. Gaole Dai, Yuhong Zhou, Jun Wang, Zhuo Li

    Hydrodynamic cloaking offers a promising approach for manipulating viscous flows by redirecting fluid around an obstacle without inducing external disturbances. By extending pseudo-conformal mappings into potential flow models, we introduce a new isobaric boundary condition that enables the construction of zero-index cloaks using isotropic and homogeneous me

  70. Congchi Yin, Yongpeng Zhang, Xuyun Wen, Piji Li

    Associative memory engages in the integration of relevant information for comprehension in the human cognition system. In this work, we seek to improve alignment between language models and human brain while processing speech information by integrating associative memory. After verifying the alignment between language model and brain by mapping language mode

  71. Yang Xiang, Canan Huang, Desheng Hu, Jingguang Tian

    Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectivenes

  72. Yunho Jin, Gu-Yeon Wei, David Brooks

    Scaling large language models (LLMs) has driven significant advancements, yet it faces diminishing returns and escalating energy demands. This work explores how test-time compute (TTC) can serve as an energy-efficient complement to conventional scaling strategies by allocating additional computational resources at inference time rather than during training.

  73. Antonio Joia Neto, Norrathep Rattanavipanon, Ivan De Oliveira Nunes

    Embedded devices are increasingly ubiquitous and vital, often supporting safety-critical functions. However, due to strict cost and energy constraints, they are typically implemented with Micro-Controller Units (MCUs) that lack advanced architectural security features. Within this space, recent efforts have created low-cost architectures capable of generatin

  74. Hannah Potgieter, Razvan C. Fetecau, Steven J. Ruuth

    We use the $p$-Laplacian with large $p$-values in order to approximate geodesic distances to features on surfaces. This differs from Fayolle and Belyaev's (2018) [1] computational results using the $p$-Laplacian for the distance-to-surface problem. Our approach appears to offer some distinct advantages over other popular PDE-based distance function approxima

  75. Yixuan Gao, Xiongkuo Min, Guangtao Zhai

    This paper explores how the image quality assessment (IQA) task affects the cognitive processes of people from the perspective of pupil size and studies the relationship between pupil size and image quality. Specifically, we first invited subjects to participate in a subjective experiment, which includes two tasks: free observation and IQA. In the free obser

  76. Zhengqing Yuan, Weixiang Sun, Yixin Liu, Huichi Zhou

    Large Language Models (LLMs) have driven significant progress, yet their growing parameter counts and context windows incur prohibitive compute, energy, and monetary costs. We introduce EfficientLLM, a novel benchmark and the first comprehensive empirical study evaluating efficiency techniques for LLMs at scale. Conducted on a production-class cluster (48xGH

  77. Zhenyu Bao, Qing Li, Guibiao Liao, Zhongyuan Zhao

    3D Gaussian Splatting (3DGS) has gained significant attention in streamable dynamic novel view synthesis (DNVS) for its photorealistic rendering capability and computational efficiency. Despite much progress in improving rendering quality and optimization strategies, 3DGS-based streamable dynamic scene reconstruction still suffers from flickering artifacts a

  78. Zhenyao Li, Shengwen Liao, Qian Zhang, Xuechun Zhang

    Integration of renewable resources is profoundly reshaping the dynamics of modern power systems. This study shows that the voltage dynamics of power systems with multiple grid-forming (GFM) converters often enjoys a desirable property called input-output monotonicity. A systematic approach for computing the derivatives of the voltage subsystem is presented f

  79. Weihao Zou, Weibing Feng, Pin Wu

    This study proposes a universal flow field prediction framework based on knowledge transfer from large language model (LLM), addressing the high computational costs of traditional computational fluid dynamics (CFD) methods and the limited cross-condition transfer capability of existing deep learning models. The framework innovatively integrates Proper Orthog

  80. Gokul Puthumanaillam, Paulo Padrao, Jose Fuentes, Leonardo Bobadilla

    Robots navigating complex environments must manage uncertainty from sensor noise, environmental changes, and incomplete information, with different tasks requiring varying levels of precision in different areas. For example, precise localization may be crucial near obstacles but less critical in open spaces. We present GUIDE (Generalized Uncertainty Integrat

  81. Jerry Tang, Ruiqi Zhang, Kaan Beyduz, Yiwei Jiang

    This paper presents Duawlfin, a drone with unified actuation for wheeled locomotion and flight operation that achieves efficient, bidirectional ground mobility. Unlike existing hybrid designs, Duawlfin eliminates the need for additional actuators or propeller-driven ground propulsion by leveraging only its standard quadrotor motors and introducing a differen

  82. Yawen Gao, Changsheng Chen, Mingbo Li, Chao Sun

    Nano-clustering occurs in the monophasic "pre-Ouzo" region of ternary liquid mixtures without the use of surfactants. This study is proposed to elucidate the nucleation and stability of multiscale nanodomains in a surfactant-free microemulsion (SFME) system composed of trans-anethol, ethanol and water, tuned by aqueous ionic environment. We examined direct-

  83. Zhi Su, Yuman Gao, Emily Lukas, Yunfei Li

    Achieving coordinated teamwork among legged robots requires both fine-grained locomotion control and long-horizon strategic decision-making. Robot soccer offers a compelling testbed for this challenge, combining dynamic, competitive, and multi-agent interactions. In this work, we present a hierarchical multi-agent reinforcement learning (MARL) framework that

  84. Ningning Yao, Huan Xi, Lang Chen, Zhe Song

    Despite policymakers deploying various tools to mitigate emissions of ozone (O\textsubscript{3}) precursors, such as nitrogen oxides (NO\textsubscript{x}), carbon monoxide (CO), and volatile organic compounds (VOCs), the effectiveness of policy combinations remains uncertain. We employ an integrated framework that couples structural break detection with mach

  85. Faizuddin Ahmed, Ahmad Al-Badawi, Izzet Sakallı

    In a recent article (Ref. \cite{AOP}), the authors obtained a static, cylindrically symmetric Anti-de Sitter (AdS) black string (BS) solutions, which are cylindrical generalizations of black holes (BHs), surrounded by a cloud of strings (CS) and the quintessence field (QF), and discussed its properties. In the present study, we present a comprehensive analys

  86. Jacob E. Turner, Timothy Dolch, Paul B. Demorest, Ryan S. Lynch

    We explore possible advantages of cyclic spectroscopy for observations of pulsars in instances where full cyclic deconvolution is not feasible. We compute cyclic merits and full-deconvolution regime boundaries for pulsars observed by NANOGrav and discuss which sources stand to benefit the most from using cyclic spectroscopy when observed with the Green Bank

  87. Zongyuan Deng, Yujie Cai, Qing Liu, Shiyao Mu

    The selection of base station sites is a critical challenge in 5G network planning, which requires efficient optimization of coverage, cost, user satisfaction, and practical constraints. Traditional manual methods, reliant on human expertise, suffer from inefficiencies and are limited to an unsatisfied planning-construction consistency. Existing AI tools, de

  88. Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai

    Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS,

  89. Xu Ge, Roman Verba, Philipp Pirro, Andrii V. Chumak

    Dipolar coupling between closely spaced magnetic waveguides enables the design of magnonic directional couplers - universal devices capable of functioning as signal combiners, power splitters, demultiplexers, and more. The wavelength-dependent coupling, combined with the weak nonlinear variation of a spin wave's wavelength at constant-frequency, introduces p

  90. Kiarash Naghavi Khanghah, Zhiling Chen, Lela Romeo, Qian Yang

    Additive manufacturing enables the fabrication of complex designs while minimizing waste, but faces challenges related to defects and process anomalies. This study presents a novel multimodal Retrieval-Augmented Generation-based framework that automates anomaly detection across various Additive Manufacturing processes leveraging retrieved information from li

  91. Yuqing Hou, Yiyin Cao, Chuangyin Dang, Yong Wang

    The sequence form, owing to its compact and holistic strategy representation, has demonstrated significant efficiency in computing normal-form perfect equilibria for two-player extensive-form games with perfect recall. Nevertheless, the examination of $n$-player games remains underexplored. To tackle this challenge, we present a sequence-form characterizatio

  92. Yafeng Chen, Chong Deng, Hui Wang, Yiheng Jiang

    Developing robust speaker verification (SV) systems without speaker labels has been a longstanding challenge. Earlier research has highlighted a considerable performance gap between self-supervised and fully supervised approaches. In this paper, we enhance the non-contrastive self-supervised framework, Self-Distillation Prototypes Network (SDPN), by introduc

  93. Anurag Mishra

    Mechanistic interpretability research seeks to reveal the inner workings of large language models, yet most work focuses on classification or generative tasks rather than summarization. This paper presents an interpretability framework for analyzing how GPT-like models adapt to summarization tasks. We conduct differential analysis between pre-trained and fin

  94. Liang-Duan Liu, Yu-Hao Zhang, Yun-Wei Yu, Ze-Xin Du

    Modeling the light curves (LCs) of luminous astronomical transients, such as supernovae, is crucial for understanding their progenitor physics, particularly with the exponential growth of survey data. However, existing methods face limitations: efficient semi-analytical models (e.g., Arnett-like) employ significant physical simplifications (like time-invaria

  95. David X. Lin, Daniel Hall, Giannis Fikioris, Siddhartha Banerjee

    We study the problem of fair online resource allocation via non-monetary mechanisms, where multiple agents repeatedly share a resource without monetary transfers. Previous work has shown that every agent can guarantee $1/2$ of their ideal utility (the highest achievable utility given their fair share of resources) robustly, i.e., under arbitrary behavior by

  96. Hiroyuki Hayashi

    We consider ruled surfaces with finite multiplicity. We study behaviors of the striction curves and the singularities of the ruled surfaces. We also give geometric meanings of invariants related to the ruled surfaces.

  97. Hikmat Khan, Ziyu Su, Huina Zhang, Yihong Wang

    Triple-negative breast cancer (TNBC) remains a major clinical challenge due to its aggressive behavior and lack of targeted therapies. Accurate early prediction of response to neoadjuvant chemotherapy (NACT) is essential for guiding personalized treatment strategies and improving patient outcomes. In this study, we present an attention-based multiple instanc

  98. Shao-Hsuan Wang, Hsin-Hsiung Huang

    Ultra-high-dimensional tensor predictors are increasingly common in neuroimaging and other biomedical studies, yet existing methods rarely integrate continuous, count, and binary responses in a single coherent model. We present a Bayesian Sparse Kronecker Product Decomposition (BSKPD) that represents each regression (or classification) coefficient tensor as

  99. Ram Mohan Rao Kadiyala, Siddhant Gupta, Jebish Purbey, Srishti Yadav

    Vision-Language Models (VLMs) have demonstrated impressive capabilities across a range of tasks, yet concerns about their potential biases exist. This work investigates the extent to which prominent VLMs exhibit cultural biases by evaluating their performance on an image-based country identification task at a country level. Utilizing the geographically diver

  100. Lucas Rosenblatt, Bin Han, Robert Wolfe, Bill Howe

    Large language models (LLMs) can leak sensitive training data through memorization and membership inference attacks. Prior work has primarily focused on strong adversarial assumptions, including attacker access to entire samples or long, ordered prefixes, leaving open the question of how vulnerable LLMs are when adversaries have only partial, unordered sampl