May 2025 arXiv papers — page 114
Showing 11,301–11,400 of 24,552 papers
Yanheng He, Jiahe Jin, Pengfei Liu
Scaling up high-quality trajectory data has long been a critical bottleneck for developing human-like computer use agents. We introduce PC Agent-E, an efficient agent training framework that significantly reduces reliance on large-scale human demonstrations. Starting with just 312 human-annotated computer use trajectories, we further augment them by synthesi
Ajitesh Bankula, Praney Bankula
Cross-lingual transfer has become a crucial aspect of multilingual NLP, as it allows for models trained on resource-rich languages to be applied to low-resource languages more effectively. Recently massively multilingual pre-trained language models (e.g., mBERT, XLM-R) demonstrate strong zero-shot transfer capabilities[14] [13]. This paper investigates cross
Junyu Luo, Yusheng Zhao, Xiao Luo, Zhiping Xiao
Unsupervised efficient domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, while maintaining low storage cost and high retrieval efficiency. However, existing methods typically fail to address potential noise in the target domain, and directly align high-level features across domains, thus resultin
XDementNET: An Explainable Attention Based Deep Convolutional Network to Detect Alzheimer Progression from MRI data
eess.IVSoyabul Islam Lincoln, Mirza Mohd Shahriar Maswood
A common neurodegenerative disease, Alzheimer's disease requires a precise diagnosis and efficient treatment, particularly in light of escalating healthcare expenses and the expanding use of artificial intelligence in medical diagnostics. Many recent studies shows that the combination of brain Magnetic Resonance Imaging (MRI) and deep neural networks have ac
Ruihan Liu, Xiaoyi Wu, Xijun Chen, Liang Hu
A comprehensive understanding of 3D scenes is essential for autonomous vehicles (AVs), and among various perception tasks, occupancy estimation plays a central role by providing a general representation of drivable and occupied space. However, most existing occupancy estimation methods rely on LiDAR or cameras, which perform poorly in degraded environments s
Ruixiao Li, Fahao Chen, Peng Li
Speculative decoding accelerates Large Language Model (LLM) inference by employing a small speculative model (SSM) to generate multiple candidate tokens and verify them using the LLM in parallel. This technique has been widely integrated into LLM inference serving systems. However, inference requests typically exhibit uncertain execution time, which poses a
Fu Luo, Xi Lin, Mengyuan Zhong, Fei Liu
Neural Combinatorial Optimisation (NCO) is a promising learning-based approach for solving Vehicle Routing Problems (VRPs) without extensive manual design. While existing constructive NCO methods typically follow an appending-based paradigm that sequentially adds unvisited nodes to partial solutions, this rigid approach often leads to suboptimal results. To
Chengyu Shen, Zhen Hao Wong, Runming He, Hao Liang
Large Language Models (LLMs) have recently achieved remarkable progress in mathematical reasoning. To enable such capabilities, many existing works distill strong reasoning models into long chains of thought or design algorithms to construct high-quality math question-answer (QA) data for training. However, these efforts primarily focus on generating correct
Naoki Hayashi, Takuro Kutsuna, Sawa Takamuku
In statistical learning, models are classified as regular or singular depending on whether the mapping from parameters to probability distributions is injective. Most models with hierarchical structures or latent variables are singular, for which conventional criteria such as the Akaike Information Criterion and the Bayesian Information Criterion are inappli
S. Alipour, A. T. Rezakhani, Alireza Tavanfar, K. Mölmer
Quantum computing employs controllable interactions to perform sequences of logical gates and entire algorithms on quantum registers. This paradigm has been widely explored, e.g., for simulating dynamics of manybody systems by decomposing their Hamiltonian evolution in a series of quantum gates. Here, we introduce a method for quantum simulation in which the
Zhanpeng Zhou, Yongyi Yang, Mahito Sugiyama, Junchi Yan
Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distinct phase transition, yet our understanding of this transition is still incomplete. In this paper, we introduce an interval-wise perspective that compares network states across a ti
Zeyu Michael Li, Hung Anh Vu, Damilola Awofisayo, Emily Wenger
Numerous works have noted similarities in how machine learning models represent the world, even across modalities. Although much effort has been devoted to uncovering properties and metrics on which these models align, surprisingly little work has explored causes of this similarity. To advance this line of inquiry, this work explores how two factors - datase
Róbert Csordás, Christopher D. Manning, Christopher Potts
Modern LLMs are increasingly deep, and depth correlates with performance, albeit with diminishing returns. However, do these models use their depth efficiently? Do they compose more features to create higher-order computations that are impossible in shallow models, or do they merely spread the same kinds of computation out over more layers? To address these
Ramon van den Akker, Bas J. M. Werker, Bo Zhou
We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized experiments across various fields. While algorithms for maximizing rewards or, equivalently, minimizing regret have received considerable attention, our focus centers on statistica
Yingwei Zhang, Ke Bu, Zhuoran Zhuang, Tao Xie
The past decades witness the significant advancements in time series forecasting (TSF) across various real-world domains, including e-commerce and disease spread prediction. However, TSF is usually constrained by the uncertainty dilemma of predicting future data with limited past observations. To settle this question, we explore the use of Cross-Future Behav
Yuning Jiang, Feiyang Shang, Freedy Tan Wei You, Huilin Wang
The dynamic landscape of cybersecurity demands precise and scalable solutions for vulnerability management in heterogeneous systems, where configuration-specific vulnerabilities are often misidentified due to inconsistent data in databases like the National Vulnerability Database (NVD). Inaccurate Common Platform Enumeration (CPE) data in NVD further leads t
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
cs.LGChenxi Liu, Tianyi Xiong, Yanshuo Chen, Ruibo Chen
The task adaptation and alignment of Large Multimodal Models (LMMs) have been significantly advanced by instruction tuning and further strengthened by recent preference optimization. Yet, most LMMs still suffer from severe modality imbalance during reasoning, i.e., outweighing language prior biases over visual inputs, which bottlenecks their generalization t
Jiangxia Cao, Pengbo Xu, Yin Cheng, Kaiwei Guo
In this paper, we provide our milestone ensemble sort work and the first-hand practical experience, Pantheon, which transforms ensemble sorting from a "human-curated art" to a "machine-optimized science". Compared with formulation-based ensemble sort, our Pantheon has the following advantages: (1) Personalized Joint Training: our Pantheon is jointly trained
Victor Valencia Torres
The quark-gluon plasma (QGP) produced in ultrarelativistic heavy-ion collisions has exhibited properties of a mostly perfect fluid. These properties can be observed through the hydrodynamic expansion of the QGP. Experimentally, this was established by measuring azimuthal anisotropies in the final state, known as elliptic flow ($v_2$) or higher order harmonic
The constraints on the stochastic gravitational wave background from cosmic strings by an electromagnetic resonance system
gr-qcJin Li, Meijin Li, Nan Yang, Li Wang
As one of the primary detection targets for contemporary gravitational wave (GW) observatories, the stochastic gravitational wave background (SGWB) holds significant potential for enhancing our understanding of the early universe's formation and evolution. Studies indicate that the SGWB spectrum from cosmic strings can span an extraordinarily broad frequency
Zhen Xiong, Yujun Cai, Zhecheng Li, Yiwei Wang
Recent advances in test-time scaling have enabled Large Language Models (LLMs) to display sophisticated reasoning abilities via extended Chain-of-Thought (CoT) generation. Despite their potential, these Reasoning LLMs (RLMs) often demonstrate counterintuitive and unstable behaviors, such as performance degradation under few-shot prompting, that challenge our
Certifiably Safe Manipulation of Deformable Linear Objects via Joint Shape and Tension Prediction
cs.ROYiting Zhang, Shichen Li
Manipulating deformable linear objects (DLOs) is challenging due to their complex dynamics and the need for safe interaction in contact-rich environments. Most existing models focus on shape prediction alone and fail to account for contact and tension constraints, which can lead to damage to both the DLO and the robot. In this work, we propose a certifiably
Ji Zhang, Shihan Wu, Xu Luo, Hao Wu
Leveraging pretrained Vision-Language Models (VLMs) to map language instruction and visual observations to raw low-level actions, Vision-Language-Action models (VLAs) hold great promise for achieving general-purpose robotic systems. Despite their advancements, existing VLAs tend to spuriously correlate task-irrelevant visual features with actions, limiting t
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
cs.AIJunyang Wang, Haiyang Xu, Xi Zhang, Ming Yan
The exponential rise in mobile device usage necessitates streamlined automation for effective task management, yet many AI frameworks fall short due to inadequate operational expertise. While manually written knowledge can bridge this gap, it is often burdensome and inefficient. We introduce Mobile-Agent-V, an innovative framework that utilizes video as a gu
Jingqi Tong, Jixin Tang, Hangcheng Li, Yurong Mou
Vision-language reinforcement learning (RL) has primarily focused on narrow domains (e.g. geometry or chart reasoning). This leaves broader training scenarios and resources underexplored, limiting the exploration and learning of Vision Language Models (VLMs) through RL. We find video games inherently provide rich visual elements and mechanics that are easy t
Huiliang Zhang, Di Wu, Arnaud Zinflou, Stephane Dellacherie
Multivariate Time Series (MTS) forecasting has a wide range of applications in both industry and academia. Recent advances in Spatial-Temporal Graph Neural Network (STGNN) have achieved great progress in modelling spatial-temporal correlations. Limited by computational complexity, most STGNNs for MTS forecasting focus primarily on short-term and local spatia
Ming-Yang Li, Chen-Xun Weng, Wen-Bo Liu, Mengya Zhu
Accurate and tamper-resistant timestamps are essential for applications demanding verifiable chronological ordering, such as legal documentation and digital intellectual property protection. Classical timestamp protocols rely on computational assumptions for security, rendering them vulnerable to quantum attacks, which is a critical limitation given the rapi
Shaocong Pei, A-Man Zhang, Chang Liu, Tianyuan Zhang
We investigate the influence of ambient temperature on the dynamics of spark-generated cavitation bubbles over a broad temperature range of 23 to 90$^\circ \text{C}$. Increasing temperature, the attenuation of collapse intensity of a bubble in a free field is quantitatively characterised through the Rayleigh factor, minimum bubble volume, and maximum collaps
Physics-Guided Sequence Modeling for Fast Simulation and Design Exploration of 2D Memristive Devices
cond-mat.mtrl-sciBenjamin Spetzler, Elizaveta Spetzler, Saba Zamankhani, Dilara Abdel
Modeling hysteretic switching dynamics in memristive devices is computationally demanding due to coupled ionic and electronic transport processes. This challenge is particularly relevant for emerging two-dimensional (2D) devices, which feature high-dimensional design spaces that remain largely unexplored. We introduce a physics-guided modeling framework that
Maysam Rabbani, Zahra Akbari
The literature documents the effects of the pandemic on birthrate, birthweight, and pregnancy complications. This study contributes to this growing body of research by examining multiple facets of the phenomenon. Using the 2012-2022 hospital inpatient discharge data of New York, we implemented fixed-effects regression models and reported three key findings.
TranSUN: A Preemptive Paradigm to Eradicate Retransformation Bias Intrinsically from Regression Models in Recommender Systems
cs.IRJiahao Yu, Haozhuang Liu, Yeqiu Yang, Lu Chen
Regression models are crucial in recommender systems. However, retransformation bias problem has been conspicuously neglected within the community. While many works in other fields have devised effective bias correction methods, all of them are post-hoc cures externally to the model, facing practical challenges when applied to real-world recommender systems.
Ziqian Wang, Xianjun Xia, Xinfa Zhu, Lei Xie
The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as speech, general audio events, and music. Furthermore, their exclusive reliance on cross-entropy loss for alignment often
Eric T. Johnson, Michael Zingale
We investigate the properties of mixed H/He flames in X-ray bursts using 2D hydrodynamic simulations. We find that as the initial hydrogen abundance of the atmosphere increases, the flame is less energetic and propagates slower. The simulation outcome, whether a flame forms and whether there's runaway burning at the base of the atmosphere, is very sensitive
Yi-Fei Xia, Zi-Xiang Xu, Yu-Ting Yan, An Chen
Non-Hermitian systems have recently shown new possibilities to manipulate wave scattering by exploiting loss, yet coherent perfect absorption at an exceptional point (CPA EP) remains elusive in acoustics. Here we demonstrate it based on a two-channel waveguide with compact lossy resonators. We realize imbalanced losses crucial for CPA EP by using active comp
Ming-Chen Sun, Shi-Han Zhao, Rui-Xuan Gao, He-Sheng Liu
Atmospheric neutrinos (ATNs) offer a paradigm for understanding neutrino properties, while it is critical to quantify uncertainties in flux modeling. Since ATNs are produced simultaneously with cosmic ray muons, precision measurements of cosmic ray muons, including arrival direction, energy spectra, and spin polarization, will help reduce ATN production unce
Yu-Sik Kim, Jae-Young Kim
Recent studies of individual track-like TeV-PeV IceCube neutrino events suggest that strongly jetted AGNs, blazars, can be plausible sources of extragalactic high-energy neutrinos. Although the broadband emission and neutrinos from such blazars can be modeled by hadronic jets with inverse Compton processes, various models show degeneracies. One of the reason
Automated Quality Evaluation of Cervical Cytopathology Whole Slide Images Based on Content Analysis
eess.IVLanlan Kang, Jian Wang, Jian QIn, Yiqin Liang
The ThinPrep Cytologic Test (TCT) is the most widely used method for cervical cancer screening, and the sample quality directly impacts the accuracy of the diagnosis. Traditional manual evaluation methods rely on the observation of pathologist under microscopes. These methods exhibit high subjectivity, high cost, long duration, and low reliability. With the
Utilizing Strategic Pre-training to Reduce Overfitting: Baguan -- A Pre-trained Weather Forecasting Model
cs.LGPeisong Niu, Ziqing Ma, Tian Zhou, Weiqi Chen
Weather forecasting has long posed a significant challenge for humanity. While recent AI-based models have surpassed traditional numerical weather prediction (NWP) methods in global forecasting tasks, overfitting remains a critical issue due to the limited availability of real-world weather data spanning only a few decades. Unlike fields like computer vision
Jingzheng Li, Tiancheng Wang, Xingyu Peng, Jiacheng Chen
Autonomous Driving (AD) systems demand the high levels of safety assurance. Despite significant advancements in AD demonstrated on open-source benchmarks like Longest6 and Bench2Drive, existing datasets still lack regulatory-compliant scenario libraries for closed-loop testing to comprehensively evaluate the functional safety of AD. Meanwhile, real-world AD
John Harding, Remi Salinas Schmeis
We provide two results. The first gives a finite graph constructed from consideration of mutually unbiased bases that occurs as a subgraph of the orthogonality space of $\mathbb{C}^3$ but not of that of $\mathbb{R}^3$. The second is a companion result to the result of Tau and Tserunyan \cite{Tau} that every countable graph occurs as an induced subgraph of th
Formula-Guided Machine Learning for Ground Vibration Propagation and Attenuation Modeling
physics.app-phPei-Yao Chen, Chen Wang, Fang Yan, Chao-Yang Zhang
Understanding the propagation and attenuation patterns of ground vibrations is critical for evaluating the impact of environmental disturbances on large-scale scientific facilities. However, complex site conditions often result in intricate vibration behaviors, limiting the accuracy of traditional predictive methods. This study proposes a hybrid iterative fi
Generic full-vector angular spectrum method for calculating diffraction of arbitrary electromagnetic fields
physics.opticsChengda Song, Jing He, Guanghui Yuan
Numerous vector angular spectrum methods have been presented to model the vectorial nature of diffractive electromagnetic field, facilitating optical field engineering in polarization-related and high numerical aperture systems. However, balancing accuracy and efficiency in state-of-the-art vector methods is challenging, especially with not well-defined inci
Maysam Rabbani, Elijah Gervais
Recent literature reports mixed evidence on whether birthweight has decreased during the pandemic. In this paper, we use New York's hospital inpatient discharge data and contribute to this ongoing debate in multiple ways. First, we corroborate that birthweight has declined during the pandemic by 7g (grams). Second, we provide the first empirical evidence tha
Yi Zhang, Wenfu Xu, Zhiqiang Tan
For sensitivity analysis against unmeasured confounding, we build on the marginal sensitivity model (MSM) and propose a new model, deMSM, by incorporating a second constraint on the shift of potential outcome distributions caused by unmeasured confounders in addition to the constraint on the shift of treatment probabilities. We show that deMSM leads to inter
Haoyu Wang, Zhi Sun, Shuangfeng Han, Xiaoyun Wang
Frequency-domain channel extrapolation is effective in reducing pilot overhead for massive multiple-input multiple-output (MIMO) systems. Recently, Deep learning (DL) based channel extrapolator has become a promising candidate for modeling complex frequency-domain dependency. Nevertheless, current DL extrapolators fail to operate in unseen environments under
Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim
Recent reasoning-focused language models achieve high accuracy by generating lengthy intermediate reasoning paths before producing final answers. While this approach is effective in solving problems that require logical thinking, long reasoning paths significantly increase memory usage and reduce throughput of token generation, limiting the practical deploym
Xue Dong, Xuexing Lu, Yu Ye
An upward planar order on an acyclic directed graph $G$ is a special linear extension of the edge poset of $G$ that satisfies the nesting condition. This order was introduced to combinatorially characterize upward plane graphs and progressive plane graphs (commonly known as plane string diagrams). In this paper, motivated by the theory of graphical calculus
Sevvandi Kandanaarachchi, Cheng Soon Ong
Social networks have a small number of large hubs, and a large number of small dense communities. We propose a generative model that captures both hub and dense structures. Based on recent results about graphons on line graphs, our model is a graphon mixture, enabling us to generate sequences of graphs where each graph is a combination of sparse and dense gr
On the distance signless Laplacian spectral radius, fractional matching and factors of graphs
math.COZ. H. Zhang, L. G. Wang
The distance signless Laplacian matrix of a graph $G$ is define as $Q(G)=$Tr$(G)+D(G)$, where Tr$(G)$ and $D(G)$ are the diagonal matrix of vertex transmissions and the distance matrix of $G$, respectively. Denote by $E_G(v)$ the set of all edges incident to a vertex $v$ in $G$. A fractional matching of a graph $G$ is a function $f:E(G) \rightarrow [0,1]$ su
Guobin Shen, Dongcheng Zhao, Linghao Feng, Xiang He
Large language models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial prompts known as jailbreaks, which can bypass safety alignment and elicit harmful outputs. Despite growing efforts in LLM safety research, existing evaluations are often fragmented, focused on isolated attack or defense techniques, and lack systematic, rep
hChain 4.0: A Secure and Scalable Permissioned Blockchain for EHR Management in Smart Healthcare
cs.CRMusharraf N. Alruwaill, Saraju P. Mohanty, Elias Kougianos
The growing utilization of Internet of Medical Things (IoMT) devices, including smartwatches and wearable medical devices, has facilitated real-time health monitoring and data analysis to enhance healthcare outcomes. These gadgets necessitate improved security measures to safeguard sensitive health data while tackling scalability issues in real-time settings
Tiancheng Jiang, Henry Wang, Md Sirajus Salekin, Parmida Atighehchian
Vision Language Models (VLMs) have demonstrated strong performance in multi-modal tasks by effectively aligning visual and textual representations. However, most video understanding VLM research has been domain-agnostic, leaving the understanding of their transfer learning capability to specialized domains under-explored. In this work, we address this by exp
Junchao Hong, Jiayan Yang, Yi Bi, Bo Yang
Coronal bright points are typical small-scale coronal brightenings that consist of a bundle of miniature coronal loops. Using the ultra-high-resolution coronal images from the Extreme Ultraviolet Image onboard Solar Obiter, we report the first observational evidence of oscillatory magnetic reconnection at a coronal bright point (CBP). The reconnection is cha
Gonzalo E. Constante-Flores, Hao Chen, Can Li
Deep learning models are increasingly deployed in safety-critical tasks where predictions must satisfy hard constraints, such as physical laws, fairness requirements, or safety limits. However, standard architectures lack built-in mechanisms to enforce such constraints, and existing approaches based on regularization or projection are often limited to simple
Tian Sun, Yuqi Chen, Baihua Zheng, Weiwei Sun
In real-world applications, GPS trajectories often suffer from low sampling rates, with large and irregular intervals between consecutive GPS points. This sparse characteristic presents challenges for their direct use in GPS-based systems. This paper addresses the task of map-constrained trajectory recovery, aiming to enhance trajectory sampling rates of GPS
Ruqin Zhou, Chenguang Dai, Wanshou Jiang, Yongsheng Zhang
Vectorized HD map is essential for autonomous driving. Remarkable work has been achieved in recent years, but there are still major issues: (1) in the generation of the BEV features, single modality-based methods are of limited perception capability, while direct concatenation-based multi-modal methods fail to capture synergies and disparities between differ
Arihant Tripathi, Liam Dugan, Charis Gao, Maggie Huan
As state-of-the-art language models continue to improve, the need for robust detection of machine-generated text becomes increasingly critical. However, current state-of-the-art machine text detectors struggle to adapt to new unseen domains and generative models. In this paper we present DoGEN (Domain Gating Ensemble Networks), a technique that allows detect
Yu Zheng
The proliferation of artificial intelligence has enabled a diversity of applications that bridge the gap between digital and physical worlds. As physical environments are too complex to model through a single information acquisition approach, it is crucial to fuse multimodal data generated by different sources, such as sensors, devices, systems, and people,
Ruihao Zheng, Jingda Deng, Zhenkun Wang
The weak Pareto boundary ($WPB$) refers to a boundary in the objective space of a multi-objective optimization problem, characterized by weak Pareto optimality rather than Pareto optimality. The $WPB$ brings severe challenges to multi-objective evolutionary algorithms (MOEAs), as it may mislead the algorithms into finding dominance-resistant solutions (DRSs)
Nonlinear Lie-Hamilton systems: $t$-Dependent curved oscillators and Kepler-Coulomb Hamiltonians
math-phRutwig Campoamor-Stursberg, Francisco J. Herranz, Javier de Lucas
The Lie-Hamilton approach for $t$-dependent Hamiltonians is extended to cover the so-called nonlinear Lie-Hamilton systems, which are no longer related to a linear $t$-dependent combination of a basis of a finite-dimensional Lie algebra of functions $\mathcal{W}$, but an arbitrary $t$-dependent function on $\mathcal{W}$. This novel formalism is accomplished
Yusheng Zhao, Chi Zhang, Yuxuan Du
Characterizing the ground state properties of quantum systems is fundamental to capturing their behavior but computationally challenging. Recent advances in AI have introduced novel approaches, with diverse machine learning (ML) and deep learning (DL) models proposed for this purpose. However, the necessity and specific role of DL models in these tasks remai
Sahil Shah, Harsh Goel, Sai Shankar Narasimhan, Minkyu Choi
Modern video understanding systems excel at tasks such as scene classification, object detection, and short video retrieval. However, as video analysis becomes increasingly central to real-world applications, there is a growing need for proactive video agents for the systems that not only interpret video streams but also reason about events and take informed
Guo-Ping Li, Meng-Qi Wu, Ke-Jian He, Qing-Quan Jiang
In this paper, based on the action of a complex scalar field minimally coupled to a gravitational field, we numerically obtain a series of massive boson star solutions in a spherically symmetric background with a quartic-order self-interaction potential. Then, considering a thin accretion flow with a certain four-velocity, we further investigate the observab
Paratat Bejrakarbum, Paolo Bertozzini, Supaporn Theesoongnern
We investigate the notion of involutive weak cubical $\omega$-categories via Penon's approach: as algebras for the monad induced by the free involutive strict $\omega$-category functor on cubical $\omega$-sets. A few examples of involutive weak cubical $\omega$-categories are provided.
Melissa Lee, Anthony Pisani
Given a permutation group $G \le \mathrm{Sym}(\Omega)$, a subset $B$ of $\Omega$ is said to be a base if its pointwise stabiliser in $G$ is trivial, and the base size $b(G)$ is the minimum size of a base. In the notable case $b(G) = 2$, Burness and Giudici define the Saxl graph of $G$ to be the graph on $\Omega$ with bases of size 2 as edges. Later work of F
Amal Raj, Vivek Balachandran
Quantum computing leverages quantum mechanics to achieve computational advantages over classical hardware, but the use of third-party quantum compilers in the Noisy Intermediate-Scale Quantum (NISQ) era introduces risks of intellectual property (IP) exposure. We address this by proposing a novel obfuscation technique that protects proprietary quantum circuit
Tianle Yang, Chengzhe Sun, Siwei Lyu, Phil Rose
This study explores the potential of using acoustic features of segmental speech sounds to detect deepfake audio. These features are highly interpretable because of their close relationship with human articulatory processes and are expected to be more difficult for deepfake models to replicate. The results demonstrate that certain segmental features commonly
Extension of Dynamic Network Biomarker using the propensity score method: Simulation of causal effects on variance and correlation coefficient
stat.MESatoru Shinoda, Hideaki Kawaguchi
In clinical biomarker studies, the Dynamic Network Biomarker (DNB) is sometimes used. DNB is a composite variable derived from the variance and the Pearson correlation coefficient of biological signals. When applying DNB to clinical data, it is important to account for confounding bias. However, little attention has been paid to statistical causal inference
Unidirectional zero-index and omnidirectional hybrid hydrodynamic cloaks constructed from isotropic media with anisotropic geometry
physics.flu-dynGaole Dai, Yuhong Zhou, Jun Wang, Zhuo Li
Hydrodynamic cloaking offers a promising approach for manipulating viscous flows by redirecting fluid around an obstacle without inducing external disturbances. By extending pseudo-conformal mappings into potential flow models, we introduce a new isobaric boundary condition that enables the construction of zero-index cloaks using isotropic and homogeneous me
Congchi Yin, Yongpeng Zhang, Xuyun Wen, Piji Li
Associative memory engages in the integration of relevant information for comprehension in the human cognition system. In this work, we seek to improve alignment between language models and human brain while processing speech information by integrating associative memory. After verifying the alignment between language model and brain by mapping language mode
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
eess.ASYang Xiang, Canan Huang, Desheng Hu, Jingguang Tian
Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectivenes
Yunho Jin, Gu-Yeon Wei, David Brooks
Scaling large language models (LLMs) has driven significant advancements, yet it faces diminishing returns and escalating energy demands. This work explores how test-time compute (TTC) can serve as an energy-efficient complement to conventional scaling strategies by allocating additional computational resources at inference time rather than during training.
Antonio Joia Neto, Norrathep Rattanavipanon, Ivan De Oliveira Nunes
Embedded devices are increasingly ubiquitous and vital, often supporting safety-critical functions. However, due to strict cost and energy constraints, they are typically implemented with Micro-Controller Units (MCUs) that lack advanced architectural security features. Within this space, recent efforts have created low-cost architectures capable of generatin
Hannah Potgieter, Razvan C. Fetecau, Steven J. Ruuth
We use the $p$-Laplacian with large $p$-values in order to approximate geodesic distances to features on surfaces. This differs from Fayolle and Belyaev's (2018) [1] computational results using the $p$-Laplacian for the distance-to-surface problem. Our approach appears to offer some distinct advantages over other popular PDE-based distance function approxima
Yixuan Gao, Xiongkuo Min, Guangtao Zhai
This paper explores how the image quality assessment (IQA) task affects the cognitive processes of people from the perspective of pupil size and studies the relationship between pupil size and image quality. Specifically, we first invited subjects to participate in a subjective experiment, which includes two tasks: free observation and IQA. In the free obser
Zhengqing Yuan, Weixiang Sun, Yixin Liu, Huichi Zhou
Large Language Models (LLMs) have driven significant progress, yet their growing parameter counts and context windows incur prohibitive compute, energy, and monetary costs. We introduce EfficientLLM, a novel benchmark and the first comprehensive empirical study evaluating efficiency techniques for LLMs at scale. Conducted on a production-class cluster (48xGH
Zhenyu Bao, Qing Li, Guibiao Liao, Zhongyuan Zhao
3D Gaussian Splatting (3DGS) has gained significant attention in streamable dynamic novel view synthesis (DNVS) for its photorealistic rendering capability and computational efficiency. Despite much progress in improving rendering quality and optimization strategies, 3DGS-based streamable dynamic scene reconstruction still suffers from flickering artifacts a
On the Input-Output Monotonicity of Voltage Dynamics of Power System with Grid-Forming Converters
eess.SYZhenyao Li, Shengwen Liao, Qian Zhang, Xuechun Zhang
Integration of renewable resources is profoundly reshaping the dynamics of modern power systems. This study shows that the voltage dynamics of power systems with multiple grid-forming (GFM) converters often enjoys a desirable property called input-output monotonicity. A systematic approach for computing the derivatives of the voltage subsystem is presented f
Weihao Zou, Weibing Feng, Pin Wu
This study proposes a universal flow field prediction framework based on knowledge transfer from large language model (LLM), addressing the high computational costs of traditional computational fluid dynamics (CFD) methods and the limited cross-condition transfer capability of existing deep learning models. The framework innovatively integrates Proper Orthog
Gokul Puthumanaillam, Paulo Padrao, Jose Fuentes, Leonardo Bobadilla
Robots navigating complex environments must manage uncertainty from sensor noise, environmental changes, and incomplete information, with different tasks requiring varying levels of precision in different areas. For example, precise localization may be crucial near obstacles but less critical in open spaces. We present GUIDE (Generalized Uncertainty Integrat
Jerry Tang, Ruiqi Zhang, Kaan Beyduz, Yiwei Jiang
This paper presents Duawlfin, a drone with unified actuation for wheeled locomotion and flight operation that achieves efficient, bidirectional ground mobility. Unlike existing hybrid designs, Duawlfin eliminates the need for additional actuators or propeller-driven ground propulsion by leveraging only its standard quadrotor motors and introducing a differen
Ionic environment-modulated nucleation and stability of multiscale nanodomains in surfactant-free microemulsions
physics.chem-phYawen Gao, Changsheng Chen, Mingbo Li, Chao Sun
Nano-clustering occurs in the monophasic "pre-Ouzo" region of ternary liquid mixtures without the use of surfactants. This study is proposed to elucidate the nucleation and stability of multiscale nanodomains in a surfactant-free microemulsion (SFME) system composed of trans-anethol, ethanol and water, tuned by aqueous ionic environment. We examined direct-
Zhi Su, Yuman Gao, Emily Lukas, Yunfei Li
Achieving coordinated teamwork among legged robots requires both fine-grained locomotion control and long-horizon strategic decision-making. Robot soccer offers a compelling testbed for this challenge, combining dynamic, competitive, and multi-agent interactions. In this work, we present a hierarchical multi-agent reinforcement learning (MARL) framework that
Effective climate policies for major emission reductions of ozone precursors: Global evidence from two decades
stat.APNingning Yao, Huan Xi, Lang Chen, Zhe Song
Despite policymakers deploying various tools to mitigate emissions of ozone (O\textsubscript{3}) precursors, such as nitrogen oxides (NO\textsubscript{x}), carbon monoxide (CO), and volatile organic compounds (VOCs), the effectiveness of policy combinations remains uncertain. We employ an integrated framework that couples structural break detection with mach
Faizuddin Ahmed, Ahmad Al-Badawi, Izzet Sakallı
In a recent article (Ref. \cite{AOP}), the authors obtained a static, cylindrically symmetric Anti-de Sitter (AdS) black string (BS) solutions, which are cylindrical generalizations of black holes (BHs), surrounded by a cloud of strings (CS) and the quintessence field (QF), and discussed its properties. In the present study, we present a comprehensive analys
Jacob E. Turner, Timothy Dolch, Paul B. Demorest, Ryan S. Lynch
We explore possible advantages of cyclic spectroscopy for observations of pulsars in instances where full cyclic deconvolution is not feasible. We compute cyclic merits and full-deconvolution regime boundaries for pulsars observed by NANOGrav and discuss which sources stand to benefit the most from using cyclic spectroscopy when observed with the Green Bank
Zongyuan Deng, Yujie Cai, Qing Liu, Shiyao Mu
The selection of base station sites is a critical challenge in 5G network planning, which requires efficient optimization of coverage, cost, user satisfaction, and practical constraints. Traditional manual methods, reliant on human expertise, suffer from inefficiencies and are limited to an unsatisfied planning-construction consistency. Existing AI tools, de
Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai
Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS,
Xu Ge, Roman Verba, Philipp Pirro, Andrii V. Chumak
Dipolar coupling between closely spaced magnetic waveguides enables the design of magnonic directional couplers - universal devices capable of functioning as signal combiners, power splitters, demultiplexers, and more. The wavelength-dependent coupling, combined with the weak nonlinear variation of a spin wave's wavelength at constant-frequency, introduces p
Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models
cs.AIKiarash Naghavi Khanghah, Zhiling Chen, Lela Romeo, Qian Yang
Additive manufacturing enables the fabrication of complex designs while minimizing waste, but faces challenges related to defects and process anomalies. This study presents a novel multimodal Retrieval-Augmented Generation-based framework that automates anomaly detection across various Additive Manufacturing processes leveraging retrieved information from li
A Sequence-Form Characterization and Differentiable Path-Following Method for Computing Normal-Form Perfect Equilibria in Extensive-Form Games
cs.GTYuqing Hou, Yiyin Cao, Chuangyin Dang, Yong Wang
The sequence form, owing to its compact and holistic strategy representation, has demonstrated significant efficiency in computing normal-form perfect equilibria for two-player extensive-form games with perfect recall. Nevertheless, the examination of $n$-player games remains underexplored. To tackle this challenge, we present a sequence-form characterizatio
Pushing the Frontiers of Self-Distillation Prototypes Network with Dimension Regularization and Score Normalization
eess.ASYafeng Chen, Chong Deng, Hui Wang, Yiheng Jiang
Developing robust speaker verification (SV) systems without speaker labels has been a longstanding challenge. Earlier research has highlighted a considerable performance gap between self-supervised and fully supervised approaches. In this paper, we enhance the non-contrastive self-supervised framework, Self-Distillation Prototypes Network (SDPN), by introduc
Anurag Mishra
Mechanistic interpretability research seeks to reveal the inner workings of large language models, yet most work focuses on classification or generative tasks rather than summarization. This paper presents an interpretability framework for analyzing how GPT-like models adapt to summarization tasks. We conduct differential analysis between pre-trained and fin
TransFit: An Efficient Framework for Transient Light-Curve Fitting with Time-Dependent Radiative Diffusion
astro-ph.HELiang-Duan Liu, Yu-Hao Zhang, Yun-Wei Yu, Ze-Xin Du
Modeling the light curves (LCs) of luminous astronomical transients, such as supernovae, is crucial for understanding their progenitor physics, particularly with the exponential growth of survey data. However, existing methods face limitations: efficient semi-analytical models (e.g., Arnett-like) employ significant physical simplifications (like time-invaria
David X. Lin, Daniel Hall, Giannis Fikioris, Siddhartha Banerjee
We study the problem of fair online resource allocation via non-monetary mechanisms, where multiple agents repeatedly share a resource without monetary transfers. Previous work has shown that every agent can guarantee $1/2$ of their ideal utility (the highest achievable utility given their fair share of resources) robustly, i.e., under arbitrary behavior by
Hiroyuki Hayashi
We consider ruled surfaces with finite multiplicity. We study behaviors of the striction curves and the singularities of the ruled surfaces. We also give geometric meanings of invariants related to the ruled surfaces.
Predicting Neoadjuvant Chemotherapy Response in Triple-Negative Breast Cancer Using Pre-Treatment Histopathologic Images
q-bio.QMHikmat Khan, Ziyu Su, Huina Zhang, Yihong Wang
Triple-negative breast cancer (TNBC) remains a major clinical challenge due to its aggressive behavior and lack of targeted therapies. Accurate early prediction of response to neoadjuvant chemotherapy (NACT) is essential for guiding personalized treatment strategies and improving patient outcomes. In this study, we present an attention-based multiple instanc
A Bayesian Sparse Kronecker Product Decomposition Framework for Tensor Predictors with Mixed-Type Responses
stat.MEShao-Hsuan Wang, Hsin-Hsiung Huang
Ultra-high-dimensional tensor predictors are increasingly common in neuroimaging and other biomedical studies, yet existing methods rarely integrate continuous, count, and binary responses in a single coherent model. We present a Bayesian Sparse Kronecker Product Decomposition (BSKPD) that represents each regression (or classification) coefficient tensor as
Ram Mohan Rao Kadiyala, Siddhant Gupta, Jebish Purbey, Srishti Yadav
Vision-Language Models (VLMs) have demonstrated impressive capabilities across a range of tasks, yet concerns about their potential biases exist. This work investigates the extent to which prominent VLMs exhibit cultural biases by evaluating their performance on an image-based country identification task at a country level. Utilizing the geographically diver
Lucas Rosenblatt, Bin Han, Robert Wolfe, Bill Howe
Large language models (LLMs) can leak sensitive training data through memorization and membership inference attacks. Prior work has primarily focused on strong adversarial assumptions, including attacker access to entire samples or long, ordered prefixes, leaving open the question of how vulnerable LLMs are when adversaries have only partial, unordered sampl