Skip to content

May 2025 arXiv papers — page 25

Showing 2,4012,500 of 24,552 papers

  1. Kazutoshi Yamazaki, Qingyuan Zhang

    We study a stochastic control problem where the underlying process follows a spectrally negative L\'{e}vy process. A controller can continuously increase the process but only decrease it at independent Poisson arrival times. We show the optimality of the double-barrier strategy, which increases the process whenever it would fall below some lower barrier and

  2. Abdul Rahman Shaikh, Maoyuan Sun, Xingchen Liu, Hamed Alhoori

    Exploring data relations across multiple views has been a common task in many domains such as bioinformatics, cybersecurity, and healthcare. To support this, various techniques (e.g., visual links and brushing and linking) are used to show related visual elements across views via lines and highlights. However, understanding the relations using these techniqu

  3. Yuu Jinnai

    Document-level text generation tasks are known to be more difficult than sentence-level text generation tasks as they require the understanding of longer context to generate high-quality texts. In this paper, we investigate the adaption of Minimum Bayes Risk (MBR) decoding for document-level text generation tasks. MBR decoding makes use of a utility function

  4. Zhennan Lin, Kaixun Huang, Wei Ren, Linju Yang

    Deep biasing improves automatic speech recognition (ASR) performance by incorporating contextual phrases. However, most existing methods enhance subwords in a contextual phrase as independent units, potentially compromising contextual phrase integrity, leading to accuracy reduction. In this paper, we propose an encoder-based phrase-level contextualized ASR m

  5. Rajkamal Sah, Sumit Sunil Tambe, Gopalan Jagadeesh

    This paper probes into the flow induced by a rotating cone-cylinder model in an enclosure. Two component particle image velocimetry measurements in the symmetry plane reveal that the rotating cone-cylinder causes an outward jet on the cylinder section, which lifts the rotating boundary layers away from the wall. A large-scale counter-rotating vortex pair set

  6. Amit Kumthekar, Zion Tilley, Henry Duong, Bhargav Patel

    Despite the growing clinical adoption of large language models (LLMs), current approaches heavily rely on single model architectures. To overcome risks of obsolescence and rigid dependence on single model systems, we present a novel framework, termed the Consensus Mechanism. Mimicking clinical triage and multidisciplinary clinical decision-making, the Consen

  7. Laszlo Gyongyosi, Sandor Imre

    The intense growth of quantum computation and communication allows the development of advanced solutions and services. Networked quantum services are provided for the users via quantum computers and quantum networking. Here, we review the fundamental concepts and recent achievements of networked quantum services. We present a comprehensive study of the state

  8. Alireza Khadem, Kamalavasan Kamalakkannan, Zhenyan Zhu, Akash Poptani

    Indirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck. Prior indirect memory access proposals, such as indirect prefetchers, runahead execution, fetchers, and decoupled access/execute architectures, primarily focus on improving memory access latency by loading data ahead of computation but still rely on th

  9. Takeshi Yoshimura, Tatsuhiro Chiba, Manish Sethi, Daniel Waddington

    The rapid increases in model parameter sizes introduces new challenges in pre-trained model loading. Currently, machine learning code often deserializes each parameter as a tensor object in host memory before copying it to device memory. We found that this approach underutilized storage throughput and significantly slowed down loading large models with a wid

  10. Peizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou

    Multi-Modal Learning (MML) integrates information from diverse modalities to improve predictive accuracy. While existing optimization strategies have made significant strides by mitigating gradient direction conflicts, we revisit MML from a gradient-based perspective to explore further improvements. Empirically, we observe an interesting phenomenon: performa

  11. Anjana Wijayawardhana, David Gunawan, Thomas Suesse

    Standard simultaneous autoregressive (SAR) models typically assume normally distributed errors, an assumption often violated in real-world datasets that frequently exhibit non-normal, skewed, or heavy-tailed characteristics. New SAR models are proposed to capture these non-Gaussian features. The spatial error model (SEM), a widely used SAR-type model, is con

  12. Arabinda Bera, Ido Regev, Alessio Zaccone, Matteo Baggioli

    Eshelby-like quadrupolar structures serve as the fundamental microscopic units for characterizing plastic instabilities in amorphous solids and play a crucial role in explaining their mechanical failure, including the formation of shear bands. However, identifying Eshelby-like plastic events in glasses remains challenging due to their inherent structural and

  13. Rui Xu, Yuzhen Niu, Yuezhou Li, Huangbiao Xu

    Existing low-light image enhancement (LLIE) and joint LLIE and deblurring (LLIE-deblur) models have made strides in addressing predefined degradations, yet they are often constrained by dynamically coupled degradations. To address these challenges, we introduce a Unified Receptance Weighted Key Value (URWKV) model with multi-state perspective, enabling flexi

  14. Rongli Huang, Changzheng Qu, Zhizhang Wang, Weifeng Wo

    We investigate the evolution of strictly convex hypersurfaces driven by the $k$-Hessian curvature flow, subject to the second boundary condition. We first explore the translating solutions corresponding to this boundary value problem. Next, we establish the long-time existence of the flow and prove that it converges to a translating solution. To overcome the

  15. Shuyin Xia, Xiaojiang Tian, Suzhen Yuan, Jeremiah D. Deng

    High time complexity is one of the biggest challenges faced by $k$-Nearest Neighbors ($k$NN). Although current classical and quantum $k$NN algorithms have made some improvements, they still have a speed bottleneck when facing large amounts of data. To address this issue, we propose an innovative algorithm called Granular-Ball based Quantum $k$NN(GB-Q$k$NN).

  16. Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang

    With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial for enhancing user experience, content understanding, and platform intelligence. Existing benchmarks primarily focus on text-centric tasks, lacking coverage of the multimodal cont

  17. Bradley Lamb, Saroj Upreti, Yunfei Wang, Daniel Struble

    The morphology of block copolymers (BCPs) critically influences material properties and applications. This work introduces a machine learning (ML)-enabled, high-throughput framework for analyzing grazing incidence small-angle X-ray scattering (GISAXS) data and atomic force microscopy (AFM) images to characterize BCP thin film morphology. A convolutional neur

  18. Denis Mamba Kabala, Adel Hafiane, Laurent Bobelin, Raphael Canals

    Crop disease detection and classification is a critical challenge in agriculture, with major implications for productivity, food security, and environmental sustainability. While deep learning models such as CNN and ViT have shown excellent performance in classifying plant diseases from images, their large-scale deployment is often limited by data privacy co

  19. Lingkai Kong, Haichuan Wang, Tonghan Wang, Guojun Xiong

    Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamics in the offline dataset differ from those encountered online. Existing approaches typically mitigate this issue by penalizing or filtering offline transitions in regions with larg

  20. Tarun Suresh, Debangshu Banerjee, Shubham Ugare, Sasa Misailovic

    Diffusion LLMs have emerged as a promising alternative to conventional autoregressive LLMs, offering significant potential for improved runtime efficiency. However, existing diffusion models lack the ability to provably enforce user-specified formal constraints, such as regular expressions, which makes them unreliable for tasks that require structured output

  21. Jeonghun Cho, Deokhyung Kang, Hyounghun Kim, Gary Geunbae Lee

    Self-correction has demonstrated potential in code generation by allowing language models to revise and improve their outputs through successive refinement. Recent studies have explored prompting-based strategies that incorporate verification or feedback loops using proprietary models, as well as training-based methods that leverage their strong reasoning ca

  22. Dohyeon Lee, Yeonseok Jeong, Seung-won Hwang

    Chain-of-Thought (CoT) prompting enables complex reasoning in large language models (LLMs), including applications in information retrieval (IR). However, it often leads to overthinking, where models produce excessively long and semantically redundant traces with little or no benefit. We identify two key challenges in IR: redundant trajectories that revisit

  23. Yutong Xie, Zhuoheng Li, Xiyuan Wang, Yijun Pan

    Despite their success in numerous fields, the potential of foundation models for modeling and understanding human behavior remains largely unexplored. We introduce Be.FM, one of the first open foundation models designed for human behavior modeling. Built upon open-source large language models and fine-tuned on a diverse range of behavioral data, Be.FM can be

  24. Jun Kigami, Yuka Ota

    We provide a rich family of self-similar sets, called locally symmetric polygon-based self-similar sets, as examples of metric spaces having conductive homogeneity, which was introduced as a sufficient condition for the construction of counterparts of "Sobolev spaces" on compact metric spaces. In particular, our results imply the existence of "Brownian motio

  25. Zijian Liu, Zhengyuan Zhou

    We study the convergence of the shuffling gradient method, a popular algorithm employed to minimize the finite-sum function with regularization, in which functions are passed to apply (Proximal) Gradient Descent (GD) one by one whose order is determined by a permutation on the indices of functions. In contrast to its easy implementation and effective perform

  26. Zhen Xiang, Aliyah R. Hsu, Austin V. Zane, Aaron E. Kornblith

    Clinical decision-making is inherently complex and fast-paced, particularly in emergency departments (EDs) where critical, rapid and high-stakes decisions are made. Clinical Decision Rules (CDRs) are standardized evidence-based tools that combine signs, symptoms, and clinical variables into decision trees to make consistent and accurate diagnoses. CDR usage

  27. Yuxuan Lin, Ruihang Chu, Zhenyu Chen, Xiao Tang

    Generative 3D reconstruction shows strong potential in incomplete observations. While sparse-view and single-image reconstruction are well-researched, partial observation remains underexplored. In this context, dense views are accessible only from a specific angular range, with other perspectives remaining inaccessible. This task presents two main challenges

  28. Wei-Hsiang Huang, Chen-Wei Ke, Wei-Ning Chiu, Yu-Xuan Su

    Large language models (LLMs) have introduced new paradigms for recommender systems by enabling richer semantic understanding and incorporating implicit world knowledge. In this study, we propose a systematic taxonomy that classifies existing approaches into two categories: (1) Pure LLM Recommenders, which rely solely on LLMs, and (2) Augmented LLM Recommende

  29. Jiarui Zhang, Xiangyu Liu, Yong Hu, Chaoyue Niu

    Retrieval-Augmented Generation (RAG) significantly improves the performance of Large Language Models (LLMs) on knowledge-intensive tasks. However, varying response quality across LLMs under RAG necessitates intelligent routing mechanisms, which select the most suitable model for each query from multiple retrieval-augmented LLMs via a dedicated router model.

  30. Ramiz Aktar, Kuo-Chuan Pan, Toru Okuda

    In this study, we investigate the effect of resistivity on the dynamics of global magnetohydrodynamic accretion flows (Res-MHD) around a spinning supermassive black hole. We perform a comparative study of 2D and 3D resistive models around black holes. We examine accretion flow dynamics considering globally uniform resistivity values, ranging from $\sim 0$ to

  31. Dexing Miao, Zijun Xu, Zhiyu Xiang, Pingcheng Liu

    A silicon microstrip detector (SSD) has been developed to have state of the art spatial resolution and a large sensitive area under stringent power constraints. The design incorporates three floating strips with their bias resistors inserted between two aluminum readout strips. Beam test measurements with the single sensor confirmed that this configuration a

  32. Tianteng Gu, Bei Liu, Bo Xiao, Ke Zeng

    Pruning is a widely used technique to compress large language models (LLMs) by removing unimportant weights, but it often suffers from significant performance degradation - especially under semi-structured sparsity constraints. Existing pruning methods primarily focus on estimating the importance of individual weights, which limits their ability to preserve

  33. Tianci Bu, Le Zhou, Wenchuan Yang, Jianhong Mou

    Trajectory data is crucial for various applications but often suffers from incompleteness due to device limitations and diverse collection scenarios. Existing imputation methods rely on sparse trajectory or travel information, such as velocity, to infer missing points. However, these approaches assume that sparse trajectories retain essential behavioral patt

  34. John L. Ball, Shon Mackie, Jacob G. van de Lindt, Willow Morrissey

    The Centrifugal Mirror Fusion Experiment (CMFX) at the University of Maryland, College Park is a rotating mirror device that utilizes a central cathode to generate a radial electric field which induces a strongly sheared azimuthal $E\times B$ flow to improve plasma confinement and stability. The fusion yield of CMFX plasmas is assessed by diagnosis of neutro

  35. Runshi Tang, Julien Chhor, Olga Klopp, Anru R. Zhang

    Canonical Polyadic (CP) tensor decomposition is a fundamental technique for analyzing high-dimensional tensor data. While the Alternating Least Squares (ALS) algorithm is widely used for computing CP decomposition due to its simplicity and empirical success, its theoretical foundation, particularly regarding statistical optimality and convergence behavior, r

  36. Chuanhao Li, Wenbo Ye, Zhen Li, Yuwei Wu

    Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V\&L) recently. Due to the multi-modal nature of V\&L tasks, the primitives composing compositions source from different modalities, resulting in multi-sourced novel compositions. However, the generaliza

  37. Yu Sheng, Jiajun Deng, Xinran Zhang, Yu Zhang

    A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a co

  38. Jihai Zhang, Tianle Li, Linjie Li, Zhengyuan Yang

    Recent advancements in unified vision-language models (VLMs), which integrate both visual understanding and generation capabilities, have attracted significant attention. The underlying hypothesis is that a unified architecture with mixed training on both understanding and generation tasks can enable mutual enhancement between understanding and generation. H

  39. Siwen Wang, Shitou Zhang, Wan-Lin Chen, Dung Truong

    Recent advancements in Large Language Models have inspired the development of foundation models across various domains. In this study, we evaluate the efficacy of Large EEG Models (LEMs) by fine-tuning LaBraM, a state-of-the-art foundation EEG model, on a real-world stress classification dataset collected in a graduate classroom. Unlike previous studies that

  40. Zhiqiang Huang, Qing-yu Cai

    This study establishes a universal mechanism for entropy production in isolated quantum systems governed by interactions that induce random-phase fluctuations. By developing a resolvent-based framework, we demonstrate that steady-state entropy generically arises from many-body interactions, independent of specific coupling details, provided the coherent accu

  41. Yihang Wu, Muhammad Owais, Reem Kateb, Ahmad Chaddad

    Deep models, such as convolutional neural networks (CNNs) and vision transformer (ViT), demonstrate remarkable performance in image classification. However, those deep models require large data to fine-tune, which is impractical in the medical domain due to the data privacy issue. Furthermore, despite the feasible performance of contrastive language image pr

  42. Kapil Vaidya, Jialin Ding, Sebastian Kosak, David Kernert

    NL2SQL (natural language to SQL) translates natural language questions into SQL queries, thereby making structured data accessible to non-technical users, serving as the foundation for intelligent data applications. State-of-the-art NL2SQL techniques typically perform translation by retrieving database-specific information, such as the database schema, and i

  43. Yuzhen Xiao, Jiahe Song, Yongxin Xu, Ruizhe Zhang

    In-Context Learning (ICL) technique based on Large Language Models (LLMs) has gained prominence in Named Entity Recognition (NER) tasks for its lower computing resource consumption, less manual labeling overhead, and stronger generalizability. Nevertheless, most ICL-based NER methods depend on large-parameter LLMs: the open-source models demand substantial c

  44. Longyin Zhang, Bowei Zou, Ai Ti Aw

    The inherent nature of social media posts, characterized by the freedom of language use with a disjointed array of diverse opinions and topics, poses significant challenges to downstream NLP tasks such as comment clustering, comment summarization, and social media opinion analysis. To address this, we propose a granular level of identifying and generating as

  45. Yuhang Dai, He Wang, Xingchen Li, Zihan Zhang

    This paper delineates AISHELL-5, the first open-source in-car multi-channel multi-speaker Mandarin automatic speech recognition (ASR) dataset. AISHLL-5 includes two parts: (1) over 100 hours of multi-channel speech data recorded in an electric vehicle across more than 60 real driving scenarios. This audio data consists of four far-field speech signals captur

  46. Hyunwoo Kim, Hanau Yi

    Machine-Facing English (MFE) is an emergent register shaped by the adaptation of everyday language to the expanding presence of AI interlocutors. Drawing on register theory (Halliday 1985, 2006), enregisterment (Agha 2003), audience design (Bell 1984), and interactional pragmatics (Giles & Ogay 2007), this study traces how sustained human-AI interaction norm

  47. Guangyi Liu, Yongqi Zhang, Xunyuan Liu, Quanming Yao

    Drug-drug interaction (DDI) prediction is critical for treatment safety. While large language models (LLMs) show promise in pharmaceutical tasks, their effectiveness in DDI prediction remains challenging. Inspired by the well-established clinical practice where physicians routinely reference similar historical cases to guide their decisions through case-base

  48. Sanchita Ghosh, Tanushree Roy

    Modern traffic management systems increasingly adopt hierarchical control strategies for improved efficiency and scalability, where a local traffic controller mode is chosen by a supervisory controller based on the changing large-scale driving patterns. Unfortunately, such local metering controllers are also vulnerable to cyberattacks that can disrupt the co

  49. Dongwoo Lee, Dong Bok Lee, Steven Adriaensen, Juho Lee

    Scaling has been a major driver of recent advancements in deep learning. Numerous empirical studies have found that scaling laws often follow the power-law and proposed several variants of power-law functions to predict the scaling behavior at larger scales. However, existing methods mostly rely on point estimation and do not quantify uncertainty, which is c

  50. Jinyi Chang, Dongliang Chang, Lei Chen, Bingyao Yu

    In recent years, Fine-Grained Visual Classification (FGVC) has achieved impressive recognition accuracy, despite minimal inter-class variations. However, existing methods heavily rely on instance-level labels, making them impractical in privacy-sensitive scenarios such as medical image analysis. This paper aims to enable accurate fine-grained recognition wit

  51. Shruti Hegde, Mabon Manoj Ninan, Jonathan R. Dillman, Shireen Hayatghaibi

    General-purpose clinical natural language processing (NLP) tools are increasingly used for the automatic labeling of clinical reports. However, independent evaluations for specific tasks, such as pediatric chest radiograph (CXR) report labeling, are limited. This study compares four commercial clinical NLP systems - Amazon Comprehend Medical (AWS), Google He

  52. Si Wu, Sebastian Bruch

    Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational methods that estimate them do so using parallel visual and semantic spaces, such as collections of image-caption pairs or multi-modal models. In t

  53. Minh Nguyen Nhat To, Paul F RWilson, Viet Nguyen, Mohamed Harmanani

    The subpopulationtion shift, characterized by a disparity in subpopulation distributibetween theween the training and target datasets, can significantly degrade the performance of machine learning models. Current solutions to subpopulation shift involve modifying empirical risk minimization with re-weighting strategies to improve generalization. This strateg

  54. Haewon Park, Gyubin Choi, Minjun Kim, Yohan Jo

    Knowledge editing (KE) methods offer an efficient way to modify knowledge in large language models. Current KE evaluations typically assess editing success by considering only the edited knowledge without any preceding contexts. In real-world applications, however, preceding contexts often trigger the retrieval of the original knowledge and undermine the int

  55. Geyue Sun, Xiao Liu, Tomas Williams, Roberto Samaniego

    We construct a novel event-level Capital Control Measures (CCM) dataset covering 196 countries from 1999 to 2023 by leveraging prompt-based large language models (LLMs). The dataset enables event study analysis and cross-country comparisons based on rich policy attributes, including action type, intensity, direction, implementing entity, and other multidimen

  56. Zhihao Wang, Wenke Huang, Tian Chen, Zekun Shi

    The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. This paper systematically investigates the behavioral differen

  57. Samaneh Goodarzi, Ali Tavanmand, Hamed Shajari

    This study evaluated the impact of Zinc (Zn) supplementation on the growth, yield, and seed quality of chickpea (Pisum sativum L.) under the semi-arid conditions of Kerman, Iran, across two growing seasons (2021-2022). A randomized complete block design was used with six treatments, including varying concentrations of zinc sulphate applied via foliar sprayin

  58. Jack Kendrick

    Kernel density estimation is a popular method for estimating unseen probability distributions. However, the convergence of these classical estimators to the true density slows down in high dimensions. Moreover, they do not define meaningful probability distributions when the intrinsic dimension of data is much smaller than its ambient dimension. We build on

  59. Yinghao Tang, Tingfeng Lan, Bo Pan, Xiuqi Huang

    Large Language Model (LLM) serving increasingly underpins online Web services such as conversational agents, Web search, and programming assistants, where requests carry heterogeneous Service Level Objectives (SLOs) such as Time to First Token (TTFT) and Time Per Output Token (TPOT). Existing LLM serving systems prioritize maximum throughput and treat all re

  60. Mrutyunjaya Sahoo, Rahul Ghosh, Bandita Das, Shishira Mahunta

    In this work we use optimal control to generate Discrete Time Crystals (DTC) in generic many-body quantum systems. We define appropriate cost functions, which, when optimized, result in the formation of DTCs. This hitherto unexplored method represents DTCs as an optimization problem, and allows us to find non-trivial realistic periodic control pulses and par

  61. Jinchuan Zhang, Lu Yin, Yan Zhou, Songlin Hu

    The acquisition of agentic capabilities has transformed LLMs from "knowledge providers" to "action executors", a trend that while expanding LLMs' capability boundaries, significantly increases their susceptibility to malicious use. Previous work has shown that current LLM-based agents execute numerous malicious tasks even without being attacked, indicating a

  62. Zeying Gong, Rong Li, Tianshuai Hu, Ronghe Qiu

    Deployable service and delivery robots struggle to navigate multi-floor buildings to reach object goals, as existing systems fail due to single-floor assumptions and requirements for offline, globally consistent maps. Multi-floor environments pose unique challenges including cross-floor transitions and vertical spatial reasoning, especially navigating unknow

  63. Haoqin Sun, Xuechen Wang, Jinghua Zhao, Shiwan Zhao

    In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have proliferated, there remains a pressing need for high-quality, comprehensive datasets tailored to the unique linguistic, cultural,

  64. Xingjian Wu, Xiangfei Qiu, Hongfan Gao, Jilin Hu

    Probabilistic Time Series Forecasting (PTSF) plays a crucial role in decision-making across various fields, including economics, energy, and transportation. Most existing methods excell at short-term forecasting, while overlooking the hurdles of Long-term Probabilistic Time Series Forecasting (LPTSF). As the forecast horizon extends, the inherent nonlinear d

  65. Aniruddh Mishra, Arthur K. Barnes, Jose E. Tabarez, Adam Mate

    Geomagnetic disturbances are a threat to the reliability and security of our national critical energy infrastructures. These events specifically result in geomagnetically induced currents, which can cause damage to transformers due to magnetic saturation. In order to mitigate these effects, blocker devices must be placed in optimal locations. Finding this pl

  66. Jinwen Chen, Hainan Zhang, Fei Sun, Qinnan Zhang

    Stealthy data poisoning during fine-tuning can backdoor large language models (LLMs), threatening downstream safety. Existing detectors either use classifier-style probability signals--ill-suited to generation--or rely on rewriting, which can degrade quality and even introduce new triggers. We address the practical need to efficiently remove poisoned example

  67. Juwei Yue, Haikuo Li, Jiawei Sheng, Xiaodong Li

    Graph neural networks (GNNs) leverage message passing mechanisms to learn the topological features of graph data. Traditional GNNs learns node features in a spatial domain unrelated to the topology, which can hardly ensure topological features. In this paper, we formulates message passing as a system of hyperbolic partial differential equations (hyperbolic P

  68. Liangkai Hang, Junjie Yao, Zhiwei Bai, Tianyi Chen

    The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their generalizability. This work demonstrates that model complexity control, conveniently implementable by adjusting the initialization rate and weight decay coefficient, improves the

  69. Shanaka Ramesh Gunasekara, Wanqing Li, Philip Ogunbona, Jack Yang

    Traditional approaches in unsupervised or self supervised learning for skeleton-based action classification have concentrated predominantly on the dynamic aspects of skeletal sequences. Yet, the intricate interaction between the moving and static elements of the skeleton presents a rarely tapped discriminative potential for action classification. This paper

  70. Oscar C. O. Dahlsten

    The Page curve is a curve of average subsystem entropy as a function of subsystem size. The curve starts from 0, rises, peaks near the maximal possible entropy when the subsystem makes up half of the total system, and then falls back to 0. We here describe subsystem entropy, averaging over quantum states, and why the curve rises and falls in that manner. We

  71. Bowen Chen, Keyan Chen, Mohan Yang, Zhengxia Zou

    High-resolution (HR) remote sensing imagery plays a vital role in a wide range of applications, including urban planning and environmental monitoring. However, due to limitations in sensors and data transmission links, the images acquired in practice often suffer from resolution degradation. Remote Sensing Image Super-Resolution (RSISR) aims to reconstruct H

  72. Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li

    Text-to-Speech (TTS) benchmarks often fail to capture how well models handle nuanced and semantically complex text. Building on $\textit{EmergentTTS}$, we introduce $\textit{EmergentTTS-Eval}$, a comprehensive benchmark covering six challenging TTS scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g. URLs, fo

  73. Jonathan Li, Zoltan Csaki, Nidhi Hiremath, Etash Guha

    Modern VLMs have achieved near-saturation accuracy in English document visual question-answering (VQA). However, this task remains challenging in lower resource languages due to a dearth of suitable training and evaluation data. In this paper we present scalable methods for curating such datasets by focusing on Hungarian, approximately the 17th highest resou

  74. Qing-Hong Cao, Masanori Tanaka, Jun-Chen Wang, Ke-Pan Xie

    We examine the idea that our universe began as a baby universe and show that this is feasible in a gauged $U(1)_{B-L}$ extension of the Standard Model with the classically conformal principle. For the first time, we define a measure to describe the probability that we reside in a baby universe, and find that it can be close to 1 in a considerable portion of

  75. Chiwan Park, Wonjun Jang, Daeryong Kim, Aelim Ahn

    The advancement of Large Language Models (LLMs) has led to significant improvements in various service domains, including search, recommendation, and chatbot applications. However, applying state-of-the-art (SOTA) research to industrial settings presents challenges, as it requires maintaining flexible conversational abilities while also strictly complying wi

  76. H. Saxena, J. Sayers, A. Gavidia, J. B. Melin

    Galaxy cluster abundance measurements are a valuable tool for constraining cosmological parameters like the mass density ($\Omega_m$) and density fluctuation amplitude ($\sigma_8$). Wide area surveys detect clusters based on observables, such as the total integrated Sunyaev-Zel'dovich effect signal ($Y_{SZ}$) in the case of Planck. Quantifying the survey sel

  77. Xinzheng Wu, Junyi Chen, Peiyi Wang, Shunxiang Chen

    In the research and development (R&D) and verification and validation (V&V) phases of autonomous driving decision-making and planning systems, it is necessary to integrate human factors to achieve decision-making and evaluation that align with human cognition. However, most existing datasets primarily focus on vehicle motion states and trajectories, neglecti

  78. Kyle R. Chickering, Bangzheng Li, Muhao Chen

    Multimodal Large Language Models (MLLMs) encode images into visual tokens, aligning visual and textual signals within a shared latent space to facilitate crossmodal representation learning. The CLIP model is a widely adopted foundational vision language model whose vision encoder has played a critical role in the development of MLLMs such as LLaVA. However,

  79. Linh Le Pham Van, Minh Hoang Nguyen, Hung Le, Hung The Tran

    Robust reinforcement learning (RL) aims to learn policies that remain effective despite uncertainties in its environment, which frequently arise in real-world applications due to variations in environment dynamics. The robust RL methods learn a robust policy by maximizing value under the worst-case models within a predefined uncertainty set. Offline robust R

  80. Zi-Kui Liu

    First, Second and Combined Laws of Thermodynamics are revised in terms of entropy change, partial entropy, partial volume, and chemical potential.

  81. Qiao Zhu, Dmitrii Chaikovskii, Bangti Jin, Ye Zhang

    Physics-informed neural network (PINN) has shown great potential in solving partial differential equations. However, it faces challenges when dealing with problems involving steep gradients. The solutions to singularly perturbed time-dependent reaction-advection-diffusion equations exhibit internal moving transition layers with sharp gradients, and thus the

  82. Yize Cheng, Wenxiao Wang, Mazda Moayeri, Soheil Feizi

    Open benchmarks are essential for evaluating and advancing large language models, offering reproducibility and transparency. However, their accessibility makes them likely targets of test set contamination. In this work, we introduce DyePack, a framework that leverages backdoor attacks to identify models that used benchmark test sets during training, without

  83. Guohua Liu, Yan Peng

    Based on previous studies, universal bounds $4\pi M \leqslant T_{min} \leqslant 6\sqrt{3}\pi M$ were conjectured to be characteristic properties of black hole spacetimes, where $M$ represents the mass of black holes and $T_{min}$ is the minimum orbital periods around black holes. In this work, we explore the minimum orbital periods of objects around Hayward

  84. Yihua Xu, Süleyman Kerimov, Sebastian Perez-Salazar

    In numerous online selection problems, decision-makers (DMs) must allocate on the fly limited resources to customers with uncertain values. The DM faces the tension between allocating resources to currently observed values and saving them for potentially better, unobserved values in the future. Addressing this tension becomes more demanding if an uncertain d

  85. Jihwan Oh

    Bargaining, a critical aspect of real-world interactions, presents challenges for large language models (LLMs) due to limitations in strategic depth and adaptation to complex human factors. Existing benchmarks often fail to capture this real-world complexity. To address this and enhance LLM capabilities in realistic bargaining, we introduce a comprehensive f

  86. Agnideep Aich, Ashit Baran Aich

    We present the Deep Copula Classifier (DCC), a class-conditional generative model that separates marginal estimation from dependence modeling using neural copula densities. DCC is interpretable, Bayes-consistent, and achieves excess-risk $O(n^{-r/(2r+d)})$ for $r$-smooth copulas. In a controlled two-class study with strong dependence ($|\rho|=0.995$), DCC le

  87. Pai Zhu, Quan Wang, Dhruuv Agarwal, Kurt Partridge

    Custom keyword spotting (KWS) allows detecting user-defined spoken keywords from streaming audio. This is achieved by comparing the embeddings from voice enrollments and input audio. State-of-the-art custom KWS models are typically trained contrastively using utterances whose keywords are randomly sampled from training dataset. These KWS models often struggl

  88. Linwei Tao, Yi-Fan Yeh, Minjing Dong, Tao Huang

    As large language models (LLMs) are increasingly deployed in high-stakes applications, robust uncertainty estimation is essential for ensuring the safe and trustworthy deployment of LLMs. We present the most comprehensive study to date of uncertainty estimation in LLMs, evaluating 80 models spanning open- and closed-source families, dense and Mixture-of-Expe

  89. Ari S. Benjamin, Kyle Daruwalla, Christian Pehle, Abdul-Malik Zekri

    One frequently wishes to learn a range of similar tasks as efficiently as possible, re-using knowledge across tasks. In artificial neural networks, this is typically accomplished by conditioning a network upon task context by injecting context as input. Brains have a different strategy: the parameters themselves are modulated as a function of various neuromo

  90. Hoang Pham, Thanh-Do Nguyen, Khac-Hoai Nam Bui

    Claim verification is a long-standing and challenging task that demands not only high accuracy but also explainability of the verification process. This task becomes an emerging research issue in the era of large language models (LLMs) since real-world claims are often complex, featuring intricate semantic structures or obfuscated entities. Traditional appro

  91. Anirban Dutta, Andrew Fullard, Wolfgang Kerzendorf, J. T. O'Brien

    Type Ia supernovae (SNe Ia) are powered by the radioactive decay of isotopes such as $^{56}$Ni and $^{56}$Co, making their $\gamma$-ray spectra useful probes of the explosion mechanism and ejecta structure. Accurate interpretation of $\gamma$-ray observables, including line ratios and continuum fluxes, requires a detailed understanding of the microphysical p

  92. Behzad Kamgar-Parsi, Behrooz Kamgar-Parsi

    Finding the number of meaningful clusters in an unlabeled dataset is important in many applications. Regularized k-means algorithm is a possible approach frequently used to find the correct number of distinct clusters in datasets. The most common formulation of the regularization function is the additive linear term $\lambda k$, where $k$ is the number of cl

  93. Pin-Han Chen, Yu-Sheng Lin, Wei-Cheng Lee, Tin-Yu Leu

    RF/Analog design is essential for bridging digital technologies with real-world signals, ensuring the functionality and reliability of a wide range of electronic systems. However, analog design procedures are often intricate, time-consuming and reliant on expert intuition, and hinder the time and cost efficiency of circuit development. To overcome the limita

  94. Francesco Di Marcantonio, Sunny Pradhan, Sofia Vallecorsa, Mari Carmen Bañuls

    We investigate the roughening transition in the pure $\mathbb{Z}_2$ lattice gauge theory in (2+1) dimensions. Using numerical simulations with matrix product states, we explore the static and dynamical properties of an electric flux string between two static charges as the coupling is varied and approaches the deconfinement phase transition from the confined

  95. Brian Seong, Paul Gebheim

    Modern blockchain applications are often constrained by a trade-off between user experience and trust. Chainless Apps present a new paradigm of application architecture that separates execution, trust, bridging, and settlement into distinct compostable layers. This enables app-specific sequencing, verifiable off-chain computation, chain-agnostic asset and me

  96. Nick Byrd

    By late 20th century, the rationality wars had launched debates about the nature and norms of intuitive and reflective thinking. Those debates drew from mid-20th century ideas such as bounded rationality, which challenged more idealized notions of rationality observed since the 19th century. Now that 21st century cognitive scientists are applying the resulti

  97. Seungjun Ahn, Eun Jeong Oh

    Network theory has proven invaluable in unraveling complex protein interactions. Previous studies have employed statistical methods rooted in network theory, including the Gaussian graphical model, to infer networks among proteins, identifying hub proteins based on key structural properties of networks such as degree centrality. However, there has been limit

  98. Masaharu Kagiyama, Tsuyoshi Okita

    This paper aims to develop an energy-efficient classifier for time-series data by introducing PatchEchoClassifier, a novel model that leverages a reservoir-based mechanism known as the Echo State Network (ESN). The model is designed for human activity recognition (HAR) using one-dimensional sensor signals and incorporates a tokenizer to extract patch-level r

  99. Guancheng Zhou, Haiping Xu, Hongkang Xu, Chenyu Li

    The popular K-means clustering algorithm potentially suffers from a major weakness for further analysis or interpretation. Some cluster may have disproportionately more (or fewer) points from one of the subpopulations in terms of some sensitive variable, e.g., gender or race. Such a fairness issue may cause bias and unexpected social consequences. This work

  100. Koki Matsuishi, Kosuke Ukita, Tsuyoshi Okita

    In recent years, the widespread adoption of wearable devices has highlighted the growing importance of behavior analysis using IMU. While applications span diverse fields such as healthcare and robotics, recent studies have increasingly focused on multimodal analysis, in addition to unimodal analysis. Several studies have proposed multimodal foundation model