Skip to content

March 2025 arXiv papers — page 181

Showing 18,00118,100 of 23,633 papers

  1. Mingxing Li, Rui Wang, Lei Sun, Yancheng Bai

    The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models (MLLMs) have shown great potential in image quality assessment (IQA) and image aesthetic assessment (IAA). Despite this

  2. Bohan Liu, Xiaosen Wang

    Transfer-based attacks pose a significant threat to real-world applications by directly targeting victim models with adversarial examples generated on surrogate models. While numerous approaches have been proposed to enhance adversarial transferability, existing works often overlook the intrinsic relationship between adversarial perturbations and input image

  3. Mingyang Song, Mao Zheng, Xuan Luo

    Pairwise LLM-as-a-judge evaluation asks the judge to identify the \emph{better} of two candidate answers. We study a one-line modification that asks for the \emph{worse} answer instead and recovers the preference by elimination, a procedure we call Goal-Reversed Prompting (GRP). GRP introduces no extra inference rounds, composes with any prompt template (dir

  4. Tadahiro Taniguchi, Yasushi Hirai, Masahiro Suzuki, Shingo Murata

    This paper introduces the System 0/1/2/3 framework as an extension of dual-process theory, employing a quad-process model of cognition. Expanding upon System 1 (fast, intuitive thinking) and System 2 (slow, deliberative thinking), we incorporate System 0, which represents pre-cognitive embodied processes, and System 3, which encompasses collective intelligen

  5. Jie He, Wanqiu Long, Deyi Xiong

    Large pre-trained neural models have achieved remarkable success in natural language process (NLP), inspiring a growing body of research analyzing their ability from different aspects. In this paper, we propose a test suite to evaluate the cohesive ability of pre-trained language models. The test suite contains multiple cohesion phenomena between adjacent an

  6. Khang Nguyen, An T. Le, Tien Pham, Manfred Huber

    Prior flow matching methods in robotics have primarily learned velocity fields to morph one distribution of trajectories into another. In this work, we extend flow matching to capture second-order trajectory dynamics, incorporating acceleration effects either explicitly in the model or implicitly through the learning objective. Unlike diffusion models, which

  7. Jian Ma, Qirong Peng, Xu Guo, Chen Chen

    Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in understanding and integrating multiple modalities. However, currently there is no straightforward and efficient framework to transfer the multimodal comprehension abilities of MLL

  8. Biplab Basak, Sourav Sarkar

    This article focuses on a class of properly edge-colored graphs, which arise from topological combinatorics, and investigates their embeddings onto surfaces. Specifically, these graphs are known as the dual graphs of balanced normal pseudomanifolds. We introduce the concept of the balanced genus, which represents the smallest genus of a surface onto which th

  9. Xiangxiang Chu, Renda Li, Yong Wang

    Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion models. However, transferring pretrained weights from vision mo

  10. Koga Okubo, Kanta Tachibana

    This study focuses on the rotation of the hips and shoulders during a baseball bat swing, analyzing the time-series changes in rotational angles, rotational velocities, and axes using marker position data obtained from a motion capture system with 12 infrared cameras. Previous studies have examined factors such as ground reaction forces, muscle activation pa

  11. Zhu-Yu Ren, Sheng-Quan Wang, Jian-Ming Shen, Xing-Gang Wu

    Theoretical calculations for event shape observables are often determined by using the conventional scale setting; i.e. the procedure defined by setting the renormalization scale to the center-of-mass energy $\mu_r=\sqrt{s}$ and evaluating theoretical uncertainties by varying the same scale $\mu_r$ in an arbitrary range. Both the event shape distributions an

  12. Jiebin Yan, Kangcheng Wu, Junjie Chen, Ziwen Tan

    Most of existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are either computationally expensive or less scalable. To solve these issues, in this paper, we present a flexible and effective

  13. Xiyang Wang, Hongyu Zhang, Shiming Zou, Zibing Bai

    Accurate momentum determination of a neutral hadron, such as a KL meson or a neutron, remains a significant challenge in particle physics and nuclear physics experiments. The Belle II experiment presents an opportunity to address this challenge through an upgrade incorporating Time-of-Flight (TOF) capability for its large KL and Muon Detector (KLM). We inves

  14. Yunrui Zheng

    We consider the evolution of contact lines for thermal convection of viscous fluids in a 2D open-top vessel. The domain is bounded above by a free moving boundary and otherwise by the solid wall of a vessel. The dynamics of the fluid are governed by the incompressible Boussinesq approximation under the influence of gravity, and the interface between fluid an

  15. Yuanlong Ruan

    We perform a complete analysis of the limiting behaviour of a class of quasilinear problems with Dirichlet boundary data g. We show that the Lipschitz constant of g plays a role in controlling the Gamma-convergence of the natural energies. However the solutions converge uniformly to solution of a limiting equation irrelevant to the Lipschitz constant of g. T

  16. Kai Yang, Zijian Bai, Yang Xiao, Xinyu Li

    3D reconstruction garners increasing attention alongside the advancement of high-level image applications, where dense stereo matching (DSM) serves as a pivotal technique. Previous studies often rely on publicly available datasets for training, focusing on modifying network architectures or incorporating specialized modules to extract domain-invariant featur

  17. Ashutosh Kumar, Sourabh Lahiri, Trilochan Bagarti, Subhashish Banerjee

    Quantum heat engines have undergone extensive studies over the last two decades. Simultaneously, the studies of the applications of stochastic resetting in various fields are on the rise. We explore the effect of stochastic resetting on the dynamics of a two-level and a three-level quantum heat engine. The extracted work is shown to increase with the resetti

  18. Liya Tang, Qirui Qu, Yiying Long, Xia Wu

    Objectives This study aimed to elucidate the potential mechanisms of electroacupuncture (EA) in restoring detrusor-bladder neck dyssynergesia (DBND) following suprasacral spinal cord injury. Methods A total of 52 adult female Sprague-Dawley rats were randomly assigned to either a sham group (n=12) or a spinal cord injury model group (n=40). In the model grou

  19. Yingna Wang, Qingqin Liu, Xiaoying Wei, Mingming Fan

    The essence of intangible cultural heritage (ICH) lies in the living knowledge and skills passed down through generations. Daily practice plays a vital role in revitalizing ICH by fostering continuous learning and improvement. However, limited resources and accessibility pose significant challenges to sustaining such practice. Virtual reality (VR) has shown

  20. Li weile, Liu Xiao

    Time series models face significant challenges in scaling to handle large and complex datasets, akin to the scaling achieved by large language models (LLMs). The unique characteristics of time series data and the computational demands of model scaling necessitate innovative approaches. While researchers have explored various architectures such as Transformer

  21. Shi-Yuan Wang, Jun-Qing Xia

    Constrained measurements of fundamental physical constants using astronomical observational data represent a powerful method for investigating potential new physics. In particular, the dispersion measure (DM) of fast radio bursts (FRBs), which probes the electron density along their propagation paths, may be influenced by the space-time variation of the fine

  22. Shinichi Tanaka, Zhao Wang, Yoichi Kato, Jun Ohya

    In this paper, we propose a unified framework that leverages a single pretrained LLM for Motion-related Multimodal Generation, referred to as MoMug. MoMug integrates diffusion-based continuous motion generation with the model's inherent autoregressive discrete text prediction capabilities by fine-tuning a pretrained LLM. This enables seamless switching betwe

  23. Xuanyu Zhang, Jiarui Meng, Zhipei Xu, Shuzhou Yang

    3D Gaussian Splatting (3DGS) has emerged as a premier method for 3D representation due to its real-time rendering and high-quality outputs, underscoring the critical need to protect the privacy of 3D assets. Traditional NeRF steganography methods fail to address the explicit nature of 3DGS since its point cloud files are publicly accessible. Existing GS steg

  24. Purbid Bambroo, Subinay Adhikary, Paheli Bhattacharya, Abhijnan Chakraborty

    Identification of rhetorical roles like facts, arguments, and final judgments is central to understanding a legal case document and can lend power to other downstream tasks like legal case summarization and judgment prediction. However, there are several challenges to this task. Legal documents are often unstructured and contain a specialized vocabulary, mak

  25. Hongjia Zhai, Boming Zhao, Hai Li, Xiaokun Pan

    Recently, neural radiance fields (NeRF) have gained significant attention in the field of visual localization. However, existing NeRF-based approaches either lack geometric constraints or require extensive storage for feature matching, limiting their practical applications. To address these challenges, we propose an efficient and novel visual localization ap

  26. Bing Chen, Xiang Liu

    Hybrid mesons occupy a unique position in the hadronic spectrum, offering a valuable insight into the non-perturbative nature of the glue field. Hybrid states with exotic $J^{PC}$ quantum numbers, such as the $1^{-+}$, $0^{+-}$ and $2^{+-}$, cannot mix with conventional $q\bar{q}$ mesons, making them critical for establishing the full spectrum of hybrid meso

  27. Ibrahim Al Azhar, Venkata Devesh Reddy, Hamed Alhoori, Akhil Pandey Akella

    The limitations sections of scientific articles play a crucial role in highlighting the boundaries and shortcomings of research, thereby guiding future studies and improving research methods. Analyzing these limitations benefits researchers, reviewers, funding agencies, and the broader academic community. We introduce LimTopic, a strategy where Topic generat

  28. Qi Zhang, Xiuyuan Chen, Ziyi He, Lianming Wu

    Cervical spondylosis, a complex and prevalent condition, demands precise and efficient diagnostic techniques for accurate assessment. While MRI offers detailed visualization of cervical spine anatomy, manual interpretation remains labor-intensive and prone to error. To address this, we developed an innovative AI-assisted Expert-based Diagnosis System that au

  29. Gozde Oney, Federico Monaco, Saptarshee Mitra, Asma Medjahed

    Aging limits lithium-ion battery lifetime and must be understood to improve durability and performance, requiring a detailed understanding of how aging alters the availability of cyclable lithium and the integrity of active particles. In this work, (de)lithiation mechanisms are examined and spatially-resolved at the microscale in aged graphite electrodes dis

  30. Hoang-Thang Ta, Anh Tran

    Kolmogorov-Arnold Networks (KANs) have inspired numerous works exploring their applications across a wide range of scientific problems, with the potential to replace Multilayer Perceptrons (MLPs). While many KANs are designed using basis and polynomial functions, such as B-splines, ReLU-KAN utilizes a combination of ReLU functions to mimic the structure of B

  31. Nikola Sandrić

    In this note, we discuss the uniform ergodicity of a diffusion process given by an It\^o stochastic differential equation. We present an integral condition in terms of the drift and diffusion coefficients that ensures the uniform ergodicity of the corresponding transition kernel with respect to the total variation distance. Applications of the obtained resul

  32. Janet Rafner, Ryan Q. Guloy, Eden W. Wen, Catherine M. Chiodo

    Generative AI (GenAI) chatbots are becoming increasingly integrated into virtual assistant technologies, yet their success hinges on the ability to gather meaningful user feedback to improve interaction quality, system outcomes, and overall user acceptance. Successful chatbot interactions can enable organizations to build long-term relationships with their c

  33. Keshu Wu, Xinyue Ye, Suphanut Jamonnak, Xin Feng

    Daily operations in large campuses depend on how efficiently people \emph{move} through space and time. In this sense, course timetables are more than administrative schedules: they act as mobility policies that orchestrate thousands of trajectories, shaping travel burden, congestion, accessibility, and the reliability of back-to-back transitions. Designing

  34. Weixuan Kong, Jinpeng Yu, Zijun Li, Hanwei Liu

    Automatic personality recognition is a research hotspot in the intersection of computer science and psychology, and in human-computer interaction, personalised has a wide range of applications services and other scenarios. In this paper, an end-to-end multimodal performance personality is established for both visual and auditory modal datarecognition network

  35. Akshat Jain

    This paper presents a novel approach to image dehazing by combining Feature Fusion Attention (FFA) networks with CycleGAN architecture. Our method leverages both supervised and unsupervised learning techniques to effectively remove haze from images while preserving crucial image details. The proposed hybrid architecture demonstrates significant improvements

  36. Kuanghong Liu, Jin Wang, Kangjian He, Dan Xu

    Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptati

  37. Xiyuan Wang, Ziang Li, Sizhe Chen, Xingxing Xing

    In-game friend recommendations significantly impact player retention and sustained engagement in online games. Balancing similarity and diversity in recommendations is crucial for fostering stronger social bonds across diverse player groups. However, automated recommendation systems struggle to achieve this balance, especially as player preferences evolve ov

  38. Syed Sajid Ullah, Li Gang, Mudassir Riaz, Ahsan Ashfaq

    Handwritten digit recognition remains a fundamental challenge in computer vision, with applications ranging from postal code reading to document digitization. This paper presents an ensemble-based approach that combines Convolutional Neural Networks (CNNs) with traditional machine learning techniques to improve recognition accuracy and robustness. We evaluat

  39. Avadhut V. Purohit, Udaysinh T. Bhosale

    We study the double kicked top (DKT), which is an extension of the standard quantum kicked top (QKT) model. The model allows us to study the transition from time-reversal symmetric to broken time-reversal symmetric dynamics. Our transformation in the kick strength parameter space $(k, k') \to (k_r, k_\theta)$ reveals interesting features. The transformed kic

  40. Naoyuki Monden, Reo Yabuguchi

    We show that for any positive integer $h$, a knot surgered elliptic surface $E(n)_{T(2,2h+1)}$ for a $(2,2h+1)$-torus knot $T(2,2h+1)$ and the elliptic surface $E(1)_{2,2h+1}$ admit handle decompositions without 1- and 3-handles using the Kirby diagrams derived from Lefschetz fibrations on them.

  41. Mingqi Yuan, Bo Li, Xin Jin, Wenjun Zeng

    Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity and computational cost. To tackle this problem, existing appr

  42. Xianjie Liu, Keren Fu, Qijun Zhao

    High-precision dichotomous image segmentation (DIS) is a task of extracting fine-grained objects from high-resolution images. Existing methods trade efficiency for accuracy: non-diffusion methods are fast but suffer from weak semantics and unstable spatial priors, causing false detections; diffusion-based methods offer high accuracy via strong generative pri

  43. Yuansong Xu, Yuheng Shao, Jiahe Dong, Shaohan Shi

    Medical education increasingly emphasizes students' ability to apply knowledge in real-world clinical settings, focusing on evidence-based clinical reasoning and differential diagnoses. Problem-based learning (PBL) addresses traditional teaching limitations by embedding learning into meaningful contexts and promoting active participation. However, current PB

  44. Xiyuan Wang, Yi-Fan Cao, Junjie Xiong, Sizhe Chen

    Indexical storytelling is gaining popularity in video games, where the narrative unfolds through fragmented clues. This approach fosters player-generated content and discussion, as story interpreters piece together the overarching narrative from these scattered elements. However, the fragmented and non-linear nature of the clues makes systematic categorizati

  45. Huinan Chen, Binbin Cai, Fei Gao, Song Lin

    The Advanced Encryption Standard (AES) is widely used and well-studied for its efficiency and strong security. This paper presents quantum circuit designs for the AES S-box by introducing the composite field \( F((2^4)^2) \) to replace the traditional field \( F(2^8) \), enabling the inversion to be decomposed into operations over \( F(2^4) \). This work red

  46. Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm

    Access to real-world healthcare data is limited by stringent privacy regulations and data imbalances, hindering advancements in research and clinical applications. Synthetic data presents a promising solution, yet existing methods often fail to ensure the realism, utility, and calibration essential for robust survival analysis. Here, we introduce Masked Clin

  47. Yong He, Hongshan Yu, Mingtao Feng, Tongjia Chen

    Diffusion probabilistic models are traditionally used to generate colors at fixed pixel positions in 2D images. Building on this, we extend diffusion models to point cloud semantic segmentation, where point positions also remain fixed, and the diffusion model generates point labels instead of colors. To accelerate the denoising process in reverse diffusion,

  48. Khoa Nguyen, Viet Huynh, Binh Tran, Tri Pham

    Bayesian Optimization (BO) is a well-established method for addressing black-box optimization problems. In many real-world scenarios, optimization often involves multiple functions, emphasizing the importance of leveraging data and learned functions from prior tasks to enhance efficiency in the current task. To expedite convergence to the global optimum, rec

  49. Lunchen Xie, Eugenio Lomurno, Matteo Gambella, Danilo Ardagna

    Differentiable Neural Architecture Search (NAS) provides a promising avenue for automating the complex design of deep learning (DL) models. However, current differentiable NAS methods often face constraints in efficiency, operation selection, and adaptability under varying resource limitations. We introduce ZO-DARTS++, a novel NAS method that effectively bal

  50. Matilde Marcolli, Richard K. Larson

    We give an explicit construction of the generating set of a colored operad that implements theta theory in the mathematical model of Minimalism in generative linguistics, in the form of a coloring algorithm for syntactic objects. We show that the coproduct operation on workspaces allows for a recursive implementation of the theta criterion. We also show that

  51. Cheng Chen, Lixuan Xu, Rong Su

    We present a simple but effective method to extract the reflectivity spectra from the interference signals with a large NA white light interferometry. Numerical simulations and experiments have demonstrated the effectiveness of our method. Furthermore, some insights are also disclosed

  52. David C. Jeong, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar

    Egocentric human body estimation allows for the inference of user body pose and shape from a wearable camera's first-person perspective. Although research has used pose estimation techniques to overcome self-occlusions and image distortions caused by head-mounted fisheye images, similar advances in 3D human mesh recovery (HMR) techniques have been limited. W

  53. Lixuan Xu, Cheng Chen, Rong Su

    In semiconductor manufacturing processes, silicon dioxide films are commonly used as barrier layers, insulating layers, and protective layers. Coherence scanning interferometry (CSI) offers thin film thickness measurements with a millimeter-scale field of view and micrometer-scale lateral resolution. When the film thickness is less than the coherence length

  54. Andrew Crawley, Adam Daigneault, Jonathan Gendron

    From 2000 to 2017, 64% of Maine's pulp and paper processing mills shut down; these closures resulted in harmful effects to communities in Maine and beyond. One question this research asks is how will key macroeconomic and related variables for Maine's forestry and logging industry change in the future? To answer this, we forecast key macroeconomic and relate

  55. Florent Foucaud, Arti Pandey, Kaustav Paul

    Given a graph $G=(V,E)$, a set $S\subseteq V$ is said to be a monitoring edge-geodetic set if the deletion of any edge in the graph results in a change in the distance between at least one pair of vertices in $S$. The minimum size of such a set in $G$ is called the monitoring edge-geodetic number of $G$ and is denoted by $meg(G)$. In this work, we compute th

  56. You Zhang, Jin Wang, Liang-Chih Yu, Dan Xu

    Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods a

  57. Yubin Wang, Xinyang Jiang, De Cheng, Xiangqian Zhao

    Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt e

  58. Manas Gupta, Xuesu Xiao

    Safety has been of paramount importance in motion planning and control techniques and is an active area of research in the past few years. Most safety research for mobile robots target at maintaining safety with the notion of collision avoidance. However, safety goes beyond just avoiding collisions, especially when robots have to navigate unstructured, verti

  59. Serena Dipierro, João Gonçalves da Silva, Giorgio Poggesi, Enrico Valdinoci

    We provide a new proof of the fractional version of the De Giorgi conjecture for the Allen-Cahn equation in $\mathbb{R}^2$ for the full range of exponents. Our proof combines a method introduced by A. Farina in 2003 with the $s$-harmonic extension of the fractional Laplacian in the half-space $\mathbb{R}^{3}_+$ introduced by L. Caffarelli and L. Silvestre in

  60. Wei Junhao, Yu Zhe, Sakuma Jun

    Model merging is a technique that combines multiple finetuned models into a single model without additional training, allowing a free-rider to cheaply inherit specialized capabilities. This study investigates methodologies to suppress unwanted model merging by free-riders. Existing methods such as model watermarking or fingerprinting can only detect merging

  61. Xueqian Sun, Shuyao Qiu, Hao Qin, Yuerui Lu

    Charge carrier transport is pivotal in advancing nanoelectronics. Despite progress in exciton transport within ultra-thin semiconductors, the intertwined transport of free carriers and excitons presents challenges. Surface Acoustic Waves (SAWs) offer a compelling solution, enabling remote, real-time control of excitonic states at room temperature via surfing

  62. Xin Zhang, Dongfang Xu, Jingjing Wang, Shenghui Song

    The reconfigurability of fluid antenna systems (FASs) and reconfigurable intelligent surfaces (RISs) provides significant flexibility in optimizing channel conditions by jointly adjusting the positions of fluid antennas and the phase shifts of RISs. However, it is challenging to acquire the instantaneous channel state information (CSI) for both fluid antenna

  63. Masaki Adachi, Masahiro Fujisawa, Michael A Osborne

    Despite the significance of probabilistic time-series forecasting models, their evaluation metrics often involve intractable integrations. The most widely used metric, the continuous ranked probability score (CRPS), is a strictly proper scoring function; however, its computation requires approximation. We found that popular CRPS estimators--specifically, the

  64. Muhammad Faraz Ul Abrar, Nicolò Michelusi

    Federated learning (FL) has emerged as a promising framework for distributed learning, enabling collaborative model training without sharing private data. Existing wireless FL works primarily adopt two communication strategies: (1) over-the-air (OTA) computation, which exploits wireless signal superposition for simultaneous gradient aggregation, and (2) digi

  65. Lin Zhang, Shengqian Han, Chenyang Yang

    The optimization of multi-user multi-input multi-output (MU-MIMO) precoders is a widely recognized challenging problem. Existing work has demonstrated the potential of graph neural networks (GNNs) in learning precoding policies. However, existing GNNs often exhibit poor generalizability for the numbers of users or antennas. In this paper, we develop a gradie

  66. Sydney Anuyah, Jack Vanschaik, Palak Jain, Sawyer Lehman

    We conduct an empirical analysis of neural network architectures and data transfer strategies for causal relation extraction. By conducting experiments with various contextual embedding layers and architectural components, we show that a relatively straightforward BioBERT-BiGRU relation extraction model generalizes better than other architectures across vary

  67. Cheng Hu, Jihao Huang, Wule Mao, Yonghao Fu

    Generating overtaking trajectories in autonomous racing is a challenging task, as the trajectory must satisfy the vehicle's dynamics and ensure safety and real-time performance running on resource-constrained hardware. This work proposes the Fast and Safe Data-Driven Planner to address this challenge. Sparse Gaussian predictions are introduced to improve bot

  68. Anil Palepu, Valentin Liévin, Wei-Hung Weng, Khaled Saab

    While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic response, and safe medication prescription - remain under-explored. We advance the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMI

  69. Xiang Lan, Feng Wu, Kai He, Qinghao Zhao

    While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between time series signals and visual ECG representations, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG

  70. Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei

    The emergence of Large Language Models (LLMs) has fundamentally transformed natural language processing, making them indispensable across domains ranging from conversational systems to scientific exploration. However, their pre-trained architectures often reveal limitations in specialized contexts, including restricted reasoning capacities, ethical uncertain

  71. Hangyu Du, Chee-Meng Chew

    In recent years, fully differentiable end-to-end autonomous driving systems have become a research hotspot in the field of intelligent transportation. Among various research directions, automatic parking is particularly critical as it aims to enable precise vehicle parking in complex environments. In this paper, we present a purely vision-based transformer m

  72. Ramin Esmzad, Farnaz Adib Yaghmaie, Hamidreza Modares

    This paper bridges optimization and control, and presents a novel closed-loop control framework based on natural gradient descent, offering a trajectory-oriented alternative to traditional cost-function tuning. By leveraging the Fisher Information Matrix, we formulate a preconditioned gradient descent update that explicitly shapes system trajectories. We sho

  73. Hiroki Aoki, Riku Higa, Ryosei Sugawara

    In this paper, we give a short and entirely elementary proof of the proposition ``For any positive integer $ N $, there exists a real number $ L $ such that for any real number $ x \geqq L $, there are at least $ N $ primes in the interval $ [kx, (k+1)x] $'' for $ k \leqq 15 $. Our proof is based on the idea of the proof by Erd\"{o}s for $ k=1 $ and its impr

  74. Hiroshi Funaki, Yuta Sekino, Hiroyuki Tajima, Shota Kisaka

    We propose a novel mechanism for angular momentum (AM) exchange between the crust and core of a neutron star (NS) via the gyromagnetic effect. Using extended hydrodynamics, we model the star by incorporating macroscopic AM and microscopic AM originating from neutron orbital and spin AM. We reveal that macroscopic dynamics in the crust can inform microscopic

  75. Daeheon Jeong, Hyehyun Chu

    Effective communication of UX considerations to stakeholders (e.g., designers and developers) is a critical challenge for UX practitioners. To explore this problem, we interviewed four UX practitioners about their communication challenges and strategies. Our study identifies that providing an example user flow-a screen sequence representing a semantic task-a

  76. Murong Yang, Shihui Ying, Xin-Jian Xu, Yue Gao

    Graph-based multi-view spectral clustering methods have achieved notable progress recently, yet they often fall short in either oversimplifying pairwise relationships or struggling with inefficient spectral decompositions in high-dimensional Euclidean spaces. In this paper, we introduce a novel approach that begins to generate hypergraphs by leveraging spars

  77. Yuchong Gao, Yinding Chi, Mohit Patel, Lishuai Jin

    Wrinkling is commonly observed as mechanical instability when a stiff thin film bound on a compliant thick substrate undergoes in-plane compression exceeding a threshold. Despite significant efforts to create a broad range of surface patterns via wrinkling, little has been studied about a dynamic and transient wrinkling process, where a suspended polymer thi

  78. Wenzhuo Du, Gerun Wang, Guancheng Chen, Hang Zhao

    With the exponential growth of user-generated content on video-sharing platforms, the challenge of facilitating efficient searching and browsing of videos has garnered significant attention. To enhance users' ability to swiftly locate and review pertinent videos, the creation of concise and informative video summaries has become increasingly important. Video

  79. Junyan Lin, Haoran Chen, Yue Fan, Yingqi Fan

    Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing meth

  80. Sanghoon Lee, Fang Wang

    In this paper, we prove a rigidity theorem for Poincar\'e-Einstein manifolds whose conformal infinity is a flat Euclidean space. The proof relies on analyzing the propagation of curvature tensors over the level sets of an adapted boundary defining function. Additionally, we provide examples of Poincar\'e-Einstein manifolds with non-compact conformal infiniti

  81. Xin-Jian Xu, Song-Jie He, Li-Jie Zhang

    In the face of an infectious disease, a key epidemiological measure is the basic reproduction number, which quantifies the average secondary infections caused by a single case in a susceptible population. In practice, the effective reproduction number, denoted as $R_t$, is widely used to assess the transmissibility of the disease at a given time $t$. Real-ti

  82. Md Sadman Sakib, Yu Sun

    Modern robotic systems, deployed across domains from industrial automation to domestic assistance, face a critical challenge: executing tasks with precision and adaptability in dynamic, unpredictable environments. To address this, we propose STAR (Smart Task Adaptation and Recovery), a novel framework that synergizes Foundation Models (FMs) with dynamically

  83. Miguel Contreras, Jessica Sena, Andrea Davidson, Jiaqing Zhang

    Acute brain dysfunction (ABD) is a common, severe ICU complication, presenting as delirium or coma and leading to prolonged stays, increased mortality, and cognitive decline. Traditional screening tools like the Glasgow Coma Scale (GCS), Confusion Assessment Method (CAM), and Richmond Agitation-Sedation Scale (RASS) rely on intermittent assessments, causing

  84. Xinyu Zhou, Yang Li, Jun Zhao

    Semantic communication (SemCom), regarded as the evolution of the traditional Shannon's communication model, stresses the transmission of semantic information instead of the data itself. Federated learning (FL), owing to its distributed learning and privacy-preserving properties, has received attention from both academia and industry. In this paper, we intro

  85. Alin Thomas Tharakan, Prince Philip, Gokulan T., Sumit Kumar

    This paper presents a low power, low cost transceiver architecture to implement radar-on-a-chip. The transceiver comprises of a full ultra-wideband (UWB) transmitter and a full UWB band receiver. A design methodology to maximize the tuning range of the voltage-controlled oscillator (VCO) is presented. At the transmitter side, a sub-harmonic mixer is used for

  86. Weixi Zheng, Aoling Huang, Jingping Yuan, Haoyu Zhao

    In histopathology, intelligent diagnosis of Whole Slide Images (WSIs) is essential for automating and objectifying diagnoses, reducing the workload of pathologists. However, diagnostic models often face the challenge of forgetting previously learned data during incremental training on datasets from different sources. To address this issue, we propose a new f

  87. Nisha Peng, John Stachurski

    New approaches to the theory of dynamic programming view dynamic programs as families of policy operators acting on partially ordered sets. In this paper, we extend these ideas by shifting from arbitrary partially ordered sets to ordered vector spaces. The integrated algebraic and order structure in such spaces leads to sharper fixed point results. These fix

  88. Suvendu Mohanty

    Recent advancements in Artificial Intelligence, particularly in Large Language Models (LLMs), have transformed natural language processing by improving generative capabilities. However, detecting biases embedded within these models remains a challenge. Subtle biases can propagate misinformation, influence decision-making, and reinforce stereotypes, raising e

  89. Runze Zhang, Guoguang Du, Xiaochuan Li, Qi Jia

    Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily focuses on either temporal or spatial consistency, or

  90. Xuexin Chen, Ruichu Cai, Zhengting Huang, Zijian Li

    Synthetic lethality (SL) is a promising gene interaction for cancer therapy. Recent SL prediction methods integrate knowledge graphs (KGs) into graph neural networks (GNNs) and employ attention mechanisms to extract local subgraphs as explanations for target gene pairs. However, attention mechanisms often lack fidelity, typically generate a single explanatio

  91. Guilherme Zeus Dantas e Moura, Olya Mandelshtam

    Non-attacking fillings are combinatorial objects central to the theory of Macdonald polynomials. A probabilistic bijection for partition-shaped non-attacking fillings was introduced by Mandelshtam (2024) to prove a compact formula for symmetric Macdonald polynomials. In this work, we generalize this probabilistic bijection to composition-shaped non-attacking

  92. Alexander Schperberg, Marcel Menner, Stefano Di Cairano

    We propose an online motion planner for legged robot locomotion with the primary objective of achieving energy efficiency. The conceptual idea is to leverage a placement set of footstep positions based on the robot's body position to determine when and how to execute steps. In particular, the proposed planner uses virtual placement sets beneath the hip joint

  93. David Benrimoh, Ryan Smith, Andreea O. Diaconescu, Timothy Friesen

    Studying psychiatric illness has often been limited by difficulties in connecting symptoms and behavior to neurobiology. Computational psychiatry approaches promise to bridge this gap by providing formal accounts of the latent information processing changes that underlie the development and maintenance of psychiatric phenomena. Models based on these theories

  94. Joshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory Shain

    Construction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis). But how much information about constructions does this distribution actually contain? Corpus-based analyses provide some answers, but text alone cannot answer counterfactual questions about what \em

  95. Wenjie Tang, Yuan Zhou, Erqiang Xu, Keyan Cheng

    Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making under uncertainty. However, common existing benchmarks either assess isolated skills, lack environmental diversity, or rely on broad overall metrics. To address these issues, we in

  96. HyunJin Kim, Xiaoyuan Yi, Jing Yao, Muhua Huang

    The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing all humans across measured domains. This gives rise to the critical research question of: As we approach ASI, how do we

  97. Zhongzhan Huang, Guoming Ling, Yupei Lin, Yandong Chen

    Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable router can significantly enhance the performance of this paradig

  98. Sung Jae Jun, Sokbae Lee

    Televised debates between presidential candidates are often regarded as the exemplar of persuasive communication. Yet, recent evidence from Le Pennec and Pons (2023) indicates that they may not sway voters as strongly as popular belief suggests. We revisit their findings through the lens of the persuasion rate and introduce a robust framework that does not r

  99. Matthew Nyaaba, Min SungEun, Mary Abiswin Apam, Kwame Owoahene Acheampong

    This study highlights the transparency and accuracy of GenAI's inductive thematic analysis, particularly using GPT-4 Turbo API integrated within a stepwise prompt-based Python script. This approach ensured a traceable and systematic coding process, generating codes with supporting statements and page references, which enhanced validation and reproducibility.

  100. Avimita Chatterjee, Archisman Ghosh, Swaroop Ghosh

    Quantum Error Correction (QEC), combined with magic state distillation, ensures fault tolerance in large-scale quantum computation. To apply QEC, a circuit must first be transformed into a non-Clifford (or T) gate set. T-depth, the number of sequential T-gate layers, determines the magic state cost, impacting both spatial and temporal overhead. Minimizing T-