March 2025 arXiv papers — page 181
Showing 18,001–18,100 of 23,633 papers
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
cs.CVMingxing Li, Rui Wang, Lei Sun, Yancheng Bai
The rapid expansion of mobile internet has resulted in a substantial increase in user-generated content (UGC) images, thereby making the thorough assessment of UGC images both urgent and essential. Recently, multimodal large language models (MLLMs) have shown great potential in image quality assessment (IQA) and image aesthetic assessment (IAA). Despite this
Bohan Liu, Xiaosen Wang
Transfer-based attacks pose a significant threat to real-world applications by directly targeting victim models with adversarial examples generated on surrogate models. While numerous approaches have been proposed to enhance adversarial transferability, existing works often overlook the intrinsic relationship between adversarial perturbations and input image
Mingyang Song, Mao Zheng, Xuan Luo
Pairwise LLM-as-a-judge evaluation asks the judge to identify the \emph{better} of two candidate answers. We study a one-line modification that asks for the \emph{worse} answer instead and recovers the preference by elimination, a procedure we call Goal-Reversed Prompting (GRP). GRP introduces no extra inference rounds, composes with any prompt template (dir
Tadahiro Taniguchi, Yasushi Hirai, Masahiro Suzuki, Shingo Murata
This paper introduces the System 0/1/2/3 framework as an extension of dual-process theory, employing a quad-process model of cognition. Expanding upon System 1 (fast, intuitive thinking) and System 2 (slow, deliberative thinking), we incorporate System 0, which represents pre-cognitive embodied processes, and System 3, which encompasses collective intelligen
Jie He, Wanqiu Long, Deyi Xiong
Large pre-trained neural models have achieved remarkable success in natural language process (NLP), inspiring a growing body of research analyzing their ability from different aspects. In this paper, we propose a test suite to evaluate the cohesive ability of pre-trained language models. The test suite contains multiple cohesion phenomena between adjacent an
Khang Nguyen, An T. Le, Tien Pham, Manfred Huber
Prior flow matching methods in robotics have primarily learned velocity fields to morph one distribution of trajectories into another. In this work, we extend flow matching to capture second-order trajectory dynamics, incorporating acceleration effects either explicitly in the model or implicitly through the learning objective. Unlike diffusion models, which
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
cs.CVJian Ma, Qirong Peng, Xu Guo, Chen Chen
Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in understanding and integrating multiple modalities. However, currently there is no straightforward and efficient framework to transfer the multimodal comprehension abilities of MLL
Biplab Basak, Sourav Sarkar
This article focuses on a class of properly edge-colored graphs, which arise from topological combinatorics, and investigates their embeddings onto surfaces. Specifically, these graphs are known as the dual graphs of balanced normal pseudomanifolds. We introduce the concept of the balanced genus, which represents the smallest genus of a surface onto which th
Xiangxiang Chu, Renda Li, Yong Wang
Recent studies have highlighted the interplay between diffusion models and representation learning. Intermediate representations from diffusion models can be leveraged for downstream visual tasks, while self-supervised vision models can enhance the convergence and generation quality of diffusion models. However, transferring pretrained weights from vision mo
Koga Okubo, Kanta Tachibana
This study focuses on the rotation of the hips and shoulders during a baseball bat swing, analyzing the time-series changes in rotational angles, rotational velocities, and axes using marker position data obtained from a motion capture system with 12 infrared cameras. Previous studies have examined factors such as ground reaction forces, muscle activation pa
Zhu-Yu Ren, Sheng-Quan Wang, Jian-Ming Shen, Xing-Gang Wu
Theoretical calculations for event shape observables are often determined by using the conventional scale setting; i.e. the procedure defined by setting the renormalization scale to the center-of-mass energy $\mu_r=\sqrt{s}$ and evaluating theoretical uncertainties by varying the same scale $\mu_r$ in an arbitrary range. Both the event shape distributions an
Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Flexible and Effective Paradigm
cs.CVJiebin Yan, Kangcheng Wu, Junjie Chen, Ziwen Tan
Most of existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, these methods are either computationally expensive or less scalable. To solve these issues, in this paper, we present a flexible and effective
R&D of KLM Upgrade for Direct Measurement of Neutral Hadron Momentum via Time-of-Flight in Belle II
physics.ins-detXiyang Wang, Hongyu Zhang, Shiming Zou, Zibing Bai
Accurate momentum determination of a neutral hadron, such as a KL meson or a neutron, remains a significant challenge in particle physics and nuclear physics experiments. The Belle II experiment presents an opportunity to address this challenge through an upgrade incorporating Time-of-Flight (TOF) capability for its large KL and Muon Detector (KLM). We inves
Yunrui Zheng
We consider the evolution of contact lines for thermal convection of viscous fluids in a 2D open-top vessel. The domain is bounded above by a free moving boundary and otherwise by the solid wall of a vessel. The dynamics of the fluid are governed by the incompressible Boussinesq approximation under the influence of gravity, and the interface between fluid an
Yuanlong Ruan
We perform a complete analysis of the limiting behaviour of a class of quasilinear problems with Dirichlet boundary data g. We show that the Lipschitz constant of g plays a role in controlling the Gamma-convergence of the natural energies. However the solutions converge uniformly to solution of a limiting equation irrelevant to the Lipschitz constant of g. T
Kai Yang, Zijian Bai, Yang Xiao, Xinyu Li
3D reconstruction garners increasing attention alongside the advancement of high-level image applications, where dense stereo matching (DSM) serves as a pivotal technique. Previous studies often rely on publicly available datasets for training, focusing on modifying network architectures or incorporating specialized modules to extract domain-invariant featur
Ashutosh Kumar, Sourabh Lahiri, Trilochan Bagarti, Subhashish Banerjee
Quantum heat engines have undergone extensive studies over the last two decades. Simultaneously, the studies of the applications of stochastic resetting in various fields are on the rise. We explore the effect of stochastic resetting on the dynamics of a two-level and a three-level quantum heat engine. The extracted work is shown to increase with the resetti
Mechanism of Electricacupuncture Treating Detrusor Bladder Neck Dyscoordination After Suprasacral Spinal Cord Injury by Proteomics
q-bio.BMLiya Tang, Qirui Qu, Yiying Long, Xia Wu
Objectives This study aimed to elucidate the potential mechanisms of electroacupuncture (EA) in restoring detrusor-bladder neck dyssynergesia (DBND) following suprasacral spinal cord injury. Methods A total of 52 adult female Sprague-Dawley rats were randomly assigned to either a sham group (n=12) or a spinal cord injury model group (n=40). In the model grou
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
cs.HCYingna Wang, Qingqin Liu, Xiaoying Wei, Mingming Fan
The essence of intangible cultural heritage (ICH) lies in the living knowledge and skills passed down through generations. Daily practice plays a vital role in revitalizing ICH by fostering continuous learning and improvement. However, limited resources and accessibility pose significant challenges to sustaining such practice. Virtual reality (VR) has shown
BlackGoose Rimer: Harnessing RWKV-7 as a Simple yet Superior Replacement for Transformers in Large-Scale Time Series Modeling
cs.LGLi weile, Liu Xiao
Time series models face significant challenges in scaling to handle large and complex datasets, akin to the scaling achieved by large language models (LLMs). The unique characteristics of time series data and the computational demands of model scaling necessitate innovative approaches. While researchers have explored various architectures such as Transformer
Constraints on Evolutions of Fundamental Constants from Clustering of Fast Radio Burst Dispersion Measure
astro-ph.COShi-Yuan Wang, Jun-Qing Xia
Constrained measurements of fundamental physical constants using astronomical observational data represent a powerful method for investigating potential new physics. In particular, the dispersion measure (DM) of fast radio bursts (FRBs), which probes the electron density along their propagation paths, may be influenced by the space-time variation of the fine
Unlocking Pretrained LLMs for Motion-Related Multimodal Generation: A Fine-Tuning Approach to Unify Diffusion and Next-Token Prediction
cs.LGShinichi Tanaka, Zhao Wang, Yoichi Kato, Jun Ohya
In this paper, we propose a unified framework that leverages a single pretrained LLM for Motion-related Multimodal Generation, referred to as MoMug. MoMug integrates diffusion-based continuous motion generation with the model's inherent autoregressive discrete text prediction capabilities by fine-tuning a pretrained LLM. This enables seamless switching betwe
Xuanyu Zhang, Jiarui Meng, Zhipei Xu, Shuzhou Yang
3D Gaussian Splatting (3DGS) has emerged as a premier method for 3D representation due to its real-time rendering and high-quality outputs, underscoring the critical need to protect the privacy of 3D assets. Traditional NeRF steganography methods fail to address the explicit nature of 3DGS since its point cloud files are publicly accessible. Existing GS steg
Purbid Bambroo, Subinay Adhikary, Paheli Bhattacharya, Abhijnan Chakraborty
Identification of rhetorical roles like facts, arguments, and final judgments is central to understanding a legal case document and can lend power to other downstream tasks like legal case summarization and judgment prediction. However, there are several challenges to this task. Legal documents are often unstructured and contain a specialized vocabulary, mak
Hongjia Zhai, Boming Zhao, Hai Li, Xiaokun Pan
Recently, neural radiance fields (NeRF) have gained significant attention in the field of visual localization. However, existing NeRF-based approaches either lack geometric constraints or require extensive storage for feature matching, limiting their practical applications. To address these challenges, we propose an efficient and novel visual localization ap
Bing Chen, Xiang Liu
Hybrid mesons occupy a unique position in the hadronic spectrum, offering a valuable insight into the non-perturbative nature of the glue field. Hybrid states with exotic $J^{PC}$ quantum numbers, such as the $1^{-+}$, $0^{+-}$ and $2^{+-}$, cannot mix with conventional $q\bar{q}$ mesons, making them critical for establishing the full spectrum of hybrid meso
LimTopic: LLM-based Topic Modeling and Text Summarization for Analyzing Scientific Articles limitations
cs.CLIbrahim Al Azhar, Venkata Devesh Reddy, Hamed Alhoori, Akhil Pandey Akella
The limitations sections of scientific articles play a crucial role in highlighting the boundaries and shortcomings of research, thereby guiding future studies and improving research methods. Analyzing these limitations benefits researchers, reviewers, funding agencies, and the broader academic community. We introduce LimTopic, a strategy where Topic generat
Qi Zhang, Xiuyuan Chen, Ziyi He, Lianming Wu
Cervical spondylosis, a complex and prevalent condition, demands precise and efficient diagnostic techniques for accurate assessment. While MRI offers detailed visualization of cervical spine anatomy, manual interpretation remains labor-intensive and prone to error. To address this, we developed an innovative AI-assisted Expert-based Diagnosis System that au
Dead, Slow and Overworked Graphite: Operando X-ray Microdiffraction Mapping of Aged Electrodes
cond-mat.mtrl-sciGozde Oney, Federico Monaco, Saptarshee Mitra, Asma Medjahed
Aging limits lithium-ion battery lifetime and must be understood to improve durability and performance, requiring a detailed understanding of how aging alters the availability of cyclable lithium and the integrity of active particles. In this work, (de)lithiation mechanisms are examined and spatially-resolved at the microscale in aged graphite electrodes dis
AF-KAN: Activation Function-Based Kolmogorov-Arnold Networks for Efficient Representation Learning
cs.LGHoang-Thang Ta, Anh Tran
Kolmogorov-Arnold Networks (KANs) have inspired numerous works exploring their applications across a wide range of scientific problems, with the potential to replace Multilayer Perceptrons (MLPs). While many KANs are designed using basis and polynomial functions, such as B-splines, ReLU-KAN utilizes a combination of ReLU functions to mimic the structure of B
Nikola Sandrić
In this note, we discuss the uniform ergodicity of a diffusion process given by an It\^o stochastic differential equation. We present an integral condition in terms of the drift and diffusion coefficients that ensures the uniform ergodicity of the corresponding transition kernel with respect to the total variation distance. Applications of the obtained resul
Janet Rafner, Ryan Q. Guloy, Eden W. Wen, Catherine M. Chiodo
Generative AI (GenAI) chatbots are becoming increasingly integrated into virtual assistant technologies, yet their success hinges on the ability to gather meaningful user feedback to improve interaction quality, system outcomes, and overall user acceptance. Successful chatbot interactions can enable organizations to build long-term relationships with their c
Keshu Wu, Xinyue Ye, Suphanut Jamonnak, Xin Feng
Daily operations in large campuses depend on how efficiently people \emph{move} through space and time. In this sense, course timetables are more than administrative schedules: they act as mobility policies that orchestrate thousands of trajectories, shaping travel burden, congestion, accessibility, and the reliability of back-to-back transitions. Designing
Multi-modal expressive personality recognition in data non-ideal audiovisual based on multi-scale feature enhancement and modal augment
cs.SDWeixuan Kong, Jinpeng Yu, Zijun Li, Hanwei Liu
Automatic personality recognition is a research hotspot in the intersection of computer science and psychology, and in human-computer interaction, personalised has a wide range of applications services and other scenarios. In this paper, an end-to-end multimodal performance personality is established for both visual and auditory modal datarecognition network
Akshat Jain
This paper presents a novel approach to image dehazing by combining Feature Fusion Attention (FFA) networks with CycleGAN architecture. Our method leverages both supervised and unsupervised learning techniques to effectively remove haze from images while preserving crucial image details. The proposed hybrid architecture demonstrates significant improvements
Kuanghong Liu, Jin Wang, Kangjian He, Dan Xu
Conventional multi-source domain few-shot adaptation (MFDA) faces the challenge of further reducing the load on edge-side devices in low-resource scenarios. Considering the native language-supervised advantage of CLIP and the plug-and-play nature of prompt to transfer CLIP efficiently, this paper introduces an uploadable multi-source few-shot domain adaptati
Prefer2SD: A Human-in-the-Loop Approach to Balancing Similarity and Diversity in In-Game Friend Recommendations
cs.HCXiyuan Wang, Ziang Li, Sizhe Chen, Xingxing Xing
In-game friend recommendations significantly impact player retention and sustained engagement in online games. Balancing similarity and diversity in recommendations is crucial for fostering stronger social bonds across diverse player groups. However, automated recommendation systems struggle to achieve this balance, especially as player preferences evolve ov
Syed Sajid Ullah, Li Gang, Mudassir Riaz, Ahsan Ashfaq
Handwritten digit recognition remains a fundamental challenge in computer vision, with applications ranging from postal code reading to document digitization. This paper presents an ensemble-based approach that combines Convolutional Neural Networks (CNNs) with traditional machine learning techniques to improve recognition accuracy and robustness. We evaluat
Avadhut V. Purohit, Udaysinh T. Bhosale
We study the double kicked top (DKT), which is an extension of the standard quantum kicked top (QKT) model. The model allows us to study the transition from time-reversal symmetric to broken time-reversal symmetric dynamics. Our transformation in the kick strength parameter space $(k, k') \to (k_r, k_\theta)$ reveals interesting features. The transformed kic
Naoyuki Monden, Reo Yabuguchi
We show that for any positive integer $h$, a knot surgered elliptic surface $E(n)_{T(2,2h+1)}$ for a $(2,2h+1)$-torus knot $T(2,2h+1)$ and the elliptic surface $E(1)_{2,2h+1}$ admit handle decompositions without 1- and 3-handles using the Kirby diagrams derived from Lefschetz fibrations on them.
ULTHO: Ultra-Lightweight yet Efficient Hyperparameter Optimization in Deep Reinforcement Learning
cs.LGMingqi Yuan, Bo Li, Xin Jin, Wenjun Zeng
Hyperparameter optimization (HPO) is a billion-dollar problem in machine learning, which significantly impacts the training efficiency and model performance. However, achieving efficient and robust HPO in deep reinforcement learning (RL) is consistently challenging due to its high non-stationarity and computational cost. To tackle this problem, existing appr
High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy
cs.CVXianjie Liu, Keren Fu, Qijun Zhao
High-precision dichotomous image segmentation (DIS) is a task of extracting fine-grained objects from high-resolution images. Existing methods trade efficiency for accuracy: non-diffusion methods are fast but suffer from weak semantics and unstable spatial priors, causing false detections; diffusion-based methods offer high accuracy via strong generative pri
Advancing Problem-Based Learning with Clinical Reasoning for Improved Differential Diagnosis in Medical Education
cs.HCYuansong Xu, Yuheng Shao, Jiahe Dong, Shaohan Shi
Medical education increasingly emphasizes students' ability to apply knowledge in real-world clinical settings, focusing on evidence-based clinical reasoning and differential diagnoses. Problem-based learning (PBL) addresses traditional teaching limitations by embedding learning into meaningful contexts and promoting active participation. However, current PB
Xiyuan Wang, Yi-Fan Cao, Junjie Xiong, Sizhe Chen
Indexical storytelling is gaining popularity in video games, where the narrative unfolds through fragmented clues. This approach fosters player-generated content and discussion, as story interpreters piece together the overarching narrative from these scattered elements. However, the fragmented and non-linear nature of the clues makes systematic categorizati
Huinan Chen, Binbin Cai, Fei Gao, Song Lin
The Advanced Encryption Standard (AES) is widely used and well-studied for its efficiency and strong security. This paper presents quantum circuit designs for the AES S-box by introducing the composite field \( F((2^4)^2) \) to replace the traditional field \( F(2^8) \), enabling the inversion to be decomposed into operations over \( F(2^4) \). This work red
Attention-Based Synthetic Data Generation for Calibration-Enhanced Survival Analysis: A Case Study for Chronic Kidney Disease Using Electronic Health Records
cs.LGNicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm
Access to real-world healthcare data is limited by stringent privacy regulations and data imbalances, hindering advancements in research and clinical applications. Synthetic data presents a promising solution, yet existing methods often fail to ensure the realism, utility, and calibration essential for robust survival analysis. Here, we introduce Masked Clin
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
cs.CVYong He, Hongshan Yu, Mingtao Feng, Tongjia Chen
Diffusion probabilistic models are traditionally used to generate colors at fixed pixel positions in 2D images. Building on this, we extend diffusion models to point cloud semantic segmentation, where point positions also remain fixed, and the diffusion model generates point labels instead of colors. To accelerate the denoising process in reverse diffusion,
Khoa Nguyen, Viet Huynh, Binh Tran, Tri Pham
Bayesian Optimization (BO) is a well-established method for addressing black-box optimization problems. In many real-world scenarios, optimization often involves multiple functions, emphasizing the importance of leveraging data and learned functions from prior tasks to enhance efficiency in the current task. To expedite convergence to the global optimum, rec
Lunchen Xie, Eugenio Lomurno, Matteo Gambella, Danilo Ardagna
Differentiable Neural Architecture Search (NAS) provides a promising avenue for automating the complex design of deep learning (DL) models. However, current differentiable NAS methods often face constraints in efficiency, operation selection, and adaptability under varying resource limitations. We introduce ZO-DARTS++, a novel NAS method that effectively bal
Matilde Marcolli, Richard K. Larson
We give an explicit construction of the generating set of a colored operad that implements theta theory in the mathematical model of Minimalism in generative linguistics, in the form of a coloring algorithm for syntactic objects. We show that the coproduct operation on workspaces allows for a recursive implementation of the theta criterion. We also show that
Cheng Chen, Lixuan Xu, Rong Su
We present a simple but effective method to extract the reflectivity spectra from the interference signals with a large NA white light interferometry. Numerical simulations and experiments have demonstrated the effectiveness of our method. Furthermore, some insights are also disclosed
David C. Jeong, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar
Egocentric human body estimation allows for the inference of user body pose and shape from a wearable camera's first-person perspective. Although research has used pose estimation techniques to overcome self-occlusions and image distortions caused by head-mounted fisheye images, similar advances in 3D human mesh recovery (HMR) techniques have been limited. W
Error Source Sensitivity Analysis in Model-Based Coherence Scanning Interferometry for Thin Film Metrology
physics.opticsLixuan Xu, Cheng Chen, Rong Su
In semiconductor manufacturing processes, silicon dioxide films are commonly used as barrier layers, insulating layers, and protective layers. Coherence scanning interferometry (CSI) offers thin film thickness measurements with a millimeter-scale field of view and micrometer-scale lateral resolution. When the film thickness is less than the coherence length
Andrew Crawley, Adam Daigneault, Jonathan Gendron
From 2000 to 2017, 64% of Maine's pulp and paper processing mills shut down; these closures resulted in harmful effects to communities in Maine and beyond. One question this research asks is how will key macroeconomic and related variables for Maine's forestry and logging industry change in the future? To answer this, we forecast key macroeconomic and relate
Florent Foucaud, Arti Pandey, Kaustav Paul
Given a graph $G=(V,E)$, a set $S\subseteq V$ is said to be a monitoring edge-geodetic set if the deletion of any edge in the graph results in a change in the distance between at least one pair of vertices in $S$. The minimum size of such a set in $G$ is called the monitoring edge-geodetic number of $G$ and is denoted by $meg(G)$. In this work, we compute th
Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian Perspective
cs.CLYou Zhang, Jin Wang, Liang-Chih Yu, Dan Xu
Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent of the impact of non-IID information and how these methods a
Yubin Wang, Xinyang Jiang, De Cheng, Xiangqian Zhao
Visual prompt tuning offers significant advantages for adapting pre-trained visual foundation models to specific tasks. However, current research provides limited insight into the interpretability of this approach, which is essential for enhancing AI reliability and enabling AI-driven knowledge discovery. In this paper, rather than learning abstract prompt e
T-CBF: Traversability-based Control Barrier Function to Navigate Vertically Challenging Terrain
cs.ROManas Gupta, Xuesu Xiao
Safety has been of paramount importance in motion planning and control techniques and is an active area of research in the past few years. Most safety research for mobile robots target at maintaining safety with the notion of collision avoidance. However, safety goes beyond just avoiding collisions, especially when robots have to navigate unstructured, verti
Serena Dipierro, João Gonçalves da Silva, Giorgio Poggesi, Enrico Valdinoci
We provide a new proof of the fractional version of the De Giorgi conjecture for the Allen-Cahn equation in $\mathbb{R}^2$ for the full range of exponents. Our proof combines a method introduced by A. Farina in 2003 with the $s$-harmonic extension of the fractional Laplacian in the half-space $\mathbb{R}^{3}_+$ introduced by L. Caffarelli and L. Silvestre in
Wei Junhao, Yu Zhe, Sakuma Jun
Model merging is a technique that combines multiple finetuned models into a single model without additional training, allowing a free-rider to cheaply inherit specialized capabilities. This study investigates methodologies to suppress unwanted model merging by free-riders. Existing methods such as model watermarking or fingerprinting can only detect merging
Optical Visualization of Carrier Surfing in 2D Monolayers Driven by Surface Acoustic Waves
physics.opticsXueqian Sun, Shuyao Qiu, Hao Qin, Yuerui Lu
Charge carrier transport is pivotal in advancing nanoelectronics. Despite progress in exciton transport within ultra-thin semiconductors, the intertwined transport of free carriers and excitons presents challenges. Surface Acoustic Waves (SAWs) offer a compelling solution, enabling remote, real-time control of excitonic states at room temperature via surfing
Fluid Antenna Meets RIS: Random Matrix Analysis and Two-Timescale Design for Multi-User Communications
cs.ITXin Zhang, Dongfang Xu, Jingjing Wang, Shenghui Song
The reconfigurability of fluid antenna systems (FASs) and reconfigurable intelligent surfaces (RISs) provides significant flexibility in optimizing channel conditions by jointly adjusting the positions of fluid antennas and the phase shifts of RISs. However, it is challenging to acquire the instantaneous channel state information (CSI) for both fluid antenna
Masaki Adachi, Masahiro Fujisawa, Michael A Osborne
Despite the significance of probabilistic time-series forecasting models, their evaluation metrics often involve intractable integrations. The most widely used metric, the continuous ranked probability score (CRPS), is a strictly proper scoring function; however, its computation requires approximation. We found that popular CRPS estimators--specifically, the
Muhammad Faraz Ul Abrar, Nicolò Michelusi
Federated learning (FL) has emerged as a promising framework for distributed learning, enabling collaborative model training without sharing private data. Existing wireless FL works primarily adopt two communication strategies: (1) over-the-air (OTA) computation, which exploits wireless signal superposition for simultaneous gradient aggregation, and (2) digi
Lin Zhang, Shengqian Han, Chenyang Yang
The optimization of multi-user multi-input multi-output (MU-MIMO) precoders is a widely recognized challenging problem. Existing work has demonstrated the potential of graph neural networks (GNNs) in learning precoding policies. However, existing GNNs often exhibit poor generalizability for the numbers of users or antennas. In this paper, we develop a gradie
Sydney Anuyah, Jack Vanschaik, Palak Jain, Sawyer Lehman
We conduct an empirical analysis of neural network architectures and data transfer strategies for causal relation extraction. By conducting experiments with various contextual embedding layers and architectural components, we show that a relatively straightforward BioBERT-BiGRU relation extraction model generalizes better than other architectures across vary
FSDP: Fast and Safe Data-Driven Overtaking Trajectory Planning for Head-to-Head Autonomous Racing Competitions
cs.ROCheng Hu, Jihao Huang, Wule Mao, Yonghao Fu
Generating overtaking trajectories in autonomous racing is a challenging task, as the trajectory must satisfy the vehicle's dynamics and ensure safety and real-time performance running on resource-constrained hardware. This work proposes the Fast and Safe Data-Driven Planner to address this challenge. Sparse Gaussian predictions are introduced to improve bot
Anil Palepu, Valentin Liévin, Wei-Hung Weng, Khaled Saab
While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic response, and safe medication prescription - remain under-explored. We advance the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMI
Xiang Lan, Feng Wu, Kai He, Qinghao Zhao
While recent multimodal large language models (MLLMs) have advanced automated ECG interpretation, they still face two key limitations: (1) insufficient multimodal synergy between time series signals and visual ECG representations, and (2) limited explainability in linking diagnoses to granular waveform evidence. We introduce GEM, the first MLLM unifying ECG
Guiyao Tie, Zeli Zhao, Dingjie Song, Fuyang Wei
The emergence of Large Language Models (LLMs) has fundamentally transformed natural language processing, making them indispensable across domains ranging from conversational systems to scientific exploration. However, their pre-trained architectures often reveal limitations in specialized contexts, including restricted reasoning capacities, ethical uncertain
TransParking: A Dual-Decoder Transformer Framework with Soft Localization for End-to-End Automatic Parking
cs.CVHangyu Du, Chee-Meng Chew
In recent years, fully differentiable end-to-end autonomous driving systems have become a research hotspot in the field of intelligent transportation. Among various research directions, automatic parking is particularly critical as it aims to enable precise vehicle parking in complex environments. In this paper, we present a purely vision-based transformer m
Ramin Esmzad, Farnaz Adib Yaghmaie, Hamidreza Modares
This paper bridges optimization and control, and presents a novel closed-loop control framework based on natural gradient descent, offering a trajectory-oriented alternative to traditional cost-function tuning. By leveraging the Fisher Information Matrix, we formulate a preconditioned gradient descent update that explicitly shapes system trajectories. We sho
Hiroki Aoki, Riku Higa, Ryosei Sugawara
In this paper, we give a short and entirely elementary proof of the proposition ``For any positive integer $ N $, there exists a real number $ L $ such that for any real number $ x \geqq L $, there are at least $ N $ primes in the interval $ [kx, (k+1)x] $'' for $ k \leqq 15 $. Our proof is based on the idea of the proof by Erd\"{o}s for $ k=1 $ and its impr
Hiroshi Funaki, Yuta Sekino, Hiroyuki Tajima, Shota Kisaka
We propose a novel mechanism for angular momentum (AM) exchange between the crust and core of a neutron star (NS) via the gyromagnetic effect. Using extended hydrodynamics, we model the star by incorporating macroscopic AM and microscopic AM originating from neutron orbital and spin AM. We reveal that macroscopic dynamics in the crust can inform microscopic
Daeheon Jeong, Hyehyun Chu
Effective communication of UX considerations to stakeholders (e.g., designers and developers) is a critical challenge for UX practitioners. To explore this problem, we interviewed four UX practitioners about their communication challenges and strategies. Our study identifies that providing an example user flow-a screen sequence representing a semantic task-a
Murong Yang, Shihui Ying, Xin-Jian Xu, Yue Gao
Graph-based multi-view spectral clustering methods have achieved notable progress recently, yet they often fall short in either oversimplifying pairwise relationships or struggling with inefficient spectral decompositions in high-dimensional Euclidean spaces. In this paper, we introduce a novel approach that begins to generate hypergraphs by leveraging spars
Geometrically Templated Dynamic Wrinkling from Suspended Poly(vinyl alcohol) Soap Films
cond-mat.softYuchong Gao, Yinding Chi, Mohit Patel, Lishuai Jin
Wrinkling is commonly observed as mechanical instability when a stiff thin film bound on a compliant thick substrate undergoes in-plane compression exceeding a threshold. Despite significant efforts to create a broad range of surface patterns via wrinkling, little has been studied about a dynamic and transient wrinkling process, where a suspended polymer thi
Wenzhuo Du, Gerun Wang, Guancheng Chen, Hang Zhao
With the exponential growth of user-generated content on video-sharing platforms, the challenge of facilitating efficient searching and browsing of videos has garnered significant attention. To enhance users' ability to swiftly locate and review pertinent videos, the creation of concise and informative video summaries has become increasingly important. Video
Junyan Lin, Haoran Chen, Yue Fan, Yingqi Fan
Multimodal Large Language Models (MLLMs) have made significant advancements in recent years, with visual features playing an increasingly critical role in enhancing model performance. However, the integration of multi-layer visual features in MLLMs remains underexplored, particularly with regard to optimal layer selection and fusion strategies. Existing meth
Sanghoon Lee, Fang Wang
In this paper, we prove a rigidity theorem for Poincar\'e-Einstein manifolds whose conformal infinity is a flat Euclidean space. The proof relies on analyzing the propagation of curvature tensors over the level sets of an adapted boundary defining function. Additionally, we provide examples of Poincar\'e-Einstein manifolds with non-compact conformal infiniti
Improved estimation of the effective reproduction number with heterogeneous transmission rates and reporting delays
physics.soc-phXin-Jian Xu, Song-Jie He, Li-Jie Zhang
In the face of an infectious disease, a key epidemiological measure is the basic reproduction number, which quantifies the average secondary infections caused by a single case in a susceptible population. In practice, the effective reproduction number, denoted as $R_t$, is widely used to assess the transmissibility of the disease at a given time $t$. Real-ti
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
cs.ROMd Sadman Sakib, Yu Sun
Modern robotic systems, deployed across domains from industrial automation to domestic assistance, face a critical challenge: executing tasks with precision and adaptability in dynamic, unpredictable environments. To address this, we propose STAR (Smart Task Adaptation and Recovery), a novel framework that synergizes Foundation Models (FMs) with dynamically
MANDARIN: Mixture-of-Experts Framework for Dynamic Delirium and Coma Prediction in ICU Patients: Development and Validation of an Acute Brain Dysfunction Prediction Model
cs.AIMiguel Contreras, Jessica Sena, Andrea Davidson, Jiaqing Zhang
Acute brain dysfunction (ABD) is a common, severe ICU complication, presenting as delirium or coma and leading to prolonged stays, increased mortality, and cognitive decline. Traditional screening tools like the Glasgow Coma Scale (GCS), Confusion Assessment Method (CAM), and Richmond Agitation-Sedation Scale (RASS) rely on intermittent assessments, causing
Xinyu Zhou, Yang Li, Jun Zhao
Semantic communication (SemCom), regarded as the evolution of the traditional Shannon's communication model, stresses the transmission of semantic information instead of the data itself. Federated learning (FL), owing to its distributed learning and privacy-preserving properties, has received attention from both academia and industry. In this paper, we intro
Alin Thomas Tharakan, Prince Philip, Gokulan T., Sumit Kumar
This paper presents a low power, low cost transceiver architecture to implement radar-on-a-chip. The transceiver comprises of a full ultra-wideband (UWB) transmitter and a full UWB band receiver. A design methodology to maximize the tuning range of the voltage-controlled oscillator (VCO) is presented. At the transmitter side, a sub-harmonic mixer is used for
Pathological Prior-Guided Multiple Instance Learning For Mitigating Catastrophic Forgetting in Breast Cancer Whole Slide Image Classification
cs.CVWeixi Zheng, Aoling Huang, Jingping Yuan, Haoyu Zhao
In histopathology, intelligent diagnosis of Whole Slide Images (WSIs) is essential for automating and objectifying diagnoses, reducing the workload of pathologists. However, diagnostic models often face the challenge of forgetting previously learned data during incremental training on datasets from different sources. To address this issue, we propose a new f
Nisha Peng, John Stachurski
New approaches to the theory of dynamic programming view dynamic programs as families of policy operators acting on partially ordered sets. In this paper, we extend these ideas by shifting from arbitrary partially ordered sets to ordered vector spaces. The integrated algebraic and order structure in such spaces leads to sharper fixed point results. These fix
Suvendu Mohanty
Recent advancements in Artificial Intelligence, particularly in Large Language Models (LLMs), have transformed natural language processing by improving generative capabilities. However, detecting biases embedded within these models remains a challenge. Subtle biases can propagate misinformation, influence decision-making, and reinforce stereotypes, raising e
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
cs.CVRunze Zhang, Guoguang Du, Xiaochuan Li, Qi Jia
Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily focuses on either temporal or spatial consistency, or
Interpretable High-order Knowledge Graph Neural Network for Predicting Synthetic Lethality in Human Cancers
cs.LGXuexin Chen, Ruichu Cai, Zhengting Huang, Zijian Li
Synthetic lethality (SL) is a promising gene interaction for cancer therapy. Recent SL prediction methods integrate knowledge graphs (KGs) into graph neural networks (GNNs) and employ attention mechanisms to extract local subgraphs as explanations for target gene pairs. However, attention mechanisms often lack fidelity, typically generate a single explanatio
Guilherme Zeus Dantas e Moura, Olya Mandelshtam
Non-attacking fillings are combinatorial objects central to the theory of Macdonald polynomials. A probabilistic bijection for partition-shaped non-attacking fillings was introduced by Mandelshtam (2024) to prove a compact formula for symmetric Macdonald polynomials. In this work, we generalize this probabilistic bijection to composition-shaped non-attacking
Alexander Schperberg, Marcel Menner, Stefano Di Cairano
We propose an online motion planner for legged robot locomotion with the primary objective of achieving energy efficiency. The conceptual idea is to leverage a placement set of footstep positions based on the robot's body position to determine when and how to execute steps. In particular, the proposed planner uses virtual placement sets beneath the hip joint
David Benrimoh, Ryan Smith, Andreea O. Diaconescu, Timothy Friesen
Studying psychiatric illness has often been limited by difficulties in connecting symptoms and behavior to neurobiology. Computational psychiatry approaches promise to bridge this gap by providing formal accounts of the latent information processing changes that underlie the development and maintenance of psychiatric phenomena. Models based on these theories
Joshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory Shain
Construction grammar posits that constructions, or form-meaning pairings, are acquired through experience with language (the distributional learning hypothesis). But how much information about constructions does this distribution actually contain? Corpus-based analyses provide some answers, but text alone cannot answer counterfactual questions about what \em
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
cs.AIWenjie Tang, Yuan Zhou, Erqiang Xu, Keyan Cheng
Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making under uncertainty. However, common existing benchmarks either assess isolated skills, lack environmental diversity, or rely on broad overall metrics. To address these issues, we in
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
cs.AIHyunJin Kim, Xiaoyuan Yi, Jing Yao, Muhua Huang
The recent leap in AI capabilities, driven by big generative models, has sparked the possibility of achieving Artificial General Intelligence (AGI) and further triggered discussions on Artificial Superintelligence (ASI)-a system surpassing all humans across measured domains. This gives rise to the critical research question of: As we approach ASI, how do we
RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs
cs.CLZhongzhan Huang, Guoming Ling, Yupei Lin, Yandong Chen
Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable router can significantly enhance the performance of this paradig
Bounding the Effect of Persuasion with Monotonicity Assumptions: Reassessing the Impact of TV Debates
econ.EMSung Jae Jun, Sokbae Lee
Televised debates between presidential candidates are often regarded as the exemplar of persuasive communication. Yet, recent evidence from Le Pennec and Pons (2023) indicates that they may not sway voters as strongly as popular belief suggests. We revisit their findings through the lens of the persuasion rate and introduce a robust framework that does not r
Optimizing Generative AI's Accuracy and Transparency in Inductive Thematic Analysis: A Human-AI Comparison
cs.HCMatthew Nyaaba, Min SungEun, Mary Abiswin Apam, Kwame Owoahene Acheampong
This study highlights the transparency and accuracy of GenAI's inductive thematic analysis, particularly using GPT-4 Turbo API integrated within a stepwise prompt-based Python script. This approach ensured a traceable and systematic coding process, generating codes with supporting statements and page references, which enhanced validation and reproducibility.
Avimita Chatterjee, Archisman Ghosh, Swaroop Ghosh
Quantum Error Correction (QEC), combined with magic state distillation, ensures fault tolerance in large-scale quantum computation. To apply QEC, a circuit must first be transformed into a non-Clifford (or T) gate set. T-depth, the number of sequential T-gate layers, determines the magic state cost, impacting both spatial and temporal overhead. Minimizing T-