March 2025 arXiv papers — page 44
Showing 4,301–4,400 of 23,633 papers
Jiabin Wang, Yunfei Xue, Li Chen, Ren Zhang
We show the possibility of simulating a dual universe in a pseudospin-1/2 Bose-Einstein condensate (BEC), wherein two phononic modes experience distinct curved spacetime metrics. Through ramping the interspecies interaction of the BEC, we observe that one universe expands, and in the mean time, the other contracts. These findings can be directly verified in
Emergent properties and the multiscale characterization challenge in condensed matter, from crystals to complex materials: a Review
cond-mat.mtrl-sciElisabetta Nocerino
The complexity of condensed matter arises from emergent behaviors that cannot be understood by analyzing individual constituents in isolation. While traditional condensed-matter approaches-developed primarily for ideal crystalline solids-have provided deep insights into symmetry, order, and electronic structure, they fall short in describing the rich, multis
Yide Di, Yun Liao, Hao Zhou, Kaijun Zhu
Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching pre-trained model (UFM) designed to address feature matching challenges across a wide spectrum of modal images. We present Mu
Fixseeker: An Empirical Driven Graph-based Approach for Detecting Silent Vulnerability Fixes in Open Source Software
cs.SEYiran Cheng, Ting Zhang, Lwin Khin Shar, Zhe Lang
Open source software vulnerabilities pose significant security risks to downstream applications. While vulnerability databases provide valuable information for mitigation, many security patches are released silently in new commits of OSS repositories without explicit indications of their security impact. This makes it challenging for software maintainers and
Revisit Time Series Classification Benchmark: The Impact of Temporal Information for Classification
cs.LGYunrui Zhang, Gustavo Batista, Salil S. Kanhere
Time series classification is usually regarded as a distinct task from tabular data classification due to the importance of temporal information. However, in this paper, by performing permutation tests that disrupt temporal information on the UCR time series classification archive, the most widely used benchmark for time series classification, we identify a
Zhihan Jiang, Junjie Huang, Zhuangbin Chen, Yichen Li
As Large Language Models (LLMs) show their capabilities across various applications, training customized LLMs has become essential for modern enterprises. However, due to the complexity of LLM training, which requires massive computational resources and extensive training time, failures are inevitable during the training process. These failures result in con
From the CDC to emerging infectious disease publics: The long-now of polarizing and complex health crises
cs.HCTawfiq Ammari, Anna Gutowska, Jacob Ziff, Casey Randazzo
This study examines how public discourse around COVID-19 unfolded on Twitter through the lens of crisis communication and digital publics. Analyzing over 275,000 tweets involving the CDC, we identify 16 distinct discourse clusters shaped by framing, sentiment, credibility, and network dynamics. We find that CDC messaging became a flashpoint for affective and
Himanshu Tiwari
The rapid increase in cybersecurity vulnerabilities necessitates automated tools for analyzing and classifying vulnerability reports. This paper presents a novel Vulnerability Report Classifier that leverages the BERT (Bidirectional Encoder Representations from Transformers) model to perform multi-label classification of Common Vulnerabilities and Exposures
Ayumi Igarashi, Frédéric Meunier
We study the problem of fairly allocating indivisible goods and chores under category constraints. Specifically, there are $n$ agents and $m$ indivisible items which are partitioned into categories with associated capacities. An allocation is considered feasible if each bundle satisfies the capacity constraints of its respective categories. For the case of t
Runbo Li
The author proves variants of Buchstab's identity on sieve functions, refining the previous work on new iteration rules of Brady. The main tool used in the proof is a special form of combinatorial identities related to the binomial coefficients. As a by--product, the author obtains better inequalities of $F_{\kappa}(s)$ and $f_{\kappa}(s)$ for dimensions $\k
Sarthak Raj, S. Sivananthan
In this article, we consider a variation of the existence of Gabor frames in a probabilistic setting, in which we consider time-frequency shifts taken over random-periodic sets. We demonstrate that the method of selecting random-periodic time-frequency shifts is successful with high probability for specific categories of well-behaved functions, notably inclu
A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
cs.CLAlejandro Lozano, Min Woo Sun, James Burgess, Jeffrey J. Nirschl
Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential. To address this gap, we introduce Biomedica, an open-source dataset derived from the PubMed Central Open Access subset, containing over 6 m
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
cs.CVJiaheng Zhou, Yanfeng Zhou, Wei Fang, Yuxing Tang
Ultrasound videos are an important form of clinical imaging data, and deep learning-based automated analysis can improve diagnostic accuracy and clinical efficiency. However, the scarcity of labeled data and the inherent challenges of video analysis have impeded the advancement of related methods. In this work, we introduce E-ViM$^3$, a data-efficient Vision
Xuying Li, Zhuo Li, Yuji Kosuga, Victor Bian
Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conflict. Reinforcement Learning from Human Feedback (RLHF) has achieved notable success in steering models, but is complex and can be unstable. Recent approaches such as Direct Prefer
Muhammed Shafi K. P., Serena Nicolazzo, Antonino Nocera, Vinod P
As Machine Learning (ML) evolves, the complexity and sophistication of security threats against this paradigm continue to grow as well, threatening data privacy and model integrity. In response, Machine Unlearning (MU) is a recent technology that aims to remove the influence of specific data from a trained model, enabling compliance with privacy regulations
Multi-dimensional anticipated backward stochastic differential equations with quadratic growth
math.PRYing Hu, Feng Li, Jiaqiang Wen
This paper is devoted to the general solvability of anticipated backward stochastic differential equations with quadratic growth by relaxing the assumptions made by Hu, Li, and Wen \cite[Journal of Differential Equations, 270 (2021), 1298--1311]{hu2021anticipated} from the one-dimensional case with bounded terminal values to the multi-dimensional situation w
MedSegNet10: A Publicly Accessible Network Repository for Split Federated Medical Image Segmentation
eess.IVChamani Shiranthika, Zahra Hafezi Kafshgari, Hadi Hadizadeh, Parvaneh Saeedi
Machine Learning (ML) and Deep Learning (DL) have shown significant promise in healthcare, particularly in medical image segmentation, which is crucial for accurate disease diagnosis and treatment planning. Despite their potential, challenges such as data privacy concerns, limited annotated data, and inadequate training data persist. Decentralized learning a
Te Li, Ping Zhang, Yibin Zhang
In this paper, we consider the asymptotic stability of the 2D Taylor-Couette flow in the exterior disk, with a small kinematic viscosity $\nu \ll 1$ and a large rotation coefficient $|B|$. Due to the degeneracy of the Taylor-Couette flow at infinity, we cannot expect the solution to decay exponentially in a space-time decoupled manner. As stated in previous
Yejin Kwon, Daeun Moon, Youngje Oh, Hyunsoo Yoon
Anomaly Detection (AD) focuses on detecting samples that differ from the standard pattern, making it a vital tool in process control. Logical anomalies may appear visually normal yet violate predefined constraints on object presence, arrangement, or quantity, depending on reasoning and explainability. We introduce LogicQA, a framework that enhances AD by pro
Particle Deflections around Microscopic Loop Quantum Black Holes with Rigorous Quantum Parameters
gr-qcHaida Li, Xiangdong Zhang
The detection of quantum gravity effects is highly limited in both macroscopic and microscopic scenarios: The small quantum parameter makes most large-scale observations practically indistinguishable from general relativity. While at the Planck scale, where the effect of quantum gravity is undoubtedly significant, the energy requirement is remarkably high fo
Junoh Heo
Dynamic simulators are computational models governed by differential equations that evolve over time. They are essential for scientific and engineering applications but remain challenging to emulate because of the unpredictable behavior of complex systems. To address this challenge, this paper introduces a fast and accurate Gaussian Process (GP)-based emulat
Mingfu Liang, Jiahuan Zhou, Xu Zou, Ying Wu
Existing progress in object keypoint estimation primarily benefits from the conventional supervised learning paradigm based on numerous data labeled with pre-defined keypoints. However, these well-trained models can hardly detect the undefined new keypoints in test time, which largely hinders their feasibility for diverse downstream tasks. To handle this, va
Utkarsh Azad, Bikash K. Behera, Houbing Song, Ahmed Farouk
Industry 5.0 depends on intelligence, automation, and hyperconnectivity operations for effective and sustainable human-machine collaboration. Pivotal technologies like the Internet of Things (IoT) enable this by facilitating connectivity and data-driven decision-making between cyber-physical devices. As IoT devices are prone to cyberattacks, they can use blo
VESTA: A Versatile SNN-Based Transformer Accelerator with Unified PEs for Multiple Computational Layers
cs.ARChing-Yao Chen, Meng-Chieh Chen, Tian-Sheuan Chang
Spiking Neural Networks (SNNs) and transformers represent two powerful paradigms in neural computation, known for their low power consumption and ability to capture feature dependencies, respectively. However, transformer architectures typically involve multiple types of computational layers, including linear layers for MLP modules and classification heads,
Chih-Chia Hsu, Tian-Sheuan Chang
Deep learning-based super-resolution (SR) is challenging to implement in resource-constrained edge devices for resolutions beyond full HD due to its high computational complexity and memory bandwidth requirements. This paper introduces an 8K@30FPS SR accelerator with edge-selective dynamic input processing. Dynamic processing chooses the appropriate subnets
Enhancing Finite State Machine Design Automation with Large Language Models and Prompt Engineering Techniques
cs.ARQun-Kai Lin, Cheng Hsu, Tian-Sheuan Chang
Large Language Models (LLMs) have attracted considerable attention in recent years due to their remarkable compatibility with Hardware Description Language (HDL) design. In this paper, we examine the performance of three major LLMs, Claude 3 Opus, ChatGPT-4, and ChatGPT-4o, in designing finite state machines (FSMs). By utilizing the instructional content pro
Matthew B. Andorf, Mei Bai, Pushpalatha Bhat, Valery Borzenets
The Linear Collider Vision calls for a Linear Collider Facility with a physics reach from a Higgs Factory to the TeV-scale with $e^+e^{-}$ collisions. One of the technologies under consideration for the accelerator is a cold-copper distributed-coupling linac capable of achieving high gradient. This technology is being pursued by the C$^3$ collaboration to un
Software Vulnerability Analysis Across Programming Language and Program Representation Landscapes: A Survey
cs.CRZhuoyun Qian, Fangtian Zhong, Qin Hu, Yili Jiang
Modern software systems are developed in diverse programming languages and often harbor critical vulnerabilities that attackers can exploit to compromise security. These vulnerabilities have been actively targeted in real-world attacks, causing substantial harm to users and cyberinfrastructure. Since many of these flaws originate from the code itself, a vari
Structural and transport properties of LiTFSI/G3 electrolyte with machine-learned molecular dynamics
cond-mat.mtrl-sciChenyang Cao, Liyi Bai, Shuo Cao, Ye Su
The lithium bis(trifluoromethylsulfonyl)azanide-triglyme electrolyte plays a critical role in the performance of lithium-ion batteries. However, its solvation structure and transport properties at the atomic scale remain incompletely understood. In this study, we develop an efficient and accurate neuroevolution potential (NEP) model by integrating bootstrap
Investigating cross-correlations between cosmic microwave background lensing and 21 cm intensity mapping
astro-ph.COAlessandro Marins, Chang Feng, Filipe B. Abdalla
The neutral hydrogen (HI) signal is a crucial probe for astrophysics and cosmology, but it is quite challenging to measure from raw data because of bright foreground contaminants at radio wavelengths. Cross-correlating the radio observations with large-scale structure tracers (LSS) could detect faint cosmological signals since they are not correlated with th
Mitsuaki Uno, Kanji Tanaka, Daiki Iwata, Yudai Noda
Object Goal Navigation (OGN) is a fundamental task for robots and AI, with key applications such as mobile robot image databases (MRID). In particular, mapless OGN is essential in scenarios involving unknown or dynamic environments. This study aims to enhance recent modular mapless OGN systems by leveraging the commonsense reasoning capabilities of large lan
Prin Phunyaphibarn, Phillip Y. Lee, Jaihoon Kim, Minhyuk Sung
Classifier-Free Guidance (CFG) is a fundamental technique in training conditional diffusion models. The common practice for CFG-based training is to use a single network to learn both conditional and unconditional noise prediction, with a small dropout rate for conditioning. However, we observe that the joint learning of unconditional noise with limited band
Ayman El Zein, Maidoun Mortada
For a non-decreasing sequence of integers $S=(s_1,s_2, \dots, s_k)$, an $S$-packing coloring of $G$ is a partition of $V(G)$ into $k$ subsets $V_1,V_2,\dots,V_k$ such that the distance between any two distinct vertices $x,y \in V_i$ is at least $s_{i}+1$, $1\leq i\leq k$. The packing chromatic number $\chi_{\rho}(G)$ of a graph $G$ is the smallest integer $p
Sandeep Joshi, Garima Rajpoot, Prashant Shukla
Quantum simulation of particle phenomena is a rapidly advancing field of research. With the widespread availability of quantum simulators, a given quantum system can be simulated in numerous ways, offering flexibility in implementation and exploration. Here, we perform quantum simulation of neutrino propagation in matter, a phenomenon that plays a crucial ro
Vineela Reddy Pippera Badguna, Aliasghar Arab, Durga Avinash Kodavalla
Collaborative robots (cobots) increasingly operate alongside humans, demanding robust real-time safeguarding. Current safety standards (e.g., ISO 10218, ANSI/RIA 15.06, ISO/TS 15066) require risk assessments but offer limited guidance for real-time responses. We propose a virtual fencing approach that detects and predicts human motion, ensuring safe cobot op
Bernhard König, Yasuo Yoshinobu
We introduce two types of variations of setwise climbability properties, which have been introduced by the second author as fragments of Jensen's square principles. We show that variations of the first type are equivalent to known principles and that they are consistent with the Proper Forcing Axiom(PFA). On the other hand, those of the second type can be ch
Ahyun Seo, Minsu Cho
Symmetry plays a vital role in understanding structural patterns, aiding object recognition and scene interpretation. This paper focuses on rotation symmetry, where objects remain unchanged when rotated around a central axis, requiring detection of rotation centers and supporting vertices. Traditional methods relied on hand-crafted feature matching, while re
Yitian Chen, Timothy L. Molloy, Iman Shames
We investigate a novel finite-horizon linear-quadratic (LQ) feedback dynamic potential game with a priori unknown cost matrices played between two players. The cost matrices are revealed to the players sequentially, with the potential for future values to be previewed over a short time window. We propose an algorithm that enables the players to predict and t
Dynamic Learning and Productivity for Data Analysts: A Bayesian Hidden Markov Model Perspective
cs.SIYue Yin
Data analysts are essential in organizations, transforming raw data into insights that drive decision-making and strategy. This study explores how analysts' productivity evolves on a collaborative platform, focusing on two key learning activities: writing queries and viewing peer queries. While traditional research often assumes static models, where performa
Wei Wang, Yujie Lin, Jianli Zhao, Moyan Zhang
Most existing contrastive learning-based sequential recommendation (SR) methods rely on random operations (e.g., crop, reorder, and substitute) to generate augmented sequences. These methods often struggle to create positive sample pairs that closely resemble the representations of the raw sequences, potentially disrupting item correlations by deleting key i
Fabian Baumann, Nipun Arora, Iyad Rahwan, Agnieszka Czaplicka
Intelligent algorithms increasingly shape the content we encounter and engage with online. TikTok's For You feed exemplifies extreme algorithm-driven curation, tailoring the stream of video content almost exclusively based on users' explicit and implicit interactions with the platform. Despite growing attention, the dynamics of content amplification on TikTo
Ugochukwu Ejike Akpudo, Yongsheng Gao, Jun Zhou, Andrew Lewis
Convolutional neural networks (CNNs) have succeeded remarkably in various computer vision tasks. However, they are not intrinsically explainable. While the feature-level understanding of CNNs reveals where the models looked, concept-based explainability methods provide insights into what the models saw. However, their assumption of linear reconstructability
Automated UI Interface Generation via Diffusion Models: Enhancing Personalization and Efficiency
cs.HCYifei Duan, Liuqingqing Yang, Tong Zhang, Zhijun Song
This study proposes a UI interface generation method based on a diffusion model, aiming to achieve high-quality, diversified, and personalized interface design through generative artificial intelligence technology. The diffusion model is based on its step-by-step denoising generation process. By combining the conditional generation mechanism, design optimiza
InfoBid: A Simulation Framework for Studying Information Disclosure in Auctions with Large Language Model-based Agents
cs.GTYue Yin
In online advertising systems, publishers often face a trade-off in information disclosure strategies: while disclosing more information can enhance efficiency by enabling optimal allocation of ad impressions, it may lose revenue potential by decreasing uncertainty among competing advertisers. Similar to other challenges in market design, understanding this
Xiao Lin, Manoj Acharya, Anirban Roy, Susmit Jha
Mitigating Trojans in Large Language Models (LLMs) is one of many tasks where alignment data is LLM specific, as different LLMs have different Trojan triggers and trigger behaviors to be removed. In this paper, we introduce TeleLoRA (Teleporting Low-Rank Adaptation), a novel framework that synergizes model-specific alignment data across multiple LLMs to enab
Advancements in Natural Language Processing: Exploring Transformer-Based Architectures for Text Understanding
cs.CLTianhao Wu, Yu Wang, Ngoc Quach
Natural Language Processing (NLP) has witnessed a transformative leap with the advent of transformer-based architectures, which have significantly enhanced the ability of machines to understand and generate human-like text. This paper explores the advancements in transformer models, such as BERT and GPT, focusing on their superior performance in text underst
Ying Ma, Shiquan Zhang, Dongju Yang, Zhanna Sarsenbayeva
Location privacy leaks can lead to unauthorised tracking, identity theft, and targeted attacks, compromising personal security and privacy. This study explores LLM-powered location privacy leaks associated with photo sharing on social media, focusing on user awareness, attitudes, and opinions. We developed and introduced an LLM-powered location privacy inter
HyungJoo Kim, Philipp Gubler, Chihiro Sasaki
We present a novel approach for investigating the spin structure of hadrons based on the two-point function in quantum field theory. In a rotating frame, we derive two independent expressions of the two-point function and identify their equivalence, which allows for a complete decomposition of the total spin of a composite system into the angular momenta of
Zhenhua Yuan, Junhao Peng, Long Gao
Transport is an important function of networks. Studying transport efficiency sheds light on the dynamic processes occurring within various underlying structures and offers a wide range of applications. To construct networks with different transport efficiencies, we focus on the networks obtained by vertex merging operation, which involves connecting multipl
Jordan Hong, Safwan Jamal, Ashish Khisti
Artificial noise (AN) transmission is a physical layer security technique in multi-antenna wireless communication systems. Synthetic noise is broadcast to all receivers except designated legitimate users via beamforming in the legitimate users' null space. We consider AN transmission employing a single RF chain and analog beamforming, where beamforming vecto
Solving 2-D Helmholtz equation in the rectangular, circular, and elliptical domains using neural networks
cs.LGD. Veerababu, Prasanta K. Ghosh
Physics-informed neural networks offered an alternate way to solve several differential equations that govern complicated physics. However, their success in predicting the acoustic field is limited by the vanishing-gradient problem that occurs when solving the Helmholtz equation. In this paper, a formulation is presented that addresses this difficulty. The p
Taorui Wang, Zitong Yu, Yong Xu
Recently, 3D Gaussian Splatting (3DGS) has emerged as a prominent framework for novel view synthesis, providing high fidelity and rapid rendering speed. However, the substantial data volume of 3DGS and its attributes impede its practical utility, requiring compression techniques for reducing memory cost. Nevertheless, the unorganized shape of 3DGS leads to d
Weijie Guo, Guofeng Zhang, Wufei Ma, Alan Yuille
Category-level 3D/6D pose estimation is a crucial step towards comprehensive 3D scene understanding, which would enable a broad range of applications in robotics and embodied AI. Recent works explored neural mesh models that approach a range of 2D and 3D tasks from an analysis-by-synthesis perspective. Despite the largely enhanced robustness to partial occlu
Tianqi Tu, Hui Wang, Jiangbo Pei, Xiaojuan Yu
Background: Renal chronicity indices (CI) have been identified as strong predictors of long-term outcomes in lupus nephritis (LN) patients. However, assessment by pathologists is hindered by challenges such as substantial time requirements, high interobserver variation, and susceptibility to fatigue. This study aims to develop an effective deep learning (DL)
Peidong Yu, Tao Lin, Ziyan Deng, Guofu Cao
The Jiangmen Underground Neutrino Observatory (JUNO) is a multi-purpose experiment under construction in southern China. JUNO aims to determine the neutrino mass ordering and precisely measure the neutrino oscillation parameters by detecting reactor neutrinos from nuclear power plants. In addition to reactor neutrinos, JUNO can study atmospheric neutrinos, s
Haiyang Liu, Zhan Xu, Fa-Ting Hong, Hsin-Ping Huang
We present Video Motion Graphs, a system designed to generate realistic human motion videos. Using a reference video and conditional signals such as music or motion tags, the system synthesizes new videos by first retrieving video clips with gestures matching the conditions and then generating interpolation frames to seamlessly connect clip boundaries. The c
Transitions of the Lyapunov spectrum and entanglement entropy in monitored quantum dynamics with homogeneous unitary gates
quant-phKen Mochizuki, Ryusuke Hamazaki
We explore the Lyapunov spectrum and entanglement entropy in systems evolved by quantum measurements and spatially homogeneous unitary gates. In models with temporally random and Floquet unitary gates, we find that the Lyapunov exponents typically converge to values independent of measurement outcomes and that spectral transitions of the Lyapunov spectrum an
Cheng Zhang, Fu-Le Hao, Shi-Pu Gu, Xing-Fu Wang
Quantum secret sharing (QSS) is a typical multipartite cryptographic primitive, which is an important part of quantum communication network. Existing QSS protocols generally require basis selection and matching, which would increase the quantum resource consumption and classical communication round, and also face weak random security vulnerabilities. We prop
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu
In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. To enable the streaming of multimodal information inputs, both audio and visual encoders utilize a block-wise proces
Spencer Gessner, Jens Osterhoff, Carl A. Lindstrøm, Kevin Cassou
This document outlines a community-driven Design Study for a 10 TeV pCM Wakefield Accelerator Collider. The 2020 ESPP Report emphasized the need for Advanced Accelerator R\&D, and the 2023 P5 Report calls for the ``delivery of an end-to-end design concept, including cost scales, with self-consistent parameters throughout." This Design Study leverages recent
Nonreciprocal spin-charge interconversion in topological insulator/ferromagnet heterostructures
cond-mat.mtrl-sciShuyuan Shi, Enlong Liu, Fanrui Hu, Guoyi Shi
The process of spin-charge interconversion is critical in modern spintronics. Nonetheless, experiments conducted on a wide variety of magnetic heterostructures consistently report that charge-to-spin and spin-to-charge conversion efficiencies can be vastly different, especially in the case of topological insulators (TI). This discrepancy between the two "rec
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
cs.CVWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao Yan
Transformer-based models have driven significant advancements in Multimodal Large Language Models (MLLMs), yet their computational costs surge drastically when scaling resolution, training data, and model parameters. A key bottleneck stems from the proliferation of visual tokens required for fine-grained image understanding. We propose Skip-Vision, a unified
Jinxu Lin, Linwei Tao, Minjing Dong, Chang Xu
Model calibration is essential for ensuring that the predictions of deep neural networks accurately reflect true probabilities in real-world classification tasks. However, deep networks often produce over-confident or under-confident predictions, leading to miscalibration. Various methods have been proposed to address this issue by designing effective loss f
A Spatial-temporal Deep Probabilistic Diffusion Model for Reliable Hail Nowcasting with Radar Echo Extrapolation
cs.LGHaonan Shi, Long Tian, Jie Tao, Yufei Li
Hail nowcasting is a considerable contributor to meteorological disasters and there is a great need to mitigate its socioeconomic effects through precise forecast that has high resolution, long lead times and local details with large landscapes. Existing medium-range weather forecasting methods primarily rely on changes in upper air currents and cloud layers
Yangyang Meng, Jinpeng Li, Guodong Lin, Yu Pu
This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source datasets to refine and optimize Dolphin's performance. The model is specifically designed to achieve notable recognition a
Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors
cs.CVWeilong Yan, Ming Li, Haipeng Li, Shuwei Shao
Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data. Previous methods either rely on synthetic inputs and pseudo-depth labels or directly apply daytime strat
Akib Karim, Shaobo Zhang, Muhammad Usman
The Variational Quantum Eigensolver (VQE) is a hybrid quantum-classical algorithm for preparing ground states in the current era of noisy devices. The classical component of the algorithm requires a large number of measurements on intermediate parameter values that are typically discarded. However, intermediate steps across many calculations can contain valu
BEAR: A Video Dataset For Fine-grained Behaviors Recognition Oriented with Action and Environment Factors
cs.CVChengyang Hu, Yuduo Chen, Lizhuang Ma
Behavior recognition is an important task in video representation learning. An essential aspect pertains to effective feature learning conducive to behavior recognition. Recently, researchers have started to study fine-grained behavior recognition, which provides similar behaviors and encourages the model to concern with more details of behaviors with effect
Liangzhi Shi, Yulin Liu, Lingqi Zeng, Bo Ai
How can robots learn dexterous grasping skills efficiently and apply them adaptively based on user instructions? This work tackles two key challenges: efficient skill acquisition from limited human demonstrations and context-driven skill selection. We introduce AdaDexGrasp, a framework that learns a library of grasping skills from a single human demonstratio
Reasoning and Learning a Perceptual Metric for Self-Training of Reflective Objects in Bin-Picking with a Low-cost Camera
cs.CVPeiyuan Ni, Chee Meng Chew, Marcelo H. Ang, Gregory S. Chirikjian
Bin-picking of metal objects using low-cost RGB-D cameras often suffers from sparse depth information and reflective surface textures, leading to errors and the need for manual labeling. To reduce human intervention, we propose a two-stage framework consisting of a metric learning stage and a self-training stage. Specifically, to automatically process data c
Manh Mai Van, Tin T. Tran
The trend of data mining using deep learning models on graph neural networks has proven effective in identifying object features through signal encoders and decoders, particularly in recommendation systems utilizing collaborative filtering methods. Collaborative filtering exploits similarities between users and items from historical data. However, it overloo
Xiao-Cheng Liao, Yi Mei, Mengjie Zhang, Xiang-Ling Chen
Appropriate traffic state representation is crucial for learning traffic signal control policies. However, most of the current traffic state representations are heuristically designed, with insufficient theoretical support. In this paper, we (1) develop a flexible, efficient, and theoretically grounded method, namely generalized phase pressure (G2P) control,
Energy transfer and budget analysis for transient process with phase-averaged reduced-order model
physics.flu-dynYuto Nakamura, Yuma Kuroda, Shintaro Sato, Naofumi Ohnishi
We derive a phase-averaged representation of transient flows based on the eigenmodes of a data-driven linear operator that approximates the Navier-Stokes dynamics. In performing phase averaging, it is assumed that, at each instant during the transient evolution, the eigenmode amplitude remains invariant, while only the complex phase angle differs among disti
Erik J. Gustafson, Henry Lamm, Diyi Liu, Edison M. Murairi
We present two deterministic algorithms to approximate single-qutrit gates. These algorithms utilize the Clifford + $\mathbf{R}$ group to find the best approximation of diagonal rotations. The first algorithm exhaustively searches over the group; while the second algorithm searches only for Householder reflections. The exhaustive search algorithm yields an a
Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao
Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGes, a novel framework that leverages large language models (LLMs) to parse speech content and generate reliable semantic gesture labels, whic
Salaheddin Alzubi, Creston Brooks, Purva Chiniya, Edoardo Contente
We introduce Open Deep Search (ODS) to close the increasing gap between the proprietary search AI solutions, such as Perplexity's Sonar Reasoning Pro and OpenAI's GPT-4o Search Preview, and their open-source counterparts. The main innovation introduced in ODS is to augment the reasoning capabilities of the latest open-source LLMs with reasoning agents that c
G. A. Valdeon Sauza
This result generalizes a previous result established in [2] where the Absolute Constants of the Koras-Russell threefold was shown to be invariant under translates in the base field to the Absolute Constants of the Koras-Russell threefold being invariant under translates of any of its absolute invariants. It must be pointed out that there is no prior reason
Mélisande Teng, Arthur Ouaknine, Etienne Laliberté, Yoshua Bengio
The potential of tree planting as a natural climate solution is often undermined by inadequate monitoring of tree planting projects. Current monitoring methods involve measuring trees by hand for each species, requiring extensive cost, time, and labour. Advances in drone remote sensing and computer vision offer great potential for mapping and characterizing
Alex Jinpeng Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang
Recent advancements in autoregressive and diffusion models have led to strong performance in image generation with short scene text words. However, generating coherent, long-form text in images, such as paragraphs in slides or documents, remains a major challenge for current generative models. We present the first work specifically focused on long text image
Zike Li, Mingwei Liu, Anji Li, Kaifeng He
Robustness is a critical factor for reliable code generation by large language models, yet most evaluations focus on correctness and overlook key issues such as missing input validation and inadequate error handling. In this work, we present the first empirical study of LLM-generated code robustness using the CoderEval benchmark. Evaluating four state-of-the
Beomseok Oh, Dongwoo Lee, Yeon-Seong Choo, Sung-Hoon Byun
Broadband underwater sound focusing in the low-frequency range is essential for various applications such as battery-free environmental monitoring and sensing. However, achieving low-frequency underwater focusing typically necessitates bulky, heavy structures that hinder practical deployment. Here, we introduce a three-dimensional underwater lens comprising
Mutual Information-Empowered Task-Oriented Communication: Principles, Applications and Challenges
cs.ITHongru Li, Songjie Xie, Jiawei Shao, Zixin Wang
Mutual information (MI)-based guidelines have recently proven to be effective for designing task-oriented communication systems, where the ultimate goal is to extract and transmit task-relevant information for downstream task. This paper provides a comprehensive overview of MI-empowered task-oriented communication, highlighting how MI-based methods can serve
Zhouhong Gu, Xingzhou Chen, Xiaoran Shi, Tao Wang
Recent advances in large language models have highlighted the critical need for precise control over model outputs through predefined constraints. While existing methods attempt to achieve this through either direct instruction-response synthesis or preferential response optimization, they often struggle with constraint understanding and adaptation. This lim
Yury Polyanskiy, Mark Sellke
We study the nonparametric maximum likelihood estimator $\widehat{\pi}$ for Gaussian location mixtures in one dimension. It has been known since (Lindsay, 1983) that given an $n$-point dataset, this estimator always returns a mixture with at most $n$ components, and more recently (Wu-Polyanskiy, 2020) gave a sharp $O(\log n)$ bound for subgaussian data. In t
A note on the asymptotics of the free energy of $1+1$ dimensional directed polymers in random environment at high temperature
math.PRMakoto Nakashima
The author gave the sharp asymptotic behavior of the free energy of $1+1$ dimensional directed polymers in random environment(DPRE) as the inverse temperature $\beta\to 0$ under the assumption that random environment satisfies a certain concentration inequality in [Nak19], \[\lim_{\beta\to0}\frac{1}{\beta^4}F(\beta)=-\frac{1}{6}. \] In this paper, we obtain
Srihas Yarlagadda, Amey Agrawal, Elton Pinto, Hakesh Darapaneni
Training large foundation models costs hundreds of millions of dollars, making deployment optimization critical. Current approaches require machine learning engineers to manually craft training recipes through error-prone trial-and-error on expensive compute clusters. To enable efficient exploration of training configurations, researchers have developed perf
Cross-Modal Prototype Allocation: Unsupervised Slide Representation Learning via Patch-Text Contrast in Computational Pathology
cs.CVYuxuan Chen, Jiawen Li, Jiali Hu, Xitong Ling
With the rapid advancement of pathology foundation models (FMs), the representation learning of whole slide images (WSIs) attracts increasing attention. Existing studies develop high-quality patch feature extractors and employ carefully designed aggregation schemes to derive slide-level representations. However, mainstream weakly supervised slide representat
Unifying and accelerating level-set and density-based topology optimization by subpixel-smoothed projection
physics.opticsAlec M. Hammond, Ardavan Oskooi, Ian M. Hammond, Mo Chen
We introduce a new "subpixel-smoothed projection" (SSP) formulation for differentiable binarization in topology optimization (TopOpt) as a drop-in replacement for previous projection schemes, which suffer from near-non-differentiability and slow convergence as binarization improves. Our new algorithm overcomes these limitations by depending on both the under
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
cs.CVXiao Guo, Xiufeng Song, Yue Zhang, Xiaohong Liu
Deepfake detection is a long-established research topic vital for mitigating the spread of malicious misinformation. Unlike prior methods that provide either binary classification results or textual explanations separately, we introduce a novel method capable of generating both simultaneously. Our method harnesses the multi-modal learning capability of the p
Pirzada Suhail, Pravesh Khaparde, Amit Sethi
In vision classification, generating inputs that elicit confident predictions is key to understanding model behavior and reliability, especially under adversarial or out-of-distribution (OOD) conditions. While traditional adversarial methods rely on perturbing existing inputs to fool a model, they are inherently input-dependent and often fail to ensure both
Vahid Mosallanejad, Wenjie Dou
Simultaneous driving by two periodic oscillations yields a practical technique for further engineering quantum systems. For quantum transport through mesoscopic systems driven by two strong periodic terms, a non-perturbative Floquet-based quantum master equation (QME) approach is developed using a set of dissipative time-dependent terms and the reduced densi
Hiroo Azuma, William J. Munro, Kae Nemoto
We explore quantum phase transitions in the multiphoton Jaynes-Cummings-Hubbard model (JCHM). Using the mean-field approximation, we demonstrate that the multiphoton JCHM exhibits quantum phase transitions between the Mott insulator (MI) phase, the superfluid phase, and an additional phase we refer to as the forbidden phase. The multiphoton JCHM MI phases ar
M. Kerem Aydin, Yi-Chun Hung, Jaclyn Pytlarz, Qi Guo
Hyperspectral cameras face harsh trade-offs between spatial, spectral, and temporal resolution in inherently low-photon conditions. Computational imaging systems break through these trade-offs with compressive sensing, but have required complex optics and/or extensive compute. We present Spectrum from Defocus (SfD), a chromatic focal sweep method that achiev
Mohammad Saif Nazir, Chayan Banerjee
Reinforcement learning (RL) often struggles with reward misalignment, where agents optimize given rewards but fail to exhibit the desired behaviors. This arises when the reward function incentivizes proxy behaviors misaligned with the true objective. While human-in-the-loop (HITL) methods can mitigate this issue, they also introduce biases, leading to incons
Hao-Dong Cai, Zhao-Sai Jia, Gang Li, Shi-Dong Liu
The dipionic transition of the $X(3872)$ to the $\eta_c$ was investigated using an effective Lagrangian approach. In this study, the $X(3872)$ was assumed to be a $D\bar{D}^\ast + \text{c.c.}$ bound state with the quantum numbers $J^{PC}=1^{++}$ and to decay via triangle and box loops. It is found that the partial decay widths arising from the box loops are
Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie
As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding their behavioral characteristics becomes essential for responsible AI development. However, existing evaluation efforts, which often adapt human psychological assessments such as the Big Five Inventory (BFI), face two signifi
Zhen-Xuan Yang, Hidefumi Matsuda, Xu-Guang Huang, Kouji Kashiwa
We study the Hamiltonian formulation of SU(2) Yang-Mills theory with staggered fermions in a (2+1)-dimensional small lattice system. We construct a gauge-invariant and finite-dimensional Hilbert space for the theory by applying the loop-string-hadron formulation and specifically map the model to a spin system. We classically emulate digital quantum simulatio
Mehdi Eddaoudi
Payne-P\'olya-Weinberger inequalities are known to be exclusive to bounded Euclidean domains with Dirichlet boundary condition. In this paper, we discuss the corresponding inequalities on Riemannian manifolds of dimension $n \geq3$, and we prove explicit bounds in terms of geometric quantities such as scalar curvature, Yamabe constant, isoperimetric constant
Hiroyuki Uchida, Koji Mori, Hiroshi Tomida, Hiroshi Nakajima
We present a summary of the in-orbit performance of the soft X-ray imaging telescope Xtend onboard the XRISM mission, based on in-flight observation data, including first-light celestial objects, calibration sources, and results from the cross-calibration campaign with other currently-operating X-ray observatories. XRISM/Xtend has a large field of view of $3
ProtoBERT-LoRA: Parameter-Efficient Prototypical Finetuning for Immunotherapy Study Identification
cs.CLShijia Zhang, Xiyu Ding, Kai Ding, Jacob Zhang
Identifying immune checkpoint inhibitor (ICI) studies in genomic repositories like Gene Expression Omnibus (GEO) is vital for cancer research yet remains challenging due to semantic ambiguity, extreme class imbalance, and limited labeled data in low-resource settings. We present ProtoBERT-LoRA, a hybrid framework that combines PubMedBERT with prototypical ne