March 2025 arXiv papers — page 172
Showing 17,101–17,200 of 23,633 papers
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
cs.SEKaiyuan Liu, Youcheng Pan, Yang Xiang, Daojing He
Recently, LLM agents have made rapid progress in improving their programming capabilities. However, existing benchmarks lack the ability to automatically evaluate from users' perspective, and also lack the explainability of the results of LLM agents' code generation capabilities. Thus, we introduce ProjectEval, a new benchmark for LLM agents project-level co
A splitting theorem for manifolds with spectral nonnegative Ricci curvature and mean-convex boundary
math.DGHan Hong, Gaoming Wang
We prove a splitting theorem for a smooth noncompact manifold with (possibly noncompact) boundary. We show that if a noncompact manifold of dimension $n\geq 2$ has $\lambda_1(-\alpha\Delta+\operatorname{Ric})\geq 0$ for some $\alpha<\frac{4}{n-1}$ and mean-convex boundary, then it is either isometric to $\Sigma\times \mathbb{R}_{\geq 0}$ for a closed manifol
SDFA: Structure Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in Video
cs.CVSania Zahan, Ghulam Mubashar Hassan, Ajmal Mian
Older people are susceptible to fall due to instability in posture and deteriorating health. Immediate access to medical support can greatly reduce repercussions. Hence, there is an increasing interest in automated fall detection, often incorporated into a smart healthcare system to provide better monitoring. Existing systems focus on wearable devices which
Explicit Solution of Tunable Input-to-State Safe-Based Controller Under High-Relative-Degree Constraints
eess.SYYan Wei, Yu Feng, Linlin Ou, Yueying Wang
This paper investigates the safety analysis and verification of nonlinear systems subject to high-relative-degree constraints and unknown disturbance. The closed-form solution of the high-order control barrier functions (HOCBF) optimization problem with and without a nominal controller is first provided, making it unnecessary to solve the quadratic program p
Shuhao Liao, Xuxin Lv, Yuhong Cao, Jeric Lew
In autonomous exploration tasks, robots are required to explore and map unknown environments while efficiently planning in dynamic and uncertain conditions. Given the significant variability of environments, human operators often have specific preference requirements for exploration, such as prioritizing certain areas or optimizing for different aspects of e
Alexander Eber, Christoph Gruber, Martin Schultze, Birgitta Bernhardt
We radically simplify coherently averaged dual-comb spectroscopy by introducing a real-time self-correction system: a radio frequency system-on-chip computes each incoming dual-comb interferogram's phase, frequency, and arrival time; calculates changes in the combs' carrier-envelope offset frequency and repetition rate difference; and immediately phase-corre
Jiaojiao Li, Shiyao Duan, Haitao XU, Rui Song
The inherent difficulty in acquiring accurately co-registered RGB-hyperspectral image (HSI) pairs has significantly impeded the practical deployment of current data-driven Hyperspectral Image Generation (HIG) networks in engineering applications. Gleichzeitig, the ill-posed nature of the aligning constraints, compounded with the complexities of mining cross-
Ruoxi Xu, Hongyu Lin, Xianpei Han, Jia Zheng
As large language models (LLMs) increasingly become central to various applications and interact with diverse user populations, ensuring their reliable and consistent performance is becoming more important. This paper explores a critical issue in assessing the reliability of LLMs: the consistency between their words and deeds. To quantitatively explore this
Jiazheng Liu, Sipeng Zheng, Börje F. Karlsson, Zongqing Lu
Multimodal large language models (MLLMs), built on large-scale pre-trained vision towers and language models, have shown great capabilities in multimodal understanding. However, most existing MLLMs are trained on single-turn vision question-answering tasks, which do not accurately reflect real-world human conversations. In this paper, we introduce MMDiag, a
Stability of Khintchine inequalities with optimal constants between the second and the $p$-th moment for $p \ge 3$
math.PRJacek Jakimiuk
We give a strengthening of the classical Khintchine inequality between the second and the $p$-th moment for $p \ge 3$ with optimal constant by adding a deficit depending on the vector of coefficients of the Rademacher sum.
Frequency-Aware Density Control via Reparameterization for High-Quality Rendering of 3D Gaussian Splatting
cs.CVZhaojie Zeng, Yuesong Wang, Lili Ju, Tao Guan
By adaptively controlling the density and generating more Gaussians in regions with high-frequency information, 3D Gaussian Splatting (3DGS) can better represent scene details. From the signal processing perspective, representing details usually needs more Gaussians with relatively smaller scales. However, 3DGS currently lacks an explicit constraint linking
Chase Hutton, Adam Melrod
Many parallel algorithms which solve basic problems in computer science use auxiliary space linear in the input to facilitate conflict-free computation. There has been significant work on improving these parallel algorithms to be in-place, that is to use as little auxiliary memory as possible. In this paper, we provide novel in-place algorithms to solve the
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
cs.CVHaoyu Zheng, Qifan Yu, Binghe Yu, Yang Dai
Diffusion models have achieved remarkable progress in image and video stylization. However, most existing methods focus on single-style transfer, while video stylization involving multiple styles necessitates seamless transitions between them. We refer to this smooth style transition between video frames as video style morphing. Current approaches often gene
PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging Sparsity
cs.LGKwanyoung Kim, Byeongsu Sim
Diffusion models have shown impressive results in generating high-quality conditional samples using guidance techniques such as Classifier-Free Guidance (CFG). However, existing methods often require additional training or neural function evaluations (NFEs), making them incompatible with guidance-distilled models. Also, they rely on heuristic approaches that
Water Quality Data Imputation via A Fast Latent Factorization of Tensors with PID-based Optimizer
cs.LGQian Liu, Lan Wang, Bing Yang, Hao Wu
Water quality data can supply a substantial decision support for water resources utilization and pollution prevention. However, there are numerous missing values in water quality data due to inescapable factors like sensor failure, thereby leading to biased result for hydrological analysis and failing to support environmental governance decision accurately.
Stylianos Zindros, Christos Chronis, Panagiotis Radoglou-Grammatikis, Vasileios Argyriou
As the security of public spaces remains a critical issue in today's world, Digital Twin technologies have emerged in recent years as a promising solution for detecting and predicting potential future threats. The applied methodology leverages a Digital Twin of a metro station in Athens, Greece, using the FlexSim simulation software. The model encompasses po
Haolin Li, Yikang Chai, Bailin Lv, Lecheng Ruan
This study introduces a unified control framework that addresses the challenge of precise quadruped locomotion with unknown payloads, named as online payload identification-based physics-informed neural network predictive control (OPI-PINNPC). By integrating online payload identification with physics-informed neural networks (PINNs), our approach embeds iden
Lei Zhang, Mukesh Ghimire, Wenlong Zhang, Zhe Xu
General-sum differential games can approximate values solved by Hamilton-Jacobi-Isaacs (HJI) equations for efficient inference when information is incomplete. However, solving such games through conventional methods encounters the curse of dimensionality (CoD). Physics-informed neural networks (PINNs) offer a scalable approach to alleviate the CoD and approx
Shihao Hou, Xinyi Shang, Shreyank N Gowda, Yang Lu
Effectively handling the co-occurrence of non-IID data and long-tailed distributions remains a critical challenge in federated learning. While fine-tuning vision-language models (VLMs) like CLIP has shown to be promising in addressing non-IID data challenges, this approach leads to severe degradation of tail classes in federated long-tailed scenarios. Under
Hanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan
High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically, existing methods mainly introduce event camera to directly fuse the spatiotemporal features between the two modalities. Howeve
Yongwoo Kim, Sungmin Cha, Donghyun Kim
Machine unlearning is a process to remove specific data points from a trained model while maintaining the performance on the retain data, addressing privacy or legal requirements. Despite its importance, existing unlearning evaluations tend to focus on logit-based metrics under small-scale scenarios. We observe that this could lead to a false sense of securi
Hyeonsoo Jo, Jongha Lee, Fanchen Bu, Kijung Shin
Time-evolving graphs, such as social and citation networks, often contain noise that distorts structural and temporal patterns, adversely affecting downstream tasks, such as node classification. Existing purification methods focus on static graphs, limiting their ability to account for critical temporal dependencies in dynamic graphs. In this work, we propos
Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, they remain vulnerable to jailbreak attacks, which exploit weaknesses in their safety alignment to generate harmful responses. Previous studies categorize jailbreaks as successful or failed based on whether responses conta
Formulas for Mutually Orthogonal Quantum States in Two-Qubit Systems: Orthogonal Schmidt Decompositions
quant-phYonghae Lee, Youngho Min, Sunghyun Bae, Youngrong Lim
We present Schmidt decomposition formulas for mutually orthogonal two-qubit pure states and classify orthonormal sets based on their entanglement structure. First, we derive explicit Schmidt decomposition formulas for any pure state and extend them to two orthogonal pure states. For three mutually orthogonal states, we provide formulas for specific cases and
Jiho Jin, Woosung Kang, Junho Myung, Alice Oh
Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark for QA (BBQ), designed to evaluate social bias in long-form generation by having LLMs generate continuations of story prom
ConcreTizer: Model Inversion Attack via Occupancy Classification and Dispersion Control for 3D Point Cloud Restoration
cs.CVYoungseok Kim, Sunwook Hwang, Hyung-Sin Kim, Saewoong Bahk
The growing use of 3D point cloud data in autonomous vehicles (AVs) has raised serious privacy concerns, particularly due to the sensitive information that can be extracted from 3D data. While model inversion attacks have been widely studied in the context of 2D data, their application to 3D point clouds remains largely unexplored. To fill this gap, we prese
Mohammed Mahfoud, Ghait Boukachab, Michał Koziarski, Alex Hernandez-Garcia
Building predictive models for tabular data presents fundamental challenges, notably in scaling consistently, i.e., more resources translating to better performance, and generalizing systematically beyond the training data distribution. Designing decision tree models remains especially challenging given the intractably large search space, and most existing m
Juncheng Wang, Chao Xu, Cheng Yu, Lei Shang
Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents how to balance the representation of mel-spectrograms in term
Jiahao Wang, Xiangyu Cao, Jiaru Zhong, Yuner Zhang
While cooperative perception can overcome the limitations of single-vehicle systems, the practical implementation of vehicle-to-vehicle and vehicle-to-infrastructure systems is often impeded by significant economic barriers. Aerial-ground cooperation (AGC), which pairs ground vehicles with drones, presents a more economically viable and rapidly deployable al
Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization
cs.LGZiqing Xu, Hancheng Min, Lachlan Ewen MacDonald, Jinqi Luo
Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix
Manjun Cui, Zhichao Zhang, Wei Yao
Graph signal processing (GSP) has emerged as a powerful framework for analyzing data on irregular domains. In recent years, many classical techniques in signal processing (SP) have been successfully extended to GSP. Among them, chirp signals play a crucial role in various SP applications. However, graph chirp signals have not been formally defined despite th
Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi
This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics (size, multimodal capabilities, architectural generation) influence grounding performance, distributional factor dependenci
Jiyong Chen, Cai Heng Li, Ci Xuan Wu, Yan Zhou Zhu
We construct connected $2$-arc-transitive covers of the Petersen graph with non-solvable transformation groups, solving the long-standing problem for the existence of such covers.
Xinyu Xi, Hua Yang, Shentai Zhang, Yijie Liu
Maritime Multi-Scene Recognition is crucial for enhancing the capabilities of intelligent marine robotics, particularly in applications such as marine conservation, environmental monitoring, and disaster response. However, this task presents significant challenges due to environmental interference, where marine conditions degrade image quality, and the compl
Andrey Akhmeteli
Previously, the author offered a plasma-like description of quantum phenomena. This article offers a new criterion of approximation of probability density functions of quantum theories by sums of $\delta$-functions with integer coefficients and a constructive approach to building such sets of $\delta$-functions.
Hao-Long Zhang, Pei-Rong Han, Fan Wu, Wen Ning
One of the most remarkable features that distinguish open systems from closed ones is the presence of exceptional points (EPs), where two or more eigenvectors of a non-Hermitian operator coalesce, accompanying the convergence of the correcponding eigenvalues. So far, EPs have been demonstrated on a number of platforms, ranging from classical optical systems
Task-Specific Knowledge Distillation from the Vision Foundation Model for Enhanced Medical Image Segmentation
cs.CVPengchen Liang, Haishan Huang, Bin Pu, Jianguo Chen
Large-scale pre-trained models, such as Vision Foundation Models (VFMs), have demonstrated impressive performance across various downstream tasks by transferring generalized knowledge, especially when target data is limited. However, their high computational cost and the domain gap between natural and medical images limit their practical application in medic
Salim Rostam
Given two affine permutations, some results of Lascoux and Deodhar, and independently Jacon-Lecouvey, allow to decide if they are comparable for the strong Bruhat order. These permutations are associated with tuples of core partitions, and the preceding problem is equivalent to compare the Young diagrams in each components for the inclusion. Using abaci, we
Yang Liu, Mengyuan Liu, Shudong Huang, Jiancheng Lv
Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modal
Xiang Liu, Zhaoxiang Liu, Huan Hu, Zezhou Chen
While conversational generative AI has shown considerable potential in enhancing decision-making for agricultural professionals, its exploration has predominantly been anchored in text-based interactions. The evolution of multimodal conversational AI, leveraging vast amounts of image-text data from diverse sources, marks a significant stride forward. However
Novel concept for low-energy antineutron production and its application for antineutron scattering experiments
nucl-exAlessandra Filippi, Hiroyuki Fujioka, Takashi Higuchi, Luca Venturelli
Extensive data of antiproton scattering cross sections with protons and nuclei have advanced our understanding of hadronic interactions with antinucleons. However, low-energy antineutron scattering data are scarce, thereby limiting our understanding of the S-wave antinucleon-nucleon and antinucleon-nucleus interactions. We present a novel production scheme o
Gideon Yoffe, Keren Duer, Tom Andre Nordheim, Itay Halevy
Europa, Jupiter's second Galilean moon, is believed to host a subsurface ocean in contact with a rocky mantle, where hydrothermal activity may drive the synthesis of organic molecules. Of these molecules, abiotic synthesis of aromatic amino acids is unlikely, and their detection on Europa could be considered a biosignature. Fluorescence from aromatic amino a
Shogen Kawanami, Kento Iseri, Tomohiro I
The Burrows-Wheeler Transform (BWT) of a string is an invertible permutation of the string, which can be used for data compression and compact indexes for string pattern matching. Ganguly et al. [SODA, 2017] introduced the parameterized BWT (pBWT) to design compact indexes for parameterized matching (p-matching), a variant of string pattern matching with par
Nihat Mugurtay
This article examines how unequal access to AI innovation creates systemic challenges for developing countries. Differential access to AI innovation results from the acute competition between domestic and global actors. While developing nations contribute significantly to AI development through data annotation labor, they face limited access to advanced AI t
Nonlinear Einstein-Power-Yang-Mills AdS Black Holes: From Quantum Tunneling to Aschenbach Effect
physics.gen-phErdem Sucu, İzzet Sakallı
This study investigates the thermodynamic and quantum properties of Einstein-Power-Yang-Mills (EPYM) black holes in an Anti-de Sitter background, focusing on the effects of the nonlinear Yang-Mills charge parameter $\gamma$. We derive the metric function, analyze Hawking radiation through boson tunneling, and calculate thermodynamic properties including temp
MaNGA DynPop. VII. A Unified Bulge-Disk-Halo Model for Explaining Diversity in Circular Velocity Curves of 6000 Spiral and Early-Type Galaxies
astro-ph.GAKai Zhu, Michele Cappellari, Shude Mao, Shengdong Lu
We derive circular velocity curves (CVCs) from stellar dynamical models for $\sim6000$ nearby galaxies in the final data release of the Sloan Digital Sky Survey-IV MaNGA survey with integral-field spectroscopy, exploring connections between the inner gravitational potential (traced by CVC amplitude/shape) and galaxy properties. The maximum circular velocity
Anna Aksamit, Kaustav Das, Ivan Guo, Kihun Nam
We consider a continuum of carbon-emitting firms who seek to maximise their stock price, and a regulator (e.g., Government) who wishes for the economy to flourish, whilst simultaneously punishing firms who behave non-green. Interpreting the regulator as a major player and the firms as the minor players, we model this setting through a mean field game with ma
Guanghao Li, Mingzhi Chen, Hao Yu, Shuting Dong
Deep learning-based denoising models have been widely employed in vision tasks, functioning as filters to eliminate noise while retaining crucial semantic information. Additionally, they play a vital role in defending against adversarial perturbations that threaten downstream tasks. However, these models can be intrinsically susceptible to adversarial attack
SeCap: Self-Calibrating and Adaptive Prompts for Cross-view Person Re-Identification in Aerial-Ground Networks
cs.CVShining Wang, Yunlong Wang, Ruiqi Wu, Bingliang Jiao
When discussing the Aerial-Ground Person Re-identification (AGPReID) task, we face the main challenge of the significant appearance variations caused by different viewpoints, making identity matching difficult. To address this issue, previous methods attempt to reduce the differences between viewpoints by critical attributes and decoupling the viewpoints. Wh
Taeyong Ahn
We investigate the intersection of positive closed currents in a general setting, employing tangent currents alongside King's residue formula. Our main result establishes a natural condition for the intersection--namely, the Dinh-Sibony product--of positive closed currents on domains and derives an integral representation of this intersection. In parallel, w
Kyungho Kim, Sunwoo Kim, Geon Lee, Jinhong Jung
Traditional recommender systems primarily rely on a single type of user-item interaction, such as item purchases or ratings, to predict user preferences. However, in real-world scenarios, users engage in a variety of behaviors, such as clicking on items or adding them to carts, offering richer insights into their interests. Multi-behavior recommender systems
Zenghao Guan, Yucan Zhou, Xiaoyan Gu
Traditional Federated Learning (FL) necessitates numerous rounds of communication between the server and clients, posing significant challenges including high communication costs, connection drop risks and susceptibility to privacy attacks. One-shot FL has become a compelling learning paradigm to overcome above drawbacks by enabling the training of a global
Hyeong-Chan Kim, Wonwoo Lee
We present a new rotating black hole solution to the Einstein equations as an extension of the Kerr spacetime. Interestingly, the solution we find may not be uniquely characterized by asymptotic parameters such as mass, angular momentum, and charge, thereby it would be the additional hair. We also analyze in detail how this additional characteristics or this
Xin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang
Pre-trained vision models (PVMs) are fundamental to modern robotics, yet their optimal configuration remains unclear. Through systematic evaluation, we find that while DINO and iBOT outperform MAE across visuomotor control and perception tasks, they struggle when trained on non-(single-)object-centric (NOC) data--a limitation strongly correlated with their d
DynTaskMAS: A Dynamic Task Graph-driven Framework for Asynchronous and Parallel LLM-based Multi-Agent Systems
cs.MAJunwei Yu, Yepeng Ding, Hiroyuki Sato
The emergence of Large Language Models (LLMs) in Multi-Agent Systems (MAS) has opened new possibilities for artificial intelligence, yet current implementations face significant challenges in resource management, task coordination, and system efficiency. While existing frameworks demonstrate the potential of LLM-based agents in collaborative problem-solving,
EnCortex: A General, Extensible and Scalable Framework for Decision Management in New-age Energy Systems
eess.SYMillend Roy, Vaibhav Balloli, Anupam Sobti, Srinivasan Iyengar
With increased global warming, there has been a significant emphasis to replace fossil fuel-dependent energy sources with clean, renewable sources. These new-age energy systems are becoming more complex with an increasing proportion of renewable energy sources (like solar and wind), energy storage systems (like batteries), and demand side control in the mix.
Chihiro Oguri, Mao Shinoda
We investigate the stability of maximizing measures for a penalty function of a two-dimensional subshift of finite type, building on the work of Gonschorowski et al. \cite{GQS}. In the one-dimensional case, such measures remain stable under Lipschitz perturbations for any subshift of finite type. However, instability arises for a penalty function of the Robi
David Darrow, George Stepaniants
This work aims to bridge the gap between pure and applied research on scalar, linear Volterra equations by examining five major classes: integral and integro-differential equations with completely monotone kernels, such as linear viscoelastic models; equations with positive definite kernels, such as partially observed quantum systems; difference equations wi
Jian Jin, Zhenbo Yu, Yang Shen, Zhenyong Fu
Customized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framewo
Zeyu Zhang, Yiran Wang, Wei Mao, Danning Li
Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models lack a mechanism to prioritize dynamic frames and body parts based on given conditions. Second, existing methods for differ
Xingye Fan, Zhongwen, Zhang, Yuri Boykov
This paper demonstrates a surprising result for segmentation with image-level targets: extending binary class tags to approximate relative object-size distributions allows off-the-shelf architectures to solve the segmentation problem. A straightforward zero-avoiding KL-divergence loss for average predictions produces segmentation accuracy comparable to the s
MERLION: Marine ExploRation with Language guIded Online iNformative Visual Sampling and Enhancement
cs.ROShrutika Vishal Thengane, Marcel Bartholomeus Prasetyo, Yu Xiang Tan, Malika Meghjani
Autonomous and targeted underwater visual monitoring and exploration using Autonomous Underwater Vehicles (AUVs) can be a challenging task due to both online and offline constraints. The online constraints comprise limited onboard storage capacity and communication bandwidth to the surface, whereas the offline constraints entail the time and effort required
Occurrence of chemically tuned spin-texture controlled large intrinsic anomalous Hall effect in epitaxial $Mn_{3+x}Pt_{1-x}$ thin Films
cond-mat.mtrl-sciIndraneel Sinha, Saurav Sachin, Shreyashi Sinha, Roumita Roy
Achieving atomically flat and stoichiometric films of chiral antiferromagnets (AFM) with two-dimensional kagome spin lattice structures are crucial for integrating these materials in both established and emerging antiferromagnetic spintronic devices. We report a systematic study of growth and anomalous Hall effect in (111)-oriented non-collinear AFM $Mn_{3+x
Xinjie Zhao, Fan Gao, Xingyu Song, Yingjian Chen
Recent advances in large language models (LLMs) have significantly improved multi-hop question answering (QA) through direct Chain-of-Thought (CoT) reasoning. However, the irreversible nature of CoT leads to error accumulation, making it challenging to correct mistakes in multi-hop reasoning. This paper introduces ReAgent: a Reversible multi-Agent collaborat
CtrlRAG: Black-box Document Poisoning Attacks for Retrieval-Augmented Generation of Large Language Models
cs.CLRunqi Sui
Retrieval-Augmented Generation (RAG) systems enhance response credibility and traceability by displaying reference contexts, but this transparency simultaneously introduces a novel black-box attack vector. Existing document poisoning attacks, where adversaries inject malicious documents into the knowledge base to manipulate RAG outputs, rely primarily on unr
Haotian Chen, Yanyu Xu, Boyan Wang, Chaoyue Zhao
In this report, we introduce our first-generation reasoning model, LexPro-1.0, a large language model designed for the highly specialized Chinese legal domain, offering comprehensive capabilities to meet diverse realistic needs. Existing legal LLMs face two primary challenges. Firstly, their design and evaluation are predominantly driven by computer science
Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo
Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverag
Aligning Instance-Semantic Sparse Representation towards Unsupervised Object Segmentation and Shape Abstraction with Repeatable Primitives
cs.CVJiaxin Li, Hongxing Wang, Jiawei Tan, Zhilong Ou
Understanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while curre
Xu-Ke Gu, Li-Zhou Tan, Franco Nori, J. Q. You
Markovian open quantum systems are governed by the Lindblad master equation where the dissipation contains two parts, i.e., the anti-Hermitian operator and the quantum jumps, which share a common dissipation rate. We generalize the Lindblad master equation via postselection to a generalized Liouvillian formalism in which the effective damping rate of the ant
Dynamic Cross-Modal Feature Interaction Network for Hyperspectral and LiDAR Data Classification
eess.IVJunyan Lin, Feng Gap, Lin Qi, Junyu Dong
Hyperspectral image (HSI) and LiDAR data joint classification is a challenging task. Existing multi-source remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel Dynamic Cross-Modal Feature Interaction Network (DCMNet
Yan Hu, Ahmad Chaddad
This study introduces the SHAP-integrated convolutional diagnostic network (SICDN), an interpretable feature selection method designed for limited datasets, to address the challenge posed by data privacy regulations that restrict access to medical datasets. The SICDN model was tested on classification tasks using pneumonia and breast cancer datasets, demonst
Zhiheng Yu, Jiancheng An, Lu Gan, Hongbin Li
Reconfigurable intelligent surfaces (RIS) can reshape the characteristics of wireless channels by intelligently regulating the phase shifts of reflecting elements. Recently, various codebook schemes have been utilized to optimize the reflection coefficients (RCs); however, the selection of the optimal codeword is usually obtained by evaluating a metric of in
Yuzhu Lei, Qiqi Xiao, Yinghui He, Guanding Yu
In massive multi-input multi-output (MIMO) systems, the main bottlenecks of location- and orientation-assisted beam alignment using deep neural networks (DNNs) are large training overhead and significant performance degradation. This paper proposes a graph neural network (GNN)-based beam selection approach that reduces the training overhead and improves the
Do JWST reionization (optical depth) puzzle, cosmological tensions, and CMB anomalies imply Harrison-Zel'dovich spectrum?
astro-ph.COHao-Hao Li, Xin-zhe Zhang, Taotao Qiu, Jun-Qing Xia
The James Webb Space Telescope (JWST) has observed massive galaxies at high redshifts, which implies an earlier epoch of reionization (EoR) compared with the cosmic microwave background (CMB) results. In this paper, based on \texttt{Planck 2020} (NPIPE release), \texttt{ACT DR4} and \texttt{SPT-3G} data, if assumed a Harrison-Zel'dovich (HZ) primordial power
CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing
cs.CVJianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng
Most research decoding brain signals into images, often using them as priors for generative models, has focused only on visual content. This overlooks the brain's natural ability to integrate auditory and visual information, for instance, sound strongly influences how we perceive visual scenes. To investigate this, we propose a new task of reconstructing con
Andy Chia, Wai-Keong Mok, Leong-Chuan Kwek, Changsuk Noh
Several important dynamical systems are in $\mathbb{R}^2$, defined by the pair of differential equations $(x',y')=(f(x,y),g(x,y))$. A question of fundamental importance is how such systems might behave quantum mechanically. In developing quantum theory, Dirac and others realized that classical Hamiltonian systems can be mapped to their quantum counterparts v
Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian
The increasing pace of population aging calls for better care and support systems. Falling is a frequent and critical problem for elderly people causing serious long-term health issues. Fall detection from video streams is not an attractive option for real-life applications due to privacy issues. Existing methods try to resolve this issue by using very low-r
Ruimeng Liu, Xinhang Xu, Shenghai Yuan, Lihua Xie
Zero-Shot Object Navigation (ZSON) requires agents to navigate to objects specified via open-ended natural language without predefined categories or prior environmental knowledge. While recent methods leverage foundation models or multi-modal maps, they often rely on 2D representations and greedy strategies or require additional training or modules with high
Zhengyang Mei, Xiaohui Song, Xueyi Guo, Xiang Li
In this paper, we introduce a method of using a double-layer resist lift-off process to prepare the capacitor dielectric layer for fabricating impedance-engineered Josephson parametric amplifiers (IMPAs). Compared with traditional techniques, this method enhances fabrication success rate, accelerates production. The IMPA we made experimentally achieves an in
Generic non-degeneracy of critical points of multiple Green functions on torus and applications to curvature equations
math.APZhijie Chen, Erjuan Fu, Chang-Shou Lin
Let $E_{\tau}:=\mathbb{C}/(\mathbb{Z}+\mathbb{Z}\tau)$ with $\operatorname{Im}\tau>0$ be a flat torus and $G(z;\tau)$ be the Green function on $E_{\tau}$ with the singularity at $0$. Consider the multiple Green function $G_{n}$ on $(E_{\tau})^{n}$: \[ G_{n}(z_{1},\cdots,z_{n};\tau):=\sum_{i<j}G(z_{i}-z_{j};\tau)-n\sum_{i=1}% ^{n}G(z_{i};\tau). \] Recently, L
Hanyu Zhou, Gim Hee Lee
Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suffer from temporal sparsity that limits language-vision tempo
W. B. Rui, Z. D. Wang
Exceptional points (EPs) are prominent non-Hermitian band degeneracies that give rise to a variety of intriguing and unconventional phenomena. Similar to Weyl and Dirac points, EPs carry topological charges and comply with the celebrated fermion doubling theorems in lattices. Beyond these characteristics, EPs exhibit more exotic topological properties, parti
Deuterium-deuterium fusion charged particle detection using CR-39 and Deep Learning Model
physics.ins-detYuxing Wang, Allan Xi Chen, Matthew Salazar, Nawar Abdalla
CR-39 solid-state nuclear track detectors are widely used in fusion research for detecting charged particles produced in fusion reactions. However, analyzing increasingly complex and large-scale CR-39 track images to extract meaningful information can be a tedious and time-consuming process, often prone to human errors and bias. To address these challenges,
Zhao Tang, Fanhao Jia, Greis J. Kim-Reyes, Yabei Wu
Color centers exhibiting deep-level states within the wide bandgap h-BN monolayer possess substantial potential for quantum applications. Uncovering precise geometric characteristics at the atomic scale is crucial for understanding defect performance. In this study, first-principles calculations were performed on the most extensively investigated CBVN and NB
Ning Ding, Jing Han, Yuchuan Tian, Chao Xu
Diffusion Transformer (DiT) has now become the preferred choice for building image generation models due to its great generation capability. Unlike previous convolution-based UNet models, DiT is purely composed of a stack of transformer blocks, which renders DiT excellent in scalability like large language models. However, the growing model size and multi-st
Yanlong Wang, Jian Xu, Shao-Lun Huang, Danny Dongning Sun
This study seeks to advance the understanding and prediction of stock market return uncertainty through the application of advanced deep learning techniques. We introduce a novel deep learning model that utilizes a Gaussian mixture distribution to capture the complex, time-varying nature of asset return distributions in the Chinese stock market. By incorpora
FinTSBridge: A New Evaluation Suite for Real-world Financial Prediction with Advanced Time Series Models
cs.LGYanlong Wang, Jian Xu, Tiantian Gao, Hongkang Zhang
Despite the growing attention to time series forecasting in recent years, many studies have proposed various solutions to address the challenges encountered in time series prediction, aiming to improve forecasting performance. However, effectively applying these time series forecasting models to the field of financial asset pricing remains a challenging issu
Yuchen Han, Yucheng Wu, Jeffrey Willard
This paper investigates a critical aspect of large language model (LLM) performance: the optimal formatting of classification task options in prompts. Through an extensive experimental study, we compared two selection formats -- bullet points and plain English -- to determine their impact on model performance. Our findings suggest that presenting options via
Complete Key Recovery of a DNA-based Encryption and Developing a Novel Stream Cipher for Color Image Encryption: Bio-SNOW
cs.CRYash Makwana, Anupama Panigrahi, Saibal K. Pal
Recent studies have explored DNA-based algorithms for IoT security and image encryption. A similar encryption algorithm was proposed by Al-Husainy et. al. in 2021 Recently, Al-Husainy et al.in 2021, proposed an encryption algorithm based on DNA processes for Internet of Things(IoT) applications. Upon finding low avalanche effect in our experiments, we first
Michael McGuire
Automatic speech recognition (ASR) has been an essential component of computer assisted language learning (CALL) and computer assisted language testing (CALT) for many years. As this technology continues to develop rapidly, it is important to evaluate the accuracy of current ASR systems for language learning applications. This study assesses five cutting-edg
Jiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen
Diffusion Transformers (DiT) have revolutionized high-fidelity image and video synthesis, yet their computational demands remain prohibitive for real-time applications. To solve this problem, feature caching has been proposed to accelerate diffusion models by caching the features in the previous timesteps and then reusing them in the following timesteps. How
Accelerated Quasi-Static FEM for Real-Time Modeling of Continuum Robots with Multiple Contacts and Large Deformation
cs.ROHao Chen, Jian Chen, Xinran Liu, Zihui Zhang
Continuum robots offer high flexibility and multiple degrees of freedom, making them ideal for navigating narrow lumens. However, accurately modeling their behavior under large deformations and frequent environmental contacts remains challenging. Current methods for solving the deformation of these robots, such as the Model Order Reduction and Gauss-Seidel (
Youngeun Kim, Seunghwan Lee, Aecheon Jung, Bogon Ryu
Model merging enables efficient multi-task models by combining task-specific fine-tuned checkpoints. However, storing multiple task-specific checkpoints requires significant memory, limiting scalability and restricting model merging to larger models and diverse tasks. In this paper, we propose quantizing task vectors (i.e., the difference between pre-trained
Chengzhi Lin, Chuyuan Wang, Annan Xie, Wuhong Wang
In video recommendation systems, user behaviors such as watch time, likes, and follows are commonly used to infer user interest. However, these behaviors are influenced by various biases, including duration bias, demographic biases, and content category biases, which obscure true user preferences. In this paper, we hypothesize that biases and user interest a
CAFusion: Controllable Anatomical Synthesis of Perirectal Lymph Nodes via SDF-guided Diffusion
eess.IVWeidong Guo, Hantao Zhang, Shouhong Wan, Bingbing Zou
Lesion synthesis methods have made significant progress in generating large-scale synthetic datasets. However, existing approaches predominantly focus on texture synthesis and often fail to accurately model masks for anatomically complex lesions. Additionally, these methods typically lack precise control over the synthesis process. For example, perirectal ly
Lingrui Ge, Yiqian Wang, Jiahao Xu
This paper solves ``The Dry Ten Martini Problem'' for $C^2$ cosine-type quasiperiodic Schr\"odinger operators with large coupling constants and Diophantine frequencies, a model originally introduced by Sinai in 1987 \cite{sinai}. This shows that the analyticity assumption on the potential is not essential for obtaining a dry Cantor spectrum and can be replac
Pranjal Awasthi, Sreenivas Gollapudi, Ravi Kumar, Kamesh Munagala
We study zeroth-order optimization where solutions must minimize a cost $d(s)$ while maintaining high probability under a complex generative prior $L(s)$ (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to $L(s) e^{-T \cdot d(s)}$. Since classical model-based optimization (MBO) lacks finite-sample guarantees for
You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed Data
cs.LGShanshan Yan, Zexi Li, Chao Wu, Meng Pang
Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic sim
Spatio-temporal characterization of nonlinear forcing and response in turbulent channel flow
physics.flu-dynYuting Huang, Simon S. Toedtli, Gregory P. Chini, Beverley J. McKeon
The quadratic convection term in the incompressible Navier-Stokes equations is considered as a nonlinear forcing to the linear resolvent operator, and it is studied in the Fourier domain through the analysis of interactions between triadically compatible wavenumber-frequency triplets. A framework to quantify the triadic contributions to the forcing and respo
Dung Xuan Nguyen, Dam Thanh Son
A low-energy neutral quasiparticle in a fractional quantum Hall system appears in the latter's energy spectrum on a sphere as a series of many-body excited states labeled by the angular momentum $L$ and whose energy is a smooth function of $L$ in the limit of large sphere radius. We argue that the signature of a nonvanishing spin (intrinsic angular momentum)