Skip to content

March 2025 arXiv papers — page 172

Showing 17,10117,200 of 23,633 papers

  1. Kaiyuan Liu, Youcheng Pan, Yang Xiang, Daojing He

    Recently, LLM agents have made rapid progress in improving their programming capabilities. However, existing benchmarks lack the ability to automatically evaluate from users' perspective, and also lack the explainability of the results of LLM agents' code generation capabilities. Thus, we introduce ProjectEval, a new benchmark for LLM agents project-level co

  2. Han Hong, Gaoming Wang

    We prove a splitting theorem for a smooth noncompact manifold with (possibly noncompact) boundary. We show that if a noncompact manifold of dimension $n\geq 2$ has $\lambda_1(-\alpha\Delta+\operatorname{Ric})\geq 0$ for some $\alpha<\frac{4}{n-1}$ and mean-convex boundary, then it is either isometric to $\Sigma\times \mathbb{R}_{\geq 0}$ for a closed manifol

  3. Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian

    Older people are susceptible to fall due to instability in posture and deteriorating health. Immediate access to medical support can greatly reduce repercussions. Hence, there is an increasing interest in automated fall detection, often incorporated into a smart healthcare system to provide better monitoring. Existing systems focus on wearable devices which

  4. Yan Wei, Yu Feng, Linlin Ou, Yueying Wang

    This paper investigates the safety analysis and verification of nonlinear systems subject to high-relative-degree constraints and unknown disturbance. The closed-form solution of the high-order control barrier functions (HOCBF) optimization problem with and without a nominal controller is first provided, making it unnecessary to solve the quadratic program p

  5. Shuhao Liao, Xuxin Lv, Yuhong Cao, Jeric Lew

    In autonomous exploration tasks, robots are required to explore and map unknown environments while efficiently planning in dynamic and uncertain conditions. Given the significant variability of environments, human operators often have specific preference requirements for exploration, such as prioritizing certain areas or optimizing for different aspects of e

  6. Alexander Eber, Christoph Gruber, Martin Schultze, Birgitta Bernhardt

    We radically simplify coherently averaged dual-comb spectroscopy by introducing a real-time self-correction system: a radio frequency system-on-chip computes each incoming dual-comb interferogram's phase, frequency, and arrival time; calculates changes in the combs' carrier-envelope offset frequency and repetition rate difference; and immediately phase-corre

  7. Jiaojiao Li, Shiyao Duan, Haitao XU, Rui Song

    The inherent difficulty in acquiring accurately co-registered RGB-hyperspectral image (HSI) pairs has significantly impeded the practical deployment of current data-driven Hyperspectral Image Generation (HIG) networks in engineering applications. Gleichzeitig, the ill-posed nature of the aligning constraints, compounded with the complexities of mining cross-

  8. Ruoxi Xu, Hongyu Lin, Xianpei Han, Jia Zheng

    As large language models (LLMs) increasingly become central to various applications and interact with diverse user populations, ensuring their reliable and consistent performance is becoming more important. This paper explores a critical issue in assessing the reliability of LLMs: the consistency between their words and deeds. To quantitatively explore this

  9. Jiazheng Liu, Sipeng Zheng, Börje F. Karlsson, Zongqing Lu

    Multimodal large language models (MLLMs), built on large-scale pre-trained vision towers and language models, have shown great capabilities in multimodal understanding. However, most existing MLLMs are trained on single-turn vision question-answering tasks, which do not accurately reflect real-world human conversations. In this paper, we introduce MMDiag, a

  10. Jacek Jakimiuk

    We give a strengthening of the classical Khintchine inequality between the second and the $p$-th moment for $p \ge 3$ with optimal constant by adding a deficit depending on the vector of coefficients of the Rademacher sum.

  11. Zhaojie Zeng, Yuesong Wang, Lili Ju, Tao Guan

    By adaptively controlling the density and generating more Gaussians in regions with high-frequency information, 3D Gaussian Splatting (3DGS) can better represent scene details. From the signal processing perspective, representing details usually needs more Gaussians with relatively smaller scales. However, 3DGS currently lacks an explicit constraint linking

  12. Chase Hutton, Adam Melrod

    Many parallel algorithms which solve basic problems in computer science use auxiliary space linear in the input to facilitate conflict-free computation. There has been significant work on improving these parallel algorithms to be in-place, that is to use as little auxiliary memory as possible. In this paper, we provide novel in-place algorithms to solve the

  13. Haoyu Zheng, Qifan Yu, Binghe Yu, Yang Dai

    Diffusion models have achieved remarkable progress in image and video stylization. However, most existing methods focus on single-style transfer, while video stylization involving multiple styles necessitates seamless transitions between them. We refer to this smooth style transition between video frames as video style morphing. Current approaches often gene

  14. Kwanyoung Kim, Byeongsu Sim

    Diffusion models have shown impressive results in generating high-quality conditional samples using guidance techniques such as Classifier-Free Guidance (CFG). However, existing methods often require additional training or neural function evaluations (NFEs), making them incompatible with guidance-distilled models. Also, they rely on heuristic approaches that

  15. Qian Liu, Lan Wang, Bing Yang, Hao Wu

    Water quality data can supply a substantial decision support for water resources utilization and pollution prevention. However, there are numerous missing values in water quality data due to inescapable factors like sensor failure, thereby leading to biased result for hydrological analysis and failing to support environmental governance decision accurately.

  16. Stylianos Zindros, Christos Chronis, Panagiotis Radoglou-Grammatikis, Vasileios Argyriou

    As the security of public spaces remains a critical issue in today's world, Digital Twin technologies have emerged in recent years as a promising solution for detecting and predicting potential future threats. The applied methodology leverages a Digital Twin of a metro station in Athens, Greece, using the FlexSim simulation software. The model encompasses po

  17. Haolin Li, Yikang Chai, Bailin Lv, Lecheng Ruan

    This study introduces a unified control framework that addresses the challenge of precise quadruped locomotion with unknown payloads, named as online payload identification-based physics-informed neural network predictive control (OPI-PINNPC). By integrating online payload identification with physics-informed neural networks (PINNs), our approach embeds iden

  18. Lei Zhang, Mukesh Ghimire, Wenlong Zhang, Zhe Xu

    General-sum differential games can approximate values solved by Hamilton-Jacobi-Isaacs (HJI) equations for efficient inference when information is incomplete. However, solving such games through conventional methods encounters the curse of dimensionality (CoD). Physics-informed neural networks (PINNs) offer a scalable approach to alleviate the CoD and approx

  19. Shihao Hou, Xinyi Shang, Shreyank N Gowda, Yang Lu

    Effectively handling the co-occurrence of non-IID data and long-tailed distributions remains a critical challenge in federated learning. While fine-tuning vision-language models (VLMs) like CLIP has shown to be promising in addressing non-IID data challenges, this approach leads to severe degradation of tail classes in federated long-tailed scenarios. Under

  20. Hanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan

    High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically, existing methods mainly introduce event camera to directly fuse the spatiotemporal features between the two modalities. Howeve

  21. Yongwoo Kim, Sungmin Cha, Donghyun Kim

    Machine unlearning is a process to remove specific data points from a trained model while maintaining the performance on the retain data, addressing privacy or legal requirements. Despite its importance, existing unlearning evaluations tend to focus on logit-based metrics under small-scale scenarios. We observe that this could lead to a false sense of securi

  22. Hyeonsoo Jo, Jongha Lee, Fanchen Bu, Kijung Shin

    Time-evolving graphs, such as social and citation networks, often contain noise that distorts structural and temporal patterns, adversely affecting downstream tasks, such as node classification. Existing purification methods focus on static graphs, limiting their ability to account for critical temporal dependencies in dynamic graphs. In this work, we propos

  23. Wenzhuo Xu, Zhipeng Wei, Xiongtao Sun, Zonghao Ying

    Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, they remain vulnerable to jailbreak attacks, which exploit weaknesses in their safety alignment to generate harmful responses. Previous studies categorize jailbreaks as successful or failed based on whether responses conta

  24. Yonghae Lee, Youngho Min, Sunghyun Bae, Youngrong Lim

    We present Schmidt decomposition formulas for mutually orthogonal two-qubit pure states and classify orthonormal sets based on their entanglement structure. First, we derive explicit Schmidt decomposition formulas for any pure state and extend them to two orthogonal pure states. For three mutually orthogonal states, we provide formulas for specific cases and

  25. Jiho Jin, Woosung Kang, Junho Myung, Alice Oh

    Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark for QA (BBQ), designed to evaluate social bias in long-form generation by having LLMs generate continuations of story prom

  26. Youngseok Kim, Sunwook Hwang, Hyung-Sin Kim, Saewoong Bahk

    The growing use of 3D point cloud data in autonomous vehicles (AVs) has raised serious privacy concerns, particularly due to the sensitive information that can be extracted from 3D data. While model inversion attacks have been widely studied in the context of 2D data, their application to 3D point clouds remains largely unexplored. To fill this gap, we prese

  27. Mohammed Mahfoud, Ghait Boukachab, Michał Koziarski, Alex Hernandez-Garcia

    Building predictive models for tabular data presents fundamental challenges, notably in scaling consistently, i.e., more resources translating to better performance, and generalizing systematically beyond the training data distribution. Designing decision tree models remains especially challenging given the intractably large search space, and most existing m

  28. Juncheng Wang, Chao Xu, Cheng Yu, Lei Shang

    Video-to-audio generation is essential for synthesizing realistic audio tracks that synchronize effectively with silent videos. Following the perspective of extracting essential signals from videos that can precisely control the mature text-to-audio generative diffusion models, this paper presents how to balance the representation of mel-spectrograms in term

  29. Jiahao Wang, Xiangyu Cao, Jiaru Zhong, Yuner Zhang

    While cooperative perception can overcome the limitations of single-vehicle systems, the practical implementation of vehicle-to-vehicle and vehicle-to-infrastructure systems is often impeded by significant economic barriers. Aerial-ground cooperation (AGC), which pairs ground vehicles with drones, presents a more economically viable and rapidly deployable al

  30. Ziqing Xu, Hancheng Min, Lachlan Ewen MacDonald, Jinqi Luo

    Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix

  31. Manjun Cui, Zhichao Zhang, Wei Yao

    Graph signal processing (GSP) has emerged as a powerful framework for analyzing data on irregular domains. In recent years, many classical techniques in signal processing (SP) have been successfully extended to GSP. Among them, chirp signals play a crucial role in various SP applications. However, graph chirp signals have not been formally defined despite th

  32. Jonghyun Lee, Dojun Park, Jiwoo Lee, Hoekeon Choi

    This study investigated whether multimodal large language models can achieve human-like sensory grounding by examining their ability to capture perceptual strength ratings across sensory modalities. We explored how model characteristics (size, multimodal capabilities, architectural generation) influence grounding performance, distributional factor dependenci

  33. Jiyong Chen, Cai Heng Li, Ci Xuan Wu, Yan Zhou Zhu

    We construct connected $2$-arc-transitive covers of the Petersen graph with non-solvable transformation groups, solving the long-standing problem for the existence of such covers.

  34. Xinyu Xi, Hua Yang, Shentai Zhang, Yijie Liu

    Maritime Multi-Scene Recognition is crucial for enhancing the capabilities of intelligent marine robotics, particularly in applications such as marine conservation, environmental monitoring, and disaster response. However, this task presents significant challenges due to environmental interference, where marine conditions degrade image quality, and the compl

  35. Andrey Akhmeteli

    Previously, the author offered a plasma-like description of quantum phenomena. This article offers a new criterion of approximation of probability density functions of quantum theories by sums of $\delta$-functions with integer coefficients and a constructive approach to building such sets of $\delta$-functions.

  36. Hao-Long Zhang, Pei-Rong Han, Fan Wu, Wen Ning

    One of the most remarkable features that distinguish open systems from closed ones is the presence of exceptional points (EPs), where two or more eigenvectors of a non-Hermitian operator coalesce, accompanying the convergence of the correcponding eigenvalues. So far, EPs have been demonstrated on a number of platforms, ranging from classical optical systems

  37. Pengchen Liang, Haishan Huang, Bin Pu, Jianguo Chen

    Large-scale pre-trained models, such as Vision Foundation Models (VFMs), have demonstrated impressive performance across various downstream tasks by transferring generalized knowledge, especially when target data is limited. However, their high computational cost and the domain gap between natural and medical images limit their practical application in medic

  38. Salim Rostam

    Given two affine permutations, some results of Lascoux and Deodhar, and independently Jacon-Lecouvey, allow to decide if they are comparable for the strong Bruhat order. These permutations are associated with tuples of core partitions, and the preceding problem is equivalent to compare the Young diagrams in each components for the inclusion. Using abaci, we

  39. Yang Liu, Mengyuan Liu, Shudong Huang, Jiancheng Lv

    Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modal

  40. Xiang Liu, Zhaoxiang Liu, Huan Hu, Zezhou Chen

    While conversational generative AI has shown considerable potential in enhancing decision-making for agricultural professionals, its exploration has predominantly been anchored in text-based interactions. The evolution of multimodal conversational AI, leveraging vast amounts of image-text data from diverse sources, marks a significant stride forward. However

  41. Alessandra Filippi, Hiroyuki Fujioka, Takashi Higuchi, Luca Venturelli

    Extensive data of antiproton scattering cross sections with protons and nuclei have advanced our understanding of hadronic interactions with antinucleons. However, low-energy antineutron scattering data are scarce, thereby limiting our understanding of the S-wave antinucleon-nucleon and antinucleon-nucleus interactions. We present a novel production scheme o

  42. Gideon Yoffe, Keren Duer, Tom Andre Nordheim, Itay Halevy

    Europa, Jupiter's second Galilean moon, is believed to host a subsurface ocean in contact with a rocky mantle, where hydrothermal activity may drive the synthesis of organic molecules. Of these molecules, abiotic synthesis of aromatic amino acids is unlikely, and their detection on Europa could be considered a biosignature. Fluorescence from aromatic amino a

  43. Shogen Kawanami, Kento Iseri, Tomohiro I

    The Burrows-Wheeler Transform (BWT) of a string is an invertible permutation of the string, which can be used for data compression and compact indexes for string pattern matching. Ganguly et al. [SODA, 2017] introduced the parameterized BWT (pBWT) to design compact indexes for parameterized matching (p-matching), a variant of string pattern matching with par

  44. Nihat Mugurtay

    This article examines how unequal access to AI innovation creates systemic challenges for developing countries. Differential access to AI innovation results from the acute competition between domestic and global actors. While developing nations contribute significantly to AI development through data annotation labor, they face limited access to advanced AI t

  45. Erdem Sucu, İzzet Sakallı

    This study investigates the thermodynamic and quantum properties of Einstein-Power-Yang-Mills (EPYM) black holes in an Anti-de Sitter background, focusing on the effects of the nonlinear Yang-Mills charge parameter $\gamma$. We derive the metric function, analyze Hawking radiation through boson tunneling, and calculate thermodynamic properties including temp

  46. Kai Zhu, Michele Cappellari, Shude Mao, Shengdong Lu

    We derive circular velocity curves (CVCs) from stellar dynamical models for $\sim6000$ nearby galaxies in the final data release of the Sloan Digital Sky Survey-IV MaNGA survey with integral-field spectroscopy, exploring connections between the inner gravitational potential (traced by CVC amplitude/shape) and galaxy properties. The maximum circular velocity

  47. Anna Aksamit, Kaustav Das, Ivan Guo, Kihun Nam

    We consider a continuum of carbon-emitting firms who seek to maximise their stock price, and a regulator (e.g., Government) who wishes for the economy to flourish, whilst simultaneously punishing firms who behave non-green. Interpreting the regulator as a major player and the firms as the minor players, we model this setting through a mean field game with ma

  48. Guanghao Li, Mingzhi Chen, Hao Yu, Shuting Dong

    Deep learning-based denoising models have been widely employed in vision tasks, functioning as filters to eliminate noise while retaining crucial semantic information. Additionally, they play a vital role in defending against adversarial perturbations that threaten downstream tasks. However, these models can be intrinsically susceptible to adversarial attack

  49. Shining Wang, Yunlong Wang, Ruiqi Wu, Bingliang Jiao

    When discussing the Aerial-Ground Person Re-identification (AGPReID) task, we face the main challenge of the significant appearance variations caused by different viewpoints, making identity matching difficult. To address this issue, previous methods attempt to reduce the differences between viewpoints by critical attributes and decoupling the viewpoints. Wh

  50. Taeyong Ahn

    We investigate the intersection of positive closed currents in a general setting, employing tangent currents alongside King's residue formula. Our main result establishes a natural condition for the intersection--namely, the Dinh-Sibony product--of positive closed currents on domains and derives an integral representation of this intersection. In parallel, w

  51. Kyungho Kim, Sunwoo Kim, Geon Lee, Jinhong Jung

    Traditional recommender systems primarily rely on a single type of user-item interaction, such as item purchases or ratings, to predict user preferences. However, in real-world scenarios, users engage in a variety of behaviors, such as clicking on items or adding them to carts, offering richer insights into their interests. Multi-behavior recommender systems

  52. Zenghao Guan, Yucan Zhou, Xiaoyan Gu

    Traditional Federated Learning (FL) necessitates numerous rounds of communication between the server and clients, posing significant challenges including high communication costs, connection drop risks and susceptibility to privacy attacks. One-shot FL has become a compelling learning paradigm to overcome above drawbacks by enabling the training of a global

  53. Hyeong-Chan Kim, Wonwoo Lee

    We present a new rotating black hole solution to the Einstein equations as an extension of the Kerr spacetime. Interestingly, the solution we find may not be uniquely characterized by asymptotic parameters such as mass, angular momentum, and charge, thereby it would be the additional hair. We also analyze in detail how this additional characteristics or this

  54. Xin Wen, Bingchen Zhao, Yilun Chen, Jiangmiao Pang

    Pre-trained vision models (PVMs) are fundamental to modern robotics, yet their optimal configuration remains unclear. Through systematic evaluation, we find that while DINO and iBOT outperform MAE across visuomotor control and perception tasks, they struggle when trained on non-(single-)object-centric (NOC) data--a limitation strongly correlated with their d

  55. Junwei Yu, Yepeng Ding, Hiroyuki Sato

    The emergence of Large Language Models (LLMs) in Multi-Agent Systems (MAS) has opened new possibilities for artificial intelligence, yet current implementations face significant challenges in resource management, task coordination, and system efficiency. While existing frameworks demonstrate the potential of LLM-based agents in collaborative problem-solving,

  56. Millend Roy, Vaibhav Balloli, Anupam Sobti, Srinivasan Iyengar

    With increased global warming, there has been a significant emphasis to replace fossil fuel-dependent energy sources with clean, renewable sources. These new-age energy systems are becoming more complex with an increasing proportion of renewable energy sources (like solar and wind), energy storage systems (like batteries), and demand side control in the mix.

  57. Chihiro Oguri, Mao Shinoda

    We investigate the stability of maximizing measures for a penalty function of a two-dimensional subshift of finite type, building on the work of Gonschorowski et al. \cite{GQS}. In the one-dimensional case, such measures remain stable under Lipschitz perturbations for any subshift of finite type. However, instability arises for a penalty function of the Robi

  58. David Darrow, George Stepaniants

    This work aims to bridge the gap between pure and applied research on scalar, linear Volterra equations by examining five major classes: integral and integro-differential equations with completely monotone kernels, such as linear viscoelastic models; equations with positive definite kernels, such as partially observed quantum systems; difference equations wi

  59. Jian Jin, Zhenbo Yu, Yang Shen, Zhenyong Fu

    Customized text-to-image generation renders user-specified concepts into novel contexts based on textual prompts. Scaling the number of concepts in customized generation meets a broader demand for user creation, whereas existing methods face challenges with generation quality and computational efficiency. In this paper, we propose LaTexBlend, a novel framewo

  60. Zeyu Zhang, Yiran Wang, Wei Mao, Danning Li

    Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models lack a mechanism to prioritize dynamic frames and body parts based on given conditions. Second, existing methods for differ

  61. Xingye Fan, Zhongwen, Zhang, Yuri Boykov

    This paper demonstrates a surprising result for segmentation with image-level targets: extending binary class tags to approximate relative object-size distributions allows off-the-shelf architectures to solve the segmentation problem. A straightforward zero-avoiding KL-divergence loss for average predictions produces segmentation accuracy comparable to the s

  62. Shrutika Vishal Thengane, Marcel Bartholomeus Prasetyo, Yu Xiang Tan, Malika Meghjani

    Autonomous and targeted underwater visual monitoring and exploration using Autonomous Underwater Vehicles (AUVs) can be a challenging task due to both online and offline constraints. The online constraints comprise limited onboard storage capacity and communication bandwidth to the surface, whereas the offline constraints entail the time and effort required

  63. Indraneel Sinha, Saurav Sachin, Shreyashi Sinha, Roumita Roy

    Achieving atomically flat and stoichiometric films of chiral antiferromagnets (AFM) with two-dimensional kagome spin lattice structures are crucial for integrating these materials in both established and emerging antiferromagnetic spintronic devices. We report a systematic study of growth and anomalous Hall effect in (111)-oriented non-collinear AFM $Mn_{3+x

  64. Xinjie Zhao, Fan Gao, Xingyu Song, Yingjian Chen

    Recent advances in large language models (LLMs) have significantly improved multi-hop question answering (QA) through direct Chain-of-Thought (CoT) reasoning. However, the irreversible nature of CoT leads to error accumulation, making it challenging to correct mistakes in multi-hop reasoning. This paper introduces ReAgent: a Reversible multi-Agent collaborat

  65. Runqi Sui

    Retrieval-Augmented Generation (RAG) systems enhance response credibility and traceability by displaying reference contexts, but this transparency simultaneously introduces a novel black-box attack vector. Existing document poisoning attacks, where adversaries inject malicious documents into the knowledge base to manipulate RAG outputs, rely primarily on unr

  66. Haotian Chen, Yanyu Xu, Boyan Wang, Chaoyue Zhao

    In this report, we introduce our first-generation reasoning model, LexPro-1.0, a large language model designed for the highly specialized Chinese legal domain, offering comprehensive capabilities to meet diverse realistic needs. Existing legal LLMs face two primary challenges. Firstly, their design and evaluation are predominantly driven by computer science

  67. Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo

    Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial alignments, limiting detection performance. To address this problem, we propose a Large Language Model (LLM) guided Progressive feature Alignment Network called LPANet, which leverag

  68. Jiaxin Li, Hongxing Wang, Jiawei Tan, Zhilong Ou

    Understanding 3D object shapes necessitates shape representation by object parts abstracted from results of instance and semantic segmentation. Promising shape representations enable computers to interpret a shape with meaningful parts and identify their repeatability. However, supervised shape representations depend on costly annotation efforts, while curre

  69. Xu-Ke Gu, Li-Zhou Tan, Franco Nori, J. Q. You

    Markovian open quantum systems are governed by the Lindblad master equation where the dissipation contains two parts, i.e., the anti-Hermitian operator and the quantum jumps, which share a common dissipation rate. We generalize the Lindblad master equation via postselection to a generalized Liouvillian formalism in which the effective damping rate of the ant

  70. Junyan Lin, Feng Gap, Lin Qi, Junyu Dong

    Hyperspectral image (HSI) and LiDAR data joint classification is a challenging task. Existing multi-source remote sensing data classification methods often rely on human-designed frameworks for feature extraction, which heavily depend on expert knowledge. To address these limitations, we propose a novel Dynamic Cross-Modal Feature Interaction Network (DCMNet

  71. Yan Hu, Ahmad Chaddad

    This study introduces the SHAP-integrated convolutional diagnostic network (SICDN), an interpretable feature selection method designed for limited datasets, to address the challenge posed by data privacy regulations that restrict access to medical datasets. The SICDN model was tested on classification tasks using pneumonia and breast cancer datasets, demonst

  72. Zhiheng Yu, Jiancheng An, Lu Gan, Hongbin Li

    Reconfigurable intelligent surfaces (RIS) can reshape the characteristics of wireless channels by intelligently regulating the phase shifts of reflecting elements. Recently, various codebook schemes have been utilized to optimize the reflection coefficients (RCs); however, the selection of the optimal codeword is usually obtained by evaluating a metric of in

  73. Yuzhu Lei, Qiqi Xiao, Yinghui He, Guanding Yu

    In massive multi-input multi-output (MIMO) systems, the main bottlenecks of location- and orientation-assisted beam alignment using deep neural networks (DNNs) are large training overhead and significant performance degradation. This paper proposes a graph neural network (GNN)-based beam selection approach that reduces the training overhead and improves the

  74. Hao-Hao Li, Xin-zhe Zhang, Taotao Qiu, Jun-Qing Xia

    The James Webb Space Telescope (JWST) has observed massive galaxies at high redshifts, which implies an earlier epoch of reionization (EoR) compared with the cosmic microwave background (CMB) results. In this paper, based on \texttt{Planck 2020} (NPIPE release), \texttt{ACT DR4} and \texttt{SPT-3G} data, if assumed a Harrison-Zel'dovich (HZ) primordial power

  75. Jianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng

    Most research decoding brain signals into images, often using them as priors for generative models, has focused only on visual content. This overlooks the brain's natural ability to integrate auditory and visual information, for instance, sound strongly influences how we perceive visual scenes. To investigate this, we propose a new task of reconstructing con

  76. Andy Chia, Wai-Keong Mok, Leong-Chuan Kwek, Changsuk Noh

    Several important dynamical systems are in $\mathbb{R}^2$, defined by the pair of differential equations $(x',y')=(f(x,y),g(x,y))$. A question of fundamental importance is how such systems might behave quantum mechanically. In developing quantum theory, Dirac and others realized that classical Hamiltonian systems can be mapped to their quantum counterparts v

  77. Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian

    The increasing pace of population aging calls for better care and support systems. Falling is a frequent and critical problem for elderly people causing serious long-term health issues. Fall detection from video streams is not an attractive option for real-life applications due to privacy issues. Existing methods try to resolve this issue by using very low-r

  78. Ruimeng Liu, Xinhang Xu, Shenghai Yuan, Lihua Xie

    Zero-Shot Object Navigation (ZSON) requires agents to navigate to objects specified via open-ended natural language without predefined categories or prior environmental knowledge. While recent methods leverage foundation models or multi-modal maps, they often rely on 2D representations and greedy strategies or require additional training or modules with high

  79. Zhengyang Mei, Xiaohui Song, Xueyi Guo, Xiang Li

    In this paper, we introduce a method of using a double-layer resist lift-off process to prepare the capacitor dielectric layer for fabricating impedance-engineered Josephson parametric amplifiers (IMPAs). Compared with traditional techniques, this method enhances fabrication success rate, accelerates production. The IMPA we made experimentally achieves an in

  80. Zhijie Chen, Erjuan Fu, Chang-Shou Lin

    Let $E_{\tau}:=\mathbb{C}/(\mathbb{Z}+\mathbb{Z}\tau)$ with $\operatorname{Im}\tau>0$ be a flat torus and $G(z;\tau)$ be the Green function on $E_{\tau}$ with the singularity at $0$. Consider the multiple Green function $G_{n}$ on $(E_{\tau})^{n}$: \[ G_{n}(z_{1},\cdots,z_{n};\tau):=\sum_{i<j}G(z_{i}-z_{j};\tau)-n\sum_{i=1}% ^{n}G(z_{i};\tau). \] Recently, L

  81. Hanyu Zhou, Gim Hee Lee

    Large multimodal models (LMMs) excel in scene understanding but struggle with fine-grained spatiotemporal reasoning due to weak alignment between linguistic and visual representations. Existing methods map textual positions and durations into the visual space encoded from frame-based videos, but suffer from temporal sparsity that limits language-vision tempo

  82. W. B. Rui, Z. D. Wang

    Exceptional points (EPs) are prominent non-Hermitian band degeneracies that give rise to a variety of intriguing and unconventional phenomena. Similar to Weyl and Dirac points, EPs carry topological charges and comply with the celebrated fermion doubling theorems in lattices. Beyond these characteristics, EPs exhibit more exotic topological properties, parti

  83. Yuxing Wang, Allan Xi Chen, Matthew Salazar, Nawar Abdalla

    CR-39 solid-state nuclear track detectors are widely used in fusion research for detecting charged particles produced in fusion reactions. However, analyzing increasingly complex and large-scale CR-39 track images to extract meaningful information can be a tedious and time-consuming process, often prone to human errors and bias. To address these challenges,

  84. Zhao Tang, Fanhao Jia, Greis J. Kim-Reyes, Yabei Wu

    Color centers exhibiting deep-level states within the wide bandgap h-BN monolayer possess substantial potential for quantum applications. Uncovering precise geometric characteristics at the atomic scale is crucial for understanding defect performance. In this study, first-principles calculations were performed on the most extensively investigated CBVN and NB

  85. Ning Ding, Jing Han, Yuchuan Tian, Chao Xu

    Diffusion Transformer (DiT) has now become the preferred choice for building image generation models due to its great generation capability. Unlike previous convolution-based UNet models, DiT is purely composed of a stack of transformer blocks, which renders DiT excellent in scalability like large language models. However, the growing model size and multi-st

  86. Yanlong Wang, Jian Xu, Shao-Lun Huang, Danny Dongning Sun

    This study seeks to advance the understanding and prediction of stock market return uncertainty through the application of advanced deep learning techniques. We introduce a novel deep learning model that utilizes a Gaussian mixture distribution to capture the complex, time-varying nature of asset return distributions in the Chinese stock market. By incorpora

  87. Yanlong Wang, Jian Xu, Tiantian Gao, Hongkang Zhang

    Despite the growing attention to time series forecasting in recent years, many studies have proposed various solutions to address the challenges encountered in time series prediction, aiming to improve forecasting performance. However, effectively applying these time series forecasting models to the field of financial asset pricing remains a challenging issu

  88. Yuchen Han, Yucheng Wu, Jeffrey Willard

    This paper investigates a critical aspect of large language model (LLM) performance: the optimal formatting of classification task options in prompts. Through an extensive experimental study, we compared two selection formats -- bullet points and plain English -- to determine their impact on model performance. Our findings suggest that presenting options via

  89. Yash Makwana, Anupama Panigrahi, Saibal K. Pal

    Recent studies have explored DNA-based algorithms for IoT security and image encryption. A similar encryption algorithm was proposed by Al-Husainy et. al. in 2021 Recently, Al-Husainy et al.in 2021, proposed an encryption algorithm based on DNA processes for Internet of Things(IoT) applications. Upon finding low avalanche effect in our experiments, we first

  90. Michael McGuire

    Automatic speech recognition (ASR) has been an essential component of computer assisted language learning (CALL) and computer assisted language testing (CALT) for many years. As this technology continues to develop rapidly, it is important to evaluate the accuracy of current ASR systems for language learning applications. This study assesses five cutting-edg

  91. Jiacheng Liu, Chang Zou, Yuanhuiyi Lyu, Junjie Chen

    Diffusion Transformers (DiT) have revolutionized high-fidelity image and video synthesis, yet their computational demands remain prohibitive for real-time applications. To solve this problem, feature caching has been proposed to accelerate diffusion models by caching the features in the previous timesteps and then reusing them in the following timesteps. How

  92. Hao Chen, Jian Chen, Xinran Liu, Zihui Zhang

    Continuum robots offer high flexibility and multiple degrees of freedom, making them ideal for navigating narrow lumens. However, accurately modeling their behavior under large deformations and frequent environmental contacts remains challenging. Current methods for solving the deformation of these robots, such as the Model Order Reduction and Gauss-Seidel (

  93. Youngeun Kim, Seunghwan Lee, Aecheon Jung, Bogon Ryu

    Model merging enables efficient multi-task models by combining task-specific fine-tuned checkpoints. However, storing multiple task-specific checkpoints requires significant memory, limiting scalability and restricting model merging to larger models and diverse tasks. In this paper, we propose quantizing task vectors (i.e., the difference between pre-trained

  94. Chengzhi Lin, Chuyuan Wang, Annan Xie, Wuhong Wang

    In video recommendation systems, user behaviors such as watch time, likes, and follows are commonly used to infer user interest. However, these behaviors are influenced by various biases, including duration bias, demographic biases, and content category biases, which obscure true user preferences. In this paper, we hypothesize that biases and user interest a

  95. Weidong Guo, Hantao Zhang, Shouhong Wan, Bingbing Zou

    Lesion synthesis methods have made significant progress in generating large-scale synthetic datasets. However, existing approaches predominantly focus on texture synthesis and often fail to accurately model masks for anatomically complex lesions. Additionally, these methods typically lack precise control over the synthesis process. For example, perirectal ly

  96. Lingrui Ge, Yiqian Wang, Jiahao Xu

    This paper solves ``The Dry Ten Martini Problem'' for $C^2$ cosine-type quasiperiodic Schr\"odinger operators with large coupling constants and Diophantine frequencies, a model originally introduced by Sinai in 1987 \cite{sinai}. This shows that the analyticity assumption on the potential is not essential for obtaining a dry Cantor spectrum and can be replac

  97. Pranjal Awasthi, Sreenivas Gollapudi, Ravi Kumar, Kamesh Munagala

    We study zeroth-order optimization where solutions must minimize a cost $d(s)$ while maintaining high probability under a complex generative prior $L(s)$ (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to $L(s) e^{-T \cdot d(s)}$. Since classical model-based optimization (MBO) lacks finite-sample guarantees for

  98. Shanshan Yan, Zexi Li, Chao Wu, Meng Pang

    Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic sim

  99. Yuting Huang, Simon S. Toedtli, Gregory P. Chini, Beverley J. McKeon

    The quadratic convection term in the incompressible Navier-Stokes equations is considered as a nonlinear forcing to the linear resolvent operator, and it is studied in the Fourier domain through the analysis of interactions between triadically compatible wavenumber-frequency triplets. A framework to quantify the triadic contributions to the forcing and respo

  100. Dung Xuan Nguyen, Dam Thanh Son

    A low-energy neutral quasiparticle in a fractional quantum Hall system appears in the latter's energy spectrum on a sphere as a series of many-body excited states labeled by the angular momentum $L$ and whose energy is a smooth function of $L$ in the limit of large sphere radius. We argue that the signature of a nonvanishing spin (intrinsic angular momentum)