Skip to content

May 2024 arXiv papers — page 61

Showing 6,0016,100 of 20,894 papers

  1. Yu Zhe, Rei Nagaike, Daiki Nishiyama, Kazuto Fukuchi

    Deep learning models are susceptible to adversarial attacks, where slight perturbations to input data lead to misclassification. Adversarial attacks become increasingly effective with access to information about the targeted classifier. In the context of multi-task learning, where a single model learns multiple tasks simultaneously, attackers may aim to expl

  2. Guy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Lior Shani

    We propose a novel approach for training large language models (LLMs) to adhere to objectives defined within a latent embedding space. Our method leverages reinforcement learning (RL), treating a pre-trained LLM as an environment. Our embedding-aligned guided language (EAGLE) agent is trained to iteratively steer the LLM's generation towards optimal regions

  3. Neehar Kondapaneni, Markus Marks, Oisin Mac Aodha, Pietro Perona

    We introduce Discovering Conceptual Network Explanations (DCNE), a new approach for generating human-comprehensible visual explanations to enhance the interpretability of deep neural image classifiers. Our method automatically finds visual explanations that are critical for discriminating between classes. This is achieved by simultaneously optimizing three c

  4. Susan Ellul, Stijn Vansteelandt, John B. Carlin, Margarita Moreno-Betancur

    Observational epidemiological studies commonly seek to estimate the causal effect of an exposure on an outcome. Adjustment for potential confounding bias in modern studies is challenging due to the presence of high-dimensional confounding, which occurs when there are many confounders relative to sample size or complex relationships between continuous confoun

  5. Jia He, Bonan Li, Ge Yang, Ziwen Liu

    Solving 3D medical inverse problems such as image restoration and reconstruction is crucial in modern medical field. However, the curse of dimensionality in 3D medical data leads mainstream volume-wise methods to suffer from high resource consumption and challenges models to successfully capture the natural distribution, resulting in inevitable volume incons

  6. Peng Kuang, Zhibo Wang, Zhixuan Chu, Jingyi Wang

    Spurious correlations in training data significantly hinder the generalization capability of machine learning models when faced with distribution shifts, leading to the proposition of numberous debiasing methods. However, it remains to be asked: \textit{Do existing benchmarks for debiasing really represent biases in the real world?} Recent works attempt to a

  7. Yuankun Yang, Li Zhang, Ziyang Xie, Zhiyuan Yuan

    Understanding the hidden mechanisms behind human's visual perception is a fundamental question in neuroscience. To that end, investigating into the neural responses of human mind activities, such as functional Magnetic Resonance Imaging (fMRI), has been a significant research vehicle. However, analyzing fMRI signals is challenging, costly, daunting, and dema

  8. Nabil Ibtehaz, Masood Mortazavi

    Electrocardiogram (ECG) signals, profiling the electrical activities of the heart, are used for a plethora of diagnostic applications. However, ECG systems require multiple leads or channels of signals to capture the complete view of the cardiac system, which limits their application in smartwatches and wearables. In this work, we propose a modally reduced r

  9. Oskar A. Sultanov

    Non-autonomous perturbations of isochronous systems in the plane are considered. It is assumed that the intensity of perturbations decays with time, and the frequency is asymptotically constant with the limiting value satisfying a resonance condition. We discuss the emergence of attracting resonant solutions with an asymptotically constant amplitude. By comb

  10. Yihang Wang, Yuying Qiu, Peng Chen, Kai Zhao

    With the growing availability of multi-domain time series data, there is an increasing demand for general forecasting models pre-trained on multi-source datasets to support diverse downstream prediction scenarios. Existing time series foundation models primarily focus on scaling up pre-training datasets and model sizes to enhance generalization performance.

  11. Christophe H. Valahu, Tomas Navickas, Michael J. Biercuk, Ting Rei Tan

    Bosonic modes are prevalent in all aspects of quantum information processing. However, existing tools for characterizing the quality, stability, and noise properties of bosonic modes are limited, especially in a driven setting. Here, we propose, demonstrate, and analyze a bosonic randomized benchmarking (BRB) protocol that uses randomized displacements of th

  12. Alvin Gonzales, Daniel Dilley, Bikun Li, Liang Jiang

    We apply the quantum error detection scheme Pauli check sandwiching (PCS) to quantum networks by turning it into a distributed multiparty protocol. PCS provides protection on the targeted qubits and generally requires less resource overhead than standard quantum error correction and detection codes. We provide analytical equations for the final fidelity and

  13. Nicolás Andruskiewitsch, Olivier Mathieu

    It is shown that if the universal enveloping algebra of a simple $\mathbb Z^n$-graded Lie algebra is Noetherian, then the Lie algebra is finite-dimensional.

  14. Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang

    Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adve

  15. Yonglong Ding

    Lattice models exhibit significant potential in investigating phase transitions, yet they encounter numerous computational challenges. To address these issues, this study introduces a Monte Carlo-based approach that transforms lattice models into a network model with intricate inter-node correlations. This framework enables a profound analysis of Ising, JQ,

  16. Run Luo, Yunshui Li, Longze Chen, Wanwei He

    The development of large language models (LLMs) has significantly advanced the emergence of large multimodal models (LMMs). While LMMs have achieved tremendous success by promoting the synergy between multimodal comprehension and creation, they often face challenges when confronted with out-of-distribution data, such as which can hardly distinguish orientati

  17. Fei Teng, Haoyang Li, Shimin Di, Lei Chen

    Cardinality Estimation (CE) for query is to estimate the number of results without execution, which is an effective index in query optimization. Recently, CE for queries over knowlege graph (KGs) with triple facts has achieved great success. To more precisely represent facts, current researchers propose hyper-relational KGs (HKGs) to represent a triple fact

  18. Long Tan Le, Han Shu, Tung-Anh Nguyen, Choong Seon Hong

    While astonishingly capable, large Language Models (LLM) can sometimes produce outputs that deviate from human expectations. Such deviations necessitate an alignment phase to prevent disseminating untruthful, toxic, or biased information. Traditional alignment methods based on reinforcement learning often struggle with the identified instability, whereas pre

  19. Massine Kelai, Stefano Reale, Roberto Robles, Jaehyun Lee

    Surface-adsorbed rare-earth nanostructures are ideal platforms to investigate the interplay between intra-atomic interactions and multi-orbital spin configurations. However, addressing these properties has posed severe experimental and theoretical challenges. Here, we use the orbital selectivity offered by X-ray absorption spectroscopy to quantify the Coulom

  20. Zhongnian Li, Jinghao Xu, Peng Ying, Meng Wei

    Pre-trained Vision-Language Models (VLMs) exhibit strong zero-shot classification abilities, demonstrating great potential for generating weakly supervised labels. Unfortunately, existing weakly supervised learning methods are short of ability in generating accurate labels via VLMs. In this paper, we propose a novel weakly supervised labeling setting, namely

  21. Adam Dai, Shubh Gupta, Grace Gao

    This work introduces Neural Elevations Models (NEMos), which adapt Neural Radiance Fields to a 2.5D continuous and differentiable terrain model. In contrast to traditional terrain representations such as digital elevation models, NEMos can be readily generated from imagery, a low-cost data source, and provide a lightweight representation of terrain through a

  22. Dominic Robe, Adrian Menzel, Andrew W Phillips, Elnaz Hajizadeh

    In this work, methods are presented to automatically generate a fully atomistic LAMMPS models of arbitrary linear multiblock polyurethane copolymers. The routine detailed here receives as parameters the number of repeat units per hard block, the number of units in a soft block, and the number of soft blocks per chain, as well as chemical formulae of three mo

  23. Yajing Liu, Shijun Zhou, Xiyao Liu, Chunhui Hao

    Single-source domain generalization (SDG) for object detection is a challenging yet essential task as the distribution bias of the unseen domain degrades the algorithm performance significantly. However, existing methods attempt to extract domain-invariant features, neglecting that the biased data leads the network to learn biased features that are non-causa

  24. Yair Litman, Venkat Kapil, Yotam M. Y. Feldman, Davide Tisi

    Atomic-scale simulations have progressed tremendously over the past decade, largely due to the availability of machine-learning interatomic potentials. These potentials combine the accuracy of electronic structure calculations with the ability to reach extensive length and time scales. The i-PI package facilitates integrating the latest developments in this

  25. Jialong Wu, Shaofeng Yin, Ningya Feng, Xu He

    World models empower model-based agents to interactively explore, reason, and plan within imagined environments for real-world decision-making. However, the high demand for interactivity poses challenges in harnessing recent advancements in video generative models for developing world models at scale. This work introduces Interactive VideoGPT (iVideoGPT), a

  26. Yanwei Zheng, Changrui Li, Chuanlin Lan, Yaling Li

    Zero-shot object navigation (ZSON) addresses situation where an agent navigates to an unseen object that does not present in the training set. Previous works mainly train agent using seen objects with known labels, and ignore the seen objects without labels. In this paper, we introduce seen objects without labels, herein termed as ``unknown objects'', into t

  27. William Misener, Matthäus Schulik, Hilke E. Schlichting, James E. Owen

    The mass loss rates of planets undergoing core-powered escape are usually modeled using an isothermal Parker-type wind at the equilibrium temperature, $T_\mathrm{eq}$. However, the upper atmospheres of sub-Neptunes may not be isothermal if there are significant differences between the opacity to incident visible and outgoing infrared radiation. We model bolo

  28. Yue-Mei Sun, Xin-Yu Wang, Liang-Jun Zhai

    In this paper, we study the critical behaviors in the non-Hermitian disorder Aubry-Andr\'{e} (DAA) model, and we assume the non-Hermiticity is introduced by nonreciprocal hopping. We employ the localization length $\xi$, the inverse participation ratio ($\rm IPR$), and the energy gap $\Delta E$ as the characteristic quantities to describe the critical proper

  29. David Senjaya, Piyabut Burikham, Tiberiu Harko

    We consider Klein-Gordon equation in the Dyonic Kerr-Sen black hole background, which is the charged rotating axially symmetric solution of the Einstein-Maxwell-Dilaton-Axion theory of gravity. The black hole incorporates electric, magnetic, dilatonic and axionic charges and is constructed in 3+1 dimensional spacetime. We begin our investigations with the co

  30. Siddhartha Shankar Das, S M Ferdous, Mahantesh M Halappanavar, Edoardo Serra

    We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend

  31. Vikas Thamizharasan, Difan Liu, Matthew Fisher, Nanxuan Zhao

    The success of denoising diffusion models in representing rich data distributions over 2D raster images has prompted research on extending them to other data representations, such as vector graphics. Unfortunately due to their variable structure and scarcity of vector training data, directly applying diffusion models on this domain remains a challenging prob

  32. Zijin Gu, Tatiana Likhomanenko, He Bai, Erik McDermott

    Language models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models (LLMs) have been applied to ASR correction, but introduce latency and hallucination concerns. We revisit ASR error correction with compact seq2seq models, trained on ASR errors fro

  33. Nobukazu Kowaki

    A sequence of combinatorial mutations of matching field polytopes preserves the property of giving rise to a toric degeneration of Grassmannians. In this paper, we find a way to check that two matching field polytopes are combinatorial mutation equivalence using tropical hyperplane arrangements, ``literally at a glance". Our way can prove that block diagonal

  34. Qingdong He, Jiangning Zhang, Jinlong Peng, Haoyang He

    Transformers have revolutionized the point cloud learning task, but the quadratic complexity hinders its extension to long sequence and makes a burden on limited computational resources. The recent advent of RWKV, a fresh breed of deep sequence models, has shown immense potential for sequence modeling in NLP tasks. In this paper, we present PointRWKV, a mode

  35. Chia-Tung Ho, Haoxing Ren

    Standard cells are essential components of modern digital circuit designs. With process technologies advancing toward 2nm, more routability issues have arisen due to the decreasing number of routing tracks, increasing number and complexity of design rules, and strict patterning rules. The state-of-the-art standard cell design automation framework is able to

  36. Sheng Yue, Xingyuan Hua, Ju Ren, Sen Lin

    In this paper, we study offline-to-online Imitation Learning (IL) that pretrains an imitation policy from static demonstration data, followed by fast finetuning with minimal environmental interaction. We find the na\"ive combination of existing offline IL and online IL methods tends to behave poorly in this context, because the initial discriminator (often u

  37. Sheng Yue, Jiani Liu, Xingyuan Hua, Ju Ren

    Offline Imitation Learning (IL) with imperfect demonstrations has garnered increasing attention owing to the scarcity of expert data in many real-world domains. A fundamental problem in this scenario is how to extract positive behaviors from noisy data. In general, current approaches to the problem select data building on state-action similarity to given exp

  38. Vibhavasu Pasumarti, Shantanu Desai

    We search for gamma-ray emission from OJ287 in the energy range from 0.1-300 GeV during 2015-2023, in coincidence with an extensive observing campaign to monitor the optical flux variability and polarization, as discussed in arXiv:2311.02372. We present results for eight segments in the aforementioned period, with each segment corresponding to an observing s

  39. Ziqing Ding, Ercai Chen, Xiaoyao Zhou

    In this paper, we first prove the variational principle for amenable packing topological pressure. Then we obtain an inequality concerning amenable packing pressure for factor maps. Finally, we show that the equality about packing topological pressure of the set of generic points when the system satisfies the almost specification property, or $\mu$ is ergodi

  40. Xiaoqun Liu, Jiacheng Liang, Muchao Ye, Zhaohan Xi

    Large language models (LLMs) are vulnerable when trained on datasets containing harmful content, which leads to potential jailbreaking attacks in two scenarios: the integration of harmful texts within crowdsourced data used for pre-training and direct tampering with LLMs through fine-tuning. In both scenarios, adversaries can compromise the safety alignment

  41. Olena Burda-Lassen, Aman Chadha, Shashank Goswami, Vinija Jain

    An image is often considered worth a thousand words, and certain images can tell rich and insightful stories. Can these stories be told via image captioning? Images from folklore genres, such as mythology, folk dance, cultural signs, and symbols, are vital to every culture. Our research compares the performance of four popular vision-language models (GPT-4V,

  42. Wei Xia, Aile Wang, Jian Yuan, Jiawei Luo

    The ferrimagnet TbMn6Sn6 has attracted vast attention, because its pristine Mn kagome lattice with strong spin-orbit coupling and out-of-plane Tb-Mn exchange supports quantum-limit Chern topological magnetism which can be described by the simple spinless Haldane model. We unveil herein that engineering the kagome lattice through partial substitution of Mn wi

  43. Sami Arja, Alexandre Marcireau, Saeed Afshar, Bharath Ramesh

    Aerial surveillance demands rapid and precise detection of moving objects in dynamic environments. Event cameras, which draw inspiration from biological vision systems, present a promising alternative to frame-based sensors due to their exceptional temporal resolution, superior dynamic range, and minimal power requirements. Unlike traditional frame-based sen

  44. Chenxi Sun, Hongzhi Zhang, Zijia Lin, Jingyuan Zhang

    Large language models have demonstrated exceptional capability in natural language understanding and generation. However, their generation speed is limited by the inherently sequential nature of their decoding process, posing challenges for real-time applications. This paper introduces Lexical Unit Decoding (LUD), a novel decoding methodology implemented in

  45. Nick Brettell, James Oxley, Charles Semple, Geoff Whittle

    A partitioned matroid $(M, \{X_1,X_2,\dots,X_n\})$ consists of a matroid $M$ and a partition $\{X_1,X_2,\dots,X_n\}$ of its ground set. As such structures arise frequently in structural matroid theory, this paper introduces a general technique for analyzing those special properties of partitioned matroids that depend solely on the values of the connectivitie

  46. Kevin S. Chen, Ying-Jen Yang

    The characterization of network and biophysical properties from neural spiking activity is an important goal in neuroscience. A framework that provides unbiased inference on causal synaptic interaction and single neural properties has been missing. Here we applied the stochastic dynamics extension of Maximum Entropy -- the Maximum Caliber Principle -- to inf

  47. Sheng Yue, Zerui Qin, Xingyuan Hua, Yongheng Deng

    Federated Reinforcement Learning (FRL) has been deemed as a promising solution for intelligent decision-making in the era of Artificial Internet of Things. However, existing FRL approaches often entail repeated interactions with the environment during local updating, which can be prohibitively expensive or even infeasible in many real-world domains. To overc

  48. Zhigao Cai, Xing-Ming Zhao

    Automatic segmentation of the fetal brain is still challenging due to the health state of fetal development, motion artifacts, and variability across gestational ages, since existing methods rely on high-quality datasets of healthy fetuses. In this work, we propose a novel cascade network called CasUNext to enhance the accuracy and generalization of fetal br

  49. Youjin Sung, Youngjin Han, Yang Liu

    Assessing fit in common factor models solely through the lens of mean and covariance structures, as is commonly done with conventional goodness-of-fit (GOF) assessments, may overlook critical aspects of misfit, potentially leading to misleading conclusions. To achieve more flexible fit assessment, we extend the theory of generalized residuals (Haberman & Sin

  50. Hyungtae Lee, Yan Zhang, Yi-Ting Shen, Heesung Kwon

    Aerial-view human detection has a large demand for large-scale data to capture more diverse human appearances compared to ground-view human detection. Therefore, synthetic data can be a good resource to expand data, but the domain gap with real-world data is the biggest obstacle to its use in training. As a common solution to deal with the domain gap, the si

  51. Yu Fu, Wen Xiao, Jia Chen, Jiachen Li

    Recent studies reveal that Large Language Models (LLMs) face challenges in balancing safety with utility, particularly when processing long texts for NLP tasks like summarization and translation. Despite defenses against malicious short questions, the ability of LLMs to safely handle dangerous long content, such as manuals teaching illicit activities, remain

  52. John Chiang

    Homomorphic encryption (HE) is a promising technique used for privacy-preserving computation. Since HE schemes only support primitive polynomial operations, homomorphic evaluation of polynomial approximations for non-polynomial functions plays an important role in privacy-preserving machine learning. In this paper, we introduce a simple solution to approxima

  53. Jie Bian, Vincent Y. F. Tan

    The Indexed Minimum Empirical Divergence (IMED) algorithm is a highly effective approach that offers a stronger theoretical guarantee of the asymptotic optimality compared to the Kullback--Leibler Upper Confidence Bound (KL-UCB) algorithm for the multi-armed bandit problem. Additionally, it has been observed to empirically outperform UCB-based algorithms and

  54. Jingyuan Zhu, Shiyu Li, Yuxuan Liu, Ping Huang

    Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain limited. This paper presents ODGEN, a novel method to gener

  55. Lianming Huang, Shangyu Wu, Yufei Cui, Ying Xiong

    Deploying large language model inference remains challenging due to their high computational overhead. Early exit optimizes model inference by adaptively reducing the number of inference layers. Current methods typically train internal classifiers or use heuristic methods to determine the exit layer. However, those methods either introduce significant traini

  56. Qiang Zou, Yunzhu Gao

    Lattice structures have been widely used in applications due to their superior mechanical properties. To fabricate such structures, a geometric processing step called triangulation is often employed to transform them into the STL format before sending them to 3D printers. Because lattice structures tend to have high geometric complexity, this step usually ge

  57. Haoxuan Qu, Zhuoling Li, Hossein Rahmani, Yujun Cai

    Recently, Gaussian Splatting, a method that represents a 3D scene as a collection of Gaussian distributions, has gained significant attention in addressing the task of novel view synthesis. In this paper, we highlight a fundamental limitation of Gaussian Splatting: its inability to accurately render discontinuities and boundaries in images due to the continu

  58. Yuta Takada

    We show that there exists an automorphism of a projective K3 surface with Picard number $2$ such that the trace of its action on the Picard lattice is $3$. Together with a result of K. Hashimoto, J. Keum and K. Lee, we determine the set of dynamical degrees of automorphisms of projective K3 surfaces with Picard number $2$.

  59. Siddhant Bhambri, Amrita Bhattacharjee, Durgesh Kalwar, Lin Guan

    Reinforcement Learning (RL) suffers from sample inefficiency in sparse reward domains, and the problem is further pronounced in case of stochastic transitions. To improve the sample efficiency, reward shaping is a well-studied approach to introduce intrinsic rewards that can help the RL agent converge to an optimal policy faster. However, designing a useful

  60. Zhuochen Fan, Yalun Cai, Zirui Liu, Jiarui Guo

    Graphs play an increasingly important role in various big data applications. However, existing graph data structures cannot simultaneously address the performance bottlenecks caused by the dynamic updates, large scale, and high query complexity of current graphs. This paper proposes a novel data structure for large-scale dynamic graphs called CuckooGraph. It

  61. Lingling Chen, Mikyoung Jun, Scott J. Cook

    Spatial point process models are widely applied to point pattern data from various applications in the social and environmental sciences. However, a serious hurdle in fitting point process models is the presence of duplicated points, wherein multiple observations share identical spatial coordinates. This often occurs because of decisions made in the geo-codi

  62. Wenxiang Pei, Qi Guo, Shi Shao, Yi He

    The stellar-to-halo mass relation (SHMR) is a fundamental relationship between galaxies and their host dark matter haloes. In this study, we examine the scatter in this relation for primary galaxies in the semi-analytic L-Galaxies model and two cosmological hydrodynamical simulations, \eagle{} and \tng{}. We find that in low-mass haloes, more massive galaxie

  63. Marie Al Ghossein, Ching-Wei Chen, Jason Tang

    Recent advances in the fields of Information Retrieval and Machine Learning have focused on improving the performance of search engines to enhance the user experience, especially in the world of online shopping. The focus has thus been on leveraging cutting-edge learning techniques and relying on large enriched datasets. This paper introduces the Shopping Qu

  64. Dong Huang, Jianbo Dai, Han Weng, Puzhen Wu

    Large language models (LLMs) have shown remarkable progress in code generation, but their generated code often suffers from inefficiency, resulting in longer execution times and higher memory consumption. To address this issue, we propose \textbf{EffiLearner}, a self-optimization framework that utilizes execution overhead profiles to improve the efficiency o

  65. Bingchen Yang, Haiyong Jiang, Hao Pan, Peter Wonka

    Reverse engineering CAD models from raw geometry is a classic but challenging research problem. In particular, reconstructing the CAD modeling sequence from point clouds provides great interpretability and convenience for editing. To improve upon this problem, we introduce geometric guidance into the reconstruction network. Our proposed model, PS-CAD, recons

  66. Tian Liu, Bo Sun, Danny H. K. Tsang

    With the increasing penetration of intermittent renewable energy sources (RESs), it becomes increasingly challenging to maintain the supply-demand balance of power systems by solely relying on the generation side. To combat the volatility led by the uncertain RESs, demand-side management by leveraging the multi-dimensional flexibility (MDF) has been recogniz

  67. Giorgos Anastasiou, Ignacio J. Araya, Daniel Ávila, Alberto Guijosa

    In the context of the holographic correspondence, we introduce a purely extrinsic renormalization prescription, exemplified with the case of a minimally-coupled scalar field in AdS space. The counterterms depend only on the field and its radial derivatives. This would seem to conflict with the Dirichlet variational principle, but we show that consistency fol

  68. Zhisheng Tang, Ke Shen, Mayank Kejriwal

    Words of estimative probability (WEPs), such as ''maybe'' or ''probably not'' are ubiquitous in natural language for communicating estimative uncertainty, compared with direct statements involving numerical probability. Human estimative uncertainty, and its calibration with numerical estimates, has long been an area of study -- including by intelligence agen

  69. Amin Sarihi, Peter Jamieson, Ahmad Patooghy, Abdel-Hameed A. Badawy

    The Hardware Trojan (HT) problem can be thought of as a continuous game between attackers and defenders, each striving to outsmart the other by leveraging any available means for an advantage. Machine Learning (ML) has recently played a key role in advancing HT research. Various novel techniques, such as Reinforcement Learning (RL) and Graph Neural Networks

  70. Valeria Rodríguez-Fajardo, Gabriela Flores-Cova, Carmelo Rosales-Guzmán, Benjamin Perez-Garcia

    In this manuscript, we put forward two new types of structured light beams, the vortex Pearcey-Gauss (VPeG) beam, with a homogeneous polarisation distribution, and the vector vortex Pearcey-Gauss (VVPeG) beam, with a non-homogeneous polarisation distribution. The later generated as a non-separable superposition of the spatial and polarisation degrees of free

  71. Peihua Mai, Ran Yan, Yan Pang

    Federated learning (FL) allows multiple devices to train a model collaboratively without sharing their data. Despite its benefits, FL is vulnerable to privacy leakage and poisoning attacks. To address the privacy concern, secure aggregation (SecAgg) is often used to obtain the aggregation of gradients on sever without inspecting individual user updates. Unfo

  72. Hidekazu Yoshioka

    Logit dynamics are dynamical systems describing transitions and equilibria of actions of interacting players under uncertainty. An uncertainty is embodied in logit dynamic as a softmax type function often called a logit function originating from a maximization problem subjected to an entropic penalization. This study provides another explanation for the gene

  73. Yang Li, Shaobo Han, Shihao Ji

    As the adoption of large language models increases and the need for per-user or per-task model customization grows, the parameter-efficient fine-tuning (PEFT) methods, such as low-rank adaptation (LoRA) and its variants, incur substantial storage and transmission costs. To further reduce stored parameters, we introduce a "divide-and-share" paradigm that brea

  74. Tao Zou, Yuhao Mao, Junchen Ye, Bowen Du

    Dynamic graph learning equips the edges with time attributes and allows multiple links between two nodes, which is a crucial technology for understanding evolving data scenarios like traffic prediction and recommendation systems. Existing works obtain the evolving patterns mainly depending on the most recent neighbor sequences. However, we argue that whether

  75. Moh. Kamalul Wafi, Milad Siami

    This paper addresses the challenge of network synchronization under limited communication, involving heterogeneous agents with different dynamics and various network topologies, to achieve consensus. We investigate the distributed adaptive control for interconnected unknown linear subsystems with a leader and followers, in the presence of input-output distur

  76. Kai Huang, Haoming Wang, Wei Gao

    Text-to-image diffusion models can be fine-tuned in custom domains to adapt to specific user preferences, but such adaptability has also been utilized for illegal purposes, such as forging public figures' portraits, duplicating copyrighted artworks and generating explicit contents. Existing work focused on detecting the illegally generated contents, but cann

  77. Sheng Yue, Xingyuan Hua, Lili Chen, Ju Ren

    Federated Reinforcement Learning (FRL) has garnered increasing attention recently. However, due to the intrinsic spatio-temporal non-stationarity of data distributions, the current approaches typically suffer from high interaction and communication costs. In this paper, we introduce a new FRL algorithm, named $\texttt{MFPO}$, that utilizes momentum, importan

  78. Yinuo Wang, Likun Wang, Yuxuan Jiang, Wenjun Zou

    Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to acquire complex policies. In response to this problem, we pr

  79. Pan Liao, Feng Yang, Di Wu, Wenhui Zhao

    Monocular 3D object detection has vast application potential across various fields. DETR-type models have shown remarkable performance in different areas, but there is still considerable room for improvement in monocular 3D detection, especially with the existing DETR-based method, MonoDETR. After addressing the query initialization issues in MonoDETR, we ex

  80. Jake McNaughton

    In this dissertation we study basic local differential geometry, projective differential geometry, and prolongations of overdetermined geometric partial differential equations. It is simple to prolong an n-th order linear ordinary differential equation into n first order equations. For partial differential equations there is a related process but it is far m

  81. Kohei Oshio, Yohichi Suzuki, Kaito Wada, Keigo Hisanaga

    In quantum computation, amplitude estimation is a fundamental subroutine that is utilized in various quantum algorithms. A general important task of such estimation problems is to characterize the estimation lower bound, which is referred to as quantum Cram\'er-Rao bound (QCRB), and to construct an optimal estimator that achieves QCRB. This paper studies the

  82. Yanshu Wang, Wenyang He, Tong Yang

    Large Language Models (LLMs) have significantly advanced natural language processing tasks such as machine translation, text generation, and sentiment analysis. However, their large size, often consisting of billions of parameters, poses challenges for storage, computation, and deployment, particularly in resource-constrained environments like mobile devices

  83. Xinan He, Yue Zhou, Shu Hu, Bin Li

    Detecting falsified faces generated by Deepfake technology is essential for safeguarding trust in digital communication and protecting individuals. However, current detectors often suffer from a dual-overfitting: they become overly specialized in both specific forgery fingerprints and particular demographic attributes. Critically, most existing methods overl

  84. Shengrong Wang, Pengfei Guo, Jingshi Xu

    In this paper, we give a approximation characterization, embedding properties and the duality of matrix weighted modulation spaces.

  85. Song Wang, Jiawei Yu, Wentong Li, Hao Shi

    Semantic scene completion aims to infer the 3D geometric structures with semantic classes from camera or LiDAR, which provide essential occupancy information in autonomous driving. Prior endeavors concentrate on constructing the network or benchmark in a fully supervised manner. While the dense occupancy grids need point-wise semantic annotations, which incu

  86. Weize Li, Zhicheng Zhao, Haochen Bai, Fei Su

    Referring Expression Segmentation (RES) has attracted rising attention, aiming to identify and segment objects based on natural language expressions. While substantial progress has been made in RES, the emergence of Generalized Referring Expression Segmentation (GRES) introduces new challenges by allowing expressions to describe multiple objects or lack spec

  87. Chanyong Park, Hanse Kim, Kyungchan Cho

    By using the braneworld model, we investigate the time evolution of microscopic and macroscopic correlations in expanding universes. To describe the FLRW cosmologies in the holographic setup, we take into account a braneworld moving in the $p$-brane gas geometry, where the radial motion of the braneworld determines the cosmology in the braneworld. We show th

  88. Ryan Thompson, Edwin V. Bonilla, Robert Kohn

    Directed acyclic graph (DAG) learning is a central task in structure discovery and causal inference. Although the field has witnessed remarkable advances over the past few years, it remains statistically and computationally challenging to learn a single (point estimate) DAG from data, let alone provide uncertainty quantification. We address the difficult tas

  89. Kyle M Sherbert, Hisham Amer, Sophia E Economou, Edwin Barnes

    In conventional variational quantum eigensolvers (VQEs), trial states are prepared by applying series of parameterized gates to a reference state, with the gate parameters being varied to minimize the energy of the target system. Recognizing that the gates are intermediates which are ultimately compiled into a set of control pulses to be applied to each qubi

  90. Catalina Gomez, Ruolin Wang, Katharina Breininger, Corinne Casey

    Primary care providers are vital for initial triage and referrals to specialty care. In glaucoma, asymptomatic and fast progression can lead to vision loss, necessitating timely referrals to specialists. However, primary eye care providers may not identify urgent cases, potentially delaying care. Artificial Intelligence (AI) offering explanations could enhan

  91. Pranjol Sen Gupta, Md Rajib Hossen, Pengfei Li, Shaolei Ren

    Freshwater scarcity is a global problem that requires collective efforts across all industry sectors. Nevertheless, a lack of access to operational water footprint data bars many applications from exploring optimization opportunities hidden within the temporal and spatial variations. To break this barrier into research in water sustainability, we build a dat

  92. Zongxi Liu, Jiacheng Chen, Yunting Xu, Ting Ma

    In communications theory, the capacity of multiple input multiple output-orthogonal frequency division multiplexing (MIMO-OFDM) systems is fundamentally determined by wireless channels, which exhibit both diversity and correlation in spatial, frequency and temporal domains. It is further envisioned to exploit the inherent nature of channels, namely represent

  93. Jing Li, Zhijie Sun, Dachao Lin, Xuan He

    Mixture-of-Experts (MoE) architectures enable efficient scaling of large language models by activating only a subset of parameters per input. However, existing MoE models suffer from two critical limitations: (1) inefficient token-to-expert routing that causes excessive communication overhead, and (2) expert homogenization that leads to redundant computation

  94. Yuanchun Wang, Jifan Yu, Zijun Yao, Jing Zhang

    Applying large language models (LLMs) for academic API usage shows promise in reducing researchers' academic information seeking efforts. However, current LLM API-using methods struggle with complex API coupling commonly encountered in academic queries. To address this, we introduce SoAy, a solution-based LLM API-using methodology for academic information se

  95. Jacob Russin, Sam Whitman McGrath, Danielle J. Williams

    Compositionality has long been considered a key explanatory property underlying human intelligence: arbitrary concepts can be composed into novel complex combinations, permitting the acquisition of an open ended, potentially infinite expressive capacity from finite learning experiences. Influential arguments have held that neural networks fail to explain thi

  96. Pouya Babahajiani, Peng Zhang, Ji Liu, Tzu-Chieh Wei

    Distributed control of multi-inverter microgrids has attracted considerable attention as it can achieve the combined goals of flexible plug-and-play architecture guaranteeing frequency and voltage regulation while preserving power sharing among nonidentical distributed energy resources (DERs). However, it turns out that cybersecurity has emerged as a serious

  97. Wen-Xu Lin, Sheng-Bang Qian, Li-Ying Zhu, Wen-Ping Liao

    Asteroseismology offers a profound window into stellar interiors and has emerged as a pivotal technique in exoplanetary research. This study harnesses the Transiting Exoplanet Survey Satellite (TESS) observations to reveal, for the first time, the asteroseismic oscillations of four exoplanet-hosting stars. Through meticulous analysis, we extracted their aste

  98. Huali Ren, Anli Yan, Chong-zhi Gao, Hongyang Yan

    Visual Prompt Learning (VPL) differs from traditional fine-tuning methods in reducing significant resource consumption by avoiding updating pre-trained model parameters. Instead, it focuses on learning an input perturbation, a visual prompt, added to downstream task data for making predictions. Since learning generalizable prompts requires expert design and

  99. Sucheng Ren, Hongru Zhu, Chen Wei, Yijiang Li

    This paper presents a new self-supervised video representation learning framework, ARVideo, which autoregressively predicts the next video token in a tailored sequence order. Two key designs are included. First, we organize autoregressive video tokens into clusters that span both spatially and temporally, thereby enabling a richer aggregation of contextual i

  100. Vrushabh Zinage, Shrenik Zinage, Srinivas Bettadpur, Efstathios Bakolas

    In this paper, we consider the problem of precise attitude control for geodetic missions, such as the GRACE Follow-on (GRACE-FO) mission. Traditional and well-established control methods, such as Proportional-Integral-Derivative (PID) controllers, have been the standard in attitude control for most space missions, including the GRACE-FO mission. Instead of s