Skip to content

May 2024 arXiv papers — page 72

Showing 7,1017,200 of 20,894 papers

  1. Yuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin

    We propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents. STAR is developed for prevalent speech foundation models based on Transformer-related architecture with auto-regressive dec

  2. Jin-Xing Hou, Alex Westström, Rui Wang, Wen-Li Yang

    Mesoscopic superconducting islands hosting Majorana zero modes (MZMs), or Majorana islands in short, offer a prototype of topological qubits. In this work we investigate theoretically the model of a generic Majorana island tunneling-coupled to a single-piece metallic substrate, hence an \textit{embedded Majorana island}. We show the crucial consequences of a

  3. Dylan Hillier, Leon Guertler, Cheston Tan, Palaash Agrawal

    The rapid advancement of large language models (LLMs) has led to significant improvements in natural language processing but also poses challenges due to their high computational and energy demands. This paper introduces a series of research efforts focused on Super Tiny Language Models (STLMs), which aim to deliver high performance with significantly reduce

  4. Boxiang Wang, Junwei Ji, Xiaoyi Shen, Dongyuan Shi

    Multichannel active noise control (ANC) systems are designed to create a large zone of quietness (ZoQ) around the error microphones, however, the placement of these microphones often presents challenges due to physical limitations. Virtual sensing technique that effectively suppresses the noise far from the physical error microphones is one of the most promi

  5. Anindya Ghatak, Narayan Rakshit, Jaydeb Sarkar, Mansi Suryawanshi

    Given a natural number $n \geq 1$, the odometer semigroup $O_n$, also known as the adding machine or the Baumslag-Solitar monoid with two generators, is a well-known object in group theory. This paper examines the odometer semigroup in relation to representations of bounded linear operators. We focus on noncommutative operators and prove that contractive rep

  6. Yuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan

    Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced with prompts in different sizes of solution spaces, LVLMs fail to always give consistent answers regarding the same knowledge point. This inconsistency of answers between different

  7. Joonsup Shim, Jinha Lim, Inki Kim, Jaeyong Jeong

    Waveguide-integrated mid-infrared (MIR) photodetectors are pivotal components for the development of molecular spectroscopy applications, leveraging mature photonic integrated circuit (PIC) technologies. Despite various strategies, critical challenges still remain in achieving broadband photoresponse, cooling-free operation, and large-scale complementary-met

  8. Feng Gu, Jie Lu, Zhen Fang, Kun Wang

    Uncertain changes in data streams present challenges for machine learning models to dynamically adapt and uphold performance in real-time. Particularly, classification boundary change, also known as real concept drift, is the major cause of classification performance deterioration. However, accurately detecting real concept drift remains challenging because

  9. Takeshi Yoshizawa

    Understanding how torsion theories are described and constructed is crucial to the study of torsion theory. Mutations of torsion theories have been studied as a method of constructing another torsion theory from a given one. We have already obtained how to mutate ordinary torsion theories into generalized torsion theories associated with a Serre subcategory.

  10. Raymond Vozzo, Stuart Johnson, Jonathan Tuke, Tanya Evans

    Active learning strategies have been widely recognised for their effectiveness in tertiary education, yet their implementation at scale, particularly in large first-year mathematics courses, presents considerable challenges. A common method for actively engaging students in large classes is through online quizzes, which may include structured peer instructio

  11. Jungyeul Park, Junrui Wang, Eunkyul Leah Jo, Angela Yoonseo Park

    We introduce an evaluation system designed to compute PARSEVAL measures, offering a viable alternative to \texttt{evalb} commonly used for constituency parsing evaluation. The widely used \texttt{evalb} script has traditionally been employed for evaluating the accuracy of constituency parsing results, albeit with the requirement for consistent tokenization a

  12. Elsayed Eshra, Konstantinos G. Papakonstantinou, Hamed Nikbakht

    This work introduces a novel framework for precisely and efficiently estimating rare event probabilities in complex, high-dimensional non-Gaussian spaces, building on our foundational Approximate Sampling Target with Post-processing Adjustment (ASTPA) approach. An unnormalized sampling target is first constructed and sampled, relaxing the optimal importance

  13. Kambhatla Akhila, Khaled R Ahmed

    Firearm Shootings and stabbings attacks are intense and result in severe trauma and threat to public safety. Technology is needed to prevent lone-wolf attacks without human supervision. Hence designing an automatic weapon detection using deep learning, is an optimized solution to localize and detect the presence of weapon objects using Neural Networks. This

  14. Oleg I. Berngardt

    This paper presents an algorithm for searching for the minimum number of neurons in fully connected layers of an arbitrary network solving given problem, which does not require multiple training of the network with different number of neurons. The algorithm is based at training the initial wide network using the cross-validation method over at least two fold

  15. Youta Noboru, Yuko Ozasa, Masayuki Tanaka

    Remote individual animal identification is important for food safety, sport, and animal conservation. Numerous existing remote individual animal identification studies have focused on RGB images. In this paper, we tackle individual penguin identification using hyperspectral (HS) images. To the best of our knowledge, it is the first work to analyze spectral d

  16. Lachlan Astfalck, Cassandra Bird, Daniel Williamson

    Motivated by big data and the vast parameter spaces in modern machine learning models, optimisation approaches to Bayesian inference have seen a surge in popularity in recent years. In this paper, we address the connection between the popular new methods termed generalised Bayesian inference and Bayes linear methods. We propose a further generalisation to Ba

  17. Jingxian Wang, Andrew G. Curtis, Mark Yim, Michael Rubenstein

    Communication and position sensing are among the most important capabilities for swarm robots to interact with their peers and perform tasks collaboratively. However, the hardware required to facilitate communication and position sensing is often too complicated, expensive, and bulky to be carried on swarm robots. Here we present Maneuverable Piccolissimo 3

  18. Renbo Zhao

    Given a finite-dimensional FTvN system $(\mathbb{V},\mathbb{W},\lambda)$, we study the convexification of the spectral set $\lambda^{-1}(\mathcal{C})$ induced by a set $\mathcal{C} \subseteq \mathbb{W}$. While the case of invariant $\mathcal{C}$ has been relatively well-studied, the results for non-invariant $\mathcal{C}$ are largely lacking in the literatur

  19. Se-eun Yoon, Hyunsik Jeon, Julian McAuley

    We introduce a multimodal dataset where users express preferences through images. These images encompass a broad spectrum of visual expressions ranging from landscapes to artistic depictions. Users request recommendations for books or music that evoke similar feelings to those captured in the images, and recommendations are endorsed by the community through

  20. Luan Thanh Nguyen

    Recent advancements in hate speech detection (HSD) in Vietnamese have made significant progress, primarily attributed to the emergence of transformer-based pre-trained language models, particularly those built on the BERT architecture. However, the necessity for specialized fine-tuned models has resulted in the complexity and fragmentation of developing a mu

  21. Chris Godsil, Xiaohong Zhang

    Let $G$ be a finite abelian group. Bridges and Mena characterized the Cayley graphs of $G$ that have only integer eigenvalues. Here we consider the $(0,1,-1)$ adjacency matrix of an oriented Cayley graph or of a signed Cayley graph $X$ on $G$. We give a characterization of when all the eigenvalues of $X$ are integer multiples of $\sqrt{\Delta}$ for some squa

  22. Xinhao Fan, Shreesh P Mysore

    Backpropagation (BP) has been pivotal in advancing machine learning and remains essential in computational applications and comparative studies of biological and artificial neural networks. Despite its widespread use, the implementation of BP in the brain remains elusive, and its biological plausibility is often questioned due to inherent issues such as the

  23. E. Annelise Bergeron, F. Sfigakis, A. Elbaroudy, A. W. M. Jordan

    We report on transport characteristics of field effect two-dimensional electron gases (2DEG) in 24 nm wide indium arsenide surface quantum wells. High quality single-subband magnetotransport with clear quantized integer quantum Hall plateaus are observed to filling factor $\nu=2$ in magnetic fields of up to B = 18 T, at electron densities up to 8$\times 10^{

  24. Jiawei Du, Jia Guo, Weihang Zhang, Shengzhu Yang

    The Vision-Language Foundation model is increasingly investigated in the fields of computer vision and natural language processing, yet its exploration in ophthalmology and broader medical applications remains limited. The challenge is the lack of labeled data for the training of foundation model. To handle this issue, a CLIP-style retinal image foundation m

  25. Yuzhang Shang, Dan Xu, Gaowen Liu, Ramana Rao Kompella

    Multi-task learning for dense prediction has emerged as a pivotal area in computer vision, enabling simultaneous processing of diverse yet interrelated pixel-wise prediction tasks. However, the substantial computational demands of state-of-the-art (SoTA) models often limit their widespread deployment. This paper addresses this challenge by introducing networ

  26. Xingchen Zou, Jiani Huang, Xixuan Hao, Yuhao Yang

    Regional socioeconomic indicators are critical across various domains, yet their acquisition can be costly. Inferring global socioeconomic indicators from a limited number of regional samples is essential for enhancing management and sustainability in urban areas and human settlements. Current inference methods typically rely on spatial interpolation based o

  27. Meng Wang, Shuijiang Zhao

    For Schr\"{o}dinger type operators in one dimension, we consider the relationship between the convergence rate and the regularity for initial data. By establishing the associated frequency-localized maximal estimates, we prove sharp results up to the endpoints. The optimal range for the wave operator in all dimensions is also obtained.

  28. Yesian Rohn

    Language's complexity is evident in the rich tapestry of slang expressions, often laden with humor and cultural nuances. This linguistic phenomenon has become increasingly prevalent, especially in digital communication. However, existing AI models, including ChatGPT-3.5, face challenges in comprehending these nuances, particularly in Chinese slang. In this s

  29. Xinyu Guo, Kai Wu, Xiaoyu Zhang, Jing Liu

    Class-imbalanced node classification tasks are prevalent in real-world scenarios. Due to the uneven distribution of nodes across different classes, learning high-quality node representations remains a challenging endeavor. The engineering of loss functions has shown promising potential in addressing this issue. It involves the meticulous design of loss funct

  30. Zexi Li, Lingzhi Gao, Chao Wu

    Generative artificial intelligence (GenAI) has made significant progress in understanding world knowledge and generating content from human languages across various modalities, like text-to-text large language models, text-to-image stable diffusion, and text-to-video Sora. While in this paper, we investigate the capability of GenAI for text-to-model generati

  31. Huy Nguyen, Pedram Akbarian, Trang Pham, Trang Nguyen

    The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better ability to mitigate the representation collapse issue, which often leads to parameter redundancy and limited representat

  32. Yassine Laguel, Yasa Syed, Necdet Serhat Aybat, Mert Gürbüzbalaban

    Stochastic smooth nonconvex minimax problems are prevalent in machine learning, e.g., GAN training, fair classification, and distributionally robust learning. Stochastic gradient descent ascent (GDA)-type methods are popular in practice due to their simplicity and single-loop nature. However, there is a significant gap between the theory and practice regardi

  33. Fei Zhao, Taotian Pang, Chunhui Li, Zhen Wu

    Multimodal Large Language Models (MLLMs) are widely regarded as crucial in the exploration of Artificial General Intelligence (AGI). The core of MLLMs lies in their capability to achieve cross-modal alignment. To attain this goal, current MLLMs typically follow a two-phase training paradigm: the pre-training phase and the instruction-tuning phase. Despite th

  34. Nikhilanj Pelluri

    Visual perception and navigation have emerged as major focus areas in the field of embodied artificial intelligence. We consider the task of image-goal navigation, where an agent is tasked to navigate to a goal specified by an image, relying only on images from an onboard camera. This task is particularly challenging since it demands robust scene understandi

  35. Mimi Dai

    We consider the electron magnetohydrodynamics (MHD) equation on the 3D torus $\mathbb T^3$. For a given smooth vector field $H$ with zero mean and zero divergence, we can construct a weak solution $B$ to the electron MHD in the space $L^\gamma_tW^{1,p}_x$ for appropriate $(\gamma, p)$ such that $B$ is arbitrarily close to $H$ in this space. The parameters $\

  36. Bum Jun Kim, Yoshinobu Kawahara, Sang Woo Kim

    Dynamical systems are often time-varying, whose modeling requires a function that evolves with respect to time. Recent studies such as the neural ordinary differential equation proposed a time-dependent neural network, which provides a neural network varying with respect to time. However, we claim that the architectural choice to build a time-dependent neura

  37. Yesian Rohn

    This paper addresses the limitations of physical models in the current field of image dehazing by proposing an innovative dehazing network (CL2S). Building on the DM2F model, it identifies issues in its ablation experiments and replaces the original logarithmic function model with a trigonometric (sine) model. This substitution aims to better fit the complex

  38. Jingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen

    Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, current evaluation benchmarks predominantly employ expert-designed contextual scenarios to assess how well LLMs align with human values. However, the labor-intensive nature of these

  39. Bin Lei, Yuchen Li, Qiuwu Chen

    We introduce AutoCoder, the first Large Language Model to surpass GPT-4 Turbo (April 2024) and GPT-4o in pass@1 on the Human Eval benchmark test ($\mathbf{90.9\%}$ vs. $\mathbf{90.2\%}$). In addition, AutoCoder offers a more versatile code interpreter compared to GPT-4 Turbo and GPT-4o. It's code interpreter can install external packages instead of limiting

  40. Minh Le, An Nguyen, Huy Nguyen, Trang Nguyen

    Exploiting the power of pre-trained models, prompt-based approaches stand out compared to other continual learning solutions in effectively preventing catastrophic forgetting, even with very few learnable parameters and without the need for a memory buffer. While existing prompt-based continual learning methods excel in leveraging prompts for state-of-the-ar

  41. Len Bos, Shayne Waldron

    We give a holomorphic quartic polynomial in the overlap variables whose zeros on the torus are precisely the Weyl-Heisenberg SICs (symmetric informationally complete positive operator valued measures). By way of comparison, all the other known systems of equations that determine a Weyl-Heisenberg SIC involve variables and their complex conjugates. We also gi

  42. Zuyuan Zhang, Mahdi Imani, Tian Lan

    Bayesian games model interactive decision-making where players have incomplete information -- e.g., regarding payoffs and private data on players' strategies and preferences -- and must actively reason and update their belief models (with regard to such information) using observation and interaction history. Existing work on counterfactual regret minimizatio

  43. Sheng-Jun Huang, Yi Li, Yiming Sun, Ying-Peng Tang

    Active learning (AL) for multiple target models aims to reduce labeled data querying while effectively training multiple models concurrently. Existing AL algorithms often rely on iterative model training, which can be computationally expensive, particularly for deep models. In this paper, we propose a one-shot AL method to address this challenge, which perfo

  44. Matthew Andres Moreno, Mark T. Holder, Jeet Sukumaran

    Contemporary bioinformatics has seen in profound new visibility into the composition, structure, and history of the natural world around us. Arguably, the central pillar of bioinformatics is phylogenetics -- the study of hereditary relatedness among organisms. Insight from phylogenetic analysis has touched nearly every corner of biology. Examples range acros

  45. Chongwei Liu, Haojie Li, Zhihui Wang, Rui Xu

    Recent advances in Multi-Object Tracking (MOT) have demonstrated significant success in short-term association within the separated tracking-by-detection online paradigm. However, long-term tracking remains challenging. While graph-based approaches address this by modeling trajectories as global graphs, these methods are unsuitable for real-time applications

  46. Sangwoo Jeon, Jihwan Kim, Duk Y. Kim, Zaeill Kim

    Microwave quantum illumination with entangled pairs of microwave signal and optical idler modes, can achieve the sub-optimal performance with joint measurement of the signal and idler modes. Here, we first propose a testbed of microwave quantum illumination with an optical memory which is simulated with a delay line in the idler mode. It provides how much an

  47. Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu

    Large language models (LLMs) store extensive factual knowledge, but the mechanisms behind how they store and express this knowledge remain unclear. The Knowledge Neuron (KN) thesis is a prominent theory for explaining these mechanisms. This theory is based on the Knowledge Localization (KL) assumption, which suggests that a fact can be localized to a few kno

  48. Xiyuan Zhao, Huijun Li, Tianyuan Miao, Xianyi Zhu

    The rapid development of collaborative robotics has provided a new possibility of helping the elderly who has difficulties in daily life, allowing robots to operate according to specific intentions. However, efficient human-robot cooperation requires natural, accurate and reliable intention recognition in shared environments. The current paramount challenge

  49. Bum Jun Kim, Sang Woo Kim

    Vision transformers (ViTs) have demonstrated remarkable performance in a variety of vision tasks. Despite their promising capabilities, training a ViT requires a large amount of diverse data. Several studies empirically found that using rich data augmentations, such as Mixup, Cutmix, and random erasing, is critical to the successful training of ViTs. Now, th

  50. Johannes Ackermann, Takayuki Osa, Masashi Sugiyama

    Current Reinforcement Learning (RL) is often limited by the large amount of data needed to learn a successful policy. Offline RL aims to solve this issue by using transitions collected by a different behavior policy. We address a novel Offline RL problem setting in which, while collecting the dataset, the transition and reward functions gradually change betw

  51. Zhusi Zhong, Jie Li, John Sollee, Scott Collins

    In response to the worldwide COVID-19 pandemic, advanced automated technologies have emerged as valuable tools to aid healthcare professionals in managing an increased workload by improving radiology report generation and prognostic analysis. This study proposes Multi-modality Regional Alignment Network (MRANet), an explainable model for radiology report gen

  52. Beniamin Goldys, Agus L. Soenjaya, Thanh Tran

    The Landau--Lifshitz--Baryakhtar (LLBar) equation perturbed by both additive and multiplicative noises is a system of fourth order stochastic PDEs which models the evolution of magnetic spin fields in ferromagnetic materials at elevated temperatures, taking into account longitudinal damping, long-range interactions, spin current, and noise-induced phenomena

  53. Yuyan Zhou, Ye Li, Lei Feng, Sheng-Jun Huang

    Recent studies showed that the generalization of neural networks is correlated with the sharpness of the loss landscape, and flat minima suggests a better generalization ability than sharp minima. In this paper, we propose a novel method called \emph{optimum shifting}, which changes the parameters of a neural network from a sharp minimum to a flatter one whi

  54. Jamie M. Taylor, David Pardo, Judit Muñoz-Matute

    Whilst the Universal Approximation Theorem guarantees the existence of approximations to Sobolev functions -- the natural function spaces for PDEs -- by Neural Networks (NNs) of sufficient size, low-regularity solutions may lead to poor approximations in practice. For example, classical fully-connected feed-forward NNs fail to approximate continuous function

  55. Ziyan Yao

    With the rapid growth and increasing complexity of industrial big data, traditional data processing methods are facing many challenges. This article takes an in-depth look at the application of cloud computing technology in industrial big data processing and explores its potential impact on improving data processing efficiency, security, and cost-effectivene

  56. Zhun-Yong Ong

    The reduced phonon specularity $p$ from boundary roughness scattering plays a major role in the lower thermal conductivity in semiconducting and insulating nanowires and films. Although the well-known Ziman formula $p=\exp(-4\sigma^{2}q_{x}^{2})$, where $\sigma$ and $q_{x}$ denote the root-mean-square boundary roughness and the normal component of the incide

  57. Alex Morehead, Nabin Giri, Jian Liu, Pawan Neupane

    The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein-ligand docking have recently been introduced, to date no prior works have syste

  58. Qian Zhu, Dakuo Wang, Shuai Ma, April Yi Wang

    As AI technology continues to advance, the importance of human-AI collaboration becomes increasingly evident, with numerous studies exploring its potential in various fields. One vital field is data science, including feature engineering (FE), where both human ingenuity and AI capabilities play pivotal roles. Despite the existence of AI-generated recommendat

  59. Meenatchi Sundaram Muthu Selva Annamalai, Emiliano De Cristofaro

    This paper presents an auditing procedure for the Differentially Private Stochastic Gradient Descent (DP-SGD) algorithm in the black-box threat model that is substantially tighter than prior work. The main intuition is to craft worst-case initial model parameters, as DP-SGD's privacy analysis is agnostic to the choice of the initial model parameters. For mod

  60. Ziyan Yao, Fei Lin, Sheng Chai, Weijie He

    In this paper, an innovative multi-modal deep learning model is proposed to deeply integrate heterogeneous information from medical images and clinical reports. First, for medical images, convolutional neural networks were used to extract high-dimensional features and capture key visual information such as focal details, texture and spatial distribution. Sec

  61. Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky

    This paper introduces distributed speculative inference (DSI), a novel inference algorithm that is provably faster than speculative inference (SI) [leviathan2023, chen2023, miao2024, sun2025, timor2025] and standard autoregressive inference (non-SI). Like other SI algorithms, DSI operates on frozen language models (LMs), requiring no training or architectura

  62. Yuehao Bai, Shunzhuang Huang, Sarah Moon, Azeem M. Shaikh

    In the context of a binary outcome, treatment, and instrument, Balke and Pearl (1993, 1997) es- tablish that the monotonicity condition of Imbens and Angrist (1994) has no identifying power beyond instrument exogeneity for average potential outcomes and average treatment effects in the sense that adding it to instrument exogeneity does not decrease the ident

  63. Yuanzhao Zhai, Zhuo Zhang, Kele Xu, Hanyang Peng

    Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward model to provide feedback for on-policy sampling responses. Recently, offline methods that directly fit responses with binary preferences in the dataset have emerged as alternative

  64. Miguel A. Romer, Layton A. Hall, Ayman F. Abouraddy

    Space-time wave packets (STWPs) are a new class of pulsed optical beams with many unique and intriguing attributes, including propagation invariance and tunable group velocity in linear optical media. STWPs are a form of spatiotemporally structured light, so their synthesis poses challenges that are not shared by conventional monochromatic structured light f

  65. Zakaria Patel, Kirill Serkh

    Diffusion models are a powerful class of generative models capable of producing high-quality images from pure noise using a simple text prompt. While most methods which introduce additional spatial constraints into the generated images (e.g., bounding boxes) require fine-tuning, a smaller and more recent subset of these methods take advantage of the models'

  66. Jeffrey S. Lee, Joe C. Yelderman, Gerald B. Cleaver

    The most pragmatic first step in the all-but-inevitable 3rd-millennium V\"olkerwanderung of humanity throughout the Solar System is the establishment of a permanent human presence on the Moon. This research examines: 1. the human, agricultural, and technical water needs of a 100-person, 500 m x 100 m x 6 m self-sustaining lunar colony; 2. choosing a strategi

  67. Chuqi Chen, Yahong Yang, Yang Xiang, Wenrui Hao

    Neural network-based approaches have recently shown significant promise in solving partial differential equations (PDEs) in science and engineering, especially in scenarios featuring complex domains or incorporation of empirical data. One advantage of the neural network methods for PDEs lies in its automatic differentiation (AD), which necessitates only the

  68. Hao Luo

    We give a continuous perspective on the Inertial Corrected Primal-Dual Proximal Splitting (IC-PDPS) proposed by Valkonen ({\it SIAM J. Optim.}, 30(2): 1391--1420, 2020) for solving saddle-point problems. The algorithm possesses nonergodic convergence rate and admits a tight preconditioned proximal point formulation which involves both inertia and additional

  69. Kuan Zhang, Yi-Kai Huo, Xiangdong Ji, Andreas Schaefer

    We analyze the gauge fixing precision dependence of some non-local quark-blinear lattice operators interesting in computing parton physics for several measurements, using 5 lattice spacings ranging from 0.032 fm to 0.121 fm. Our results show that gauge dependent non-local measurements are significantly more sensitive to the precision of gauge fixing than ant

  70. Wenrui Hao, Xinliang Liu, Yahong Yang

    Solving nonlinear partial differential equations (PDEs) with multiple solutions using neural networks has found widespread applications in various fields such as physics, biology, and engineering. However, classical neural network methods for solving nonlinear PDEs, such as Physics-Informed Neural Networks (PINN), Deep Ritz methods, and DeepONet, often encou

  71. Layton A. Hall, Ayman F. Abouraddy

    All linear, propagation-invariant, paraxial pulsed beams are spatiotemporally X-shaped (conical waves) in absence of group-velocity dispersion (GVD), or in presence of normal GVD. It is known, however, that such conical waves become O-shaped in presence of anomalous GVD, resulting in a field profile that is circularly symmetric in space and time. To date, ex

  72. Rubén Ballester, Pablo Hernández-García, Mathilde Papillon, Claudio Battiloro

    Topological Deep Learning seeks to enhance the predictive performance of neural network models by harnessing topological structures in input data. Topological neural networks operate on spaces such as cell complexes and hypergraphs, that can be seen as generalizations of graphs. In this work, we introduce the Cellular Transformer (CT), a novel architecture t

  73. Yuang Wang, Pengfei Jin, Siyeop Yoon, Matthew Tivnan

    Score-based diffusion models are frequently employed as structural priors in inverse problems. However, their iterative denoising process, initiated from Gaussian noise, often results in slow inference speeds. The Image-to-Image Schr\"odinger Bridge (I$^2$SB), which begins with the corrupted image, presents a promising alternative as a prior for addressing i

  74. Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao

    Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and vision-language models (VLMs), a new category of multimodal models -- referred to as vision-language-action (VLA) models

  75. Zhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan

    Intrinsic self-correct was a method that instructed large language models (LLMs) to verify and correct their responses without external feedback. Unfortunately, the study concluded that the LLMs could not self-correct reasoning yet. We find that a simple yet effective verification method can unleash inherent capabilities of the LLMs. That is to mask a key co

  76. Armaan V. Goyal, Songhu Wang

    The ubiquity of "peas-in-a-pod" architectural patterns and the existence of the radius valley each present a striking population-level trend for planets with $R_{p} \leq 4 R_{\oplus}$ that serves to place powerful constraints on the formation and evolution of these subgiant worlds. As it has yet to be determined whether the strength of this peas-in-a-pod uni

  77. Jingchi Jiang, Rujia Shen, Boran Wang, Yi Guan

    Type 1 diabetes mellitus (T1D) is characterized by insulin deficiency and blood glucose (BG) control issues. The state-of-the-art solution for continuous BG control is reinforcement learning (RL), where an agent can dynamically adjust exogenous insulin doses in time to maintain BG levels within the target range. However, due to the lack of action guidance, t

  78. Rosario Messana, Rui Chen, Andrea Lodi, Alberto Ceselli

    We consider solving a combinatorial optimization problem with unknown knapsack constraints using a membership oracle for each unknown constraint such that, given a solution, the oracle determines whether the constraint is satisfied or not with absolute certainty. The goal of the decision maker is to find the best possible solution subject to a budget on the

  79. Siba Smarak Panigrahi, Arnab Kumar Mondal

    This work introduces a novel approach to achieving architecture-agnostic equivariance in deep learning, particularly addressing the limitations of traditional layerwise equivariant architectures and the inefficiencies of the existing architecture-agnostic methods. Building equivariant models using traditional methods requires designing equivariant versions o

  80. Aymane El Firdoussi, Mohamed El Amine Seddik

    This paper provides theoretical insights into high-dimensional binary classification with class-conditional noisy labels. Specifically, we study the behavior of a linear classifier with a label noisiness aware loss function, when both the dimension of data $p$ and the sample size $n$ are large and comparable. Relying on random matrix theory by supposing a Ga

  81. JuAe Song

    We prove that the congruence on the tropical rational function semifield in $n$-variables associated with a subset $V$ of $\boldsymbol{R}^n$ is finitely generated if and only if the closure of $V$ is a finite union of $\boldsymbol{R}$-rational polyhedral sets. With this fact, we characterize rational function semifields of tropical curves.

  82. Kang Liu, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong

    The automated generation of imaging reports proves invaluable in alleviating the workload of radiologists. A clinically applicable reports generation algorithm should demonstrate its effectiveness in producing reports that accurately describe radiology findings and attend to patient-specific indications. In this paper, we introduce a novel method, \textbf{S}

  83. Zhi Li, Guoxin Wei

    In this paper, we obtain several classification results of $2$-dimensional complete Lagrangian translators and lagrangian self-expanders with constant squared norm $|\vec{H}|^{2}$ of the mean curvature vector in $\mathbb{C}^{2}$ by using a new Omori-Yau type maximum principle which was proved by Chen and Qiu \cite{CQ}. The same idea is also used to give a si

  84. Shezheng Song, Shasha Li, Shan Zhao, Chengyu Wang

    Multimodal aspect-based sentiment analysis (MABSA) aims to understand opinions in a granular manner, advancing human-computer interaction and other fields. Traditionally, MABSA methods use a joint prediction approach to identify aspects and sentiments simultaneously. However, we argue that joint models are not always superior. Our analysis shows that joint m

  85. Subhash Kantamneni, Ziming Liu, Max Tegmark

    How do transformers model physics? Do transformers model systems with interpretable analytical solutions, or do they create "alien physics" that are difficult for humans to decipher? We take a step in demystifying this larger puzzle by investigating the simple harmonic oscillator (SHO), $\ddot{x}+2\gamma \dot{x}+\omega_0^2x=0$, one of the most fundamental sy

  86. Goutam Paul, Nirupam Basak, Soumya Das

    This work presents two significant contributions from the perspectives of quantum random number generator (QRNG) manufacturers and users. For manufacturers, the conventional method of assessing the quantumness of single-photon-based QRNGs through mean and variance comparisons of photon counts is statistically unreliable due to finite sample sizes. Given the

  87. Yu Fu, Roy Zhao

    Given a correspondence $V$ between a connected Shimura variety $S$, a commutative connected algebraic group $G$, and $n \in \mathbb{N}$, we prove that the $V$-images of any $n$ special points on $S$ outside a proper Zariski closed subset are algebraically independent. Our result unifies previous unlikely intersection results on multiplicative independence an

  88. Doosung Park

    We define the log motivic nearby cycles functor. We show that this sends the motive of a proper smooth scheme over the fraction field of a DVR to the motive of the boundary of a log smooth model assuming absolute purity, which is unconditional in the equal characteristic case. In characteristic $0$, we show that the $\infty$-categories of motives over the st

  89. Junghyuk Yeom, Yonghyeon Jo, Jungmo Kim, Sanghyeon Lee

    Constraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limitation in existing offline RL methods with penalized value function, indicating the potential for underestimation bias due to unnecessary bia

  90. P. Liu, D. Wu, D. W. Yuan, G. Zhao

    Magnetized collisionless shocks drive particle acceleration broadly in space and astrophysics. We perform the first large-scale particle-in-cell simulations with realistic laboratory parameters (density, temperature, and velocity) to investigate the magnetized shock in head-on colliding plasmas with an applied magnetic field of tens of Tesla. It is shown tha

  91. Vladislav V. Serov, Anatoli S. Kheifets

    Two-photon atomic ionization driven by time-locked XUV and IR pulses allows to study dynamics of Fano resonances in time and energy domains. Different time evolution of the two interfering pathways leading to a Fano resonance can be exploited to turn the Fano profile of the two-photon XUV/IR ionization into a symmetric Gaussian once the directly ejected phot

  92. Dingyi Zhuang, Qingyi Wang, Yunhan Zheng, Xiaotong Guo

    Transportation mode share analysis is important to various real-world transportation tasks as it helps researchers understand the travel behaviors and choices of passengers. A typical example is the prediction of communities' travel mode share by accounting for their sociodemographics like age, income, etc., and travel modes' attributes (e.g. travel cost and

  93. Han-Dong Lim, Donghwan Lee

    Multi-agent reinforcement learning (MARL) has witnessed a remarkable surge in interest, fueled by the empirical success achieved in applications of single-agent reinforcement learning (RL). In this study, we consider a distributed Q-learning scenario, wherein a number of agents cooperatively solve a sequential decision making problem without access to the ce

  94. Rongyi Zhu, Zeliang Zhang, Susan Liang, Zhuo Liu

    Adversarial examples, crafted by adding perturbations imperceptible to humans, can deceive neural networks. Recent studies identify the adversarial transferability across various models, \textit{i.e.}, the cross-model attack ability of adversarial samples. To enhance such adversarial transferability, existing input transformation-based methods diversify inpu

  95. S. Yu. Orevkov, V. Florens

    We define the Witt coindex of a link with non-trivial Alexander polynomial, as a concordance invariant from the Seifert form. We show that it provides an upper bound for the (locally flat) slice Euler characteristic of the link, extending the work of Levine on algebraically slice knots and Taylor on the genera of knots. Then we extend the techniques by Levin

  96. Chengkun Cai, Xu Zhao, Yucheng Du, Haoliang Liu

    Large Language Models (LLMs) have emerged as powerful tools in artificial intelligence, especially in complex decision-making scenarios, but their static problem-solving strategies often limit their adaptability to dynamic environments. We explore the enhancement of reasoning capabilities in LLMs through Temperature Tree ($T^2$) prompting via a heuristic alg

  97. Lav Gupta, Guoxing Yao

    Researchers are exploring the integration of IoT and the cloud continuum, together with AI to enhance the cost-effectiveness and efficiency of critical infrastructure (CI) systems. This integration, however, increases susceptibility of CI systems to cyberattacks, potentially leading to disruptions like power outages, oil spills, or even a nuclear mishap. CI

  98. Chengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu

    Designing generalizable agents capable of adapting to diverse embodiments has achieved significant attention in Reinforcement Learning (RL), which is critical for deploying RL agents in various real-world applications. Previous Cross-Embodiment RL approaches have focused on transferring knowledge across embodiments within specific tasks. These methods often

  99. Bence Bakó, Dániel T. R. Nagy, Péter Hága, Zsófia Kallus

    Leveraging the intrinsic probabilistic nature of quantum systems, generative quantum machine learning (QML) offers the potential to outperform classical learning models. Current generative QML algorithms mostly rely on general-purpose models that, while being very expressive, face several training challenges. One potential way to address these setbacks is by

  100. Hung Viet Chu, Zachary Louis Vasseur

    For a finite set $A\subset\mathbb{N}$ and $k\in \mathbb{N}$, let $\omega_k(A) = \sum_{i\in A, i\neq k}1$. For each $n\in \mathbb{N}$, define $$a_{k, n}\ =\ |\{E\subset \mathbb{N}\,:\, E = \emptyset\mbox{ or } \omega_k(E) < \min E\leqslant \max E\leqslant n\}|.$$ First, we prove that $$a_{k,k+\ell} \ =\ 2F_{k+\ell},\mbox{ for all }\ell\geqslant 0\mbox{ and }k