Skip to content

May 2023 arXiv papers — page 83

Showing 8,2018,300 of 19,695 papers

  1. BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson

    The processes $e^{+}e^{-} \to \Delta^{++}\bar{\Delta}^{--}$ and $e^{+}e^{-}\to \Delta^{++} \bar{p} \pi^{-} + c.c.$ are studied for the first time with $179~{\rm pb}^{-1}$ of $e^{+}e^{-}$ annihilation data collected with the BESIII detector at center-of-mass energies from $2.3094$ GeV to $2.6464$ GeV. No significant signal for the $e^{+}e^{-}\to \Delta^{++}\b

  2. Shixin Xu, Robert Eisenberg, Zilong Song, Huaxiong Huang

    Chemical reactions involve the movement of charges, and this work presents a mathematical model for describing chemical reactions in electrolytes. The model is developed using an energy variational method that aligns with classical thermodynamics principles. It encompasses both electrostatics and chemical reactions within consistently defined energetic and d

  3. Edoardo Otranto, Luca Scaffidi Domianello

    Markov Switching models have had increasing success in time series analysis due to their ability to capture the existence of unobserved discrete states in the dynamics of the variables under study. This result is generally obtained thanks to the inference on states derived from the so--called Hamilton filter. One of the open problems in this framework is the

  4. Han Wang, Ana B. Villas Bôas, William R. Young, Jacques Vanneste

    The refraction of surface gravity waves by currents leads to spatial modulations in the wave field and, in particular, in the significant wave height. We examine this phenomenon in the case of waves scattered by a localised current feature, assuming (i) the smallness of the ratio between current velocity and wave group speed, and (ii) a swell-like, highly di

  5. Zhijian Duan, Haoran Sun, Yurong Chen, Xiaotie Deng

    Automated auction design aims to find empirically high-revenue mechanisms through machine learning. Existing works on multi item auction scenarios can be roughly divided into RegretNet-like and affine maximizer auctions (AMAs) approaches. However, the former cannot strictly ensure dominant strategy incentive compatibility (DSIC), while the latter faces scala

  6. Denis Gessert, Henrik Christiansen, Wolfhard Janke

    One key aspect of coarsening following a quench below the critical temperature is domain growth. For the non-conserved Ising model a power-law growth of domains of like spins with exponent $\alpha = 1/2$ is predicted. Including recent work, it was not possible to clearly observe this growth law in the special case of a zero-temperature quench in the three-di

  7. Kelsey Linnell, Mikaela Fudolig, Laura Bloomfield, Thomas McAndrew

    A large and growing body of research demonstrates the value of local parks to mental and physical well-being. Recently, researchers have begun using passive digital data sources to investigate equity in usage; exactly who is benefiting from parks? Early studies suggest that park visitation differs according to demographic features, and that the demographic c

  8. Jera Hensel, Jürgen Giesl

    There are many techniques and tools to prove termination of C programs, but up to now these tools were not very powerful for fully automated termination proofs of programs whose termination depends on recursive data structures like lists. We present the first approach that extends powerful techniques for termination analysis of C programs (with memory alloca

  9. Ibrahim Ahmed, Marcos Quinones-Grueiro, Gautam Biswas

    In this paper, we leverage ideas from model-based control to address the sample efficiency problem of reinforcement learning (RL) algorithms. Accelerating learning is an active field of RL highly relevant in the context of time-varying systems. Traditional transfer learning methods propose to use prior knowledge of the system behavior to devise a gradual or

  10. Susana Perez Blazquez, Inas Hipolito

    This paper argues that Machine Learning (ML) algorithms must be educated. ML-trained algorithms moral decisions are ubiquitous in human society. Sometimes reverting the societal advances governments, NGOs and civil society have achieved with great effort in the last decades or are yet on the path to be achieved. While their decisions have an incommensurable

  11. Niklas Hörnedal, Ole Sönnerborn

    Geometric phase is a concept of central importance in virtually every branch of physics. In this paper, we show that the evolution time of a cyclically evolving quantum system is restricted by the system's energy resources and the geometric phase acquired by the state. Specifically, we derive and examine three tight lower bounds on the time required to gener

  12. Youcef Kehal, Khireddine Nouicer, Hamza Boumaza

    We study the existence and structure of static and slowly rotating neutron stars (NSs) in a particular truncation of scalar torsion theory with a scalar field $ \phi $ non-minimally coupled to the torsion scalar, and a potential of the form $ V(\phi)=-\mu^2\phi^2/2 +\lambda \phi^4 /4 $. We derive the hydrostatic equilibrium equations in the static case and s

  13. Dhruba Prakash Biswas, Priti Sharma, Sandip Jana

    In this paper we have found a necessary and sufficient condition for equivalence of two norms on a linear space using the theory of exponential vector space. Exponential vector space is an ordered algebraic structure which can be considered as an algebraic ordered extension of vector space. This structure is axiomatised on the basis of the intrinsic properti

  14. Thomas Nagel, Tymofiy Gerasimov, Dominik Kern

    This paper is intended to serve as a low-hurdle introduction to non-locality for graduate students and researchers with an engineering mechanics or physics background who did not have a formal introduction to the underlying mathematical basis. We depart from simple examples motivated by structural mechanics to form a physical intuition and demonstrate non-lo

  15. Jingyi Wang, Wu You, Yuheng Jiao, Yanhong Zhu

    The flexible endoscope is a minimally invasive tool in clinical settings, but most of them rely on exogenous staining for diagnosis to provide qualitative information. Here, we demonstrated a flexible endoscopic microscopy (FEM) with diffracted gradient light for quantitative phase imaging of unlabeled thick samples. Our instrument features a small form fact

  16. Dominik Stammbach, Vilém Zouhar, Alexander Hoyle, Mrinmaya Sachan

    Topic models are used to make sense of large text collections. However, automatically evaluating topic model output and determining the optimal number of topics both have been longstanding challenges, with no effective automated solutions to date. This paper proposes using large language models to evaluate such output. We find that large language models appr

  17. Sabrina C. Shen, Nicolas A. Lee, William J. Lockett, Aliai D. Acuil

    Fungal mycelium, a living network of filamentous threads, thrives on lignocellulosic waste and exhibits rapid growth, hydrophobicity, and intrinsic regeneration, offering a potential means to create next-generation sustainable and functional composites. However, existing hybrid-living mycelium composites (myco-composites) are tremendously constrained by conv

  18. Wen-Lei Zhao, Chao Han, Han Ke, Jie Liu

    We investigate the out-of-time ordered correlators and Loschmidt echo in a non-Hermitian interacting system governed by a Gross-Pitaevskii map model, which incorporates a periodically modulated complex strength of the nonlinear interaction as delta kicks. We uncover that the time evolutions of the out-of-time ordered correlators follow that of the Loschmidt

  19. Florentin Coeurdoux, Nicolas Dobigeon, Pierre Chainais

    Normalizing flows (NF) use a continuous generator to map a simple latent (e.g. Gaussian) distribution, towards an empirical target distribution associated with a training data set. Once trained by minimizing a variational objective, the learnt map provides an approximate generative model of the target distribution. Since standard NF implement differentiable

  20. Man Yao, Yuhong Chou, Guangshe Zhao, Xiawu Zheng

    The Lottery Ticket Hypothesis (LTH) states that a randomly-initialized large neural network contains a small sub-network (i.e., winning tickets) which, when trained in isolation, can achieve comparable performance to the large network. LTH opens up a new path for network pruning. Existing proofs of LTH in Artificial Neural Networks (ANNs) are based on contin

  21. Hanmeng Liu, Zhiyang Teng, Leyang Cui, Chaoli Zhang

    Generative Pre-trained Transformer 4 (GPT-4) demonstrates impressive chain-of-thought reasoning ability. Recent work on self-instruction tuning, such as Alpaca, has focused on enhancing the general proficiency of models. These instructions enable the model to achieve performance comparable to GPT-3.5 on general tasks like open-domain text generation and para

  22. Zhaokun Jiang, Ziyin Zhang

    Hedges are widely studied across registers and disciplines, yet research on the translation of hedges in political texts is extremely limited. This contrastive study is dedicated to investigating whether there is a diachronic change in the frequencies of hedging devices in the target texts, to what extent the changing frequencies of translated hedges through

  23. Zhangyang Gao, Xingran Chen, Cheng Tan, Stan Z. Li

    Is there a unified framework for graph-based retrosynthesis prediction? Through analysis of full-, semi-, and non-template retrosynthesis methods, we discovered that they strive to strike an optimal balance between combinability and consistency: \textit{Should atoms be combined as motifs to simplify the molecular editing process, or should motifs be broken d

  24. Sylvy Anscombe, Philip Dittmann, Franziska Jahnke

    We study the model theory of finitely ramified henselian valued fields of fixed initial ramification, obtaining versions of the Ax-Kochen-Ershov principle as follows. We identify the induced structure on the residue field and show that once the residue field is endowed with this structure, the theory of the valued field is determined by the theories of the e

  25. Yufeng He, Zefan Cai, Xu Gan, Baobao Chang

    Current image captioning works usually focus on generating descriptions in an autoregressive manner. However, there are limited works that focus on generating descriptions non-autoregressively, which brings more decoding diversity. Inspired by the success of diffusion models on generating natural-looking images, we propose a novel method DiffCap to apply con

  26. Sophie Blum, Raoul Koudijs, Ana Ozaki, Samia Touileb

    We investigate an approach for extracting knowledge from trained neural networks based on Angluin's exact learning model with membership and equivalence queries to an oracle. In this approach, the oracle is a trained neural network. We consider Angluin's classical algorithm for learning Horn theories and study the necessary changes to make it applicable to l

  27. Kai Ren

    Credit risk in the China's bond market has become increasingly evident, creating a progressively escalating risk of default for credit bond investors. Given the current incomplete and inaccurate bond information disclosure, timely tracking and forecasting the individual credit bond default risks have become essential to maintain market stability and ensure h

  28. Zhangyang Gao, Cheng Tan, Stan Z. Li

    Recent studies have shown competitive performance in protein design that aims to find the amino acid sequence folding into the desired structure. However, most of them disregard the importance of predictive confidence, fail to cover the vast protein space, and do not incorporate common protein knowledge. After witnessing the great success of pretrained model

  29. Ryusei Okaniwa, Takumi Mikawa, Yuichiro Matsuzaki, Tatsuma Yamaguchi

    The nitrogen-vacancy (NV) center is a promising candidate to realize practical quantum sensors with high sensitivity and high spatial resolution, even at room temperature and atmospheric pressure. In conventional high-frequency AC magnetometry with NV centers, the setup requires a pulse sequence with an appropriate time synchronization and strong microwave p

  30. Zihao Yue, Qi Zhang, Anwen Hu, Liang Zhang

    To help the visually impaired enjoy movies, automatic movie narrating systems are expected to narrate accurate, coherent, and role-aware plots when there are no speaking lines of actors. Existing works benchmark this challenge as a normal video captioning task via some simplifications, such as removing role names and evaluating narrations with ngram-based me

  31. Felix Strehle, Juan E. Machado, Michele Cucuzzella, Albertus J. Malan

    A fundamental precondition for the secure and efficient operation of district heating networks (DHNs) is a stable hydraulic behavior. However, the ongoing transition towards a sustainable heat supply, especially the rising integration of distributed heat sources and the increasingly meshed topologies, introduce complex and potentially destabilizing hydraulic

  32. Yao Du, Qing Li, Huawei Fan, Meng Zhan

    Power systems dominated by renewable energy encounter frequently large, random disturbances, and a critical challenge faced in power-system management is how to anticipate accurately whether the perturbed systems will return to the functional state after the transient or collapse. Whereas model-based studies show that the key to addressing the challenge lies

  33. Zixi Chen, Federico Renda, Alexia Le Gall, Lorenzo Mocellin

    Soft robots show compliance and have infinite degrees of freedom. Thanks to these properties, such robots can be leveraged for surgery, rehabilitation, biomimetics, unstructured environment exploring, and industrial grippers. In this case, they attract scholars from a variety of areas. However, nonlinearity and hysteresis effects also bring a burden to robot

  34. Md Zubair Ebne Rafique, Ali Basiri, Jing Bai, Jiawei Zuo

    Exploring novel materials with enhanced optical nonlinearities at low power levels with ultrafast response and small footprints is of great interests for information processing, communication, sensing and quantum systems. Recent progress on nonlinear metamaterials and metasurfaces suggests promising solutions to overcome the limitations of nonlinear material

  35. Jinsong Liu, Zheng-yi Lu, Ting Zhou

    Let $\{(p_n, \mathcal{D}_n, L_n)\}$ be a sequence of Hadamard triples on $\mathbb{R}$. Suppose that the associated Cantor-Moran measure $$ \mu_{\{p_n,\mathcal{D}_n\}}=\delta_{p_1^{-1}\mathcal{D}_1}\ast\delta_{(p_2p_1)^{-1}\mathcal{D}_2}\ast\cdots, $$ where $\sup_n\{|p_n^{-1}d|:d\in \mathcal{D}_n\}<\infty$ and $\sup\#\mathcal{D}_n<\infty$. It has been observe

  36. Alex Iacob, Pedro P. B. Gusmão, Nicholas D. Lane, Armand K. Koupai

    Human Activity Recognition (HAR) training data is often privacy-sensitive or held by non-cooperative entities. Federated Learning (FL) addresses such concerns by training ML models on edge clients. This work studies the impact of privacy in federated HAR at a user, environment, and sensor level. We show that the performance of FL for HAR depends on the assum

  37. Xiaolong Li, Zhi-Qin John Xu, Zhongwang Zhang

    In this work, we investigate the mechanism underlying loss spikes observed during neural network training. When the training enters a region with a lower-loss-as-sharper (LLAS) structure, the training becomes unstable, and the loss exponentially increases once the loss landscape is too sharp, resulting in the rapid ascent of the loss spike. The training stab

  38. Boxin Wang, Yibo Jacky Zhang, Yuan Cao, Bo Li

    We study (differentially) private federated learning (FL) of language models. The language models in cross-device FL are relatively small, which can be trained with meaningful formal user-level differential privacy (DP) guarantees when massive parallelism in training is enabled by the participation of a moderate size of users. Recently, public data has been

  39. Yuanyu Wan, Chang Yao, Yitao Ma, Mingli Song

    Although online convex optimization (OCO) under arbitrary delays has received increasing attention recently, previous studies focus on stationary environments with the goal of minimizing static regret. In this paper, we investigate the delayed OCO in non-stationary environments, and choose dynamic regret with respect to any sequence of comparators as the per

  40. Minrui Xu, Dusit Niyato, Hongliang Zhang, Jiawen Kang

    With the rapid development of artificial general intelligence (AGI), various multimedia services based on pretrained foundation models (PFMs) need to be effectively deployed. With edge servers that have cloud-level computing power, edge intelligence can extend the capabilities of AGI to mobile edge networks. However, compared with cloud data centers, resourc

  41. Chen Zhang, Yang Yang, Jiahao Liu, Jingang Wang

    Pretrained language models (LMs) have shown compelling performance on various downstream tasks, but unfortunately they require a tremendous amount of inference compute. Knowledge distillation finds a path to compress LMs to small ones with a teacher-student paradigm. However, when the capacity gap between the teacher and the student is large, a curse of capa

  42. Iryna Banakh, Taras Banakh, Maria Kolinko, Alex Ravsky

    A subset $X$ of an Abelian group $G$ is called $midconvex$ if for every $x,y\in X$ the set $\frac{x+y}2=\{z\in G:2z=x+y\}$ is a subset of $X$. We prove that a subset $X$ of an Abelian group $G$ is midconvex if and only if for every $g\in G$ and $x\in X$, the set $\{n\in\mathbb Z:x+ng\in X\}$ is equal to $C\cap H$ for some order-convex set $C\subseteq \mathbb

  43. Aleksei Petrenko, Arthur Allshire, Gavriel State, Ankur Handa

    In this work, we propose algorithms and methods that enable learning dexterous object manipulation using simulated one- or two-armed robots equipped with multi-fingered hand end-effectors. Using a parallel GPU-accelerated physics simulator (Isaac Gym), we implement challenging tasks for these robots, including regrasping, grasp-and-throw, and object reorient

  44. Maria G. Dainotti, Ritwik Sharma, Aditya Narendra, Delina Levine

    Gamma-Ray Bursts (GRBs), being observed at high redshift (z = 9.4), vital to cosmological studies and investigating Population III stars. To tackle these studies, we need correlations among relevant GRB variables with the requirement of small uncertainties on their variables. Thus, we must have good coverage of GRB light curves (LCs). However, gaps in the LC

  45. Arunselvan Ramaswamy, Shalabh Bhatnagar, Naman Saxena

    We present a novel algorithm for training deep neural networks in supervised (classification and regression) and unsupervised (reinforcement learning) scenarios. This algorithm combines the standard stochastic gradient descent and the gradient clipping method. The output layer is updated using clipped gradients, the rest of the neural network is updated usin

  46. Yuechun Song

    Using solar wind observation near PSP perihelions as constraints, we have investigated the parameters in various PFSS model methods. It's found that the interplanetary magnetic field extrapolation with source surface height $R_\mathrm{SS} = 2\,Rs$ is better than that with $R_\mathrm{SS} = 2.5\,Rs$. HMI and GONG magnetograms show similar performance in the si

  47. Ting Wu, Rui Zheng, Tao Gui, Qi Zhang

    Models trained with empirical risk minimization (ERM) are revealed to easily rely on spurious correlations, resulting in poor generalization. Group distributionally robust optimization (group DRO) can alleviate this problem by minimizing the worst-case loss over pre-defined groups. While promising, in practice factors like expensive annotations and privacy p

  48. Tatiana Shulman, Adam Skalski

    By Bekka's theorem the group C*-algebra of an amenable group $G$ is residually finite dimensional (RFD) if and only if $G$ is maximally almost periodic (MAP). We generalize this result in two directions of dynamical flavour. Firstly, we completely characterize the RFD property for crossed products by amenable actions of discrete groups on C*-algebras in term

  49. Jia Qi Yip, Tuan Truong, Dianwen Ng, Chong Zhang

    In this paper, we propose ACA-Net, a lightweight, global context-aware speaker embedding extractor for Speaker Verification (SV) that improves upon existing work by using Asymmetric Cross Attention (ACA) to replace temporal pooling. ACA is able to distill large, variable-length sequences into small, fixed-sized latents by attending a small query to large key

  50. Mostafa Eslami, Afshin Banazadeh

    Flight dynamics involve uncertainties in parameters, aerodynamic derivatives, and engine thrust. These uncertainties can be categorized into three types: known-predictable, known-unpredictable, and unknown. While advanced control systems typically rely on high-fidelity dynamical models in dealing with known-predictable uncertainties, simplified approaches ar

  51. Nima Anari, Moses Charikar, Prasanna Ramakrishnan

    Suppose that we have $n$ agents and $n$ items which lie in a shared metric space. We would like to match the agents to items such that the total distance from agents to their matched items is as small as possible. However, instead of having direct access to distances in the metric, we only have each agent's ranking of the items in order of distance. Given th

  52. Yu-Yu Wu, Hung-Jui Wang, Shang-Tse Chen

    In standard adversarial training, models are optimized to fit one-hot labels within allowable adversarial perturbation budgets. However, the ignorance of underlying distribution shifts brought by perturbations causes the problem of robust overfitting. To address this issue and enhance adversarial robustness, we analyze the characteristics of robust models an

  53. Peyman Alipour

    In this paper we apply the boundary elements method (BEM) and the dual reciprocity boundary elements method (DRBEM) for the numerical solution of two-dimensional time-fractional partial differential equations (TFPDEs). The fractional derivative of problem is described in the Caputo sense. In BEM, the main equation deduces to solving the Helmholtz equation in

  54. Jingtian Shi, A. H. MacDonald

    Single layer $\alpha$-ruthenium trichloride ($\rm\alpha-RuCl_3$) has been proposed as a potential quantum spin liquid. Graphene/$\rm RuCl_3$ heterobilayers have been extensively studied with a focus on the large interlayer electron transfer that dopes both materials. Here we examine the interplay between the competing magnetic state of $\rm RuCl_3$ layer and

  55. Mamta Gautam, Nitesh Jaiswal, Ankit Gill

    We study spread complexity and the statistics of work done for quenches in the three-spin interacting Ising model, the XY spin chain, and the Su-Schrieffer-Heeger model. We study these models without quench and for different schemes of quenches, such as sudden quench and multiple sudden quenches. We employ the Floquet operator technique to investigate all th

  56. Mingjie Cai, Zhishan Wu, Qingguo Li, Feng Xu

    Currently, density-based clustering algorithms are widely applied because they can detect clusters with arbitrary shapes. However, they perform poorly in measuring global density, determining reasonable cluster centers or structures, assigning samples accurately and handling data with large density differences among clusters. To overcome their drawbacks, thi

  57. Sundance O. Bilson-Thompson, Scott L. Todd, James Read, Valentina Baccetti

    In sonic models of special relativity, the fact that the sonic medium violates (ordinary) Lorentz symmetry is apparent to observers external to the sonic medium but not to a class of observers existing within the medium itself. We show that the situation is symmetric: internal observers will judge physics in the external laboratory to violate their own sonic

  58. Takeru K. Suzuki

    By performing ideal magnetohydrodynamical (MHD) simulations with weak vertical magnetic fields in unstratified cylindrical shearing boxes with modified boundary treatment, we investigate MHD turbulence excited by magnetorotational instability. The cylindrical simulation exhibits extremely large temporal variation in the magnetic activity compared to the simu

  59. Xiao-Min Zeng, Yan Song, Zhu Zhuo, Yu Zhou

    In this paper, we propose a joint generative and contrastive representation learning method (GeCo) for anomalous sound detection (ASD). GeCo exploits a Predictive AutoEncoder (PAE) equipped with self-attention as a generative model to perform frame-level prediction. The output of the PAE together with original normal samples, are used for supervised contrast

  60. Jinfeng Song

    We prove that the duals of the quantum Frobenius morphisms and their splittings by Lusztig are compatible with quantum cluster monomials. After specialisation, we deduce that the canonical Frobenius splittings on flag varieties are compatible with cluster algebra structures on Schubert cells.

  61. Yuwei Sun

    Meta-learning aims to develop algorithms that can learn from other learning algorithms to adapt to new and changing environments. This requires a model of how other learning algorithms operate and perform in different contexts, which is similar to representing and reasoning about mental states in the theory of mind. Furthermore, the problem of uncertainty in

  62. Atul Atul, Majid Ahmadi, Panagiotis Koutsogiannis, Heng Zhang

    The metal-insulator transition (MIT) observed in vanadium dioxide (VO2) has been a topic of great research interest for past decades, with the underlying physics yet not fully understood due to the complex electron interactions and structures involved. The ability to understand and tune the MIT behaviour is of vital importance from the perspective of both un

  63. Yi Zhong, Chen Zhang, Xule Liu, Chenxi Sun

    While Current TTS systems perform well in synthesizing high-quality speech, producing highly expressive speech remains a challenge. Emphasis, as a critical factor in determining the expressiveness of speech, has attracted more attention nowadays. Previous works usually enhance the emphasis by adding intermediate features, but they can not guarantee the overa

  64. Longkang Peng, Tao Wei, Xuehong Chen, Xiaobei Chen

    Convolutional neural networks (ConvNets) have been successfully applied to satellite image scene classification. Human-labeled training datasets are essential for ConvNets to perform accurate classification. Errors in human-annotated training datasets are unavoidable due to the complexity of satellite images. However, the distribution of real-world human-ann

  65. Chuan-gang Kang

    The Kaczmarz method is a popular iterative method for solving consistent, overdetermined linear system such as medical imaging in computerized tomography. The Kaczmarz's iteration repeatedly scans all equations in order, which leads to lower computational efficiency especially in solving a large scale problem. The standard form of Kaczmarz-Tanabe's iteration

  66. Taichi Kato

    Using Asteroid Terrestrial-impact Last Alert System (ATLAS) and Zwicky Transient Facility (ZTF) data, I found that the SW Sex star V1315 Aql entered a low state early in 2023. As far as I know, this is the first such an event since the discovery of this object with observations dating back to 1948. This object is renowned for its nova shell and the nova expl

  67. Eun-Ho Lee

    This paper proposes a modeling structure for the relativistic constitutive equations of inelastic deformation in materials moving at high speeds. While the theory of relativity has successfully approximated material motion in space-time, most existing models only consider elastic behavior, neglecting inelastic deformation. A comprehensive relativistic inelas

  68. Bita Banihashemi, Giuseppe De Giacomo, Yves Lespérance

    We develop a general framework for abstracting the behavior of an agent that operates in a nondeterministic domain, i.e., where the agent does not control the outcome of the nondeterministic actions, based on the nondeterministic situation calculus and the ConGolog programming language. We assume that we have both an abstract and a concrete nondeterministic

  69. Benjamin Coleman, Wang-Cheng Kang, Matthew Fahrbach, Ruoxi Wang

    Learning high-quality feature embeddings efficiently and effectively is critical for the performance of web-scale machine learning systems. A typical model ingests hundreds of features with vocabularies on the order of millions to billions of tokens. The standard approach is to represent each feature value as a d-dimensional embedding, introducing hundreds o

  70. Mohammad Lari

    A new method for capacity and spectral efficiency increases is a full-duplex (FD) communication, where sending and receiving are done simultaneously. Hence, severe interference leaked from the transmitter to the receiver, which can disrupt the system's operation completely. For interference reduction, the transceiver tries to estimate the interfering symbols

  71. Simone Bombari, Marco Mondelli

    Deep learning models are known to overfit and memorize spurious features in the training dataset. While numerous empirical studies have aimed at understanding this phenomenon, a rigorous theoretical framework to quantify it is still missing. In this paper, we consider spurious features that are uncorrelated with the learning task, and we provide a precise ch

  72. Xiangyu Gao, Yaping Sun, Hao Chen, Xiaodong Xu

    To support future 6G mobile applications, the mobile edge computing (MEC) network needs to be jointly optimized for computing, pushing, and caching to reduce transmission load and computation cost. To achieve this, we propose a framework based on deep reinforcement learning that enables the dynamic orchestration of these three activities for the MEC network.

  73. Jingjing Lu, Youjun Hu, Nong Xiang, Youwen Sun

    A deep neural network is developed and trained on magnetic measurements (input) and EFIT poloidal magnetic flux (output) on the EAST tokamak. In optimizing the network architecture, we use automatic optimization in searching for the best hyperparameters, which helps the model generalize better. We compare the inner magnetic surfaces and last-closed-flux surf

  74. Zongbo Bao, Penghui Yao

    We consider the problems of testing and learning quantum $k$-junta channels, which are $n$-qubit to $n$-qubit quantum channels acting non-trivially on at most $k$ out of $n$ qubits and leaving the rest of qubits unchanged. We show the following. 1. An $O\left(k\right)$-query algorithm to distinguish whether the given channel is $k$-junta channel or is far fr

  75. Neeraj Varshney, Mihir Parmar, Nisarg Patel, Divij Handa

    Pre-training on large corpora of text enables the language models to acquire a vast amount of factual and commonsense knowledge which allows them to achieve remarkable performance on a variety of language understanding tasks. They typically acquire this knowledge by learning from the pre-training text and capturing certain patterns from it. However, real-wor

  76. Wang Xue, Tian Zhou, Qingsong Wen, Jinyang Gao

    Recent studies have demonstrated the great power of Transformer models for time series forecasting. One of the key elements that lead to the transformer's success is the channel-independent (CI) strategy to improve the training robustness. However, the ignorance of the correlation among different channels in CI would limit the model's forecasting capacity. I

  77. Junchang Sun, Shuai Ma, Shiyin Li

    Integrated positioning and communication (IPAC) system and reconfigurable intelligent surface (RIS) are both considered to be key technologies for future wireless networks. Therefore, in this paper, we propose a RIS-enabled IPAC scheme with the millimeter wave system. First, we derive the explicit expressions of the time-of-arrival (ToA)-based Cram\'er-Rao b

  78. Gianni Cataldi, Yuri Aikawa, Kazunari Iwasaki, Sebastian Marino

    The origin and evolution of gas in debris disks is still not well understood. Secondary gas production from cometary material or a primordial origin have been proposed. So far, observations have mostly concentrated on CO, with only few C observations available. We create an overview of the C and CO content of debris disk gas and use it test state-of-the-art

  79. Mike Zhang, Rob van der Goot, Barbara Plank

    The increasing number of benchmarks for Natural Language Processing (NLP) tasks in the computational job market domain highlights the demand for methods that can handle job-related tasks such as skill extraction, skill classification, job title classification, and de-identification. While some approaches have been developed that are specific to the job marke

  80. Chao Zhao, Spandana Gella, Seokhwan Kim, Di Jin

    Task-oriented Dialogue (TOD) Systems aim to build dialogue systems that assist users in accomplishing specific goals, such as booking a hotel or a restaurant. Traditional TODs rely on domain-specific APIs/DBs or external factual knowledge to generate responses, which cannot accommodate subjective user requests (e.g., "Is the WIFI reliable?" or "Does the rest

  81. Wenyue Hua, Yingqiang Ge, Shuyuan Xu, Jianchao Ji

    Recent advances in Foundation Models such as Large Language Models (LLMs) have propelled them to the forefront of Recommender Systems (RS). Despite their utility, there is a growing concern that LLMs might inadvertently perpetuate societal stereotypes, resulting in unfair recommendations. Since fairness is critical for RS as many users take it for decision-m

  82. Roberto de A. Capistrano Filho, Luan S. de Sousa, Fernando A. Gallego

    Control properties of the Kawahara equation are considered when the equation is posed on an unbounded domain. Precisely, the paper's main results are related to an approximation theorem that ensures the exact (internal) controllability in $(0,+\infty)$. Following Rosier SIAM Simon (2000), the problem is reduced to prove an approximate theorem which is achiev

  83. Minhyeok Lee

    In this paper, we navigate the intricate domain of reviewer rewards in open-access academic publishing, leveraging the precision of mathematics and the strategic acumen of game theory. We conceptualize the prevailing voucher-based reviewer reward system as a two-player game, subsequently identifying potential shortcomings that may incline reviewers towards b

  84. Gang Liu, Tong Zhao, Eric Inae, Tengfei Luo

    Data imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rar

  85. Jonathan Li, Will Aitken, Rohan Bhambhoria, Xiaodan Zhu

    Parameter-efficient tuning aims to mitigate the large memory requirements of adapting pretrained language models for downstream tasks. For example, one popular method, prefix-tuning, prepends trainable tokens to sequences while freezing the rest of the model's parameters. Although such models attain comparable performance with fine-tuning when applied to seq

  86. Shiyu Liu, Linsen Wei, Shaogao Lv, Ming Li

    Graph convolutional networks (GCN) are viewed as one of the most popular representations among the variants of graph neural networks over graph data and have shown powerful performance in empirical experiments. That $\ell_2$-based graph smoothing enforces the global smoothness of GCN, while (soft) $\ell_1$-based sparse graph learning tends to promote signal

  87. Vivek Verma, Nicholas Tomlin, Dan Klein

    The uniform information density (UID) hypothesis states that humans tend to distribute information roughly evenly across an utterance or discourse. Early evidence in support of the UID hypothesis came from Genzel & Charniak (2002), which proposed an entropy rate constancy principle based on the probability of English text under n-gram language models. We re-

  88. Muhammad Abdullah Naeem, Miroslav Pajic

    We study the problem of identification of linear dynamical system from a single trajectory, via excitations of isotropic Gaussian. In stark contrast with previously reported results, Ordinary Least Squares (OLS) estimator for even \emph{stable} dynamical system contains non-vanishing error in \emph{high dimensions}; which stems from the fact that realization

  89. Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong

    Text-to-image generative models such as Stable Diffusion and DALL$\cdot$E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safety filters are often adopted to prevent the generation of NSFW images. In this work, we propose SneakyPrompt, the first automated attack fra

  90. Zifeng Wang, Chufan Gao, Cao Xiao, Jimeng Sun

    Tabular data prediction has been employed in medical applications such as patient health risk prediction. However, existing methods usually revolve around the algorithm design while overlooking the significance of data engineering. Medical tabular datasets frequently exhibit significant heterogeneity across different sources, with limited sample sizes per so

  91. Subhadeep Roy

    The article reports a numerical investigation of the breakdown of a disordered system considering the effect of local stress concentration under the action of an external tensile force. The statistics of the record-breaking magnitudes of emitted energies during the failure process, as well as the waiting time to achieve those record events, show rich behavio

  92. Gerdus Benadè, Ariel D. Procaccia, Jamie Tucker-Foltz

    The design of algorithms for political redistricting generally takes one of two approaches: optimize an objective such as compactness or, drawing on fair division, construct a protocol whose outcomes guarantee partisan fairness. We aim to have the best of both worlds by optimizing an objective subject to a binary fairness constraint. As the fairness constrai

  93. Sangchul Oh, Sabre Kais

    How fast a state of a system converges to a stationary state is one of the fundamental questions in science. Some Markov chains and random walks on finite groups are known to exhibit the non-asymptotic convergence to a stationary distribution, called the cutoff phenomenon. Here, we examine how quickly a random quantum circuit could transform a quantum state

  94. Kaige Xie, Tong Yu, Haoliang Wang, Junda Wu

    In real-world scenarios, labeled samples for dialogue summarization are usually limited (i.e., few-shot) due to high annotation costs for high-quality dialogue summaries. To efficiently learn from few-shot samples, previous works have utilized massive annotated data from other downstream tasks and then performed prompt transfer in prompt tuning so as to enab

  95. Yi-Shuai Niu

    We are interested in solving the Asymmetric Eigenvalue Complementarity Problem (AEiCP) by accelerated Difference-of-Convex (DC) algorithms. Two novel hybrid accelerated DCA: the Hybrid DCA with Line search and Inertial force (HDCA-LI) and the Hybrid DCA with Nesterov's extrapolation and Inertial force (HDCA-NI), are established. We proposed three DC programm

  96. I. I. Gontchar, M. V. Chushnyakova

    The process of fusion of complex nuclei is of significant interest as an example of the collective nuclear motion of large amplitude as well as a route for synthesis of new superheavy chemical elements. This process is accompanied by the dissipation of the energy of collective motion, at least at the last stage. The dissipative nature of fusion is accounted

  97. Weifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li

    Many text mining models are constructed by fine-tuning a large deep pre-trained language model (PLM) in downstream tasks. However, a significant challenge nowadays is maintaining performance when we use a lightweight model with limited labelled samples. We present DisCo, a semi-supervised learning (SSL) framework for fine-tuning a cohort of small student mod

  98. Minhyeok Lee

    Selecting the most suitable activation function is a critical factor in the effectiveness of deep learning models, as it influences their learning capacity, stability, and computational efficiency. In recent years, the Gaussian Error Linear Unit (GELU) activation function has emerged as a dominant method, surpassing traditional functions such as the Rectifie

  99. Weizhi Nie, Chen Zhang, Dan Song, Lina Zhao

    The chest X-ray (CXR) is one of the most common and easy-to-get medical tests used to diagnose common diseases of the chest. Recently, many deep learning-based methods have been proposed that are capable of effectively classifying CXRs. Even though these techniques have worked quite well, it is difficult to establish whether what these algorithms actually le

  100. Junzhe Cao, Sha Liu, Sirui Yang, Chengwen Zhong

    In this work, the Navier-Stokes (NS) solver is combined with the Direct simulation Monte Carlo (DSMC) solver in a direct way, under the wave-particle formulation [J. Comput. Phys. 401, 108977 (2020)]. Different from the classical domain decomposition method with buffer zone for overlap, in the proposed direct unified wave-particle (DUWP) method, the NS solve