Skip to content

May 2022 arXiv papers — page 38

Showing 3,7013,800 of 15,811 papers

  1. Marcio Fonseca, Yftah Ziser, Shay B. Cohen

    We argue that disentangling content selection from the budget used to cover salient content improves the performance and applicability of abstractive summarizers. Our method, FactorSum, does this disentanglement by factorizing summarization into two steps through an energy function: (1) generation of abstractive summary views; (2) combination of these views

  2. Aman Madaan, Dheeraj Rajagopal, Niket Tandon, Yiming Yang

    Conditional set generation learns a mapping from an input sequence of tokens to a set. Several NLP tasks, such as entity typing and dialogue emotion tagging, are instances of set generation. Seq2Seq models, a popular choice for set generation, treat a set as a sequence and do not fully leverage its key properties, namely order-invariance and cardinality. We

  3. Pedram Hosseini, Christopher R. Wolfe, Mona Diab, David A. Broniatowski

    Decision making theories such as Fuzzy-Trace Theory (FTT) suggest that individuals tend to rely on gist, or bottom-line meaning, in the text when making decisions. In this work, we delineate the process of developing GisPy, an open-source tool in Python for measuring the Gist Inference Score (GIS) in text. Evaluation of GisPy on documents in three benchmarks

  4. Berend Zwartsenberg, Ryan P. Day, Elia Razzoli, Matteo Michiardi

    Sr$_{2}$IrO$_{4}$ has often been described via a simple, one-band pseudo-spin 1/2 model, subject to electron-electron interactions, on a square lattice, fostering analogies with cuprate superconductors, believed to be well described by a similar model. In this work we argue - based on a detailed study of the low-energy electronic structure by circularly pola

  5. Marcel Dengler

    In this work the following energy is considered $I(u)=\int\limits_B{\frac{1}{2}|\nabla u|^2+\rho(\det\nabla u)\;dx},$ where $B\subset\mathbb{R}^2$ denotes the unit ball, $u\in W^{1,2}(B,\mathbb{R}^2),$ and $\rho:\mathbb{R}\rightarrow\mathbb{R}_0^+$ smooth and convex with $\rho(s)=0$ for all $s\le0$ and $\rho$ becomes affine when $s$ exceeds some value $s_0>0

  6. Xuchen You, Shouvanik Chakrabarti, Xiaodi Wu

    The Variational Quantum Eigensolver (VQE) is a promising candidate for quantum applications on near-term Noisy Intermediate-Scale Quantum (NISQ) computers. Despite a lot of empirical studies and recent progress in theoretical understanding of VQE's optimization landscape, the convergence for optimizing VQE is far less understood. We provide the first rigorou

  7. Dongmei Zhang, Fangyang Zheng

    In 1984, Gauduchon considered the functional of $L^2$-norm of his torsion $1$-form on a compact Hermitian manifold. He obtained the Euler-Lagrange equation for this functional, and showed that in dimension $2$ the critical metrics must be balanced (namely with vanishing torsion $1$-form). In this note we extend his result to higher dimensions, and show that

  8. Chuan-Xin Cui, Mamiya Kawaguchi, Jin-Yang Li, Shinya Matsuzaki

    Violation of the $U(1)$ axial symmetry in QCD is stricter than the chiral $SU(2)$ breaking, simply because of the presence of the quantum axial anomaly. If the QCD gauge coupling is sent to zero, the strength of the $U(1)$ axial breaking coincides with that of the chiral $SU(2)$ breaking, which we shall in short call an axial-chiral coincidence. This coincid

  9. Dewi Amaliah

    Surveys provide important evidence for policymaking, decision-making, and understanding of society. However, conducting the large surveys required to provide subpopulation level estimates is expensive and time-consuming. Multilevel Regression and Poststratification (MRP) is a promising method to provide reliable estimates for subpopulations from surveys with

  10. Wei Liu, Jingyu Li, Tan Lee

    The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is expected when applying an automatic speech recognition (ASR) system trained on adult speech to child speech directly, as a result of domain mismatch. The present study is focused on

  11. Yixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb

    Neural attention models have achieved significant improvements on many natural language processing tasks. However, the quadratic memory complexity of the self-attention module with respect to the input length hinders their applications in long text summarization. Instead of designing more efficient attention modules, we approach this problem by investigating

  12. Xiangyang Li, Xiang Long, Yu Xia, Sujian Li

    Text style transfer (TST) without parallel data has achieved some practical success. However, most of the existing unsupervised text style transfer methods suffer from (i) requiring massive amounts of non-parallel data to guide transferring different text styles. (ii) colossal performance degradation when fine-tuning the model in new domains. In this work, w

  13. Sukeerthi Mandyam, Shanmuga Priya, Shalini Suresh, Kavitha Srinivasan

    There are numerous geo-climatic and human factors that contribute to the occurrence of natural disasters in the real-world scenario. Besides the study of causes and preconditions of such calamities, post-disaster analysis is essential for the efficient management of the disaster situation. This process needs timely and accurate data in light of the increasin

  14. Quanzhi Ye, Jérémie Vaubaillon

    The encounter of the meteoric material from 73P/Schmassmann--Wachmann~3 produced during the comet's 1995 outburst in May 2022 provides a rare and valuable opportunity to understand a fragmenting comet. Here we explore various ejection configurations and their impact on the meteor outburst detected in the early hours of UT 2022 May 31. We show that the dust m

  15. Manoj Mathews

    Anyone who looks into the circuitry world will be familiar with the three fundamental circuit elements - capacitor, resistor, and inductor. These circuit elements are defined by the relation between two of the four fundamental circuit variables current, voltage, charge, and flux. However, in 1971, Prof. Leon Chua proposed on the grounds of symmetry that ther

  16. Yukun Huang, Kun Qian, Zhou Yu

    Prompt tuning (PT) is an effective approach to adapting pre-trained language models to downstream tasks. Without a good initialization, prompt tuning doesn't perform well under few-shot settings. So pre-trained prompt tuning (PPT) is proposed to initialize prompts by leveraging pre-training data. We propose MetaPT (Meta-learned Prompt Tuning) to further impr

  17. Yicong Zhu, Changnian Han, Peng Zhang, Guojing Cong

    We have developed an AI-aided multiple time stepping (AI-MTS) algorithm and multiscale modeling framework (AI-MSM) and implemented them on the Summit-like supercomputer, AIMOS. AI-MSM is the first of its kind to integrate multi-physics, including intra-platelet, inter-platelet, and fluid-platelet interactions, into one system. It has simulated a record-setti

  18. Ryan Baker, John Garvey, Mitchell Kraft, Manoj Mathews

    Our application of command and control is the Aegis Combat System. Major components of this system include missile guidance and missile tracking. To look further into some of the aspects of these systems, an extremely simplified model of the Aegis Combat System will be designed. In this simplified model, a small-scale car will autonomously follow a small-sca

  19. Suzanna Sia, Anton Belyy, Amjad Almahairi, Madian Khabsa

    Evaluating an explanation's faithfulness is desired for many reasons such as trust, interpretability and diagnosing the sources of model's errors. In this work, which focuses on the NLI task, we introduce the methodology of Faithfulness-through-Counterfactuals, which first generates a counterfactual hypothesis based on the logical predicates expressed in the

  20. Lixiang Lin, Jianke Zhu, Yisu Zhang

    Although having achieved the promising results on shape and color recovery through self-supervision, the multi-layer perceptrons-based methods usually suffer from heavy computational cost on learning the deep implicit surface representation. Since rendering each pixel requires a forward network inference, it is very computational intensive to synthesize a wh

  21. Linyong Nan, Lorenzo Jaime Yu Flores, Yilun Zhao, Yixin Liu

    Unfaithful text generation is a common problem for text generation systems. In the case of Data-to-Text (D2T) systems, the factuality of the generated text is particularly crucial for any real-world applications. We introduce R2D2, a training framework that addresses unfaithful Data-to-Text generation by training a system both as a generator and a faithfulne

  22. Chong Ma, Lin Zhao, Yuzhong Chen, Lu Zhang

    Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning the meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical imaging, where the clinical data (e.g., MR images with patholog

  23. Xuan Kan, Hejie Cui, Joshua Lukemire, Ying Guo

    Functional magnetic resonance imaging (fMRI) is one of the most common imaging modalities to investigate brain functions. Recent studies in neuroscience stress the great potential of functional brain networks constructed from fMRI data for clinical predictions. Traditional functional brain networks, however, are noisy and unaware of downstream prediction tas

  24. Yoones Rezaei, Stephen Lee

    Three-dimensional (3D) urban models have gained interest because of their applications in many use-cases such as urban planning and virtual reality. However, generating these 3D representations requires LiDAR data, which are not always readily available. Thus, the applicability of automated 3D model generation algorithms is limited to a few locations. In thi

  25. Jae-Hwan Choi, Ildoo Kim

    We obtain the existence, uniqueness, and regularity estimates of the following Cauchy problem \begin{equation}\label{ab eqn} \begin{cases} \partial_t u(t,x)=\psi(t,-i\nabla)u(t,x)+f(t,x),\quad &(t,x)\in(0,T)\times\mathbb{R}^d,\\ u(0,x)=0,\quad & x\in\mathbb{R}^d \end{cases} \end{equation} in (Muckenhoupt) weighted $L_p$-spaces with time-measurable pseudo-dif

  26. Yuting Yang, Yuke Li, Binbin Du

    The CTC-based automatic speech recognition (ASR) models without the external language model usually lack the capacity to model conditional dependencies and textual interactions. In this paper, we present a Gated Interlayer Collaboration (GIC) mechanism to improve the performance of CTC-based models, which introduces textual information into the model and thu

  27. Jianhan Wu, Shijing Si, Jianzong Wang, Jing Xiao

    Deep neural networks have become popular in many supervised learning tasks, but they may suffer from overfitting when the training dataset is limited. To mitigate this, many researchers use data augmentation, which is a widely used and effective method for increasing the variety of datasets. However, the randomness introduced by data augmentation causes inev

  28. Liyun Zeng, Hao Helen Zhang

    Multiclass probability estimation is the problem of estimating conditional probabilities of a data point belonging to a class given its covariate information. It has broad applications in statistical analysis and data science. Recently a class of weighted Support Vector Machines (wSVMs) has been developed to estimate class probabilities through ensemble lear

  29. Zhiqiang Gong, Ping Zhong, Jiahao Qi, Panhe Hu

    Deep Neural Networks have been successfully applied in hyperspectral image classification. However, most of prior works adopt general deep architectures while ignore the intrinsic structure of the hyperspectral image, such as the physical noise generation. This would make these deep models unable to generate discriminative features and provide impressive cla

  30. Guodong Sun, Yang Zhou, Huilin Pan, Bo Wu

    Real-time vision-based system of fault detection (RVBS-FD) for freight trains is an essential part of ensuring railway transportation safety. Most existing vision-based methods still have high computational costs based on convolutional neural networks. The computational cost is mainly reflected in the backbone, neck, and post-processing, i.e., non-maximum su

  31. Sergei M. Grudsky, Egor A. Maximenko, Alejandro Soto-González

    In this paper we study the eigenvalues of the laplacian matrices of the cyclic graphs with one edge of weight $\alpha$ and the others of weight $1$. We denote by $n$ the order of the graph and suppose that $n$ tends to infinity. We notice that the characteristic polynomial and the eigenvalues depend only on $\operatorname{Re}(\alpha)$. After that, through th

  32. Shramay Palta, Haozhe An, Yifan Yang, Shuaiyi Huang

    Retrieval based open-domain QA systems use retrieved documents and answer-span selection over retrieved documents to find best-answer candidates. We hypothesize that multilingual Question Answering (QA) systems are prone to information inconsistency when it comes to documents written in different languages, because these documents tend to provide a model wit

  33. S. Rebeca Juárez Wysozka, Piotr Kielanowski, Liliana Vazquez Mercado

    The angles of all unitarity triangles of the Cabibbo-Kobayashi-Maskawa matrix are determined from the experimental data. Our analysis is independent of the parameterization of the CKM matrix and it is based on the predictions of the unitarity for the angles and the areas of the unitarity triangles. We note that the lengths of the sides of the four unitarity

  34. Ladislav Rampášek, Mikhail Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu

    We propose a recipe on how to build a general, powerful, scalable (GPS) graph Transformer with linear complexity and state-of-the-art results on a diverse set of benchmarks. Graph Transformers (GTs) have gained popularity in the field of graph representation learning with a variety of recent publications but they lack a common foundation about what constitut

  35. Mozhdeh Gheini, Xuezhe Ma, Jonathan May

    A recent family of techniques, dubbed lightweight fine-tuning methods, facilitates parameter-efficient transfer learning by updating only a small set of additional parameters while keeping the parameters of the pretrained language model frozen. While proven to be an effective method, there are no existing studies on if and how such knowledge of the downstrea

  36. Daniel Campos, Alexandre Marques, Tuan Nguyen, Mark Kurtz

    Large Language Models have become the core architecture upon which most modern natural language processing (NLP) systems build. These models can consistently deliver impressive accuracy and robustness across tasks and domains, but their high computational overhead can make inference difficult and expensive. To make using these models less costly, recent work

  37. Linfeng Zhang, Xin Chen, Runpei Dong, Kaisheng Ma

    Recent progress in image-to-image translation has witnessed the success of generative adversarial networks (GANs). However, GANs usually contain a huge number of parameters, which lead to intolerant memory and computation consumption and limit their deployment on edge devices. To address this issue, knowledge distillation is proposed to transfer the knowledg

  38. Seungkwon Kim, Chaeheon Gwak, Dohyun Kim, Kwangho Lee

    Cartoon domain has recently gained increasing popularity. Previous studies have attempted quality portrait stylization into the cartoon domain; however, this poses a great challenge since they have not properly addressed the critical constraints, such as requiring a large number of training images or the lack of support for abstract cartoon faces. Recently,

  39. Stephanie Milani, Zhicheng Zhang, Nicholay Topin, Zheyuan Ryan Shi

    Many recent breakthroughs in multi-agent reinforcement learning (MARL) require the use of deep neural networks, which are challenging for human experts to interpret and understand. On the other hand, existing work on interpretable reinforcement learning (RL) has shown promise in extracting more interpretable decision tree-based policies from neural networks,

  40. Muhammad Abdullah Naeem, Miroslav Pajic

    We study the concentration phenomenon for discrete-time random dynamical systems with an unbounded state space. We develop a heuristic approach towards obtaining exponential concentration inequalities for dynamical systems using an entirely functional analytic framework. We also show that existence of exponential-type Lyapunov function, compared to the purel

  41. Santiago R. Balseiro, Shangzhou Xia

    We study a dynamic allocation problem in which $T$ sequentially arriving divisible resources are to be allocated to a number of agents with linear utilities. The marginal utilities of each resource to the agents are drawn stochastically from a known joint distribution, independently and identically across time, and the central planner makes immediate and irr

  42. Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang

    We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic

  43. Akash Doshi, Manan Gupta, Jeffrey G. Andrews

    Future wireless systems are trending towards higher carrier frequencies that offer larger communication bandwidth but necessitate the use of large antenna arrays. Existing signal processing techniques for channel estimation do not scale well to this "high-dimensional" regime in terms of performance and pilot overhead. Meanwhile, training deep learning based

  44. A. Legros, K. W. Post, Prashant Chauhan, D. G. Rickel

    The recent observation of cyclotron resonance in optimally-doped La$_{2-x}$Sr$_x$CuO$_4$ using time-domain THz spectroscopy in high magnetic field has given new possibilities for the study of cuprate superconductors. One can measure the cyclotron mass in more disordered cuprates possesing short scattering times therefore expanding the study to materials and

  45. Kaiyu Yang, Jia Deng, Danqi Chen

    Reasoning over natural language is a challenging problem in NLP. In this work, we focus on proof generation: Given a hypothesis and a set of supporting facts, the model generates a proof tree indicating how to derive the hypothesis from supporting facts. Compared to generating the entire proof in one shot, stepwise generation can better exploit the compositi

  46. Peng Chen, Zhimin Chen, Pu Miao, Yun Chen

    Reconfigurable Intelligent Surfaces (RIS) emerge as promising technologies in future radar and wireless communication domains. This letter addresses the passive sensing issue utilizing wireless communication signals and RIS amidst interference from wireless access points (APs). We introduce an atomic norm minimization (ANM) approach to leverage spatial domai

  47. Santosh Pandey, Yunsoo Park, Ankita Ankita, Gregory J. Phillips

    A hallmark of bacterial populations cultured in vitro is their homogeneity of growth, where the majority of cells display identical growth rate, cell size and content. Recent insights, however, have revealed that even cells growing in exponential growth phase can be heterogeneous with respect to variables typically used to measure cell growth. Bacterial hete

  48. Donglei Du

    We propose a two-phase systematical framework for approximation algorithm design and analysis via Lyapunov function. The first phase consists of using Lyapunov function as an input and outputs a continuous-time approximation algorithm with a provable approximation ratio. The second phase then converts this continuous-time algorithm to a discrete-time algorit

  49. Ivan Hu, Zihuai Lin

    The ability to visually verify some element of a remotely controlled agricultural automation system through a photograph is valuable in many cases, not only in the operational phase of the system, but especially in the design and implementation phases. Owing to the remote location of many of the application sites, cellular technology is one enabling medium t

  50. Fam Le Kien, Sile Nic Chormaic, Thomas Busch

    We study the transfer of angular momentum of guided photons to a two-level atom with an electric quadrupole transition near an optical nanofiber. We show that the generation of the axial orbital torque of the driving guided field on the atom is governed by the internal-state selection rules for the quadrupole transition and by the angular momentum conservati

  51. Michael J. Mossinghoff, Christopher Pinner

    Newman showed that for primes $p\geq 5$ an integral circulant determinant of prime power order $p^t$ cannot take the value $p^{t+1}$ once $t\geq 2.$ We show that many other values are also excluded. In particular, we show that $p^{2t}$ is the smallest power of $p$ attained for any $t\geq 3$, $p\geq 3.$ We demonstrate the complexity involved by giving a compl

  52. Upender Kalwa, Christopher Legner, Taejoon Kong, Santosh Pandey

    Among the different types of skin cancer, melanoma is considered to be the deadliest and is difficult to treat at advanced stages. Detection of melanoma at earlier stages can lead to reduced mortality rates. Desktop-based computer-aided systems have been developed to assist dermatologists with early diagnosis. However, there is significant interest in develo

  53. Rafael Villarroel-Flores

    A collection of sets is intersecting, if any pair of sets in the collection has nonempty intersection. A collection of sets \(\mathcal{C}\) has the Helly property if any intersecting subcollection has nonempty intersection. A graph is /Helly/ if the collection of maximal complete subgraphs of \(G\) has the Helly property. We prove that if \(G\) is a \(k\)-re

  54. Upender Kalwa, Christopher Legner, Elizabeth Wlezien, Gregory Tylka

    The soybean cyst nematode (SCN), Heterodera glycines, is the most damaging pathogen of soybeans in the United States. To assess the severity of nematode infestations in the field, SCN egg population densities are determined. Cysts (dead females) of the nematode must be extracted from soil samples and then ground to extract the eggs within. Sucrose centrifuga

  55. Kendric Schefers

    Let $Z$ be an l.c.i. scheme over $\mathbb{C}$. In this paper, we introduce a Kashiwara--Schapira-style functor of derived microlocalization, which we use to define a perverse sheaf $\mu_{Z}$ on the $-1$-shifted cotangent bundle, $T^*[-1]Z$. The sheaf $\mu_{Z}$ is designed to be a refinement of the microlocal homology of $Z$: a family of invariants introduced

  56. Erjuan Fu

    Let $X$ be a closed Riemann surface. When $X$ is embedded into a projective space, the first rational cohomology group can be concretely obtained from the monodromy in the family of its smooth hyperplane sections by C. Schnell's tube mapping. We generalize this result to the first integral homology group by relating the tube mapping with the topological Abel

  57. Chris G. Antonopoulos, Mohammad H. Akram, Vasileios Basios, Anouchah Latifi

    The slogan "nobody is safe until everybody is safe" is a dictum to raise awareness that in an interconnected world, pandemics such as COVID-19, require a global approach. Motivated by the ongoing COVID-19 pandemic, we model here the spread of a virus in interconnected communities and explore different vaccination scenarios, assuming that the efficacy of the

  58. Ying-Chieh Lin, Jay Chu, John M. Hong, Hsin-Yi Lee

    In this paper, we study the global existence and asymptotic behavior of classical solutions near vacuum for the initial-boundary value problem modeling isentropic supersonic flows through divergent ducts. The governing equations are the compressible Euler equations with a small parameter, which can be written as a hyperbolic system in terms of the Riemann in

  59. Nana Wei, Yating Nie, Lin Liu, Xiaoqi Zheng

    Identifying cell clusters is a critical step for single-cell transcriptomics study. Despite the numerous clustering tools developed recently, the rapid growth of scRNA-seq volumes prompts for a more (computationally) efficient clustering method. Here, we introduce Secuer, a Scalable and Efficient speCtral clUstERing algorithm for scRNA-seq data. By employing

  60. Wanshan Li, Daren Wang, Alessandro Rinaldo

    The Bradley-Terry-Luce (BTL) model is a classic and very popular statistical approach for eliciting a global ranking among a collection of items using pairwise comparison data. In applications in which the comparison outcomes are observed as a time series, it is often the case that data are non-stationary, in the sense that the true underlying ranking change

  61. Yunhao Yang, Parham Gohari, Ufuk Topcu

    We study the privacy risks that are associated with training a neural network's weights with self-supervised learning algorithms. Through empirical evidence, we show that the fine-tuning stage, in which the network weights are updated with an informative and often private dataset, is vulnerable to privacy attacks. To address the vulnerabilities, we design a

  62. Makiya Nakashima, Inyeop Jang, Ramesh Basnet, Mitchel Benovoy

    Training deep learning models on cardiac magnetic resonance imaging (CMR) can be a challenge due to the small amount of expert generated labels and inherent complexity of data source. Self-supervised contrastive learning (SSCL) has recently been shown to boost performance in several medical imaging tasks. However, it is unclear how much the pre-trained repre

  63. Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh, Tianda Li

    Knowledge Distillation (KD) is a prominent neural model compression technique that heavily relies on teacher network predictions to guide the training of a student model. Considering the ever-growing size of pre-trained language models (PLMs), KD is often adopted in many NLP tasks involving PLMs. However, it is evident that in KD, deploying the teacher netwo

  64. Shang Liu, Jiashuo Jiang, Xiaocheng Li

    In this paper, we study the problem of bandits with knapsacks (BwK) in a non-stationary environment. The BwK problem generalizes the multi-arm bandit (MAB) problem to model the resource consumption associated with playing each arm. At each time, the decision maker/player chooses to play an arm, and s/he will receive a reward and consume certain amount of res

  65. Keisuke Suzuki

    In this paper, we propose a novel uniform generalization bound on the time and inverse temperature for stochastic gradient Langevin dynamics (SGLD) in a non-convex setting. While previous works derive their generalization bounds by uniform stability, we use Rademacher complexity to make our generalization bound independent of the time and inverse temperature

  66. Wensen Wei, Jin Tang, Yaodong Wu, Yihao Wang

    Topological magnetic charge Q is a fundamental parameter that describes the magnetic domains and determines their intriguing electromagnetic properties. The ability to switch Q in a controlled way by electrical methods allows for flexible manipulation of electromagnetic behavior in future spintronic devices. Here we report the room-temperature current-contro

  67. Shadaj Laddad, Conor Power, Mae Milano, Alvin Cheung

    Conflict-free replicated data types (CRDTs) are a promising tool for designing scalable, coordination-free distributed systems. However, constructing correct CRDTs is difficult, posing a challenge for even seasoned developers. As a result, CRDT development is still largely the domain of academics, with new designs often awaiting peer review and a manual proo

  68. Hazim Hanif, Sergio Maffeis

    This paper presents VulBERTa, a deep learning approach to detect security vulnerabilities in source code. Our approach pre-trains a RoBERTa model with a custom tokenisation pipeline on real-world code from open-source C/C++ projects. The model learns a deep knowledge representation of the code syntax and semantics, which we leverage to train vulnerability de

  69. Naofumi Hama, Masayoshi Mase, Art B. Owen

    A basic task in explainable AI (XAI) is to identify the most important features behind a prediction made by a black box function $f$. The insertion and deletion tests of Petsiuk et al. (2018) can be used to judge the quality of algorithms that rank pixels from most to least important for a classification. Motivated by regression problems we establish a formu

  70. Michaela Cully-Hugill, Adrian W. Dudek

    This paper gives an explicit version of Selberg's 1943 mean-value estimate for the prime number theorem in intervals under the Riemann hypothesis. Two applications are given: for primes in short intervals, and Goldbach numbers (sums of two primes) in short intervals. Under the Riemann hypothesis, we show there exists a prime in $(y,y+32277\log^2 y]$ for at l

  71. Ruiqi Zhong, Charlie Snell, Dan Klein, Jason Eisner

    Can non-programmers annotate natural language utterances with complex programs that represent their meaning? We introduce APEL, a framework in which non-programmers select among candidate programs generated by a seed semantic parser (e.g., Codex). Since they cannot understand the candidate programs, we ask them to select indirectly by examining the programs'

  72. Akiyoshi Kawamoto, Tomohiro I

    Let $S_{T}(k)$ denote the set of distinct substrings of length $k$ in a string $T$, then the $k$-th substring complexity is defined by its cardinality $|S_{T}(k)|$. Recently, $\delta = \max \{ |S_{T}(k)| / k : k \ge 1 \}$ is shown to be a good compressibility measure of highly-repetitive strings. In this paper, given $T$ of length $n$ in the run-length compr

  73. Te-Lin Wu, Caiqi Zhang, Qingyuan Hu, Alex Spangher

    The ability to infer pre- and postconditions of an action is vital for comprehending complex instructions, and is essential for applications such as autonomous instruction-guided agents and assistive AI that supports humans to perform physical tasks. In this work, we propose a task dubbed action condition inference, and collecting a high-quality, human annot

  74. Shady E. Ahmed, Omer San, Adil Rasheed, Traian Iliescu

    We propose a new physics guided machine learning (PGML) paradigm that leverages the variational multiscale (VMS) framework and available data to dramatically increase the accuracy of reduced order models (ROMs) at a modest computational cost. The hierarchical structure of the ROM basis and the VMS framework enable a natural separation of the resolved and unr

  75. Jiawei Huang, Li Zhao, Tao Qin, Wei Chen

    We propose a new learning framework that captures the tiered structure of many real-world user-interaction applications, where the users can be divided into two groups based on their different tolerance on exploration risks and should be treated separately. In this setting, we simultaneously maintain two policies $\pi^{\text{O}}$ and $\pi^{\text{E}}$: $\pi^{

  76. Mari Ohfuchi, Akihiko Sekine, Manabu Ohtomo, Kenichi Kawaguchi

    Monolayer WTe2 stripes are quantum spin Hall (QSH) insulators. Density functional theory was used for investigating the electronic properties of the stripes and steps in bilayer Td-WTe2. For the stripes oriented along the dimer chains of W atoms (x direction), the hybridization between the two layers suppresses the QSH states. However, the QSH nature can be

  77. Dheeraj Rajagopal, Siamak Shakeri, Cicero Nogueira dos Santos, Eduard Hovy

    Abstractive summarization systems based on pretrained language models often generate coherent but factually inconsistent sentences. In this paper, we present a counterfactual data augmentation approach where we augment data with perturbed summaries that increase the training data diversity. Specifically, we present three augmentation approaches based on repl

  78. Pak Hung Au, Mark Whitmeyer

    We consider a model of oligopolistic competition in a market with search frictions, in which competing firms with products of unknown quality advertise how much information a consumer's visit will glean. In the unique symmetric equilibrium of this game, the countervailing incentives of attraction and persuasion yield a payoff function for each firm that

  79. Oliver Pechenik, Matthew Satriano

    The ring of symmetric functions occupies a central place in algebraic combinatorics, with a particularly notable role in Schubert calculus, where the standard cell decompositions of Grassmannians yield the celebrated family of Schur functions and the cohomology ring is governed by Littlewood-Richardson rules. The past 50 years have seen an analogous developm

  80. COHERENT Collaboration, D. Akimov, P. An, C. Awe

    We use data from the COHERENT CsI[Na] scintillation detector to constrain sub-GeV leptophobic dark matter models. This detector was built to observe low-energy nuclear recoils from coherent elastic neutrino-nucleus scattering. These capabilities enable searches for dark matter particles produced at the Spallation Neutron Source mediated by a vector portal pa

  81. Luis A. Anchordoqui

    For decades, new physics searches in collider experiments have focused on the high-$p_T$ region. However, it has recently become evident that the LHC physics potential has not been fully exploited. To be specific, forward collisions, which produce particles along the beamline with enormous rates, have been almost completely ignored. For all practical purpose

  82. Jiankai Sun, Xin Yang, Yuanshun Yao, Junyuan Xie

    Federated learning has gained great attention recently as a privacy-enhancing tool to jointly train a machine learning model by multiple parties. As a sub-category, vertical federated learning (vFL) focuses on the scenario where features and labels are split into different parties. The prior work on vFL has mostly studied how to protect label privacy during

  83. Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc

    It is widely accepted in the mode connectivity literature that when two neural networks are trained similarly on the same data, they are connected by a path through parameter space over which test set accuracy is maintained. Under some circumstances, including transfer learning from pretrained models, these paths are presumed to be linear. In contrast to exi

  84. Yaqing Wang, Sahaj Agarwal, Subhabrata Mukherjee, Xiaodong Liu

    Standard fine-tuning of large pre-trained language models (PLMs) for downstream tasks requires updating hundreds of millions to billions of parameters, and storing a large copy of the PLM weights for every task resulting in increased cost for storing, sharing and serving the models. To address this, parameter-efficient fine-tuning (PEFT) techniques were intr

  85. Dan Chen, Xiaojin Zhang

    Let $\Lambda$ be a radical square zero algebra of a Dynkin quiver and let $\Gamma$ be the Auslander algebra of $\Lambda$. Then the number of tilting right $\Gamma$-modules is $2^{m-1}$ if $\Lambda$ is of $A_{m}$ type for $m\geq 1$. Otherwise, the number of tilting right $\Gamma$-modules is $2^{m-3}\times14$ if $\Lambda$ is either of $D_{m}$ type for $m\geq 4

  86. Bruce A. Corliss, Yaotian Wang, Heman Shakeri, Philip E. Bourne

    With limited resources, scientific inquiries must be prioritized for further study, funding, and translation based on their practical significance: whether the effect size is large enough to be meaningful in the real world. Doing so must evaluate a result's effect strength, defined as a conservative assessment of practical significance. We propose the least

  87. Payam Pourashraf, Bamshad Mobasher

    American local newspapers have been experiencing a large loss of reader retention and business within the past 15 years due to the proliferation of online news sources. Local media companies are starting to shift from an advertising-supported business model to one based on subscriptions to mitigate this problem. With this subscription model, there is a need

  88. Alexander Pondaven, Märt Bakler, Donghu Guo, Hamzah Hashim

    The widespread availability of satellite images has allowed researchers to model complex systems such as disease dynamics. However, many satellite images have missing values due to measurement defects, which render them unusable without data imputation. For example, the scanline corrector for the LANDSAT 7 satellite broke down in 2003, resulting in a loss of

  89. Hui Gao, Yihan Yang

    In online advertising, it is highly important to predict the probability and the value of a conversion (e.g., a purchase). It not only impacts user experience by showing relevant ads, but also affects ROI of advertisers and revenue of marketplaces. Unlike clicks, which often occur within minutes after impressions, conversions are expected to happen over a lo

  90. Rodrigo Sandoval-Orozco, Celia Escamilla-Rivera

    The current cosmic time evolution of the Universe is described by the General Relativity theory when a cosmological principle is considered under a flat space time landscape. The set of known as Friedmann equations, contain the principles that lead to the construction of the standard $\Lambda$CDM model. However, the current state-of-art regarding these equat

  91. Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda Muresan

    Figurative language understanding has been recently framed as a recognizing textual entailment (RTE) task (a.k.a. natural language inference, or NLI). However, similar to classical RTE/NLI datasets, the current benchmarks suffer from spurious correlations and annotation artifacts. To tackle this problem, work on NLI has built explanation-based datasets such

  92. Chloe LeGendre, Lukas Lepicovsky, Paul Debevec

    While the LED panels used in virtual production systems can display vibrant imagery with a wide color gamut, they produce problematic color shifts when used as lighting due to their peaky spectral output from narrow-band red, green, and blue LEDs. In this work, we present an improved color calibration process for virtual production stages which ameliorates t

  93. Christopher E. Denniston, Yun Chang, Andrzej Reinke, Kamak Ebadi

    Multi-robot SLAM systems in GPS-denied environments require loop closures to maintain a drift-free centralized map. With an increasing number of robots and size of the environment, checking and computing the transformation for all the loop closure candidates becomes computationally infeasible. In this work, we describe a loop closure module that is able to p

  94. Xinran Liang, Katherine Shu, Kimin Lee, Pieter Abbeel

    Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL methods are able to learn a more flexible reward model based on human preferences by actively incorporating human feedback, i.e. teacher's preferences between two clips of behaviors. However, poor feedback-efficiency still rema

  95. Kazuma Furukawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi

    In this study, we propose a head-to-head type (H2H-type) inter-personal multimodal Dirichlet mixture (Inter-MDM) by modifying the original Inter-MDM, which is a probabilistic generative model that represents the symbol emergence between two agents as multiagent multimodal categorization. A Metropolis--Hastings method-based naming game based on the Inter-MDM

  96. Julian K. Nauth, Vladimir M. Stojanovic

    Using the quantum-brachistochrone formalism, we address the problem of finding the fastest possible (time-optimal) deterministic conversion between $W$ and Greenberger-Horne-Zeilinger (GHZ) states in a system of three identical and equidistant neutral atoms that are acted upon by four external laser pulses. Assuming that all four pulses are close to being re

  97. James Lee-Thorp, Joshua Ainslie

    We combine the capacity of sparsely gated Mixture-of-Experts (MoE) with the speed and stability of linear, mixing transformations to design the Sparse Mixer encoder model. Sparse Mixer slightly outperforms (<1%) BERT on GLUE and SuperGLUE, but more importantly trains 65% faster and runs inference 61% faster. We also present a faster variant, prosaically name

  98. Teagan E Bate, Megan E Varney, Ezra H Taylor, Joshua H Dickie

    Active fluids have applications in micromixing, but little is known about the mixing kinematics of systems with spatiotemporally-varying activity. To investigate, UV-activated caged ATP was used to activate controlled regions of microtubule-kinesin active fluid and the mixing process was observed with fluorescent tracers and molecular dyes. At low P\'eclet n

  99. Pingakshya Goswami, Dinesh Bhatia

    Machine learning (ML) has been widely used to improve the predictability of EDA tools. The use of CAD tools that express designs at higher levels of abstraction makes machine learning even more important to highlight the performance of various design steps. Behavioral descriptions used during the high-level synthesis (HLS) are completely technology independe

  100. Yijun Tian, Chuxu Zhang, Zhichun Guo, Yihong Ma

    Learning effective recipe representations is essential in food studies. Unlike what has been developed for image-based recipe retrieval or learning structural text embeddings, the combined effect of multi-modal information (i.e., recipe images, text, and relation data) receives less attention. In this paper, we formalize the problem of multi-modal recipe rep