May 2023 arXiv papers — page 56
Showing 5,501–5,600 of 19,695 papers
Moonseok Choi, Hyungi Lee, Giung Nam, Juho Lee
Given the ever-increasing size of modern neural networks, the significance of sparse architectures has surged due to their accelerated inference speeds and minimal memory demands. When it comes to global pruning techniques, Iterative Magnitude Pruning (IMP) still stands as a state-of-the-art algorithm despite its simple nature, particularly in extremely spar
Pengfei He, Han Xu, Jie Ren, Yingqian Cui
Recent research has highlighted the vulnerability of Deep Neural Networks (DNNs) against data poisoning attacks. These attacks aim to inject poisoning samples into the models' training dataset such that the trained models have inference failures. While previous studies have executed different types of attacks, one major challenge that greatly limits their ef
On the well-posedness of a nonlocal (two-place) FORQ equation via a two-component peakon system
math.APKenneth Karlsen, Yan Rybalko
We investigate the Cauchy problem for a nonlocal (two-place) FORQ equation. By interpreting this equation as a special case of a two-component peakon system (exhibiting a cubic nonlinearity), we convert the Cauchy problem into a system of ordinary differential equations in a Banach space. Using this approach, we are able to demonstrate local well-posedness i
Taesun Yeom, Minhyeok Lee
Class-conditional image generation using generative adversarial networks (GANs) has been investigated through various techniques; however, it continues to face challenges such as mode collapse, training instability, and low-quality output in cases of datasets with high intra-class variation. Furthermore, most GANs often converge in larger iterations, resulti
Mareike Dressler, Salma Kuhlmann, Moritz Schick
In this article, we combine sums of squares (SOS) and sums of nonnegative circuit (SONC) forms, two independent nonnegativity certificates for real homogeneous polynomials. We consider the convex cone SOS+SONC of forms that decompose into a sum of an SOS and a SONC form and study it from a geometric point of view. We show that the SOS+SONC cone is proper and
Anisha Gunjal, Greg Durrett
Past work has studied event prediction and event language modeling, sometimes mediated through structured representations of knowledge in the form of event schemas. Such schemas can lead to explainable predictions and forecasting of unseen events given incomplete information. In this work, we look at the process of creating such schemas to describe complex e
Introducing Competition to Boost the Transferability of Targeted Adversarial Examples through Clean Feature Mixup
cs.CVJunyoung Byun, Myung-Joon Kwon, Seungju Cho, Yoonji Kim
Deep neural networks are widely known to be susceptible to adversarial examples, which can cause incorrect predictions through subtle input modifications. These adversarial examples tend to be transferable between models, but targeted attacks still have lower attack success rates due to significant variations in decision boundaries. To enhance the transferab
Tiberiu Harko
We consider a generalization of the quintessence type scalar field cosmological models, by adding a multiplicative dissipative term in the scalar field Lagrangian, which is represented in an exponential form. The generalized dissipative Klein-Gordon equation is obtained from the variational principle in a covariant form. The energy-momentum tensor of the dis
Bruno Ebner, Norbert Henze, Simos Meintanis
We propose a general and relatively simple method for the construction of goodness-of-fit tests on the sphere and the hypersphere. The method is based on the characterization of probability distributions via their characteristic function, and it leads to test criteria that are convenient regarding applications and consistent against arbitrary deviations from
Hanxu Hu, Frank Keller
Current pre-trained vison-language models (PVLMs) achieve excellent performance on a range of multi-modal datasets. Recent work has aimed at building multilingual models, and a range of novel multilingual multi-modal datasets have been proposed. Current PVLMs typically perform poorly on these datasets when used for multi-modal zero-shot or few-shot cross-lin
Karthick Prasad Gunasekaran
Sentiment analysis (SA) is the automated process of detecting and understanding the emotions conveyed through written text. Over the past decade, SA has gained significant popularity in the field of Natural Language Processing (NLP). With the widespread use of social media and online platforms, SA has become crucial for companies to gather customer feedback
Deep Learning-based Bio-Medical Image Segmentation using UNet Architecture and Transfer Learning
eess.IVNima Hassanpour, Abouzar Ghavami
Image segmentation is a branch of computer vision that is widely used in real world applications including biomedical image processing. With recent advancement of deep learning, image segmentation has achieved at a very high level performance. Recently, UNet architecture is found as the core of novel deep learning segmentation methods. In this paper we imple
Hong Wang, Su Yang, Xiaoke Huang, Weishan Zhang
Token filtering to reduce irrelevant tokens prior to self-attention is a straightforward way to enable efficient vision Transformer. This is the first work to view token filtering from a feature selection perspective, where we weigh the importance of a token according to how much it can change the loss once masked. If the loss changes greatly after masking a
Yunshui Li, Binyuan Hui, ZhiChao Yin, Min Yang
Perceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence. Pre-training is commonly regarded as an effective approach for multi-modal dialogue. However, due to the limited availability of multi-modal dialogue data, there is still scarce research on multi-modal dialogue pre-training. Yet another int
Chenyang Le, Yao Qian, Long Zhou, Shujie Liu
Joint speech-language training is challenging due to the large demand for training data and GPU consumption, as well as the modality gap between speech and language. We present ComSL, a speech-language model built atop a composite architecture of public pretrained speech-only and language-only models and optimized data-efficiently for spoken language tasks.
Using the Uniqueness of Global Identifiers to Determine the Provenance of Python Software Source Code
cs.SEYiming Sun, Daniel M. German, Stefano Zacchiroli
We consider the problem of identifying the provenance of free/open source software (FOSS) and specifically the need of identifying where reused source code has been copied from. We propose a lightweight approach to solve the problem based on software identifiers-such as the names of variables, classes, and functions chosen by programmers. The proposed approa
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao
We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues. Compared to traditional VQA tasks, VQA in autonomous driving scenario presents more challenges. Firstly, the raw visual data are multi-modal, including images and point clouds captured by came
Haopeng Zhang, Xiao Liu, Jiawei Zhang
Text summarization systems have made significant progress in recent years, but typically generate summaries in one single step. However, the one-shot summarization setting is sometimes inadequate, as the generated summary may contain hallucinations or overlook essential details related to the reader's interests. This paper addresses this limitation by propos
Dhananjay Ashok, Zachary C. Lipton
In a surprising turn, Large Language Models (LLMs) together with a growing arsenal of prompt-based heuristics now offer powerful off-the-shelf approaches providing few-shot solutions to myriad classic NLP problems. However, despite promising early results, these LLM-based few-shot methods remain far from the state of the art in Named Entity Recognition (NER)
Jeremy L. Smallwood
In light of the recent confirmation of an eccentric orbit giant planet, $\beta$ Pic c, I revisit the formation and evolution of the warped debris disc in the system. $\beta$ Pic c is interior to $\beta$ Pic b, and the debris disc is exterior to both planets. Previous $N$-body simulations have shown that $\beta$ Pic b is responsible for exciting the inclinati
Stable photon orbits in stationary axisymmetric spacetimes with an electromagnetic field and a cosmological constant
gr-qcJake O. Shipley
Stable light rings, which are associated with spacetime instabilities, are known to exist in four-dimensional stationary axisymmetric spacetimes that solve the Einstein-Maxwell equations (so-called electrovacuum solutions, with Faraday tensor $F_{\mu \nu} \neq 0$); however, they are not permitted in pure vacuum ($F_{\mu \nu} = 0$). In this work, we extend th
Karol Capała, Bartłomiej Dybiec
Stochastic restarting is a strategy of starting anew. Incorporation of the resetting to the random walks can result in the decrease of the mean first passage time, due to the ability to limit unfavorably meandering, sub-optimal trajectories. In the following manuscript we examine how stochastic resetting influences escape dynamics from the $(-\infty,1)$ inte
Zhiwen Yan, Chen Li, Gim Hee Lee
Dynamic neural radiance fields (dynamic NeRFs) have demonstrated impressive results in novel view synthesis on 3D dynamic scenes. However, they often require complete video sequences for training followed by novel view synthesis, which is similar to playing back the recording of a dynamic 3D scene. In contrast, we propose OD-NeRF to efficiently train and ren
Bin Chen, Weidong Wang, Xia Zhao, Peibiao Zhao
In [Calc. Var., 57:5 (2018)], Hong-Ye-Zhang proposed the $p$-capacitary Orlicz-Minkowski problem and proved the existence of convex solutions to this problem by variational method for $p\in(1,n)$. However, the smoothness and uniqueness of solutions are still open. Notice that the $p$-capacitary Orlicz-Minkowski problem can be converted equivalently to a Mong
Deakin RF-Sensing: Experiments on Correlated Knowledge Distillation for Monitoring Human Postures with Radios
cs.CVShiva Raj Pokhrel, Jonathan Kua, Deol Satish, Philip Williams
In this work, we propose and develop a simple experimental testbed to study the feasibility of a novel idea by coupling radio frequency (RF) sensing technology with Correlated Knowledge Distillation (CKD) theory towards designing lightweight, near real-time and precise human pose monitoring systems. The proposed CKD framework transfers and fuses pose knowled
Towards Few-shot Entity Recognition in Document Images: A Graph Neural Network Approach Robust to Image Manipulation
cs.CLPrashant Krishnan, Zilong Wang, Yangkun Wang, Jingbo Shang
Recent advances of incorporating layout information, typically bounding box coordinates, into pre-trained language models have achieved significant performance in entity recognition from document images. Using coordinates can easily model the absolute position of each token, but they might be sensitive to manipulations in document images (e.g., shifting, rot
Mujeen Sung, James Gung, Elman Mansimov, Nikolaos Pappas
Intent classification (IC) plays an important role in task-oriented dialogue systems. However, IC models often generalize poorly when training without sufficient annotated examples for each user intent. We propose a novel pre-training method for text encoders that uses contrastive learning with intent psuedo-labels to produce embeddings that are well-suited
Xuhong Wang, Ding Wang, Liang Chen, Yilun Lin
Efficient traffic management is crucial for maintaining urban mobility, especially in densely populated areas where congestion, accidents, and delays can lead to frustrating and expensive commutes. However, existing prediction methods face challenges in terms of optimizing a single objective and understanding the complex composition of the transportation sys
Xiaojuan Tang, Zilong Zheng, Jiaqi Li, Fanxu Meng
The emergent few-shot reasoning capabilities of Large Language Models (LLMs) have excited the natural language and machine learning community over recent years. Despite of numerous successful applications, the underlying mechanism of such in-context capabilities still remains unclear. In this work, we hypothesize that the learned \textit{semantics} of langua
Michael J. Q. Zhang, Eunsol Choi
While large language models are able to retain vast amounts of world knowledge seen during pretraining, such knowledge is prone to going out of date and is nontrivial to update. Furthermore, these models are often used under temporal misalignment, tasked with answering questions about the present, despite having only been trained on data collected in the pas
Ian D. Roberts, Toby Brown, Nikki Zabel, Christine D. Wilson
We analyze cold-gas distributions in Virgo cluster galaxies using resolved CO(2-1) (tracing molecular hydrogen, H2) and HI observations from the Virgo Environment Traced In CO (VERTICO) and the VLA Imaging of Virgo in Atomic Gas (VIVA) surveys. From a theoretical perspective, it is expected that environmental processes in clusters will have a stronger influe
Niva Elkin-Koren, Uri Hacohen, Roi Livni, Shay Moran
There is a growing concern that generative AI models will generate outputs closely resembling the copyrighted materials for which they are trained. This worry has intensified as the quality and complexity of generative models have immensely improved, and the availability of extensive datasets containing copyrighted material has expanded. Researchers are acti
Attila Szolnoki, Xiaojie Chen
Competing strategies in an evolutionary game model, or species in a biosystem, can easily form a larger unit which protects them from the invasion of an external actor. Such a defensive alliance may have two, three, four or even more members. But how effective can be such formation against an alternative group composed by other competitors? To address this q
A new discretely divergence-free positivity-preserving high-order finite volume method for ideal MHD equations
math.NAShengrong Ding, Kailiang Wu
This paper proposes and analyzes a novel efficient high-order finite volume method for the ideal magnetohydrodynamics (MHD). As a distinctive feature, the method simultaneously preserves a discretely divergence-free (DDF) constraint on the magnetic field and the positivity-preserving (PP) property, which ensures the positivity of density, pressure, and inter
Chen Gong, Yvon Maday
Molecular representation learning (MRL) has long been crucial in the fields of drug discovery and materials science, and it has made significant progress due to the development of natural language processing (NLP) and graph neural networks (GNNs). NLP treats the molecules as one dimensional sequential tokens while GNNs treat them as two dimensional topology
Koyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh
We study the problem of optimizing data storage and access costs on the cloud while ensuring that the desired performance or latency is unaffected. We first propose an optimizer that optimizes the data placement tier (on the cloud) and the choice of compression schemes to apply, for given data partitions with temporal access predictions. Secondly, we propose
Jun Hui See Toh, Mengxin Du, Xinxin Tang, Ying Su
Understanding the interplay of interactions and disorder in quantum transport poses long-standing scientific challenges, with many-body quantum transport phenomena in high-dimensional disordered systems remaining largely unexplored experimentally. We utilize a momentum space lattice platform using quasi-periodically kicked ultracold atomic gases to experimen
Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee
In this paper, we investigate the problem of offline Preference-based Reinforcement Learning (PbRL) with human feedback where feedback is available in the form of preference between trajectory pairs rather than explicit rewards. Our proposed algorithm consists of two main steps: (1) estimate the implicit reward using Maximum Likelihood Estimation (MLE) with
Areej Alsini, Du Q. Huynh, Amitava Datta
Automatic evaluation of hashtag recommendation models is a fundamental task in many online social network systems. In the traditional evaluation method, the recommended hashtags from an algorithm are firstly compared with the ground truth hashtags for exact correspondences. The number of exact matches is then used to calculate the hit rate, hit ratio, precis
Dung Thai, Dhruv Agarwal, Mudit Chaudhary, Wenlong Zhao
We present an accurate and interpretable method for answer extraction in machine reading comprehension that is reminiscent of case-based reasoning (CBR) from classical AI. Our method (CBR-MRC) builds upon the hypothesis that contextualized answers to similar questions share semantic similarities with each other. Given a test question, CBR-MRC first retrieves
What functions can Graph Neural Networks compute on random graphs? The role of Positional Encoding
cs.LGNicolas Keriven, Samuel Vaiter
We aim to deepen the theoretical understanding of Graph Neural Networks (GNNs) on large graphs, with a focus on their expressive power. Existing analyses relate this notion to the graph isomorphism problem, which is mostly relevant for graphs of small sizes, or studied graph classification or regression tasks, while prediction tasks on nodes are far more rel
Yuhang Zang, Kaiyang Zhou, Chen Huang, Chen Change Loy
This paper focuses on long-tailed object detection in the semi-supervised learning setting, which poses realistic challenges, but has rarely been studied in the literature. We propose a novel pseudo-labeling-based detector called CascadeMatch. Our detector features a cascade network architecture, which has multi-stage detection heads with progressive confide
Samson Clymton, Hyun-Chul Kim
We investigate the dynamical generation of the $b_1$ meson in the $\pi\omega$ interaction, using the fully off-mass-shell coupled-channel formalism with the $\pi\omega$, $\eta\rho$, $\pi\phi$, and $K\bar{K}^*$ channels included. We first construct the Feynman amplitudes for the sixteen different kernel amplitudes, considering only the $t$ and $u$ channels. S
Magali Hersant
In this article, we identify the didactic conditions for learning mathematics in kindergarten. To do so, we rely on the framework of the theory of didactic situations (Brousseau, 1998) and the notion of problem-situation (Douady, 1984). We first explain what constitutes for us the stakes of teaching mathematics in kindergarten and then, based on examples, we
Pengfei Zhang
In this work, we extend the recent study of entropy dynamics induced by an external impulse in open quantum systems, where the entropy response follows the Page curve. For small system-bath coupling $\kappa$, we expect that the entropy first increases exponentially $\kappa^2 e^{\varkappa t}$ in the early-time regime $t\lesssim |\log \kappa|$ due to quantum m
Jiongnan Liu, Zhicheng Dou, Guoyu Tang, Sulong Xu
Recently, personalized product search attracts great attention and many models have been proposed. To evaluate the effectiveness of these models, previous studies mainly utilize the simulated Amazon recommendation dataset, which contains automatically generated queries and excludes cold users and tail products. We argue that evaluating with such a dataset ma
Proposition of Augmenting V2X Roadside Unit to Enhance Cooperative Awareness of Heterogeneously Connected Road Users
cs.NIKeyvan Ansari, Khondokar Fida Hasan
Intelligent transportation and autonomous mobility solutions rely on cooperative awareness developed by exchanging proximity and mobility data among road users. To maintain pervasive awareness on roads, all vehicles and vulnerable road users must be identified, either cooperatively, where road users equipped with wireless capabilities of Vehicle-to-Everythin
Yuwei Zhang, Zhi Jin, Zejun Wang, Ying Xing
Generating meaningful assert statements is one of the key challenges in automated test case generation, which requires understanding the intended functionality of the tested code. Recently, deep learning-based models have shown promise in improving the performance of assert statement generation. However, existing models only rely on the test prefixes along w
Gil Weinberg, Uri Weiss, Ori Katz
Fiber-based confocal endomicroscopy has shown great promise for minimally-invasive deep-tissue imaging. Despite its advantages, confocal fiber-bundle endoscopy inherently suffers from undersampling due to the spacing between fiber cores, and low collection efficiency when the target is not in proximity to the distal fiber facet. Here, we demonstrate an adapt
AWESOME: GPU Memory-constrained Long Document Summarization using Memory Mechanism and Global Salient Content
cs.CLShuyang Cao, Lu Wang
Long document summarization systems are critical for domains with lengthy and jargonladen text, yet they present significant challenges to researchers and developers with limited computing resources. Existing solutions mainly focus on efficient attentions or divide-and-conquer strategies. The former reduces theoretical time complexity, but is still memory-he
Provably convergent Newton-Raphson methods for recovering primitive variables with applications to physical-constraint-preserving Hermite WENO schemes for relativistic hydrodynamics
math.NAChaoyi Cai, Jianxian Qiu, Kailiang Wu
The relativistic hydrodynamics (RHD) equations have three crucial intrinsic physical constraints on the primitive variables: positivity of pressure and density, and subluminal fluid velocity. However, numerical simulations can violate these constraints, leading to nonphysical results or even simulation failure. Designing genuinely physical-constraint-preserv
Jean-Marie Malherbe, Florence Cornu, Isabelle Bualé
Systematic observations of the chromosphere and the photosphere started in Meudon Observatory 115 years ago with Deslandres spectroheliograph. An exceptional collection of more than 100 000 monochromatic images in CaII K and H$\alpha$ spanning more than 10 solar cycles is proposed to the international community by the BASS2000 solar database. We started in 2
Gilles Dowek
The induction principle for natural numbers expresses that when a property holds for some natural number a and is hereditary, then it holds for all numbers greater than or equal to a. We present a similar principle for real numbers.
Harvey Yiyun Fu, Qinyuan Ye, Albert Xu, Xiang Ren
Large Language Models (LLMs) have the impressive ability to perform in-context learning (ICL) from only a few examples, but the success of ICL varies widely from task to task. Thus, it is important to quickly determine whether ICL is applicable to a new task, but directly evaluating ICL accuracy can be expensive in situations where test data is expensive to
A sufficient condition for the lower semicontinuity of nonlocal supremal functionals in the vectorial case
math.APGiuliano Gargiulo, Elvira Zappale
In this note we present a sufficient condition ensuring lower semicontinuity for nonlocal supremal functionals of the type $$W^{1,\infty}(\Omega;\mathbb R^d)\ni u \mapsto \sup{\rm ess}_{(x,y)\in \Omega} W(x,y, \nabla u(x),\nabla u(y)),$$ where $\Omega$ is a bounded open subset of $\mathbb R^N$ and $W:\Omega \times \Omega \times \mathbb R^{d \times N}\times \
Xu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen
After discovering that Language Models (LMs) can be good in-context few-shot learners, numerous strategies have been proposed to optimize in-context sequence configurations. Recently, researchers in Vision-Language (VL) domains also develop their few-shot learners, while they only use the simplest way, ie., randomly sampling, to configure in-context image-te
Hoang Tien Nguyen, Young-Jin Kim, Dae-Hyun Choi
A surrogate model that accurately predicts distribution system voltages is crucial for reliable smart grid planning and operation. This letter proposes a fixed-point data-driven surrogate modeling method that employs a limited dataset to learn the power-voltage relationship of an unbalanced three-phase distribution system. The proposed surrogate model is des
Ying Cui, Junyi Liu, Jong-Shi Pang
There are many significant applied contexts that require the solution of discontinuous optimization problems in finite dimensions. Yet these problems are very difficult, both computationally and analytically. With the functions being discontinuous and a minimizer (local or global) of the problems, even if it exists, being impossible to verifiably compute, a
Multi-Abstractive Neural Controller: An Efficient Hierarchical Control Architecture for Interactive Driving
cs.ROXiao Li, Igor Gilitschenski, Guy Rosman, Sertac Karaman
As learning-based methods make their way from perception systems to planning/control stacks, robot control systems have started to enjoy the benefits that data-driven methods provide. Because control systems directly affect the motion of the robot, data-driven methods, especially black box approaches, need to be used with caution considering aspects such as
J. C. J. Koelemeij, H. Dun, C. E. V. Diouf, E. F. Dierikx
Global navigation satellite systems (GNSS) are widely used for navigation and time distribution, features indispensable for critical infrastructure such as mobile communication networks, as well as emerging technologies like automated driving and sustainable energy grids. While GNSS can provide centimetre-level precision, GNSS receivers are prone to many-met
Shunji Nishimura
In sequential circuits, the current output may depend on both past and current inputs. However, certain kinds of sequential circuits do not refer to all of the past inputs to generate the current output; they only refer to a subset of past inputs. This paper investigates which subset of past inputs a sequential circuit refers to, and proposes a new classific
Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts
The information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option. This has recently given rise to a range of techniques for injecting new facts through updating model weights. Current evaluation paradigms are extremely limited, mainly validating the recall of edited facts, but changing one f
Debiasing Made State-of-the-art: Revisiting the Simple Seed-based Weak Supervision for Text Classification
cs.CLChengyu Dong, Zihan Wang, Jingbo Shang
Recent advances in weakly supervised text classification mostly focus on designing sophisticated methods to turn high-level human heuristics into quality pseudo-labels. In this paper, we revisit the seed matching-based method, which is arguably the simplest way to generate pseudo-labels, and show that its power was greatly underestimated. We show that the li
Zhuoer Wang, Marcus Collins, Nikhita Vedula, Simone Filice
Methods to generate text from structured data have advanced significantly in recent years, primarily due to fine-tuning of pre-trained language models on large datasets. However, such models can fail to produce output faithful to the input data, particularly on out-of-domain data. Sufficient annotated data is often not available for specific domains, leading
ACE: Adversarial Correspondence Embedding for Cross Morphology Motion Retargeting from Human to Nonhuman Characters
cs.ROTianyu Li, Jungdam Won, Alexander Clegg, Jeonghwan Kim
Motion retargeting is a promising approach for generating natural and compelling animations for nonhuman characters. However, it is challenging to translate human movements into semantically equivalent motions for target characters with different morphologies due to the ambiguous nature of the problem. This work presents a novel learning-based motion retarge
Yongqi Li, Mayi Xu, Xin Miao, Shen Zhou
Large language models (LLMs) have made remarkable progress in a wide range of natural language understanding and generation tasks. However, their ability to generate counterfactuals has not been examined systematically. To bridge this gap, we present a comprehensive evaluation framework on various types of NLU tasks, which covers all key factors in determini
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
cs.CLFeng Jiang, Weihao Liu, Xiaomin Chu, Peifeng Li
Topic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and understand the overall context of the document from a higher l
Junyi Xie, Xinyi Yuan
We introduce a new approach to the geometric Bombieri--Lang conjecture for hyperbolic varieties in characteristic 0. The main idea is to construct an entire curve on a special fiber of a variety over a complex function field from an infinite sequence of rational points of the variety. The construction relies on the classical Brody lemma in complex geometry a
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen
Transformer-based language models (LMs) are powerful and widely-applicable tools, but their usefulness is constrained by a finite context window and the expensive computational cost of processing long text documents. We propose to adapt pre-trained LMs into AutoCompressors. These language models are capable of compressing long contexts into compact summary v
Michael Baltaxe, Tomer Pe'er, Dan Levi
Autonomous driving and advanced driver-assistance systems rely on a set of sensors and algorithms to perform the appropriate actions and provide alerts as a function of the driving scene. Typically, the sensors include color cameras, radar, lidar and ultrasonic sensors. Strikingly however, although light polarization is a fundamental property of light, it is
Takuya Aoyama, Kenya Ohgushi
We examined the piezomagnetic effect in an antiferromagnet composed of MnTe, which is a candidate material for altermagnetism with a high critical temperature. We observed that the magnetization develops with the application of stress and revealed that the piezomagnetic coefficient Q is 1.38$\times10^{-8}$ ${\mu}$B/MPa at 300 K. The poling-field dependence o
Victoria Basmov, Yoav Goldberg, Reut Tsarfaty
We evaluate LLMs' language understanding capacities on simple inference tasks that most humans find trivial. Specifically, we target (i) grammatically-specified entailments, (ii) premises with evidential adverbs of uncertainty, and (iii) monotonicity entailments. We design evaluation sets for these tasks and conduct experiments in both zero-shot and chain-of
Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, Ashwin Kalyan
Anthropomorphization is the tendency to attribute human-like traits to non-human entities. It is prevalent in many social contexts -- children anthropomorphize toys, adults do so with brands, and it is a literary device. It is also a versatile tool in science, with behavioral psychology and evolutionary biology meticulously documenting its consequences. With
Zihong Liang, Xiaojun Quan, Qifan Wang
Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts. Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phonetic representations with character representations, which tends to weaken the representation effect of normal texts. In this work, we propo
Pengyuan Lu, Michele Caprio, Eric Eaton, Insup Lee
Algorithms that balance the stability-plasticity trade-off are well studied in the Continual Learning literature. However, only a few focus on obtaining models for specified trade-off preferences. When solving the problem of continual learning under specific trade-offs (CLuST), state-of-the-art techniques leverage rehearsal-based learning, which requires ret
Accelerated Nonconvex ADMM with Self-Adaptive Penalty for Rank-Constrained Model Identification
math.OCQingyuan Liu, Zhengchao Huang, Hao Ye, Dexian Huang
The alternating direction method of multipliers (ADMM) has been widely adopted in low-rank approximation and low-order model identification tasks; however, the performance of nonconvex ADMM is highly reliant on the choice of penalty parameter. To accelerate ADMM for solving rank-constrained identification problems, this paper proposes a new self-adaptive str
F. H. Haydarov
In this paper, we shall discuss the extendability of probability and non-probability measures on Cayley trees to a $\sigma$-additive measure on Borel fields which has a fundamental role in the theory of Gibbs measures.
Bo Xue, Kayode Adedotun Oyesina, Alex M. H. Wong
Electric dipoles and magnetic dipoles are the most fundamental particles in electromagnetic theory. Huygens and Janus sources, formed by the orthogonal combination of electric and magnetic dipoles, both show good directionality in the near field. Although the Huygens source has been widely used in antennas and metasurfaces, the applications of Janus source a
Nikita Srivatsan, Sofia Samaniego, Omar Florez, Taylor Berg-Kirkpatrick
In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that d
Xiyuan Wang, Fangyuan Wang, Bo Xu, Liang Xu
Typically, the Time-Delay Neural Network (TDNN) and Transformer can serve as a backbone for Speaker Verification (SV). Both of them have advantages and disadvantages from the perspective of global and local feature modeling. How to effectively integrate these two style features is still an open issue. In this paper, we explore a Parallel-coupled TDNN/Transfo
Jaemoo Choi, Jaewoong Choi, Myungjoo Kang
Optimal Transport (OT) problem investigates a transport map that bridges two distributions while minimizing a given cost function. In this regard, OT between tractable prior distribution and data has been utilized for generative modeling tasks. However, OT-based methods are susceptible to outliers and face optimization challenges during training. In this pap
Yuchen Ding, Lilu Zhao
Let $k\ge 2$ be a positive integer and $P^+(n)$ the greatest prime factor of a positive integer $n$ with convention $P^+(1)=1$. For any $\theta\in \left[\frac 1{2k},\frac{17}{32k}\right)$, set $$T_{k,\theta}(x)=\sum_{\substack{p_1\cdot\cdot\cdot p_k\le x\\ P^+(\gcd(p_1-1,...,p_k-1))\ge (p_1\cdot\cdot\cdot p_k)^\theta}}1,$$ where the $p'$s are primes. It is p
Amirhossein Kazemnejad, Mehdi Rezagholizadeh, Prasanna Parthasarathi, Sarath Chandar
While pre-trained language models (PLMs) have shown evidence of acquiring vast amounts of knowledge, it remains unclear how much of this parametric knowledge is actually usable in performing downstream tasks. We propose a systematic framework to measure parametric knowledge utilization in PLMs. Our framework first extracts knowledge from a PLM's parameters a
M. C. Gordillo, J. Boronat
Using a diffusion Monte Carlo (DMC) technique, we calculated the phase diagrams of $^4$He and H$_2$ adsorbed on a single (5,5) carbon nanotube, one of the narrowest that can be obtained experimentally. For a single monolayer, when the adsorbate density increases, both species undergo a series of first order solid-solid phase transitions between incommensurat
Hogyun Kim, Gilhwan Kang, Seokhwan Jeong, Seungjun Ma
Place recognition using SOund Navigation and Ranging (SONAR) images is an important task for simultaneous localization and mapping(SLAM) in underwater environments. This paper proposes a robust and efficient imaging SONAR based place recognition, SONAR context, and loop closure method. Unlike previous methods, our approach encodes geometric information based
A Question Answering Framework for Decontextualizing User-facing Snippets from Scientific Documents
cs.CLBenjamin Newman, Luca Soldaini, Raymond Fok, Arman Cohan
Many real-world applications (e.g., note taking, search) require extracting a sentence or paragraph from a document and showing that snippet to a human outside of the source document. Yet, users may find snippets difficult to understand as they lack context from the original document. In this work, we use language models to rewrite snippets from scientific d
David helps Goliath: Inference-Time Collaboration Between Small Specialized and Large General Diffusion LMs
cs.CLXiaochuang Han, Sachin Kumar, Yulia Tsvetkov, Marjan Ghazvininejad
Diffusion-based language models are emerging as a promising alternative to autoregressive LMs: they approach the competence of autoregressive LMs while offering nuanced controllability at inference time. While autoregressive LMs have benefited immensely from scaling and instruction-based learning, existing studies of diffusion LMs have been conducted on a sm
Manya Wadhwa, Jifan Chen, Junyi Jessy Li, Greg Durrett
The rise of large language models (LLMs) has brought a critical need for high-quality human-labeled data, particularly for processes like human feedback and evaluation. A common practice is to label data via consensus annotation over human judgments. However, annotators' judgments for subjective tasks can differ in many ways: they may reflect different quali
Raúl Fuentes-Azcatl, Minerva González-Melchor
In this work, a new force field is presented for the ionic liquid 1-butyl-3-methylimidazolium-bis(trifluoromethylsulfonyl)imide, [Bmim][Nf$_2$T]. As a part of the \epsilon force field, the acronym IL/\epsilon is used to refer to this ionic liquid. This new force field reproduces the dielectric constant, the density, and the entalphy of vaporization, with an
Zhengkai Jiang, Liang Liu, Jiangning Zhang, Yabiao Wang
This paper introduces a novel attention mechanism, called dual attention, which is both efficient and effective. The dual attention mechanism consists of two parallel components: local attention generated by Convolutional Neural Networks (CNNs) and long-range attention generated by Vision Transformers (ViTs). To address the high computational complexity and
Interpretation and visualization of distance covariance through additive decomposition of correlations formula
stat.MEAndi Wang, Hao Yan, Juan Du
Distance covariance is a widely used statistical methodology for testing the dependency between two groups of variables. Despite the appealing properties of consistency and superior testing power, the testing results of distance covariance are often hard to be interpreted. This paper presents an elementary interpretation of the mechanism of distance covarian
Hao Sun, Xiao Liu, Yeyun Gong, Yan Zhang
With the advance of large language models (LLMs), the research field of LLM applications becomes more and more popular and the idea of constructing pipelines to accomplish complex tasks by stacking LLM API calls come true. However, this kind of methods face two limitations: narrow information coverage and low fault tolerance. In this work, we propose a novel
Insung Kong, Dongyoon Yang, Jongjin Lee, Ilsang Ohn
Bayesian approaches for learning deep neural networks (BNN) have been received much attention and successfully applied to various applications. Particularly, BNNs have the merit of having better generalization ability as well as better uncertainty quantification. For the success of BNN, search an appropriate architecture of the neural networks is an importan
Detection of Non-uniformity in Parameters for Magnetic Domain Pattern Generation by Machine Learning
cond-mat.mtrl-sciNaoya Mamada, Masaichiro Mizumaki, Ichiro Akai, Toru Aonishi
We estimate the spatial distribution of heterogeneous physical parameters involved in the formation of magnetic domain patterns of polycrystalline thin films by using convolutional neural networks. We propose a method to obtain a spatial map of physical parameters by estimating the parameters from patterns within a small subregion window of the full magnetic
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou
The escalating debate on AI's capabilities warrants developing reliable metrics to assess machine "intelligence". Recently, many anecdotal examples were used to suggest that newer large language models (LLMs) like ChatGPT and GPT-4 exhibit Neural Theory-of-Mind (N-ToM); however, prior work reached conflicting conclusions regarding those abilities. We investi
Taishi Kurahashi, Yuta Sato
We study the finite frame property of some extensions of Fitting, Marek, and Truszczy\'nski's pure logic of necessitation $\mathbf{N}$. For any natural numbers $m, n$, we introduce the logic $\mathbf{N}^+\mathbf{A}_{m, n}$ by adding the single axiom scheme $\Box^n \varphi \to \Box^m \varphi$ and the rule $\dfrac{\neg \Box \varphi}{\neg \Box \Box \varphi}$ (R
Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque
Charts are very popular for analyzing data, visualizing key insights and answering complex reasoning questions about data. To facilitate chart-based data analysis using natural language, several downstream tasks have been introduced recently such as chart question answering and chart summarization. However, most of the methods that solve these tasks use pret
Bi-Drop: Enhancing Fine-tuning Generalization via Synchronous sub-net Estimation and Optimization
cs.CLShoujie Tong, Heming Xia, Damai Dai, Runxin Xu
Pretrained language models have achieved remarkable success in natural language understanding. However, fine-tuning pretrained models on limited training data tends to overfit and thus diminish performance. This paper presents Bi-Drop, a fine-tuning strategy that selectively updates model parameters using gradients from various sub-nets dynamically generated
Combining direct and indirect sparse data for learning generalizable turbulence models
physics.flu-dynXin-Lei Zhang, Heng Xiao, Xiaodong Luo, Guowei He
Learning turbulence models from observation data is of significant interest in discovering a unified model for a broad range of practical flow applications. Either the direct observation of Reynolds stress or the indirect observation of velocity has been used to improve the predictive capacity of turbulence models. In this work, we propose combining the dire
Tianlun Zheng, Zhineng Chen, BingChen Huang, Wei Zhang
Multilingual text recognition (MLTR) systems typically focus on a fixed set of languages, which makes it difficult to handle newly added languages or adapt to ever-changing data distribution. In this paper, we propose the Incremental MLTR (IMLTR) task in the context of incremental learning (IL), where different languages are introduced in batches. IMLTR is p