May 2023 arXiv papers — page 61
Showing 6,001–6,100 of 19,695 papers
Yongtao Liu, Anna N. Morozovska, Ayana Ghosh, Kyle P. Kelley
Nanoscale ferroelectric 2D materials offer unique opportunity to investigate curvature and strain effects on materials functionalities. Among these, CuInP2S6 (CIPS) has attracted tremendous research interest in recent years due to combination of room temperature ferroelectricity, scalability to a few layers thickness, and unique ferrielectric properties due
Haldun Özgür Bayındır
Using the root adjunction formalism developed in an earlier work and logarithmic THH, we obtain a simplified computation of $T(2)_*\text{K}(ku)$ for $p>3$. Through this, we also produce a new algebraic $K$-theory computation; namely we obtain $T(2)_*\text{K}(ku/p)$, where $ku/p$ is the $2$-periodic Morava $K$-theory spectrum of height $1$.
Debiasing should be Good and Bad: Measuring the Consistency of Debiasing Techniques in Language Models
cs.CLRobert Morabito, Jad Kabbara, Ali Emami
Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods that yield not only desirable results, but are also consistent with their mechanisms and specifications. For example, we a
Hierarchical Adaptive Voxel-guided Sampling for Real-time Applications in Large-scale Point Clouds
cs.CVJunyuan Ouyang, Xiao Liu, Haoyao Chen
While point-based neural architectures have demonstrated their efficacy, the time-consuming sampler currently prevents them from performing real-time reasoning on scene-level point clouds. Existing methods attempt to overcome this issue by using random sampling strategy instead of the commonly-adopted farthest point sampling~(FPS), but at the expense of lowe
Esteban González, Kimet Jusufi, Genly Leon, Emmanuel N. Saridakis
We confront Yukawa modified cosmology, proposed in arXiv:2304.11492 [Jusufi et al. arXiv:2304.11492], with data from Supernovae Type Ia (SNe Ia) and Hubble parameter (OHD) observations. Yukawa cosmology is obtained from a Yukawa-like gravitational potential, with coupling parameter $\alpha$ and wavelength parameter $\lambda$, which gives rise to modified Fri
Fang Zhang, Xing Zhu, Rui Chao, Cupjin Huang
Scaling bottlenecks the making of digital quantum computers, posing challenges from both the quantum and the classical components. We present a classical architecture to cope with a comprehensive list of the latter challenges {\em all at once}, and implement it fully in an end-to-end system by integrating a multi-core RISC-V CPU with our in-house control ele
Yilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi
People primarily consult tables to conduct data analysis or answer specific questions. Text generation systems that can provide accurate table summaries tailored to users' information needs can facilitate more efficient access to relevant data insights. Motivated by this, we define a new query-focused table summarization task, where text generation models ha
Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu
Computer Vision (CV), Natural Language Processing (NLP), and Recommender Systems (RecSys) are three prominent AI applications that have traditionally developed independently, resulting in disparate modeling and engineering methodologies. This has impeded the ability for these fields to directly benefit from each other's advancements. With the recent developm
Fangda Li, Zhiqiang Hu, Wen Chen, Avinash Kak
Hematoxylin and Eosin (H&E) staining is a widely used sample preparation procedure for enhancing the saturation of tissue sections and the contrast between nuclei and cytoplasm in histology images for medical diagnostics. However, various factors, such as the differences in the reagents used, result in high variability in the colors of the stains actually re
Orr Fischer, Merav Parter
In their seminal PODC 1991 paper, Ostrovsky and Yung introduced the study of distributed computation in the presence of mobile adversaries which can dynamically appear throughout the network. Over the years, this setting has been studied mostly under the assumption that the communication graph is fully-connected. Resilient CONGEST algorithms for general grap
Minsik Oh, Jiwei Li, Guoyin Wang
Learning high quality sentence embeddings from dialogues has drawn increasing attentions as it is essential to solve a variety of dialogue-oriented tasks with low annotation cost. Annotating and gathering utterance relationships in conversations are difficult, while token-level annotations, \eg, entities, slots and templates, are much easier to obtain. Other
En Yu, Tiancai Wang, Zhuoling Li, Yuang Zhang
Although end-to-end multi-object trackers like MOTR enjoy the merits of simplicity, they suffer from the conflict between detection and association seriously, resulting in unsatisfactory convergence dynamics. While MOTRv2 partly addresses this problem, it demands an additional detection network for assistance. In this work, we serve as the first to reveal th
Thomas Izgin, David I. Ketcheson, Andreas Meister
In recent years, many positivity-preserving schemes for initial value problems have been constructed by modifying a Runge--Kutta (RK) method by weighting the right-hand side of the system of differential equations with solution-dependent factors. These include the classes of modified Patankar--Runge--Kutta (MPRK) and Geometric Conservative (GeCo) methods. Co
Kundan Krishna, Prakhar Gupta, Sanjana Ramprasad, Byron C. Wallace
While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived benchmark, complemented by a rich set of crowd-sourced annotations, that supports $8$ interrelated tasks: (i) extractive summa
Francesco Giovanni Celiberto
In this review we discuss and extend the study of the inclusive production of vector quarkonia, $J/\psi$ and $\Upsilon$, emitted with large transverse momenta and rapidities at the LHC. We adopt the novel ZCW19$^+$ determination to depict the quarkonium production mechanism at the next-to-leading level of perturbative QCD. This approach is based on the nonre
Alessandro Sinibaldi, Clemens Giuliani, Giuseppe Carleo, Filippo Vicentini
We analyze the accuracy and sample complexity of variational Monte Carlo approaches to simulate the dynamics of many-body quantum systems classically. By systematically studying the relevant stochastic estimators, we are able to: (i) prove that the most used scheme, the time-dependent Variational Monte Carlo (tVMC), is affected by a systematic statistical bi
On the motion of satellite around the natural moons of planets using the concept of ER3BP with variable eccentricity
physics.gen-phSergey Ershkov, Dmytro Leshchenko, Evgeniy Prosviryakov
In the current study, we explore stability of motion of satellite around the natural moons of planets in Solar system using the novel concept of ER3BP with variable eccentricity. This concept was introduced earlier when novel type of ER3BP (Sun-planet-satellite) was investigated with variable spin state of secondary planet correlated implicitly to the motion
Chenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos
Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata knowledge base. However, models trained only on Wikipedia have limitations when applied to web domains, which often contain
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
cs.CLSina J. Semnani, Violet Z. Yao, Heidi C. Zhang, Monica S. Lam
This paper presents the first few-shot LLM-based chatbot that almost never hallucinates and has high conversationality and low latency. WikiChat is grounded on the English Wikipedia, the largest curated free-text corpus. WikiChat generates a response from an LLM, retains only the grounded facts, and combines them with additional information it retrieves from
Nicholas Deas, Jessi Grieser, Shana Kleiner, Desmond Patton
We evaluate how well LLMs understand African American Language (AAL) in comparison to their performance on White Mainstream English (WME), the encouraged "standard" form of English taught in American classrooms. We measure LLM performance using automatic metrics and human judgments for two tasks: a counterpart generation task, where a model generates AAL (or
Karol Gietka, Christoph Hotter, Helmut Ritsch
Squeezing is essential to many quantum technologies and our understanding of quantum physics. Here we develop a theory of steady-state squeezing that can be generated in the closed and open quantum Rabi as well as Dicke model. To this end, we eliminate the spin dynamics which effectively leads to an abstract harmonic oscillator whose eigenstates are squeezed
Xili Yi, Nima Fazeli
In this paper, we discuss the mechanics and planning algorithms to slide an object on a horizontal planar surface via frictional patch contact made with its top surface. Here, we propose an asymmetric dual limit surface model to determine slip boundary conditions for both the top and bottom contact. With this model, we obtain a range of twists that can keep
Chenxi Whitehouse, Monojit Choudhury, Alham Fikri Aji
This paper explores the potential of leveraging Large Language Models (LLMs) for data augmentation in multilingual commonsense reasoning datasets where the available training data is extremely limited. To achieve this, we utilise several LLMs, namely Dolly-v2, StableVicuna, ChatGPT, and GPT-4, to augment three datasets: XCOPA, XWinograd, and XStoryCloze. Sub
Max Weinreich
We introduce an algebraic formulation of billiards on plane curves over algebraically closed fields, extending Glutsyuk's complex billiards. For any smooth algebraic curve $C$ of degree $d \geq 2$, algebraic billiards is a rational $(d-1)$-to-$(d-1)$ surface correspondence on the space of unit tangent vectors based on $C$. We prove that the dynamical degree
Koen Minartz, Yoeri Poels, Simon Koop, Vlado Menkovski
Neural networks are emerging as a tool for scalable data-driven simulation of high-dimensional dynamical systems, especially in settings where numerical methods are infeasible or computationally expensive. Notably, it has been shown that incorporating domain symmetries in deterministic neural simulators can substantially improve their accuracy, sample effici
Matteo Piccolini, Vittorio Giovannetti, Rosario Lo Franco
We propose a procedure for the robust preparation of maximally entangled states of identical fermionic qubits, studying the role played by particle statistics in the process. The protocol exploits externally activated noisy channels to reset the system to a known state. The subsequent interference effects generated at a beam splitter result in a mixture of m
Transport properties of hybrid single-bilayer graphene interfaces in magnetic field
cond-mat.mes-hallNadia Benlakhouy, Ahmed Jellal, Michael Schreiber
We investigate the electronic properties of a hybrid system that comprises single-bilayer graphene structures subjected to a perpendicular magnetic field. Specifically, our focus is on the behavior exhibited by the zigzag boundaries of the junction, namely Zigzag-1 (ZZ1) and Zigzag-2 (ZZ2), using the continuum Dirac model for rigorous analysis. Our findings
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao
Large Language Models (LLMs) play powerful, black-box readers in the retrieve-then-read pipeline, making remarkable progress in knowledge-intensive tasks. This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting. Unlike prior studies foc
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song
Automatically evaluating the quality of language generation is critical. Although recent learned metrics show high correlation with human judgement, these metrics can not explain their verdict or associate the scores with defects in generated text. To address this limitation, we present InstructScore, an explainable evaluation metric for text generation. By
Emanuele Bugliarello, Aida Nematzadeh, Lisa Anne Hendricks
Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations. In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data. In particular, we propose two pretraining approaches to contextualise vi
Elizabeth Salesky, Neha Verma, Philipp Koehn, Matt Post
We introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations. We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings. We explore various properties of pixel representations such as parameter sha
Angelica Chen, Jason Phang, Alicia Parrish, Vishakh Padmakumar
Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consistency. We argue that self-consistency is an important criteria for valid multi-step reasoning in tasks where the solution is composed of the answers to multiple sub-steps. We propos
Combined funnel, concentrator, and particle valve functional element for magnetophoretic bead transport based on engineered magnetic domain patterns
physics.app-phRico Huhnstock, Lukas Paetzold, Maximilian Merkel, Piotr Kuświk
Controlled actuation of superparamagnetic beads (SPBs) within a microfluidic environment using tailored dynamic magnetic field landscapes (MFLs) is a potent approach for the realization of point-of-care diagnostics within Lab-on-a-chip (LOC) systems. Making use of an engineered magnetic domain pattern as the MFL source, a functional LOC-element with combined
Ammar Ahmed, Anwar Said, Mudassir Shabbir, Xenofon Koutsoukos
Vulnerability identification constitutes a task of high importance for cyber security. It is quite helpful for locating and fixing vulnerable functions in large applications. However, this task is rather challenging owing to the absence of reliable and adequately managed datasets and learning models. Existing solutions typically rely on human expertise to an
Emergent correlated phases in rhombohedral trilayer graphene induced by proximity spin-orbit and exchange coupling
cond-mat.str-elYaroslav Zhumagulov, Denis Kochan, Jaroslav Fabian
The impact of proximity-induced spin-orbit and exchange coupling on the correlated phase diagram of rhombohedral trilayer graphene (RTG) is investigated theoretically. By employing \emph{ab initio}-fitted effective models of RTG encapsulated by transition metal dichalcogenides (spin-orbit proximity effect) and ferromagnetic Cr$_2$Ge$_2$Te$_6$ (exchange proxi
Ada Chan, Peter Sin
In a continuous-time quantum walk on a network of qubits, pretty good state transfer is the phenomenon of state transfer between two vertices with fidelity arbitrarily close to 1. We construct families of graphs to demonstrate that there is no bound on the size of a set of vertices that admit pretty good state transfer between any two vertices of the set.
Yash Patel, Declan McNamara, Jackson Loper, Jeffrey Regier
Amortized variational inference is an often employed framework in simulation-based inference that produces a posterior approximation that can be rapidly computed given any new observation. Unfortunately, there are few guarantees about the quality of these approximate posteriors. We propose Conformalized Amortized Neural Variational Inference (CANVI), a proce
B. V. Rajarama Bhat, Purbayan Chakraborty, Uwe Franz
The Weyl operators give a convenient basis of $M_n(\mathbb{C})$ which is also orthonormal with respect to the Hilbert-Schmidt inner product. The properties of such a basis can be generalised to the notion of a nice error basis(NEB), as introduced by E. Knill. We can use an NEB of $M_n(\mathbb{C})$ to construct an NEB for $Lin(M_n(\mathbb{C}))$, the space of
Hyun Jeong, Kohei Kamada, Alexei A. Starobinsky, Jun'ichi Yokoyama
Post-inflationary evolution and (re)heating of the viable inflationary model, the $R^2$ one, is made more realistic by including the leptogenesis scenario into it. For this purpose, right-handed Majorana neutrinos with a large mass are added to the matter sector of the Standard Model to explain the neutrino oscillation experiments and the baryon asymmetry of
Kyle DeBry, Jasmine Sinanan-Singh, Colin D. Bruzewicz, David Reens
We present experimental demonstrations of accurate and unambiguous single-shot discrimination between three quantum channels using a single trapped $^{40}\text{Ca}^{+}$ ion. The three channels cannot be distinguished unambiguously using repeated single channel queries, the natural classical analogue. We develop techniques for using the 6-dimensional $\text{D
Fast and energy-efficient non-volatile III-V-on-silicon photonic phase shifter based on memristors
physics.opticsZhuoran Fang, Bassem Tossoun, Antoine Descos, Di Liang
Silicon photonics has evolved from lab research to commercial products in the past decade as it plays an increasingly crucial role in data communication for next-generation data centers and high performance computing1. Recently, programmable silicon photonics has also found new applications in quantum2 and classical 3 information processing. A key component
NCC: Natural Concurrency Control for Strictly Serializable Datastores by Avoiding the Timestamp-Inversion Pitfall
cs.DCHaonan Lu, Shuai Mu, Siddhartha Sen, Wyatt Lloyd
Strictly serializable datastores greatly simplify the development of correct applications by providing strong consistency guarantees. However, existing techniques pay unnecessary costs for naturally consistent transactions, which arrive at servers in an order that is already strictly serializable. We find these transactions are prevalent in datacenter worklo
Giulia Rizzoli, Donald Shenaj, Pietro Zanuttigh
With the increasing availability of depth sensors, multimodal frameworks that combine color information with depth data are gaining interest. However, ground truth data for semantic segmentation is burdensome to provide, thus making domain adaptation a significant research area. Yet most domain adaptation methods are not able to effectively handle multimodal
Zi-Yi Dou, Feng Gao, Nanyun Peng
Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited availability of training data, which hinders the models' ability to generalize effectively. Previous approaches have attempted to address this issue by introducing additional superv
Martin Gonzalez, Nelson Fernandez, Thuy Tran, Elies Gherbi
A potent class of generative models known as Diffusion Probabilistic Models (DPMs) has become prominent. A forward diffusion process adds gradually noise to data, while a model learns to gradually denoise. Sampling from pre-trained DPMs is obtained by solving differential equations (DE) defined by the learnt model, a process which has shown to be prohibitive
Dmitry Svintsov, Georgy Alymov
Despite numerous applications of two-dimensional plasmons for electromagnetic energy manipulation at the nanoscale, their quantitative refraction and reflection laws (analogs of Fresnel formulas in optics) have not yet been established. This fact can be traced down to the strong non-locality of equations governing the 2d plasmon propagation. Here, we tackle
Timothy B. Armstrong, Patrick Kline, Liyang Sun
Empirical research typically involves a robustness-efficiency tradeoff. A researcher seeking to estimate a scalar parameter can invoke strong assumptions to motivate a restricted estimator that is precise but may be heavily biased, or they can relax some of these assumptions to motivate a more robust, but variable, unrestricted estimator. When a bound on the
Katerina Margatina, Timo Schick, Nikolaos Aletras, Jane Dwivedi-Yu
The remarkable advancements in large language models (LLMs) have significantly enhanced the performance in few-shot learning settings. By using only a small number of labeled examples, referred to as demonstrations, LLMs can effectively grasp the task at hand through in-context learning. However, the process of selecting appropriate demonstrations has receiv
LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages
cs.CLMilind Agarwal, Md Mahfuz Ibn Alam, Antonios Anastasopoulos
Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems. Language identification is a well-studied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, current systems cannot accurately identify most of the world'
A partially stripped massive star in a Be binary at low metallicity: A missing link towards Be X-ray binaries and double neutron star mergers
astro-ph.SRV. Ramachandran, J. Klencki, A. A. C. Sander, D. Pauli
Standard binary evolutionary models predict a significant population of core helium-burning stars that lost their hydrogen-rich envelope after mass transfer via Roche-lobe overflow. However, there is a scarcity of observations of such stripped stars in the intermediate mass regime (~1.5 - 8$ M_{\odot}$), which are thought to be prominent progenitors of SN Ib
V. M. Jiménez, M. De León
In this paper, we study internal properties of a Cosserat media. In fact, by using groupoids and smooth distributions, we obtain a three canonical equations. The \textit{non-holonomic material equation for Cosserat media} characterizes the uniformity of the material. The \textit{holonomic material equation for Cosserat media} permits us to study when a Cosse
Yue Fan, Jing Gu, Kaizhi Zheng, Xin Eric Wang
Intelligent navigation-helper agents are critical as they can navigate users in unknown areas through environmental awareness and conversational ability, serving as potential accessibility tools for individuals with disabilities. In this work, we first introduce a novel benchmark, Respond to Help Requests (R2H), to promote the development of multi-modal navi
Qingyun Wang, Doug Downey, Heng Ji, Tom Hope
We explore and enhance the ability of neural language models to generate novel scientific directions grounded in literature. Work on literature-based hypothesis generation has traditionally focused on binary link prediction--severely limiting the expressivity of hypotheses. This line of work also does not focus on optimizing novelty. We take a dramatic depar
Zheng Xie, Yu Liu, Hao-Yuan He, Ming Li
Since acquiring perfect supervision is usually difficult, real-world machine learning tasks often confront inaccurate, incomplete, or inexact supervision, collectively referred to as weak supervision. In this work, we present WSAUC, a unified framework for weakly supervised AUC optimization problems, which covers noisy label learning, positive-unlabeled lear
Abishek Sridhar, Robert Lo, Frank F. Xu, Hao Zhu
Large language models (LLMs) struggle on processing complicated observations in interactive decision making tasks. To alleviate this issue, we propose a simple hierarchical prompting approach. Diverging from previous prompting approaches that always put the full observation (e.g. a web page) to the prompt, we propose to first construct an action-aware observ
Oleg Vasilyev, Fumika Isono, John Bohannon
Semantics of a sentence is defined with much less ambiguity than semantics of a single word, and we assume that it should be better preserved by translation to another language. If multilingual sentence embeddings intend to represent sentence semantics, then the similarity between embeddings of any two sentences must be invariant with respect to translation.
Tanchumin Xu, Yunshu Zhang, Shu Yang
Propensity score matching (PSM) and augmented inverse propensity weighting (AIPW) are widely used in observational studies to estimate causal effects. The two approaches present complementary features. The AIPW estimator is doubly robust and locally efficient but can be unstable when the propensity scores are close to zero or one due to weighting by the inve
Yiyun Fan, John Billingham, Kristoffer van der Zee
We develop a shape-Newton method for solving generic free-boundary problems where one of the free-boundary conditions is governed by the Bernoulli equation. The Newton-like scheme is developed by employing shape derivatives in the weak forms, which allows us to update the position of the free surface and the potential on the free boundary by solving a bounda
Christopher Hughes, Greg Martin, Andrew Pearce-Crump
Shanks conjectured that $\zeta ' (\rho)$, where $\rho$ ranges over non-trivial zeros of the Riemann zeta function, is real and positive in the mean. We present a history of this problem, including a generalisation to all higher-order derivatives $\zeta^{(n)}(s)$, for which the sign of the mean alternatives between positive for odd $n$ and negative for even $
Quantum Kolmogorov complexity and quantum correlations in deterministic-control quantum Turing machines
quant-phMariano Lemus, Ricardo Faleiro, Paulo Mateus, Nikola Paunković
This work presents a study of Kolmogorov complexity for general quantum states from the perspective of deterministic-control quantum Turing Machines (dcq-TM). We extend the dcq-TM model to incorporate mixed state inputs and outputs, and define dcq-computable states as those that can be approximated by a dcq-TM. Moreover, we introduce (conditional) Kolmogorov
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis
Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate, and (2) human evaluation is time-consuming and costly. In this paper, we introduce FACTSCORE, a new evaluation th
Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson
While large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs". This lack of interpretability is a growing impediment to widespread use of LLMs. To address this, our goals are to make model beliefs and their inferential relationships explicit, and to resolve
Blake Jackson
This paper proves the existence of an intersection-dimension formula for preprojective modules over path algebras of type $\widetilde{D}_n$. Identical intersection-dimension formulas have previously been provided for modules over path algebras of type $A_n, D_n,$ and $\widetilde{A}_n$ due to Schiffler as well as He, Zhou, and Zhu. These modules can be repres
Improved rates of convergence for the multivariate Central Limit Theorem in Wasserstein distance
math.PRThomas Bonis
We provide new bounds for the rate of convergence of the multivariate Central Limit Theorem in Wasserstein distances of order $p \geq 2$. In particular, we obtain what we conjecture to be the asymptotically optimal rate whenever the density of the summands admits a non-zero continuous component and has a non-zero third moment.
Evaluation of the MACE Force Field Architecture: from Medicinal Chemistry to Materials Science
physics.chem-phDavid Peter Kovacs, Ilyes Batatia, Eszter Sara Arany, Gabor Csanyi
The MACE architecture represents the state of the art in the field of machine learning force fields for a variety of in-domain, extrapolation and low-data regime tasks. In this paper, we further evaluate MACE by fitting models for published benchmark datasets. We show that MACE generally outperforms alternatives for a wide range of systems from amorphous car
Jocelyn Shen, Maarten Sap, Pedro Colon-Hernandez, Hae Won Park
The most meaningful connections between people are often fostered through expression of shared vulnerability and emotional experiences in personal narratives. We introduce a new task of identifying similarity in personal stories based on empathic resonance, i.e., the extent to which two people empathize with each others' experiences, as opposed to raw semant
Gregorio García-Valladares, Carlos A. Plata, Antonio Prados, Alessandro Manacorda
In many physical situations, there appears the problem of reaching a single target that is spatially distributed. Here we analyse how stochastic resetting, also spatially distributed, can be used to improve the search process when the target location is quenched, i.e. it does not evolve in time. More specifically, we consider a model with minimal but suffici
Shengchao Chen, Guodong Long, Tao Shen, Jing Jiang
On-device intelligence for weather forecasting uses local deep learning models to analyze weather patterns without centralized cloud computing, holds significance for supporting human activates. Federated Learning is a promising solution for such forecasting by enabling collaborative model training without sharing raw data. However, it faces three main chall
Manuel Tran, Yashin Dicente Cid, Amal Lahiani, Fabian J. Theis
Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer are datasets that align all three modalities at once. Critical domains such as healthcare, infrastructure, or transportation are particularl
Keisuke Harigaya, Keisuke Inomata, Takahiro Terada
Rotations of axion fields in the early universe can produce dark matter and the matter-antimatter asymmetry of the universe. We point out that the rotation can generate an observable amount of a stochastic gravitational-wave (GW) background. It can be doubly enhanced in a class of models in which the equation of state of the rotations rapidly changes from a
Saba Etezad-Razavi, Lucien Hardy
A pair of interferometers can be coupled by allowing one path from each to overlap such that if the particles meet in this overlap region, they annihilate. It was shown by one of us over thirty years ago that such annihilation-coupled interferometers can exhibit apparently paradoxical behaviour. More recently, Bose et al. and Marletto and Vedral have conside
Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan
Machine Translation (MT) has been widely used for cross-lingual classification, either by translating the test set into English and running inference with a monolingual model (translate-test), or translating the training set into the target languages and finetuning a multilingual model (translate-train). However, most research in the area focuses on the mult
Yixin Liu, Kejian Shi, Katherine S He, Longtian Ye
Recent studies have found that summaries generated by large language models (LLMs) are favored by human annotators over the original reference summaries in commonly used summarization datasets. Therefore, we study an LLM-as-reference learning setting for smaller text summarization models to investigate whether their performance can be substantially improved.
Rocco D'Agostino, Matteo Califano, Nicola Menadeo, Daniele Vernieri
This paper investigates the effects of nonvanishing spatial curvature on the propagation of primordial gravitational waves produced during inflation. In particular, we consider tensor perturbations over a homogeneous and isotropic background, and describe the propagation of gravitational waves in the de Sitter phase with spatially curved geometries. We thus
Wenting Zhao, Justin T. Chiu, Claire Cardie, Alexander M. Rush
Explainable multi-hop question answering (QA) not only predicts answers but also identifies rationales, i. e. subsets of input sentences used to derive the answers. This problem has been extensively studied under the supervised setting, where both answer and rationale annotations are given. Because rationale annotations are expensive to collect and not alway
Lingteng Qiu, Guanying Chen, Jiapeng Zhou, Mutian Xu
Reconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality dynamic clothed human reconstruction results from monocular video, but these methods cannot separate the garment surface f
Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata
Multilingual Large Language Models (LLMs) have recently shown great capabilities in a wide range of tasks, exhibiting state-of-the-art performance through zero-shot or few-shot prompting methods. While there have been extensive studies on their abilities in monolingual tasks, the investigation of their potential in the context of code-switching (CSW), the pr
Benny Avelin, Mingyi Hou, Kaj Nyström
In this paper, we develop a Galerkin-type approximation, with quantitative error estimates, for weak solutions to the Cauchy problem for kinetic Fokker-Planck equations in the domain $(0, T) \times D \times \mathbb{R}^d$, where $D$ is either $\mathbb{T}^d$ or $\mathbb{R}^d$. Our approach is based on a Hermite expansion in the velocity variable only, with a h
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin
Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance. This paper aims to improve the upper bound of open-source models further. We first provide a
Yu Zhang, Hao Cheng, Zhihong Shen, Xiaodong Liu
Scientific literature understanding tasks have gained significant attention due to their potential to accelerate scientific discovery. Pre-trained language models (LMs) have shown effectiveness in these tasks, especially when tuned via contrastive learning. However, jointly utilizing pre-training data across multiple heterogeneous tasks (e.g., extreme multi-
Yuchen Guo, Jian-Hao Zhang, Zhen Bi, Shuo Yang
We investigate the phase diagram at the boundary of an infinite two-dimensional cluster state subject to bulk measurements using tensor network methods. The state is subjected to uniform measurements $M = \cos{\theta}Z+\sin{\theta}X$ on the lower boundary qubits and in all bulk qubits. Our results show that the boundary of the system exhibits volume-law enta
Neha Verma, Kenton Murray, Kevin Duh
Multilingual machine translation has proven immensely useful for both parameter efficiency and overall performance across many language pairs via complete multilingual parameter sharing. However, some language pairs in multilingual models can see worse performance than in bilingual models, especially in the one-to-many translation setting. Motivated by their
Jack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf
Learning structured representations of the visual world in terms of objects promises to significantly improve the generalization abilities of current machine learning models. While recent efforts to this end have shown promising empirical progress, a theoretical account of when unsupervised object-centric representation learning is possible is still lacking.
Matthias Stiefenhofer
We give conditions for local diagonalization of analytic operator families acting between real or complex Banach spaces. The transformations are constructed from an operator Toeplitz matrix obtained from Jordan chains of increasing length. The basic assumption is given by stabilization of the Jordan chains at length k in the sense that no root elements with
Transmutations from the Covariant Transform on the Heisenberg Group and an Extended Umbral Principle
math.APVladimir V. Kisil
We discuss several seemingly assorted objects: the umbral calculus, generalised translations and associated transmutations, symbolic calculus of operators. The common framework for them is representations of the Weyl algebra of the Heisenberg group by ladder operators. Transporting various properties between different implementations we review some classic r
Maximilian Schumacher, Gernot Alber
Entanglement detection by local measurements, which can possibly be performed by far distant observers, are of particular interest for applications in quantum key distribution and quantum communication. In this paper sufficient conditions for arbitrary dimensional bipartite entanglement detection based on correlation matrices and joint probability distributi
Kung-Hsiang Huang, Hou Pong Chan, Kathleen McKeown, Heng Ji
Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant chal
Jonas Pfeiffer, Francesco Piccinno, Massimo Nicosia, Xinyi Wang
Multilingual sequence-to-sequence models perform poorly with increased language coverage and fail to consistently generate text in the correct target language in few-shot settings. To address these challenges, we propose mmT5, a modular multilingual sequence-to-sequence model. mmT5 utilizes language-specific modules during pre-training, which disentangle lan
Max Olan Smith, Michael P. Wellman
Game-based decision-making involves reasoning over both world dynamics and strategic interactions among the agents. Typically, empirical models capturing these respective aspects are learned and used separately. We investigate the potential gain from co-learning these elements: a world model for dynamics and an empirical game for strategic interactions. Empi
Xinyu Zhu, Cheng Yang, Bei Chen, Siheng Li
Question answering plays a pivotal role in human daily life because it involves our acquisition of knowledge about the world. However, due to the dynamic and ever-changing nature of real-world facts, the answer can be completely different when the time constraint in the question changes. Recently, Large Language Models (LLMs) have shown remarkable intelligen
Jiayin Dong, Daniel Foreman-Mackey
Stellar obliquity, the angle between a planet's orbital axis and its host star's spin axis, traces the formation and evolution of a planetary system. In transiting exoplanet observations, only the sky-projected stellar obliquity can be measured, but this can be de-projected using an estimate of the stellar obliquity. In this paper, we introduce a flexible, h
Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Mohammed Khan
Visual document understanding is a complex task that involves analyzing both the text and the visual elements in document images. Existing models often rely on manual feature engineering or domain-specific pipelines, which limit their generalization ability across different document types and languages. In this paper, we propose DUBLIN, which is pretrained o
Amine Marrakchi, Stefaan Vaes
Since the early days of Tomita-Takesaki theory, it is known that a von Neumann algebra $M$ that admits a state $\varphi$ with trivial centralizer $M_\varphi$ must be a type III$_1$ factor, but the converse remained open. We solve this problem and prove that such ergodic states form a dense $G_\delta$ set among all faithful normal states on any III$_1$ factor
Chengbin Xuan, Feng Zhang, Faliang Yin, Hak-Keung Lam
The problem of constrained reinforcement learning (CRL) holds significant importance as it provides a framework for addressing critical safety satisfaction concerns in the field of reinforcement learning (RL). However, with the introduction of constraint satisfaction, the current CRL methods necessitate the utilization of second-order optimization or primal-
Chang-You Tai, Ziru Chen, Tianshu Zhang, Xiang Deng
In-context learning with large language models (LLMs) has recently caught increasing attention due to its superior few-shot performance on various tasks. However, its performance on text-to-SQL parsing still has much room for improvement. In this paper, we hypothesize that a crucial aspect of LLMs to improve for text-to-SQL parsing is their multi-step reason
Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulić
While many languages possess processes of joining two or more words to create compound words, previous studies have been typically limited only to languages with excessively productive compound formation (e.g., German, Dutch) and there is no public dataset containing compound and non-compound words across a large number of languages. In this work, we systema
Griffin M. Kearney, Kasey M. Laurent, Reece V. Kearney
Lagrangian Particle Tracking (LPT) enables practitioners to study various concepts in turbulence by measuring particle positions in flows of interest. This data is subject to measurement errors, and filtering techniques are applied to mitigate these errors and improve the accuracy of analyses utilizing the data. We develop a new type of position filter throu
A. Bahri, M. Bendersky, F. R. Cohen, S. Gitler
We give a geometric method for determining the cohomology groups of a polyhedral product under suitable freeness conditions or with coefficients taken in a field. This is done by considering first the special case for which the pairs of spaces are wedge decomposable. We derive a decomposition for these polyhedral products which resembles a Cartan formula. Th
Minjun Zhu, Yixuan Weng, Shizhu He, Kang Liu
In Textual question answering (TQA) systems, complex questions often require retrieving multiple textual fact chains with multiple reasoning steps. While existing benchmarks are limited to single-chain or single-hop retrieval scenarios. In this paper, we propose to conduct Graph-Hop -- a novel multi-chains and multi-hops retrieval and reasoning paradigm in c
Shengnan An, Bo Zhou, Zeqi Lin, Qiang Fu
In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples. Few-shot selection -- selecting appropriate examples for each test instance separately -- is important for in-context learning. In this paper, we propose Skill-KNN, a skill-based few-shot selection method for in-context learning. The key adv