Skip to content

May 2023 arXiv papers — page 61

Showing 6,0016,100 of 19,695 papers

  1. Yongtao Liu, Anna N. Morozovska, Ayana Ghosh, Kyle P. Kelley

    Nanoscale ferroelectric 2D materials offer unique opportunity to investigate curvature and strain effects on materials functionalities. Among these, CuInP2S6 (CIPS) has attracted tremendous research interest in recent years due to combination of room temperature ferroelectricity, scalability to a few layers thickness, and unique ferrielectric properties due

  2. Haldun Özgür Bayındır

    Using the root adjunction formalism developed in an earlier work and logarithmic THH, we obtain a simplified computation of $T(2)_*\text{K}(ku)$ for $p>3$. Through this, we also produce a new algebraic $K$-theory computation; namely we obtain $T(2)_*\text{K}(ku/p)$, where $ku/p$ is the $2$-periodic Morava $K$-theory spectrum of height $1$.

  3. Robert Morabito, Jad Kabbara, Ali Emami

    Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods that yield not only desirable results, but are also consistent with their mechanisms and specifications. For example, we a

  4. Junyuan Ouyang, Xiao Liu, Haoyao Chen

    While point-based neural architectures have demonstrated their efficacy, the time-consuming sampler currently prevents them from performing real-time reasoning on scene-level point clouds. Existing methods attempt to overcome this issue by using random sampling strategy instead of the commonly-adopted farthest point sampling~(FPS), but at the expense of lowe

  5. Esteban González, Kimet Jusufi, Genly Leon, Emmanuel N. Saridakis

    We confront Yukawa modified cosmology, proposed in arXiv:2304.11492 [Jusufi et al. arXiv:2304.11492], with data from Supernovae Type Ia (SNe Ia) and Hubble parameter (OHD) observations. Yukawa cosmology is obtained from a Yukawa-like gravitational potential, with coupling parameter $\alpha$ and wavelength parameter $\lambda$, which gives rise to modified Fri

  6. Fang Zhang, Xing Zhu, Rui Chao, Cupjin Huang

    Scaling bottlenecks the making of digital quantum computers, posing challenges from both the quantum and the classical components. We present a classical architecture to cope with a comprehensive list of the latter challenges {\em all at once}, and implement it fully in an end-to-end system by integrating a multi-core RISC-V CPU with our in-house control ele

  7. Yilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi

    People primarily consult tables to conduct data analysis or answer specific questions. Text generation systems that can provide accurate table summaries tailored to users' information needs can facilitate more efficient access to relevant data insights. Motivated by this, we define a new query-focused table summarization task, where text generation models ha

  8. Shijie Geng, Juntao Tan, Shuchang Liu, Zuohui Fu

    Computer Vision (CV), Natural Language Processing (NLP), and Recommender Systems (RecSys) are three prominent AI applications that have traditionally developed independently, resulting in disparate modeling and engineering methodologies. This has impeded the ability for these fields to directly benefit from each other's advancements. With the recent developm

  9. Fangda Li, Zhiqiang Hu, Wen Chen, Avinash Kak

    Hematoxylin and Eosin (H&E) staining is a widely used sample preparation procedure for enhancing the saturation of tissue sections and the contrast between nuclei and cytoplasm in histology images for medical diagnostics. However, various factors, such as the differences in the reagents used, result in high variability in the colors of the stains actually re

  10. Orr Fischer, Merav Parter

    In their seminal PODC 1991 paper, Ostrovsky and Yung introduced the study of distributed computation in the presence of mobile adversaries which can dynamically appear throughout the network. Over the years, this setting has been studied mostly under the assumption that the communication graph is fully-connected. Resilient CONGEST algorithms for general grap

  11. Minsik Oh, Jiwei Li, Guoyin Wang

    Learning high quality sentence embeddings from dialogues has drawn increasing attentions as it is essential to solve a variety of dialogue-oriented tasks with low annotation cost. Annotating and gathering utterance relationships in conversations are difficult, while token-level annotations, \eg, entities, slots and templates, are much easier to obtain. Other

  12. En Yu, Tiancai Wang, Zhuoling Li, Yuang Zhang

    Although end-to-end multi-object trackers like MOTR enjoy the merits of simplicity, they suffer from the conflict between detection and association seriously, resulting in unsatisfactory convergence dynamics. While MOTRv2 partly addresses this problem, it demands an additional detection network for assistance. In this work, we serve as the first to reveal th

  13. Thomas Izgin, David I. Ketcheson, Andreas Meister

    In recent years, many positivity-preserving schemes for initial value problems have been constructed by modifying a Runge--Kutta (RK) method by weighting the right-hand side of the system of differential equations with solution-dependent factors. These include the classes of modified Patankar--Runge--Kutta (MPRK) and Geometric Conservative (GeCo) methods. Co

  14. Kundan Krishna, Prakhar Gupta, Sanjana Ramprasad, Byron C. Wallace

    While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived benchmark, complemented by a rich set of crowd-sourced annotations, that supports $8$ interrelated tasks: (i) extractive summa

  15. Francesco Giovanni Celiberto

    In this review we discuss and extend the study of the inclusive production of vector quarkonia, $J/\psi$ and $\Upsilon$, emitted with large transverse momenta and rapidities at the LHC. We adopt the novel ZCW19$^+$ determination to depict the quarkonium production mechanism at the next-to-leading level of perturbative QCD. This approach is based on the nonre

  16. Alessandro Sinibaldi, Clemens Giuliani, Giuseppe Carleo, Filippo Vicentini

    We analyze the accuracy and sample complexity of variational Monte Carlo approaches to simulate the dynamics of many-body quantum systems classically. By systematically studying the relevant stochastic estimators, we are able to: (i) prove that the most used scheme, the time-dependent Variational Monte Carlo (tVMC), is affected by a systematic statistical bi

  17. Sergey Ershkov, Dmytro Leshchenko, Evgeniy Prosviryakov

    In the current study, we explore stability of motion of satellite around the natural moons of planets in Solar system using the novel concept of ER3BP with variable eccentricity. This concept was introduced earlier when novel type of ER3BP (Sun-planet-satellite) was investigated with variable spin state of secondary planet correlated implicitly to the motion

  18. Chenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos

    Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata knowledge base. However, models trained only on Wikipedia have limitations when applied to web domains, which often contain

  19. Sina J. Semnani, Violet Z. Yao, Heidi C. Zhang, Monica S. Lam

    This paper presents the first few-shot LLM-based chatbot that almost never hallucinates and has high conversationality and low latency. WikiChat is grounded on the English Wikipedia, the largest curated free-text corpus. WikiChat generates a response from an LLM, retains only the grounded facts, and combines them with additional information it retrieves from

  20. Nicholas Deas, Jessi Grieser, Shana Kleiner, Desmond Patton

    We evaluate how well LLMs understand African American Language (AAL) in comparison to their performance on White Mainstream English (WME), the encouraged "standard" form of English taught in American classrooms. We measure LLM performance using automatic metrics and human judgments for two tasks: a counterpart generation task, where a model generates AAL (or

  21. Karol Gietka, Christoph Hotter, Helmut Ritsch

    Squeezing is essential to many quantum technologies and our understanding of quantum physics. Here we develop a theory of steady-state squeezing that can be generated in the closed and open quantum Rabi as well as Dicke model. To this end, we eliminate the spin dynamics which effectively leads to an abstract harmonic oscillator whose eigenstates are squeezed

  22. Xili Yi, Nima Fazeli

    In this paper, we discuss the mechanics and planning algorithms to slide an object on a horizontal planar surface via frictional patch contact made with its top surface. Here, we propose an asymmetric dual limit surface model to determine slip boundary conditions for both the top and bottom contact. With this model, we obtain a range of twists that can keep

  23. Chenxi Whitehouse, Monojit Choudhury, Alham Fikri Aji

    This paper explores the potential of leveraging Large Language Models (LLMs) for data augmentation in multilingual commonsense reasoning datasets where the available training data is extremely limited. To achieve this, we utilise several LLMs, namely Dolly-v2, StableVicuna, ChatGPT, and GPT-4, to augment three datasets: XCOPA, XWinograd, and XStoryCloze. Sub

  24. Max Weinreich

    We introduce an algebraic formulation of billiards on plane curves over algebraically closed fields, extending Glutsyuk's complex billiards. For any smooth algebraic curve $C$ of degree $d \geq 2$, algebraic billiards is a rational $(d-1)$-to-$(d-1)$ surface correspondence on the space of unit tangent vectors based on $C$. We prove that the dynamical degree

  25. Koen Minartz, Yoeri Poels, Simon Koop, Vlado Menkovski

    Neural networks are emerging as a tool for scalable data-driven simulation of high-dimensional dynamical systems, especially in settings where numerical methods are infeasible or computationally expensive. Notably, it has been shown that incorporating domain symmetries in deterministic neural simulators can substantially improve their accuracy, sample effici

  26. Matteo Piccolini, Vittorio Giovannetti, Rosario Lo Franco

    We propose a procedure for the robust preparation of maximally entangled states of identical fermionic qubits, studying the role played by particle statistics in the process. The protocol exploits externally activated noisy channels to reset the system to a known state. The subsequent interference effects generated at a beam splitter result in a mixture of m

  27. Nadia Benlakhouy, Ahmed Jellal, Michael Schreiber

    We investigate the electronic properties of a hybrid system that comprises single-bilayer graphene structures subjected to a perpendicular magnetic field. Specifically, our focus is on the behavior exhibited by the zigzag boundaries of the junction, namely Zigzag-1 (ZZ1) and Zigzag-2 (ZZ2), using the continuum Dirac model for rigorous analysis. Our findings

  28. Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao

    Large Language Models (LLMs) play powerful, black-box readers in the retrieve-then-read pipeline, making remarkable progress in knowledge-intensive tasks. This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting. Unlike prior studies foc

  29. Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song

    Automatically evaluating the quality of language generation is critical. Although recent learned metrics show high correlation with human judgement, these metrics can not explain their verdict or associate the scores with defects in generated text. To address this limitation, we present InstructScore, an explainable evaluation metric for text generation. By

  30. Emanuele Bugliarello, Aida Nematzadeh, Lisa Anne Hendricks

    Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations. In this work, we take a step further and explore how we can tap into supervision from small-scale visual relation data. In particular, we propose two pretraining approaches to contextualise vi

  31. Elizabeth Salesky, Neha Verma, Philipp Koehn, Matt Post

    We introduce and demonstrate how to effectively train multilingual machine translation models with pixel representations. We experiment with two different data settings with a variety of language and script coverage, demonstrating improved performance compared to subword embeddings. We explore various properties of pixel representations such as parameter sha

  32. Angelica Chen, Jason Phang, Alicia Parrish, Vishakh Padmakumar

    Large language models (LLMs) have achieved widespread success on a variety of in-context few-shot tasks, but this success is typically evaluated via correctness rather than consistency. We argue that self-consistency is an important criteria for valid multi-step reasoning in tasks where the solution is composed of the answers to multiple sub-steps. We propos

  33. Rico Huhnstock, Lukas Paetzold, Maximilian Merkel, Piotr Kuświk

    Controlled actuation of superparamagnetic beads (SPBs) within a microfluidic environment using tailored dynamic magnetic field landscapes (MFLs) is a potent approach for the realization of point-of-care diagnostics within Lab-on-a-chip (LOC) systems. Making use of an engineered magnetic domain pattern as the MFL source, a functional LOC-element with combined

  34. Ammar Ahmed, Anwar Said, Mudassir Shabbir, Xenofon Koutsoukos

    Vulnerability identification constitutes a task of high importance for cyber security. It is quite helpful for locating and fixing vulnerable functions in large applications. However, this task is rather challenging owing to the absence of reliable and adequately managed datasets and learning models. Existing solutions typically rely on human expertise to an

  35. Yaroslav Zhumagulov, Denis Kochan, Jaroslav Fabian

    The impact of proximity-induced spin-orbit and exchange coupling on the correlated phase diagram of rhombohedral trilayer graphene (RTG) is investigated theoretically. By employing \emph{ab initio}-fitted effective models of RTG encapsulated by transition metal dichalcogenides (spin-orbit proximity effect) and ferromagnetic Cr$_2$Ge$_2$Te$_6$ (exchange proxi

  36. Ada Chan, Peter Sin

    In a continuous-time quantum walk on a network of qubits, pretty good state transfer is the phenomenon of state transfer between two vertices with fidelity arbitrarily close to 1. We construct families of graphs to demonstrate that there is no bound on the size of a set of vertices that admit pretty good state transfer between any two vertices of the set.

  37. Yash Patel, Declan McNamara, Jackson Loper, Jeffrey Regier

    Amortized variational inference is an often employed framework in simulation-based inference that produces a posterior approximation that can be rapidly computed given any new observation. Unfortunately, there are few guarantees about the quality of these approximate posteriors. We propose Conformalized Amortized Neural Variational Inference (CANVI), a proce

  38. B. V. Rajarama Bhat, Purbayan Chakraborty, Uwe Franz

    The Weyl operators give a convenient basis of $M_n(\mathbb{C})$ which is also orthonormal with respect to the Hilbert-Schmidt inner product. The properties of such a basis can be generalised to the notion of a nice error basis(NEB), as introduced by E. Knill. We can use an NEB of $M_n(\mathbb{C})$ to construct an NEB for $Lin(M_n(\mathbb{C}))$, the space of

  39. Hyun Jeong, Kohei Kamada, Alexei A. Starobinsky, Jun'ichi Yokoyama

    Post-inflationary evolution and (re)heating of the viable inflationary model, the $R^2$ one, is made more realistic by including the leptogenesis scenario into it. For this purpose, right-handed Majorana neutrinos with a large mass are added to the matter sector of the Standard Model to explain the neutrino oscillation experiments and the baryon asymmetry of

  40. Kyle DeBry, Jasmine Sinanan-Singh, Colin D. Bruzewicz, David Reens

    We present experimental demonstrations of accurate and unambiguous single-shot discrimination between three quantum channels using a single trapped $^{40}\text{Ca}^{+}$ ion. The three channels cannot be distinguished unambiguously using repeated single channel queries, the natural classical analogue. We develop techniques for using the 6-dimensional $\text{D

  41. Zhuoran Fang, Bassem Tossoun, Antoine Descos, Di Liang

    Silicon photonics has evolved from lab research to commercial products in the past decade as it plays an increasingly crucial role in data communication for next-generation data centers and high performance computing1. Recently, programmable silicon photonics has also found new applications in quantum2 and classical 3 information processing. A key component

  42. Haonan Lu, Shuai Mu, Siddhartha Sen, Wyatt Lloyd

    Strictly serializable datastores greatly simplify the development of correct applications by providing strong consistency guarantees. However, existing techniques pay unnecessary costs for naturally consistent transactions, which arrive at servers in an order that is already strictly serializable. We find these transactions are prevalent in datacenter worklo

  43. Giulia Rizzoli, Donald Shenaj, Pietro Zanuttigh

    With the increasing availability of depth sensors, multimodal frameworks that combine color information with depth data are gaining interest. However, ground truth data for semantic segmentation is burdensome to provide, thus making domain adaptation a significant research area. Yet most domain adaptation methods are not able to effectively handle multimodal

  44. Zi-Yi Dou, Feng Gao, Nanyun Peng

    Vision-and-language navigation (VLN) agents are trained to navigate in real-world environments by following natural language instructions. A major challenge in VLN is the limited availability of training data, which hinders the models' ability to generalize effectively. Previous approaches have attempted to address this issue by introducing additional superv

  45. Martin Gonzalez, Nelson Fernandez, Thuy Tran, Elies Gherbi

    A potent class of generative models known as Diffusion Probabilistic Models (DPMs) has become prominent. A forward diffusion process adds gradually noise to data, while a model learns to gradually denoise. Sampling from pre-trained DPMs is obtained by solving differential equations (DE) defined by the learnt model, a process which has shown to be prohibitive

  46. Dmitry Svintsov, Georgy Alymov

    Despite numerous applications of two-dimensional plasmons for electromagnetic energy manipulation at the nanoscale, their quantitative refraction and reflection laws (analogs of Fresnel formulas in optics) have not yet been established. This fact can be traced down to the strong non-locality of equations governing the 2d plasmon propagation. Here, we tackle

  47. Timothy B. Armstrong, Patrick Kline, Liyang Sun

    Empirical research typically involves a robustness-efficiency tradeoff. A researcher seeking to estimate a scalar parameter can invoke strong assumptions to motivate a restricted estimator that is precise but may be heavily biased, or they can relax some of these assumptions to motivate a more robust, but variable, unrestricted estimator. When a bound on the

  48. Katerina Margatina, Timo Schick, Nikolaos Aletras, Jane Dwivedi-Yu

    The remarkable advancements in large language models (LLMs) have significantly enhanced the performance in few-shot learning settings. By using only a small number of labeled examples, referred to as demonstrations, LLMs can effectively grasp the task at hand through in-context learning. However, the process of selecting appropriate demonstrations has receiv

  49. Milind Agarwal, Md Mahfuz Ibn Alam, Antonios Anastasopoulos

    Knowing the language of an input text/audio is a necessary first step for using almost every NLP tool such as taggers, parsers, or translation systems. Language identification is a well-studied problem, sometimes even considered solved; in reality, due to lack of data and computational challenges, current systems cannot accurately identify most of the world'

  50. V. Ramachandran, J. Klencki, A. A. C. Sander, D. Pauli

    Standard binary evolutionary models predict a significant population of core helium-burning stars that lost their hydrogen-rich envelope after mass transfer via Roche-lobe overflow. However, there is a scarcity of observations of such stripped stars in the intermediate mass regime (~1.5 - 8$ M_{\odot}$), which are thought to be prominent progenitors of SN Ib

  51. V. M. Jiménez, M. De León

    In this paper, we study internal properties of a Cosserat media. In fact, by using groupoids and smooth distributions, we obtain a three canonical equations. The \textit{non-holonomic material equation for Cosserat media} characterizes the uniformity of the material. The \textit{holonomic material equation for Cosserat media} permits us to study when a Cosse

  52. Yue Fan, Jing Gu, Kaizhi Zheng, Xin Eric Wang

    Intelligent navigation-helper agents are critical as they can navigate users in unknown areas through environmental awareness and conversational ability, serving as potential accessibility tools for individuals with disabilities. In this work, we first introduce a novel benchmark, Respond to Help Requests (R2H), to promote the development of multi-modal navi

  53. Qingyun Wang, Doug Downey, Heng Ji, Tom Hope

    We explore and enhance the ability of neural language models to generate novel scientific directions grounded in literature. Work on literature-based hypothesis generation has traditionally focused on binary link prediction--severely limiting the expressivity of hypotheses. This line of work also does not focus on optimizing novelty. We take a dramatic depar

  54. Zheng Xie, Yu Liu, Hao-Yuan He, Ming Li

    Since acquiring perfect supervision is usually difficult, real-world machine learning tasks often confront inaccurate, incomplete, or inexact supervision, collectively referred to as weak supervision. In this work, we present WSAUC, a unified framework for weakly supervised AUC optimization problems, which covers noisy label learning, positive-unlabeled lear

  55. Abishek Sridhar, Robert Lo, Frank F. Xu, Hao Zhu

    Large language models (LLMs) struggle on processing complicated observations in interactive decision making tasks. To alleviate this issue, we propose a simple hierarchical prompting approach. Diverging from previous prompting approaches that always put the full observation (e.g. a web page) to the prompt, we propose to first construct an action-aware observ

  56. Oleg Vasilyev, Fumika Isono, John Bohannon

    Semantics of a sentence is defined with much less ambiguity than semantics of a single word, and we assume that it should be better preserved by translation to another language. If multilingual sentence embeddings intend to represent sentence semantics, then the similarity between embeddings of any two sentences must be invariant with respect to translation.

  57. Tanchumin Xu, Yunshu Zhang, Shu Yang

    Propensity score matching (PSM) and augmented inverse propensity weighting (AIPW) are widely used in observational studies to estimate causal effects. The two approaches present complementary features. The AIPW estimator is doubly robust and locally efficient but can be unstable when the propensity scores are close to zero or one due to weighting by the inve

  58. Yiyun Fan, John Billingham, Kristoffer van der Zee

    We develop a shape-Newton method for solving generic free-boundary problems where one of the free-boundary conditions is governed by the Bernoulli equation. The Newton-like scheme is developed by employing shape derivatives in the weak forms, which allows us to update the position of the free surface and the potential on the free boundary by solving a bounda

  59. Christopher Hughes, Greg Martin, Andrew Pearce-Crump

    Shanks conjectured that $\zeta ' (\rho)$, where $\rho$ ranges over non-trivial zeros of the Riemann zeta function, is real and positive in the mean. We present a history of this problem, including a generalisation to all higher-order derivatives $\zeta^{(n)}(s)$, for which the sign of the mean alternatives between positive for odd $n$ and negative for even $

  60. Mariano Lemus, Ricardo Faleiro, Paulo Mateus, Nikola Paunković

    This work presents a study of Kolmogorov complexity for general quantum states from the perspective of deterministic-control quantum Turing Machines (dcq-TM). We extend the dcq-TM model to incorporate mixed state inputs and outputs, and define dcq-computable states as those that can be approximated by a dcq-TM. Moreover, we introduce (conditional) Kolmogorov

  61. Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis

    Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate, and (2) human evaluation is time-consuming and costly. In this paper, we introduce FACTSCORE, a new evaluation th

  62. Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson

    While large language models (LLMs) are proficient at question-answering (QA), it is not always clear how (or even if) an answer follows from their latent "beliefs". This lack of interpretability is a growing impediment to widespread use of LLMs. To address this, our goals are to make model beliefs and their inferential relationships explicit, and to resolve

  63. Blake Jackson

    This paper proves the existence of an intersection-dimension formula for preprojective modules over path algebras of type $\widetilde{D}_n$. Identical intersection-dimension formulas have previously been provided for modules over path algebras of type $A_n, D_n,$ and $\widetilde{A}_n$ due to Schiffler as well as He, Zhou, and Zhu. These modules can be repres

  64. Thomas Bonis

    We provide new bounds for the rate of convergence of the multivariate Central Limit Theorem in Wasserstein distances of order $p \geq 2$. In particular, we obtain what we conjecture to be the asymptotically optimal rate whenever the density of the summands admits a non-zero continuous component and has a non-zero third moment.

  65. David Peter Kovacs, Ilyes Batatia, Eszter Sara Arany, Gabor Csanyi

    The MACE architecture represents the state of the art in the field of machine learning force fields for a variety of in-domain, extrapolation and low-data regime tasks. In this paper, we further evaluate MACE by fitting models for published benchmark datasets. We show that MACE generally outperforms alternatives for a wide range of systems from amorphous car

  66. Jocelyn Shen, Maarten Sap, Pedro Colon-Hernandez, Hae Won Park

    The most meaningful connections between people are often fostered through expression of shared vulnerability and emotional experiences in personal narratives. We introduce a new task of identifying similarity in personal stories based on empathic resonance, i.e., the extent to which two people empathize with each others' experiences, as opposed to raw semant

  67. Gregorio García-Valladares, Carlos A. Plata, Antonio Prados, Alessandro Manacorda

    In many physical situations, there appears the problem of reaching a single target that is spatially distributed. Here we analyse how stochastic resetting, also spatially distributed, can be used to improve the search process when the target location is quenched, i.e. it does not evolve in time. More specifically, we consider a model with minimal but suffici

  68. Shengchao Chen, Guodong Long, Tao Shen, Jing Jiang

    On-device intelligence for weather forecasting uses local deep learning models to analyze weather patterns without centralized cloud computing, holds significance for supporting human activates. Federated Learning is a promising solution for such forecasting by enabling collaborative model training without sharing raw data. However, it faces three main chall

  69. Manuel Tran, Yashin Dicente Cid, Amal Lahiani, Fabian J. Theis

    Training multimodal foundation models is challenging due to the limited availability of multimodal datasets. While many public datasets pair images with text, few combine images with audio or text with audio. Even rarer are datasets that align all three modalities at once. Critical domains such as healthcare, infrastructure, or transportation are particularl

  70. Keisuke Harigaya, Keisuke Inomata, Takahiro Terada

    Rotations of axion fields in the early universe can produce dark matter and the matter-antimatter asymmetry of the universe. We point out that the rotation can generate an observable amount of a stochastic gravitational-wave (GW) background. It can be doubly enhanced in a class of models in which the equation of state of the rotations rapidly changes from a

  71. Saba Etezad-Razavi, Lucien Hardy

    A pair of interferometers can be coupled by allowing one path from each to overlap such that if the particles meet in this overlap region, they annihilate. It was shown by one of us over thirty years ago that such annihilation-coupled interferometers can exhibit apparently paradoxical behaviour. More recently, Bose et al. and Marletto and Vedral have conside

  72. Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan

    Machine Translation (MT) has been widely used for cross-lingual classification, either by translating the test set into English and running inference with a monolingual model (translate-test), or translating the training set into the target languages and finetuning a multilingual model (translate-train). However, most research in the area focuses on the mult

  73. Yixin Liu, Kejian Shi, Katherine S He, Longtian Ye

    Recent studies have found that summaries generated by large language models (LLMs) are favored by human annotators over the original reference summaries in commonly used summarization datasets. Therefore, we study an LLM-as-reference learning setting for smaller text summarization models to investigate whether their performance can be substantially improved.

  74. Rocco D'Agostino, Matteo Califano, Nicola Menadeo, Daniele Vernieri

    This paper investigates the effects of nonvanishing spatial curvature on the propagation of primordial gravitational waves produced during inflation. In particular, we consider tensor perturbations over a homogeneous and isotropic background, and describe the propagation of gravitational waves in the de Sitter phase with spatially curved geometries. We thus

  75. Wenting Zhao, Justin T. Chiu, Claire Cardie, Alexander M. Rush

    Explainable multi-hop question answering (QA) not only predicts answers but also identifies rationales, i. e. subsets of input sentences used to derive the answers. This problem has been extensively studied under the supervised setting, where both answer and rationale annotations are given. Because rationale annotations are expensive to collect and not alway

  76. Lingteng Qiu, Guanying Chen, Jiapeng Zhou, Mutian Xu

    Reconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality dynamic clothed human reconstruction results from monocular video, but these methods cannot separate the garment surface f

  77. Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata

    Multilingual Large Language Models (LLMs) have recently shown great capabilities in a wide range of tasks, exhibiting state-of-the-art performance through zero-shot or few-shot prompting methods. While there have been extensive studies on their abilities in monolingual tasks, the investigation of their potential in the context of code-switching (CSW), the pr

  78. Benny Avelin, Mingyi Hou, Kaj Nyström

    In this paper, we develop a Galerkin-type approximation, with quantitative error estimates, for weak solutions to the Cauchy problem for kinetic Fokker-Planck equations in the domain $(0, T) \times D \times \mathbb{R}^d$, where $D$ is either $\mathbb{T}^d$ or $\mathbb{R}^d$. Our approach is based on a Hermite expansion in the velocity variable only, with a h

  79. Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin

    Fine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT. Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance. This paper aims to improve the upper bound of open-source models further. We first provide a

  80. Yu Zhang, Hao Cheng, Zhihong Shen, Xiaodong Liu

    Scientific literature understanding tasks have gained significant attention due to their potential to accelerate scientific discovery. Pre-trained language models (LMs) have shown effectiveness in these tasks, especially when tuned via contrastive learning. However, jointly utilizing pre-training data across multiple heterogeneous tasks (e.g., extreme multi-

  81. Yuchen Guo, Jian-Hao Zhang, Zhen Bi, Shuo Yang

    We investigate the phase diagram at the boundary of an infinite two-dimensional cluster state subject to bulk measurements using tensor network methods. The state is subjected to uniform measurements $M = \cos{\theta}Z+\sin{\theta}X$ on the lower boundary qubits and in all bulk qubits. Our results show that the boundary of the system exhibits volume-law enta

  82. Neha Verma, Kenton Murray, Kevin Duh

    Multilingual machine translation has proven immensely useful for both parameter efficiency and overall performance across many language pairs via complete multilingual parameter sharing. However, some language pairs in multilingual models can see worse performance than in bilingual models, especially in the one-to-many translation setting. Motivated by their

  83. Jack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf

    Learning structured representations of the visual world in terms of objects promises to significantly improve the generalization abilities of current machine learning models. While recent efforts to this end have shown promising empirical progress, a theoretical account of when unsupervised object-centric representation learning is possible is still lacking.

  84. Matthias Stiefenhofer

    We give conditions for local diagonalization of analytic operator families acting between real or complex Banach spaces. The transformations are constructed from an operator Toeplitz matrix obtained from Jordan chains of increasing length. The basic assumption is given by stabilization of the Jordan chains at length k in the sense that no root elements with

  85. Vladimir V. Kisil

    We discuss several seemingly assorted objects: the umbral calculus, generalised translations and associated transmutations, symbolic calculus of operators. The common framework for them is representations of the Weyl algebra of the Heisenberg group by ladder operators. Transporting various properties between different implementations we review some classic r

  86. Maximilian Schumacher, Gernot Alber

    Entanglement detection by local measurements, which can possibly be performed by far distant observers, are of particular interest for applications in quantum key distribution and quantum communication. In this paper sufficient conditions for arbitrary dimensional bipartite entanglement detection based on correlation matrices and joint probability distributi

  87. Kung-Hsiang Huang, Hou Pong Chan, Kathleen McKeown, Heng Ji

    Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant chal

  88. Jonas Pfeiffer, Francesco Piccinno, Massimo Nicosia, Xinyi Wang

    Multilingual sequence-to-sequence models perform poorly with increased language coverage and fail to consistently generate text in the correct target language in few-shot settings. To address these challenges, we propose mmT5, a modular multilingual sequence-to-sequence model. mmT5 utilizes language-specific modules during pre-training, which disentangle lan

  89. Max Olan Smith, Michael P. Wellman

    Game-based decision-making involves reasoning over both world dynamics and strategic interactions among the agents. Typically, empirical models capturing these respective aspects are learned and used separately. We investigate the potential gain from co-learning these elements: a world model for dynamics and an empirical game for strategic interactions. Empi

  90. Xinyu Zhu, Cheng Yang, Bei Chen, Siheng Li

    Question answering plays a pivotal role in human daily life because it involves our acquisition of knowledge about the world. However, due to the dynamic and ever-changing nature of real-world facts, the answer can be completely different when the time constraint in the question changes. Recently, Large Language Models (LLMs) have shown remarkable intelligen

  91. Jiayin Dong, Daniel Foreman-Mackey

    Stellar obliquity, the angle between a planet's orbital axis and its host star's spin axis, traces the formation and evolution of a planetary system. In transiting exoplanet observations, only the sky-projected stellar obliquity can be measured, but this can be de-projected using an estimate of the stellar obliquity. In this paper, we introduce a flexible, h

  92. Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Mohammed Khan

    Visual document understanding is a complex task that involves analyzing both the text and the visual elements in document images. Existing models often rely on manual feature engineering or domain-specific pipelines, which limit their generalization ability across different document types and languages. In this paper, we propose DUBLIN, which is pretrained o

  93. Amine Marrakchi, Stefaan Vaes

    Since the early days of Tomita-Takesaki theory, it is known that a von Neumann algebra $M$ that admits a state $\varphi$ with trivial centralizer $M_\varphi$ must be a type III$_1$ factor, but the converse remained open. We solve this problem and prove that such ergodic states form a dense $G_\delta$ set among all faithful normal states on any III$_1$ factor

  94. Chengbin Xuan, Feng Zhang, Faliang Yin, Hak-Keung Lam

    The problem of constrained reinforcement learning (CRL) holds significant importance as it provides a framework for addressing critical safety satisfaction concerns in the field of reinforcement learning (RL). However, with the introduction of constraint satisfaction, the current CRL methods necessitate the utilization of second-order optimization or primal-

  95. Chang-You Tai, Ziru Chen, Tianshu Zhang, Xiang Deng

    In-context learning with large language models (LLMs) has recently caught increasing attention due to its superior few-shot performance on various tasks. However, its performance on text-to-SQL parsing still has much room for improvement. In this paper, we hypothesize that a crucial aspect of LLMs to improve for text-to-SQL parsing is their multi-step reason

  96. Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulić

    While many languages possess processes of joining two or more words to create compound words, previous studies have been typically limited only to languages with excessively productive compound formation (e.g., German, Dutch) and there is no public dataset containing compound and non-compound words across a large number of languages. In this work, we systema

  97. Griffin M. Kearney, Kasey M. Laurent, Reece V. Kearney

    Lagrangian Particle Tracking (LPT) enables practitioners to study various concepts in turbulence by measuring particle positions in flows of interest. This data is subject to measurement errors, and filtering techniques are applied to mitigate these errors and improve the accuracy of analyses utilizing the data. We develop a new type of position filter throu

  98. A. Bahri, M. Bendersky, F. R. Cohen, S. Gitler

    We give a geometric method for determining the cohomology groups of a polyhedral product under suitable freeness conditions or with coefficients taken in a field. This is done by considering first the special case for which the pairs of spaces are wedge decomposable. We derive a decomposition for these polyhedral products which resembles a Cartan formula. Th

  99. Minjun Zhu, Yixuan Weng, Shizhu He, Kang Liu

    In Textual question answering (TQA) systems, complex questions often require retrieving multiple textual fact chains with multiple reasoning steps. While existing benchmarks are limited to single-chain or single-hop retrieval scenarios. In this paper, we propose to conduct Graph-Hop -- a novel multi-chains and multi-hops retrieval and reasoning paradigm in c

  100. Shengnan An, Bo Zhou, Zeqi Lin, Qiang Fu

    In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples. Few-shot selection -- selecting appropriate examples for each test instance separately -- is important for in-context learning. In this paper, we propose Skill-KNN, a skill-based few-shot selection method for in-context learning. The key adv