Skip to content

May 2024 arXiv papers — page 65

Showing 6,4016,500 of 20,894 papers

  1. Qi Zhang, Narendra N. Hegade, Alejandro Gomez Cadavid, Lucas Lassablière

    We propose analog counterdiabatic quantum computing (ACQC) to tackle combinatorial optimization problems on neutral-atom quantum processors. While these devices allow for the use of hundreds of qubits, adiabatic quantum computing struggles with non-adiabatic errors, which are inevitable due to the hardware's restricted coherence time. We design counterdiabat

  2. Katherine Xu, Lingzhi Zhang, Jianbo Shi

    Recent advances in text-to-image (T2I) diffusion models have facilitated creative and photorealistic image synthesis. By varying the random seeds, we can generate many images for a fixed text prompt. Technically, the seed controls the initial noise and, in multi-step diffusion inference, the noise used for reparameterization at intermediate timesteps in the

  3. Tianshu Wen, Matthew J. Zahr

    We present an augmented Lagrangian trust-region method to efficiently solve constrained optimization problems governed by large-scale nonlinear systems with application to partial differential equation-constrained optimization. At each major augmented Lagrangian iteration, the expensive optimization subproblem involving the full nonlinear system is replaced

  4. Victor B. Valera, Damiano F. G. Fiorillo, Ivan Esteban, Mauricio Bustamante

    Since neutrinos have mass differences, they could decay into one another. But their lifetimes are likely long, even when shortened by new physics, so decay likely impacts neutrinos only during long trips. This makes high-energy astrophysical neutrinos, traveling for up to billions of light-years, sensitive probes of decay. However, their sensitivity must be

  5. Gergő Roósz

    In the present paper, a global Lindbladian ansatz is constructed which leads to thermalization at temperature $T$ to the Gibs state of the investigated system. This ansatz connects every two eigenstates of the Hamiltonian and leads to a simple master equation known in the literature as the relaxation time approximation (RTA). The main message of this paper i

  6. Shiyao Xu, Caiyun Liu, Yuantao Chen, Zhenxin Zhu

    Camera relocalization is a crucial problem in computer vision and robotics. Recent advancements in neural radiance fields (NeRFs) have shown promise in synthesizing photo-realistic images. Several works have utilized NeRFs for refining camera poses, but they do not account for lighting changes that can affect scene appearance and shadow regions, causing a de

  7. Tongtian Ren, Michael A. Garrett, Andrew P. V. Siemion

    Project Hephaistos recently identified seven M-dwarfs as possible Dyson Spheres (DS) candidates. We have cross-matched three of these candidates (A, B \& G) with radio sources detected in various all-sky surveys. The radio sources are offset from the Gaia stellar positions by $\sim 4.9$, $\sim 0.4$ and $\sim 5.0$ arcseconds for candidates A, B, and G respect

  8. Emmanuel Junior Wafo Wembe, Adnane Saoud

    This paper delves into the problem of computing robust controlled invariants for monotone continuous-time systems, with a specific focus on lower-closed specifications. We consider the classes of state monotone (SM) and control-state monotone (CSM) systems, we provide the structural properties of robust controlled invariants for these classes of systems and

  9. Kota Hashimoto, Tomonori Tanaka, Yoshihiro Gohda

    We propose a method to evaluate the Gibbs free energy from constant-volume first-principles phonon calculations. The volume integral of the pressure is performed by determining the volume and the bulk modulus in equilibrium at finite temperatures, where the pressure and its volume derivative are evaluated utilizing first-principles calculations of the Gr\"{u

  10. Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Yuhta Takida

    The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on downsampled data, distilling the pretrained diffusion, and p

  11. Aleksa Deric, Kyle Mitard, Shahin Tajik, Daniel Holcomb

    Driven by a need for ever increasing chip performance and inclusion of innovative features, a growing number of semiconductor companies are opting for all-inclusive System-on-Chip (SoC) architectures. Although Moore's Law has been able to keep up with the demand for more complex logic, manufacturing large dies still poses a challenge. Increasingly the soluti

  12. Merel L. R. van 't Hoff, Edwin A. Bergin, Penelope Riley, Sanil Mittal

    The low carbon content of Earth and primitive meteorites compared to the Sun and interstellar grains suggests that carbon-rich grains were destroyed in the inner few astronomical units of the young solar system. A promising mechanism to selectively destroy carbonaceous grains is thermal sublimation within the soot line at $\gtrsim$ 300 K. To address whether

  13. Davide Addona, Davide Augusto Bignamini

    Let $U,H$ be two separable Hilbert spaces and $T>0$. We consider an SDE which evolves in the Hilbert space $H$ of the form \begin{align} dX(t)=AX(t)dt+\widetilde{\mathscr L}B(X(t))dt+GdW(t), \quad t\in[0,T], \quad X(0)=x \in H, \end{align} where $A:D(A)\subseteq H\to H$ is the infinitesimal generator of a strongly continuous semigroup $(e^{tA})_{t\geq0}$, $W

  14. Beibin Li, Yi Zhang, Sébastien Bubeck, Jeevan Pathuri

    We study the efficacy of Small Language Models (SLMs) in facilitating application usage through natural language interactions. Our focus here is on a particular internal application used in Microsoft for cloud supply chain fulfilment. Our experiments show that small models can outperform much larger ones in terms of both accuracy and running time, even when

  15. Alexandra G. Hanselman, Aditya Vijaykumar, Maya Fishbach, Daniel E. Holz

    The detection of GW170817 and the measurement of its redshift from the associated electromagnetic counterpart provided the first gravitational wave determination of the Hubble constant ($H_0$), demonstrating the potential power of standard-siren cosmology. In contrast to this bright siren approach, the dark siren approach can be utilized for gravitational-wa

  16. Alexander R. Muñoz, W. Adam Phelan, Matthew S. Cook, Greta L. Chappell

    Plutonium's phase diagram is host to complex structures and interactions that make the description of its ground state properties elusive. Using all-electron density functional theory, we study the thermodynamic properties of $\alpha$-Pu. To do this, we build on recent work in the literature by introducing a novel noncollinear magnetic ansatz for $\alpha$-Pu

  17. Raymond Wang, Nicholas R. Record, D. Whitney King, Tahiya Chowdhury

    Marine debris poses a significant ecological threat to birds, fish, and other animal life. Traditional methods for assessing debris accumulation involve labor-intensive and costly manual surveys. This study introduces a framework that utilizes aerial imagery captured by drones to conduct remote trash surveys. Leveraging computer vision techniques, our approa

  18. Zakariya El-Machachi, Damyan Frantzov, A. Nijamudheen, Tigany Zarrouk

    Graphene oxide (GO) materials are widely studied, and yet their atomic-scale structures remain to be fully understood. Here we show that the chemical and configurational space of GO can be rapidly explored by advanced machine-learning methods, combining on-the-fly acceleration for first-principles molecular dynamics with message-passing neural-network potent

  19. Gianluca Amato, Mary DeMarco, James Lipton

    This paper introduces a model theory for resolution on Higher Order Hereditarily Harrop formulae (HOHH), the logic underlying the Lambda-Prolog programming language, and proves soundness and completeness of resolution. The semantics and the proof of completeness of the formal system is shown in several ways, suitably adapted to deal with the impredicativity

  20. Tim Large, Yang Liu, Minyoung Huh, Hyojin Bahng

    To improve performance in contemporary deep learning, one is interested in scaling up the neural network in terms of both the number and the size of the layers. When ramping up the width of a single layer, graceful scaling of training has been linked to the need to normalize the weights and their updates in the "natural norm" particular to that layer. In thi

  21. Pedro Fittipaldi de Castro, Wladimir Alejandro Benalcazar

    The nonlinear Schrodinger equation supports solitons -- self-interacting, localized states that behave as nearly independent objects. We exhibit solitons with self-induced nonreciprocal dynamics in a discrete nonlinear Schrodinger equation. This nonreciprocal behavior, dependent on soliton power, arises from the interplay between linear and nonlinear terms i

  22. Shomik Jain, D Calacci, Ashia Wilson

    We investigate the phenomenon of norm inconsistency: where LLMs apply different norms in similar situations. Specifically, we focus on the high-risk application of deciding whether to call the police in Amazon Ring home surveillance videos. We evaluate the decisions of three state-of-the-art LLMs -- GPT-4, Gemini 1.0, and Claude 3 Sonnet -- in relation to th

  23. Minxuan Wang, Xiaoyu Wang, Oskar Vafek

    We study electron-electron interaction induced states of twisted bilayer MoTe$_2$ in an out-of-plane magnetic field $B\hat{\bf z}$ near one hole per moir\'e unit cell filling. The 3D phase diagram showing the evolution of competing phases with $B$, interaction strength and an out-of-plane electric field is presented at electron fillings that follow the Dioph

  24. Simon Brandt, Mihai Alexandru Petrovici, Walter Senn, Katharina Anna Wilmes

    Brains can process sensory information from different modalities at astonishing speed; this is surprising as the integration of inputs through the membrane of each individual neuron already causes a delayed response. Neuronal recordings {\em in vitro} reveal a possible explanation for this fast processing, in terms of individual neurons advancing their outpu

  25. K. Green, E. Elmer, D. T. Maltby, O. Almaini

    In this work, we use 8 years of deep near-infrared imaging to select and study a new set of 601 active galaxies identified through long-term near-infrared (NIR) variability in the UKIDSS Ultra Deep Survey (UDS). These objects are compared to 710 X-ray bright AGN detected by the Chandra X-ray observatory. We show that infrared variability and X-ray emission s

  26. Zhijing Jin, Nils Heil, Jiarui Liu, Shehzaad Dhuliawala

    Implicit Personalization (IP) is a phenomenon of language models inferring a user's background from the implicit cues in the input prompts and tailoring the response based on this inference. While previous work has touched upon various instances of this problem, there lacks a unified framework to study this behavior. This work systematically studies IP throu

  27. Achyut Dhar, Valery I. Levitas, K. K. Pandey, Changyong Park

    Plastic strain-induced phase transformations (PTs) and chemical reactions under high pressure are broadly spread in modern technologies, friction and wear, geophysics, and astrogeology. However, because of very heterogeneous fields of plastic strain $\mathbf{E}^{p}$ and stress $\mathbf{\sigma}$ tensors and volume fraction $c$ of phases in a sample compressed

  28. Qian-Wei Wang, Yuqiu Xie, Letian Zhang, Zimo Liu

    Pre-trained vision-language models learn massive data to model unified representations of images and natural languages, which can be widely applied to downstream machine learning tasks. In addition to zero-shot inference, in order to better adapt pre-trained models to the requirements of downstream tasks, people usually use methods such as few-shot or parame

  29. Jonas Spinner, Victor Bresó, Pim de Haan, Tilman Plehn

    Extracting scientific understanding from particle-physics experiments requires solving diverse learning problems with high precision and good data efficiency. We propose the Lorentz Geometric Algebra Transformer (L-GATr), a new multi-purpose architecture for high-energy physics. L-GATr represents high-energy data in a geometric algebra over four-dimensional

  30. Linda V. Alegria, William W. Menasco

    From classical knot theory we know that every knot in $S^3$ is the boundary of an oriented, embedded surface. A standard demonstration of this fact achieved by elementary technique comes from taking a regular projection of any knot and employing Seifert's constructive algorithm. In this note we give a natural generalization of Seifert's algorithm to any clos

  31. Yao Lai, Sungyoung Lee, Guojin Chen, Souradip Poddar

    Analog circuit design is a significant task in modern chip technology, focusing on the selection of component types, connectivity, and parameters to ensure proper circuit functionality. Despite advances made by Large Language Models (LLMs) in digital circuit design, the complexity and scarcity of data in analog circuitry pose significant challenges. To mitig

  32. Xin Xu, Tong Xiao, Zitong Chao, Zhenya Huang

    Math Word Problems (MWPs) play a vital role in assessing the capabilities of Large Language Models (LLMs), yet current research primarily focuses on questions with concise contexts. The impact of longer contexts on mathematical reasoning remains under-explored. This study pioneers the investigation of Context Length Generalizability (CoLeG), which refers to

  33. Andreas Schlatter

    We explicitly calculate the value of the cosmological constant, based on the recently developed theory connecting entropic gravity with quantum-events, induced by transactions, called transactional gravity. We suggest a novel interpretation of the cosmological constant and rigorously show its inverse proportionality to the squared radius of the causal univer

  34. Hongxu Jiang, Muhammad Imran, Teng Zhang, Yuyin Zhou

    Denoising diffusion probabilistic models (DDPMs) have achieved unprecedented success in computer vision. However, they remain underutilized in medical imaging, a field crucial for disease diagnosis and treatment planning. This is primarily due to the high computational cost associated with (1) the use of large number of time steps (e.g., 1,000) in diffusion

  35. Jihang Zhu, Tessa Cookmeyer, Sankar Das Sarma

    We investigate the instability of layer pseudospin paramagnetic (PSP) state to the formation of pseudospin density wave (PSDW) in two-dimensional (2D) electron bilayers, analogous to the formation of Overhauser spin density wave (SDW) in a single-layer 2D electron gas (2DEG) with spin 1/2. Our comprehensive study on phase diagrams, based on the self-consiste

  36. Shengfang Zhai, Huanran Chen, Yinpeng Dong, Jiajun Li

    Text-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential auditing method for detecting unauthorized data usage. While some efforts have been made on diffusion models, th

  37. Gabriella V. Ambrósio, Michelly S. Andrade, Paulo R. F. Alves, Cleber N. Costa

    We investigate the description of black-hole thermodynamics in terms of a recently proposed modified version for Kaniadakis entropy. We discuss the role of that proposal within the Modified Newtonian Dynamics (MOND) theory, a generalization of Newton's second law aimed at explaining galaxy rotation curves without resorting to dark matter. We posit a conjectu

  38. Ezra Getzler

    Using a homotopy introduced by de Wilde and Lecomte and homological perturbation theory for $A_\infty$-algebras, we give an explicit proof that the universal enveloping algebra $UL$ of a differential graded Lie algebra $L$ is Koszul, via an explicit contracting homotopy from the cobar construction $\Omega CL$ of the Chevalley-Eilenberg chain coalgebra $CL$ o

  39. Xuan Xuan Xiao, Xin Zhang

    We use circle method prove an asymptotic local-global theorem on the heights of point orbits of thin subgroups of Bianchi groups in $\mathbb H^3$.

  40. Mohamed Debbagh, Yixue Liu, Zhouzhou Zheng, Xintong Jiang

    A plant growth simulation can be characterized as a reconstructed visual representation of a plant or plant system. The phenotypic characteristics and plant structures are controlled by the scene environment and other contextual attributes. Considering the temporal dependencies and compounding effects of various factors on growth trajectories, we formulate a

  41. Noga Alon, Colin Defant, Noah Kravitz

    A rainbow stacking of $r$-edge-colorings $\chi_1, \ldots, \chi_m$ of the complete graph on $n$ vertices is a way of superimposing $\chi_1, \ldots, \chi_m$ so that no edges of the same color are superimposed on each other. We determine a sharp threshold for $r$ (as a function of $m$ and $n$) governing the existence and nonexistence of rainbow stackings of ran

  42. Qiaoyi Chen, Siyu Liu, Kaihui Huang, Xingbo Wang

    Reading and repeatedly retelling a short story is a common and effective approach to learning the meanings and usages of target words. However, learners often struggle with comprehending, recalling, and retelling the story contexts of these target words. Inspired by the Cognitive Theory of Multimedia Learning, we propose a computational workflow to generate

  43. Yihan Wang, Lahav Lipson, Jia Deng

    We introduce SEA-RAFT, a more simple, efficient, and accurate RAFT for optical flow. Compared with RAFT, SEA-RAFT is trained with a new loss (mixture of Laplace). It directly regresses an initial flow for faster convergence in iterative refinements and introduces rigid-motion pre-training to improve generalization. SEA-RAFT achieves state-of-the-art accuracy

  44. Richard J. Szabo, Michelangelo Tirelli

    Given a homomorphism $\tau$ from a suitable finite group $\mathsf{\Gamma}$ to $\mathsf{SU}(4)$ with image $\mathsf{\Gamma}^\tau$, we construct a cohomological gauge theory on a noncommutative resolution of the quotient singularity $\mathbb{C}^4/\mathsf{\Gamma}^\tau$ whose BRST fixed points are $\mathsf{\Gamma}$-invariant tetrahedron instantons on a generally

  45. Royson Lee, Javier Fernandez-Marques, Shell Xu Hu, Da Li

    Federated learning (FL) has enabled distributed learning of a model across multiple clients in a privacy-preserving manner. One of the main challenges of FL is to accommodate clients with varying hardware capacities; clients have differing compute and memory requirements. To tackle this challenge, recent state-of-the-art approaches leverage the use of early

  46. Jinxin Liu, Xinghong Guo, Zifeng Zhuang, Donglin Wang

    In this paper, we propose a novel approach called DIffusion-guided DIversity (DIDI) for offline behavioral generation. The goal of DIDI is to learn a diverse set of skills from a mixture of label-free offline data. We achieve this by leveraging diffusion probabilistic models as priors to guide the learning process and regularize the policy. By optimizing a j

  47. Seungha Um, Tracy Sweet, Samrachana Adhikari

    Researchers have focused on understanding how individual's behavior is influenced by the behaviors of their peers in observational studies of social networks. Identifying and estimating causal peer influence, however, is challenging due to confounding by homophily, where people tend to connect with those who share similar characteristics with them. Moreover,

  48. Theodoros Pissas, Pablo Márquez-Neila, Sebastian Wolf, Martin Zinkernagel

    This work explores the effectiveness of masked image modelling for learning representations of retinal OCT images. To this end, we leverage Masked Autoencoders (MAE), a simple and scalable method for self-supervised learning, to obtain a powerful and general representation for OCT images by training on 700K OCT images from 41K patients collected under real w

  49. Pernilla Ekborg-Tanner, Petter Rosander, Erik Fransson, Paul Erhart

    Crystalline alloys and related mixed systems make up a large family of materials with high tunability which have been proposed as the solution to a large number of energy related materials design problems. Due to the presence of chemical order and disorder in these systems, neither experimental efforts nor ab-initio computational methods alone are sufficient

  50. Emmanuel Lorin, Arian Novruzi

    In this paper we develop a non-diffusive neural network (NDNN) algorithm for accurately solving weak solutions to hyperbolic conservation laws. The principle is to construct these weak solutions by computing smooth local solutions in subdomains bounded by discontinuity lines (DLs), the latter defined from the Rankine-Hugoniot jump conditions. The proposed ap

  51. Alex S. Polanski, Jack Lubin, Corey beard, Jospeh M. Akana Murphy

    The Transiting Exoplanet Survey Satellite (TESS) has discovered hundreds of new worlds, with TESS planet candidates now outnumbering the total number of confirmed planets from $\textit{Kepler}$. Owing to differences in survey design, TESS continues to provide planets that are better suited for subsequent follow-up studies, including mass measurement through

  52. Ling Yang, Bohan Zeng, Jiaming Liu, Hong Li

    Diffusion models have significantly improved the performance of image editing. Existing methods realize various approaches to achieve high-quality image editing, including but not limited to text control, dragging operation, and mask-and-inpainting. Among these, instruction-based editing stands out for its convenience and effectiveness in following human ins

  53. Yiyu Xia, Zhongdong Han, Kenji Watanabe, Takashi Taniguchi

    Moir\'e materials have enabled the realization of flat electron bands and quantum phases that are driven by strong correlations associated with flat bands. Superconductivity has been observed, but solely, in graphene moir\'e materials. The absence of robust superconductivity in moir\'e materials beyond graphene, such as semiconductor moir\'e materials, has r

  54. Beyza Dabak, Major Glenn, Jingyang Liu, Alexander Buck

    Energy is a primary constraint in processor design, and much of that energy is consumed in on-chip communication. Communication can be intra-core (e.g., from a register file to an ALU) or inter-core (e.g., over the on-chip network). In this paper, we use the on-chip network (OCN) as a case study for saving on-chip communication energy. We have identified a n

  55. Nay Myat Min, Long H. Pham, Jun Sun

    Deep neural networks have achieved remarkable success across various applications; however, their vulnerability to backdoor attacks poses severe security risks -- especially in situations where only a limited set of clean samples is available for defense. In this work, we address this critical challenge by proposing ULRL (UnLearn and ReLearn for backdoor rem

  56. Kacper Kapuśniak, Peter Potaptchik, Teodora Reu, Leo Zhang

    Matching objectives underpin the success of modern generative models and rely on constructing conditional paths that transform a source distribution into a target distribution. Despite being a fundamental building block, conditional paths have been designed principally under the assumption of Euclidean geometry, resulting in straight interpolations. However,

  57. Cristian García-Romero, Miquel Esplà-Gomis, Felipe Sánchez-Martínez

    Crawling parallel texts -- texts that are mutual translations -- from the Internet is usually done following a brute-force approach: documents are massively downloaded in an unguided process, and only a fraction of them end up leading to actual parallel content. In this work we propose a smart crawling method that guides the crawl towards finding parallel co

  58. Dimitri Meunier, Zikai Shen, Mattes Mollenhauer, Arthur Gretton

    We study theoretical properties of a broad class of regularized algorithms with vector-valued output. These spectral algorithms include kernel ridge regression, kernel principal component regression, various implementations of gradient descent and many more. Our contributions are twofold. First, we rigorously confirm the so-called saturation effect for ridge

  59. Cong Li, Mengli Hu, Zhilin Li, Yang Wang

    Altermagnets constitute a novel, third fundamental class of collinear magnetic ordered materials, alongside with ferro- and antiferromagnets. They share with conventional antiferromagnets the feature of a vanishing net magnetization. At the same time they show a spin-splitting of electronic bands, just as in ferromagnets, caused by the atomic exchange intera

  60. Supriyo Ghosh, Sheng Zhang, Chen Cheng, Gia-Wei Chern

    We present a scalable machine learning (ML) force-field model for the adiabatic dynamics of cooperative Jahn-Teller (JT) systems. Large scale dynamical simulations of the JT model also shed light on the orbital ordering dynamics in colossal magnetoresistance manganites. The JT effect in these materials describes the distortion of local oxygen octahedra drive

  61. Karl Kunisch, Fredi Troeltzsch

    For a nonlinear ordinary differential equation with time delay, the differentiation of the solution with respect to the delay is investigated. Special emphasis is laid on the second-order derivative. The results are applied to an associated optimization problem for the time delay. A first- and second-order sensitivity analysis is performed including an adjoi

  62. Timon Mede, Samir El Shawish

    A simple micromechanical model of polycrystalline materials is proposed, which enables us to swiftly produce grain-boundary-stress distributions induced by the uniform external loading (in the elastic strain regime). Such statistical knowledge of local stresses is a necessary prerequisite to assess the probability for intergranular cracking initiation. Model

  63. Francesco Giovanni Celiberto, Gabriele Gatto, Alessandro papa

    We investigate the inclusive production of fully charmed tetraquarks, $T_{4c}(0^{++})$ or $T_{4c}(2^{++})$ radial excitations, in high-energy proton collisions. We build our study upon the collinear fragmentation of a single parton in a variable-flavor number scheme, suited to describe the tetraquark formation mechanism from moderate to large transverse-mome

  64. Maria Blum, Christian Döding, Patrick Henning

    This paper considers minimizers of the Ginzburg-Landau energy functional in special multiscale spaces that are based on finite elements. The spaces are constructed by localized orthogonal decomposition techniques and their usage for solving the Ginzburg-Landau equation was first suggested in [D\"orich, Henning, SINUM 2024]. In this work we further explore th

  65. Mabrouk Sghaier, Francisco Marcellán

    Let $\mathcal{T}_{\mu}$ be the Dunkl operator. A pair of symmetric measures $(u, v)$ supported on a symmetric subset of the real line is said to be a symmetric Dunkl-coherent pair if the corresponding sequences of monic orthogonal polynomials $\{P_n\}_{n\geq 0}$ and $\{R_n\}_{n\geq 0}$ (resp.) satisfy $$ R_{n}(x)=\frac{\mathcal{T}_{\mu}P_{n+1} (x)}{\mu_{n+1}

  66. Shuo Han, Yongshun Xu, Dayang Wang, Bahareh Morovati

    Cardiac computed tomography (CT) has emerged as a major imaging modality for the diagnosis and monitoring of cardiovascular diseases. High temporal resolution is essential to ensure diagnostic accuracy. Limited-angle data acquisition can reduce scan time and improve temporal resolution, but typically leads to severe image degradation and motivates for improv

  67. Andi Peng, Yuying Sun, Tianmin Shu, David Abel

    Humans use social context to specify preferences over behaviors, i.e. their reward functions. Yet, algorithms for inferring reward models from preference data do not take this social learning view into account. Inspired by pragmatic human communication, we study how to extract fine-grained data regarding why an example is preferred that is useful for learnin

  68. Peng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu

    Large language models (LLMs) need knowledge updates to meet the ever-growing world facts and correct the hallucinated responses, facilitating the methods of lifelong model editing. Where the updated knowledge resides in memories is a fundamental question for model editing. In this paper, we find that editing either long-term memory (direct model parameters)

  69. Hongyang Yang, Boyu Zhang, Neng Wang, Cheng Guo

    As financial institutions and professionals increasingly incorporate Large Language Models (LLMs) into their workflows, substantial barriers, including proprietary data and specialized knowledge, persist between the finance sector and the AI community. These challenges impede the AI community's ability to enhance financial tasks effectively. Acknowledging fi

  70. Joshua Harris, Timothy Laurence, Leo Loman, Fan Grayson

    Advances in Large Language Models (LLMs) have led to significant interest in their potential to support human experts across a range of domains, including public health. In this work we present automated evaluations of LLMs for public health tasks involving the classification and extraction of free text. We combine six externally annotated datasets with seve

  71. Yanlin Chen, András Gilyén, Ronald de Wolf

    Finding a good approximation of the top eigenvector of a given $d\times d$ matrix $A$ is a basic and important computational problem, with many applications. We give two different quantum algorithms that, given query access to the entries of a Hermitian matrix $A$ and assuming a constant eigenvalue gap, output a classical description of a good approximation

  72. Ronald Gamble, Jordan Forman, Amethyst Barnes, Gokul Srinivasaragavan

    Multi-Messenger observations and theory of astrophysical objects is fast becoming a critical research area in the astrophysics scientific community. In particular, point-like objects like that of BL Lac, flat spectrum radio quasars (FSRQ), and blazar candidates of uncertain type (BCU) are of distinct interest among those who look at the synchrotron, Compton,

  73. Francisco Guillén-González, Giordano Tierra

    In this work we present two new numerical schemes to approximate the Navier-Stokes-Cahn-Hilliard system with degenerate mobility using finite differences in time and finite elements in space. The proposed schemes are conservative, energy-stable and preserve the maximum principle approximately (the amount of the phase variable being outside of the interval [0

  74. Nicholas Gao, Stephan Günnemann

    Neural wave functions accomplished unprecedented accuracies in approximating the ground state of many-electron systems, though at a high computational cost. Recent works proposed amortizing the cost by learning generalized wave functions across different structures and compounds instead of solving each problem independently. Enforcing the permutation antisym

  75. Natasha Latouf, Emma Schwartzman, Jeffrey McKaig, Sara Doan

    Effective and ethical mentorship practices are crucial to improving recruitment and retention especially for historically minoritized groups (HMGs). Spectrum is a diversity, inclusion, equity, and accessibility (DEIA) grassroots organization committed to empowering equitable excellence through sustainable change. By improving transparency and DEIA within the

  76. Masoud Ganji, Cristina Giannotti, Gerd Schmalz, Andrea Spiro

    We classify the Ricci flat Lorentzian $n$-manifolds satisfying three particular conditions, encoding and combining some crucial features of the Kerr metrics and the Robinson-Trautman optical structures. We prove that: (a) If $n>4$, there is no Lorentzian manifold satisfying the considered Kerr type conditions, in unexpected contrast with what occurs for the

  77. Tehila Dahan, Kfir Y. Levy

    In this paper, we investigate the challenging framework of Byzantine-robust training in distributed machine learning (ML) systems, focusing on enhancing both efficiency and practicality. As distributed ML systems become integral for complex ML tasks, ensuring resilience against Byzantine failures-where workers may contribute incorrect updates due to malice o

  78. Luise Ge, Daniel Halpern, Evi Micha, Ariel D. Procaccia

    In the context of reinforcement learning from human feedback (RLHF), the reward function is generally derived from maximum likelihood estimation of a random utility model based on pairwise comparisons made by humans. The problem of learning a reward function is one of preference aggregation that, we argue, largely falls within the scope of social choice theo

  79. CMS Collaboration

    A search for violation of Lorentz invariance in the production of top quark pairs ($\mathrm{t\bar{t}}$) is presented. The measured normalized differential $\mathrm{t\bar{t}}$ production cross section, as function of the sidereal time, is examined for potential modulations induced by Lorentz-invariance breaking operators in an effective field theory extension

  80. Emilia Mezzetti, Rosa M. Miró-Roig

    In this paper, we determine the maximum $h_{max}$ and the minimum $h_{min}$ of the Hilbert vectors of Perazzo algebras $A_F$, where $F$ is a Perazzo polynomial of degree $d$ in $n+m+1$ variables. These algebras always fail the Strong Lefschetz Property. We determine the integers $n,m,d$ such that $h_{max}$ (resp. $h_{min}$) is unimodal, and we prove that $A_

  81. Sarah Alnegheimish, Linh Nguyen, Laure Berti-Equille, Kalyan Veeramachaneni

    Recent studies have shown the ability of large language models to perform a variety of tasks, including time series forecasting. The flexible nature of these models allows them to be used for many applications. In this paper, we present a novel study of large language models used for the challenging task of time series anomaly detection. This problem entails

  82. A. Herreros-Martínez, R. Magdalena-Benedicto, J. Vila-Francés, A. J. Serrano-López

    In a context of a continuous digitalisation of processes, organisations must deal with the challenge of detecting anomalies that can reveal suspicious activities upon an increasing volume of data. To pursue this goal, audit engagements are carried out regularly, and internal auditors and purchase specialists are constantly looking for new methods to automate

  83. Wei Huang, Haotong Qin, Yangdong Liu, Yawei Li

    Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs). However, while uniform-precision quantization is computationally efficient, it often compromises model performance. To address this, we propose SliM-LLM, a salience-driven mixed-precision quantization framework that allocates bit-widths at the group-wise.

  84. Aral de Moor, Arie van Deursen, Maliheh Izadi

    Transformer-based language models are highly effective for code completion, with much research dedicated to enhancing the content of these completions. Despite their effectiveness, these models come with high operational costs and can be intrusive, especially when they suggest too often and interrupt developers who are concentrating on their work. Current re

  85. Alex Gnech, Alessandro Lovato, Noemi Rocco

    We compute ground-state and dynamical properties of $^4$He and $^{16}$O nuclei using as input high-resolution, phenomenological nucleon-nucleon and three-nucleon forces that are local in coordinate space. The nuclear Schr\"odinger equation for both nuclei is accurately solved employing the auxiliary-field diffusion Monte Carlo approach. For the $^4$He nucleu

  86. Yutaro Iiyama, Wonho Jang, Naoki Kanazawa, Ryu Sawada

    The fidelity of certain gates on noisy quantum computers may be improved when they are implemented using more than two levels of the involved transmons. The main impediments to achieving this potential are the dynamic gate phase errors that cannot be corrected via calibration. The standard tool for countering such phase errors in two-level qubits is the echo

  87. Peiyuan Feng, Yichen He, Guanhua Huang, Yuan Lin

    We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks with users, leveraging LLMs, memory, tools, and interactions with experts. The agent possesses capabilities beyond conversation, including reflection, tool usage, and expert consu

  88. Juyoung Yun, Jungmin Shin

    Solar flares, especially C, M, and X class, pose significant risks to satellite operations, communication systems, and power grids. We present a novel approach for predicting extreme solar flares using HMI intensitygrams and magnetograms. By detecting sunspots from intensitygrams and extracting magnetic field patches from magnetograms, we train a Residual Ne

  89. Ismail Lotfi, Marwa Qaraqe, Ali Ghrayeb, Niyato Dusit

    In this paper, we tackle the issue of moral hazard within the realm of the vehicular Metaverse. A pivotal facilitator of the vehicular Metaverse is the effective orchestration of its market elements, primarily comprised of sensing internet of things (SIoT) devices. These SIoT devices play a critical role by furnishing the virtual service provider (VSP) with

  90. Minheng Xiao, Xian Yu, Lei Ying

    Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL (DRL) seeks to estimate the entire distribution of it, which leads to a unified framework for handling different risk measures. Howe

  91. Georgios Chatzigeorgakidis, Konstantinos Lentzos, Dimitrios Skoutas

    Predicting future values in multivariate time series is vital across various domains. This work explores the use of large language models (LLMs) for this task. However, LLMs typically handle one-dimensional data. We introduce MultiCast, a zero-shot LLM-based approach for multivariate time series forecasting. It allows LLMs to receive multivariate time series

  92. Yanping Fu, Wenbin Liao, Xinyuan Liu, Hang xu

    As an emerging task that integrates perception and reasoning, topology reasoning in autonomous driving scenes has recently garnered widespread attention. However, existing work often emphasizes "perception over reasoning": they typically boost reasoning performance by enhancing the perception of lanes and directly adopt MLP to learn lane topology from lane q

  93. Michele Cattelan, Jemma Bennett, Sheir Yarkoni, Wolfgang Lechner

    One of the main bottlenecks in solving combinatorial optimization problems with quantum annealers is the qubit connectivity in the hardware. A possible solution for larger connectivty is minor embedding. This techniques makes the geometrical properties of the combinatorial optimization problem, encoded as a Hamiltonian, match the properties of the quantum an

  94. Doheon Han, Nuno Moniz, Nitesh V Chawla

    Many evaluation metrics can be used to assess the performance of models in binary classification tasks. However, most of them are derived from a confusion matrix in a non-differentiable form, making it very difficult to generate a differentiable loss function that could directly optimize them. The lack of solutions to bridge this challenge not only hinders o

  95. Xuan Liu, Jie Zhang, Haoyang Shang, Song Guo

    Large language models (LLMs) have been shown to face hallucination issues due to the data they trained on often containing human bias; whether this is reflected in the decision-making process of LLM Agents remains under-explored. As LLM Agents are increasingly employed in intricate social environments, a pressing and natural question emerges: Can we utilize

  96. Kaihua Ding, Jingsong Cui, Mohammad Soltani, Jing Jin

    The field of causal Machine Learning (ML) has made significant strides in recent years. Notable breakthroughs include methods such as meta learners (arXiv:1706.03461v6) and heterogeneous doubly robust estimators (arXiv:2004.14497) introduced in the last five years. Despite these advancements, the field still faces challenges, particularly in managing tightly

  97. Zhuo Xu, Lu Bai, Lixin Cui, Ming Li

    Graph Auto-Encoders (GAEs) are powerful tools for graph representation learning. In this paper, we develop a novel Hierarchical Cluster-based GAE (HC-GAE), that can learn effective structural characteristics for graph data analysis. To this end, during the encoding process, we commence by utilizing the hard node assignment to decompose a sample graph into a

  98. Huajie Qian, Donghao Ying, Henry Lam, Wotao Yin

    Ensemble learning is a popular technique to improve the accuracy of machine learning models. It traditionally hinges on the rationale that aggregating multiple weak models can lead to better models with lower variance and hence higher stability, especially for discontinuous base learners. In this paper, we provide a new perspective on ensembling. By selectin

  99. Amavi Dossa, El Mehdi Amhoud

    In the current context of massive IoT, the Pure-Aloha scheme used in LoRaWAN is reaching its limit, and Slotted-Aloha is being considered as an alternative, as it offers twice Pure-Aloha's packet success rate. It however requires synchronization across the nodes. In this paper, we propose a new slot structure adapted to devices with low quality clock, and a

  100. Chongjie Si, Xuehui Wang, Xue Yang, Zhengqin Xu

    Adapting pre-trained foundation models for various downstream tasks has been prevalent in artificial intelligence. Due to the vast number of tasks and high costs, adjusting all parameters becomes unfeasible. To mitigate this, several fine-tuning techniques have been developed to update the pre-trained model weights in a more resource-efficient manner, such a