Skip to content

May 2023 arXiv papers — page 71

Showing 7,0017,100 of 19,695 papers

  1. Jiading Fang, Shengjie Lin, Igor Vasiljevic, Vitor Guizilini

    A practical benefit of implicit visual representations like Neural Radiance Fields (NeRFs) is their memory efficiency: large scenes can be efficiently stored and shared as small neural nets instead of collections of images. However, operating on these implicit visual data structures requires extending classical image-based vision techniques (e.g., registrati

  2. Abhijit Biswas, Gustavo A. Alvarez, Tao Li, Joyce Christiansen-Salameh

    Heterostructures based on ultrawide-bandgap (UWBG) semiconductors (bandgap >4.0 eV), boron nitride (BN) and diamond are important for next-generation high-power electronics. However, in-situ hetero-epitaxy of BN/diamond or vice-versa remains extremely challenging, due to their non-trivial growth kinetics. Here, we have grown BN thin film on (100) single crys

  3. Zheng-Wei Liu, Friedrich K. Roepke, Zhanwen Han

    SNe Ia play a key role in the fields of astrophysics and cosmology. It is widely accepted that SNe Ia arise from thermonuclear explosions of WDs in binaries. However, there is no consensus on the fundamental aspects of the nature of SN Ia progenitors and their explosion mechanism. This fundamentally flaws our understanding of these important astrophysical ob

  4. Wangchunshu Zhou, Yuchen Eleanor Jiang, Peng Cui, Tiannan Wang

    The fixed-size context of Transformer makes GPT models incapable of generating arbitrarily long text. In this paper, we introduce RecurrentGPT, a language-based simulacrum of the recurrence mechanism in RNNs. RecurrentGPT is built upon a large language model (LLM) such as ChatGPT and uses natural language to simulate the Long Short-Term Memory mechanism in a

  5. Jannis Vamvas, Rico Sennrich

    Automatically highlighting words that cause semantic differences between two documents could be useful for a wide range of applications. We formulate recognizing semantic differences (RSD) as a token-level regression task and study three unsupervised approaches that rely on a masked language model. To assess the approaches, we begin with basic English senten

  6. Abdullatif Köksal, Omer Faruk Yalcin, Ahmet Akbiyik, M. Tahir Kilavuz

    Pretrained language models (PLMs) are key components in NLP, but they contain strong social biases. Quantifying these biases is challenging because current methods focusing on fill-the-mask objectives are sensitive to slight changes in input. To address this, we propose a bias probing technique called LABDet, for evaluating social bias in PLMs with a robust

  7. Kevin Black, Michael Janner, Yilun Du, Ilya Kostrikov

    Diffusion models are a class of flexible generative models trained with an approximation to the log-likelihood objective. However, most use cases of diffusion models are not concerned with likelihoods, but instead with downstream objectives such as human-perceived image quality or drug effectiveness. In this paper, we investigate reinforcement learning metho

  8. Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou

    By providing external information to large language models (LLMs), tool augmentation (including retrieval augmentation) has emerged as a promising solution for addressing the limitations of LLMs' static parametric memory. However, how receptive are LLMs to such external evidence, especially when the evidence conflicts with their parametric memory? We present

  9. Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng

    In-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood. We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features

  10. Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li

    In this paper, we propose DiffusionNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans. During training, DiffusionNER gradually adds noises to the golden entity boundaries by a fixed forward diffusion process and learns a reverse diffusion process to recover the

  11. Shashank Sonkar, Richard G. Baraniuk

    This paper investigates the key role of Feed-Forward Networks (FFNs) in transformer models by utilizing the Parallel Attention and Feed-Forward Net Design (PAF) architecture, and comparing it to their Series Attention and Feed-Forward Net Design (SAF) counterparts. Central to the effectiveness of PAF are two main assumptions regarding the FFN block and the a

  12. Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang

    Large language models (LLMs) such as ChatGPT have seen widespread adoption due to their strong instruction-following abilities. Developing these LLMs involves a complex yet poorly understood workflow requiring training with human feedback. Replicating and understanding this instruction-following requires tackling three major challenges: the high cost of data

  13. Dongwei Pan, Long Zhuo, Jingtan Piao, Huiwen Luo

    Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still face great obstacles in real-world scenarios. One of the vital causes is inadequate datasets -- 1) current public datasets can only support researchers to explore high-fidelity head a

  14. Jun Li, Devkishen Sisodia, Yebo Feng, Lumin Shi

    Despite the proliferation of traffic filtering capabilities throughout the Internet, attackers continue to launch distributed denial-of-service (DDoS) attacks to successfully overwhelm the victims with DDoS traffic. In this paper, we introduce a distributed filtering system that leverages nodes distributed along the paths of DDoS traffic to filter the DDoS t

  15. Chen Firestein, Amir Shlivinski, Yakir Hadad

    The classical scenario where a \emph{single plane-wave} field impinge a Dallenbach absorber is well studied both theoretically and experimentally. However, occasionally a \emph{spectrum of plane-waves} impinges the absorber. Such a scenario occurs for example if an antenna is located adjacent to the absorbing layer. In this paper, for this scenario we obtain

  16. Yunqi Li, Lanjing Zhang, Yongfeng Zhang

    Understanding and addressing unfairness in LLMs are crucial for responsible AI deployment. However, there is a limited number of quantitative analyses and in-depth studies regarding fairness evaluations in LLMs, especially when applying LLMs to high-stakes fields. This work aims to fill this gap by providing a systematic evaluation of the effectiveness and f

  17. Michael Herrmann, Katia Kleine

    We consider a one-dimensional peridynamical medium and show the existence of solitary waves with small amplitudes and long wavelength. Our proof uses nonlinear Bochner integral operators and characterizes their asymptotic properties in a singular scaling limit.

  18. Adam Lechowicz, Rik Sengupta, Bo Sun, Shahin Kamali

    The online knapsack problem is a classic problem in the field of online algorithms. Its canonical version asks how to pack items of different values and weights arriving online into a capacity-limited knapsack so as to maximize the total value of the admitted items. Although optimal competitive algorithms are known for this problem, they may be fundamentally

  19. Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu

    With the exponential growth of video data, there is an urgent need for automated technology to analyze and comprehend video content. However, existing video understanding models are often task-specific and lack a comprehensive capability of handling diverse tasks. The success of large language models (LLMs) like GPT has demonstrated their impressive abilitie

  20. Prafull Sharma, Julien Philip, Michaël Gharbi, William T. Freeman

    Separating an image into meaningful underlying components is a crucial first step for both editing and understanding images. We present a method capable of selecting the regions of a photograph exhibiting the same material as an artist-chosen area. Our proposed approach is robust to shading, specular highlights, and cast shadows, enabling selection in real i

  21. Katharina Ott, Michael Tiemann, Philipp Hennig

    Neural ordinary differential equations (ODEs) are an emerging class of deep learning models for dynamical systems. They are particularly useful for learning an ODE vector field from observed trajectories (i.e., inverse problems). We here consider aspects of these models relevant for their application in science and engineering. Scientific predictions general

  22. Yue Wang, Jinjun Xiong, Shaofeng Zou

    Offline reinforcement learning aims to learn from pre-collected datasets without active exploration. This problem faces significant challenges, including limited data availability and distributional shifts. Existing approaches adopt a pessimistic stance towards uncertainty by penalizing rewards of under-explored state-action pairs to estimate value functions

  23. John Cardy

    The Yang-Lee edge singularity is a prototypical example of the application of renormalization group ideas to critical behavior, and one to which Michael Fisher made several important contributions. Moreover it has connections to several other problems such as the statistics of branched polymers, and its scaling limit in two dimensions provides a simple examp

  24. Aldin Vehabovic, Hadi Zanddizari, Nasir Ghani, Farooq Shaikh

    Researchers have proposed a wide range of ransomware detection and analysis schemes. However, most of these efforts have focused on older families targeting Windows 7/8 systems. Hence there is a critical need to develop efficient solutions to tackle the latest threats, many of which may have relatively fewer samples to analyze. This paper presents a machine

  25. Rochelle Choenni, Dan Garrette, Ekaterina Shutova

    Multilingual large language models (MLLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languages' data. Impressive performance on zero-shot cross-lingual transfer shows that these models are capable of exploiting data from other languages. Yet, it remains unclear to what ex

  26. Kamal Oudrhiri, James M. Kohel, Nate Harvey, James R. Kellogg

    NASA's Cold Atom Laboratory (CAL) is a multi-user science facility for studying quantum gases in the microgravity environment of the International Space Station. The persistent microgravity environment of the ISS enables research with ultracold atoms in a temperature regime and force-free environment inaccessible to terrestrial laboratories, unlocking the po

  27. Kowshik Thopalli, Rakshith Subramanyam, Pavan Turaga, Jayaraman J. Thiagarajan

    In this paper, we address the problem of adapting models from a source domain to a target domain, a task that has become increasingly important due to the brittle generalization of deep neural networks. While several test-time adaptation techniques have emerged, they typically rely on synthetic toolbox data augmentations in cases of limited target data avail

  28. Zhentao Zhang

    Uncharged parallel plates in the misaligned system can experience a tangential Casimir force between them. We consider the role of magnetic response in this effect by extending the tangential force to magnetic media, and the extension is realized by working out the total zero-point energy of multilayered magnetodielectrics. Then we investigate the tangential

  29. Flavio Chierichetti, Mirko Giacchini, Ravi Kumar, Alessandro Panconesi

    In this work we consider the problem of fitting Random Utility Models (RUMs) to user choices. Given the winner distributions of the subsets of size $k$ of a universe, we obtain a polynomial-time algorithm that finds the RUM that best approximates the given distribution on average. Our algorithm is based on a linear program that we solve using the ellipsoid m

  30. Rheeya Uppaal, Junjie Hu, Yixuan Li

    Out-of-distribution (OOD) detection is a critical task for reliable predictions over text. Fine-tuning with pre-trained language models has been a de facto procedure to derive OOD detectors with respect to in-distribution (ID) data. Despite its common use, the understanding of the role of fine-tuning and its necessity for OOD detection is largely unexplored.

  31. Roi Cohen, May Hamri, Mor Geva, Amir Globerson

    A prominent weakness of modern language models (LMs) is their tendency to generate factually incorrect text, which hinders their usability. A natural question is whether such factual errors can be detected automatically. Inspired by truth-seeking mechanisms in law, we propose a factuality evaluation framework for LMs that is based on cross-examination. Our k

  32. Jaeuk Kim, Salvatore Torquato

    Disordered stealthy hyperuniform dielectric composites exhibit novel electromagnetic wave transport properties in two and three dimensions. Here, we carry out the first study of the electromagnetic properties of one-dimensional (1D) disordered stealthy hyperuniform layered media. From an exact nonlocal theory, we derive an approximation formula for the effec

  33. Vivek Sridhar, Michael Breuß

    Sampling is a basic operation in image processing. In classic literature, a morphological sampling theorem has been established, which shows how sampling interacts by morphological operations with image reconstruction. Many aspects of morphological sampling have been investigated for binary images, but only some of them have been explored for grey-value imag

  34. F. R. Klinkhamer

    We have recently discovered a smooth vacuum-wormhole solution of the first-order equations of general relativity. Here, we obtain the corresponding multiple-vacuum-wormhole solution. Assuming that our world is essentially Minkowski spacetime with a large number of these vacuum-defect wormholes inserted, there is then another flat spacetime with opposite spat

  35. Corinne Stucker, Vivien Sainte Fare Garnot, Konrad Schindler

    Satellite image time series in the optical and infrared spectrum suffer from frequent data gaps due to cloud cover, cloud shadows, and temporary sensor outages. It has been a long-standing problem of remote sensing research how to best reconstruct the missing pixel values and obtain complete, cloud-free image sequences. We approach that problem from the pers

  36. Mithun Das, Saurabh Kumar Pandey, Animesh Mukherjee

    Hate speech is a severe issue that affects many online platforms. So far, several studies have been performed to develop robust hate speech detection systems. Large language models like ChatGPT have recently shown a great promise in performing several tasks, including hate speech detection. However, it is crucial to comprehend the limitations of these models

  37. Ziaullah Momand, Debajyoti Pal, Pornchai Mongkolnam, Jonathan H. Chan

    Child dehydration is a significant health concern, especially among children under 5 years of age who are more susceptible to diarrhea and vomiting. In Afghanistan, severe diarrhea contributes to child mortality due to dehydration. However, there is no evidence of research exploring the potential of machine learning techniques in diagnosing dehydration in Af

  38. Zhenwen Liang, Wenhao Yu, Tanmay Rajpurohit, Peter Clark

    In this paper, we present a novel approach for distilling math word problem solving capabilities from large language models (LLMs) into smaller, more efficient student models. Our approach is designed to consider the student model's weaknesses and foster a tailored learning experience by generating targeted exercises aligned with educational science principl

  39. Nikhil Mahajan, Marten H. van Kerkwijk

    Giant pulses emitted by PSR B1937+21 are bright, intrinsically impulsive bursts. Thus, the observed signal from a giant pulse is a noisy but direct measurement of the impulse response from the ionized interstellar medium. We use this fact to detect 13,025 giant pulses directly in the baseband data of two observations of PSR B1937+21. Using the giant pulse si

  40. Michal Vyvlecka, Lennart Jehle, Cornelius Nawrath, Francesco Giorgino

    Building a quantum internet requires efficient and reliable quantum hardware, from photonic sources to quantum repeaters and detectors, ideally operating at telecommunication wavelengths. Thanks to their high brightness and single-photon purity, quantum dot (QD) sources hold the promise to achieve high communication rates for quantum-secured network applicat

  41. Shashank Sonkar, Naiming Liu, Debshila Basu Mallick, Richard G. Baraniuk

    We present a design framework called Conversational Learning with Analytical Step-by-Step Strategies (CLASS) for building advanced Intelligent Tutoring Systems (ITS) powered by high-performance Large Language Models (LLMs). The CLASS framework empowers ITS with two key capabilities. First, through a carefully curated scaffolding dataset, CLASS equips ITS wit

  42. Charles Arnal, Felix Hensel, Mathieu Carrière, Théo Lacombe

    Despite their successful application to a variety of tasks, neural networks remain limited, like other machine learning methods, by their sensitivity to shifts in the data: their performance can be severely impacted by differences in distribution between the data on which they were trained and that on which they are deployed. In this article, we propose a ne

  43. Rajeev Gupta, Gadadhar Misra, Samya Kumar Ray

    We investigate a Grothendieck-type inequality for pairs of Banach spaces $E,F$ assuming $E$ is finite-dimensional and study the associated Grothendieck-type constant. We prove that if there is a $C >0$ such that $\|A\otimes \operatorname{id}_{F}\|_{E_m\check{\otimes}F\to E_n^*\hat{\otimes}F}\leqslant C \|A\|_{E_m\to E_n^*}$ for all $m,n\in\mathbb{N},$ where

  44. Xingxuan Li, Ruochen Zhao, Yew Ken Chia, Bosheng Ding

    We present chain-of-knowledge (CoK), a novel framework that augments large language models (LLMs) by dynamically incorporating grounding information from heterogeneous sources. It results in more factual rationales and reduced hallucination in generation. Specifically, CoK consists of three stages: reasoning preparation, dynamic knowledge adapting, and answe

  45. Yueting Yang, Xintong Zhang, Wenjuan Han

    Pre-trained visual language models (VLM) have shown excellent performance in image caption tasks. However, it sometimes shows insufficient reasoning ability. In contrast, large language models (LLMs) emerge with powerful reasoning capabilities. Therefore, we propose a method called TReE, which transfers the reasoning ability of a large language model to a vi

  46. Takeo Uramoto

    This paper is a sequel to our previous work, where we proved the ``modularity theorem'' for algebraic Witt vectors over imaginary quadratic fields. This theorem states that, in the case of imaginary quadratic fields $K$, the algebraic Witt vectors over $K$ are precisely those generated by the modular vectors whose components are given by special values of de

  47. Jennifer Hu, Roger Levy

    Prompting is now a dominant method for evaluating the linguistic knowledge of large language models (LLMs). While other methods directly read out models' probability distributions over strings, prompting requires models to access this internal information by processing linguistic input, thereby implicitly testing a new type of emergent ability: metalinguisti

  48. Christopher Mitcheltree, Christian J. Steinmetz, Marco Comunità, Joshua D. Reiss

    Low frequency oscillator (LFO) driven audio effects such as phaser, flanger, and chorus, modify an input signal using time-varying filters and delays, resulting in characteristic sweeping or widening effects. It has been shown that these effects can be modeled using neural networks when conditioned with the ground truth LFO signal. However, in most cases, th

  49. Jiseong Noh, Donghwan Kwon, Soohwan Cho, Neo C. K. Yiu

    The comparative analysis examined eleven Proof-of-Stake (PoS) consensus-based blockchain networks to assess their openness based on five indicative metrics. These metrics include those of decentralization-related aspects, such as the number of validators and capital concentration, and participation-related aspects, including entry capital requirements and ec

  50. David Herron, Ernesto Jiménez-Ruiz, Giacomo Tarroni, Tillman Weyde

    NeSy4VRD is a multifaceted resource designed to support the development of neurosymbolic AI (NeSy) research. NeSy4VRD re-establishes public access to the images of the VRD dataset and couples them with an extensively revised, quality-improved version of the VRD visual relationship annotations. Crucially, NeSy4VRD provides a well-aligned, companion OWL ontolo

  51. Yixin Liu, Hongsheng Hu, Xun Chen, Xuyun Zhang

    Substantial research works have shown that deep models, e.g., pre-trained models, on the large corpus can learn universal language representations, which are beneficial for downstream NLP tasks. However, these powerful models are also vulnerable to various privacy attacks, while much sensitive information exists in the training dataset. The attacker can easi

  52. Joongwon Kim, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi

    Recent work in NLP has shown promising results in training models on large amounts of tasks to achieve better generalization. However, it is not well-understood how tasks are related, and how helpful training tasks can be chosen for a new task. In this work, we investigate whether knowing task relationships via pairwise task transfer improves choosing one or

  53. Leon A. Luxemburg, Steven B. Damelin

    In this paper we study the scale-space classification of signals via the maximal set of kernels. We use a geometric approach which arises naturally when we consider parameter variations in scale-space. We derive the Fourier transform formulas for quick and efficient computation of zero-crossings and the corresponding classifying trees. General theory of conv

  54. Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie

    Large Language Models (LLMs) may hallucinate and generate fake information, despite pre-training on factual data. Inspired by the journalistic device of "according to sources", we propose according-to prompting: directing LLMs to ground responses against previously observed text. To quantify this grounding, we propose a novel evaluation metric (QUIP-Score) t

  55. Xiaofan Zhou, Xunzhu Tang

    Electronic Health Record (EHR) coding involves automatically classifying EHRs into diagnostic codes. While most previous research treats this as a multi-label classification task, generating probabilities for each code and selecting those above a certain threshold as labels, these approaches often overlook the challenge of identifying complex diseases. In th

  56. Katharina Ott, Michael Tiemann, Philipp Hennig, François-Xavier Briol

    Bayesian probabilistic numerical methods for numerical integration offer significant advantages over their non-Bayesian counterparts: they can encode prior information about the integrand, and can quantify uncertainty over estimates of an integral. However, the most popular algorithm in this class, Bayesian quadrature, is based on Gaussian process models and

  57. Moshe Babaioff, Shahar Dobzinski, Shiri Ron

    We explore the performance of polynomial-time incentive-compatible mechanisms in single-crossing domains. Single-crossing domains were extensively studied in the economics literature. Roughly speaking, a domain is single crossing if monotonicity characterizes incentive compatibility. That is, single-crossing domains are the standard mathematical formulation

  58. Zekun Wang, Ge Zhang, Kexin Yang, Ning Shi

    Interactive Natural Language Processing (iNLP) has emerged as a novel paradigm within the field of NLP, aimed at addressing limitations in existing frameworks while aligning with the ultimate goals of artificial intelligence. This paradigm considers language models as agents capable of observing, acting, and receiving feedback iteratively from external entit

  59. Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy

    Multi-query attention (MQA), which only uses a single key-value head, drastically speeds up decoder inference. However, MQA can lead to quality degradation, and moreover it may not be desirable to train a separate model just for faster inference. We (1) propose a recipe for uptraining existing multi-head language model checkpoints into models with MQA using

  60. Leon A. Luxemburg, Steven B. Damelin

    In this article we construct a maximal set of kernels for a multi-parameter linear scale-space that allow us to construct trees for classification and recognition of one-dimensional continuous signals similar the Gaussian linear scale-space approach. Fourier transform formulas are provided and used for quick and efficient computations. A number of useful pro

  61. Javier Cuerda, Jani M. Taskinen, Nicki Källman, Leo Grabitz

    We theoretically predict the full quantum geometric tensor, comprising the quantum metric and the Berry curvature, for a square lattice of plasmonic nanoparticles. The gold nanoparticles act as dipole or multipole antenna radiatively coupled over long distances. The photonic-plasmonic eigenfunctions and energies of the system depend on momentum and polarizat

  62. Jason Blocklove, Siddharth Garg, Ramesh Karri, Hammond Pearce

    Modern hardware design starts with specifications provided in natural language. These are then translated by hardware engineers into appropriate Hardware Description Languages (HDLs) such as Verilog before synthesizing circuit elements. Automating this translation could reduce sources of human error from the engineering process. But, it is only recently that

  63. Yafu Li, Qintong Li, Leyang Cui, Wei Bi

    Large language models (LLMs) have achieved human-level text generation, emphasizing the need for effective AI-generated text detection to mitigate risks like the spread of fake news and plagiarism. Existing research has been constrained by evaluating detection methods on specific domains or particular language models. In practical scenarios, however, the det

  64. Ben L. Titzer

    Compilers face an intrinsic tradeoff between compilation speed and code quality. The tradeoff is particularly stark in a dynamic setting where JIT compilation time contributes to application runtime. Many systems now employ multiple compilation tiers, where one tier offers fast compile speed while another has much slower compile speed but produces higher qua

  65. Mark J. Arildsen, Ji-Yao Chen, Norbert Schuch, Andreas W. W. Ludwig

    We address the key question of representation of chiral topological quantum states in (2+1) dimensions (i.e., with non-zero chiral central charge) by Projected Entangled Pair States (PEPS). A noted result (due to Wahl, Tu, Schuch, and Cirac [Phys. Rev. Lett. 111, 236805 (2013)], and Dubail and Read [Phys. Rev. B 92, 205307 (2015)]) says that this is possible

  66. Andreas Galanis, Leslie Ann Goldberg, Paulina Smolarova

    We consider the performance of Glauber dynamics for the random cluster model with real parameter $q>1$ and temperature $\beta>0$. Recent work by Helmuth, Jenssen and Perkins detailed the ordered/disordered transition of the model on random $\Delta$-regular graphs for all sufficiently large $q$ and obtained an efficient sampling algorithm for all temperatures

  67. Hanlin Li, Nicholas Vincent, Stevie Chancellor, Brent Hecht

    Many recent technological advances (e.g. ChatGPT and search engines) are possible only because of massive amounts of user-generated data produced through user interactions with computing systems or scraped from the web (e.g. behavior logs, user-generated content, and artwork). However, data producers have little say in what data is captured, how it is used,

  68. Yuhao Zhou, Xiaohong Li, Jie Hong, Rony Keppens

    Observations have shown that some filaments appear and disappear in the H$\alpha$ line wing images periodically. There have been no attempts to model these "winking filaments" thus far. The evaporation--condensation mechanism is widely used to explain the formation of solar filaments. Here, we demonstrate, for the first time, how multi-dimensional evaporatio

  69. Vahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah Muzahid

    Neural network training is inherently sequential where the layers finish the forward propagation in succession, followed by the calculation and back-propagation of gradients (based on a loss function) starting from the last layer. The sequential computations significantly slow down neural network training, especially the deeper ones. Prediction has been succ

  70. Jesus Solano, Mardhiyah Sanni, Oana-Maria Camburu, Pasquale Minervini

    Models that generate natural language explanations (NLEs) for their predictions have recently gained increasing interest. However, this approach usually demands large datasets of human-written NLEs for the ground-truth answers at training time, which can be expensive and potentially infeasible for some applications. When only a few NLEs are available (a few-

  71. Brandon Park Coy, Conor A. Nixon, Naomi Rowe-Gurney, Richard Achterberg

    In this work we present, for the first time, infrared spectra of Titan from the Spitzer Space Telescope ($2004-2009$). The data are from both the short wavelength-low resolution (SL, $5.13-14.29\mathrm{\mu m}, R\sim60-127$) and short wavelength-high resolution channels (SH, $9.89 - 19.51\mathrm{\mu m}, R\sim600$) showing the emissions of CH$_{4}$, C$_{2}$H$_

  72. Peter Wirnsberger, Borja Ibarz, George Papamakarios

    We present a machine-learning model based on normalizing flows that is trained to sample from the isobaric-isothermal ensemble. In our approach, we approximate the joint distribution of a fully-flexible triclinic simulation box and particle coordinates to achieve a desired internal pressure. This novel extension of flow-based sampling to the isobaric-isother

  73. Muzhou Yu, Linfeng Zhang, Kaisheng Ma

    The excellent performance of deep neural networks is usually accompanied by a large number of parameters and computations, which have limited their usage on the resource-limited edge devices. To address this issue, abundant methods such as pruning, quantization and knowledge distillation have been proposed to compress neural networks and achieved significant

  74. Anna Erschler, Josh Frisch, Mark Rychnovsky

    We prove that finite entropy random walks on the torsion-free Baumslag group in dimension $d=2$ have non-trivial Poisson boundary. This is in contrast with the torsion case where the situation for simple random walks on Baumslag groups is the same as for the lamplighter groups of the same dimension. Our proof uses the realization of the Baumslag group as a l

  75. Fuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng

    Recent research has highlighted the importance of dataset size in scaling language models. However, large language models (LLMs) are notoriously token-hungry during pre-training, and high-quality text data on the web is approaching its scaling limit for LLMs. To further enhance LLMs, a straightforward approach is to repeat the pre-training data for additiona

  76. Svante Janson

    Serfozo (2009, Theorem 2.65) gives a useful central limit theorem for processes with regenerative increments. Unfortunately, there is a gap in the proof. We fill this gap, and at the same time we weaken the assumptions. Furthermore, we give conditions for moment convergence in this setting. We give also further results complementing results in Serfozo (2009)

  77. Nguyen Manh Linh

    The descent method is one of the approaches to study the Brauer--Manin obstruction to the local--global principle and to weak approximation on varieties over number fields, by reducing the problem to ``descent varieties''. In recent lecture notes by Wittenberg, he formulated a ``descent conjecture'' for torsors under linear algebraic groups. The present arti

  78. Orhan Koc

    Price feeds of securities is a critical component for many financial services, allowing for collateral liquidation, margin trading, derivative pricing and more. With the advent of blockchain technology, value in reporting accurate prices without a third party has become apparent. There have been many attempts at trying to calculate prices without a third par

  79. Sean Paulsen, Michael Casey

    We present a sequential transfer learning framework for transformers on functional Magnetic Resonance Imaging (fMRI) data and demonstrate its significant benefits for decoding musical timbre. In the first of two phases, we pre-train our stacked-encoder transformer architecture on Next Thought Prediction, a self-supervised task of predicting whether or not on

  80. Ehsan Qasemi, Amani R. Maina-Kilaas, Devadutta Dash, Khalid Alsaggaf

    Humans can infer the affordance of objects by extracting related contextual preconditions for each scenario. For example, upon seeing an image of a broken cup, we can infer that this precondition prevents the cup from being used for drinking. Reasoning with preconditions of commonsense is studied in NLP where the model explicitly gets the contextual precondi

  81. Yue Zhang, Leyang Cui, Deng Cai, Xinting Huang

    Proprietary Large Language Models (LLMs), such as ChatGPT, have garnered significant attention due to their exceptional capabilities in handling a diverse range of tasks. Recent studies demonstrate that open-sourced smaller foundational models, such as 7B-size LLaMA, can also display remarkable proficiency in tackling diverse tasks when fine-tuned using inst

  82. Ryoichiro Noda

    In this paper, it is shown that if a sequence of resistance metric spaces equipped with measures converges with respect to the local Gromov-Hausdorff-vague topology, and certain non-explosion and metric-entropy conditions are satisfied, then the associated stochastic processes and their local times also converge. The metric-entropy condition can be checked b

  83. Felix Denzinger, Michael Wels, Oliver Taubmann, Florian Kordon

    Coronary artery disease (CAD) is often treated minimally invasively with a catheter being inserted into the diseased coronary vessel. If a patient exhibits a Shepherd's Crook (SC) Right Coronary Artery (RCA) - an anatomical norm variant of the coronary vasculature - the complexity of this procedure is increased. Automated reporting of this variant from coron

  84. Shuoyang Wang, Guanqun Cao

    The intrinsically infinite-dimensional features of the functional observations over multidimensional domains render the standard classification methods effectively inapplicable. To address this problem, we introduce a novel multiclass functional deep neural network (mfDNN) classifier as an innovative data mining and classification tool. Specifically, we cons

  85. Yousef K. Chahine, Ian R. Nemitz, John D. Lekki

    Multi-photon emissions constitute a fundamental source of noise in quantum repeaters and other quantum communication protocols when probabilistic photon sources are employed. In this paper, it is shown that by alternating the Bell state measurement (BSM) basis in concatenated entanglement swapping links one can automatically identify and discard many errors

  86. Rômulo M. Vermersch

    For a homeomorphism $T$ on a compact metric space $X$, a $T$-invariant Borel probability measure $\mu$ on $X$ and a measure-theoretic quasifactor $\widetilde{\mu}$ of $\mu$, we study the relationship between the local entropy of the system $(X,\mu,T)$ and of its induced system $(\mathcal{M}(X),\widetilde{\mu},\widetilde{T})$, where $\widetilde{T}$ is the hom

  87. Ahmad Qazza, Rania Saadeh

    This article presents the solution of the fractional SIR epidemic model using the Laplace residual power series method. We introduce the fractional SIR model in the sense of Caputo's derivative, it is presented by three fractional differential equations, in which the third one depends on the first coupled equations. The Laplace residual power series method i

  88. Sudipto Saha, Jonathan R. Bradley

    Additive spatial statistical models with weakly stationary process assumptions have become standard in spatial statistics. However, one disadvantage of such models is the computation time, which rapidly increases with the number of data points. The goal of this article is to apply an existing subsampling strategy to standard spatial additive models and to de

  89. Wei Dong, Chris Choy, Charles Loop, Or Litany

    Indoor scene reconstruction from monocular images has long been sought after by augmented reality and robotics developers. Recent advances in neural field representations and monocular priors have led to remarkable results in scene-level surface reconstructions. The reliance on Multilayer Perceptrons (MLP), however, significantly limits speed in training and

  90. William Johnston, Rebecca G. Wahl

    This paper extends topics in linear algebra and operator theory for linear transformations on complex vector spaces to those on bicomplex Hilbert and Banach spaces. For example, Definition 3 for the first time defines a bicomplex vector space, its dimension, and its basis in terms of a corresponding vectorial idempotent representation, and the paper shows ho

  91. Lucia Absalom Bautista, Timotej Hrga, Janez Povh, Shudian Zhao

    The clustering of data is one of the most important and challenging topics in data science. The minimum sum-of-squares clustering (MSSC) problem asks to cluster the data points into $k$ clusters such that the sum of squared distances between the data points and their cluster centers (centroids) is minimized. This problem is NP-hard, but there exist exact sol

  92. Kyuil Cho, M. Konczykowski, M. A. Tanatar, I. I. Mazin

    Low-temperature variable-energy electron irradiation was used to induce non-magnetic disorder in a single crystal of hole-doped iron-based superconductor, Ba$_{1-x}$K$_x$Fe$_2$As$_2$, $x=$0.80. To avoid systematic errors, the beam energy was adjusted non-consequently for five values between 1.0 and 2.5 MeV, whence sample resistance was measured in-situ at 22

  93. Theodore R. Gull, Henrik Hartman, Mairan Teodoro, D. John Hillier

    Previous STIS long-slit observations of Eta Carinae identified numerous absorption features in both the stellar spectrum, and in the adjacent nebular spectra, along our line-of-sight. The absorption features became temporarily stronger when the ionizing FUV radiation field was reduced by the periastron passage of the secondary star. Subsequently, dissipation

  94. Kamal Basulaiman, Masoud Barati

    Power system state forecasting has gained more attention in real-time operations recently. Unique challenges to energy systems are emerging with the massive deployment of renewable energy resources. As a result, power system state forecasting are becoming more crucial for monitoring, operating and securing modern power systems. This paper proposes an end-to-

  95. Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Oana-Maria Camburu

    With recent advances, neural models can achieve human-level performance on various natural language tasks. However, there are no guarantees that any explanations from these models are faithful, i.e. that they reflect the inner workings of the model. Atomic inference overcomes this issue, providing interpretable and faithful model decisions. This approach inv

  96. Somenath Pal, Guruprasad Kadam, Abhijit Bhattacharyya

    We investigate the effect of repulsive interaction between hadrons on the susceptibilities of conserved charges, namely baryon number (B), electric charge (Q) and strangeness (S). We estimate second fourth and sixth-order susceptibilities of conserved charges, their differences, ratios, and correlations within the ambit of the mean-field hadron resonance gas

  97. Inês Nolasco, Shubhr Singh, Veronica Morfi, Vincent Lostanlen

    Automatic detection and classification of animal sounds has many applications in biodiversity monitoring and animal behaviour. In the past twenty years, the volume of digitised wildlife sound available has massively increased, and automatic classification through deep learning now shows strong results. However, bioacoustics is not a single task but a vast ra

  98. Arun Ganesh, Mahdi Haghifam, Thomas Steinke, Abhradeep Thakurta

    Differentially private (stochastic) gradient descent is the workhorse of DP private machine learning in both the convex and non-convex settings. Without privacy constraints, second-order methods, like Newton's method, converge faster than first-order methods like gradient descent. In this work, we investigate the prospect of using the second-order informatio

  99. Sayed Erfan Arefin, Tasnia Ashrafi Heya, Jia Uddin

    Establishing a communication bridge by transferring data driven from different embedded sensors via internet or reconcilable network protocols between enormous number of distinctively addressable objects or "things", is known as the Internet of Things (IoT). IoT can be amalgamated with multitudinous objects such as thermostats, cars, lights, refrigerators, a

  100. Jannis Weil, Johannes Czech, Tobias Meuser, Kristian Kersting

    In combination with Reinforcement Learning, Monte-Carlo Tree Search has shown to outperform human grandmasters in games such as Chess, Shogi and Go with little to no prior domain knowledge. However, most classical use cases only feature up to two players. Scaling the search to an arbitrary number of players presents a computational challenge, especially if d