Skip to content

May 2023 arXiv papers — page 54

Showing 5,3015,400 of 19,695 papers

  1. Veniamin Veselovsky, Manoel Horta Ribeiro, Akhil Arora, Martin Josifoski

    Large Language Models (LLMs) have democratized synthetic data generation, which in turn has the potential to simplify and broaden a wide gamut of NLP tasks. Here, we tackle a pervasive problem in synthetic data generation: its generative distribution often differs from the distribution of real-world data researchers care about (in other words, it is unfaithf

  2. Yotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald

    The field of Natural Language Generation (NLG) suffers from a severe shortage of labeled data due to the extremely expensive and time-consuming process involved in manual annotation. A natural approach for coping with this problem is active learning (AL), a well-known machine learning technique for improving annotation efficiency by selectively choosing the

  3. E. B. Balbutsev, I. V. Molodtsova

    The solution of time dependent Hartree-Fock-Bogoliubov equations by the Wigner function moments method predicts four low-lying $1^+$ states. Three of them are known as various scissors modes. Fourth state is disposed below all scissors modes and has the electrical nature. It is found that it represents one of three branches of $2^+$ state which can exist in

  4. Liying Cheng, Xingxuan Li, Lidong Bing

    As large language models (LLMs) have demonstrated their powerful capabilities in plenty of domains and tasks, including context understanding, code generation, language generation, data storytelling, etc., many data analysts may raise concerns if their jobs will be replaced by artificial intelligence (AI). This controversial topic has drawn great attention i

  5. Zhengxue Ren, Serdar Elhatisari, Timo A. Lähde, Dean Lee

    We present a systematic ab initio study of clustering in hot dilute nuclear matter using nuclear lattice effective field theory with an SU(4)-symmetric interaction. We introduce a method called light-cluster distillation to determine the abundances of dimers, trimers, and alpha clusters as a function of density and temperature. Our lattice results are compar

  6. Junchen Fu, Fajie Yuan, Yu Song, Zheng Yuan

    Adapters, a plug-in neural network module with some tunable parameters, have emerged as a parameter-efficient transfer learning technique for adapting pre-trained models to downstream tasks, especially for natural language processing (NLP) and computer vision (CV) fields. Meanwhile, learning recommendation models directly from raw item modality features -- e

  7. Wei-Lin Chen, Cheng-Kuang Wu, Yun-Nung Chen, Hsin-Hsi Chen

    Large language models (LLMs) have exhibited striking in-context learning (ICL) ability to adapt to target tasks with a few input-output demonstrations. For better ICL, different methods are proposed to select representative demonstrations from existing training corpora. However, such settings are not aligned with real-world practices, as end-users usually qu

  8. Adam Kubica, Katarzyna Ryszewska, Rico Zacher

    We study the regularity of weak solutions to evolution equations with distributed order fractional time derivative. We prove a weak Harnack inequality for nonnegative weak supersolutions and H\"older continuity of weak solutions to this problem. Our results substantially generalise analogous known results for the problem with single order fractional time der

  9. Zekun Wang, Jingchang Chen, Wangchunshu Zhou, Haichao Zhu

    Despite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, significantly hampering their efficiency in real-world applications. Moreover, the degree of redundancy in token representations and model parameters, such as attention heads, varies signi

  10. Xinpeng Wang, Leonie Weissweiler, Hinrich Schütze, Barbara Plank

    Recently, various intermediate layer distillation (ILD) objectives have been shown to improve compression of BERT models via Knowledge Distillation (KD). However, a comprehensive evaluation of the objectives in both task-specific and task-agnostic settings is lacking. To the best of our knowledge, this is the first work comprehensively evaluating distillatio

  11. Jonas Beyrer, Fanny Kassel

    For any integers $p\geq 2$ and $q\geq 1$, let $\mathbb{H}^{p,q}$ be the pseudo-Riemannian hyperbolic space of signature $(p,q)$. We prove that if $\Gamma$ is the fundamental group of a closed aspherical $p$-manifold, then the set of representations of $\Gamma$ to $\mathrm{PO}(p,q+1)$ which are convex cocompact in $\mathbb{H}^{p,q}$ is a union of connected co

  12. Javier Molina-Vilaplana

    We use a post-Gaussian variational approach to non-perturbatively study a general class of interacting bosonic quantum field theories with generalized dipole symmetries and fractonic behaviour. We find that while a Gaussian approach allows to carry out a consistent renormalization group (RG) flow analysis of these theories, this only grasps the interaction t

  13. Shilv Cai, Liqun Chen, Sheng Zhong, Luxin Yan

    Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vital for meeting the requirements of storage and transmission i

  14. Jamie Le Signe, Thomas McDermott, Eros Mariani

    The coupling of superconducting systems to mechanical resonators is an emerging field, with wide reaching implications including high precision sensing and metrology. Experimental signatures of this coupling have so far been small, seldom and often reliant on high frequency AC electronics. To overcome this limitation, in this work we consider a mechanical re

  15. Heming Xia, Qingxiu Dong, Lei Li, Jingjing Xu

    Recently, Large Language Models (LLMs) have been serving as general-purpose interfaces, posing a significant demand for comprehensive visual knowledge. However, it remains unclear how well current LLMs and their visually augmented counterparts (VaLMs) can master visual commonsense knowledge. To investigate this, we propose ImageNetVC, a human-annotated datas

  16. Veit David Wild, Sahra Ghalebikesabi, Dino Sejdinovic, Jeremias Knoblauch

    We establish the first mathematically rigorous link between Bayesian, variational Bayesian, and ensemble methods. A key step towards this it to reformulate the non-convex optimisation problem typically encountered in deep learning as a convex optimisation in the space of probability measures. On a technical level, our contribution amounts to studying general

  17. Rodrigo Valerio, Joao Bordalo, Michal Yarom, Yonatan Bitton

    Text to image generation methods (T2I) are widely popular in generating art and other creative artifacts. While visual hallucinations can be a positive factor in scenarios where creativity is appreciated, such artifacts are poorly suited for cases where the generated image needs to be grounded in complex natural language without explicit visual elements. In

  18. Tianyu Yang, Thy Thy Tran, Iryna Gurevych

    Current variational dialog models have employed pre-trained language models (PLMs) to parameterize the likelihood and posterior distributions. However, the Gaussian assumption made on the prior distribution is incompatible with these distributions, thus restricting the diversity of generated responses. These models also suffer from posterior collapse, i.e.,

  19. Biao Zhao, Weiqiang Jin, Javier Del Ser, Guang Yang

    In the era of sustainable smart agriculture, a massive amount of agricultural news text is being posted on the Internet, in which massive agricultural knowledge has been accumulated. In this context, it is urgent to explore effective text classification techniques for users to access the required agricultural knowledge with high efficiency. Mainstream deep l

  20. Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen

    Recently, growing interest has been aroused in extending the multimodal capability of large language models (LLMs), e.g., vision-language (VL) learning, which is regarded as the next milestone of artificial general intelligence. However, existing solutions are prohibitively expensive, which not only need to optimize excessive parameters, but also require ano

  21. Annie Gray, Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley

    In this paper we offer a new perspective on the well established agglomerative clustering algorithm, focusing on recovery of hierarchical structure. We recommend a simple variant of the standard algorithm, in which clusters are merged by maximum average dot product and not, for example, by minimum distance or within-cluster variance. We demonstrate that the

  22. Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang

    Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work, we introduce EmbodiedGPT, an end-to-end multi-modal foundation model for embodied AI, empowering embodied agents with multi-modal understanding and execution capabilities. To ach

  23. Asahi Ushio, Yi Zhou, Jose Camacho-Collados

    Multilingual language model (LM) have become a powerful tool in NLP especially for non-English languages. Nevertheless, model parameters of multilingual LMs remain large due to the larger embedding matrix of the vocabulary covering tokens in different languages. On the contrary, monolingual LMs can be trained in a target language with the language-specific v

  24. Anurag Dey, Probal Chaudhuri

    Several well known estimators of finite population mean and its functions are investigated under some standard sampling designs. Such functions of mean include the variance, the correlation coefficient and the regression coefficient in the population as special cases. We compare the performance of these estimators under different sampling designs based on th

  25. Jian Leng, Fan Yang, Xiang-Bin Wang

    The decomposition for controlled-$ZX$ gate in [Phys. Rev. A, 87, 062318 (2013)] has a shallow circuit depth $8n-20$ with no ancilla. Here we modify this decomposition to decompose $n$-qubit Toffoli gate with only $2n-3$ additional single-qubit gates. The circuit depth is unchanged and no ancilla is needed. We explicitly show that the circuit after decomposit

  26. Marek Kadlčík, Michal Štefánik, Ondřej Sotolář, Vlastimil Martinek

    Despite outstanding performance in many tasks, language models are notoriously inclined to make factual errors in tasks requiring arithmetic computation. We address this deficiency by creating Calc-X, a collection of datasets that demonstrates the appropriate use of a calculator in reasoning chains. Calc-X is suitable for teaching language models to offload

  27. Kostis Gourgoulias, Najah Ghalyan, Maxime Labonne, Yash Satsangi

    This paper introduces an unsupervised method to estimate the class separability of text datasets from a topological point of view. Using persistent homology, we demonstrate how tracking the evolution of embedding manifolds during training can inform about class separability. More specifically, we show how this technique can be applied to detect when the trai

  28. Daniel Reich, Felix Putze, Tanja Schultz

    Metrics for Visual Grounding (VG) in Visual Question Answering (VQA) systems primarily aim to measure a system's reliance on relevant parts of the image when inferring an answer to the given question. Lack of VG has been a common problem among state-of-the-art VQA systems and can manifest in over-reliance on irrelevant image parts or a disregard for the visu

  29. Xingxuan Li, Liying Cheng, Qingyu Tan, Hwee Tou Ng

    The temporal aspect is a significant dimension of our reality. We notice the challenge that large language models (LLMs) face when engaging in temporal reasoning. Our preliminary experiments show that methods involving the generation of intermediate reasoning steps, such as chain-of-thought and program-aided language models, do not consistently boost the per

  30. A. D. Alhaidari, A. Laradji

    An algebraic system is introduced, which is very useful for doing scattering calculations in quantum field theory. It is the set of all real numbers greater than or equal to -m^2 with parity designation and a special rule for addition and subtraction, where m is the rest mass of the scattered particle.

  31. Linxuan Pan, Shenghui Song

    With multiple iterations of updates, local statistical gradient descent (L-SGD) has been proven to be very effective in distributed machine learning schemes such as federated learning. In fact, many innovative works have shown that L-SGD with independent and identically distributed (IID) data can even outperform SGD. As a result, extensive efforts have been

  32. Jitendra Joshi, Mir Alimuddin, T S Mahesh, Manik Banik

    The phenomenon of quantum entanglement underlies several important protocols that enable emerging quantum technologies. Entangled states, however, are extremely delicate and often get perturbed by tiny fluctuations in their external environment. Certification of entanglement is therefore immensely crucial for the successful implementation of protocols involv

  33. Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji

    Instruction tuning has shown great promise in improving the performance of large language models. However, research on multilingual instruction tuning has been limited due to the scarcity of high-quality instruction-response datasets across different languages. To bridge this gap, we present Bactrian-X, a comprehensive multilingual parallel dataset of 3.4 mi

  34. Hongbo Zhang, Xiang Wan, Benyou Wang

    Pre-trained language models (PLMs) were considered to be able to store relational knowledge present in the training data. However, some relational knowledge seems to be discarded unsafely in PLMs due to \textbf{report bias}: low-frequency relational knowledge might be underexpressed compared to high-frequency one in PLMs. This gives us a hint that relational

  35. Elena Losero, Valentin Goblot, Yuchun Zhu, Hossein Babashah

    Single-crystal diamond substrates presenting a high concentration of negatively charged nitrogen-vacancy centers (NV-) are on high demand for the development of optically pumped solid-state sensors such as magnetometers, thermometers or electrometers. While nitrogen impurities can be easily incorporated during crystal growth, the creation of vacancies requir

  36. Aman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu

    LLM-powered chatbots are becoming widely adopted in applications such as healthcare, personal assistants, industry hiring decisions, etc. In many of these cases, chatbots are fed sensitive, personal information in their prompts, as samples for in-context learning, retrieved records from a database, or as part of the conversation. The information provided in

  37. Jacopo Giordano, Angelo Cenedese

    In this paper, a robust control solution for a satellite equipped with a robotic manipulator is presented. First, the dynamic model of the system is derived based on quaternions to describe the evolution of the attitude of the base satellite. Then, a non-singular terminal sliding mode controller that employs quaternions for attitude control, is proposed for

  38. Michael Gebauer, Faraz Maschhur, Nicola Leschke, Elias Grünewald

    Machine-readable representations of privacy policies are door openers for a broad variety of novel privacy-enhancing and, in particular, transparency-enhancing technologies (TETs). In order to generate such representations, transparency information needs to be extracted from written privacy policies. However, respective manual annotation and extraction proce

  39. Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Jialin Pan

    Sentiment analysis (SA) has been a long-standing research area in natural language processing. It can offer rich insights into human sentiments and opinions and has thus seen considerable interest from both academia and industry. With the advent of large language models (LLMs) such as ChatGPT, there is a great potential for their employment on SA problems. H

  40. Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng

    Generated texts from large language models (LLMs) are remarkably close to high-quality human-authored text, raising concerns about their potential misuse in spreading false information and academic misconduct. Consequently, there is an urgent need for a highly practical detection tool capable of accurately identifying the source of a given text. However, exi

  41. Ashwin George, Luciano Cavalcante Siebert, David Abbink, Arkady Zgonnikov

    Modelling causal responsibility in multi-agent spatial interactions is crucial for safety and efficiency of interactions of humans with autonomous agents. However, current formal metrics and models of responsibility either lack grounding in ethical and philosophical concepts of responsibility, or cannot be applied to spatial interactions. In this work we pro

  42. Asahi Ushio, Jose Camacho Collados, Steven Schockaert

    Relations such as "is influenced by", "is known for" or "is a competitor of" are inherently graded: we can rank entity pairs based on how well they satisfy these relations, but it is hard to draw a line between those pairs that satisfy them and those that do not. Such graded relations play a central role in many applications, yet they are typically not cover

  43. Aleksandar Stanić, Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber

    Current state-of-the-art object-centric models use slots and attention-based routing for binding. However, this class of models has several conceptual limitations: the number of slots is hardwired; all slots have equal capacity; training has high computational cost; there are no object-level relational factors within slots. Synchrony-based models in principl

  44. Christian Berger, Lívio Rodrigues, Hans P. Reiser, Vinicius Cogo

    Blockchain technology sparked renewed interest in planetary-scale Byzantine fault-tolerant (BFT) state machine replication (SMR). While recent works predominantly focused on improving the scalability and throughput of these protocols, few of them addressed latency. We present Mercury, a novel transformation to autonomously optimize the latency of quorum-base

  45. Jingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu

    Chain-of-Thought (CoT) prompting enables large language models to solve complex reasoning problems by generating intermediate steps. However, confined by its inherent single-pass and sequential generation process, CoT heavily relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers. In contrast, humans adopt r

  46. Saba Ahmadi, Aishwarya Agrawal

    Recently, reference-free metrics such as CLIPScore (Hessel et al., 2021), UMIC (Lee et al., 2021), and PAC-S (Sarto et al., 2023) have been proposed for automatic reference-free evaluation of image captions. Our focus lies in evaluating the robustness of these metrics in scenarios that require distinguishing between two captions with high lexical overlap but

  47. Zhaowei Chang, Jianhua Zhang, Pan Tang, Lei Tian

    Terahertz (THz) communication is envisioned as one of the possible technologies for the sixth-generation (6G) communication system due to its rich spectrum. To evaluate the performance of THz communication, it is essential to propose THz channel models within the common framework of the geometry-based stochastic model (GBSM) in the 3rd Generation Partnership

  48. Chongjian Yue, Xinrun Xu, Xiaojun Ma, Lun Du

    Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains underexplored. In this research, we specialize in harnessing the potential of LLMs to comprehend critical information from financial r

  49. Shaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan

    We present ACL OCL, a scholarly corpus derived from the ACL Anthology to assist Open scientific research in the Computational Linguistics domain. Integrating and enhancing the previous versions of the ACL Anthology, the ACL OCL contributes metadata, PDF files, citation graphs and additional structured full texts with sections, figures, and links to a large k

  50. Benjamin Arras

    In these notes, we obtain new stability estimates for centered non-degenerate selfdecomposable probability measures on $\mathbb{R}^d$ with finite second moment and for non-degenerate symmetric $\alpha$-stable probability measures on $\mathbb{R}^d$ with $\alpha \in [1,2)$. These new results are refinements of the corresponding ones available in the literature

  51. Dongjie Yang, Ruifeng Yuan, Yuantao Fan, Yifei Yang

    Large Language Models (LLMs) have attained the impressive capability to resolve a wide range of NLP tasks by fine-tuning high-quality instruction data. However, collecting human-written data of high quality, especially multi-turn dialogues, is expensive and unattainable for most people. Though previous studies have used powerful LLMs to generate the dialogue

  52. Sweta Agrawal, Marine Carpuat

    Text simplification (TS) systems rewrite text to make it more readable while preserving its content. However, what makes a text easy to read depends on the intended readers. Recent work has shown that pre-trained language models can simplify text using a wealth of techniques to control output simplicity, ranging from specifying only the desired reading grade

  53. Andrew T. Hyman

    The theory of point-particles in classical electrodynamics has a well-known problem of infinite self-energy, and the same is true of quantum electrodynamics. Instead of concluding that there is no such thing as a true point-particle, it is shown here how to remove the infinities by supposing that the electromagnetic field tensor has a symmetric part. This do

  54. Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong

    Large language models (LLMs) have shown remarkable reasoning capabilities, especially when prompted to generate intermediate reasoning steps (e.g., Chain-of-Thought, CoT). However, LLMs can still struggle with problems that are easy for humans, such as generating action plans for executing tasks in a given environment, or performing complex math, logical, an

  55. Taelin Karidi, Leshem Choshen, Gal Patel, Omri Abend

    We propose a novel methodology (namely, MuLER) that transforms any reference-based evaluation metric for text generation, such as machine translation (MT) into a fine-grained analysis tool. Given a system and a metric, MuLER quantifies how much the chosen metric penalizes specific error types (e.g., errors in translating names of locations). MuLER thus enabl

  56. Akhil Reddy Peeketi, Edwin Joseph, Narasimhan Swaminathan, Ratna Kumar Annabattula

    We use molecular dynamics simulations to unravel the physics underpinning the light-induced density changes caused by the dynamic trans-cis-trans isomerization cycles of azo-mesogens embedded in a liquid crystal polymer network, an intriguing experimental observation reported in the literature. We employ two approaches, cyclic and probabilistic switching of

  57. Hua Cai, Xuli Shen, Qing Xu, Weilin Shen

    In empathetic conversations, individuals express their empathy towards others. Previous work has mainly focused on generating empathetic responses by utilizing the speaker's emotion. Besides, external commonsense knowledge has been applied to enhance the system's understandings of the speaker's situation. However, given an event, commonsense knowledge base c

  58. El Moatez Billah Nagoudi, AbdelRahim Elmadany, Ahmed El-Shangiti, Muhammad Abdul-Mageed

    We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range of 13 different NLG tasks, including dialogue generation, question answering, machine translation, summarization, among

  59. Shraddha Rajkhowa, Nipen Saikia

    We deduce $q$-continued fractions $S_{1}(q)$, $S_{2}(q)$ and $S_{3}(q)$ of order fourteen, and continued fractions $V_{1}(q)$, $V_{2}(q)$ and $V_{3}(q)$ of order twenty-eight from a general continued fraction identity of Ramanujan. We establish some theta-function identities for the continued fractions and derive some colour partition identities as applicati

  60. Yilun Zhao, Haowei Zhang, Shengyun Si, Linyong Nan

    Tabular data is prevalent across various industries, necessitating significant time and effort for users to understand and manipulate for their information-seeking purposes. The advancements in large language models (LLMs) have shown enormous potential to improve user efficiency. However, the adoption of LLMs in real-world applications for table information

  61. Gorana Gojić, Vladimir Vincan, Ognjen Kundačina, Dragiša Mišković

    Non-adversarial robustness, also known as natural robustness, is a property of deep learning models that enables them to maintain performance even when faced with distribution shifts caused by natural variations in data. However, achieving this property is challenging because it is difficult to predict in advance the types of distribution shifts that may occ

  62. Manuel Glöckler, Michael Deistler, Jakob H. Macke

    Bayesian inference usually requires running potentially costly inference procedures separately for every new observation. In contrast, the idea of amortized Bayesian inference is to initially invest computational cost in training an inference network on simulated data, which can subsequently be used to rapidly perform inference (i.e., to return estimates of

  63. Benjamin Lowe, Bernard Field, Jack Hellerstedt, Julian Ceddia

    Electron-electron interactions in materials lead to exotic many-body quantum phenomena including Mott metal-insulator transitions (MITs), magnetism, quantum spin liquids, and superconductivity. These phases depend on electronic band occupation and can be controlled via the chemical potential. Flat bands in two-dimensional (2D) and layered materials with a ka

  64. Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain

    Recent advancements in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. Despite this progress, these models lack specific benchmarking against state-of-the-art (SOTA) models tailored to particular languages and tasks. LAraBench addresses this gap for Arabic Natural Language Processing (NLP) and Speech

  65. Tanay Dixit, Fei Wang, Muhao Chen

    Improving factual consistency of abstractive summarization has been a widely studied topic. However, most of the prior works on training factuality-aware models have ignored the negative effect it has on summary quality. We propose EFACTSUM (i.e., Effective Factual Summarization), a candidate summary generation and ranking technique to improve summary factua

  66. Yuehong Chen, Yu Dai, Mingde Ding

    Recent observations in extreme-ultraviolet (EUV) wavelengths reveal an EUV late phase in some solar flares, which is characterized by a second peak in the warm coronal emissions (about 3 MK) occurring several tens of minutes to a few hours after the corresponding main flare peak. We aim to clarify the physical origin of an atypical plateau-like EUV late phas

  67. Gabriel Kasmi, Laurent Dubus, Yves-Marie Saint Drenan, Philippe Blanc

    Neural networks have shown remarkable performance in computer vision, but their deployment in numerous scientific and technical fields is challenging due to their black-box nature. Scientists and practitioners need to evaluate the reliability of a decision, i.e., to know simultaneously if a model relies on the relevant features and whether these features are

  68. Nathanael Bosch, Philipp Hennig, Filip Tronarp

    Probabilistic solvers provide a flexible and efficient framework for simulation, uncertainty quantification, and inference in dynamical systems. However, like standard solvers, they suffer performance penalties for certain stiff systems, where small steps are required not for reasons of numerical accuracy but for the sake of stability. This issue is greatly

  69. Florian Heidecker, Ahmad El-Khateeb, Bernhard Sick

    The examination of uncertainty in the predictions of machine learning (ML) models is receiving increasing attention. One uncertainty modeling technique used for this purpose is Monte-Carlo (MC)-Dropout, where repeated predictions are generated for a single input. Therefore, clustering is required to describe the resulting uncertainty, but only through effici

  70. Kaimin Wang, Haoran Liu, Peng Li, Mingzhe Liu

    The publicly accessible dataset includes neutron and gamma-ray pulse signals for conducting pulse shape discrimination experiments. Several traditional and recently proposed pulse shape discrimination algorithms are utilized to evaluate the performance of pulse shape discrimination under raw pulse signals and noise-enhanced datasets. These algorithms compris

  71. Md Tawkat Islam Khondaker, Abdul Waheed, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed

    ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT's capabilities on Arabic

  72. Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma

    A trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to an expert in cases of low-confidence predictions. Recent studies have shown that unsupervised pre-training produces large language models (LMs

  73. David Kappel, Khaleelulla Khan Nazeer, Cabrel Teguemne Fokam, Christian Mayr

    The ubiquitous backpropagation algorithm requires sequential updates through the network introducing a locking problem. In addition, back-propagation relies on the transpose of forward weight matrices to compute updates, introducing a weight transport problem across the network. Locking and weight transport are problems because they prevent efficient paralle

  74. Jiazheng Li, Runcong Zhao, Yongxin Yang, Yulan He

    The remarkable performance of pre-trained large language models has revolutionised various natural language processing applications. Due to huge parametersizes and extensive running costs, companies or organisations tend to transfer the models to the target task by zero-shot prompting techniques. However, the prohibitive costs of tokens and time have hindere

  75. Ciprian-Octavian Truică, Neculai-Ovidiu Istrate, Elena-Simona Apostol

    Automatic Term Recognition is used to extract domain-specific terms that belong to a given domain. In order to be accurate, these corpus and language-dependent methods require large volumes of textual data that need to be processed to extract candidate terms that are afterward scored according to a given metric. To improve text preprocessing and candidate te

  76. Nicholas G. Polson, Vadim Sokolov

    Bayesian Generative AI (BayesGen-AI) methods are developed and applied to Bayesian computation. BayesGen-AI reconstructs the posterior distribution by directly modeling the parameter of interest as a mapping (a.k.a. deep learner) from a large simulated dataset. This provides a generator that we can evaluate at the observed data and provide draws from the pos

  77. M. Caramazza, B. Stelzer, E. Magaudda, St. Raetz

    We have embarked in a systematic study of the X-ray emission in a volume-limited sample of M dwarf stars, in order to explore the full range of activity levels present in their coronae and, thus, to understand the conditions in their outer atmospheres and their possible impact on the circumstellar environment. We identify in a recent catalog of the Gaia obje

  78. Tianqing Fang, Zhaowei Wang, Wenxuan Zhou, Hongming Zhang

    Event temporal reasoning aims at identifying the temporal relations between two or more events from narratives. However, knowledge conflicts arise when there is a mismatch between the actual temporal relations of events in the context and the prior knowledge or biases learned by the model. In this paper, we propose to detect knowledge-conflict examples in ev

  79. Yichen Yan, Xingjian He, Wenxuan Wan, Jing Liu

    Referring image segmentation aims to segment an object referred to by natural language expression from an image. However, this task is challenging due to the distinct data properties between text and image, and the randomness introduced by diverse objects and unrestricted language expression. Most of previous work focus on improving cross-modal feature fusio

  80. Simon Wegener, Kris K. Nikov, Jose Nunez-Yanez, Kerstin Eder

    This paper presents EnergyAnalyzer, a code-level static analysis tool for estimating the energy consumption of embedded software based on statically predictable hardware events. The tool utilises techniques usually used for worst-case execution time (WCET) analysis together with bespoke energy models developed for two predictable architectures - the ARM Cort

  81. M. A. Moreno-Frías, J. C. Rosales

    Let $S$ be a numerical semigroup. We will say that $h\in {\mathbb{N}} \backslash S$ is an {\it isolated gap }of $S$ if $\{h-1,h+1\}\subseteq S.$ A numerical semigroup without isolated gaps is called perfect numerical semigroup. Denote by ${\mathrm m}(S)$ the multiplicity of a numerical semigroup $S$. A covariety is a nonempty family ${\mathscr{C}}$ of numeri

  82. Sanoopkumar P. S., Stephen McWade, Arman Farhang

    Orthogonal time frequency space (OTFS) is a promising candidate waveform for the next generation wireless communication systems. OTFS places data in the delay-Doppler (DD) domain, which simplifies channel estimation in highmobility scenarios. However, due to the 2-D convolution effect of the time-varying channel in the DD domain, equalization is still a chal

  83. Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya

    Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these a

  84. Sagi Pendzel, Nir Lotan, Alon Zoizner, Einat Minkov

    The rise of social media has been argued to intensify uncivil and hostile online political discourse. Yet, to date, there is a lack of clarity on what incivility means in the political sphere. In this work, we utilize a multidimensional perspective of political incivility, developed in the fields of political science and communication, that differentiates be

  85. Yau-Shian Wang, Ta-Chung Chi, Ruohong Zhang, Yiming Yang

    We present PESCO, a novel contrastive learning framework that substantially improves the performance of zero-shot text classification. We formulate text classification as a neural text matching problem where each document is treated as a query, and the system learns the mapping from each query to the relevant class labels by (1) adding prompts to enhance lab

  86. Christoph Auer, Ahmed Nassar, Maksym Lysak, Michele Dolfi

    Transforming documents into machine-processable representations is a challenging task due to their complex structures and variability in formats. Recovering the layout structure and content from PDF files or scanned material has remained a key problem for decades. ICDAR has a long tradition in hosting competitions to benchmark the state-of-the-art and encour

  87. Simon Wiegrebe, Philipp Kopper, Raphael Sonabend, Bernd Bischl

    The influx of deep learning (DL) techniques into the field of survival analysis in recent years has led to substantial methodological progress; for instance, learning from unstructured or high-dimensional data such as images, text or omics data. In this work, we conduct a comprehensive systematic review of DL-based methods for time-to-event analysis, charact

  88. Jingwen Ma, Ding Jia, Li Zhang, Yi-jun Guan

    As a hypothetical topological defect in the geometry of spacetime, vortex strings play a crucial role in shaping the clusters of galaxies that exist today, and their distinct features can provide observable clues about the early universe's evolution. A key feature of vortex strings is that they can interact with Weyl fermionic modes and support topological c

  89. Omid Esrafilian, Rajeev Gangula, David Gesbert

    In this paper, we investigate the problem of UAV-aided user localization in wireless networks. Unlike the existing works, we do not assume perfect knowledge of the UAV location, hence we not only need to localize the users but also to track the UAV location. To do so, we utilize the time-of-arrival along with received signal strength radio measurements colle

  90. François-Xavier Coudert

    We reproduced the simulations described in Wang et al [J. Non-Cryst. Sol., 498 (2018) 294-304] and found we could not obtain the results reported. The root cause was identified to be incorrect atom masses in the original simulation files. As a consequence, the potential does not reproduce the experimental glass density -- and presumably, other structural pro

  91. Alessandro Ursi, Nicolò Parmiggiani, Mauro Messerotti, Alberto Pellizzoni

    We report the Astrorivelatore Gamma ad Immagini LEggero (AGILE) observations of solar flares, detected by the on board anticoincidence system in the 80-200 keV energy range, from 2007 May 1st to 2022 August 31st. In more than 15 yr, AGILE detected 5003 X-ray, minute-lasting transients, compatible with a solar origin. A cross-correlation of these transients w

  92. Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao

    Editing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023). However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer. Commonsense knowledge with multiple correct answers, e.g., an apple can be gree

  93. Jiayi Zhu, Xuebin Qin, Abdulmotaleb Elsaddik

    In this paper, we introduce Divide-and-Conquer into the salient object detection (SOD) task to enable the model to learn prior knowledge that is for predicting the saliency map. We design a novel network, Divide-and-Conquer Network (DC-Net) which uses two encoders to solve different subtasks that are conducive to predicting the final saliency map, here is to

  94. Valeria Giunta, Thomas Hillen, Mark A. Lewis, Jonathan R. Potts

    Nonlocal interactions are ubiquitous in nature and play a central role in many biological systems. In this paper, we perform a bifurcation analysis of a widely-applicable advection-diffusion model with nonlocal advection terms describing the species movements generated by inter-species interactions. We use linear analysis to assess the stability of the const

  95. Alexander F. Zakharov

    In May 2022 ICRANet organized the Workshop dedicated to the 80th anniversary of Professor Ruffini. This paper is based on the talk delivered at the meeting. Professor Ruffini was well known for Soviet scientific community not only due to his publications in leading journals but also due Russian translations of his books where he was an author or a contributo

  96. Waleed El Hanafy, Adel Awad

    It has been shown that the nonminimal coupling between geometry and matter can provide models for massive compact stars that are consistent with the conformal bound on the sound speed, $0\leqslant {c}_{s}^{2}\leqslant {c}^{2}/3$, where the core density approaches a few times the nuclear saturation density. We impose the conformal upper bound on the sound spe

  97. Shahar Lutati, Itamar Zimerman, Lior Wolf

    We present a new layer in which dynamic (i.e.,input-dependent) Infinite Impulse Response (IIR) filters of order two are used to process the input sequence prior to applying conventional attention. The input is split into chunks, and the coefficients of these filters are determined based on previous chunks to maintain causality. Despite their relatively low o

  98. Jue Liu

    To solve the problem of pose distortion in the forward propagation of pose features in existing methods, this pa-per proposes a Dual-Side Feature Fusion Network for pose transfer (DSFFNet). Firstly, a fixed-length pose code is extracted from the source mesh by a pose encoder and combined with the target vertices to form a mixed feature; Then, a Feature Fusio

  99. Jiongxiao Wang, Zichen Liu, Keun Hee Park, Zhuojun Jiang

    With the emergence of more powerful large language models (LLMs), such as ChatGPT and GPT-4, in-context learning (ICL) has gained significant prominence in leveraging these models for specific tasks by utilizing data-label pairs as precondition prompts. While incorporating demonstrations can greatly enhance the performance of LLMs across various tasks, it ma

  100. Qi Gou, Zehua Xia, Wenzhe Du

    This paper proposes a framework to address the issue of data scarcity in Document-Grounded Dialogue Systems(DGDS). Our model leverages high-resource languages to enhance the capability of dialogue generation in low-resource languages. Specifically, We present a novel pipeline CLEM (Cross-Lingual Enhanced Model) including adversarial training retrieval (Retri