May 2023 arXiv papers — page 54
Showing 5,301–5,400 of 19,695 papers
Generating Faithful Synthetic Data with Large Language Models: A Case Study in Computational Social Science
cs.CLVeniamin Veselovsky, Manoel Horta Ribeiro, Akhil Arora, Martin Josifoski
Large Language Models (LLMs) have democratized synthetic data generation, which in turn has the potential to simplify and broaden a wide gamut of NLP tasks. Here, we tackle a pervasive problem in synthetic data generation: its generative distribution often differs from the distribution of real-world data researchers care about (in other words, it is unfaithf
Yotam Perlitz, Ariel Gera, Michal Shmueli-Scheuer, Dafna Sheinwald
The field of Natural Language Generation (NLG) suffers from a severe shortage of labeled data due to the extremely expensive and time-consuming process involved in manual annotation. A natural approach for coping with this problem is active learning (AL), a well-known machine learning technique for improving annotation efficiency by selectively choosing the
E. B. Balbutsev, I. V. Molodtsova
The solution of time dependent Hartree-Fock-Bogoliubov equations by the Wigner function moments method predicts four low-lying $1^+$ states. Three of them are known as various scissors modes. Fourth state is disposed below all scissors modes and has the electrical nature. It is found that it represents one of three branches of $2^+$ state which can exist in
Liying Cheng, Xingxuan Li, Lidong Bing
As large language models (LLMs) have demonstrated their powerful capabilities in plenty of domains and tasks, including context understanding, code generation, language generation, data storytelling, etc., many data analysts may raise concerns if their jobs will be replaced by artificial intelligence (AI). This controversial topic has drawn great attention i
Zhengxue Ren, Serdar Elhatisari, Timo A. Lähde, Dean Lee
We present a systematic ab initio study of clustering in hot dilute nuclear matter using nuclear lattice effective field theory with an SU(4)-symmetric interaction. We introduce a method called light-cluster distillation to determine the abundances of dimers, trimers, and alpha clusters as a function of density and temperature. Our lattice results are compar
Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical Insights
cs.IRJunchen Fu, Fajie Yuan, Yu Song, Zheng Yuan
Adapters, a plug-in neural network module with some tunable parameters, have emerged as a parameter-efficient transfer learning technique for adapting pre-trained models to downstream tasks, especially for natural language processing (NLP) and computer vision (CV) fields. Meanwhile, learning recommendation models directly from raw item modality features -- e
Wei-Lin Chen, Cheng-Kuang Wu, Yun-Nung Chen, Hsin-Hsi Chen
Large language models (LLMs) have exhibited striking in-context learning (ICL) ability to adapt to target tasks with a few input-output demonstrations. For better ICL, different methods are proposed to select representative demonstrations from existing training corpora. However, such settings are not aligned with real-world practices, as end-users usually qu
Holder continuity of weak solutions to evolution equations with distributed order fractional time derivative
math.APAdam Kubica, Katarzyna Ryszewska, Rico Zacher
We study the regularity of weak solutions to evolution equations with distributed order fractional time derivative. We prove a weak Harnack inequality for nonnegative weak supersolutions and H\"older continuity of weak solutions to this problem. Our results substantially generalise analogous known results for the problem with single order fractional time der
Zekun Wang, Jingchang Chen, Wangchunshu Zhou, Haichao Zhu
Despite achieving remarkable performance on various vision-language tasks, Transformer-based Vision-Language Models (VLMs) suffer from redundancy in inputs and parameters, significantly hampering their efficiency in real-world applications. Moreover, the degree of redundancy in token representations and model parameters, such as attention heads, varies signi
How to Distill your BERT: An Empirical Study on the Impact of Weight Initialisation and Distillation Objectives
cs.CLXinpeng Wang, Leonie Weissweiler, Hinrich Schütze, Barbara Plank
Recently, various intermediate layer distillation (ILD) objectives have been shown to improve compression of BERT models via Knowledge Distillation (KD). However, a comprehensive evaluation of the objectives in both task-specific and task-agnostic settings is lacking. To the best of our knowledge, this is the first work comprehensively evaluating distillatio
Jonas Beyrer, Fanny Kassel
For any integers $p\geq 2$ and $q\geq 1$, let $\mathbb{H}^{p,q}$ be the pseudo-Riemannian hyperbolic space of signature $(p,q)$. We prove that if $\Gamma$ is the fundamental group of a closed aspherical $p$-manifold, then the set of representations of $\Gamma$ to $\mathrm{PO}(p,q+1)$ which are convex cocompact in $\mathbb{H}^{p,q}$ is a union of connected co
Javier Molina-Vilaplana
We use a post-Gaussian variational approach to non-perturbatively study a general class of interacting bosonic quantum field theories with generalized dipole symmetries and fractonic behaviour. We find that while a Gaussian approach allows to carry out a consistent renormalization group (RG) flow analysis of these theories, this only grasps the interaction t
Shilv Cai, Liqun Chen, Sheng Zhong, Luxin Yan
Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement is usually adopted. Besides, lossy image compression is vital for meeting the requirements of storage and transmission i
Jamie Le Signe, Thomas McDermott, Eros Mariani
The coupling of superconducting systems to mechanical resonators is an emerging field, with wide reaching implications including high precision sensing and metrology. Experimental signatures of this coupling have so far been small, seldom and often reliant on high frequency AC electronics. To overcome this limitation, in this work we consider a mechanical re
Heming Xia, Qingxiu Dong, Lei Li, Jingjing Xu
Recently, Large Language Models (LLMs) have been serving as general-purpose interfaces, posing a significant demand for comprehensive visual knowledge. However, it remains unclear how well current LLMs and their visually augmented counterparts (VaLMs) can master visual commonsense knowledge. To investigate this, we propose ImageNetVC, a human-annotated datas
Veit David Wild, Sahra Ghalebikesabi, Dino Sejdinovic, Jeremias Knoblauch
We establish the first mathematically rigorous link between Bayesian, variational Bayesian, and ensemble methods. A key step towards this it to reformulate the non-convex optimisation problem typically encountered in deep learning as a convex optimisation in the space of probability measures. On a technical level, our contribution amounts to studying general
Rodrigo Valerio, Joao Bordalo, Michal Yarom, Yonatan Bitton
Text to image generation methods (T2I) are widely popular in generating art and other creative artifacts. While visual hallucinations can be a positive factor in scenarios where creativity is appreciated, such artifacts are poorly suited for cases where the generated image needs to be grounded in complex natural language without explicit visual elements. In
Tianyu Yang, Thy Thy Tran, Iryna Gurevych
Current variational dialog models have employed pre-trained language models (PLMs) to parameterize the likelihood and posterior distributions. However, the Gaussian assumption made on the prior distribution is incompatible with these distributions, thus restricting the diversity of generated responses. These models also suffer from posterior collapse, i.e.,
Biao Zhao, Weiqiang Jin, Javier Del Ser, Guang Yang
In the era of sustainable smart agriculture, a massive amount of agricultural news text is being posted on the Internet, in which massive agricultural knowledge has been accumulated. In this context, it is urgent to explore effective text classification techniques for users to access the required agricultural knowledge with high efficiency. Mainstream deep l
Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen
Recently, growing interest has been aroused in extending the multimodal capability of large language models (LLMs), e.g., vision-language (VL) learning, which is regarded as the next milestone of artificial general intelligence. However, existing solutions are prohibitively expensive, which not only need to optimize excessive parameters, but also require ano
Annie Gray, Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley
In this paper we offer a new perspective on the well established agglomerative clustering algorithm, focusing on recovery of hierarchical structure. We recommend a simple variant of the standard algorithm, in which clusters are merged by maximum average dot product and not, for example, by minimum distance or within-cluster variance. We demonstrate that the
Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang
Embodied AI is a crucial frontier in robotics, capable of planning and executing action sequences for robots to accomplish long-horizon tasks in physical environments. In this work, we introduce EmbodiedGPT, an end-to-end multi-modal foundation model for embodied AI, empowering embodied agents with multi-modal understanding and execution capabilities. To ach
Asahi Ushio, Yi Zhou, Jose Camacho-Collados
Multilingual language model (LM) have become a powerful tool in NLP especially for non-English languages. Nevertheless, model parameters of multilingual LMs remain large due to the larger embedding matrix of the vocabulary covering tokens in different languages. On the contrary, monolingual LMs can be trained in a target language with the language-specific v
Anurag Dey, Probal Chaudhuri
Several well known estimators of finite population mean and its functions are investigated under some standard sampling designs. Such functions of mean include the variance, the correlation coefficient and the regression coefficient in the population as special cases. We compare the performance of these estimators under different sampling designs based on th
Jian Leng, Fan Yang, Xiang-Bin Wang
The decomposition for controlled-$ZX$ gate in [Phys. Rev. A, 87, 062318 (2013)] has a shallow circuit depth $8n-20$ with no ancilla. Here we modify this decomposition to decompose $n$-qubit Toffoli gate with only $2n-3$ additional single-qubit gates. The circuit depth is unchanged and no ancilla is needed. We explicitly show that the circuit after decomposit
Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems
cs.LGMarek Kadlčík, Michal Štefánik, Ondřej Sotolář, Vlastimil Martinek
Despite outstanding performance in many tasks, language models are notoriously inclined to make factual errors in tasks requiring arithmetic computation. We address this deficiency by creating Calc-X, a collection of datasets that demonstrates the appropriate use of a calculator in reasoning chains. Calc-X is suitable for teaching language models to offload
Kostis Gourgoulias, Najah Ghalyan, Maxime Labonne, Yash Satsangi
This paper introduces an unsupervised method to estimate the class separability of text datasets from a topological point of view. Using persistent homology, we demonstrate how tracking the evolution of embedding manifolds during training can inform about class separability. More specifically, we show how this technique can be applied to detect when the trai
Daniel Reich, Felix Putze, Tanja Schultz
Metrics for Visual Grounding (VG) in Visual Question Answering (VQA) systems primarily aim to measure a system's reliance on relevant parts of the image when inferring an answer to the given question. Lack of VG has been a common problem among state-of-the-art VQA systems and can manifest in over-reliance on irrelevant image parts or a disregard for the visu
Unlocking Temporal Question Answering for Large Language Models with Tailor-Made Reasoning Logic
cs.CLXingxuan Li, Liying Cheng, Qingyu Tan, Hwee Tou Ng
The temporal aspect is a significant dimension of our reality. We notice the challenge that large language models (LLMs) face when engaging in temporal reasoning. Our preliminary experiments show that methods involving the generation of intermediate reasoning steps, such as chain-of-thought and program-aided language models, do not consistently boost the per
A. D. Alhaidari, A. Laradji
An algebraic system is introduced, which is very useful for doing scattering calculations in quantum field theory. It is the set of all real numbers greater than or equal to -m^2 with parity designation and a special rule for addition and subtraction, where m is the rest mass of the scattered particle.
Linxuan Pan, Shenghui Song
With multiple iterations of updates, local statistical gradient descent (L-SGD) has been proven to be very effective in distributed machine learning schemes such as federated learning. In fact, many innovative works have shown that L-SGD with independent and identically distributed (IID) data can even outperform SGD. As a result, extensive efforts have been
Jitendra Joshi, Mir Alimuddin, T S Mahesh, Manik Banik
The phenomenon of quantum entanglement underlies several important protocols that enable emerging quantum technologies. Entangled states, however, are extremely delicate and often get perturbed by tiny fluctuations in their external environment. Certification of entanglement is therefore immensely crucial for the successful implementation of protocols involv
Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji
Instruction tuning has shown great promise in improving the performance of large language models. However, research on multilingual instruction tuning has been limited due to the scarcity of high-quality instruction-response datasets across different languages. To bridge this gap, we present Bactrian-X, a comprehensive multilingual parallel dataset of 3.4 mi
Injecting Knowledge into Biomedical Pre-trained Models via Polymorphism and Synonymous Substitution
cs.CLHongbo Zhang, Xiang Wan, Benyou Wang
Pre-trained language models (PLMs) were considered to be able to store relational knowledge present in the training data. However, some relational knowledge seems to be discarded unsafely in PLMs due to \textbf{report bias}: low-frequency relational knowledge might be underexpressed compared to high-frequency one in PLMs. This gives us a hint that relational
Elena Losero, Valentin Goblot, Yuchun Zhu, Hossein Babashah
Single-crystal diamond substrates presenting a high concentration of negatively charged nitrogen-vacancy centers (NV-) are on high demand for the development of optically pumped solid-state sensors such as magnetometers, thermometers or electrometers. While nitrogen impurities can be easily incorporated during crystal growth, the creation of vacancies requir
Are Chatbots Ready for Privacy-Sensitive Applications? An Investigation into Input Regurgitation and Prompt-Induced Sanitization
cs.CLAman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu
LLM-powered chatbots are becoming widely adopted in applications such as healthcare, personal assistants, industry hiring decisions, etc. In many of these cases, chatbots are fed sensitive, personal information in their prompts, as samples for in-context learning, retrieved records from a database, or as part of the conversation. The information provided in
Quaternion-based non-singular terminal sliding mode control for a satellite-mounted space manipulator
eess.SYJacopo Giordano, Angelo Cenedese
In this paper, a robust control solution for a satellite equipped with a robotic manipulator is presented. First, the dynamic model of the system is derived based on quaternions to describe the evolution of the attitude of the base satellite. Then, a non-singular terminal sliding mode controller that employs quaternions for attitude control, is proposed for
A Human-in-the-Loop Approach for Information Extraction from Privacy Policies under Data Scarcity
cs.CYMichael Gebauer, Faraz Maschhur, Nicola Leschke, Elias Grünewald
Machine-readable representations of privacy policies are door openers for a broad variety of novel privacy-enhancing and, in particular, transparency-enhancing technologies (TETs). In order to generate such representations, transparency information needs to be extracted from written privacy policies. However, respective manual annotation and extraction proce
Wenxuan Zhang, Yue Deng, Bing Liu, Sinno Jialin Pan
Sentiment analysis (SA) has been a long-standing research area in natural language processing. It can offer rich insights into human sentiments and opinions and has thus seen considerable interest from both academia and industry. With the advent of large language models (LLMs) such as ChatGPT, there is a great potential for their employment on SA problems. H
Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng
Generated texts from large language models (LLMs) are remarkably close to high-quality human-authored text, raising concerns about their potential misuse in spreading false information and academic misconduct. Consequently, there is an urgent need for a highly practical detection tool capable of accurately identifying the source of a given text. However, exi
Feasible Action-Space Reduction as a Metric of Causal Responsibility in Multi-Agent Spatial Interactions
cs.MAAshwin George, Luciano Cavalcante Siebert, David Abbink, Arkady Zgonnikov
Modelling causal responsibility in multi-agent spatial interactions is crucial for safety and efficiency of interactions of humans with autonomous agents. However, current formal metrics and models of responsibility either lack grounding in ethical and philosophical concepts of responsibility, or cannot be applied to spatial interactions. In this work we pro
Asahi Ushio, Jose Camacho Collados, Steven Schockaert
Relations such as "is influenced by", "is known for" or "is a competitor of" are inherently graded: we can rank entity pairs based on how well they satisfy these relations, but it is hard to draw a line between those pairs that satisfy them and those that do not. Such graded relations play a central role in many applications, yet they are typically not cover
Aleksandar Stanić, Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber
Current state-of-the-art object-centric models use slots and attention-based routing for binding. However, this class of models has several conceptual limitations: the number of slots is hardwired; all slots have equal capacity; training has high computational cost; there are no object-level relational factors within slots. Synchrony-based models in principl
Christian Berger, Lívio Rodrigues, Hans P. Reiser, Vinicius Cogo
Blockchain technology sparked renewed interest in planetary-scale Byzantine fault-tolerant (BFT) state machine replication (SMR). While recent works predominantly focused on improving the scalability and throughput of these protocols, few of them addressed latency. We present Mercury, a novel transformation to autonomously optimize the latency of quorum-base
Jingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu
Chain-of-Thought (CoT) prompting enables large language models to solve complex reasoning problems by generating intermediate steps. However, confined by its inherent single-pass and sequential generation process, CoT heavily relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers. In contrast, humans adopt r
Saba Ahmadi, Aishwarya Agrawal
Recently, reference-free metrics such as CLIPScore (Hessel et al., 2021), UMIC (Lee et al., 2021), and PAC-S (Sarto et al., 2023) have been proposed for automatic reference-free evaluation of image captions. Our focus lies in evaluating the robustness of these metrics in scenarios that require distinguishing between two captions with high lexical overlap but
3GPP-Like GBSM THz Channel Characterization, Modeling, and Simulation Based on Experimental Observations
eess.SPZhaowei Chang, Jianhua Zhang, Pan Tang, Lei Tian
Terahertz (THz) communication is envisioned as one of the possible technologies for the sixth-generation (6G) communication system due to its rich spectrum. To evaluate the performance of THz communication, it is essential to propose THz channel models within the common framework of the geometry-based stochastic model (GBSM) in the 3rd Generation Partnership
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
cs.CLChongjian Yue, Xinrun Xu, Xiaojun Ma, Lun Du
Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains underexplored. In this research, we specialize in harnessing the potential of LLMs to comprehend critical information from financial r
Shaurya Rohatgi, Yanxia Qin, Benjamin Aw, Niranjana Unnithan
We present ACL OCL, a scholarly corpus derived from the ACL Anthology to assist Open scientific research in the Computational Linguistics domain. Integrating and enhancing the previous versions of the ACL Anthology, the ACL OCL contributes metadata, PDF files, citation graphs and additional structured full texts with sections, figures, and links to a large k
Some Notes on Quantitative Generalized CLTs with Self-Decomposable Limiting Laws by Spectral Methods
math.PRBenjamin Arras
In these notes, we obtain new stability estimates for centered non-degenerate selfdecomposable probability measures on $\mathbb{R}^d$ with finite second moment and for non-degenerate symmetric $\alpha$-stable probability measures on $\mathbb{R}^d$ with $\alpha \in [1,2)$. These new results are refinements of the corresponding ones available in the literature
Dongjie Yang, Ruifeng Yuan, Yuantao Fan, Yifei Yang
Large Language Models (LLMs) have attained the impressive capability to resolve a wide range of NLP tasks by fine-tuning high-quality instruction data. However, collecting human-written data of high quality, especially multi-turn dialogues, is expensive and unattainable for most people. Though previous studies have used powerful LLMs to generate the dialogue
Sweta Agrawal, Marine Carpuat
Text simplification (TS) systems rewrite text to make it more readable while preserving its content. However, what makes a text easy to read depends on the intended readers. Recent work has shown that pre-trained language models can simplify text using a wealth of techniques to control output simplicity, ranging from specifying only the desired reading grade
Andrew T. Hyman
The theory of point-particles in classical electrodynamics has a well-known problem of infinite self-energy, and the same is true of quantum electrodynamics. Instead of concluding that there is no such thing as a true point-particle, it is shown here how to remove the infinities by supposing that the electromagnetic field tensor has a symmetric part. This do
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong
Large language models (LLMs) have shown remarkable reasoning capabilities, especially when prompted to generate intermediate reasoning steps (e.g., Chain-of-Thought, CoT). However, LLMs can still struggle with problems that are easy for humans, such as generating action plans for executing tasks in a given environment, or performing complex math, logical, an
Taelin Karidi, Leshem Choshen, Gal Patel, Omri Abend
We propose a novel methodology (namely, MuLER) that transforms any reference-based evaluation metric for text generation, such as machine translation (MT) into a fine-grained analysis tool. Given a system and a metric, MuLER quantifies how much the chosen metric penalizes specific error types (e.g., errors in translating names of locations). MuLER thus enabl
Photo-activated dynamic isomerization induced large density changes in liquid crystal polymers: A molecular dynamics study
cond-mat.softAkhil Reddy Peeketi, Edwin Joseph, Narasimhan Swaminathan, Ratna Kumar Annabattula
We use molecular dynamics simulations to unravel the physics underpinning the light-induced density changes caused by the dynamic trans-cis-trans isomerization cycles of azo-mesogens embedded in a liquid crystal polymer network, an intriguing experimental observation reported in the literature. We employ two approaches, cyclic and probabilistic switching of
Hua Cai, Xuli Shen, Qing Xu, Weilin Shen
In empathetic conversations, individuals express their empathy towards others. Previous work has mainly focused on generating empathetic responses by utilizing the speaker's emotion. Besides, external commonsense knowledge has been applied to enhance the system's understandings of the speaker's situation. However, given an event, commonsense knowledge base c
El Moatez Billah Nagoudi, AbdelRahim Elmadany, Ahmed El-Shangiti, Muhammad Abdul-Mageed
We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range of 13 different NLG tasks, including dialogue generation, question answering, machine translation, summarization, among
Some Identities of Ramanujan's q-Continued Fractions of Order Fourteen and Twenty-Eight, and Vanishing Coefficients
math.NTShraddha Rajkhowa, Nipen Saikia
We deduce $q$-continued fractions $S_{1}(q)$, $S_{2}(q)$ and $S_{3}(q)$ of order fourteen, and continued fractions $V_{1}(q)$, $V_{2}(q)$ and $V_{3}(q)$ of order twenty-eight from a general continued fraction identity of Ramanujan. We establish some theta-function identities for the continued fractions and derive some colour partition identities as applicati
Investigating Table-to-Text Generation Capabilities of LLMs in Real-World Information Seeking Scenarios
cs.CLYilun Zhao, Haowei Zhang, Shengyun Si, Linyong Nan
Tabular data is prevalent across various industries, necessitating significant time and effort for users to understand and manipulate for their information-seeking purposes. The advancements in large language models (LLMs) have shown enormous potential to improve user efficiency. However, the adoption of LLMs in real-world applications for table information
Gorana Gojić, Vladimir Vincan, Ognjen Kundačina, Dragiša Mišković
Non-adversarial robustness, also known as natural robustness, is a property of deep learning models that enables them to maintain performance even when faced with distribution shifts caused by natural variations in data. However, achieving this property is challenging because it is difficult to predict in advance the types of distribution shifts that may occ
Manuel Glöckler, Michael Deistler, Jakob H. Macke
Bayesian inference usually requires running potentially costly inference procedures separately for every new observation. In contrast, the idea of amortized Bayesian inference is to initially invest computational cost in training an inference network on simulated data, which can subsequently be used to rapidly perform inference (i.e., to return estimates of
Local gate control of Mott metal-insulator transition in a 2D metal-organic framework
cond-mat.str-elBenjamin Lowe, Bernard Field, Jack Hellerstedt, Julian Ceddia
Electron-electron interactions in materials lead to exotic many-body quantum phenomena including Mott metal-insulator transitions (MITs), magnetism, quantum spin liquids, and superconductivity. These phases depend on electronic band occupation and can be controlled via the chemical potential. Flat bands in two-dimensional (2D) and layered materials with a ka
Ahmed Abdelali, Hamdy Mubarak, Shammur Absar Chowdhury, Maram Hasanain
Recent advancements in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. Despite this progress, these models lack specific benchmarking against state-of-the-art (SOTA) models tailored to particular languages and tasks. LAraBench addresses this gap for Arabic Natural Language Processing (NLP) and Speech
Tanay Dixit, Fei Wang, Muhao Chen
Improving factual consistency of abstractive summarization has been a widely studied topic. However, most of the prior works on training factuality-aware models have ignored the negative effect it has on summary quality. We propose EFACTSUM (i.e., Effective Factual Summarization), a candidate summary generation and ranking technique to improve summary factua
An Atypical Plateau-like Extreme-ultraviolet Late-phase Solar Flare Driven by the Non-radial Eruption of a Magnetic Flux Rope
astro-ph.SRYuehong Chen, Yu Dai, Mingde Ding
Recent observations in extreme-ultraviolet (EUV) wavelengths reveal an EUV late phase in some solar flares, which is characterized by a second peak in the warm coronal emissions (about 3 MK) occurring several tens of minutes to a few hours after the corresponding main flare peak. We aim to clarify the physical origin of an atypical plateau-like EUV late phas
Assessment of the Reliablity of a Model's Decision by Generalizing Attribution to the Wavelet Domain
cs.CVGabriel Kasmi, Laurent Dubus, Yves-Marie Saint Drenan, Philippe Blanc
Neural networks have shown remarkable performance in computer vision, but their deployment in numerous scientific and technical fields is challenging due to their black-box nature. Scientists and practitioners need to evaluate the reliability of a decision, i.e., to know simultaneously if a model relies on the relevant features and whether these features are
Nathanael Bosch, Philipp Hennig, Filip Tronarp
Probabilistic solvers provide a flexible and efficient framework for simulation, uncertainty quantification, and inference in dynamical systems. However, like standard solvers, they suffer performance penalties for certain stiff systems, where small steps are required not for reasons of numerical accuracy but for the sake of stability. This issue is greatly
Florian Heidecker, Ahmad El-Khateeb, Bernhard Sick
The examination of uncertainty in the predictions of machine learning (ML) models is receiving increasing attention. One uncertainty modeling technique used for this purpose is Monte-Carlo (MC)-Dropout, where repeated predictions are generated for a single input. Therefore, clustering is required to describe the resulting uncertainty, but only through effici
Kaimin Wang, Haoran Liu, Peng Li, Mingzhe Liu
The publicly accessible dataset includes neutron and gamma-ray pulse signals for conducting pulse shape discrimination experiments. Several traditional and recently proposed pulse shape discrimination algorithms are utilized to evaluate the performance of pulse shape discrimination under raw pulse signals and noise-enhanced datasets. These algorithms compris
Md Tawkat Islam Khondaker, Abdul Waheed, El Moatez Billah Nagoudi, Muhammad Abdul-Mageed
ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT's capabilities on Arabic
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
cs.CLKatherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma
A trustworthy real-world prediction system should produce well-calibrated confidence scores; that is, its confidence in an answer should be indicative of the likelihood that the answer is correct, enabling deferral to an expert in cases of low-confidence predictions. Recent studies have shown that unsupervised pre-training produces large language models (LMs
David Kappel, Khaleelulla Khan Nazeer, Cabrel Teguemne Fokam, Christian Mayr
The ubiquitous backpropagation algorithm requires sequential updates through the network introducing a locking problem. In addition, back-propagation relies on the transpose of forward weight matrices to compute updates, introducing a weight transport problem across the network. Locking and weight transport are problems because they prevent efficient paralle
Jiazheng Li, Runcong Zhao, Yongxin Yang, Yulan He
The remarkable performance of pre-trained large language models has revolutionised various natural language processing applications. Due to huge parametersizes and extensive running costs, companies or organisations tend to transfer the models to the target task by zero-shot prompting techniques. However, the prohibitive costs of tokens and time have hindere
A Distributed Automatic Domain-Specific Multi-Word Term Recognition Architecture using Spark Ecosystem
cs.CLCiprian-Octavian Truică, Neculai-Ovidiu Istrate, Elena-Simona Apostol
Automatic Term Recognition is used to extract domain-specific terms that belong to a given domain. In order to be accurate, these corpus and language-dependent methods require large volumes of textual data that need to be processed to extract candidate terms that are afterward scored according to a given metric. To improve text preprocessing and candidate te
Nicholas G. Polson, Vadim Sokolov
Bayesian Generative AI (BayesGen-AI) methods are developed and applied to Bayesian computation. BayesGen-AI reconstructs the posterior distribution by directly modeling the parameter of interest as a mapping (a.k.a. deep learner) from a large simulated dataset. This provides a generator that we can evaluate at the observed data and provide draws from the pos
Complete X-ray census of Mdwarfs in the solar Neighborhood I. GJ 745 AB: Coronal-hole Stars in the 10 pc Sample
astro-ph.SRM. Caramazza, B. Stelzer, E. Magaudda, St. Raetz
We have embarked in a systematic study of the X-ray emission in a volume-limited sample of M dwarf stars, in order to explore the full range of activity levels present in their coronae and, thus, to understand the conditions in their outer atmospheres and their possible impact on the circumstellar environment. We identify in a recent catalog of the Gaia obje
Getting Sick After Seeing a Doctor? Diagnosing and Mitigating Knowledge Conflicts in Event Temporal Reasoning
cs.CLTianqing Fang, Zhaowei Wang, Wenxuan Zhou, Hongming Zhang
Event temporal reasoning aims at identifying the temporal relations between two or more events from narratives. However, knowledge conflicts arise when there is a mismatch between the actual temporal relations of events in the context and the prior knowledge or biases learned by the model. In this paper, we propose to detect knowledge-conflict examples in ev
Yichen Yan, Xingjian He, Wenxuan Wan, Jing Liu
Referring image segmentation aims to segment an object referred to by natural language expression from an image. However, this task is challenging due to the distinct data properties between text and image, and the randomness introduced by diverse objects and unrestricted language expression. Most of previous work focus on improving cross-modal feature fusio
EnergyAnalyzer: Using Static WCET Analysis Techniques to Estimate the Energy Consumption of Embedded Applications
cs.SESimon Wegener, Kris K. Nikov, Jose Nunez-Yanez, Kerstin Eder
This paper presents EnergyAnalyzer, a code-level static analysis tool for estimating the energy consumption of embedded software based on statically predictable hardware events. The tool utilises techniques usually used for worst-case execution time (WCET) analysis together with bespoke energy models developed for two predictable architectures - the ARM Cort
M. A. Moreno-Frías, J. C. Rosales
Let $S$ be a numerical semigroup. We will say that $h\in {\mathbb{N}} \backslash S$ is an {\it isolated gap }of $S$ if $\{h-1,h+1\}\subseteq S.$ A numerical semigroup without isolated gaps is called perfect numerical semigroup. Denote by ${\mathrm m}(S)$ the multiplicity of a numerical semigroup $S$. A covariety is a nonempty family ${\mathscr{C}}$ of numeri
Sanoopkumar P. S., Stephen McWade, Arman Farhang
Orthogonal time frequency space (OTFS) is a promising candidate waveform for the next generation wireless communication systems. OTFS places data in the delay-Doppler (DD) domain, which simplifies channel estimation in highmobility scenarios. However, due to the 2-D convolution effect of the time-varying channel in the DD domain, equalization is still a chal
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya
Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited studies have been conducted to formalize and analyze these a
Sagi Pendzel, Nir Lotan, Alon Zoizner, Einat Minkov
The rise of social media has been argued to intensify uncivil and hostile online political discourse. Yet, to date, there is a lack of clarity on what incivility means in the political sphere. In this work, we utilize a multidimensional perspective of political incivility, developed in the fields of political science and communication, that differentiates be
Yau-Shian Wang, Ta-Chung Chi, Ruohong Zhang, Yiming Yang
We present PESCO, a novel contrastive learning framework that substantially improves the performance of zero-shot text classification. We formulate text classification as a neural text matching problem where each document is treated as a query, and the system learns the mapping from each query to the relevant class labels by (1) adding prompts to enhance lab
Christoph Auer, Ahmed Nassar, Maksym Lysak, Michele Dolfi
Transforming documents into machine-processable representations is a challenging task due to their complex structures and variability in formats. Recovering the layout structure and content from PDF files or scanned material has remained a key problem for decades. ICDAR has a long tradition in hosting competitions to benchmark the state-of-the-art and encour
Simon Wiegrebe, Philipp Kopper, Raphael Sonabend, Bernd Bischl
The influx of deep learning (DL) techniques into the field of survival analysis in recent years has led to substantial methodological progress; for instance, learning from unstructured or high-dimensional data such as images, text or omics data. In this work, we conduct a comprehensive systematic review of DL-based methods for time-to-event analysis, charact
Jingwen Ma, Ding Jia, Li Zhang, Yi-jun Guan
As a hypothetical topological defect in the geometry of spacetime, vortex strings play a crucial role in shaping the clusters of galaxies that exist today, and their distinct features can provide observable clues about the early universe's evolution. A key feature of vortex strings is that they can interact with Weyl fermionic modes and support topological c
Omid Esrafilian, Rajeev Gangula, David Gesbert
In this paper, we investigate the problem of UAV-aided user localization in wireless networks. Unlike the existing works, we do not assume perfect knowledge of the UAV location, hence we not only need to localize the users but also to track the UAV location. To do so, we utilize the time-of-arrival along with received signal strength radio measurements colle
Failure to Reproduce the Results of "A new transferable interatomic potential for molecular dynamics simulations of borosilicate glasses''
physics.comp-phFrançois-Xavier Coudert
We reproduced the simulations described in Wang et al [J. Non-Cryst. Sol., 498 (2018) 294-304] and found we could not obtain the results reported. The root cause was identified to be incorrect atom masses in the original simulation files. As a consequence, the potential does not reproduce the experimental glass density -- and presumably, other structural pro
Alessandro Ursi, Nicolò Parmiggiani, Mauro Messerotti, Alberto Pellizzoni
We report the Astrorivelatore Gamma ad Immagini LEggero (AGILE) observations of solar flares, detected by the on board anticoincidence system in the 80-200 keV energy range, from 2007 May 1st to 2022 August 31st. In more than 15 yr, AGILE detected 5003 X-ray, minute-lasting transients, compatible with a solar origin. A cross-correlation of these transients w
Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao
Editing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023). However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer. Commonsense knowledge with multiple correct answers, e.g., an apple can be gree
Jiayi Zhu, Xuebin Qin, Abdulmotaleb Elsaddik
In this paper, we introduce Divide-and-Conquer into the salient object detection (SOD) task to enable the model to learn prior knowledge that is for predicting the saliency map. We design a novel network, Divide-and-Conquer Network (DC-Net) which uses two encoders to solve different subtasks that are conducive to predicting the final saliency map, here is to
Valeria Giunta, Thomas Hillen, Mark A. Lewis, Jonathan R. Potts
Nonlocal interactions are ubiquitous in nature and play a central role in many biological systems. In this paper, we perform a bifurcation analysis of a widely-applicable advection-diffusion model with nonlocal advection terms describing the species movements generated by inter-species interactions. We use linear analysis to assess the stability of the const
Alexander F. Zakharov
In May 2022 ICRANet organized the Workshop dedicated to the 80th anniversary of Professor Ruffini. This paper is based on the talk delivered at the meeting. Professor Ruffini was well known for Soviet scientific community not only due to his publications in leading journals but also due Russian translations of his books where he was an author or a contributo
Implications of the Conformal Constraint on Sound Speed on the Radius of PSR J0952-0607 within Rastall Gravity
astro-ph.HEWaleed El Hanafy, Adel Awad
It has been shown that the nonminimal coupling between geometry and matter can provide models for massive compact stars that are consistent with the conformal bound on the sound speed, $0\leqslant {c}_{s}^{2}\leqslant {c}^{2}/3$, where the core density approaches a few times the nuclear saturation density. We impose the conformal upper bound on the sound spe
Shahar Lutati, Itamar Zimerman, Lior Wolf
We present a new layer in which dynamic (i.e.,input-dependent) Infinite Impulse Response (IIR) filters of order two are used to process the input sequence prior to applying conventional attention. The input is split into chunks, and the coefficients of these filters are determined based on previous chunks to maintain causality. Despite their relatively low o
Jue Liu
To solve the problem of pose distortion in the forward propagation of pose features in existing methods, this pa-per proposes a Dual-Side Feature Fusion Network for pose transfer (DSFFNet). Firstly, a fixed-length pose code is extracted from the source mesh by a pose encoder and combined with the target vertices to form a mixed feature; Then, a Feature Fusio
Jiongxiao Wang, Zichen Liu, Keun Hee Park, Zhuojun Jiang
With the emergence of more powerful large language models (LLMs), such as ChatGPT and GPT-4, in-context learning (ICL) has gained significant prominence in leveraging these models for specific tasks by utilizing data-label pairs as precondition prompts. While incorporating demonstrations can greatly enhance the performance of LLMs across various tasks, it ma
Qi Gou, Zehua Xia, Wenzhe Du
This paper proposes a framework to address the issue of data scarcity in Document-Grounded Dialogue Systems(DGDS). Our model leverages high-resource languages to enhance the capability of dialogue generation in low-resource languages. Specifically, We present a novel pipeline CLEM (Cross-Lingual Enhanced Model) including adversarial training retrieval (Retri