May 2023 arXiv papers — page 53
Showing 5,201–5,300 of 19,695 papers
Huikang Liu, Xiao Li, Anthony Man-Cho So
This work presents ReSync, a Riemannian subgradient-based algorithm for solving the robust rotation synchronization problem, which arises in various engineering applications. ReSync solves a least-unsquared minimization formulation over the rotation group, which is nonsmooth and nonconvex, and aims at recovering the underlying rotations directly. We provide
Field-free all-optical switching and electrical read-out of Tb/Co-based magnetic tunnel junctions
physics.app-phD. Salomoni, Y. Peng, L. Farcis, S. Auffret
Switching of magnetic tunnel junction using femto-second laser enables a possible path for THz frequency memory operation, which means writing speeds 2 orders of magnitude faster than alternative electrical approaches based on spin transfer or spin orbit torque. In this work we demonstrate successful field-free 50fs single laser pulse driven magnetization re
Jinjin Gu, Xianzheng Ma, Xiangtao Kong, Yu Qiao
Deep deraining networks consistently encounter substantial generalization issues when deployed in real-world applications, although they are successful in laboratory benchmarks. A prevailing perspective in deep learning encourages using highly complex data for training, with the expectation that richer image background content will facilitate overcoming the
Katie Ansaldi, Gabriel Cowley, Eric Green, Kihyun Kim
An exact r-coloring of a set $S$ is a surjective function $c:S \rightarrow \{1, 2, \ldots,r\}$. A rainbow solution to an equation over $S$ is a solution such that all components are a different color. We prove that every 3-coloring of $\mathbb{N}$ with an upper density greater than $(4^s-1)/(3 \cdot 4^s)$ contains a rainbow solution to $x-y=z^k$. The rainbow
Rooted Almost-binary Phylogenetic Networks for which the Maximum Covering Subtree Problem is Solvable in Linear Time
math.COTakatora Suzuki, Han Guo, Momoko Hayamizu
Phylogenetic networks are a flexible model of evolution that can represent reticulate evolution and handle complex data. Tree-based networks, which are phylogenetic networks that have a spanning tree with the same root and leaf-set as the network itself, have been well studied. However, not all networks are tree-based. Francis-Semple-Steel (2018) thus introd
Sakshi Shukla, Praveen C. Srivastava, Kosuke Nomura, Larry Zamick
We report systematic large-scale shell-model calculation for Po isotopes with $A=$ 200 to 210. We have performed calculations using KHH7B interaction in the model space $Z$ = 58-114 and $N$ = 100-164 around doubly-magic $^{208}$Pb. We allow valence neutrons to occupy in the $1f_{5/2}$, $2p_{3/2}$, $2p_{1/2}$, and $0i_{13/2}$ orbitals, while two valence proto
Saku Sugawara, Shun Tsugita
Natural language understanding (NLU) studies often exaggerate or underestimate the capabilities of systems, thereby limiting the reproducibility of their findings. These erroneous evaluations can be attributed to the difficulty of defining and testing NLU adequately. In this position paper, we reconsider this challenge by identifying two types of researcher
Quasi-static responses of marine mussel plaques attached to deformable wet substrates under directional tensions
physics.bio-phYong Pang, Tao Liu
Quantifying the response of marine mussel plaque attachment on wet surfaces remains a significant challenge to a mechanistic understanding of plaque adhesion. Here, we developed a customised microscopy system combined with two-dimensional (2D) in-situ digital image correlation (DIC) to quantify the in-plane deformation of a deformable substrate that interact
Age of Information in Reservation Multi-Access Networks with Stochastic Arrivals: Analysis and Optimization
cs.ITQian Wang, He Chen
This paper analyzes and optimizes the average Age of Information (AAoI) of Frame Slotted ALOHA with Reservation and Data slots (FSA-RD) in a multi-access network, where multiple users transmit their randomly generated status updates to a common access point in a framed manner. Each frame consists of one reservation slot and several data slots. The reservatio
PLCMOS -- a data-driven non-intrusive metric for the evaluation of packet loss concealment algorithms
cs.SDLorenz Diener, Marju Purin, Sten Sootla, Ando Saabas
Speech quality assessment is a problem for every researcher working on models that produce or process speech. Human subjective ratings, the gold standard in speech quality assessment, are expensive and time-consuming to acquire in a quantity that is sufficient to get reliable data, while automated objective metrics show a low correlation with gold standard r
K. Kohno, S. Fujimoto, A. Tsujita, V. Kokorev
The ALMA lensing cluster survey (ALCS) is a 96-hr large program dedicated to uncovering and characterizing intrinsically faint continuum sources and line emitters with the assistance of gravitational lensing. All 33 cluster fields were selected from HST/Spitzer treasury programs including CLASH, Hubble Frontier Fields, and RELICS, which also have Herschel an
Kazuo Murota, Akihisa Tamura
The Shapley-Folkman theorem is a statement about the Minkowski sum of (non-convex) sets, expressing the closeness of the Minkowski sum to convexity in a quantitative manner. This paper establishes similar theorems for integrally convex sets, L-natural-convex sets, and M-natural-convex sets, which are major classes of discrete convex sets in discrete convex a
Anurag Dey, Probal Chaudhuri
The Horvitz-Thompson (HT), the Rao-Hartley-Cochran (RHC) and the generalized regression (GREG) estimators of the finite population mean are considered, when the observations are from an infinite dimensional space. We compare these estimators based on their asymptotic distributions under some commonly used sampling designs and some superpopulations satisfying
Manas Kulkarni, Satya N. Majumdar
We provide a general framework to compute the probability distribution $F_r(t)$ of the first detection time of a 'state of interest' in a generic quantum system subjected to random projective measurements. In our 'quantum resetting' protocol, resetting of a state is not implemented by an additional classical stochastic move, but rather by the random projecti
An atomistically informed multiscale approach to the intrusion and extrusion of water in hydrophobic nanopores
cond-mat.softGonçalo Paulo, Alberto Gubbiotti, Alberto Giacomello
Understanding intrusion and extrusion in nanoporous materials is a challenging multiscale problem of utmost importance for applications ranging from energy storage and dissipation to water desalination and hydrophobic gating in ion channels. Including atomistic details in simulations is required to predict the overall behavior of such systems, because the st
Hugo Thimonier, Fabrice Popineau, Arpad Rimmel, Bich-Liên Doan
Anomaly detection is vital in many domains, such as finance, healthcare, and cybersecurity. In this paper, we propose a novel deep anomaly detection method for tabular data that leverages Non-Parametric Transformers (NPTs), a model initially proposed for supervised tasks, to capture both feature-feature and sample-sample dependencies. In a reconstruction-bas
Chantal David, Patrick Meisner
We compute the expected value of Dirichlet $L$-functions defined over $\mathbb{F}_q[T]$ attached to cubic characters evaluated at an arbitrary $s \in (0,1)$. We find a transition term at the point $s=\frac{1}{3}$, reminiscent of the transition at the point $s=\frac{1}{2}$ of the bound for the size of an $L$-function implied by the Lindel\"of hypothesis. We s
Alberto Muñoz-Ortiz, David Vilares
The usefulness of part-of-speech tags for parsing has been heavily questioned due to the success of word-contextualized parsers. Yet, most studies are limited to coarse-grained tags and high quality written content; while we know little about their influence when it comes to models in production that face lexical errors. We expand these setups and design an
Atakan Coban
During the periods of sudden transition to online education, the opportunity to make applications that might attract students' attention to the course has decreased even more. Although this deficiency was tried to be eliminated with videos and simulations, it was not possible to ensure active participation of students in some cases. In this study, the Algodo
Marwa El Halabi, Federico Fusco, Ashkan Norouzi-Fard, Jakab Tardos
Streaming submodular maximization is a natural model for the task of selecting a representative subset from a large-scale dataset. If datapoints have sensitive attributes such as gender or race, it becomes important to enforce fairness to avoid bias and discrimination. This has spurred significant interest in developing fair machine learning algorithms. Rece
Christian Herglotz, Werner Robitza, Alexander Raake, Tobias Hossfeld
This paper uses a crowdsourced dataset of online video streaming sessions to investigate opportunities to reduce the power consumption while considering QoE. For this, we base our work on prior studies which model both the end-user's QoE and the end-user device's power consumption with the help of high-level video features such as the bitrate, the frame rate
Dominik Thönnes, Ulrich Rüde
In this work, we present how code generation techniques significantly improve the performance of the computational kernels in the HyTeG software framework. This HPC framework combines the performance and memory advantages of matrix-free multigrid solvers with the flexibility of unstructured meshes. The pystencils code generation toolbox is used to replace th
Yubao Tang, Ruqing Zhang, Jiafeng Guo, Jiangui Chen
Recently, a new paradigm called Differentiable Search Index (DSI) has been proposed for document retrieval, wherein a sequence-to-sequence model is learned to directly map queries to relevant document identifiers. The key idea behind DSI is to fully parameterize traditional ``index-retrieve'' pipelines within a single neural model, by encoding all documents
Thinking Twice: Clinical-Inspired Thyroid Ultrasound Lesion Detection Based on Feature Feedback
cs.CVLingtao Wang, Jianrui Ding, Fenghe Tang, Chunping Ning
Accurate detection of thyroid lesions is a critical aspect of computer-aided diagnosis. However, most existing detection methods perform only one feature extraction process and then fuse multi-scale features, which can be affected by noise and blurred features in ultrasound images. In this study, we propose a novel detection network based on a feature feedba
Simon Schindler, Martin Uray, Stefan Huber
Reinforcement Learning (RL) is a powerful machine learning paradigm that has been applied in various fields such as robotics, natural language processing and game playing achieving state-of-the-art results. Targeted to solve sequential decision making problems, it is by design able to learn from experience and therefore adapt to changing dynamic environments
Shivam Bajpeyi, Dhiraj Patel, S. Sivananthan
In this paper, we address the random sampling problem for the class of Mellin band-limited functions BT which is concentrated on a bounded cube. It is established that any function in BT can be approximated by an element in a finite-dimensional subspace of BT. Utilizing the notion of covering number and Bernstein's inequality to the sum of independent random
Elise Özalp, Georgios Margazoglou, Luca Magri
The forecasting and computation of the stability of chaotic systems from partial observations are tasks for which traditional equation-based methods may not be suitable. In this computational paper, we propose data-driven methods to (i) infer the dynamics of unobserved (hidden) chaotic variables (full-state reconstruction); (ii) time forecast the evolution o
Hoang Ky Nguyen, Mustapha Azreg-Aïnou
It is known that the formation of a wormhole typically involves a violation of the Weak Energy Condition (WEC), but the reverse is not necessarily true. In the context of Brans-Dicke gravity, the $\textit{generalized}$ Campanelli-Lousto solution, which we shall unveil in this paper, demonstrates a WEC violation that coincides with the appearance of $\textit{
Jan Kretinsky, Tobias Meggendorfer, Maximilian Prokop, Sabine Rieder
We provide a learning-based technique for guessing a winning strategy in a parity game originating from an LTL synthesis problem. A cheaply obtained guess can be useful in several applications. Not only can the guessed strategy be applied as best-effort in cases where the game's huge size prohibits rigorous approaches, but it can also increase the scalabilit
Debayan Banerjee, Pranav Ajit Nair, Ricardo Usbeck, Chris Biemann
In this work, we analyse the role of output vocabulary for text-to-text (T2T) models on the task of SPARQL semantic parsing. We perform experiments within the the context of knowledge graph question answering (KGQA), where the task is to convert questions in natural language to the SPARQL query language. We observe that the query vocabulary is distinct from
Sven-Erik Ekström, David Meadon
Consider the Toeplitz matrix $T_n(f)$ generated by the symbol $f(\theta)=\hat{f}_r e^{\mathbf{i}r\theta}+\hat{f}_0+\hat{f}_{-s} e^{-\mathbf{i}s\theta}$, where $\hat{f}_r, \hat{f}_0, \hat{f}_{-s} \in \mathbb{C}$ and $0<r<n,~0<s<n$. For $r=s=1$ we have the classical tridiagonal Toeplitz matrices, for which the eigenvalues and eigenvectors are known. Similarly,
Study of anomalous $W^-W^+\gamma/Z$ couplings using polarizations and spin correlations in $e^-e^+\to W^-W^+$ with polarized beams
hep-phAmir Subba, Ritesh K. Singh
We study the anomalous $W^-W^+\gamma/Z$ couplings in $e^-e^+\to W^-W^+$ followed by semileptonic decay using a complete set of polarization and spin correlation observables of $W$ boson with the longitudinally polarized beam. We consider a complete set of dimension-six operators affecting $W^-W^+\gamma/Z$ vertex, which are $SU(2)\times U(1)$ gauge invariant.
Massimo Bianchi, Giorgio Di Russo, Alfredo Grillo, Jose Francisco Morales
Topological stars, or top stars for brevity, are smooth horizonless static solutions of Einstein-Maxwell theory in 5-d that reduce to spherically symmetric solutions of Einstein-Maxwell-Dilaton theory in 4-d. We study linear scalar perturbations of top stars and argue for their stability and deformability. We tackle the problem with different techniques incl
Yican Sun, Hongfei Fu, Krishnendu Chatterjee, Amir Kafshdar Goharshady
Probabilistic recurrence relations (PRRs) are a standard formalism for describing the runtime of a randomized algorithm. Given a PRR and a time limit $\kappa$, we consider the classical concept of tail probability $\Pr[T \ge \kappa]$, i.e., the probability that the randomized runtime $T$ of the PRR exceeds the time limit $\kappa$. Our focus is the formal ana
Andrea Seppi, Graham Smith, Jérémy Toulisse
We provide a full classification of complete maximal $p$-dimensional spacelike submanifolds in the pseudo-hyperbolic space $\mathbf{H}^{p,q}$, and we study its applications to Teichm\"uller theory and to the theory of Anosov representations of hyperbolic groups in $\mathsf{PO}(p,q+1)$.
The Calculation of Kinetic and Static Friction Coefficient and Friction Graph Analysis Using Arduino
physics.ed-phAtakan Coban, Seher Boyacı
In this study, with the help of the Arduino UNO and Load-cell force sensor, a simple experimental material has been developed to calculate kinetic and static friction coefficients and analyse the friction force in detail. The system with a force sensor mounted on it is placed on the plane. With the help of the rope attached to the force sensor, a force, whos
Analysis of modular CMA-ES on strict box-constrained problems in the SBOX-COST benchmarking suite
cs.NEDiederick Vermetten, Manuel López-Ibáñez, Olaf Mersmann, Richard Allmendinger
Box-constraints limit the domain of decision variables and are common in real-world optimization problems, for example, due to physical, natural or spatial limitations. Consequently, solutions violating a box-constraint may not be evaluable. This assumption is often ignored in the literature, e.g., existing benchmark suites, such as COCO/BBOB, allow the opti
Felix Joos, Jonathan Schrodt
Let $T$ be an oriented tree on $n$ vertices with maximum degree at most $e^{o(\sqrt{\log n})}$. If $G$ is a digraph on $n$ vertices with minimum semidegree $\delta^0(G)\geq(\frac12+o(1))n$, then $G$ contains $T$ as a spanning tree, as recently shown by Kathapurkar and Montgomery (in fact, they only require maximum degree $o(n/\log n)$). This generalizes the
Distinguishing the nanohertz gravitational-wave sources by the observations of compact dark matter subhalos
astro-ph.COJing Liu
The latest pulsar timing array data reveals evidence of nanohertz gravitational waves (GWs), which have been explained by both cosmological and astrophysical sources. However, current observations lack the precision needed to differentiate between different models from the spectral index. We find that the cosmological GW sources, including bubble collisions,
Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
cs.CLZiwei He, Meng Yang, Minwei Feng, Jingcheng Yin
The transformer model is known to be computationally demanding, and prohibitively costly for long sequences, as the self-attention module uses a quadratic time and space complexity with respect to sequence length. Many researchers have focused on designing new forms of self-attention or introducing new parameters to overcome this limitation, however a large
Michael Tang, Shunyu Yao, John Yang, Karthik Narasimhan
We propose Referral-Augmented Retrieval (RAR), a simple technique that concatenates document indices with referrals, i.e. text from other documents that cite or link to the given document, to provide significant performance gains for zero-shot information retrieval. The key insight behind our method is that referrals provide a more complete, multi-view repre
Jiesheng Yang, Andreas Wilde, Karsten Menzel, Md Zubair Sheikh
Construction progress monitoring (CPM) is essential for effective project management, ensuring on-time and on-budget delivery. Traditional CPM methods often rely on manual inspection and reporting, which are time-consuming and prone to errors. This paper proposes a novel approach for automated CPM using state-of-the-art object detection algorithms. The propo
Zachary Ankner, Naomi Saphra, Davis Blalock, Jonathan Frankle
Most works on transformers trained with the Masked Language Modeling (MLM) objective use the original BERT model's fixed masking rate of 15%. We propose to instead dynamically schedule the masking rate throughout training. We find that linearly decreasing the masking rate over the course of pretraining improves average GLUE accuracy by up to 0.46% and 0.25%
David Viennot
We study the fuzzy spaces (as special examples of noncommutative manifolds) with their quasicoherent states in order to find their pertinent metrics. We show that they are naturally endowed with two natural "quantum metrics" which are associated with quantum fluctuations of "paths". The first one provides the length the mean path whereas the second one provi
Dongqing Wang, Tong Zhang, Alaa Abboud, Sabine Süsstrunk
We propose InNeRF360, an automatic system that accurately removes text-specified objects from 360-degree Neural Radiance Fields (NeRF). The challenge is to effectively remove objects while inpainting perceptually consistent content for the missing regions, which is particularly demanding for existing NeRF models due to their implicit volumetric representatio
Ameet Deshpande, Carlos E. Jimenez, Howard Chen, Vishvak Murahari
Semantic textual similarity (STS), a cornerstone task in NLP, measures the degree of similarity between a pair of sentences, and has broad application in fields such as information retrieval and natural language understanding. However, sentence similarity can be inherently ambiguous, depending on the specific aspect of interest. We resolve this ambiguity by
Philipp Wiesner, Ramin Khalili, Dennis Grinwald, Pratik Agrawal
Federated Learning (FL) is an emerging machine learning technique that enables distributed model training across data silos or edge devices without data sharing. Yet, FL inevitably introduces inefficiencies compared to centralized model training, which will further increase the already high energy usage and associated carbon emissions of machine learning in
Modeling Complex Object Changes in Satellite Image Time-Series: Approach based on CSP and Spatiotemporal Graph
cs.CVZouhayra Ayadi, Wadii Boulila, Imed Riadh Farah
This paper proposes a method for automatically monitoring and analyzing the evolution of complex geographic objects. The objects are modeled as a spatiotemporal graph, which separates filiation relations, spatial relations, and spatiotemporal relations, and is analyzed by detecting frequent sub-graphs using constraint satisfaction problems (CSP). The process
STAR: Boosting Low-Resource Information Extraction by Structure-to-Text Data Generation with Large Language Models
cs.CLMingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham
Information extraction tasks such as event extraction require an in-depth understanding of the output structure and sub-task dependencies. They heavily rely on task-specific training data in the form of (passage, target structure) pairs to obtain reasonable performance. However, obtaining such data through human annotation is costly, leading to a pressing ne
Wiktoria Zajkowska, Jakub Turczynski, Boguslawa Kurowska, Henryk Teisseyre
We present Al2O3-ZnAl2O4-ZnO nanostructure, which could be a prominent candidate for optoelectronics, mechanical and sensing applications. While ZnO and ZnAl2O4 composites are mostly synthesized by sol-gel technique, we propose a solid-vapor growth mechanism. To produce Al2O3-ZnAl2O4-ZnO nanostructure, we conduct ZnO:C powder heating resulting in ZnO nanowir
Anton A. Kutsenko
For the density of Galton-Watson processes in the Schr\"oder case, we derive a complete left tail asymptotic series consisting of power terms multiplied by periodic factors.
Petar Ivanov, Ivan Koychev, Momchil Hardalov, Preslav Nakov
Developing tools to automatically detect check-worthy claims in political debates and speeches can greatly help moderators of debates, journalists, and fact-checkers. While previous work on this problem has focused exclusively on the text modality, here we explore the utility of the audio modality as an additional input. We create a new multimodal dataset (t
Pento-DIARef: A Diagnostic Dataset for Learning the Incremental Algorithm for Referring Expression Generation from Examples
cs.CLPhilipp Sadler, David Schlangen
NLP tasks are typically defined extensionally through datasets containing example instantiations (e.g., pairs of image i and text t), but motivated intensionally through capabilities invoked in verbal descriptions of the task (e.g., "t is a description of i, for which the content of i needs to be recognised and understood"). We present Pento-DIARef, a diagno
Beomsu Kim, Gihyun Kwon, Kwanyoung Kim, Jong Chul Ye
Diffusion models are a powerful class of generative models which simulate stochastic differential equations (SDEs) to generate data from noise. While diffusion models have achieved remarkable progress, they have limitations in unpaired image-to-image (I2I) translation tasks due to the Gaussian prior assumption. Schr\"{o}dinger Bridge (SB), which learns an SD
Yoshiki Aibara, Yoshimichi Ueda
We explain how Pusz--Woronowicz's idea of their functional calculus fits the theory of Lebesgue decomposition for positive operators on Hilbert spaces initially developed by Ando. In this way, we reconstruct the essential and fundamental part of the theory.
Błażej Leporowski, Arian Bakhtiarnia, Nicole Bonnici, Adrian Muscat
We introduce the first audio-visual dataset for traffic anomaly detection taken from real-world scenes, called MAVAD, with a diverse range of weather and illumination conditions. In addition, we propose a novel method named AVACA that combines visual and audio features extracted from video sequences by means of cross-attention to detect anomalies. We demonst
Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
cs.CLJiahuan Li, Hao Zhou, Shujian Huang, Shanbo Cheng
Large-scale Pretrained Language Models (LLMs), such as ChatGPT and GPT4, have shown strong abilities in multilingual translations, without being explicitly trained on parallel corpora. It is interesting how the LLMs obtain their ability to carry out translation instructions for different languages. In this paper, we present a detailed analysis by finetuning
Rahul Rao, Ryan Selhorst, Jie Jiang, Benjamin S. Conner
CuInP2S6 (CIPS) is an emerging layered ferroelectric material with a TC above room temperature. When synthesized with Cu deficiencies (i.e., Cu1-xIn1+x/3P2S6), the material segregates into CIPS and In4/3P2S6 (IPS) self-assembled heterostructures within the same single crystal. This segregation results in significant in-plane and out-of-plane strains between
Petrosian Vah/'e, Maria Giovanna Dainotti
Bimodal distribution of the observed duration of gamma-ray bursts (GRBs) has led to two distinct progenitors; compact star mergers, either two neutron stars (NSs) or a NS and a black hole (BH), for short GRBs (SGRBs), and so-called collapsars for long GRBs (LGRBs). It is therefore expected that formation rate (FR) of LGRBs should be similar to the cosmic sta
Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models
cs.CLGeewook Kim, Hodong Lee, Daehee Kim, Haeji Jung
Recent advances in Large Language Models (LLMs) have stimulated a surge of research aimed at extending their applications to the visual domain. While these models exhibit promise in generating abstract image captions and facilitating natural conversations, their performance on text-rich images still requires improvement. In this paper, we introduce Contrasti
Yinguo Yang, Yiling Ye, Zhuoxiao Cheng, Guangchun Ruan
Battery storage is essential to enhance the flexibility and reliability of electric power systems by providing auxiliary services and load shifting. Storage owners typically gains incentives from quick responses to auxiliary service prices, but frequent charging and discharging also reduce its lifetime. Therefore, this paper embeds the battery degradation co
Yunfan LU, Guoqiang Liang, Yusheng Wang, Lin Wang
Video frames captured by rolling shutter (RS) cameras during fast camera movement frequently exhibit RS distortion and blur simultaneously. Naturally, recovering high-frame-rate global shutter (GS) sharp frames from an RS blur frame must simultaneously consider RS correction, deblur, and frame interpolation. A naive way is to decompose the whole process into
Junlei Zhang, Zhenzhong Lan, Junxian He
Contrastive learning has been the dominant approach to train state-of-the-art sentence embeddings. Previous studies have typically learned sentence embeddings either through the use of human-annotated natural language inference (NLI) data or via large-scale unlabeled sentences in an unsupervised manner. However, even in the case of unlabeled data, their acqu
Nathan Hu, Eric Mitchell, Christopher D. Manning, Chelsea Finn
Large language models encode impressively broad world knowledge in their parameters. However, the knowledge in static language models falls out of date, limiting the model's effective "shelf life." While online fine-tuning can reduce this degradation, we find that naively fine-tuning on a stream of documents leads to a low level of information uptake. We hyp
Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu
In this paper, we present HuatuoGPT, a large language model (LLM) for medical consultation. The core recipe of HuatuoGPT is to leverage both \textit{distilled data from ChatGPT} and \textit{real-world data from doctors} in the supervised fine-tuned stage. The responses of ChatGPT are usually detailed, well-presented and informative while it cannot perform li
Daman Arora, Himanshu Gaurav Singh, Mausam
The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past years. In response, we present JEEBench, a considerably more challenging benchmark dataset for evaluating the problem solving abilities of LLMs. We curate 515 challenging pre-engineering mathematics, physics and chemistry problems from th
Robustness of Quantum Random Walk Search Algorithm in Hypercube when only first or both first and second neighbors are measured
quant-phHristo Tonchev, Petar Danev
In this work we study the robustness of two modifications of quantum random walk search algorithm on hypercube. In the first previously suggested modification, on each even iteration only quantum walk is applied. And in the second, the closest neighbors of the solution are measured classically. In our approach the traversing coin is constructed by both gener
PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology
cs.CVYuxuan Sun, Chenglu Zhu, Sunyi Zheng, Kai Zhang
As advances in large language models (LLMs) and multimodal techniques continue to mature, the development of general-purpose multimodal large language models (MLLMs) has surged, offering significant applications in interpreting natural images. However, the field of pathology has largely remained untapped, particularly in gathering high-quality data and desig
Towards Cyber Security for Low-Carbon Transportation: Overview, Challenges and Future Directions
cs.ITYue Cao, Sifan Li, Chenchen Lv, Di Wang
In recent years, low-carbon transportation has become an indispensable part as sustainable development strategies of various countries, and plays a very important responsibility in promoting low-carbon cities. However, the security of low-carbon transportation has been threatened from various ways. For example, denial of service attacks pose a great threat t
Annotation Imputation to Individualize Predictions: Initial Studies on Distribution Dynamics and Model Predictions
cs.CLLondon Lowmanstone, Ruyuan Wan, Risako Owan, Jaehyung Kim
Annotating data via crowdsourcing is time-consuming and expensive. Due to these costs, dataset creators often have each annotator label only a small subset of the data. This leads to sparse datasets with examples that are marked by few annotators. The downside of this process is that if an annotator doesn't get to label a particular example, their perspectiv
Lucas Giroto de Oliveira, Elizabeth Bekker, Axel Diewald, Benjamin Nuss
This article introduces adaptations to the conventional frame structure in binary phase-modulated continuous wave (PMCW) radars with sequence generation via linear-feedbck shift registers and additional processing steps to enable joint radar-communication (RadCom) operation. In this context, a preamble structure based on pseudorandom binary sequences (PRBSs)
ToMChallenges: A Principle-Guided Dataset and Diverse Evaluation Tasks for Exploring Theory of Mind
cs.CLXiaomeng Ma, Lingyu Gao, Qihui Xu
Theory of Mind (ToM), the capacity to comprehend the mental states of distinct individuals, is essential for numerous practical applications. With the development of large language models (LLMs), there is a heated debate about whether they are able to perform ToM tasks. Previous studies have used different tasks and prompts to test the ToM on LLMs and the re
Tianyi Tang, Hongyuan Lu, Yuchen Eleanor Jiang, Haoyang Huang
Most research about natural language generation (NLG) relies on evaluation benchmarks with limited references for a sample, which may result in poor correlations with human judgements. The underlying reason is that one semantic meaning can actually be expressed in different forms, and the evaluation with a single or few references may not accurately reflect
GPT4Graph: Can Large Language Models Understand Graph Structured Data ? An Empirical Evaluation and Benchmarking
cs.AIJiayan Guo, Lun Du, Hengyu Liu, Mengyu Zhou
Large language models~(LLM) like ChatGPT have become indispensable to artificial general intelligence~(AGI), demonstrating excellent performance in various natural language processing tasks. In the real world, graph data is ubiquitous and an essential part of AGI and prevails in domains like social network analysis, bioinformatics and recommender systems. Th
Ximing Lu, Faeze Brahman, Peter West, Jaehun Jang
While extreme-scale language models have demonstrated exceptional performance on a variety of language tasks, the degree of control over these language models through pure prompting can often be limited. Directly fine-tuning such language models can be effective for tailoring them, but it can be either extremely costly (e.g., GPT-3) or not even feasible for
Siqi Ouyang, Lei Li
Recent large language models (LLMs) are promising for making decisions in grounded environments. However, LLMs frequently fail in complex decision-making tasks due to the misalignment between the pre-trained knowledge in LLMs and the actual rules in the environment. Existing methods require either costly gradient computation or lengthy in-context demonstrati
Riccardo W. Maffucci
We consider the graph degree sequences such that every realisation is a polyhedron. It turns out that there are exactly eight of them. All of these are unigraphic, in the sense that each is realised by exactly one polyhedron. This is a revisitation of a Theorem of Rao about sequences that are realised by only planar graphs. Our proof yields additional geomet
Quzhe Huang, Mingxu Tao, Chen Zhang, Zhenwei An
Large Language Models (LLMs), like LLaMA, have exhibited remarkable performance across various tasks. Nevertheless, when deployed to specific domains such as law or medicine, the models still confront the challenge of a deficiency in domain-specific knowledge and an inadequate capability to leverage that knowledge to resolve domain-related problems. In this
Javier Ureña-Carrion, Fariba Karimi, Gerardo Iñiguez, Mikko Kivelä
Core-periphery is a key feature of large-scale networks underlying a wide range of social, biological, and transportation phenomena. Despite its prevalence in empirical data, it is unclear whether this property is a consequence of more fundamental network evolution processes. While preferential attachment can create degree heterogeneity indistinguishable fro
Jiacheng Yao, Zhaohui Yang, Wei Xu, Mingzhe Chen
Due to the dynamics of wireless environment and limited bandwidth, wireless federated learning (FL) is challenged by frequent transmission errors and incomplete aggregation from devices. In order to overcome these challenges, we propose a global model reuse strategy (GoMORE) that reuses the outdated global model to replace the local model parameters once a t
Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong
Since the remarkable generation performance of large language models raised ethical and legal concerns, approaches to detect machine-generated text by embedding watermarks are being developed. However, we discover that the existing works fail to function appropriately in code generation tasks due to the task's nature of having low entropy. Extending a logit-
Bernard Boigelot, Pascal Fontaine, Baptiste Vergain
First-order logic fragments mixing quantifiers, arithmetic, and uninterpreted predicates are often undecidable, as is, for instance, Presburger arithmetic extended with a single uninterpreted unary predicate. In the SMT world, difference logic is a quite popular fragment of linear arithmetic which is less expressive than Presburger arithmetic. Difference log
Bistatic OFDM-based Joint Radar-Communication: Synchronization, Data Communication and Sensing
eess.SPLucas Giroto de Oliveira, David Brunner, Axel Diewald, Charlotte Muth
This article introduces a bistatic joint radar-communication (RadCom) system based on orthogonal frequency-division multiplexing (OFDM). In this context, the adopted OFDM frame structure is described and system model encompassing time, frequency, and sampling synchronization mismatches between the transmitter and receiver of the bistatic system is outlined.
Tianyu Liu, Afra Amini, Mrinmaya Sachan, Ryan Cotterell
Tasks that model the relation between pairs of tokens in a string are a vital part of understanding natural language. Such tasks, in general, require exhaustive pair-wise comparisons of tokens, thus having a quadratic runtime complexity in the length of the string. We show that these exhaustive comparisons can be avoided, and, moreover, the complexity of suc
Jiajie Zhang, Shulin Cao, Tingjia Zhang, Xin Lv
Explainable question answering (XQA) aims to answer a given question and provide an explanation why the answer is selected. Existing XQA methods focus on reasoning on a single knowledge source, e.g., structured knowledge bases, unstructured corpora, etc. However, integrating information from heterogeneous knowledge sources is essential to answer complex ques
Mayank Kumar Singh, Naoya Takahashi, Onoe Naoyuki
Many existing works on voice conversion (VC) tasks use automatic speech recognition (ASR) models for ensuring linguistic consistency between source and converted samples. However, for the low-data resource domains, training a high-quality ASR remains to be a challenging task. In this work, we propose a novel iterative way of improving both the ASR and VC mod
Existence of ground state solution of Nehari-Poho\v{z}aev type for a quasilinear Schr\"{o}dinger system
math.APJianqing Chen, Qian Zhang
This paper is concerned with the following quasilinear Schr\"{o}dinger system in the entire space $\mathbb R^{N}$($N\geq3$): $$\left\{\begin{align} &-\Delta u+A(x)u-\frac{1}{2}\triangle(u^{2})u = \frac{2\alpha}{\alpha+\beta}|u|^{\alpha-2}u|v|^{\beta},\\ &-\Delta v+Bv-\frac{1}{2}\triangle(v^{2})v=\frac{2\beta}{\alpha+\beta}|u|^{\alpha}|v|^{\beta-2}v.\end{alig
A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
cs.CLAlessandro Stolfo, Yonatan Belinkov, Mrinmaya Sachan
Mathematical reasoning in large language models (LMs) has garnered significant attention in recent work, but there is a limited understanding of how these models process and store information related to arithmetic tasks within their architecture. In order to improve our understanding of this aspect of language models, we present a mechanistic interpretation
Kevin Lin, Kyle Lo, Joseph E. Gonzalez, Dan Klein
When re-finding items, users who forget or are uncertain about identifying details often rely on creative strategies for expressing their information needs -- complex queries that describe content elements (e.g., book characters or events), information beyond the document text (e.g., descriptions of book covers), or personal context (e.g., when they read a b
Rayan Moukhader, Davi Rodrigues, Eleonora Raimondo, Vito Puliafito
Magnetic solitons are promising for applications due to their intrinsic properties such as small size, topological stability, ultralow power manipulation and potentially ultrafast operations. To date, research has focused on the manipulation of skyrmions, domain walls, and vortices by applied currents. The discovery of new methods to control magnetic paramet
Erica Cai, Brendan O'Connor
Current social science efforts automatically populate event databases of "who did what to whom?" tuples, by applying event extraction (EE) to text such as news. The event databases are used to analyze sociopolitical dynamics between actor pairs (dyads) in, e.g., international relations. While most EE methods heavily rely on rules or supervised learning, \emp
Study of the long-term $BVR_{c}I_{c}$ photometric variability of eight PMS stars in the young open cluster Trumpler 37
astro-ph.SRSunay Ibryamov, Gabriela Zidarova, Evgeni Semkov, Stoyanka Peneva
This paper reports results from our long-term $BV(RI)_{c}$ photometric CCD observations of eight pre-main-sequence stars collected from June 2008 to October 2022. These stars are located in the young open cluster Trumpler 37, in the field of GM Cephei. The observational data indicate that all stars from our study exhibit variability in all-optical passbands,
Mulyanto, Fiki Taufik Akbar, Bobby Eka Gunara
In this paper, we prove the decay estimate of Maxwell-Higgs system on four dimensional Schwarzschild spacetimes. We show that if the field equations support a Morawetz type estimate supported around the trapped surface, the uniform decay properties in the entire exterior of the Schwarzschild black holes can be obtained by using Sobolev inequalities and energ
Mete Sertkan, Sophia Althammer, Sebastian Hofstätter
In this paper, we introduce Ranger - a toolkit to facilitate the easy use of effect-size-based meta-analysis for multi-task evaluation in NLP and IR. We observed that our communities often face the challenge of aggregating results over incomparable metrics and scenarios, which makes conclusions and take-away messages less reliable. With Ranger, we aim to add
Vivek Verma, Eve Fleisig, Nicholas Tomlin, Dan Klein
We introduce Ghostbuster, a state-of-the-art system for detecting AI-generated text. Our method works by passing documents through a series of weaker language models, running a structured search over possible combinations of their features, and then training a classifier on the selected features to predict whether documents are AI-generated. Crucially, Ghost
Initial-boundary value problems for Poiseuille flow of nematic liquid crystal via full Ericksen-Leslie model
math.APGeng Chen, Yanbo Hu, Qingtian Zhang
In this paper, we study the initial-boundary value problem for the Poiseuille flow of hyperbolic-parabolic Ericksen-Leslie model of nematic liquid crystals in one space dimension. Due to the quasilinearity, the solution of this model in general forms cusp singularity. We prove the global existence of H\"older continuous solution, which may include cusp singu
Xiyan Fu, Anette Frank
We propose SETI (Systematicity Evaluation of Textual Inference), a novel and comprehensive benchmark designed for evaluating pre-trained language models (PLMs) for their systematicity capabilities in the domain of textual inference. Specifically, SETI offers three different NLI tasks and corresponding datasets to evaluate various types of systematicity in re
Xiao Pu, Mingqi Gao, Xiaojun Wan
Research on automated text summarization relies heavily on human and automatic evaluation. While recent work on human evaluation mainly adopted intrinsic evaluation methods, judging the generic quality of text summaries, e.g. informativeness and coherence, our work focuses on evaluating the usefulness of text summaries with extrinsic methods. We carefully de
Jacob Holford, Myoungkyu Lee, Yongyun Hwang
A data-driven implementation of a quasi-linear approximation is presented, extending a minimal quasi-linear approximation (MQLA) (Hwang & Ekchardt, J. Fluid Mech., 2020, 894:A23) to incorporate non-zero streamwise Fourier modes. A data-based approach is proposed, matching the two-dimensional wavenumber spectra for a fixed spanwise wavenumber between a direct
Zaccharie Ramzi, Pierre Ablin, Gabriel Peyré, Thomas Moreau
Implicit deep learning has recently gained popularity with applications ranging from meta-learning to Deep Equilibrium Networks (DEQs). In its general formulation, it relies on expressing some components of deep learning pipelines implicitly, typically via a root equation called the inner problem. In practice, the solution of the inner problem is approximate