Skip to content

May 2023 arXiv papers — page 26

Showing 2,5012,600 of 19,695 papers

  1. Yu Fei, Yifan Hou, Zeming Chen, Antoine Bosselut

    Various design settings for in-context learning (ICL), such as the choice and order of the in-context examples, can bias a model toward a particular prediction without being reflective of an understanding of the task. While many studies discuss these design choices, there have been few systematic investigations into categorizing them and mitigating their imp

  2. Zhenya Zhang, Jie An, Paolo Arcaini, Ichiro Hasuo

    Online monitoring is an effective validation approach for hybrid systems, that, at runtime, checks whether the (partial) signals of a system satisfy a specification in, e.g., Signal Temporal Logic (STL). The classic STL monitoring is performed by computing a robustness interval that specifies, at each instant, how far the monitored signals are from violating

  3. Petter Holme

    This is a comment on a recent review article about reputation and reciprocity as mechanisms promoting cooperation. I also discuss the necessary changes for the currently game-theory-based cooperation studies to become a complete theory of cooperation in our contemporary society.

  4. Lorenzo Baldassari, Ali Siahkoohi, Josselin Garnier, Knut Solna

    Since their initial introduction, score-based diffusion models (SDMs) have been successfully applied to solve a variety of linear inverse problems in finite-dimensional vector spaces due to their ability to efficiently approximate the posterior distribution. However, using SDMs for inverse problems in infinite-dimensional function spaces has only been addres

  5. Yue Liu, Ke Liang, Jun Xia, Sihang Zhou

    Deep graph clustering, which aims to group the nodes of a graph into disjoint clusters with deep neural networks, has achieved promising progress in recent years. However, the existing methods fail to scale to the large graph with million nodes. To solve this problem, a scalable deep graph clustering method (Dink-Net) is proposed with the idea of dilation an

  6. Nasser Heydari, Kazuo Muroi

    This article studies the application of the Pythagorean theorem in the Susa Mathematical Texts (\textbf{SMT}) and we discuss those texts whose problems and related calculations demonstrate its use. Among these texts, \textbf{SMT No.\,1} might be the most important as it contains a geometric application of the Pythagorean theorem.

  7. Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu

    As large language models continue to be widely developed, robust uncertainty quantification techniques will become crucial for their safe deployment in high-stakes scenarios. In this work, we explore how conformal prediction can be used to provide uncertainty quantification in language models for the specific task of multiple-choice question-answering. We fi

  8. A. S. Umar, K. Godbey, C. Simenel

    We employ the constrained density functional theory to investigate cluster phenomena for the $^{12}$C nucleus. The proton and neutron densities are generated from the placement of three $^{4}$He nuclei (alpha particles) geometrically. These densities are then used in a density constrained Hartree-Fock calculation that produces an antisymmetrized state with t

  9. Yuanwei Liu, Zhaolin Wang, Jiaqi Xu, Chongjun Ouyang

    Extremely large-scale antenna arrays, tremendously high frequencies, and new types of antennas are three clear trends in multi-antenna technology for supporting the sixth-generation (6G) networks. To properly account for the new characteristics introduced by these three trends in communication system design, the near-field spherical-wave propagation model ne

  10. Hristu Culetu

    The (4+1) dimensional conformally flat Eisenhart geometry is investigated in this work, stressing the contribution of the stress tensor generating its curvature. The energy-momentum tensor $T^{a}_{~b}$ is traceless and has only one nonzero component. It could be written as an anisotropic fluid with null transversal pressures and nonzero energy fluxes. The nu

  11. Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang

    Large Language Models (LLMs), such as LLaMA and T5, have shown exceptional performance across various tasks through fine-tuning. Although low-rank adaption (LoRA) has emerged to cheaply fine-tune these LLMs on downstream tasks, their deployment is still hindered by the vast model scale and computational costs. Post-training model pruning offers a way to comp

  12. Ella Rabinovich, Matan Vetzler, Samuel Ackerman, Ateret Anaby-Tavor

    Data drift is the change in model input data that is one of the key factors leading to machine learning models performance degradation over time. Monitoring drift helps detecting these issues and preventing their harmful consequences. Meaningful drift interpretation is a fundamental step towards effective re-training of the model. In this study we propose an

  13. Yongchao Huang, Yuhang He, Hong Ge

    In this work, we introduce a novel framework which combines physics and machine learning methods to analyse acoustic signals. Three methods are developed for this task: a Bayesian inference approach for inferring the spectral acoustics characteristics, a neural-physical model which equips a neural network with forward and backward physical losses, and the no

  14. Shreyas Malakarjun Patil, Loizos Michael, Constantine Dovrolis

    Natural target functions and tasks typically exhibit hierarchical modularity -- they can be broken down into simpler sub-functions that are organized in a hierarchy. Such sub-functions have two important features: they have a distinct set of inputs (input-separability) and they are reused as inputs higher in the hierarchy (reusability). Previous studies have

  15. Jawad Ettayb

    This paper deals with the condition pseudospectrum and essential condition pseudospectrum of operator pencils on n.a Banach spaces. We give a characterization of the condition pseudospectrum of operator pencils on n.a Banach spaces, the relation between the condition pseudospectrum of $(A,B)$ and the usual spectrum in a n.a valued field is investigated. Fina

  16. Subhajit Maity, Ram Kumar Karsh

    Tamper detection using image hash is a very common problem of modern days. Several research and advancements have already been done to address this problem. However, most of the existing methods lack the accuracy of tamper detection when the tampered area is low, as well as requiring long image hashes. In this paper, we propose a novel method objectively to

  17. Xiaojin Zhang, Yan Kang, Lixin Fan, Kai Chen

    Trustworthy Federated Learning (TFL) typically leverages protection mechanisms to guarantee privacy. However, protection mechanisms inevitably introduce utility loss or efficiency reduction while protecting data privacy. Therefore, protection mechanisms and their parameters should be carefully chosen to strike an optimal tradeoff between \textit{privacy leak

  18. Svetlana Gavrilova, Leonid Petrov

    We study probability measures on partitions based on symmetric Grothendieck polynomials. These deformations of Schur polynomials introduced in the K-theory of Grassmannians share many common properties. Our Grothendieck measures are analogs of the Schur measures on partitions introduced by Okounkov (arXiv:math/9907127 [math.RT]). Despite the similarity of de

  19. Wenjie Zhuo, Yifan Sun, Xiaohan Wang, Linchao Zhu

    This paper presents a whitening-based contrastive learning method for sentence embedding learning (WhitenedCSE), which combines contrastive learning with a novel shuffled group whitening. Generally, contrastive learning pulls distortions of a single sample (i.e., positive samples) close and push negative samples far away, correspondingly facilitating the ali

  20. Ran Chen, Baogang Xu

    Let $P_t$ and $C_t$ be a path and a cycle on $t$ vertices, respectively. In 2021, Choudum {\em et al.} [Disc. Math. 344 (2021) 112244] determined the structures of $(P_7,C_7,C_4$, diamond)-free and $(P_7,C_7,C_4$, gem)-frees, and gave correspondingly tight upper bounds to the chromatic numbers of these graphs. In this paper, we study the structure of $(P_7,

  21. Naichen Shi, Raed Al Kontar, Salar Fattahi

    In myriad statistical applications, data are collected from related but heterogeneous sources. These sources share some commonalities while containing idiosyncratic characteristics. One of the most fundamental challenges in such scenarios is to recover the shared and source-specific factors. Despite the existence of a few heuristic approaches, a generic algo

  22. J. C. Phillips

    Rhodopsin is a G-protein coupled receptor found in retinal rod cells, where it mediates monocrhromatic vision in dim light. It is one of the most studied proteins with thousands of reviewed entries in Uniprot. It has seven transmembrane segments, here examined for their hydrophobic character, and how that has evolved from chickens to humans. Elastic features

  23. David Smith, Joseph Samuel Myers, Craig S. Kaplan, Chaim Goodman-Strauss

    The recently discovered "hat" aperiodic monotile mixes unreflected and reflected tiles in every tiling it admits, leaving open the question of whether a single shape can tile aperiodically using translations and rotations alone. We show that a close relative of the hat -- the equilateral member of the continuum to which it belongs -- is a weakly chiral aperi

  24. Masaki Waga

    We present an algorithm to learn a deterministic timed automaton (DTA) via membership and equivalence queries. Our algorithm is an extension of the L* algorithm with a Myhill-Nerode style characterization of recognizable timed languages, which is the class of timed languages recognizable by DTAs. We first characterize the recognizable timed languages with a

  25. David McCune, Adam Graham-Squire

    Single Transferable Vote (STV) is a voting method used to elect multiple candidates in ranked-choice elections. One weakness of STV is that it fails multiple fairness criteria related to monotonicity and no show paradoxes. We analyze 1,079 local government STV elections in Scotland to estimate the frequency of such monotonicity anomalies in real-world electi

  26. Somnath Kumar, Vaibhav Balloli, Mercy Ranjit, Kabir Ahuja

    Large language models (LLMs) have revolutionized various domains but still struggle with non-Latin scripts and low-resource languages. This paper addresses the critical challenge of improving multilingual performance without extensive fine-tuning. We introduce a novel dynamic learning approach that optimizes prompt strategy, embedding model, and LLM per quer

  27. Lin Zhang, Xin Wang, Erica Cooper, Nicholas Evans

    Spoof localization, also called segment-level detection, is a crucial task that aims to locate spoofs in partially spoofed audio. The equal error rate (EER) is widely used to measure performance for such biometric scenarios. Although EER is the only threshold-free metric, it is usually calculated in a point-based way that uses scores and references with a pr

  28. Amir Joudaki, Hadi Daneshmand, Francis Bach

    In this paper, we explore the structure of the penultimate Gram matrix in deep neural networks, which contains the pairwise inner products of outputs corresponding to a batch of inputs. In several architectures it has been observed that this Gram matrix becomes degenerate with depth at initialization, which dramatically slows training. Normalization layers,

  29. Indrakshi Dey, Nicola Marchetti

    Industrial Internet-of-Things (IIoT) involve multiple groups of sensors, each group sending its observations on a particular phenomenon to a central computing platform over a multiple access channel (MAC). The central platform incorporates a decision fusion center (DFC) that arrives at global decisions regarding each set of phenomena by combining the receive

  30. Thomas Fernique, Olga Mikhailovna Sizova

    We provide a complete description of the edge-to-edge tilings with a regular triangle and a shield-shaped hexagon with no right angle. The case of a hexagon with a right angle is also briefly discussed.

  31. A. V. Kotikov

    We review field theoretical studies dedicated to understanding the effects of electron-electron interaction in graphene, which is characterized by gapless bands, strong electron-electron interactions, and emerging Lorentz invariance deep in the infrared. We consider the influence of interactions on the transport properties of the system as well as their supp

  32. Subhajit Barman, Indranil Chakraborty, Sajal Mukherjee

    In the present article, we study how different gravitational wave (GW) burst profiles in linearized gravity, with and without the asymptotic memory, may influence the harvesting between two static Unruh-DeWitt detectors. To this end, we investigate the following burst profiles -- Gaussian, sech-squared, Heaviside step function, and tanh. Out of these, the fi

  33. Jiaqi Miao, Siqi Sun

    Soft robots have demonstrated superior flexibility and functionality than conventional rigid robots. These versatile devices can respond to a wide range of external stimuli (including light, magnetic field, heat, electric field, etc.), and can perform sophisticated tasks. Notably, soft magnetic robots exhibit unparalleled advantages over numerous soft robots

  34. Hao Yang, Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi

    Pre-trained speech encoders have been central to pushing state-of-the-art results across various speech understanding and generation tasks. Nonetheless, the capabilities of these encoders in low-resource settings are yet to be thoroughly explored. To address this, we conduct a comprehensive set of experiments using a representative set of 3 state-of-the-art

  35. Kun Song, Yi Ren, Yi Lei, Chunfeng Wang

    Direct speech-to-speech translation (S2ST) has gradually become popular as it has many advantages compared with cascade S2ST. However, current research mainly focuses on the accuracy of semantic translation and ignores the speech style transfer from a source language to a target language. The lack of high-fidelity expressive parallel data makes such style tr

  36. Kazuma Sawaya, Yoshimasa Uematsu, Masaaki Imaizumi

    We develop a statistical inference method for generalized linear models (GLMs) in high-dimensional settings, where the number of unknown coefficients $p$ is of the same order as the sample size $n$. In this regime, constructing confidence intervals requires estimating unknown hyperparameters, such as the signal strength. However, existing estimators for the

  37. Matteo Beccaria, Alejandro Cabo-Bizet

    We consider the Schur index of $\mathcal N=4$ $U(N)$ SYM theory in 4d and its holographic giant graviton-type expansion at finite $N$. We compute the world-volume brane superconformal index by a recently proposed definition of the gauge holonomy integral as a multivariate residue. This is evaluated by a novel deformation algorithm that avoids Gr\"obner basis

  38. Henry Weld, Sijia Hu, Siqu Long, Josiah Poon

    Natural language understanding typically maps single utterances to a dual level semantic frame, sentence level intent and slot labels at the word level. The best performing models force explicit interaction between intent detection and slot filling. We present a novel tri-level joint natural language understanding approach, adding domain, and explicitly exch

  39. Mohammad Noormohammadi, Mehdi Khakian Ghomi, Hossein Haghi

    A combination of two unsupervised machine learning algorithms, DBSCAN and GMM are used to find members with a high probability of twelve open clusters, M38, NGC2099, Coma Ber, NGC752, M67, NGC2243, Alessi01, Bochum04, M34, M35, M41, and M48, based on Gaia DR3. These clusters have different ages, distances, and numbers of members which makes a suitable cover

  40. Hang Chen, Bingyu Liao, Jing Luo, Wenjing Zhu

    Reasoning, a crucial aspect of NLP research, has not been adequately addressed by prevailing models including Large Language Model. Conversation reasoning, as a critical component of it, remains largely unexplored due to the absence of a well-designed cognitive model. In this paper, inspired by intuition theory on conversation cognition, we develop a convers

  41. Reda Tiani, Uwe C. Täuber

    We numerically and analytically investigate the behavior of a non-equilibrium phase transition in the second Schl\"ogl autocatalytic reaction scheme. Our model incorporates both an interaction-induced phase separation and a bifurcation in the reaction kinetics, with these critical lines coalescing at a bicritical point in the macroscopic limit. We construct

  42. F. V. Difonzo, M. Roubalik, J. Marecek

    Virtual power plants and load aggregation are becoming increasingly common. There, one regulates the aggregate power output of an ensemble of distributed energy resources (DERs). Marecek et al. [Automatica, Volume 147, January 2023, 110743, arXiv:2110.03001] recently suggested that long-term averages of prices or incentives offered should exist and be indepe

  43. Sewade Ogun, Vincent Colotte, Emmanuel Vincent

    Flow-based generative models are widely used in text-to-speech (TTS) systems to learn the distribution of audio features (e.g., Mel-spectrograms) given the input tokens and to sample from this distribution to generate diverse utterances. However, in the zero-shot multi-speaker TTS scenario, the generated utterances lack diversity and naturalness. In this pap

  44. Subhadip Kumar

    Today information technology is a data-driven environment. The role of data is to empower business leaders to make decisions based on facts, trends, and statistical numbers. SAP is no exception. In modern days many companies use business suites like SAP on HANA S/4 or ERP or SAP Business Warehouse and other non-SAP applications and run those on HANA database

  45. Manuel Brack, Felix Friedrich, Patrick Schramowski, Kristian Kersting

    Text-conditioned image generation models have recently achieved astonishing results in image quality and text alignment and are consequently employed in a fast-growing number of applications. Since they are highly data-driven, relying on billion-sized datasets randomly scraped from the web, they also reproduce inappropriate human behavior. Specifically, we d

  46. G. S. Bisnovatyi-Kogan, A. M. Nikishin

    It is accepted in modern cosmology that the scalar field responsible for the inflationary stage of the early Universe is completely transformed into matter. It is assumed that the accelerated expansion is currently driven by dark energy (DE), which is likely determined by Einstein's cosmological constant. We consider a cosmological model where DE can have tw

  47. Hongqiu Wu, Shaohua Zhang, Yuchen Zhang, Hai Zhao

    In this paper, we study Chinese Spelling Correction (CSC) as a joint decision made by two separate models: a language model and an error model. Through empirical analysis, we find that fine-tuning BERT tends to over-fit the error model while under-fit the language model, resulting in poor generalization to out-of-distribution error patterns. Given that BERT

  48. Cem Suulker, Sophie Skach, Kaspar Althoefer

    The elastic bands integrated using the ruffles technique proved to be effective in enhancing the performance of the soft robotic structures. In the actuator application, the elastic bands greatly increased the bending capability and force capability of the structure, while in the eversion robot cap application, the elastic bands improved the performance slig

  49. Aysun Bozanta, Fuad Bayrak, Ayse Basar

    Social media platforms influence the way political campaigns are run and therefore they have become an increasingly important tool for politicians to directly interact with citizens. Previous elections in various countries have shown that social media data may significantly impact election results. In this study, we aim to predict the vote shares of parties

  50. Jinhua Liang, Xubo Liu, Haohe Liu, Huy Phan

    We presented the Treff adapter, a training-efficient adapter for CLAP, to boost zero-shot classification performance by making use of a small set of labelled data. Specifically, we designed CALM to retrieve the probability distribution of text-audio clips over classes using a set of audio-label pairs and combined it with CLAP's zero-shot classification resul

  51. Noam Rotstein, David Bensaid, Shaked Brody, Roy Ganz

    The advent of vision-language pre-training techniques enhanced substantial progress in the development of models for image captioning. However, these models frequently produce generic captions and may omit semantically important image details. This limitation can be traced back to the image-text datasets; while their captions typically offer a general descri

  52. Yonatan Gutman, Michael Levin, Tom Meyerovitch

    We prove an equivariant version of the classical Menger-Nobeling theorem regarding topological embeddings: Whenever a group $G$ acts on a finite-dimensional compact metric space $X$, a generic continuous equivariant function from $X$ into $([0,1]^r)^G$ is a topological embedding, provided that for every positive integer $N$ the space of points in $X$ with or

  53. Xuanqi Liu, Zhuotao Liu

    The community explored to build private inference frameworks for transformer-based large language models (LLMs) in a server-client setting, where the server holds the model parameters and the client inputs its private data (or prompt) for inference. However, these frameworks impose significant overhead when the private inputs are forward propagated through t

  54. Haobo Yang, Wenyu Wang, Ze Cao, Zhekai Duan

    This paper introduces a novel approach to evaluating deep learning models' capacity for in-diagram logic interpretation. Leveraging the intriguing realm of visual illusions, we establish a unique dataset, InDL, designed to rigorously test and benchmark these models. Deep learning has witnessed remarkable progress in domains such as computer vision and natura

  55. Minki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi

    Large Language Models (LLMs) have shown promising performance in knowledge-intensive reasoning tasks that require a compound understanding of knowledge. However, deployment of the LLMs in real-world applications can be challenging due to their high computational requirements and concerns on data privacy. Previous studies have focused on building task-specifi

  56. Zhuowei Sun, Hongyuan Cao, Li Chen, Jason P. Fine

    In linear models, omitting a covariate that is orthogonal to covariates in the model does not result in biased coefficient estimation. This in general does not hold for longitudinal data, where additional assumptions are needed to get unbiased coefficient estimation in addition to the orthogonality between omitted longitudinal covariates and longitudinal cov

  57. Amit Moryossef, Mathias Müller, Anne Göhring, Zifan Jiang

    Sign language translation systems are complex and require many components. As a result, it is very hard to compare methods across publications. We present an open-source implementation of a text-to-gloss-to-pose-to-video pipeline approach, demonstrating conversion from German to Swiss German Sign Language, French to French Sign Language of Switzerland, and I

  58. Marco Pegoraro, Clémentine Dominé, Emanuele Rodolà, Petar Veličković

    Antibody-antigen interactions play a crucial role in identifying and neutralizing harmful foreign molecules. In this paper, we investigate the optimal representation for predicting the binding sites in the two molecules and emphasize the importance of geometric information. Specifically, we compare different geometric deep learning methods applied to protein

  59. Mirko Consiglio

    Preparing the Gibbs state of an interacting quantum many-body system on noisy intermediate-scale quantum (NISQ) devices is a crucial task for exploring the thermodynamic properties in the quantum regime. It encompasses understanding protocols such as thermalization and out-of-equilibrium thermodynamics, as well as sampling from faithfully prepared Gibbs stat

  60. Matthias J. Ehrhardt, Silvia Gazzola, Sebastian J. Scott

    Variational regularization is commonly used to solve linear inverse problems, and involves augmenting a data fidelity by a regularizer. The regularizer is used to promote a priori information and is weighted by a regularization parameter. Selection of an appropriate regularization parameter is critical, with various choices leading to very different reconstr

  61. Christian Rohrbeck, Deborah A Costain

    Regression analysis under the assumption of monotonicity is a well-studied statistical problem and has been used in a wide range of applications. However, there remains a lack of a broadly applicable methodology that permits information borrowing, for efficiency gains, when jointly estimating multiple monotonic regression functions. We introduce such a metho

  62. Wentao Chao, Fuqing Duan, Xuechun Wang, Yingqian Wang

    Light field (LF) depth estimation is a crucial task with numerous practical applications. However, mainstream methods based on the multi-view stereo (MVS) are resource-intensive and time-consuming as they need to construct a finer cost volume. To address this issue and achieve a better trade-off between accuracy and efficiency, we propose an occlusion-aware

  63. Gongbo Tang, Christian Hardmeier

    Coreference resolution is the task of finding expressions that refer to the same entity in a text. Coreference models are generally trained on monolingual annotated data but annotating coreference is expensive and challenging. Hardmeier et al.(2013) have shown that parallel data contains latent anaphoric knowledge, but it has not been explored in end-to-end

  64. Hao Liu, Yanlin Wang, Zhao Wei, Yong Xu

    Refactoring is an indispensable practice of improving the quality and maintainability of source code in software evolution. Rename refactoring is the most frequently performed refactoring that suggests a new name for an identifier to enhance readability when the identifier is poorly named. However, most existing works only identify renaming activities betwee

  65. Ara Ghukasyan, Jack S. Baker, Oktay Goktas, Juan Carrasquilla

    As quantum computers become increasingly practical, so does the prospect of using quantum computation to improve upon traditional algorithms. Kernel methods in machine learning is one area where such improvements could be realized in the near future. Paired with kernel methods like support-vector machines, small and noisy quantum computers can evaluate class

  66. Ying Shi, Dong Wang, Lantian Li, Jiqing Han

    Most existing keyword spotting research focuses on conditions with slight or moderate noise. In this paper, we try to tackle a more challenging task: detecting keywords buried under strong interfering speech (10 times higher than the keyword in amplitude), and even worse, mixed with other keywords. We propose a novel Mix Training (MT) strategy that encourage

  67. Eyad Shaklab, Areg Karapetyan, Arjun Sharma, Murad Mebrahtu

    In addition to its crucial impact on customer satisfaction, last-mile delivery (LMD) is notorious for being the most time-consuming and costly stage of the shipping process. Pressing environmental concerns combined with the recent surge of e-commerce sales have sparked renewed interest in automation and electrification of last-mile logistics. To address the

  68. Hyeokjun Kwon, Sung Joon Maeng, Ismail Guvenc

    Advancements in unmanned aerial vehicle (UAV) technology have led to their increased utilization in various commercial and military applications. One such application is signal source search and localization (SSSL) using UAVs, which offers significant benefits over traditional ground-based methods due to improved RF signal reception at higher altitudes and i

  69. Shigeki Kawai, Orlando J. Silveira, Lauri Kurki, Zhangyu Yuan

    Synthesis of one-dimensional molecular arrays with tailored stereoisomers is challenging yet has a great potential for application in molecular opto-, electronic- and magnetic-devices, where the local array structure plays a decisive role in the functional properties. Here, we demonstrate construction and characterization of dehydroazulene isomer and diradic

  70. Stephan Rabanser, Anvith Thudi, Abhradeep Thakurta, Krishnamurthy Dvijotham

    Training reliable deep learning models which avoid making overconfident but incorrect predictions is a longstanding challenge. This challenge is further exacerbated when learning has to be differentially private: protection provided to sensitive data comes at the price of injecting additional randomness into the learning process. In this work, we conduct a t

  71. Indrakshi Dey, Nicola Marchetti

    Internet-of-Things (IoT) devices are low size, weight and power (SWaP), low complexity and include sensors, meters, wearables and trackers. Transmitting information with high signal power is exacting on device battery life, therefore an efficient link and network configuration is absolutely crucial to avoid signal power enhancement in interference-rich envir

  72. Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim

    Large language models (LLMs) learn not only natural text generation abilities but also social biases against different demographic groups from real-world data. This poses a critical risk when deploying LLM-based applications. Existing research and resources are not readily applicable in South Korea due to the differences in language and culture, both of whic

  73. James H Hepworth, Hendrik D Mouton

    This project aimed to design, simulate, and implement a two-axis inertially stabilised platform (ISP) for use in astronomical applications. It aimed to approximate the stabilisation of a Meade ETX-90 3.5" compound telescope at low-cost using a mechanical assembly designed to geometrically and inertially model the telescope. A set of system specifications was

  74. Yutao Mou, Xiaoshuai Song, Keqing He, Chen Zeng

    Generalized intent discovery aims to extend a closed-set in-domain intent classifier to an open-world intent set including in-domain and out-of-domain intents. The key challenges lie in pseudo label disambiguation and representation learning. Previous methods suffer from a coupling of pseudo label disambiguation and representation learning, that is, the reli

  75. Lei Li, Kai Fan, Lingyu Yang, Hongjia Li

    Existing wisdom demonstrates the significance of syntactic knowledge for the improvement of neural machine translation models. However, most previous works merely focus on leveraging the source syntax in the well-known encoder-decoder framework. In sharp contrast, this paper proposes an end-to-end translation architecture from the (graph \& sequence) structu

  76. Benjamin Brück, Robin J. Sroka

    In this note we present an alternative proof of a theorem of Gunnells, which states that the Steinberg module of $\operatorname{Sp_{2n}}(\mathbb{Q})$ is a cyclic $\operatorname{Sp_{2n}}(\mathbb{Z})$-module, generated by integral apartment classes.

  77. Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim

    The potential social harms that large language models pose, such as generating offensive content and reinforcing biases, are steeply rising. Existing works focus on coping with this concern while interacting with ill-intentioned users, such as those who explicitly make hate speech or elicit harmful responses. However, discussions on sensitive issues can beco

  78. Eun Jung Yeo, Kwanghee Choi, Sunhee Kim, Minhwa Chung

    This paper proposes an improved Goodness of Pronunciation (GoP) that utilizes Uncertainty Quantification (UQ) for automatic speech intelligibility assessment for dysarthric speech. Current GoP methods rely heavily on neural network-driven overconfident predictions, which is unsuitable for assessing dysarthric speech due to its significant acoustic difference

  79. Ori Nizan, Ayellet Tal

    Anomaly detection aims at identifying images that deviate significantly from the norm. We focus on algorithms that embed the normal training examples in space and when given a test image, detect anomalies based on the features distance to the k-nearest training neighbors. We propose a new operator that takes into account the varying structure & importance of

  80. Sabin Roman

    A previously overlooked relation governing planetary surface temperatures in terms of solar irradiance and top-of-atmosphere Bond albedo is identified. It reproduces the observed climates of Venus, Earth, and Titan, predicts condensation-level temperatures in the gas giants Jupiter, Saturn, Uranus, and Neptune, and extends naturally to rocky planets and larg

  81. Vasiliki Kougia, Simon Fetzel, Thomas Kirchmair, Erion Çano

    Memes are a popular form of communicating trends and ideas in social media and on the internet in general, combining the modalities of images and text. They can express humor and sarcasm but can also have offensive content. Analyzing and classifying memes automatically is challenging since their interpretation relies on the understanding of visual elements,

  82. Andrei Dumitrasc, Carola Kruse, Ulrich Ruede

    Deflation techniques are typically used to shift isolated clusters of small eigenvalues in order to obtain a tighter distribution and a smaller condition number. Such changes induce a positive effect in the convergence behavior of Krylov subspace methods, which are among the most popular iterative solvers for large sparse linear systems. We develop a deflati

  83. Uzi Pereg

    Entanglement assistance can improve communication rates significantly. Yet, its generation is susceptible to failure. The unreliable assistance model accounts for those challenges. Previous work provided an asymptotic formula that outlines the tradeoff between the unassisted and excess rates from entanglement assistance. We derive a full characterization for

  84. Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Chaojun Xiao

    This work examines the presence of modularity in pre-trained Transformers, a feature commonly found in human brains and thought to be vital for general intelligence. In analogy to human brains, we consider two main characteristics of modularity: (1) functional specialization of neurons: we evaluate whether each neuron is mainly specialized in a certain funct

  85. Zhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang

    Injecting external knowledge can improve the performance of pre-trained language models (PLMs) on various downstream NLP tasks. However, massive retraining is required to deploy new knowledge injection methods or knowledge bases for downstream tasks. In this work, we are the first to study how to improve the flexibility and efficiency of knowledge injection

  86. Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad, Aneesh Bose

    This paper presents HaVQA, the first multimodal dataset for visual question-answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, the dataset provides 12,044 gold standard English-Hausa paralle

  87. Mansour Zoubeirou A Mayaki, Michel Riveill

    Anomaly detection or more generally outliers detection is one of the most popular and challenging subject in theoretical and applied machine learning. The main challenge is that in general we have access to very few labeled data or no labels at all. In this paper, we present a new semi-supervised anomaly detection method called \textbf{AnoRand} by combining

  88. Wen-Han Dong, Jinbo Pan, Jia-Tao Sun, Shixuan Du

    Phonons have provided an ideal platform for a variety of intriguing physical states, such as non-Abelian braiding and the Haldane model. It is promising that phonons will realize the complicated nodal states accompanying unusual quantum phenomena. Here, we propose the hybrid nodal surface and nodal line (NS+NL) phonons beyond the single-genre nodal phonons.

  89. Zhanhao Hu, Jun Zhu, Bo Zhang, Xiaolin Hu

    Recent works found that deep neural networks (DNNs) can be fooled by adversarial examples, which are crafted by adding adversarial noise on clean inputs. The accuracy of DNNs on adversarial examples will decrease as the magnitude of the adversarial noise increase. In this study, we show that DNNs can be also fooled when the noise is very small under certain

  90. Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos

    We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorithm, quantile temporal-difference learning (QTD), for this task. We reach the surprising conclusion that even if a practitioner has no interest in the return distribution beyond the

  91. Mohammad Lataifeh, Xavier Carrasco, Ashraf Elnagar, Naveed Ahmed

    Recent advances in Generative Adversarial Networks (GANs) continue to attract the attention of researchers in different fields due to the wide range of applications devised to take advantage of their key features. Most recent GANs are focused on realism, however, generating hyper-realistic output is not a priority for some domains, as in the case of this wor

  92. Nan Wu, Zhuohan Li, Wanzhou Zhang

    In recent years, developing unsupervised machine learning for identifying phase transition is a research direction. In this paper, we introduce a two-times clustering method that can help select perfect configurations from a set of degenerate samples and assign the configuration with labels in a manner of unsupervised machine learning. These perfect configur

  93. Angelina Agabin, J. Xavier Prochaska

    This thesis presents a new algorithm to mitigate cloud masking in the analysis of sea surface temperature (SST) data generated by remote sensing technologies, e.g., Clouds interfere with the analysis of all remote sensing data using wavelengths shorter than 12 microns, significantly limiting the quantity of usable data and creating a biased geographical dist

  94. Zi-Hao Chen, YiJing Yan

    The dissipaton equations of motion (DEOM) method is one of the most popular methods for simulating quantum impurity systems. In this article, we use DOEM theory to deal with the Kondo problem of the double quantum dots (DQDs) impurity system. We focus on the impurity spectral function and the total noise spectral function, this two function will be used to d

  95. Toshiaki Maeno, Satoshi Naito, Daisuke Sagaki

    In our previous paper, we gave a presentation of the torus-equivariant quantum $K$-theory ring $QK_{H}(Fl_{n+1})$ of the (full) flag manifold $Fl_{n+1}$ of type $A_{n}$ as a quotient of a polynomial ring by an explicit ideal. In this paper, we prove that quantum double Grothendieck polynomials, introduced by Lenart-Maeno, represent the corresponding (opposit

  96. Shinichiro Yamano, Takaya Matsuura, Yui Kuramochi, Toshihiko Sasaki

    Continuous Variable (CV) quantum key distribution (QKD) is a promising candidate for practical implementations due to its compatibility with the existing communication technology. A trusted device scenario assuming that an adversary has no access to imperfections such as electronic noises in the detector is expected to provide significant improvement in the

  97. Rok Cestnik, Erik A. Martens

    We present an exact dimensionality reduction for dynamics of an arbitrary array of globally coupled complex-valued Riccati equations. It generalizes the Watanabe-Strogatz theory [Phys. Rev. Lett. 70, 2391 (1993)] for sinusoidally coupled phase oscillators and seamlessly includes quadratic integrate-and-fire neurons as the real-valued special case. This simpl

  98. Guangtao Zeng, Peiyuan Zhang, Wei Lu

    Fine-tuning pre-trained language models for multiple tasks tends to be expensive in terms of storage. To mitigate this, parameter-efficient transfer learning (PETL) methods have been proposed to address this issue, but they still require a significant number of parameters and storage when being applied to broader ranges of tasks. To achieve even greater stor

  99. Hao Guo, Wanxin Li, Mark Nejad

    Blockchain-based IoT systems can manage IoT devices and achieve a high level of data integrity, security, and provenance. However, incorporating existing consensus protocols in many IoT systems limits scalability and leads to high computational cost and consensus latency. In addition, location-centric characteristics of many IoT applications paired with limi

  100. Han Wang, Ming Shan Hee, Md Rabiul Awal, Kenny Tsu Wei Choo

    Recent research has focused on using large language models (LLMs) to generate explanations for hate speech through fine-tuning or prompting. Despite the growing interest in this area, these generated explanations' effectiveness and potential limitations remain poorly understood. A key concern is that these explanations, generated by LLMs, may lead to erroneo