Skip to content

May 2023 arXiv papers — page 67

Showing 6,6016,700 of 19,695 papers

  1. Chenguang Lu

    A new trend in deep learning, represented by Mutual Information Neural Estimation (MINE) and Information Noise Contrast Estimation (InfoNCE), is emerging. In this trend, similarity functions and Estimated Mutual Information (EMI) are used as learning and objective functions. Coincidentally, EMI is essentially the same as Semantic Mutual Information (SeMI) pr

  2. Hiroki Ono, Yoshitaka Umeda, Kaito Yoshida, Kenzaburo Tsutsui

    We have investigated structural, electronic and magnetic properties of H$_2$Pc on Fe$_2$N/Fe using low-energy electron diffraction and soft x-ray absorption spectroscopy/x-ray magnetic circular dichroism. Element specific magnetization curves reveal that the magnetic coupling with H$_2$Pc enhances the perpendicular magnetic anisotropy of Fe$_2$N/Fe at the H$

  3. Ying Xiao, Shangwen Wang, Sicen Liu, Dingyuan Xue

    Software built on top of machine learning algorithms is becoming increasingly prevalent in a variety of fields, including college admissions, healthcare, insurance, and justice. The effectiveness and efficiency of these systems heavily depend on the quality of the training datasets. Biased datasets can lead to unfair and potentially harmful outcomes, particu

  4. Naveed Akhtar, Muhammad A. A. K. Jalwana

    Originally inspired by game-theory, path attribution framework stands out among the post-hoc model interpretation tools due to its axiomatic nature. However, recent developments show that this framework can still suffer from counter-intuitive results. Moreover, specifically for deep visual models, the existing path-based methods also fall short on conforming

  5. Melih İs, İsmet Karaca

    In this paper, we transfer the problem of measuring navigational complexity in topological spaces to the nearness theory. We investigate the most important component of this problem, the topological complexity number (denoted by TC), with its different versions including relative and higher TC, on the proximal Schwarz genus as well as the proximal (higher) h

  6. Raghav Gupta, Renat Aksitov, Samrat Phatale, Simral Chaudhary

    Conversational recommendation systems (CRS) aim to recommend suitable items to users through natural language conversation. However, most CRS approaches do not effectively utilize the signal provided by these conversations. They rely heavily on explicit external knowledge e.g., knowledge graphs to augment the models' understanding of the items and attributes

  7. Yuki Saito, Shinnosuke Takamichi, Eiji Iimori, Kentaro Tachibana

    We propose ChatGPT-EDSS, an empathetic dialogue speech synthesis (EDSS) method using ChatGPT for extracting dialogue context. ChatGPT is a chatbot that can deeply understand the content and purpose of an input prompt and appropriately respond to the user's request. We focus on ChatGPT's reading comprehension and introduce it to EDSS, a task of synthesizing s

  8. Yunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang

    Weakly-supervised text classification trains a classifier using the label name of each target class as the only supervision, which largely reduces human annotation efforts. Most existing methods first use the label names as static keyword-based features to generate pseudo labels, which are then used for final classifier training. While reasonable, such a com

  9. C. Di Bello, A. V. Chechkin, A. K. Hartmann, Z. Palmowski

    Stochastic resetting is a rapidly developing topic in the field of stochastic processes and their applications. It denotes the occasional reset of a diffusing particle to its starting point and effects, inter alia, optimal first-passage times to a target. Recently the concept of partial resetting, in which the particle is reset to a given fraction of the cur

  10. Hyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu

    Dialogue systems are frequently updated to accommodate new services, but naively updating them by continually training with data for new services in diminishing performance on previously learnt services. Motivated by the insight that dialogue state tracking (DST), a crucial component of dialogue systems that estimates the user's goal as a conversation procee

  11. F. N. Arzikulov, I. A. Karimjanov, S. M. Umrzaqov

    In the present paper automorphisms, local and 2-local automorphisms of $n$-dimensional null-filiform and filiform associative algebras are studied. Namely, a common form of the matrix of automorphisms and local automorphisms of these algebras is clarified. It turns out that the common form of the matrix of an automorphism on these algebras does not coincide

  12. Bobir Toshmatov, Zdeněk Stuchlík, Bobomurat Ahmedov

    One of the eminent generalizations of theory of general relativity is the Rastall gravity which was {constructed} based on the assumption of the non-conserved energy-momentum tensor of the matter field. Despite in the literature several solutions of black holes in the Rastall gravity coupled to the electromagnetic field have been presented, in the current pa

  13. Fangkai Jiao, Zhiyang Teng, Bosheng Ding, Zhengyuan Liu

    Existing efforts to improve logical reasoning ability of language models have predominantly relied on supervised fine-tuning, hindering generalization to new domains and/or tasks. The development of Large Langauge Models (LLMs) has demonstrated the capacity of compressing abundant knowledge into a single proxy, enabling them to tackle multiple tasks effectiv

  14. Alejandro Silva

    The problem of detecting chirps is present in many applications of Signal Processing. Proper denoising, which involves filtering the signals after their acquisition, improves the efficacy of their detection. This manuscript describes how a recently-published method of Time-Frequency Analysis (TFA) with reassignment, namely the Newton Time-Extracting Wavelet

  15. Yuhao Liang, Fan Yu, Yangze Li, Pengcheng Guo

    The recently proposed serialized output training (SOT) simplifies multi-talker automatic speech recognition (ASR) by generating speaker transcriptions separated by a special token. However, frequent speaker changes can make speaker change prediction difficult. To address this, we propose boundary-aware serialized output training (BA-SOT), which explicitly in

  16. Insung Kong, Yuha Park, Joonhyuk Jung, Kwonsang Lee

    Weighting methods in causal inference have been widely used to achieve a desirable level of covariate balancing. However, the existing weighting methods have desirable theoretical properties only when a certain model, either the propensity score or outcome regression model, is correctly specified. In addition, the corresponding estimators do not behave well

  17. A. V. Valov, E. V. Dontsov, A. N. Baykin, S. V. Golovin

    The capability to simulate a hydraulic fracturing process is an essential tool that can be used to optimize treatment design and increase the efficiency of field operations. In most practical cases, hydraulic fractures propagate in a multi-layered rock formation. As a result, there is a need to incorporate the effect of such heterogeneities in fracturing mod

  18. Yuki Saito, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana

    We present CALLS, a Japanese speech corpus that considers phone calls in a customer center as a new domain of empathetic spoken dialogue. The existing STUDIES corpus covers only empathetic dialogue between a teacher and student in a school. To extend the application range of empathetic dialogue speech synthesis (EDSS), we designed our corpus to include the s

  19. Kirill Rivkin

    Spin wave computing device where an algorithm can be encoded by recording a corresponding magnetization pattern onto a hard magnetic material was previously proposed1 and a particular implementation of a vector-matrix algorithm was demonstrated. In the present article we analyze the conditions allowing for implementation of complex algorithms which can combi

  20. Ashwin Viswanathan Kannan, Goutam Mylavarapu, Johnson P Thomas

    In this study, we build a computational model of Prefrontal Cortex (PFC) using Spiking Neural Networks (SNN) to understand how neurons adapt and respond to tasks switched under short and longer duration of stimulus changes. We also explore behavioral deficits arising out of the PFC lesions by simulating lesioned states in our Spiking architecture model. Alth

  21. Alfonso Amayuelas, Kyle Wong, Liangming Pan, Wenhu Chen

    This paper investigates the capabilities of Large Language Models (LLMs) in the context of understanding their knowledge and uncertainty over questions. Specifically, we focus on addressing known-unknown questions, characterized by high uncertainty due to the absence of definitive answers. To facilitate our study, we collect a new dataset with Known-Unknown

  22. Yen-Ting Lin, Yun-Nung Chen

    We propose LLM-Eval, a unified multi-dimensional automatic evaluation method for open-domain conversations with large language models (LLMs). Existing evaluation methods often rely on human annotations, ground-truth responses, or multiple LLM prompts, which can be expensive and time-consuming. To address these issues, we design a single prompt-based evaluati

  23. Qingyang Wu, Deema Alnuhait, Derek Chen, Zhou Yu

    Traditional end-to-end task-oriented dialogue systems have been built with a modularized design. However, such design often causes misalignment between the agent response and external knowledge, due to inadequate representation of information. Furthermore, its evaluation metrics emphasize assessing the agent's pre-lexicalization response, neglecting the qual

  24. Muhammad Yasir, Xia Tiecheng, Faisal Javed, G. Mustafa

    This study examines a recently hypothesized black hole, which is a perfect solution of metric-affine gravity with a positive cosmological constant, and its thermodynamic features as well as the Joule-Thomson expansion. We develop some thermodynamical quantities, such as volume, Gibbs free energy, and heat capacity, using the entropy and Hawking temperature.

  25. Yoshihiro Mizuta, Tetsu Shimomura

    In this paper, we study Sobolev type inequalities for fractional maximal functions $M_{{\mathbb H},\nu}f$ and Riesz potentials $I_{{\mathbb H},\alpha} f$ of functions in weighted Morrey spaces of the double phase functional $\Phi(x,t) = t^{p} + (b(x) t)^{q}$ in the half space, where $1<p<q$ and $b(\cdot)$ is non-negative, bounded and H\"older continuous of o

  26. Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai

    Language models have graduated from being research prototypes to commercialized products offered as web APIs, and recent works have highlighted the multilingual capabilities of these products. The API vendors charge their users based on usage, more specifically on the number of ``tokens'' processed or generated by the underlying language models. What constit

  27. Jiazheng Chen, Wanchun Liu, Daniel Quevedo, Yonghui Li

    For cyber-physical systems in the 6G era, semantic communications connecting distributed devices for dynamic control and remote state estimation are required to guarantee application-level performance, not merely focus on communication-centric performance. Semantics here is a measure of the usefulness of information transmissions. Semantic-aware transmission

  28. Lijun Li, Li'an Zhuo, Bang Zhang, Liefeng Bo

    Hand mesh reconstruction from the monocular image is a challenging task due to its depth ambiguity and severe occlusion, there remains a non-unique mapping between the monocular image and hand mesh. To address this, we develop DiffHand, the first diffusion-based framework that approaches hand mesh reconstruction as a denoising diffusion process. Our one-stag

  29. Thejan Wijesinghe, Chamath Abeysinghe, Chanuka Wijayakoon, Lahiru Jayathilake

    We develop an automated video colorization framework that minimizes the flickering of colors across frames. If we apply image colorization techniques to successive frames of a video, they treat each frame as a separate colorization task. Thus, they do not necessarily maintain the colors of a scene consistently across subsequent frames. The proposed solution

  30. EunJeong Hwang, Vered Shwartz

    Memes are a widely popular tool for web users to express their thoughts using visual metaphors. Understanding memes requires recognizing and interpreting visual metaphors with respect to the text inside or around the meme, often while employing background knowledge and reasoning abilities. We present the task of meme captioning and release a new dataset, Mem

  31. Kouta Kondou, Shinji Miwa, Daigo Miyajima

    Chirality-induced spin selectivity (CISS) has been extensively studied over the past two decades. While current-induced spin polarization in chiral molecules is widely recognized as the fundamental principle of the CISS, only a few studies have been reported on bias-current-free CISS, where there is no bias electric current in chiral molecules. Recent studie

  32. Chenglong Wang, Jiangyan Yi, Jianhua Tao, Chuyuan Zhang

    Current fake audio detection relies on hand-crafted features, which lose information during extraction. To overcome this, recent studies use direct feature extraction from raw audio signals. For example, RawNet is one of the representative works in end-to-end fake audio detection. However, existing work on RawNet does not optimize the parameters of the Sinc-

  33. Chenglong Wang, Jiangyan Yi, Jianhua Tao, Chuyuan Zhang

    Existing fake audio detection systems perform well in in-domain testing, but still face many challenges in out-of-domain testing. This is due to the mismatch between the training and test data, as well as the poor generalizability of features extracted from limited views. To address this, we propose multi-view features for fake audio detection, which aim to

  34. Peng Zhang, Fa Ge, Yuhong Liu

    Multi-signature aggregates signatures from multiple users on the same message into a joint signature, which is widely applied in blockchain to reduce the percentage of signatures in blocks and improve the throughput of transactions. The $k$-sum attacks are one of the major challenges to design secure multi-signature schemes. In this work, we address $k$-sum

  35. Frederick Riemenschneider, Anette Frank

    Recent advances in NLP have led to the creation of powerful language models for many languages including Ancient Greek and Latin. While prior work on Classical languages unanimously uses BERT, in this work we create four language models for Ancient Greek that vary along two dimensions to study their versatility for tasks of interest for Classical languages:

  36. Hao Yang, Can Gao, Hao Líu, Xinyan Xiao

    Vision-and-language (VL) pre-training, which aims to learn a general representation of image-text pairs that can be transferred to various vision-and-language tasks. Compared with modeling uni-modal data, the main challenge of the VL model is: how to learn the cross-modal interaction from multimodal data, especially the fine-grained interaction. Existing wor

  37. Khang Nhut Lam, Thieu Gia Doan, Khang Thua Pham, Jugal Kalita

    Summary sentences produced by abstractive summarization models may be coherent and comprehensive, but they lack control and rely heavily on reference summaries. The BRIO training paradigm assumes a non-deterministic distribution to reduce the model's dependence on reference summaries, and improve model performance during inference. This paper presents a stra

  38. Samuel Barnier, Pierre-Olivier Petrucci, Jonathan Ferreira, Gregoire Marcel

    The non linear correlation between the UV and X-ray emission observed in Active Galactic Nuclei remains a puzzling question that challenged accretion models. While the UV emission originates from the cold disk, the X-ray emission is emitted by a hot corona whose physical characteristics and geometry are still highly debated. The Jet Emitting Disk - Standard

  39. Chunxiang Wang, Ran Li, Huanyuan Shan, Weiwei Xu

    The galaxy-galaxy lensing technique allows us to measure the subhalo mass of satellite galaxies, studying their mass loss and evolution within galaxy clusters and providing direct observational validation for theories of galaxy formation. In this study, we use the weak gravitational lensing observations from DECaLS DR8, in combination with the redMaPPer gala

  40. Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong

    Evaluating multi-document summarization (MDS) quality is difficult. This is especially true in the case of MDS for biomedical literature reviews, where models must synthesize contradicting evidence reported across different documents. Prior work has shown that rather than performing the task, models may exploit shortcuts that are difficult to detect using st

  41. A. Buzulutskov, E. Frolov, E. Borisova, V. Nosov

    Our recent studies of electroluminescence (EL) properties in two-phase argon detectors for dark matter searches have revealed the presence of unusual delayed pulses in the EL signal in the form of two slow components with time constants of about 5 and 50 $\mu$s. These components were shown to be present in the charge signal itself, which clearly indicates th

  42. Mingda Chen, Xilun Chen, Wen-tau Yih

    Few-shot learning for open domain multi-hop question answering typically relies on the incontext learning capability of large language models (LLMs). While powerful, these LLMs usually contain tens or hundreds of billions of parameters, making them rather inefficient at inference time. To improve performance of smaller language models, we propose a data synt

  43. Yue Feng, Hossein A. Rahmani, Aldo Lipani, Emine Yilmaz

    Task-oriented dialogue systems aim at providing users with task-specific services. Users of such systems often do not know all the information about the task they are trying to accomplish, requiring them to seek information about the task. To provide accurate and personalized task-oriented information seeking results, task-oriented dialogue systems need to a

  44. Charles Lovering, Ellie Pavlick

    Text-to-image models can often generate some relations, i.e., "astronaut riding horse", but fail to generate other relations composed of the same basic parts, i.e., "horse riding astronaut". These failures are often taken as evidence that models rely on training priors rather than constructing novel images compositionally. This paper tests this intuition on

  45. Utku Ozbulak, Hyun Jung Lee, Beril Boga, Esla Timothy Anzaku

    Although supervised learning has been highly successful in improving the state-of-the-art in the domain of image-based computer vision in the past, the margin of improvement has diminished significantly in recent years, indicating that a plateau is in sight. Meanwhile, the use of self-supervised learning (SSL) for the purpose of natural language processing (

  46. Tieying Li, Kan Wu, Xujia Zhang, Minglu Cai

    Different from the Kerr effect,stimulated Raman scattering (SRS) is a delayed response to molecular vibrations in materials. In microcavities, when driven in an anomalous group velocity dispersion (GVD) regime, SRS typically leads to self-frequency shift of solitons and generation of breather solitons which have been verified both theoretically and experimen

  47. Ivan Jeliazkov, Shubham Karnawat, Mohammad Arshad Rahman, Angela Vossmeyer

    This article develops a random effects quantile regression model for panel data that allows for increased distributional flexibility, multivariate heterogeneity, and time-invariant covariates in situations where mean regression may be unsuitable. Our approach is Bayesian and builds upon the generalized asymmetric Laplace distribution to decouple the modeling

  48. Ye-Xin Lu, Yang Ai, Zhen-Hua Ling

    This paper proposes MP-SENet, a novel Speech Enhancement Network which directly denoises Magnitude and Phase spectra in parallel. The proposed MP-SENet adopts a codec architecture in which the encoder and decoder are bridged by convolution-augmented transformers. The encoder aims to encode time-frequency representations from the input noisy magnitude and pha

  49. Jiachang Liu, Qi Zhang, Chongyang Shi, Usman Naseem

    Abstractive related work generation has attracted increasing attention in generating coherent related work that better helps readers grasp the background in the current research. However, most existing abstractive models ignore the inherent causality of related work generation, leading to low quality of generated related work and spurious correlations that a

  50. Peiqin Lin, Chengzhi Hu, Zheyu Zhang, André F. T. Martins

    Recent multilingual pretrained language models (mPLMs) have been shown to encode strong language-specific signals, which are not explicitly provided during pretraining. It remains an open question whether it is feasible to employ mPLMs to measure language similarity, and subsequently use the similarity results to select source languages for boosting cross-li

  51. Shijie Chen, Ziru Chen, Huan Sun, Yu Su

    Despite remarkable progress in text-to-SQL semantic parsing in recent years, the performance of existing parsers is still far from perfect. Specifically, modern text-to-SQL parsers based on deep learning are often over-confident, thus casting doubt on their trustworthiness when deployed for real use. In this paper, we propose a parser-independent error detec

  52. Maximilian J. Schilcher, David J. Abramovitch, Matthew Z. Mayers, Liang Z. Tan

    Halide pervoskites are an important class of semiconducting materials which hold great promise for optoelectronic applications. In this work we investigate the relationship between vibrational anharmonicity and dynamic disorder in this class of solids. Via a multi-scale model parameterized from first-principles calculations, we demonstrate that the non-Gauss

  53. Weiye Zhao, Yifan Sun, Feihan Li, Rui Chen

    Due to the trial-and-error nature, it is typically challenging to apply RL algorithms to safety-critical real-world applications, such as autonomous driving, human-robot interaction, robot manipulation, etc, where such errors are not tolerable. Recently, safe RL (i.e. constrained RL) has emerged rapidly in the literature, in which the agents explore the envi

  54. Eng Lieh Ouh, Benjamin Kok Siew Gan, Kyong Jin Shim, Swavek Wlodkowski

    In this study, we assess the efficacy of employing the ChatGPT language model to generate solutions for coding exercises within an undergraduate Java programming course. ChatGPT, a large-scale, deep learning-driven natural language processing model, is capable of producing programming code based on textual input. Our evaluation involves analyzing ChatGPT-gen

  55. M. Grzeszczyk, K. Vaklinova, K. Watanabe, T. Taniguchi

    Defect centers in wide-band-gap crystals attracted considerable attention due to the realisations of qubits, sensors, or single photon emitters at room temperature. The family of these centers is constantly growing, including well-known examples such as nitrogen-vacancy centers in diamond, silicon-vacancy in silicon carbide, chromium substitutions in alumini

  56. Minchan Kwon, Kangil Kim

    In real life, adversarial attack to deep learning models is a fatal security issue. However, the issue has been rarely discussed in a widely used class-incremental continual learning (CICL). In this paper, we address problems of applying adversarial training to CICL, which is well-known defense method against adversarial attack. A well-known problem of CICL

  57. Chu Fei Luo, Rohan Bhambhoria, Xiaodan Zhu, Samuel Dahan

    Hate speech causes widespread and deep-seated societal issues. Proper enforcement of hate speech laws is key for protecting groups of people against harmful and discriminatory language. However, determining what constitutes hate speech is a complex task that is highly open to subjective interpretations. Existing works do not align their systems with enforcea

  58. Hannah L. Weaver, Cora M. Went, Joeson Wong, Dipti Jasrasaria

    Since dissipative processes are ubiquitous in semiconductors, characterizing how electronic and thermal energy transduce and transport at the nanoscale is vital for understanding and leveraging their fundamental properties. For example, in low-dimensional transition metal dichalcogenides (TMDCs), excess heat generation upon photoexcitation is difficult to av

  59. Tim Schott, Daniel Furman, Shreshta Bhat

    In this work, we assess the ability of foundation models to recall encyclopedic knowledge across a wide range of linguistic contexts. To support this, we: 1) produce a 20-language dataset that contains 303k factual associations paired with counterfactuals, 2) evaluate 5 models in a multilingual test, and 3) benchmark a diverse set of 24 models in an English-

  60. Christopher G. Bailey, Lara V. Gillan, Minwoo Lee, Nicholas Sloane

    The organic spacer cation plays a crucial role in determining the exciton fine structure in two-dimensional (2D) perovskites. Here, we use low-temperature magneto-optical spectroscopy to gain insight into the influence of the organic spacer on dark excitons in Ruddlesden-Popper (RP) perovskites. We show that by using modest magnetic field strengths (<1.5 T),

  61. Zeyuan Allen-Zhu, Yuanzhi Li

    Transformer-based language models are effective but complex, and understanding their inner workings and reasoning mechanisms is a significant challenge. Previous research has primarily explored how these models handle simple tasks like name copying or selection, and we extend this by investigating how these models perform recursive language structure reasoni

  62. Elahe Vedadi, Joshua V. Dillon, Philip Andrew Mansfield, Karan Singhal

    Conventional federated learning algorithms train a single global model by leveraging all participating clients' data. However, due to heterogeneity in client generative distributions and predictive models, these approaches may not appropriately approximate the predictive process, converge to an optimal state, or generalize to new clients. We study personaliz

  63. Jun Tong, Bing Duan, Yanfeng Luo

    In this paper, we introduce a combinatorial path model of representation of the quantum affine algebra of type $D_n$, inspired by Mukhin and Young's combinatorial path models of representations of the quantum affine algebras of types $A_n$ and $B_n$. In particular, we give a combinatorial formula for $q$-characters of fundamental modules of type $D_{n}$ by a

  64. Ping Li, Xiao Yang, Qing-Song Jiang, Yin-Zhong Wu

    The valley-related multiple topological phase transitions attracted significant attention due to their providing significant opportunities for fundamental research and practical applications. However, unfortunately, to date there is no real material that can realize valley-related multiple topological phase transitions. Here, through first-principles calcula

  65. Shuo Zhang, Liangming Pan, Junzhou Zhao, William Yang Wang

    Large language models often necessitate grounding on external knowledge to generate faithful and reliable answers. Yet even with the correct groundings in the reference, they can ignore them and rely on wrong groundings or their inherent biases to hallucinate when users, being largely unaware of the specifics of the stored information, pose questions that mi

  66. Sadaf Ghaffari, Nikhil Krishnaswamy

    We present a novel method for using agent experiences gathered through an embodied simulation to ground contextualized word vectors to object representations. We use similarity learning to make comparisons between different object types based on their properties when interacted with, and to extract common features pertaining to the objects' behavior. We then

  67. Chenxin An, Jiangtao Feng, Fei Huang, Xipeng Qiu

    Non-autoregressive Transformers (NATs) reduce the inference latency of Autoregressive Transformers (ATs) by predicting words all at once rather than in sequential order. They have achieved remarkable progress in machine translation as well as many other applications. However, a long-standing challenge for NATs is the learning of multi-modality data distribut

  68. Wanli Xing, Xuan-Gong Wang, Anthony W. Thomas

    Recent experimental studies have led to the suggestion that short-range correlations may be a major contributor to the nuclear EMC effect. This hypothesis requires that the structure function for nucleons involved in short-range correlations should be heavily suppressed compared to that of a free nucleon. Based on calculations performed within an AdS/QCD mot

  69. Linwei Tao, Minjing Dong, Chang Xu

    The use of deep neural networks in real-world applications require well-calibrated networks with confidence scores that accurately reflect the actual probability. However, it has been found that these networks often provide over-confident predictions, which leads to poor calibration. Recent efforts have sought to address this issue by focal loss to reduce ov

  70. Achraf Bahamou, Donald Goldfarb

    We propose a new per-layer adaptive step-size procedure for stochastic first-order optimization methods for minimizing empirical loss functions in deep learning, eliminating the need for the user to tune the learning rate (LR). The proposed approach exploits the layer-wise stochastic curvature information contained in the diagonal blocks of the Hessian in de

  71. Lexiong Huang, Ruihua Han, Guoliang Li, He Li

    Autonomous parking (AP) is an emering technique to navigate an intelligent vehicle to a parking space without any human intervention. Existing AP methods based on mathematical optimization or machine learning may lead to potential failures due to either excessive execution time or lack of generalization. To fill this gap, this paper proposes an integrated co

  72. Eng Lieh Ouh, Benjamin Kok Siew Gan

    Cloud Computing skills have been increasing in demand. Many software engineers are learning these skills and taking cloud certification examinations to be job competitive. Preparing undergraduates to be cloud-certified remains challenging as cloud computing is a relatively new topic in the computing curriculum, and many of these certifications require workin

  73. Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov

    In this paper, we comprehensively investigate the potential misuse of modern Large Language Models (LLMs) for generating credible-sounding misinformation and its subsequent impact on information-intensive applications, particularly Open-Domain Question Answering (ODQA) systems. We establish a threat model and simulate potential misuse scenarios, both uninten

  74. Xiao Yu, Maximillian Chen, Zhou Yu

    Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress. Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS). However, this training often requires abundant annotated data, which creates challenges wh

  75. F. J. Vaquero-Caballero, Gernot Goeger, Fabio Pittala, Yabin Ye

    In this letter we experimentally validate, for the first time, the Stokes space algorithm (SSA) equalizer for space division multiplexing (SDM) transmission systems. We introduce the frequency domain (FD)-SSA and FD least-mean square (FD-LMS) algorithms, and evaluate their performance for different frequency offsets by computer simulations. Our simulations s

  76. Aihua Zheng, Zhiqi Ma, Zi Wang, Chenglong Li

    Multi-spectral vehicle re-identification aims to address the challenge of identifying vehicles in complex lighting conditions by incorporating complementary visible and infrared information. However, in harsh environments, the discriminative cues in RGB and NIR modalities are often lost due to strong flares from vehicle lamps or sunlight, and existing multi-

  77. Farhan Samir, Miikka Silfverberg

    Data augmentation techniques are widely used in low-resource automatic morphological inflection to overcome data sparsity. However, the full implications of these techniques remain poorly understood. In this study, we aim to shed light on the theoretical aspects of the prominent data augmentation strategy StemCorrupt (Silfverberg et al., 2017; Anastasopoulos

  78. Md Mahadi Hassan, Alex Knipper, Shubhra Kanti Karmaker Santu

    The rise of big data has amplified the need for efficient, user-friendly automated machine learning (AutoML) tools. However, the intricacy of understanding domain-specific data and defining prediction tasks necessitates human intervention making the process time-consuming while preventing full automation. Instead, envision an intelligent agent capable of ass

  79. Zexi Huang, Mert Kosan, Arlei Silva, Ambuj Singh

    Link prediction, which consists of predicting edges based on graph features, is a fundamental task in many graph applications. As for several related problems, Graph Neural Networks (GNNs), which are based on an attribute-centric message-passing paradigm, have become the predominant framework for link prediction. GNNs have consistently outperformed tradition

  80. Long Lian, Boyi Li, Adam Yala, Trevor Darrell

    Recent advancements in text-to-image diffusion models have yielded impressive results in generating realistic and diverse images. However, these models still struggle with complex prompts, such as those that involve numeracy and spatial reasoning. This work proposes to enhance prompt understanding capabilities in diffusion models. Our method leverages a pret

  81. Oscar Chew, Hsuan-Tien Lin, Kai-Wei Chang, Kuan-Hao Huang

    Recent research has revealed that machine learning models have a tendency to leverage spurious correlations that exist in the training set but may not hold true in general circumstances. For instance, a sentiment classifier may erroneously learn that the token "performances" is commonly associated with positive movie reviews. Relying on these spurious correl

  82. Yang Bai, Min Cao, Daming Gao, Ziqiang Cao

    Text-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivit

  83. Z. Li, Z. Q. Zhao, X. H. Yang, G. B. Zhang

    Quasi-isentropic compression is an effective method to achieve high-density and high-temperature implosion in laser-driven inertial confinement fusion (ICF). However, it requires precise matching between the laser profile and the target structure. Designing the optimal laser profile and the corresponding target for ICF is a challenge due to the large number

  84. Ying Huang, Liang Chen

    In this note, we consider the problem of robust learning mixtures of linear regressions. We connect mixtures of linear regressions and mixtures of Gaussians with a simple thresholding, so that a quasi-polynomial time algorithm can be obtained under some mild separation condition. This algorithm has significantly better robustness than the previous result.

  85. Jan Silovsky, Liuhui Deng, Arturo Argueta, Tresi Arvizo

    Voice technology has become ubiquitous recently. However, the accuracy, and hence experience, in different languages varies significantly, which makes the technology not equally inclusive. The availability of data for different languages is one of the key factors affecting accuracy, especially in training of all-neural end-to-end automatic speech recognition

  86. Zhiyi Dong, Yongyi Mao

    Adversarial attacks pose significant challenges to the robustness of modern deep neural networks in computer vision, and defending these networks against adversarial attacks has attracted intense research efforts. Among various defense strategies, preprocessing-based defenses are practically appealing since there is no need to train the network under protect

  87. Saba Ghaffari, Ehsan Saleh, Alexander G. Schwing, Yu-Xiong Wang

    Protein design, a grand challenge of the day, involves optimization on a fitness landscape, and leading methods adopt a model-based approach where a model is trained on a training set (protein sequences and fitness) and proposes candidates to explore next. These methods are challenged by sparsity of high-fitness samples in the training set, a problem that ha

  88. Jacob Sillman

    An analog implementation of the Softmax activation function is presented. A modular design is proposed, scaling linearly with the number of inputs and outputs. The circuit behaves similarly using both a BJT and NMOS design scheme. Experimental results extracted from a BJT breadboard prototype presents computational accuracy within 4.2% margin of error. Simul

  89. Jiayi Wang, Ke Wang, Yuqi Zhang, Yu Zhao

    Non-parametric, k-nearest-neighbor algorithms have recently made inroads to assist generative models such as language models and machine translation decoders. We explore whether such non-parametric models can improve machine translation models at the fine-tuning stage by incorporating statistics from the kNN predictions to inform the gradient updates for a b

  90. Zhixuan Zhang, Yuheng Huang, Dan Ou, Sen Li

    E-commerce search systems such as Taobao Search, the largest e-commerce searching system in China, aim at providing users with the most preferred items (e.g., products). Due to the massive data and limited time for response, a typical industrial ranking system consists of three or more modules, including matching, pre-ranking, and ranking. The pre-ranking is

  91. Sinan Rasiya Koya, Kanak Kanti Kar, Shivendra Srivastava, Tsegaye Tadesse

    In several regions across the globe, snow has a significant impact on hydrology. The amounts of water that infiltrate the ground and flow as runoff are driven by the melting of snow. Therefore, it is crucial to study the magnitude and effect of snowmelt. Snow droughts, resulting from reduced snow storage, can drastically impact the water supplies in basins w

  92. Weiwen Xu, Xin Li, Wai Lam, Lidong Bing

    We present multilingual Pre-trained Machine Reader (mPMR), a novel method for multilingual machine reading comprehension (MRC)-style pre-training. mPMR aims to guide multilingual pre-trained language models (mPLMs) to perform natural language understanding (NLU) including both sequence classification and span extraction in multiple languages. To achieve cros

  93. Jing Wang, Hairun Xie, Miao Zhang, Hui Xu

    Transonic buffet is a flow instability phenomenon that arises from the interaction between the shock wave and the separated boundary layer. This flow phenomenon is considered to be highly detrimental during flight and poses a significant risk to the structural strength and fatigue life of aircraft. Up to now, there has been a lack of an accurate, efficient,

  94. Jacob Sillman, Ajay Suresh

    This report investigates the potential impact of a Trojan attack on power conversion circuits, specifically a switching signal attack designed to trigger a locking of the pulse width modulation (PWM) signal that goes to a power field-effect transistor (FET). The first simulation shows that this type of attack can cause severe overvoltage, potentially leading

  95. Wadim Gerner

    In the present work we present a general framework which guarantees the existence of optimal domains for isoperimetric problems within the class of $C^{1,1}$-regular domains satisfying a uniform ball condition as long as the desired objective function satisfies certain properties. We then verify that the helicity isoperimetric problem studied in [J. Cantarel

  96. Jiadong Dan, Moaz Waqar, Ivan Erofeev, Kui Yao

    A continuing challenge in atomic resolution microscopy is to identify significant structural motifs and their assembly rules in synthesized materials with limited observations. Here we propose and validate a simple and effective hybrid generative model capable of predicting unseen domain boundaries in a potassium sodium niobate thin film from only a small nu

  97. Vanessa Liao, Syed Shariyar Murtaza, Yifan Nie, Jimmy Lin

    A common way to use large pre-trained language models for downstream tasks is to fine tune them using additional layers. This may not work well if downstream domain is a specialized domain whereas the large language model has been pre-trained on a generic corpus. In this paper, we discuss the use of regular expression patterns employed as features for domain

  98. Abhijnan Nath, Sheikh Mannan, Nikhil Krishnaswamy

    Despite their successes in NLP, Transformer-based language models still require extensive computing resources and suffer in low-resource or low-compute settings. In this paper, we present AxomiyaBERTa, a novel BERT model for Assamese, a morphologically-rich low-resource language (LRL) of Eastern India. AxomiyaBERTa is trained only on the masked language mode

  99. Mitsuhiro Nishijima

    We consider a wide class of closed convex cones $K$ in the space of real $n\times n$ symmetric matrices and establish the existence of a chain of faces of $K$, the length of which is maximized at $\frac{n(n+1)}{2} + 1$. Examples of such cones include, but are not limited to, the completely positive and the copositive cones. Using this chain, we prove that th

  100. Yuta Kambe

    The signatures of polynomials were originally introduced by Faug\`{e}re for the efficient computation of Gr\"obner bases [Fau02], and redefined by Arri-Perry [AP11] as the standard monomials modulo the module of syzygies. Since it is difficult to determine signatures, Vaccon-Yokoyama [VY17] introduced an alternative object called guessed signatures. In this