Skip to content

May 2023 arXiv papers — page 47

Showing 4,6014,700 of 19,695 papers

  1. Wenhao Cheng, Junbo Yin, Wei Li, Ruigang Yang

    This paper addresses the problem of 3D referring expression comprehension (REC) in autonomous driving scenario, which aims to ground a natural language to the targeted region in LiDAR point clouds. Previous approaches for REC usually focus on the 2D or 3D-indoor domain, which is not suitable for accurately predicting the location of the queried 3D region in

  2. Aihua Zheng, Chaobin Zhang, Weijun Zhang, Chenglong Li

    Existing vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance of vehicle Re-ID in complicated surveillance networks. In this paper, we propose a more realistic and easily accessible task, called multi-query vehicle Re-ID, which leverages mult

  3. Tai Kai Ng

    Motivated by the discovery of the anomalous metal state in thin film systems and suggestions that coexistence of superconducting and metallic components is crucial to the formation of the state, we study in this paper a model of mixed metallic and superconducting grains coupled by electron tunneling - the metallic grains are expected to become superconductin

  4. Aihua Zheng, Ziling He, Zi Wang, Chenglong Li

    Many existing multi-modality studies are based on the assumption of modality integrity. However, the problem of missing arbitrary modalities is very common in real life, and this problem is less studied, but actually important in the task of multi-modality person re-identification (Re-ID). To this end, we design a novel dynamic enhancement network (DENet), w

  5. Sipu Ruan, Weixiao Liu, Xiaoli Wang, Xin Meng

    This paper proposes a learning-from-demonstration method using probability densities on the workspaces of robot manipulators. The method, named "PRobabilistically-Informed Motion Primitives (PRIMP)", learns the probability distribution of the end effector trajectories in the 6D workspace that includes both positions and orientations. It is able to adapt to n

  6. Tahir Javed, Sakshi Joshi, Vignesh Nagarajan, Sai Sundaresan

    India is the second largest English-speaking country in the world with a speaker base of roughly 130 million. Thus, it is imperative that automatic speech recognition (ASR) systems for English should be evaluated on Indian accents. Unfortunately, Indian speakers find a very poor representation in existing English ASR benchmarks such as LibriSpeech, Switchboa

  7. Michael F. Liu, Saiyue Lyu, Margarita Vinaroz, Mijung Park

    Diffusion models (DMs) are one of the most widely used generative models for producing high quality images. However, a flurry of recent papers points out that DMs are least private forms of image generators, by extracting a significant number of near-identical replicas of training images from DMs. Existing privacy-enhancing techniques for DMs, unfortunately,

  8. Rawad Melhem, Assef Jafar, Oumayma Al Dakkak

    Speech separation is very important in real-world applications such as human-machine interaction, hearing aids devices, and automatic meeting transcription. In recent years, a significant improvement occurred towards the solution based on deep learning. In fact, much attention has been drawn to supervised learning methods using synthetic mixtures datasets de

  9. Zi Liang, Pinghui Wang, Ruofei Zhang, Shuo Zhang

    Recent years have seen increasing concerns about the unsafe response generation of large-scale dialogue systems, where agents will learn offensive or biased behaviors from the real-world corpus. Some methods are proposed to address the above issue by detecting and replacing unsafe training examples in a pipeline style. Though effective, they suffer from a hi

  10. Zhiming Mao, Huimin Wang, Yiming Du, Kam-fai Wong

    Prior study has shown that pretrained language models (PLM) can boost the performance of text-based recommendation. In contrast to previous works that either use PLM to encode user history as a whole input text, or impose an additional aggregation network to fuse multi-turn history representations, we propose a unified local- and global-attention Transformer

  11. Arunava Naha, Subhrakanti Dey

    This paper studies a deep deterministic policy gradient (DDPG) based actor critic (AC) reinforcement learning (RL) technique to control a linear discrete-time system with a quadratic control cost while ensuring a constraint on the probability of potentially risky or undesirable events. The proposed methodology can be applied to both known and unknown system

  12. Tomoya Wakayama, Masaaki Imaizumi

    In high-dimensional Bayesian statistics, various methods have been developed, including prior distributions that induce parameter sparsity to handle many parameters. Yet, these approaches often overlook the rich spectral structure of the covariate matrix, which can be crucial when true signals are not sparse. To address this gap, we introduce a data-adaptive

  13. Weizhi Nie, Ruidong Chen, Weijie Wang, Bruno Lepri

    In recent years, 3D models have been utilized in many applications, such as auto-driver, 3D reconstruction, VR, and AR. However, the scarcity of 3D model data does not meet its practical demands. Thus, generating high-quality 3D models efficiently from textual descriptions is a promising but challenging way to solve this problem. In this paper, inspired by t

  14. R. Casadio, R. da Rocha

    The minimal geometric deformation (MGD) paradigm is here employed to survey axion stars on fluid branes. The finite value of the brane tension provides beyond-general relativity corrections to the density, compactness, radius, and asymptotic limit of the gravitational mass function of axion stars, in a MGD background. The brane tension also enhances the effe

  15. Lu Wang, Xiang Lyu, Zhengwu Zhang, Lexin Li

    There is increasing interest in modeling high-dimensional longitudinal outcomes in applications such as developmental neuroimaging research. Growth curve model offers a useful tool to capture both the mean growth pattern across individuals, as well as the dynamic changes of outcomes over time within each individual. However, when the number of outcomes is la

  16. Liheng Bian, Daoyu Li, Shuoguang Wang, Chunyang Teng

    Millimeter-wave (MMW) imaging is emerging as a promising technique for safe security inspection. It achieves a delicate balance between imaging resolution, penetrability and human safety, resulting in higher resolution compared to low-frequency microwave, stronger penetrability compared to visible light, and stronger safety compared to X ray. Despite of rece

  17. Rustem Yeshpanov, Saida Mussakhojayeva, Yerbolat Khassanov

    This work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek. We specifically target the zero-shot learning scenario, where a TTS model trained using the data of one language is applied to synthesise speech for oth

  18. Cheng Luo, Siyang Song, Weicheng Xie, Micol Spitale

    In dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an interpolation or fitting problem, emphasizing deterministic outcomes but ignoring the diversity and uncertainty of human facial reac

  19. Jiaxing Xu, Aihu Zhang, Qingtian Bian, Vijay Prakash Dwivedi

    Graph Neural Networks (GNNs) are widely used for graph representation learning in many application domains. The expressiveness of vanilla GNNs is upper-bounded by 1-dimensional Weisfeiler-Leman (1-WL) test as they operate on rooted subtrees through iterative message passing. In this paper, we empower GNNs by injecting neighbor-connectivity information extrac

  20. wala Draidi Areed, Aiden Price, Kathryn Arnett, Helen Thompson

    The research explores the influence of preschool attendance (one year before full-time school) on the development of children during their first year of school. Using data collected by the Australian Early Development Census, the findings show that areas with high proportions of preschool attendance tended to have lower proportions of children with at least

  21. Kha-Dinh Luong, Mert Kosan, Arlei Lopes Da Silva, Ambuj Singh

    Explaining the decisions made by machine learning models for high-stakes applications is critical for increasing transparency and guiding improvements to these decisions. This is particularly true in the case of models for graphs, where decisions often depend on complex patterns combining rich structural and attribute data. While recent work has focused on d

  22. Meseret Asrat

    In this paper, we show that for one sign of the deformation coupling single-trace $T{\bar T}$ deformation moves the holographic screen in G\"{o}del universe radially inward. For the other sign of the coupling it moves the holographic screen radially outward. We (thus) argue, on general grounds, that in holography (single-trace) $T{\bar T}$ deformation can be

  23. Ding Wang, Xuhong Wang, Liang Chen, Shengyue Yao

    Traffic simulation is a crucial tool for transportation decision-making and policy development. However, achieving realistic simulations in the face of the high dimensionality and heterogeneity of traffic environments is a longstanding challenge. In this paper, we present TransWordNG, a traffic simulator that uses Data-driven algorithms and Graph Computing t

  24. Shenghao Wu, Wenbin Zhou, Minshuo Chen, Shixiang Zhu

    Estimating the counterfactual outcome of treatment is essential for decision-making in public health and clinical science, among others. Often, treatments are administered in a sequential, time-varying manner, leading to an exponentially increased number of possible counterfactual outcomes. Furthermore, in modern applications, the outcomes are high-dimension

  25. C. L. Liu, C. P. Sun

    We study the task of coherence filtration under strictly incoherent operations in this paper. The aim of this task is to transform a given state $\rho$ into another one $\rho^\prime$ whose fidelity with the maximally coherent state is maximal by using stochastic strictly incoherent operations. We find that the maximal fidelity between $\rho^\prime$ and the m

  26. Gwantae Kim, Seonghyeok Noh, Insung Ham, Hanseok Ko

    When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and gene

  27. Nikolaj Glazunov

    We investigate lattice packings of Minkowski's balls and domains, as well as the distribution of lattice points on Minkowski's curves which are boundaries of Minkowski's balls. By results of the proof of Minkowski's conjecture about the critical determinant we devide the balls and domains on 3 classes: Minkowski, Davis and Chebyshev-Cohn balls. The optimal l

  28. Peter Gartland, Daniel Lokshtanov, Tomáš Masařík, Marcin Pilipczuk

    We show that the Maximum Weight Independent Set problem (MWIS) can be solved in quasi-polynomial time on $H$-free graphs (graphs excluding a fixed graph $H$ as an induced subgraph) for every $H$ whose every connected component is a path or a subdivided claw (i.e., a tree with at most three leaves). This completes the dichotomy of the complexity of MWIS in $\

  29. Feng Lin, Valery Levitas, Krishan Pandey, Sorb Yesudhas

    Rough diamond anvils (rough-DA) are introduced to intensify all occurring processes during an in-situ study of heterogeneous compression of strongly pre-deformed Zr in diamond anvil cell (DAC). Crystallite size and dislocation density of Zr are getting pressure-, plastic strain tensor- and strain-path-independent during {\alpha}-{\omega} phase transformation

  30. Arka Banerjee, Subinoy Das, Anshuman Maharana, Ethan O. Nadler

    We present small-scale structure constraints on sterile dark matter produced from a heavy mediator particle, inspired by models of moduli decay. Dark matter particles produced through this mechanism can contribute to the entire dark matter energy density but the particles have a non-thermal phase-space distribution; however, we show that the resulting linear

  31. Yun Zhu, Kangkang Zhang, Yuncai Zhu, Jinming Zhou

    Most MPC (Model Predictive Control) algorithms used in industries and studied in the control academia use a two-term QP (quadratic programming), where the first term is the weighted norm of the output errors, and the second term is that of the input increments. In this work, a DMC (Dynamic Matrix Control) algorithm that uses three-term QP is studied, where t

  32. Hyeongrok Han, Siwon Kim, Hyun-Soo Choi, Sungroh Yoon

    Several recent studies have elucidated why knowledge distillation (KD) improves model performance. However, few have researched the other advantages of KD in addition to its improving model performance. In this study, we have attempted to show that KD enhances the interpretability as well as the accuracy of models. We measured the number of concept detectors

  33. Imra Aqeel, Abdul Majid

    The COVID-19 pandemic has initiated a global health emergency, with an exigent need for effective cure. Progressively, drug repurposing is emerging a promise solution as it saves the time, cost and labor. However, the number of drug candidates that have been identified as being repurposed for the treatment of COVID-19 are still insufficient, so more effectiv

  34. Hao-Jie Lin, Tao Zhu, Shao-Jun Zhang, Anzhong Wang

    It is well-known that parity symmetry is broken in the weak interaction but conserved for Einstein's general relativity and Maxwell's electromagnetic theory. Nevertheless, parity symmetry could also be violated in the gravitational/electromagnetic sectors if a fundamental scalar field couples to the parity-violating gravitational/electromagnetic curvature te

  35. Yu-Chen Wang, Zhen-Zhao Tao, Zhi-Song Zhang, Cheqiu Lyu

    The "search for extraterrestrial intelligence" (SETI) commensal surveys aim to scan the sky to find possible technosignatures from the extraterrestrial intelligence (ETI). The mitigation of radio frequency interference (RFI) is an important step, especially for the most sensitive Five-hundred-meter Aperture Spherical radio Telescope (FAST), which can detect

  36. Ming Gao, YanWu Xu, Yang Zhao, Tingbo Hou

    In this paper, we propose a novel language-guided 3D arbitrary neural style transfer method (CLIP3Dstyler). We aim at stylizing any 3D scene with an arbitrary style from a text description, and synthesizing the novel stylized view, which is more flexible than the image-conditioned style transfer. Compared with the previous 2D method CLIPStyler, we are able t

  37. Jiancheng An, Chau Yuen, Chongwen Huang, Merouane Debbah

    Holographic multiple-input multiple-output (HMIMO) technology, which uses spatially continuous surfaces for signal transmission and reception, is envisioned to be a promising solution for improving the data rate and coverage of wireless networks. In Parts I and II of this three-part tutorial on HMIMO communications, we provided an overview of channel modelin

  38. Jiancheng An, Chau Yuen, Chongwen Huang, Merouane Debbah

    As Part II of a three-part tutorial on holographic multiple-input multiple-output (HMIMO), this Letter focuses on the state-of-the-art in performance analysis and on holographic beamforming for HMIMO communications. We commence by discussing the spatial degrees of freedom (DoF) and ergodic capacity of a point-to-point HMIMO system, based on the channel model

  39. Junfeng Chen, Zili Tang, Meng Guo

    Coalition is an important mean of multi-robot systems to collaborate on common tasks. An adaptive coalition strategy is essential for the online performance in dynamic and unknown environments. In this work, the problem of territory defense by large-scale heterogeneous robotic teams is considered. The tasks include exploration, capture of dynamic targets, an

  40. Jiancheng An, Chau Yuen, Chongwen Huang, Merouane Debbah

    By integrating a nearly infinite number of reconfigurable elements into a finite space, a spatially continuous array aperture is formed for holographic multiple-input multiple-output (HMIMO) communications. This three-part tutorial aims for providing an overview of the latest advances in HMIMO communications. As Part I of the tutorial, this letter first intr

  41. Zhiwen Fan, Panwang Pan, Peihao Wang, Yifan Jiang

    Despite the significant progress in six degrees-of-freedom (6DoF) object pose estimation, existing methods have limited applicability in real-world scenarios involving embodied agents and downstream 3D vision tasks. These limitations mainly come from the necessity of 3D models, closed-category detection, and a large number of densely annotated support views.

  42. Mitsuharu Uemoto, Masaki Nishiura, Tomoya Ono

    Valleytronics, which makes use of the two valleys in graphenes, attracts considerable attention and a valley filter is expected to be the central component in valleytronics. We propose the application of the graphene valley filter using blister defects to the investigation of the valley-dependent transport properties of the Stone--Wales and blister defects o

  43. Fangwei Zhu, Jifan Yu, Hailong Jin, Juanzi Li

    Entity linking models have achieved significant success via utilizing pretrained language models to capture semantic features. However, the NIL prediction problem, which aims to identify mentions without a corresponding entity in the knowledge base, has received insufficient attention. We categorize mentions linking to NIL into Missing Entity and Non-Entity

  44. Sailesh Ranjan Mohanty, Sayantan Ghosh, Pinku Routaray, H. C. Das

    This study presents a universal relation for anisotropic neutron stars, called the $I-f-C$ relation, which accounts for the local anisotropic pressure using the Quasi-Local (QL) Model proposed by Horvat et al. \cite{QL_Model} to describe the anisotropy inside the neutron star. This study analyzes approximately 60 unified tabulated EoS-ensembles, spanning fro

  45. Yangsibo Huang, Haotian Jiang, Daogao Liu, Mohammad Mahdian

    In this paper, we study the setting in which data owners train machine learning models collaboratively under a privacy notion called joint differential privacy [Kearns et al., 2018]. In this setting, the model trained for each data owner $j$ uses $j$'s data without privacy consideration and other owners' data with differential privacy guarantees. This settin

  46. Aryan Patil, Varad Patwardhan, Abhishek Phaltankar, Gauri Takawane

    The term "Code Mixed" refers to the use of more than one language in the same text. This phenomenon is predominantly observed on social media platforms, with an increasing amount of adaptation as time goes on. It is critical to detect foreign elements in a language and process them correctly, as a considerable number of individuals are using code-mixed langu

  47. Ritesh Goenka, Pardis Semnani, Chi Hoi Yip

    We show that there are $O(n \cdot 4^{n/11})$ planar graphs on $n$ vertices which do not admit a simultaneous straight-line embedding on any $n$-point set in the plane. In particular, this improves the best known bound $O(n!)$ significantly.

  48. George Zerveas, Navid Rekabsaz, Carsten Eickhoff

    Sparse annotation poses persistent challenges to training dense retrieval models; for example, it distorts the training signal when unlabeled relevant documents are used spuriously as negatives in contrastive learning. To alleviate this problem, we introduce evidence-based label smoothing, a novel, computationally efficient method that prevents penalizing th

  49. Max W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin

    Recent progress in music generation has been remarkably advanced by the state-of-the-art MusicLM, which comprises a hierarchy of three LMs, respectively, for semantic, coarse acoustic, and fine acoustic modelings. Yet, sampling with the MusicLM requires processing through these LMs one by one to obtain the fine-grained acoustic tokens, making it computationa

  50. Yichong Huang, Xiaocheng Feng, Xinwei Geng, Baohang Li

    Multilingual neural machine translation has witnessed remarkable progress in recent years. However, the long-tailed distribution of multilingual corpora poses a challenge of Pareto optimization, i.e., optimizing for some languages may come at the cost of degrading the performance of others. Existing balancing training strategies are equivalent to a series of

  51. Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng

    An emerging method to cheaply improve a weaker language model is to finetune it on outputs from a stronger model, such as a proprietary system like ChatGPT (e.g., Alpaca, Self-Instruct, and others). This approach looks to cheaply imitate the proprietary model's capabilities using a weaker open-source model. In this work, we critically analyze this approach.

  52. Sota Inoue, Yasuyuki Kimura, Yuki Uematsu

    Bubble solutions are of growing interest because of various technological applications in surface cleaning, water treatment, and agriculture. However, their physicochemical properties such as the stability and interfacial charge of bubbles are not fully understood yet. In this study, the kinetics of radii in aqueous microbubble solutions are experimentally i

  53. Jian-Kang Li, Yu Chen, Zhen-Zhao Tao, Xiao-Hang Luan

    In this paper, we propose a novel method for distinguishing extraterrestrial intelligence (ETI) signals from radio frequency interference (RFI) by leveraging polarization features. We exploit the sinusoidal variation of the linearly polarized components of Stokes parameters with the parallactic angle as a characteristic signature of ETI signals, while such l

  54. Yuki Uematsu

    The nonlinear electrokinetic response of ionic solutions is important in nanofluidics. However, quantitatively understanding the mechanisms is still a challenging problem because of a lack of analytic approaches. Here, a general framework for calculating the nonlinear electrokinetic coefficients of strongly confined electrolytes is constructed using a pertur

  55. Yufeng Wang, Ripeng Luo, Jian Chen, Xuefeng Zhou

    Proton tunneling is believed to be non-local in ice but has never been shown experimentally. Here we measured thermal conductivity of ice under pressure up to 50 GPa and found it to increase with pressure until 20 GPa but decrease at higher pressures. We attribute this anomalous drop of thermal conductivity to the collective tunneling of protons at high pres

  56. Tao Huang, Yuan Zhang, Mingkai Zheng, Shan You

    The representation gap between teacher and student is an emerging topic in knowledge distillation (KD). To reduce the gap and improve the performance, current methods often resort to complicated training schemes, loss functions, and feature alignments, which are task-specific and feature-specific. In this paper, we state that the essence of these methods is

  57. Shota Goto, Kang Kim, Nobuyuki Matubayasi

    Ring polymers are an intriguing class of polymers with unique physical properties, and understanding their behavior is important for developing accurate theoretical models. In this study, we investigate the effect of chain stiffness and monomer density on static and dynamic behaviors of ring polymer melts using molecular dynamics simulations. Our first focus

  58. Linfeng Liang, Yao Deng, Yang Zhang, Jianchao Lu

    Discrepancies in decision-making between Autonomous Driving Systems (ADS) and human drivers underscore the need for intuitive human gaze predictors to bridge this gap, thereby improving user trust and experience. Existing gaze datasets, despite their value, suffer from noise that hampers effective training. Furthermore, current gaze prediction models exhibit

  59. Xianghao Jiao, Yaohua Liu, Jiaxin Gao, Xinyuan Chu

    In light of the significant progress made in the development and application of semantic segmentation tasks, there has been increasing attention towards improving the robustness of segmentation models against natural degradation factors (e.g., rain streaks) or artificially attack factors (e.g., adversarial attack). Whereas, most existing methods are designed

  60. Daniel Wesego, Pedram Rooshenas

    Multimodal Variational Autoencoders (VAEs) represent a promising group of generative models that facilitate the construction of a tractable posterior within the latent space given multiple modalities. Previous studies have shown that as the number of modalities increases, the generative quality of each modality declines. In this study, we explore an alternat

  61. Yuki Uematsu, Hiroyuki Ohshima

    Water-in-oil emulsions and droplets exhibit completely different physico-chemical properties compared to oil-in-water emulsions and droplets. Thus, directly applying a standard theoretical model to water-in-oil systems cannot describe these anomalous properties. Here, the electrophoretic mobility of a water-in-oil droplet is analytically investigated using D

  62. Jiahao Tan, Yipeng Zhou, Gang Liu, Jessie Hui Wang

    The federated learning (FL) paradigm emerges to preserve data privacy during model training by only exposing clients' model parameters rather than original data. One of the biggest challenges in FL lies in the non-IID (not identical and independently distributed) data (a.k.a., data heterogeneity) distributed on clients. To address this challenge, various per

  63. Xia Feng, Jiao-Kai Chen, Jia-Qi Xie

    The concept of diquark is important for understanding hadron structure and high-energy particle reactions. We attempt to apply the Regge trajectory approach to the doubly heavy diquarks. We present a method for determining the parameters in the diquark Regge trajectory. The spectra of diquarks $(cc)$, $(bb)$, and $(bc)$ are obtained by using the {\rt} approa

  64. Uijeong Jang, Shuvomoy Das Gupta, Ernest K. Ryu

    The accelerated composite optimization method FISTA (Beck, Teboulle 2009) is suboptimal by a constant factor, and we present a new method OptISTA that improves FISTA by a constant factor of 2. The performance estimation problem (PEP) has recently been introduced as a new computer-assisted paradigm for designing optimal first-order methods. In this work, we p

  65. Kaiwen Wang, Kevin Zhou, Runzhe Wu, Nathan Kallus

    While distributional reinforcement learning (DistRL) has been empirically effective, the question of when and why it is better than vanilla, non-distributional RL has remained unanswered. This paper explains the benefits of DistRL through the lens of small-loss bounds, which are instance-dependent bounds that scale with optimal achievable cost. Particularly,

  66. Fernando Iniguez, Mark Srednicki

    We consider the properties of an observable (such as a single spin component that squares to the identity) when expressed as a matrix in the basis of energy eigenstates, and then truncated to a microcanonical slice of energies of varying width. For a quantum chaotic system, we model the unitary or orthogonal matrix that relates the spin basis to the energy b

  67. Jiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng

    Temporal action localization (TAL), which involves recognizing and locating action instances, is a challenging task in video understanding. Most existing approaches directly predict action classes and regress offsets to boundaries, while overlooking the discrepant importance of each frame. In this paper, we propose an Action Sensitivity Learning framework (A

  68. Yixuan Su, Tian Lan, Huayang Li, Jialu Xu

    We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation, writing stories inspired by videos, and answering questions about audios. More interestingly, PandaGPT can take multimo

  69. Thanh-Dat Truong, Hoang-Quan Nguyen, Bhiksha Raj, Khoa Luu

    Continual semantic segmentation aims to learn new classes while maintaining the information from the previous classes. Although prior studies have shown impressive progress in recent years, the fairness concern in the continual semantic segmentation needs to be better addressed. Meanwhile, fairness is one of the most vital factors in deploying the deep learn

  70. Thanh-Dat Truong, Khoa Luu

    Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learning-based action recognition models remains difficult. Transferring knowledge learned from the large-scale exocentric data to the egocentric

  71. Zi Wang, Jihye Choi, Ke Wang, Somesh Jha

    Motivated by the success of traditional software testing, numerous diversity measures have been proposed for testing deep neural networks (DNNs). In this study, we propose a shift in perspective, advocating for the consideration of DNN testing as directed testing problems rather than diversity-based testing tasks. We note that the objective of testing DNNs i

  72. Nan Jiang, Yexiang Xue

    Learning symbolic expressions directly from experiment data is a vital step in AI-driven scientific discovery. Nevertheless, state-of-the-art approaches are limited to learning simple expressions. Regressing expressions involving many independent variables still remain out of reach. Motivated by the control variable experiments widely utilized in science, we

  73. Siping Shi, Bihai Zhang, Dan Wang

    Recently, inference privacy has attracted increasing attention. The inference privacy concern arises most notably in the widely deployed edge-cloud video analytics systems, where the cloud needs the videos captured from the edge. The video data can contain sensitive information and subject to attack when they are transmitted to the cloud for inference. Many

  74. Jesse Cummings, Elías Snorrason, Jonas Mueller

    We present a straightforward statistical test to detect certain violations of the assumption that the data are Independent and Identically Distributed (IID). The specific form of violation considered is common across real-world applications: whether the examples are ordered in the dataset such that almost adjacent examples tend to have more similar feature v

  75. Xiaoyu Chen, Shenao Zhang, Pushi Zhang, Li Zhao

    With strong capabilities of reasoning and a broad understanding of the world, Large Language Models (LLMs) have demonstrated immense potential in building versatile embodied decision-making agents capable of executing a wide array of tasks. Nevertheless, when deployed in unfamiliar environments, we show that LLM agents encounter challenges in efficiently gat

  76. Liang Peng, Junkai Xu, Haoran Cheng, Zheng Yang

    Monocular 3D detection is a challenging task due to the lack of accurate 3D information. Existing approaches typically rely on geometry constraints and dense depth estimates to facilitate the learning, but often fail to fully exploit the benefits of three-dimensional feature extraction in frustum and 3D space. In this paper, we propose \textbf{OccupancyM3D},

  77. Aaron Shih, Mathias Casiulis, Stefano Martiniani

    Media with correlated disorder display unexpected transport properties, but it is still a challenge to design structures with desired spectral features at scale. In this work, we introduce an optimal formulation of this inverse problem by means of the non-uniform fast Fourier transform, thus arriving at an algorithm capable of generating systems with arbitra

  78. Zihan Wang, Yang Yang, Zhi Liu, Yifan Zheng

    Currently, video behavior recognition is one of the most foundational tasks of computer vision. The 2D neural networks of deep learning are built for recognizing pixel-level information such as images with RGB, RGB-D, or optical flow formats, with the current increasingly wide usage of surveillance video and more tasks related to human action recognition. Th

  79. Eric Mbakop

    This paper presents an algorithm that generates the conditional moment inequalities that characterize the identified set of the common parameter of various semi-parametric panel multinomial choice models. I consider both static and dynamic models, and consider various weak stochastic restrictions on the distribution of observed and unobserved components of t

  80. H. Narayanan

    In this paper, we study the ideas of composition and decomposition in the context of vector spaces, graphs and matroids. For vector spaces $\V_{AB},$ treated as collection of row vectors, with specified column set $A\uplus B,$ we define $\V_{SP}\lrarv \V_{PQ}, S\cap Q= \emptyset, $ to be the collection of all vectors $(f_S,f_Q)$ such that $(f_S,f_P)\in \V_{S

  81. Adithya Kulkarni, Mohna Chakraborty, Yonas Sium, Sai Charishma Valluri

    In this paper, we explore the feasibility of finding algorithm implementations from code. Successfully matching code and algorithms can help understand unknown code, provide reference implementations, and automatically collect data for learning-based program synthesis. To achieve the goal, we designed a new language named p-language to specify the algorithms

  82. Mohna Chakraborty, Adithya Kulkarni, Qi Li

    Recent studies have demonstrated that natural-language prompts can help to leverage the knowledge learned by pre-trained language models for the binary sentence-level sentiment classification task. Specifically, these methods utilize few-shot learning settings to fine-tune the sentiment classification model using manual or automatically generated prompts. Ho

  83. Jiqing Zhang, Yuanchen Wang, Wenxi Liu, Meng Li

    Most existing RGB-based trackers target low frame rate benchmarks of around 30 frames per second. This setting restricts the tracker's functionality in the real world, especially for fast motion. Event-based cameras as bioinspired sensors provide considerable potential for high frame rate tracking due to their high temporal resolution. However, event-based c

  84. Zhi Jiang, Hui Liu

    Pareschi showed that an ample divisor of degree $d$ on a simple abelian variety of dimension $g$ has mild singularities when $d<g$. We extend his result to general polarized abelian varieties. We also show that ample divisors of degree $3$ or $4$ on an abelian variety have mild singularities.

  85. Chunlin Sun, Linyu Liu, Xiaocheng Li

    Contextual optimization, also known as predict-then-optimize or prescriptive analytics, considers an optimization problem with the presence of covariates (context or side information). The goal is to learn a prediction model (from the training data) that predicts the objective function from the covariates, and then in the test phase, solve the optimization p

  86. Lei Shu, Liangchen Luo, Jayakumar Hoskere, Yun Zhu

    Large Language Models (LLMs) have demonstrated impressive capabilities in creative tasks such as storytelling and E-mail generation. However, as LLMs are primarily trained on final text results rather than intermediate revisions, it might be challenging for them to perform text rewriting tasks. Most studies in the rewriting tasks focus on a particular transf

  87. Huawen Feng, Zhenxi Lin, Qianli Ma

    In text classification, the traditional attention mechanisms usually focus too much on frequent words, and need extensive labeled data in order to learn. This paper proposes a perturbation-based self-supervised attention approach to guide attention learning without any annotation overhead. Specifically, we add as much noise as possible to all the words in th

  88. Zheng Fan, Dan Long, Xuan Mao, Guo-Qing Qin

    Dynamic encircling a second-order exception point (EP) exhibit chiral state transfer, while there is few research on dynamic encircling multiple and higher-order EPs. Here, we study proximity-encirclement of the EPs in a multimode optomechanical system to understand the closed path evolution of high-order non-Hermitian systems. The optomechanical system has

  89. Katsuyuki Naoi

    The extended $T$-systems are a number of short exact sequences in the category of finite-dimensional modules over the quantum affine algebras of types $A_n^{(1)}$ and $B_n^{(1)}$, introduced by Mukhin and Young as a generalization of the $T$-systems. In this paper we establish the extended $T$-systems for more general modules, which are constructed from an a

  90. Shangshang Shi, Zhimin Wang, Jiaxin Li, Yanan Li

    The self-attention mechanism (SAM) has demonstrated remarkable success in various applications. However, training SAM on classical computers becomes computationally challenging as the number of trainable parameters grows. Quantum neural networks (QNNs) have been developed as a novel learning model that promises to provide speedup for pattern recognition usin

  91. Zhenhua Liu, Feipeng Ma, Tianyi Wang, Fengyun Rao

    With the development of multimedia technology, Video Copy Detection has been a crucial problem for social media platforms. Meta AI hold Video Similarity Challenge on CVPR 2023 to push the technology forward. In this report, we share our winner solutions on Matching Track. We propose a Similarity Alignment Model(SAM) for video copy segment matching. Our SAM e

  92. Michael J. Ryan, Tarek Naous, Wei Xu

    Recent advancements in high-quality, large-scale English resources have pushed the frontier of English Automatic Text Simplification (ATS) research. However, less work has been done on multilingual text simplification due to the lack of a diverse evaluation benchmark that covers complex-simple sentence pairs in many languages. This paper introduces the Multi

  93. Dong Liang, Martin Guay, Shimin Wang

    In this paper, a bipartite output regulation problem is solved for a class of nonlinear multi-agent systems subject to static signed communication networks. A nonlinear distributed observer is proposed for a nonlinear exosystem with cooperation-competition interactions to address the problem. Sufficient conditions are provided to guarantee its existence and

  94. Yuejiao Fei, Leyang Cui, Sen Yang, Wai Lam

    Grammatical error correction systems improve written communication by detecting and correcting language mistakes. To help language learners better understand why the GEC system makes a certain correction, the causes of errors (evidence words) and the corresponding error types are two key factors. To enhance GEC systems with explanations, we introduce EXPECT,

  95. Abbas Javan Jafari, Diego Elias Costa, Emad Shihab, Rabe Abdalkareem

    Managing project dependencies is a key maintenance issue in software development. Developers need to choose an update strategy that allows them to receive important updates and fixes while protecting them from breaking changes. Semantic Versioning was proposed to address this dilemma but many have opted for more restrictive or permissive alternatives. This e

  96. Miranda Christ, Sam Gunn, Or Zamir

    Recent advances in the capabilities of large language models such as GPT-4 have spurred increasing concern about our ability to detect AI-generated text. Prior works have suggested methods of embedding watermarks in model outputs, by noticeably altering the output distribution. We ask: Is it possible to introduce a watermark without incurring any detectable

  97. Meng-Yao Zhang, Hao Chen, Hassan Hassanabadi, Zheng-Wen Long

    The classification of critical points of charged topological black holes (TBHs) in anti-de Sitter spacetime (AdS) under the Power Maxwell Invariant (PMI)-massive gravity is accomplished within the framework of black hole chemistry (BHC). Considering the grand canonical ensemble (GCE), we show that $d=4$ black hole have only one topological class, whereas $d\

  98. Rui Liu, Jinhua Zhang, Guanglai Gao, Haizhou Li

    Audio Deepfake Detection (ADD) aims to detect the fake audio generated by text-to-speech (TTS), voice conversion (VC) and replay, etc., which is an emerging topic. Traditionally we take the mono signal as input and focus on robust feature extraction and effective classifier design. However, the dual-channel stereo information in the audio signal also include

  99. Aakas Zhiyuli, Yanfang Chen, Xuan Zhang, Xun Liang

    With the continuous development and change exhibited by large language model (LLM) technology, represented by generative pretrained transformers (GPTs), many classic scenarios in various fields have re-emerged with new opportunities. This paper takes ChatGPT as the modeling object, incorporates LLM technology into the typical book resource understanding and

  100. Hu Sun, Zuofeng Shang, Yang Chen

    We develop a new methodology for forecasting matrix-valued time series with historical matrix data and auxiliary vector time series data. We focus on a time series of matrices defined on a static 2-D spatial grid and an auxiliary time series of non-spatial vectors. The proposed model, Matrix AutoRegression with Auxiliary Covariates (MARAC), contains an autor