Skip to content

December 2023 arXiv papers — page 117

Showing 11,60111,700 of 18,165 papers

  1. Daniel Grumiller, Romain Ruzziconi, Céline Zwikel

    Leaky boundary conditions in asymptotically AdS spacetimes are relevant to discuss black hole evaporation and the evolution of the Page curve via the island formula. We explore the consequences of leaky boundary conditions on the one-loop partition function of gravity. We focus on JT gravity minimally coupled to a scalar field whose normalizable and non-norm

  2. James Bonifacio, Kurt Hinterbichler

    We study extended shift symmetries that arise for fermionic fields on anti-de Sitter (AdS) space and de Sitter (dS) space for particular values of the mass relative to the curvature scale. We classify these symmetries for general mixed-symmetry fermionic fields in arbitrary dimension and describe how fields with these symmetries arise as the decoupled longit

  3. Luis Colmenarez, Ze-Min Huang, Sebastian Diehl, Markus Müller

    Quantum error correcting (QEC) codes protect quantum information from decoherence, as long as error rates fall below critical error thresholds. In general, obtaining thresholds implies simulating the QEC procedure using, in general, sub-optimal decoding strategies. In a few cases and for sufficiently simple noise models, optimal decoding of QEC codes can be

  4. Ziyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu

    The increased demand for 3D data in AR/VR, robotics and gaming applications, gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS) algorithm to optimize a 3D representation such that the rendered image maintains a high likelihood as evaluated by a pre-tra

  5. Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu

    We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a win

  6. Bharath Raj Nagoor Kani, Hsin-Ying Lee, Sergey Tulyakov, Shubham Tulsiani

    We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate information from input views, but are not robust in-the-wild when such

  7. Chong Zhou, Xiangtai Li, Chen Change Loy, Bo Dai

    This paper presents EdgeSAM, an accelerated variant of the Segment Anything Model (SAM), optimized for efficient execution on edge devices with minimal compromise in performance. Our approach involves distilling the original ViT-based SAM image encoder into a purely CNN-based architecture, better suited for edge devices. We carefully benchmark various distil

  8. Andrea Angiuli, Jean-Pierre Fouque, Mathieu Laurière, Mengrui Zhang

    We establish the convergence of the unified two-timescale Reinforcement Learning (RL) algorithm presented in a previous work by Angiuli et al. This algorithm provides solutions to Mean Field Game (MFG) or Mean Field Control (MFC) problems depending on the ratio of two learning rates, one for the value function and the other for the mean field term. Our proof

  9. Yuzhe Yang, Haoran Zhang, Judy W Gichoya, Dina Katabi

    As artificial intelligence (AI) rapidly approaches human-level performance in medical imaging, it is crucial that it does not exacerbate or propagate healthcare disparities. Prior research has established AI's capacity to infer demographic data from chest X-rays, leading to a key concern: do models using demographic shortcuts have unfair predictions across s

  10. Alex Kulesza, Ananda Theertha Suresh, Yuyan Wang

    Differential privacy is often studied under two different models of neighboring datasets: the add-remove model and the swap model. While the swap model is frequently used in the academic literature to simplify analysis, many practical applications rely on the more conservative add-remove model, where obtaining tight results can be difficult. Here, we study t

  11. Ka Leong Cheng, Qiuyu Wang, Zifan Shi, Kecheng Zheng

    Neural radiance fields, which represent a 3D scene as a color field and a density field, have demonstrated great progress in novel view synthesis yet are unfavorable for editing due to the implicitness. This work studies the task of efficient 3D editing, where we focus on editing speed and user interactivity. To this end, we propose to learn the color field

  12. Ghazaleh Asghari, Jani A. Virtanen, Zhangjian Hu

    Using the notion of integral distance to analytic functions, we give a characterization of Schatten class Hankel operators acting on doubling Fock spaces on the complex plane and use it to show that for $f\in L^{\infty}$, if $H_{f}$ is Hilbert-Schmidt, then so is $H_{\bar{f}}$. This property is known as the Berger-Coburn phenomenon. When $0<p\le 1$, we show

  13. Fangfu Liu, Diankun Wu, Yi Wei, Yongming Rao

    Recently, 3D content creation from text prompts has demonstrated remarkable progress by utilizing 2D and 3D diffusion models. While 3D diffusion models ensure great multi-view consistency, their ability to generate high-quality and diverse 3D assets is hindered by the limited 3D data. In contrast, 2D diffusion models find a distillation approach that achieve

  14. Ava Pun, Gary Sun, Jingkang Wang, Yun Chen

    Different outdoor illumination conditions drastically alter the appearance of urban scenes, and they can harm the performance of image-based robot perception systems if not seen during training. Camera simulation provides a cost-effective solution to create a large dataset of images captured under different lighting conditions. Towards this goal, we propose

  15. Neerja Thakkar, Karttikeya Mangalam, Andrea Bajcsy, Jitendra Malik

    Human trajectory prediction is typically posed as a zero-shot generalization problem: a predictor is learnt on a dataset of human motion in training scenes, and then deployed on unseen test scenes. While this paradigm has yielded tremendous progress, it fundamentally assumes that trends in human behavior within the deployment scene are constant over time. As

  16. Shabaz Patel, Hassan Kane, Rayhan Patel

    Large Language Models (LLMs) have demonstrated remarkable performance across numerous natural language understanding use cases. However, this impressive performance comes with inherent limitations, such as the tendency to perpetuate stereotypical biases or fabricate non-existent facts. In the context of Islam and its representation, accurate and factual repr

  17. Junbum Cha, Wooyoung Kang, Jonghwan Mun, Byungseok Roh

    In Multimodal Large Language Models (MLLMs), a visual projector plays a crucial role in bridging pre-trained vision encoders with LLMs, enabling profound visual understanding while harnessing the LLMs' robust capabilities. Despite the importance of the visual projector, it has been relatively less explored. In this study, we first identify two essential proj

  18. Pratul P. Srinivasan, Stephan J. Garbin, Dor Verbin, Jonathan T. Barron

    Existing UV mapping algorithms are designed to operate on well-behaved meshes, instead of the geometry representations produced by state-of-the-art 3D reconstruction and generation techniques. As such, applying these methods to the volume densities recovered by neural radiance fields and related techniques (or meshes triangulated from such fields) results in

  19. Wenbo Sun

    This paper is the first part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we prove a quantitative equidistribution theorem for polynomial sequences

  20. Wenbo Sun

    This paper is the second part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we study additive combinatorial properties for shifted modules, i.e. the

  21. Wenbo Sun

    This paper is the fourth and the last part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the Geometric Ramsey Conjecture in the finite field setting. In this paper, we proof a conjecture of Graham on the Remsey properties fo

  22. Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu

    Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks. When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g. document, passage, or sentence. We discover that the retrieval unit ch

  23. David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar, Teresa Yeo

    Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly versatile models in computer vision. In this paper, we take a step in this direction and propose a multimodal training scheme

  24. Junwei Deng, Xirui Jiang, Shiyuan Zhang, Shichang Zhang

    The rapid rise of generative AI has intensified copyright and economic tensions in creative industries, particularly in music. Current approaches addressing this challenge often focus on preventing infringement or establishing one-time licensing, which fail to provide the sustainable, recurring economic incentives necessary to maintain creative ecosystems. T

  25. Teodora Popordanoska, Aleksei Tiulpin, Matthew B. Blaschko

    Despite their impressive predictive performance in various computer vision tasks, deep neural networks (DNNs) tend to make overly confident predictions, which hinders their widespread use in safety-critical applications. While there have been recent attempts to calibrate DNNs, most of these efforts have primarily been focused on classification tasks, thus ne

  26. Rao Fu, Zehao Wen, Zichen Liu, Srinath Sridhar

    Inspired by cognitive theories, we introduce AnyHome, a framework that translates any text into well-structured and textured indoor scenes at a house-scale. By prompting Large Language Models (LLMs) with designed templates, our approach converts provided textual narratives into amodal structured representations. These representations guarantee consistent and

  27. Pooja Prajod, Matteo Lavit Nicora, Marta Mondellini, Giovanni Tauro

    Collaborative robots (cobots) are widely used in industrial applications, yet extensive research is still needed to enhance human-robot collaborations and operator experience. A potential approach to improve the collaboration experience involves adapting cobot behavior based on natural cues from the operator. Inspired by the literature on human-human interac

  28. Yixing Lao, Xiaogang Xu, Zhipeng Cai, Xihui Liu

    Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed by off-the-shelf methods to supervise NeRF training. We desi

  29. Vijeth Hebbar, Cedric Langbort

    In many online sequential decision-making scenarios, a learner's choices affect not just their current costs but also the future ones. In this work, we look at one particular case of such a situation where the costs depend on the time average of past decisions over a history horizon. We first recast this problem with history dependent costs as a problem of d

  30. Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo

    Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomn

  31. Ruihan Yang, Yejin Kim, Rose Hendrix, Aniruddha Kembhavi

    Recent advancements in robotics have enabled robots to navigate complex scenes or manipulate diverse objects independently. However, robots are still impotent in many household tasks requiring coordinated behaviors such as opening doors. The factorization of navigation and manipulation, while effective for some tasks, fails in scenarios requiring coordinated

  32. Lev V. Utkin, Danila Y. Eremenko, Andrei V. Konstantinov

    A new method called the Survival Beran-based Neural Importance Model (SurvBeNIM) is proposed. It aims to explain predictions of machine learning survival models, which are in the form of survival or cumulative hazard functions. The main idea behind SurvBeNIM is to extend the Beran estimator by incorporating the importance functions into its kernels and by im

  33. Seongjoon Kang

    Due to the high complexity of geometry-deterministic wireless channel modeling and the difficulty in its implementation, geometry-based stochastic channel modeling (GBSM) approaches have been used to evaluate wireless systems. This paper introduces a new method to model any GBSM by training a generative neural network using images formed by channel parameter

  34. Wenbo Sun

    This paper is the third part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we prove an inverse theorem over finite field for spherical Gowers norms,

  35. Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda

    Transformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness an

  36. Wentao Tang

    This work proposes a data-driven approach for bifurcation analysis in nonlinear systems when the governing differential equations are not available. Specifically, regularized regression with barrier terms is used to learn a homeomorphism that transforms the underlying system to a reference linear dynamics -- either an explicit reference model with desired qu

  37. Kevin Coakley, Christine R. Kirkpatrick, Odd Erik Gundersen

    Reproducing published deep learning papers to validate their conclusions can be difficult due to sources of irreproducibility. We investigate the impact that implementation factors have on the results and how they affect reproducibility of deep learning studies. Three deep learning experiments were ran five times each on 13 different hardware environments an

  38. Jiyan He, Weitao Feng, Yaosen Min, Jingwei Yi

    The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks like creation of harmful substances, or circumvention of established regulations. In this study, we aim to raise awaren

  39. Digvijay Wadekar, Javier Roulet, Tejaswi Venumadhav, Ajit Kumar Mehta

    Nearly all of the previous gravitational wave (GW) searches in the LIGO-Virgo data included GW waveforms with only the dominant quadrupole harmonic, i.e., omitting higher-order harmonics which are predicted by general relativity. We improved the IAS pipeline by efficiently introducing higher harmonics in the GW templates using the techniques in Wadekar et al

  40. Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang

    Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggregation of datasets to enhance data volume and diversity. Ho

  41. Angsuman Das

    In this paper, we introduce and study the iterates of the following family of functions $\varphi_k$ defined on natural numbers which exhibits nice properties. $$\varphi_k(x)=\left\lbrace \begin{array}{ll} x+k, & \mbox{ if $x$ is prime;}\\ \mbox{largest prime divisor of $x$,} & \mbox{ if $x$ is composite;} \end{array} \right.$$ In particular, we study the per

  42. Denis I. Saveliev

    We consider a certain class of infinitary rules of inference, called here restriction rules, using of which allows us to deduce complete theories of given models. The first instance of such rules was the $\omega$-rule introduced by Hilbert, and generalizations of the $\omega$-rule were first considered by Henkin. Later on Barwise showed that within countable

  43. Giordano De Marzo, Luciano Pietronero, David Garcia

    Scale-free networks are one of the most famous examples of emergent behavior and are ubiquitous in social systems, especially online social media in which users can follow each other. By analyzing the interactions of multiple generative agents using GPT3.5-turbo as a language model, we demonstrate their ability to not only mimic individual human linguistic b

  44. Mu Tian, Qinzhu Yang, Yi Gao

    The success of deep networks in medical image segmentation relies heavily on massive labeled training data. However, acquiring dense annotations is a time-consuming process. Weakly-supervised methods normally employ less expensive forms of supervision, among which scribbles started to gain popularity lately thanks to its flexibility. However, due to lack of

  45. Georgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos

    Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being conditioned on audio clips. However, having the ability to synthesize talking humans from text transcriptions rather than audio i

  46. Lior Gishboliner, Zhihan Jin, Benny Sudakov

    Many well-studied problems in extremal combinatorics deal with the maximum possible size of a family of objects in which every pair of objects satisfies a given restriction. One problem of this type was recently raised by Alon, Gujgiczer, K\"orner, Milojevi\'c and Simonyi. They asked to determine the maximum size of a family $\mathcal{G}$ of graphs on $[n]$,

  47. Vladislav Makarov, Marat Movsin

    GF(2)-grammars are a somewhat recently introduced grammar family that have some unusual algebraic properties and are closely connected to unambiguous grammars. In "Bounded languages described by GF(2)-grammars", Makarov proved a necessary condition for subsets of $a_1^* a_2^* \cdots a_k^*$ to be described by some GF(2)-grammar. By extending these methods fur

  48. Matthew S. Schmitt, Maciej Koch-Janusz, Michel Fruchart, Daniel S. Seara

    The dynamics of many-body systems can often be captured in terms of only a few relevant variables. Mathematical and numerical approaches exist to identify these variables by exploiting a separation of time scales between slow relevant and fast irrelevant variables, but such a separation of scales is not always obvious or even available. In this work, we intr

  49. Haoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen

    Reconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods might face challenges related to the preservation of image cate

  50. Chris A. J. Klaassen

    For a random variable with a unimodal distribution and finite second moment Gau\ss \, (1823) proved a sharp bound on the probability of the random variable to be outside a symmetric interval around its mode. An alternative proof for it is given based on Khintchine's representation of unimodal random variables. Analogously, a sharp inequality is proved for th

  51. Jinming Li, Shihao Wu, Chengyu Cui, Gongjun Xu

    Latent space models are powerful statistical tools for modeling and understanding network data. While the importance of accounting for uncertainty in network analysis has been well recognized, the current literature predominantly focuses on point estimation and prediction, leaving the statistical inference of latent space models an open question. This work a

  52. Jyoti Prakash Saha

    Let $\Gamma$ be a Cayley graph, or a Cayley sum graph, or a twisted Cayley graph, or a twisted Cayley sum graph, or a vertex-transitive graph. Suppose $\Gamma$ is undirected and non-bipartite. Let $\mu$ (resp. $\mu_2$) denote the smallest (resp. the second largest) eigenvalue of the normalized adjacency operator of $\Gamma$, and $d$ denote the degree of $\Ga

  53. Tirasan Khandhawit, Puttipong Pongtanapaisan, Athibadee Wasun

    Two isomorphic graphs can have inequivalent spatial embeddings in 3-space. In this way, an isomorphism class of graphs contains many spatial graph types. A common way to measure the complexity of a spatial graph type is to count the minimum number of straight sticks needed for its construction in 3-space. In this paper, we give estimates of this quantity by

  54. Paolo M Bassani, Joao Magueijo

    In unimodular-like theories, the constants of nature are demoted from pre-given parameters to phase space variables. Their canonical duals provide physical time variables. We investigate how this interacts with an alternative approach to varying constants, where they are replaced by dynamical scalar fields. Specifically we investigate the Brans-Dicke theory

  55. Zheng-Duo Fan, Xiao-Qi Sun, Jing-Yuan Chen

    Thermal transport has been used to probe the nature of $\alpha$-RuCl$_3$, an important candidate of Kitaev material. Two remarkable observations were made under applied magnetic fields at low temperatures, and have stimulated extensive discussions. One is a sizable thermal Hall effect, and the other is an apparent "oscillation" of the longitudinal thermal co

  56. Mikkel Bennedsen, Eric Hillebrand, Jingying Zhou Lykke

    We propose a non-linear state-space model to examine the relationship between CO$_2$ emissions, energy sources, and macroeconomic activity, using data from 1971 to 2019. CO$_2$ emissions are modeled as a weighted sum of fossil fuel use, with emission conversion factors that evolve over time to reflect technological changes. GDP is expressed as the outcome of

  57. M. Poisson, M. López Fuentes, C. H. Mandrini, F. Grings

    Active regions (ARs) appear in the solar atmosphere as a consequence of the emergence of magnetic flux ropes (FRs). Due to the presence of twist, the photospheric line-of-sight (LOS) magnetograms of emerging ARs show an elongation of the polarities known as magnetic tongues. These tongues can affect the estimation of tilt angles during their emergence phase.

  58. Guglielmo Camporese, Alessandro Bergamo, Xunyu Lin, Joseph Tighe

    Early action recognition is an important and challenging problem that enables the recognition of an action from a partially observed video stream where the activity is potentially unfinished or even not started. In this work, we propose a novel model that learns a prototypical representation of the full action for each class and uses it to regularize the arc

  59. Rishabh, Hadi Zadeh-Haghighi, Christoph Simon

    Weak magnetic field exposure can affect many biological processes across a wide range of living organisms. Recently, it has been observed that weak magnetic fields can modulate reactive oxygen species (ROS) concentration, affecting regeneration in planaria. These effects show unusual nonlinear dependence on magnetic field strength, including a sign change. I

  60. Tejes Gaertner, Jared Reiten

    In this work, we construct a new data type for hadronic jets in which the traditional point-cloud representation is transformed into a simplicial complex consisting of vertices, or 0-simplexes. An angular resolution scale, $r$, is then drawn about each vertex, forming balls about hadrons. As $r$ grows, the overlap of balls form 2- and 3-point connections, th

  61. Stefano Meda, Federico Santagati

    We introduce the centred and the uncentred triangular maximal operators $\mathcal T$ and $\mathcal U$, respectively, on any locally finite tree in which each vertex has at least three neighbours. We prove that both $\mathcal T$ and $\mathcal U$ are bounded on $L^p$ for every $p$ in $(1,\infty]$, that $\mathcal T$ is also bounded on $L^1(\mathfrak T)$, and th

  62. Aditya Prakash, Arjun Gupta, Saurabh Gupta

    Objects undergo varying amounts of perspective distortion as they move across a camera's field of view. Models for predicting 3D from a single image often work with crops around the object of interest and ignore the location of the object in the camera's field of view. We note that ignoring this location information further exaggerates the inherent ambiguity

  63. Patrick M. Duerr, Abigail Holmes Mills

    This paper critically examines the models of scientific discovery propounded by Kuhn, McArthur, Hudson, and Schindler. As an alternative, we proffer the $\unicode{x201c}$change$\unicode{x2013}$driver model$\unicode{x201c}$. It conceives of discoveries as problems or solutions to problems that have epistemically advanced science. Here we take a problem to be

  64. Aaron L. Sarvet, Mats J. Stensrud, Lan Wen

    We formalize an interpretational error that is common in statistical causal inference, termed identity slippage. This formalism is used to describe historically-recognized fallacies, and analyse a fast-growing literature in statistics and applied fields. We conducted a systematic review of natural language claims in the literature on stochastic mediation par

  65. Thomas Foster, Ioana Croitoru, Robert Dorfman, Christoffer Edlund

    In this work, we address in-context learning (ICL) for the task of image segmentation, introducing a novel approach that adapts a modern Video Object Segmentation (VOS) technique for visual in-context learning. This adaptation is inspired by the VOS method's ability to efficiently and flexibly learn objects from a few examples. Through evaluations across a r

  66. Anish Chakrabarty, Arkaprabha Basu, Swagatam Das

    Variational Autoencoders (VAEs) have been a pioneering force in the realm of deep generative models. Amongst its legions of progenies, Wasserstein Autoencoders (WAEs) stand out in particular due to the dual offering of heightened generative quality and a strong theoretical backbone. WAEs consist of an encoding and a decoding network forming a bottleneck with

  67. Alejandro Cárdenas-Avendaño, Aaron Held

    General relativity's prediction that all black holes are described by the Kerr metric, irrespective of their size, can now be empirically tested using electromagnetic observations of supermassive black holes and gravitational waves from mergers of stellar-mass black holes. In this work, we focus on the electromagnetic side of this test and quantify the const

  68. Alexander Roth

    The decarbonization of buildings requires the phase-out of fossil fuel heating systems. Heat pumps are considered a crucial technology to supply a substantial part of heating energy for buildings. Yet, their introduction is not without challenges, as heat pumps generate additional electricity demand as well as peak loads. To better understand these challenge

  69. L. A. Lessa, R. V. Maluf, J. E. G. Silva, C. A. S. Almeida

    Einstenian cubic gravity (ECG) is a modified theory of gravity constructed with cubic contractions of the curvature tensor. This theory has the remarkable feature of having the same two propagating degrees of freedom of Einstein gravity (EG), at the perturbative level on maximally symmetric spacetimes. The additional unstable modes steaming from the higher o

  70. Yao Sun, Yi Wang, Michael Eineder

    Quick and automated earthquake-damaged building detection from post-event satellite imagery is crucial, yet it is challenging due to the scarcity of training data required to develop robust algorithms. This letter presents the first dataset dedicated to detecting earthquake-damaged buildings from post-event very high resolution (VHR) Synthetic Aperture Radar

  71. Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, Andrew J. Davison

    We present the first application of 3D Gaussian Splatting in monocular SLAM, the most fundamental but the hardest setup for Visual SLAM. Our method, which runs live at 3fps, utilises Gaussians as the only 3D representation, unifying the required representation for accurate, efficient tracking, mapping, and high-quality rendering. Designed for challenging mon

  72. Mauro Barbieri

    This document details the first public data release of the HARPS radial velocities catalog. This data release aims to provide the astronomical community with a catalog of radial velocities obtained with spectroscopic observations acquired from 2003 to 2023 with the High Accuracy Radial Velocity Planet Searcher (HARPS) spectrograph installed at the ESO 3.6m t

  73. Avi Singh, John D. Co-Reyes, Rishabh Agarwal, Ankesh Anand

    Fine-tuning language models~(LMs) on human-generated data remains a prevalent practice. However, the performance of such models is often limited by the quantity and diversity of high-quality human data. In this paper, we explore whether we can go beyond human data on tasks where we have access to scalar feedback, for example, on math problems where one can v

  74. Aybike Çatal-Özer, Keremcan Doğan, Cem Yetişmişoğlu

    We extend the notion of Lie bialgebroids for more general bracket structures used in string and M theories. We formalize the notions of calculus and dual calculi on algebroids. We achieve this by reinterpreting the main results of the matched pairs of Leibniz algebroids. By examining a rather general set of fundamental algebroid axioms, we present the compat

  75. Aditya Prakash, Ruisen Tu, Matthew Chang, Saurabh Gupta

    3D hand pose estimation in everyday egocentric images is challenging for several reasons: poor visual signal (occlusion from the object of interaction, low resolution & motion blur), large perspective distortion (hands are close to the camera), and lack of 3D annotations outside of controlled settings. While existing methods often use hand crops as input to

  76. U. Singh, A. Al-Durra

    Hosting capacity analysis is essential for effective integration of distributed energy resources into distribution systems. This paper discusses hosting capacity analysis with emphasis on various aspects affecting the process. This paper addresses key research gaps, aiming to improve the accuracy, scalability, and practicality of hosting capacity estimation.

  77. Dashiell Stander, Qinan Yu, Honglu Fan, Stella Biderman

    The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep neural networks, but all have substantial limitations. Algorithmic tasks have proven to be a fruitful test ground for interpreting a neural network end-to-end. Building on previous

  78. Ruochen Dai, Michael Lee, Patrick Hoey, Weimin Fu

    As the complexity of logic designs increase, new avenues for testing digital hardware becomes necessary. Fuzz Testing (fuzzing) has recently received attention as a potential candidate for input vector generation on hardware designs. Using this technique, a fuzzer is used to generate an input to a logic design. Using a simulation engine, the logic design is

  79. Samyukta Sethuraman, Ankur Bansal, Setareh Mardan, Mauricio G. C. Resende

    Amazon Locker is a self-service delivery or pickup location where customers can pick up packages and drop off returns. A basic first-come-first-served policy for accepting package delivery requests to lockers results in lockers becoming full with standard shipping speed (3-5 day shipping) packages, and leaving no space left for expedited packages which are m

  80. Zhezheng Hao, Feiping Nie, Rong Wang

    Support Vector Machine (SVM) stands out as a prominent machine learning technique widely applied in practical pattern recognition tasks. It achieves binary classification by maximizing the "margin", which represents the minimum distance between instances and the decision boundary. Although many efforts have been dedicated to expanding SVM for multi-class cas

  81. Max Hallgren

    We show that any tangent cone of a singular shrinking K\"ahler-Ricci soliton is a normal affine algebraic variety. Moreover, the regular set of such a tangent cone in the metric sense coincides with the regular set in the algebraic sense. Along the way, we give a parabolic proof of H\"ormander's $L^{2}$ estimate, which can be used to solve the $\overline{\pa

  82. Kushal Bose, Swagatam Das

    Graph Transformers (GTs) facilitate the comprehension of complex relationships on graph-structured data by leveraging self-attention of the possible pairs of nodes. The structural information or inductive bias of the input graph is provided as positional encodings to the GT. The positional encodings are mostly Euclidean and are not able to capture the comple

  83. Zhen Xu, Tao Xie, Sida Peng, Haotong Lin

    Volumetric video is a technology that digitally records dynamic events such as artistic performances, sporting events, and remote conversations. When acquired, such volumography can be viewed from any viewpoint and timestamp on flat screens, 3D displays, or VR headsets, enabling immersive viewing experiences and more flexible content creation in a variety of

  84. Lioba Heimbach, Quentin Kniep, Yann Vonlanthen, Roger Wattenhofer

    Ethereum introduced Transaction Access Lists (TALs) in 2020 to optimize gas costs during transaction execution. In this work, we present a comprehensive analysis of TALs in Ethereum, focusing on adoption, quality, and gas savings. Analyzing a full month of mainnet data with 31,954,474 transactions, we found that only 1.46% of transactions included a TAL, eve

  85. Denis Zavadski, Johann-Friedrich Feiden, Carsten Rother

    The field of image synthesis has made tremendous strides forward in the last years. Besides defining the desired output image with text-prompts, an intuitive approach is to additionally use spatial guidance in form of an image, such as a depth map. In state-of-the-art approaches, this guidance is realized by a separate controlling model that controls a pre-t

  86. SHINE Collaboration, H. Adhikary, P. Adrich, K. K. Allison

    Strong interactions preserve an approximate isospin symmetry between up ($u$) and down ($d$) quarks, part of the more general flavor symmetry. In the case of $K$ meson production, if this isospin symmetry were exact, it would result in equal numbers of charged ($K^+$ and $K^-$) and neutral ($K^0$ and $\overline K^{\,0}$) mesons in the final state. Here, we r

  87. Takahide Yoshida, Atsushi Masumori, Takashi Ikegami

    We report the development of Alter3, a humanoid robot capable of generating spontaneous motion using a Large Language Model (LLM), specifically GPT-4. This achievement was realized by integrating GPT-4 into our proprietary android, Alter3, thereby effectively grounding the LLM with Alter's bodily movement. Typically, low-level robot control is hardware-depen

  88. Maciej Grzeszczuk, Kinga Skorupska

    In this article, we report the pilot results of a survey study (N=1036) related to social attitudes towards the early digital heritage. On the basis of the answers, we consider what constitutes early digital artifacts (EDA) and outline how knowledge about them can be useful. We explore attitudes toward the historical and cultural importance of various EDAs a

  89. M. Majid Butt, Nitin R. Mangalvedhe, Nuno K. Pratas, Johannes Harrebek

    Ambient internet of things (IoT) is the network of devices which harvest energy from ambient sources for powering their communication. After decades of research on operation of these devices, Third Generation Partnership Project (3GPP) has started discussing energy harvesting technology in cellular networks to support massive deployment of IoT devices at low

  90. Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz

    We present Monocular Neural Parametric Head Models (MonoNPHM) for dynamic 3D head reconstructions from monocular RGB videos. To this end, we propose a latent appearance space that parameterizes a texture field on top of a neural parametric model. We constrain predicted color values to be correlated with the underlying geometry such that gradients from RGB ef

  91. Yuzhou Huang, Liangbin Xie, Xintao Wang, Ziyang Yuan

    Current instruction-based editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach to instruction-based image editing that leverages Multimodal Large Language Models (

  92. Shufan Li, Harkanwar Singh, Aditya Grover

    The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction tuning with text-based prompts and multi-modal conditioning. However, these works make one or more unnatural assumptions

  93. Subhajit Dutta Chowdhury, Zhiyu Ni, Qingyuan Peng, Souvik Kundu

    Graph Lottery Tickets (GLTs), comprising a sparse adjacency matrix and a sparse graph neural network (GNN), can significantly reduce the inference latency and compute footprint compared to their dense counterparts. Despite these benefits, their performance against adversarial structure perturbations remains to be fully explored. In this work, we first invest

  94. Tarik Dzanic, Luigi Martinelli

    This work explores the feasibility of performing three-dimensional molecular gas dynamics simulations of hypersonic flows such as re-entry vehicles through directly solving the six-dimensional nonlinear Boltzmann equation closed with the BGK (Bhatnagar-Gross-Krook) collision model. Through the combination of high-order unstructured spatial discretizations an

  95. Guanyu Huang, Roger K. Moore

    Personalisation is essential to achieve more acceptable and effective results in human-robot interaction. Placing users in the central role, many studies have focused on enhancing the abilities of social robots to perceive and understand users. However, little is known about improving user perceptions and interpretation of a social robot in spoken interactio

  96. Luca Marannino

    We generalize and simplify the constructions of Darmon-Rotger and Hsieh of an unbalanced triple product $p$-adic $L$-function $\mathscr{L}_p^f(\boldsymbol{f},\boldsymbol{g},\boldsymbol{h})$ attached to a triple $(\boldsymbol{f},\boldsymbol{g},\boldsymbol{h})$ of $p$-adic families of modular forms, allowing more flexibility for the choice of $\boldsymbol{g}$

  97. Francesco Leofante, Nico Potyka

    Counterfactual explanations shed light on the decisions of black-box models by explaining how an input can be altered to obtain a favourable decision from the model (e.g., when a loan application has been rejected). However, as noted recently, counterfactual explainers may lack robustness in the sense that a minor change in the input can cause a major change

  98. Richard Kadison, Simon Levin, Zhe Liu

    Based on the success of a well-known method for solving higher order linear differential equations, a study of two of the most important mathematical features of that method, viz. the null spaces and commutativity of the product of unbounded linear operators on a Hilbert space is carried out. A principle is proved describing solutions for the product of such

  99. Adrian de Wynter, Xun Wang, Qilong Gu, Si-Qing Chen

    Modern large language models (LLMs) are capable of interpreting input strings as instructions, or prompts, and carry out tasks based on them. Unlike traditional learners, LLMs cannot use back-propagation to obtain feedback, and condition their output in situ in a phenomenon known as in-context learning (ICL). Many approaches to prompting and pre-training the

  100. Hong-Xing Yu, Yang Zheng, Yuan Gao, Yitong Deng

    We study recovering fluid density and velocity from sparse multiview videos. Existing neural dynamic reconstruction methods predominantly rely on optical flows; therefore, they cannot accurately estimate the density and uncover the underlying velocity due to the inherent visual ambiguities of fluid velocity, as fluids are often shapeless and lack stable visu