December 2023 arXiv papers — page 117
Showing 11,601–11,700 of 18,165 papers
Daniel Grumiller, Romain Ruzziconi, Céline Zwikel
Leaky boundary conditions in asymptotically AdS spacetimes are relevant to discuss black hole evaporation and the evolution of the Page curve via the island formula. We explore the consequences of leaky boundary conditions on the one-loop partition function of gravity. We focus on JT gravity minimally coupled to a scalar field whose normalizable and non-norm
James Bonifacio, Kurt Hinterbichler
We study extended shift symmetries that arise for fermionic fields on anti-de Sitter (AdS) space and de Sitter (dS) space for particular values of the mass relative to the curvature scale. We classify these symmetries for general mixed-symmetry fermionic fields in arbitrary dimension and describe how fields with these symmetries arise as the decoupled longit
Luis Colmenarez, Ze-Min Huang, Sebastian Diehl, Markus Müller
Quantum error correcting (QEC) codes protect quantum information from decoherence, as long as error rates fall below critical error thresholds. In general, obtaining thresholds implies simulating the QEC procedure using, in general, sub-optimal decoding strategies. In a few cases and for sufficiently simple noise models, optimal decoding of QEC codes can be
Ziyu Wan, Despoina Paschalidou, Ian Huang, Hongyu Liu
The increased demand for 3D data in AR/VR, robotics and gaming applications, gave rise to powerful generative pipelines capable of synthesizing high-quality 3D objects. Most of these models rely on the Score Distillation Sampling (SDS) algorithm to optimize a 3D representation such that the rendered image maintains a high likelihood as evaluated by a pre-tra
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu
We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified latent space, enabling training and generation across modalities. Second, for memory and training efficiency, we use a win
Bharath Raj Nagoor Kani, Hsin-Ying Lee, Sergey Tulyakov, Shubham Tulsiani
We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically rely on camera poses to geometrically aggregate information from input views, but are not robust in-the-wild when such
Chong Zhou, Xiangtai Li, Chen Change Loy, Bo Dai
This paper presents EdgeSAM, an accelerated variant of the Segment Anything Model (SAM), optimized for efficient execution on edge devices with minimal compromise in performance. Our approach involves distilling the original ViT-based SAM image encoder into a purely CNN-based architecture, better suited for edge devices. We carefully benchmark various distil
Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems
math.OCAndrea Angiuli, Jean-Pierre Fouque, Mathieu Laurière, Mengrui Zhang
We establish the convergence of the unified two-timescale Reinforcement Learning (RL) algorithm presented in a previous work by Angiuli et al. This algorithm provides solutions to Mean Field Game (MFG) or Mean Field Control (MFC) problems depending on the ratio of two learning rates, one for the value function and the other for the mean field term. Our proof
Yuzhe Yang, Haoran Zhang, Judy W Gichoya, Dina Katabi
As artificial intelligence (AI) rapidly approaches human-level performance in medical imaging, it is crucial that it does not exacerbate or propagate healthcare disparities. Prior research has established AI's capacity to infer demographic data from chest X-rays, leading to a key concern: do models using demographic shortcuts have unfair predictions across s
Alex Kulesza, Ananda Theertha Suresh, Yuyan Wang
Differential privacy is often studied under two different models of neighboring datasets: the add-remove model and the swap model. While the swap model is frequently used in the academic literature to simplify analysis, many practical applications rely on the more conservative add-remove model, where obtaining tight results can be difficult. Here, we study t
Ka Leong Cheng, Qiuyu Wang, Zifan Shi, Kecheng Zheng
Neural radiance fields, which represent a 3D scene as a color field and a density field, have demonstrated great progress in novel view synthesis yet are unfavorable for editing due to the implicitness. This work studies the task of efficient 3D editing, where we focus on editing speed and user interactivity. To this end, we propose to learn the color field
Ghazaleh Asghari, Jani A. Virtanen, Zhangjian Hu
Using the notion of integral distance to analytic functions, we give a characterization of Schatten class Hankel operators acting on doubling Fock spaces on the complex plane and use it to show that for $f\in L^{\infty}$, if $H_{f}$ is Hilbert-Schmidt, then so is $H_{\bar{f}}$. This property is known as the Berger-Coburn phenomenon. When $0<p\le 1$, we show
Fangfu Liu, Diankun Wu, Yi Wei, Yongming Rao
Recently, 3D content creation from text prompts has demonstrated remarkable progress by utilizing 2D and 3D diffusion models. While 3D diffusion models ensure great multi-view consistency, their ability to generate high-quality and diverse 3D assets is hindered by the limited 3D data. In contrast, 2D diffusion models find a distillation approach that achieve
Ava Pun, Gary Sun, Jingkang Wang, Yun Chen
Different outdoor illumination conditions drastically alter the appearance of urban scenes, and they can harm the performance of image-based robot perception systems if not seen during training. Camera simulation provides a cost-effective solution to create a large dataset of images captured under different lighting conditions. Towards this goal, we propose
Neerja Thakkar, Karttikeya Mangalam, Andrea Bajcsy, Jitendra Malik
Human trajectory prediction is typically posed as a zero-shot generalization problem: a predictor is learnt on a dataset of human motion in training scenes, and then deployed on unseen test scenes. While this paradigm has yielded tremendous progress, it fundamentally assumes that trends in human behavior within the deployment scene are constant over time. As
Building Domain-Specific LLMs Faithful To The Islamic Worldview: Mirage or Technical Possibility?
cs.AIShabaz Patel, Hassan Kane, Rayhan Patel
Large Language Models (LLMs) have demonstrated remarkable performance across numerous natural language understanding use cases. However, this impressive performance comes with inherent limitations, such as the tendency to perpetuate stereotypical biases or fabricate non-existent facts. In the context of Islam and its representation, accurate and factual repr
Junbum Cha, Wooyoung Kang, Jonghwan Mun, Byungseok Roh
In Multimodal Large Language Models (MLLMs), a visual projector plays a crucial role in bridging pre-trained vision encoders with LLMs, enabling profound visual understanding while harnessing the LLMs' robust capabilities. Despite the importance of the visual projector, it has been relatively less explored. In this study, we first identify two essential proj
Pratul P. Srinivasan, Stephan J. Garbin, Dor Verbin, Jonathan T. Barron
Existing UV mapping algorithms are designed to operate on well-behaved meshes, instead of the geometry representations produced by state-of-the-art 3D reconstruction and generation techniques. As such, applying these methods to the volume densities recovered by neural radiance fields and related techniques (or meshes triangulated from such fields) results in
Spherical higher order Fourier analysis over finite fields I: equidistribution for nilsequences
math.NTWenbo Sun
This paper is the first part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we prove a quantitative equidistribution theorem for polynomial sequences
Spherical higher order Fourier analysis over finite fields II: additive combinatorics for shifted ideals
math.ACWenbo Sun
This paper is the second part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we study additive combinatorial properties for shifted modules, i.e. the
Spherical higher order Fourier analysis over finite fields IV: an application to the Geometric Ramsey Conjecture
math.NTWenbo Sun
This paper is the fourth and the last part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the Geometric Ramsey Conjecture in the finite field setting. In this paper, we proof a conjecture of Graham on the Remsey properties fo
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu
Dense retrieval has become a prominent method to obtain relevant context or world knowledge in open-domain NLP tasks. When we use a learned dense retriever on a retrieval corpus at inference time, an often-overlooked design choice is the retrieval unit in which the corpus is indexed, e.g. document, passage, or sentence. We discover that the retrieval unit ch
David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar, Teresa Yeo
Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly versatile models in computer vision. In this paper, we take a step in this direction and propose a multimodal training scheme
Junwei Deng, Xirui Jiang, Shiyuan Zhang, Shichang Zhang
The rapid rise of generative AI has intensified copyright and economic tensions in creative industries, particularly in music. Current approaches addressing this challenge often focus on preventing infringement or establishing one-time licensing, which fail to provide the sustainable, recurring economic incentives necessary to maintain creative ecosystems. T
Beyond Classification: Definition and Density-based Estimation of Calibration in Object Detection
cs.CVTeodora Popordanoska, Aleksei Tiulpin, Matthew B. Blaschko
Despite their impressive predictive performance in various computer vision tasks, deep neural networks (DNNs) tend to make overly confident predictions, which hinders their widespread use in safety-critical applications. While there have been recent attempts to calibrate DNNs, most of these efforts have primarily been focused on classification tasks, thus ne
Rao Fu, Zehao Wen, Zichen Liu, Srinath Sridhar
Inspired by cognitive theories, we introduce AnyHome, a framework that translates any text into well-structured and textured indoor scenes at a house-scale. By prompting Large Language Models (LLMs) with designed templates, our approach converts provided textual narratives into amodal structured representations. These representations guarantee consistent and
Gaze Detection and Analysis for Initiating Joint Activity in Industrial Human-Robot Collaboration
cs.ROPooja Prajod, Matteo Lavit Nicora, Marta Mondellini, Giovanni Tauro
Collaborative robots (cobots) are widely used in industrial applications, yet extensive research is still needed to enhance human-robot collaborations and operator experience. A potential approach to improve the collaboration experience involves adapting cobot behavior based on natural cues from the operator. Inspired by the literature on human-human interac
Yixing Lao, Xiaogang Xu, Zhipeng Cai, Xihui Liu
Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed by off-the-shelf methods to supervise NeRF training. We desi
Vijeth Hebbar, Cedric Langbort
In many online sequential decision-making scenarios, a learner's choices affect not just their current costs but also the future ones. In this work, we look at one particular case of such a situation where the costs depend on the time average of past decisions over a history horizon. We first recast this problem with history dependent costs as a problem of d
Shangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo
Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains challenging due to the high demands for output fidelity and temporal consistency, which is complicated by the inherent randomn
Ruihan Yang, Yejin Kim, Rose Hendrix, Aniruddha Kembhavi
Recent advancements in robotics have enabled robots to navigate complex scenes or manipulate diverse objects independently. However, robots are still impotent in many household tasks requiring coordinated behaviors such as opening doors. The factorization of navigation and manipulation, while effective for some tasks, fails in scenarios requiring coordinated
Lev V. Utkin, Danila Y. Eremenko, Andrei V. Konstantinov
A new method called the Survival Beran-based Neural Importance Model (SurvBeNIM) is proposed. It aims to explain predictions of machine learning survival models, which are in the form of survival or cumulative hazard functions. The main idea behind SurvBeNIM is to extend the Beran estimator by incorporating the importance functions into its kernels and by im
Seongjoon Kang
Due to the high complexity of geometry-deterministic wireless channel modeling and the difficulty in its implementation, geometry-based stochastic channel modeling (GBSM) approaches have been used to evaluate wireless systems. This paper introduces a new method to model any GBSM by training a generative neural network using images formed by channel parameter
Spherical higher order Fourier analysis over finite fields III: a spherical Gowers inverse theorem
math.NTWenbo Sun
This paper is the third part of the series "Spherical higher order Fourier analysis over finite fields", aiming to develop the higher order Fourier analysis method along spheres over finite fields, and to solve the geometric Ramsey conjecture in the finite field setting. In this paper, we prove an inverse theorem over finite field for spherical Gowers norms,
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda
Transformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness an
Wentao Tang
This work proposes a data-driven approach for bifurcation analysis in nonlinear systems when the governing differential equations are not available. Specifically, regularized regression with barrier terms is used to learn a homeomorphism that transforms the underlying system to a reference linear dynamics -- either an explicit reference model with desired qu
Kevin Coakley, Christine R. Kirkpatrick, Odd Erik Gundersen
Reproducing published deep learning papers to validate their conclusions can be difficult due to sources of irreproducibility. We investigate the impact that implementation factors have on the results and how they affect reproducibility of deep learning studies. Three deep learning experiments were ran five times each on 13 different hardware environments an
Jiyan He, Weitao Feng, Yaosen Min, Jingwei Yi
The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks like creation of harmful substances, or circumvention of established regulations. In this study, we aim to raise awaren
New black hole mergers in the LIGO-Virgo O3 data from a gravitational wave search including higher-order harmonics
gr-qcDigvijay Wadekar, Javier Roulet, Tejaswi Venumadhav, Ajit Kumar Mehta
Nearly all of the previous gravitational wave (GW) searches in the LIGO-Virgo data included GW waveforms with only the dominant quadrupole harmonic, i.e., omitting higher-order harmonics which are predicted by general relativity. We improved the IAS pipeline by efficiently introducing higher harmonics in the GW templates using the techniques in Wadekar et al
Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang
Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggregation of datasets to enhance data volume and diversity. Ho
Angsuman Das
In this paper, we introduce and study the iterates of the following family of functions $\varphi_k$ defined on natural numbers which exhibits nice properties. $$\varphi_k(x)=\left\lbrace \begin{array}{ll} x+k, & \mbox{ if $x$ is prime;}\\ \mbox{largest prime divisor of $x$,} & \mbox{ if $x$ is composite;} \end{array} \right.$$ In particular, we study the per
Denis I. Saveliev
We consider a certain class of infinitary rules of inference, called here restriction rules, using of which allows us to deduce complete theories of given models. The first instance of such rules was the $\omega$-rule introduced by Hilbert, and generalizations of the $\omega$-rule were first considered by Henkin. Later on Barwise showed that within countable
Giordano De Marzo, Luciano Pietronero, David Garcia
Scale-free networks are one of the most famous examples of emergent behavior and are ubiquitous in social systems, especially online social media in which users can follow each other. By analyzing the interactions of multiple generative agents using GPT3.5-turbo as a language model, we demonstrate their ability to not only mimic individual human linguistic b
AttenScribble: Attentive Similarity Learning for Scribble-Supervised Medical Image Segmentation
cs.CVMu Tian, Qinzhu Yang, Yi Gao
The success of deep networks in medical image segmentation relies heavily on massive labeled training data. However, acquiring dense annotations is a time-consuming process. Weakly-supervised methods normally employ less expensive forms of supervision, among which scribbles started to gain popularity lately thanks to its flexibility. However, due to lack of
Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism
cs.CVGeorgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos
Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being conditioned on audio clips. However, having the ability to synthesize talking humans from text transcriptions rather than audio i
Lior Gishboliner, Zhihan Jin, Benny Sudakov
Many well-studied problems in extremal combinatorics deal with the maximum possible size of a family of objects in which every pair of objects satisfies a given restriction. One problem of this type was recently raised by Alon, Gujgiczer, K\"orner, Milojevi\'c and Simonyi. They asked to determine the maximum size of a family $\mathcal{G}$ of graphs on $[n]$,
Vladislav Makarov, Marat Movsin
GF(2)-grammars are a somewhat recently introduced grammar family that have some unusual algebraic properties and are closely connected to unambiguous grammars. In "Bounded languages described by GF(2)-grammars", Makarov proved a necessary condition for subsets of $a_1^* a_2^* \cdots a_k^*$ to be described by some GF(2)-grammar. By extending these methods fur
Matthew S. Schmitt, Maciej Koch-Janusz, Michel Fruchart, Daniel S. Seara
The dynamics of many-body systems can often be captured in terms of only a few relevant variables. Mathematical and numerical approaches exist to identify these variables by exploiting a separation of time scales between slow relevant and fast irrelevant variables, but such a separation of scales is not always obvious or even available. In this work, we intr
Haoyang He, Jiangning Zhang, Hongxu Chen, Xuhai Chen
Reconstruction-based approaches have achieved remarkable outcomes in anomaly detection. The exceptional image reconstruction capabilities of recently popular diffusion models have sparked research efforts to utilize them for enhanced reconstruction of anomalous images. Nonetheless, these methods might face challenges related to the preservation of image cate
A Markovian Gau\ss \, Inequality for Asymmetric Deviations from the Mode of Symmetric Unimodal Distributions
math.PRChris A. J. Klaassen
For a random variable with a unimodal distribution and finite second moment Gau\ss \, (1823) proved a sharp bound on the probability of the random variable to be outside a symmetric interval around its mode. An alternative proof for it is given based on Khintchine's representation of unimodal random variables. Analogously, a sharp inequality is proved for th
Jinming Li, Shihao Wu, Chengyu Cui, Gongjun Xu
Latent space models are powerful statistical tools for modeling and understanding network data. While the importance of accounting for uncertainty in network analysis has been well recognized, the current literature predominantly focuses on point estimation and prediction, leaving the statistical inference of latent space models an open question. This work a
Jyoti Prakash Saha
Let $\Gamma$ be a Cayley graph, or a Cayley sum graph, or a twisted Cayley graph, or a twisted Cayley sum graph, or a vertex-transitive graph. Suppose $\Gamma$ is undirected and non-bipartite. Let $\mu$ (resp. $\mu_2$) denote the smallest (resp. the second largest) eigenvalue of the normalized adjacency operator of $\Gamma$, and $d$ denote the degree of $\Ga
Tirasan Khandhawit, Puttipong Pongtanapaisan, Athibadee Wasun
Two isomorphic graphs can have inequivalent spatial embeddings in 3-space. In this way, an isomorphism class of graphs contains many spatial graph types. A common way to measure the complexity of a spatial graph type is to count the minimum number of straight sticks needed for its construction in 3-space. In this paper, we give estimates of this quantity by
Paolo M Bassani, Joao Magueijo
In unimodular-like theories, the constants of nature are demoted from pre-given parameters to phase space variables. Their canonical duals provide physical time variables. We investigate how this interacts with an alternative approach to varying constants, where they are replaced by dynamical scalar fields. Specifically we investigate the Brans-Dicke theory
Zheng-Duo Fan, Xiao-Qi Sun, Jing-Yuan Chen
Thermal transport has been used to probe the nature of $\alpha$-RuCl$_3$, an important candidate of Kitaev material. Two remarkable observations were made under applied magnetic fields at low temperatures, and have stimulated extensive discussions. One is a sizable thermal Hall effect, and the other is an apparent "oscillation" of the longitudinal thermal co
Mikkel Bennedsen, Eric Hillebrand, Jingying Zhou Lykke
We propose a non-linear state-space model to examine the relationship between CO$_2$ emissions, energy sources, and macroeconomic activity, using data from 1971 to 2019. CO$_2$ emissions are modeled as a weighted sum of fossil fuel use, with emission conversion factors that evolve over time to reflect technological changes. GDP is expressed as the outcome of
M. Poisson, M. López Fuentes, C. H. Mandrini, F. Grings
Active regions (ARs) appear in the solar atmosphere as a consequence of the emergence of magnetic flux ropes (FRs). Due to the presence of twist, the photospheric line-of-sight (LOS) magnetograms of emerging ARs show an elongation of the polarities known as magnetic tongues. These tongues can affect the estimation of tilt angles during their emergence phase.
Guglielmo Camporese, Alessandro Bergamo, Xunyu Lin, Joseph Tighe
Early action recognition is an important and challenging problem that enables the recognition of an action from a partially observed video stream where the activity is potentially unfinished or even not started. In this work, we propose a novel model that learns a prototypical representation of the full action for each class and uses it to regularize the arc
Radical pairs and superoxide amplification can explain magnetic field effects on planarian regeneration
physics.bio-phRishabh, Hadi Zadeh-Haghighi, Christoph Simon
Weak magnetic field exposure can affect many biological processes across a wide range of living organisms. Recently, it has been observed that weak magnetic fields can modulate reactive oxygen species (ROS) concentration, affecting regeneration in planaria. These effects show unusual nonlinear dependence on magnetic field strength, including a sign change. I
Tejes Gaertner, Jared Reiten
In this work, we construct a new data type for hadronic jets in which the traditional point-cloud representation is transformed into a simplicial complex consisting of vertices, or 0-simplexes. An angular resolution scale, $r$, is then drawn about each vertex, forming balls about hadrons. As $r$ grows, the overlap of balls form 2- and 3-point connections, th
Stefano Meda, Federico Santagati
We introduce the centred and the uncentred triangular maximal operators $\mathcal T$ and $\mathcal U$, respectively, on any locally finite tree in which each vertex has at least three neighbours. We prove that both $\mathcal T$ and $\mathcal U$ are bounded on $L^p$ for every $p$ in $(1,\infty]$, that $\mathcal T$ is also bounded on $L^1(\mathfrak T)$, and th
Aditya Prakash, Arjun Gupta, Saurabh Gupta
Objects undergo varying amounts of perspective distortion as they move across a camera's field of view. Models for predicting 3D from a single image often work with crops around the object of interest and ignore the location of the object in the camera's field of view. We note that ignoring this location information further exaggerates the inherent ambiguity
The Change-Driver Account of Scientific Discovery: Philosophical and Historical Dimensions of the Discovery of the Expanding Universe
physics.hist-phPatrick M. Duerr, Abigail Holmes Mills
This paper critically examines the models of scientific discovery propounded by Kuhn, McArthur, Hudson, and Schindler. As an alternative, we proffer the $\unicode{x201c}$change$\unicode{x2013}$driver model$\unicode{x201c}$. It conceives of discoveries as problems or solutions to problems that have epistemically advanced science. Here we take a problem to be
Aaron L. Sarvet, Mats J. Stensrud, Lan Wen
We formalize an interpretational error that is common in statistical causal inference, termed identity slippage. This formalism is used to describe historically-recognized fallacies, and analyse a fast-growing literature in statistics and applied fields. We conducted a systematic review of natural language claims in the literature on stochastic mediation par
Thomas Foster, Ioana Croitoru, Robert Dorfman, Christoffer Edlund
In this work, we address in-context learning (ICL) for the task of image segmentation, introducing a novel approach that adapts a modern Video Object Segmentation (VOS) technique for visual in-context learning. This adaptation is inspired by the VOS method's ability to efficiently and flexibly learn objects from a few examples. Through evaluations across a r
Anish Chakrabarty, Arkaprabha Basu, Swagatam Das
Variational Autoencoders (VAEs) have been a pioneering force in the realm of deep generative models. Amongst its legions of progenies, Wasserstein Autoencoders (WAEs) stand out in particular due to the dual offering of heightened generative quality and a strong theoretical backbone. WAEs consist of an encoding and a decoding network forming a bottleneck with
Alejandro Cárdenas-Avendaño, Aaron Held
General relativity's prediction that all black holes are described by the Kerr metric, irrespective of their size, can now be empirically tested using electromagnetic observations of supermassive black holes and gravitational waves from mergers of stellar-mass black holes. In this work, we focus on the electromagnetic side of this test and quantify the const
Alexander Roth
The decarbonization of buildings requires the phase-out of fossil fuel heating systems. Heat pumps are considered a crucial technology to supply a substantial part of heating energy for buildings. Yet, their introduction is not without challenges, as heat pumps generate additional electricity demand as well as peak loads. To better understand these challenge
L. A. Lessa, R. V. Maluf, J. E. G. Silva, C. A. S. Almeida
Einstenian cubic gravity (ECG) is a modified theory of gravity constructed with cubic contractions of the curvature tensor. This theory has the remarkable feature of having the same two propagating degrees of freedom of Einstein gravity (EG), at the perturbative level on maximally symmetric spacetimes. The additional unstable modes steaming from the higher o
QuickQuakeBuildings: Post-earthquake SAR-Optical Dataset for Quick Damaged-building Detection
eess.IVYao Sun, Yi Wang, Michael Eineder
Quick and automated earthquake-damaged building detection from post-event satellite imagery is crucial, yet it is challenging due to the scarcity of training data required to develop robust algorithms. This letter presents the first dataset dedicated to detecting earthquake-damaged buildings from post-event very high resolution (VHR) Synthetic Aperture Radar
Hidenobu Matsuki, Riku Murai, Paul H. J. Kelly, Andrew J. Davison
We present the first application of 3D Gaussian Splatting in monocular SLAM, the most fundamental but the hardest setup for Visual SLAM. Our method, which runs live at 3fps, utilises Gaussians as the only 3D representation, unifying the required representation for accurate, efficient tracking, mapping, and high-quality rendering. Designed for challenging mon
Mauro Barbieri
This document details the first public data release of the HARPS radial velocities catalog. This data release aims to provide the astronomical community with a catalog of radial velocities obtained with spectroscopic observations acquired from 2003 to 2023 with the High Accuracy Radial Velocity Planet Searcher (HARPS) spectrograph installed at the ESO 3.6m t
Avi Singh, John D. Co-Reyes, Rishabh Agarwal, Ankesh Anand
Fine-tuning language models~(LMs) on human-generated data remains a prevalent practice. However, the performance of such models is often limited by the quantity and diversity of high-quality human data. In this paper, we explore whether we can go beyond human data on tasks where we have access to scalar feedback, for example, on math problems where one can v
Aybike Çatal-Özer, Keremcan Doğan, Cem Yetişmişoğlu
We extend the notion of Lie bialgebroids for more general bracket structures used in string and M theories. We formalize the notions of calculus and dual calculi on algebroids. We achieve this by reinterpreting the main results of the matched pairs of Leibniz algebroids. By examining a rather general set of fundamental algebroid axioms, we present the compat
Aditya Prakash, Ruisen Tu, Matthew Chang, Saurabh Gupta
3D hand pose estimation in everyday egocentric images is challenging for several reasons: poor visual signal (occlusion from the object of interaction, low resolution & motion blur), large perspective distortion (hands are close to the camera), and lack of 3D annotations outside of controlled settings. While existing methods often use hand crops as input to
Implementing hosting capacity analysis in distribution networks: Practical considerations, advancements and future directions
eess.SYU. Singh, A. Al-Durra
Hosting capacity analysis is essential for effective integration of distributed energy resources into distribution systems. This paper discusses hosting capacity analysis with emphasis on various aspects affecting the process. This paper addresses key research gaps, aiming to improve the accuracy, scalability, and practicality of hosting capacity estimation.
Dashiell Stander, Qinan Yu, Honglu Fan, Stella Biderman
The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep neural networks, but all have substantial limitations. Algorithmic tasks have proven to be a fruitful test ground for interpreting a neural network end-to-end. Building on previous
Ruochen Dai, Michael Lee, Patrick Hoey, Weimin Fu
As the complexity of logic designs increase, new avenues for testing digital hardware becomes necessary. Fuzz Testing (fuzzing) has recently received attention as a potential candidate for input vector generation on hardware designs. Using this technique, a fuzzer is used to generate an input to a logic design. Using a simulation engine, the logic design is
Samyukta Sethuraman, Ankur Bansal, Setareh Mardan, Mauricio G. C. Resende
Amazon Locker is a self-service delivery or pickup location where customers can pick up packages and drop off returns. A basic first-come-first-served policy for accepting package delivery requests to lockers results in lockers becoming full with standard shipping speed (3-5 day shipping) packages, and leaving no space left for expedited packages which are m
Zhezheng Hao, Feiping Nie, Rong Wang
Support Vector Machine (SVM) stands out as a prominent machine learning technique widely applied in practical pattern recognition tasks. It achieves binary classification by maximizing the "margin", which represents the minimum distance between instances and the decision boundary. Although many efforts have been dedicated to expanding SVM for multi-class cas
Max Hallgren
We show that any tangent cone of a singular shrinking K\"ahler-Ricci soliton is a normal affine algebraic variety. Moreover, the regular set of such a tangent cone in the metric sense coincides with the regular set in the algebraic sense. Along the way, we give a parabolic proof of H\"ormander's $L^{2}$ estimate, which can be used to solve the $\overline{\pa
Kushal Bose, Swagatam Das
Graph Transformers (GTs) facilitate the comprehension of complex relationships on graph-structured data by leveraging self-attention of the possible pairs of nodes. The structural information or inductive bias of the input graph is provided as positional encodings to the GT. The positional encodings are mostly Euclidean and are not able to capture the comple
Zhen Xu, Tao Xie, Sida Peng, Haotong Lin
Volumetric video is a technology that digitally records dynamic events such as artistic performances, sporting events, and remote conversations. When acquired, such volumography can be viewed from any viewpoint and timestamp on flat screens, 3D displays, or VR headsets, enabling immersive viewing experiences and more flexible content creation in a variety of
Lioba Heimbach, Quentin Kniep, Yann Vonlanthen, Roger Wattenhofer
Ethereum introduced Transaction Access Lists (TALs) in 2020 to optimize gas costs during transaction execution. In this work, we present a comprehensive analysis of TALs in Ethereum, focusing on adoption, quality, and gas savings. Analyzing a full month of mainnet data with 31,954,474 transactions, we found that only 1.46% of transactions included a TAL, eve
ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems
cs.CVDenis Zavadski, Johann-Friedrich Feiden, Carsten Rother
The field of image synthesis has made tremendous strides forward in the last years. Besides defining the desired output image with text-prompts, an intuitive approach is to additionally use spatial guidance in form of an image, such as a depth map. In state-of-the-art approaches, this guidance is realized by a separate controlling model that controls a pre-t
SHINE Collaboration, H. Adhikary, P. Adrich, K. K. Allison
Strong interactions preserve an approximate isospin symmetry between up ($u$) and down ($d$) quarks, part of the more general flavor symmetry. In the case of $K$ meson production, if this isospin symmetry were exact, it would result in equal numbers of charged ($K^+$ and $K^-$) and neutral ($K^0$ and $\overline K^{\,0}$) mesons in the final state. Here, we r
Takahide Yoshida, Atsushi Masumori, Takashi Ikegami
We report the development of Alter3, a humanoid robot capable of generating spontaneous motion using a Large Language Model (LLM), specifically GPT-4. This achievement was realized by integrating GPT-4 into our proprietary android, Alter3, thereby effectively grounding the LLM with Alter's bodily movement. Typically, low-level robot control is hardware-depen
Maciej Grzeszczuk, Kinga Skorupska
In this article, we report the pilot results of a survey study (N=1036) related to social attitudes towards the early digital heritage. On the basis of the answers, we consider what constitutes early digital artifacts (EDA) and outline how knowledge about them can be useful. We explore attitudes toward the historical and cultural importance of various EDAs a
M. Majid Butt, Nitin R. Mangalvedhe, Nuno K. Pratas, Johannes Harrebek
Ambient internet of things (IoT) is the network of devices which harvest energy from ambient sources for powering their communication. After decades of research on operation of these devices, Third Generation Partnership Project (3GPP) has started discussing energy harvesting technology in cellular networks to support massive deployment of IoT devices at low
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz
We present Monocular Neural Parametric Head Models (MonoNPHM) for dynamic 3D head reconstructions from monocular RGB videos. To this end, we propose a latent appearance space that parameterizes a texture field on top of a neural parametric model. We constrain predicted color values to be correlated with the underlying geometry such that gradients from RGB ef
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models
cs.CVYuzhou Huang, Liangbin Xie, Xintao Wang, Ziyang Yuan
Current instruction-based editing methods, such as InstructPix2Pix, often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this, this paper introduces SmartEdit, a novel approach to instruction-based image editing that leverages Multimodal Large Language Models (
Shufan Li, Harkanwar Singh, Aditya Grover
The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two directions: instruction tuning with text-based prompts and multi-modal conditioning. However, these works make one or more unnatural assumptions
Subhajit Dutta Chowdhury, Zhiyu Ni, Qingyuan Peng, Souvik Kundu
Graph Lottery Tickets (GLTs), comprising a sparse adjacency matrix and a sparse graph neural network (GNN), can significantly reduce the inference latency and compute footprint compared to their dense counterparts. Despite these benefits, their performance against adversarial structure perturbations remains to be fully explored. In this work, we first invest
Direct molecular gas dynamics simulations of re-entry vehicles via the Boltzmann equation
physics.flu-dynTarik Dzanic, Luigi Martinelli
This work explores the feasibility of performing three-dimensional molecular gas dynamics simulations of hypersonic flows such as re-entry vehicles through directly solving the six-dimensional nonlinear Boltzmann equation closed with the BGK (Bhatnagar-Gross-Krook) collision model. Through the combination of high-order unstructured spatial discretizations an
Guanyu Huang, Roger K. Moore
Personalisation is essential to achieve more acceptable and effective results in human-robot interaction. Placing users in the central role, many studies have focused on enhancing the abilities of social robots to perceive and understand users. However, little is known about improving user perceptions and interpretation of a social robot in spoken interactio
Luca Marannino
We generalize and simplify the constructions of Darmon-Rotger and Hsieh of an unbalanced triple product $p$-adic $L$-function $\mathscr{L}_p^f(\boldsymbol{f},\boldsymbol{g},\boldsymbol{h})$ attached to a triple $(\boldsymbol{f},\boldsymbol{g},\boldsymbol{h})$ of $p$-adic families of modular forms, allowing more flexibility for the choice of $\boldsymbol{g}$
Francesco Leofante, Nico Potyka
Counterfactual explanations shed light on the decisions of black-box models by explaining how an input can be altered to obtain a favourable decision from the model (e.g., when a loan application has been rejected). However, as noted recently, counterfactual explainers may lack robustness in the sense that a minor change in the input can cause a major change
Richard Kadison, Simon Levin, Zhe Liu
Based on the success of a well-known method for solving higher order linear differential equations, a study of two of the most important mathematical features of that method, viz. the null spaces and commutativity of the product of unbounded linear operators on a Hilbert space is carried out. A principle is proved describing solutions for the product of such
Adrian de Wynter, Xun Wang, Qilong Gu, Si-Qing Chen
Modern large language models (LLMs) are capable of interpreting input strings as instructions, or prompts, and carry out tasks based on them. Unlike traditional learners, LLMs cannot use back-propagation to obtain feedback, and condition their output in situ in a phenomenon known as in-context learning (ICL). Many approaches to prompting and pre-training the
Hong-Xing Yu, Yang Zheng, Yuan Gao, Yitong Deng
We study recovering fluid density and velocity from sparse multiview videos. Existing neural dynamic reconstruction methods predominantly rely on optical flows; therefore, they cannot accurately estimate the density and uncover the underlying velocity due to the inherent visual ambiguities of fluid velocity, as fluids are often shapeless and lack stable visu