Skip to content

December 2024 arXiv papers — page 171

Showing 17,00117,100 of 20,868 papers

  1. Jice Zeng, Yuanzhe Wang, Alexandre M. Tartakovsky, David Barajas-Solano

    We present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary

  2. Saeed Mahdavifar, Hadi Cheraghi, Kourosh Afrousheh

    We investigate the spin-squeezing behavior under thermal effects in a one-dimensional transverse field XY model with spin-1/2. The exact solution of the model helps us to compute the spin-squeezing parameter as a function of temperature and also in all excited states with higher energy than the ground state. We find that below the thermal factorized field, h

  3. Atit Pokharel, Pratik Sapkota, Dilip Sapkota, Shashank Dahal

    LoRa technology has garnered significant interest in the Information and Communications Technology (ICT) field in recent years due to its ability to operate at low power while maintaining effective communication. Despite gaining attention, LoRa technology faces challenges in effectively facilitating communication in rural settings due to specific transmissio

  4. Oton Vázquez Doce, Dimitar Mihaylov, Laura Fabbietti

    The femtoscopy correlation between positively charged kaons and deuterons measured in pp collisions at $\sqrt{s}=$ 13 TeV at the LHC with the ALICE experiment is employed to determine the upper limit for the time delay of the deuteron emission with respect to all other hadrons. Two scenarios are considered: the first assumes that deuterons form following the

  5. Francisco Divi, Jeff Murugan, Dario Rosa

    We investigate the charging dynamics of Sachdev-Ye-Kitaev (SYK) models as quantum batteries, highlighting their capacity to achieve quantum charging advantages. By analytically deriving the scaling of the charging power in SYK batteries, we identify the two key mechanisms underlying this advantage: the use of operators scaling extensively with system size $N

  6. Erwin T. Lau, Daisuke Nagai, Ákos Bogdán, Isabel Medlock

    The circumgalactic medium (CGM) around massive galaxies plays a crucial role in regulating star formation and feedback. Using the CAMELS simulation suite, we develop emulators for the X-ray surface brightness profile and the X-ray luminosity--stellar mass scaling relation to investigate how stellar and AGN feedback shape the X-ray properties of the hot CGM.

  7. Maria Mihaela Trusca, Mingxiao Li, Marie-Francine Moens

    Text-based image editing is typically approached as a static task that involves operations such as inserting, deleting, or modifying elements of an input image based on human instructions. Given the static nature of this task, in this paper, we aim to make this task dynamic by incorporating actions. By doing this, we intend to modify the positions or posture

  8. Ivo Labbe, Jenny E. Greene, Jorryt Matthee, Helena Treiber

    We present a detailed exploration of the most optically-luminous Little Red Dot ($L_{H\alpha}=10^{44}$erg/s, $L_V=10^{45}$erg/s, F444W=22AB) found to date. Located in the Abell 2744 field, source A744-45924 was observed by NIRSpec/PRISM with ultradeep spectroscopy reaching SNR$\sim$100pix$^{-1}$, high-resolution 3-4 micron NIRCam/Grism spectroscopy, and NIRC

  9. Purnendu Das, Sumilan Banerjee

    Control and manipulation of quantum states by measurements and bath engineering in open quantum systems have emerged as new paradigms in many-body physics. Here, taking a prototypical example of Josephson junction arrays (JJAs), we show how repetitive monitoring through continuous weak measurements and feedback can transform an insulating state in these syst

  10. Subash Katel, Haoyang Li, Zihan Zhao, Raghav Kansal

    In high energy physics, self-supervised learning (SSL) methods have the potential to aid in the creation of machine learning models without the need for labeled datasets for a variety of tasks, including those related to jets -- narrow sprays of particles produced by quarks and gluons in high energy particle collisions. This study introduces an approach to l

  11. C. Vargas, L. Pereira, A. Delgado

    Quantum state estimation is important for various quantum information processes, including quantum communications, computation, and metrology, which require the characterization of quantum states for evaluation and optimization. We present a three-stage adaptive method for estimating arbitrary $d$-dimensional pure quantum states using locally informationally

  12. Marek Gluza, Jeongrak Son, Bi Hong Tiang, René Zander

    Efficiently preparing approximate ground-states of large, strongly correlated systems on quantum hardware is challenging and yet nature is innately adept at this. This has motivated the study of thermodynamically inspired approaches to ground-state preparation that aim to replicate cooling processes via imaginary-time evolution. However, synthesizing quantum

  13. Antonio J. Porras-Valverde, John C. Forbes

    As star-forming galaxies approach or exceed a stellar mass around $10^{11} M_\odot$, they are increasingly likely to be quenched in a process generically called mass quenching. Central galaxies, which are quenched via mass rather than environmental quenching, therefore accumulate in a peak around this characteristic mass. While a number of processes may infl

  14. Luke Finnerty, Yinzi Xin, Jerry W. Xuan, Julie Inglis

    We present Keck/KPIC phase II $K$-band observations of the non-transiting hot Jupiter HD 143105 b. Using a cross-correlation approach, we make the first detection of the planetary atmosphere at $K_p = 185^{+11}_{-13}\rm km\ s^{-1}$ and an inferior conjunction time 2.5 hours before the previously-published ephemeris. The retrieved $K_p$ value, in combination

  15. Maxime Gadioux, Hangzhi Wang

    In recent years there have been many studies on exactly solvable black hole mergers, based on a model by Emparan and Martinez where the mass of one black hole is blown up to infinity. Here we replace the large black hole by a cosmological horizon, and study how it merges with a black hole in the Schwarzschild-de Sitter spacetime by considering an observer po

  16. Nicolás Bernal, Chee Sheng Fong, Óscar Zapata

    The parameter space of freeze-in dark matter (DM) with mass $m_\chi$ through light dark photon (``minimal freeze-in DM'') is currently being probed by direct detection experiments through electron and nuclear recoil. Exploring the DM production in the mass range $10^{-2}~{\rm MeV} < m_\chi < 10^3$ TeV, we quantify the impact of quantum statistics and the reh

  17. Wolfgang Altmannshofer, Admir Greljo

    The flavor puzzles remain among the most compelling open questions in particle physics. The striking hierarchies observed in the masses and mixing of charged fermions define the Standard Model (SM) flavor puzzle, a profound structural enigma pointing to physics beyond the SM. Simultaneously, the absence of deviations from SM predictions in precision measurem

  18. Jianwei Lyu, George H. Rieke, Meredith Stone, Jane Morrison

    The majority of most luminous quasars during the epoch of reionization accrete near or above the Eddington limit, marking the vigorous growth of primitive supermassive black holes (SMBHs). However, their subsequent evolution and environmental impact remain poorly characterized. We present JWST/NIRSpec prism IFU observations of HSC J2239+0207, a low-luminosit

  19. Pierre A. Pantaleon, Zhen Zhan, Siddhartha E. Morales, Gerardo G. Naumis

    We investigate the electronic properties of two-dimensional electron gases (2DEGs) subjected to a periodic patterned gate. By incorporating the superlattice (SL) potential induced by patterning into the Schrodinger equation, we develop a methodology for obtaining exact analytical solutions. These solutions enable us to construct a comprehensive phase diagram

  20. Eshita Banerjee, Sowgat Muzahid, Joop Schaye, Sebastiano Cantalupo

    According to modern cosmological models, galaxies are embedded within cosmic filaments, which supply a continuous flow of pristine gas, fueling star formation and driving their evolution. However, due to their low density, the direct detection of diffuse gas in cosmic filaments remains elusive. Here, we report the discovery of an extremely metal-poor ($[ X/H

  21. Matthias Heinz, Martin Hoferichter, Takayuki Miyagi, Frederic Noël

    The rate for $\mu\to e$ conversion in nuclei is set to provide the most stringent test of lepton-flavor symmetry and a window into physics beyond the Standard Model. However, to disentangle new lepton-flavor-violating interactions, in combination with information from $\mu\to e\gamma$ and $\mu\to 3e$, it is critical that uncertainties at each step of the ana

  22. Anjishnu Bose, Arun Paramekanti

    Recent work has shown that the honeycomb lattice spin-$1/2$ $J_1$-$J_3$ XY model, with nearest-neighbor ferromagnetic exchange $J_1$ and frustration induced by third-neighbor antiferromagnetic exchange $J_3$, may be relevant to a wide range of cobaltate materials. We explore a variational Monte Carlo study of Gutzwiller projected wavefunctions for this model

  23. V. A. Bronner, F. R. N. Schneider, Ph. Podsiadlowski, F. K. Roepke

    One-dimensional (1D) methods for simulating the common-envelope (CE) phase offer advantages over three-dimensional (3D) simulations regarding their computational speed and feasibility. We present the 1D CE method from Bronner et al. (2024), including the results of the CE simulations of an asymptotic giant branch star donor. We further test this method in th

  24. Isaac H. Laseter, Michael V. Maseda, Charlotte Simmonds, Ryan Endsley

    Early JWST photometric studies discovered a population of UV faint ($\rm <L^{*}_{UV}$) $z \sim 6.5-8$ Lyman break galaxies with spectral energy distributions implying young ages ($\sim10$ Myr) yet relatively weak H$\beta$+[OIII] equivalent widths ($\rm EW_{H\beta+[OIII]} \approx 400$\r{A}). These galaxies seemingly contradict the implicit understanding that

  25. Kimihiko Nakajima, Masami Ouchi, Yuki Isobe, Yi Xu

    Using the Subaru/FOCAS IFU capability, we examine the spatially resolved relationships between gas-phase metallicity, stellar mass, and star-formation rate surface densities (Sigma_* and Sigma_SFR, respectively) in extremely metal-poor galaxies (EMPGs) in the local universe. Our analysis includes 24 EMPGs, comprising 9,177 spaxels, which span a unique parame

  26. G. Yucel, V Bakis, R. Canbay, N. Alan

    In this study we present a detailed analysis of CN Lyn, an overlooked triple star system, by combining spectroscopic data from the literature, photometric \textit{TESS} data, and kinematic techniques. We updated the fundamental parameters of the known eclipsing components in the system with high precision. The chemical composition of both eclipsing component

  27. Luca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia

    We introduce Stereo Anywhere, a novel stereo-matching framework that combines geometric constraints with robust priors from monocular depth Vision Foundation Models (VFMs). By elegantly coupling these complementary worlds through a dual-branch architecture, we seamlessly integrate stereo matching with learned contextual cues. Following this design, our frame

  28. Vinayak Gupta, Yunze Man, Yu-Xiong Wang

    Recent advances in diffusion models have revolutionized 2D and 3D content creation, yet generating photorealistic dynamic 4D scenes remains a significant challenge. Existing dynamic 4D generation methods typically rely on distilling knowledge from pre-trained 3D generative models, often fine-tuned on synthetic object datasets. Consequently, the resulting sce

  29. Hanzhe Hu, Tianwei Yin, Fujun Luan, Yiwei Hu

    We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Te

  30. Sharath Girish, Tianye Li, Amrita Mazumdar, Abhinav Shrivastava

    Online free-viewpoint video (FVV) streaming is a challenging problem, which is relatively under-explored. It requires incremental on-the-fly updates to a volumetric representation, fast training and rendering to satisfy real-time constraints and a small memory footprint for efficient transmission. If achieved, it can enhance user experience by enabling novel

  31. Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang

    Visual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention. This paper introduces NVILA, a family of open VLMs designed to jointly optimize efficiency and accuracy. Building on top of VILA, we improve its model architecture by first scaling up the spatial and temporal r

  32. Senqiao Yang, Yukang Chen, Zhuotao Tian, Chengyao Wang

    Recent advancements in vision-language models have enhanced performance by increasing the length of visual tokens, making them much longer than text tokens and significantly raising computational costs. However, we observe that the visual tokens generated by popular vision encoders, such as CLIP and SigLIP, contain significant redundancy. To address this, we

  33. Sophie Greenwood, Sudalakshmee Chiniah, Nikhil Garg

    In the basic recommendation paradigm, the most (predicted) relevant item is recommended to each user. This may result in some items receiving lower exposure than they "should"; to counter this, several algorithmic approaches have been developed to ensure item fairness. These approaches necessarily degrade recommendations for some users to improve outcomes fo

  34. Zhongkai Shangguan, Zanming Huang, Eshed Ohn-Bar, Ola Ozernov-Palchik

    Models for student reading performance can empower educators and institutions to proactively identify at-risk students, thereby enabling early and tailored instructional interventions. However, there are no suitable publicly available educational datasets for modeling and predicting future reading performance. In this work, we introduce the Enhanced Core Rea

  35. Chang Liu, Viraj Shah, Aiyu Cui, Svetlana Lazebnik

    This paper introduces UnZipLoRA, a method for decomposing an image into its constituent subject and style, represented as two distinct LoRAs (Low-Rank Adaptations). Unlike existing personalization techniques that focus on either subject or style in isolation, or require separate training sets for each, UnZipLoRA disentangles these elements from a single imag

  36. Ben Kaye, Tomas Jakab, Shangzhe Wu, Christian Rupprecht

    The choice of data representation is a key factor in the success of deep learning in geometric tasks. For instance, DUSt3R recently introduced the concept of viewpoint-invariant point maps, generalizing depth prediction and showing that all key problems in the 3D reconstruction of static scenes can be reduced to predicting such point maps. In this paper, we

  37. Zhengqi Li, Richard Tucker, Forrester Cole, Qianqian Wang

    We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large amounts of parallax. Such methods tend to produce erroneous

  38. Chaoyang Wang, Peiye Zhuang, Tuan Duc Ngo, Willi Menapace

    We propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames from the same viewpoint. We propose a novel two-stream architecture. One stream performs viewpoint updates on columns, an

  39. Denis S. Krotov

    For the Hamming graph $H(n,q)$, where a $q$ is a constant prime power and $n$ grows, we construct perfect colorings without non-essential arguments such that $n$ depends exponentially on the off-diagonal part of the quotient matrix. In particular, we construct unbalanced Boolean ($q=2$) functions such that the number of essential arguments depends exponentia

  40. Yusuf Dalva, Yijun Li, Qing Liu, Nanxuan Zhao

    Large-scale diffusion models have achieved remarkable success in generating high-quality images from textual descriptions, gaining popularity across various applications. However, the generation of layered content, such as transparent images with foreground and background layers, remains an under-explored area. Layered content generation is crucial for creat

  41. Cheng Sun, Jaesung Choe, Charles Loop, Wei-Chiu Ma

    We propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are two key contributions coupled with the proposed system. The first is to adaptively and explicitly allocate sparse voxels to different levels of detail within scenes, faithfully repro

  42. Justin Lazarow, David Griffiths, Gefen Kohavi, Francisco Crespo

    We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respect to both data and modeling. First, we establish that existing datasets have significant limitations to scale, accuracy, and diversity of objects. As a result, we introduce the Cub

  43. Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy, Runfeng Li

    Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is interest in enabling view synthesis for dynamic scenes using only monocular input data -- an ill-posed and challenging problem. The fast pace of work in this area has produced multipl

  44. Yuto Matsubara, Ko Nishino

    We introduce a novel method for human shape and pose recovery that can fully leverage multiple static views. We target fixed-multiview people monitoring, including elderly care and safety monitoring, in which calibrated cameras can be installed at the corners of a room or an open space but whose configuration may vary depending on the environment. Our key id

  45. Enshen Zhou, Qi Su, Cheng Chi, Zhizheng Zhang

    Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones proactively. To this end, we propose Code-as-Monitor (CaM), a novel paradigm leveraging the vision-language model (VLM) for bo

  46. Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu

    Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning via in

  47. An-Chieh Cheng, Yandong Ji, Zhaojing Yang, Zaitian Gongye

    This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We prop

  48. Mohammed Suhail, Carlos Esteves, Leonid Sigal, Ameesh Makadia

    Latent variable generative models have emerged as powerful tools for generative tasks including image and video synthesis. These models are enabled by pretrained autoencoders that map high resolution data into a compressed lower dimensional latent space, where the generative models can subsequently be developed while requiring fewer computational resources.

  49. Mohammed Abouzaid, Shaoyun Bai

    We adapt algorithms for resolving the singularities of complex algebraic varieties to prove that the natural map of homology theories from complex bordism to the bordism theory of complex derived orbifolds splits. In equivariant stable homotopy theory, our techniques yield a splitting of homology theories for the map from bordism to the equivariant bordism t

  50. Liheng Yao, Robert L. Jack

    We analyze motility-induced phase separation and bubbly phase separation in a two-dimensional lattice model of self-propelled particles. We compare systems where the dense (liquid) phase has slab and droplet geometries. We find that interfacial fluctuations of the slab are well-described by capillary wave theory, despite the existence of bubbles in the dense

  51. Jun Zhang, Desen Meng, Zhengming Zhang, Zhenpeng Huang

    Despite the remarkable performance of multimodal large language models (MLLMs) across diverse tasks, the substantial training and inference costs impede their advancement. In this paper, we propose p-MoD, an efficient MLLM architecture that significantly reduces training and inference costs while maintaining model performance. The majority of computation in

  52. Philip Easo, Franco Severo, Vincent Tassion

    We prove two results concerning percolation on general graphs. - We establish the converse of the classical Peierls argument: if the critical parameter for (uniform) percolation satisfies $p_c<1$, then the number of minimal cutsets of size $n$ separating a given vertex from infinity is bounded above exponentially in $n$. This resolves a conjecture of Babson

  53. Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan

    Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing natural, audio-aligned expressions in generated talking videos remain significant challenges. To address these challenges, w

  54. Lu Qiu, Yi Chen, Yuying Ge, Yixiao Ge

    The advent of Multimodal Large Language Models, leveraging the power of Large Language Models, has recently demonstrated superior multimodal understanding and reasoning abilities, heralding a new era for artificial general intelligence. However, achieving AGI necessitates more than just comprehension and reasoning. A crucial capability required is effective

  55. Yizhuo Li, Yuying Ge, Yixiao Ge, Ying Shan

    Videos are inherently temporal sequences by their very nature. In this work, we explore the potential of modeling videos in a chronological and scalable manner with autoregressive (AR) language models, inspired by their success in natural language processing. We introduce DiCoDe, a novel approach that leverages Diffusion-Compressed Deep Tokens to generate vi

  56. Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li

    Recent developments in Large Language Models pre-trained on extensive corpora have shown significant success in various natural language processing tasks with minimal fine-tuning. This success offers new promise for robotics, which has long been constrained by the high cost of action-labeled data. We ask: given the abundant video data containing interaction-

  57. Daniel C. Hackett, Michael L. Wagman

    Recent work introduced a new framework for analyzing correlation functions with improved convergence and signal-to-noise properties, as well as rigorous quantification of excited-state effects, based on the Lanczos algorithm and spurious eigenvalue filtering with the Cullum-Willoughby test. Here, we extend this framework to the analysis of correlation-functi

  58. Tristan Hoellinger, Florent Leclercq

    The next generation of galaxy surveys has the potential to substantially deepen our understanding of the Universe. This potential hinges on our ability to rigorously address systematic uncertainties. Until now, diagnosing systematic effects prior to inferring cosmological parameters has been out of reach in field-based implicit likelihood cosmological infere

  59. Emma Finn, T. Anderson Keller, Emmanouil Theodosis, Demba E. Ba

    Despite nearly a decade of literature on style transfer, there is no undisputed definition of artistic style. State-of-the-art models produce impressive results but are difficult to interpret since, without a coherent definition of style, the problem of style transfer is inherently ill-posed. Early work framed style-transfer as an optimization problem but tr

  60. Kaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin

    Text-to-video generation models have shown significant progress in the recent years. However, they still struggle with generating complex dynamic scenes based on compositional text prompts, such as attribute binding for multiple objects, temporal dynamics associated with different objects, and interactions between objects. Our key motivation is that complex

  61. L. F. Lozano, I. Ledoino, B. J. Plohr, D. Marchesin

    Undercompressive shocks are a special type of discontinuities that satisfy the viscous profile criterion rather than the Lax inequalities. These shocks can appear as a solution to systems of two or more conservation laws. This paper presents the construction of the undercompressive shock surface for two types of diffusion matrices. The first type is the iden

  62. Jace Rusznak, Xian-Yu Wang, Malena Rice, Songhu Wang

    We present a pattern emerging from stellar obliquity measurements in single-star systems: planets with high planet-to-star mass ratios ($M_{\rm p}/M{_*}$$>$ $2\times10^{-3}$) -- such as super-Jupiters, brown dwarf companions, and M-dwarfs hosting Jupiter-like planets -- tend to be aligned, even around hot stars. This alignment represents a 3.7$\sigma$ deviat

  63. May Sela, Jake P. Solomon

    We define a normed matrix factorization category and a notion of bounding cochains for objects of this category. We classify bounding cochains up to gauge equivalence for spherical objects and use this classification to define numerical invariants. These invariants are expected to correspond under mirror symmetry to the open Gromov-Witten invariants with onl

  64. Dejene Zewdie, Roberto J. Assef, Trystan Lambert, Chiara Mazzucchelli

    Hot dust-obscured galaxies (Hot DOGs), are a family of hyper-luminous, heavily obscured quasars. A number of studies have shown that these objects reside in significantly overdense regions of the Universe based on the identification of companions at optical through far-IR wavelengths. Here we present further characterization of their environments by studying

  65. Jungbin Kim

    We prove the exact worst-case convergence rate of gradient descent for smooth strongly convex optimization, with respect to the performance criterion $\Vert \nabla f(x_N)\Vert^2/(f(x_0)-f_*)$. The proof differs from the previous one by Rotaru \emph{et al.} [RGP24], and is based on the performance estimation methodology [DT14].

  66. Bin Yan, Martin Sundermeyer, David Joseph Tan, Huchuan Lu

    In this paper, we address the challenge of performing open-vocabulary video instance segmentation (OV-VIS) in real-time. We analyze the computational bottlenecks of state-of-the-art foundation models that performs OV-VIS, and propose a new method, TROY-VIS, that significantly improves processing speed while maintaining high accuracy. We introduce three key t

  67. Shota Sasaki, Jane Wu, Ko Nishino

    This paper introduces a novel clothed human model that can be learned from multiview RGB videos, with a particular emphasis on recovering physically accurate body and cloth movements. Our method, Position Based Dynamic Gaussians (PBDyG), realizes ``movement-dependent'' cloth deformation via physical simulation, rather than merely relying on ``pose-dependent'

  68. Yuying Ge, Yizhuo Li, Yixiao Ge, Ying Shan

    In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The core challenge lies in developing a versatile video tokenizer that captures both the spatial characteristics and temporal

  69. Jian Han, Jinlai Liu, Yi Jiang, Bin Yan

    We present Infinity, a Bitwise Visual AutoRegressive Modeling capable of generating high-resolution, photorealistic images following language instruction. Infinity redefines visual autoregressive model under a bitwise token prediction framework with an infinite-vocabulary tokenizer & classifier and bitwise self-correction mechanism, remarkably improving the

  70. Xianzhe TZ Tang, Dillon Brout, Tanvi Karwal, Chihway Chang

    Recent results from Type Ia Supernovae (SNe), baryon acoustic oscillations (BAO), and the cosmic microwave background (CMB) indicate 1) potentially discrepant measurements of the matter density $\Omega_m$ and Hubble constant $ H_0 $ in $\Lambda$CDM model when analyzed individually, and 2) hints of dynamical dark energy in a $w_0w_a$CDM model when data are co

  71. Shaunak Halbe, Junjiao Tian, K J Joseph, James Seale Smith

    Vision-language models (VLMs) like CLIP have been cherished for their ability to perform zero-shot visual recognition on open-vocabulary concepts. This is achieved by selecting the object category whose textual representation bears the highest similarity with the query image. While successful in some domains, this method struggles with identifying fine-grain

  72. Ira Z. Rothstein, Michael Saavedra

    We derive a systematic Lagrangian approach for quantum gravity in the super-Planckian limit where $s\gg M_{pl}^2\gg t$. The action can be used to calculate to arbitrary accuracy in the quantum and classical expansion parameters $\alpha_Q= \frac{t}{M_{pl}^2}$ and $\alpha_C= \frac{st}{M_{pl}^4}$, respectively, for the scattering of massless particles. The pert

  73. Jungbin Kim

    We prove the exact worst-case convergence rate of gradient descent for smooth strongly convex optimization on $\mathbb{R}^d$. Concretely, assuming that the objective function $f$ is $\mu$-strongly convex and $L$-smooth, we identify the smallest possible value of $\tau$ for which the inequality $f(x_{N})-f_{*}\leq\tau\|x_{0}-x_{*}\|^{2}$ always holds. The res

  74. Keru Chen, Honghao Wei, Zhigang Deng, Sen Lin

    The high costs and risks involved in extensive environment interactions hinder the practical application of current online safe reinforcement learning (RL) methods. While offline safe RL addresses this by learning policies from static datasets, the performance therein is usually limited due to reliance on data quality and challenges with out-of-distribution

  75. Yen-Ju Lu, Jing Liu, Thomas Thebaud, Laureano Moro-Velazquez

    We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from earlier layers, making the SSL model aware of the current l

  76. Jiuhai Chen, Jianwei Yang, Haiping Wu, Dianqi Li

    We present Florence-VL, a new family of multimodal large language models (MLLMs) with enriched visual representations produced by Florence-2, a generative vision foundation model. Unlike the widely used CLIP-style vision transformer trained by contrastive learning, Florence-2 can capture different levels and aspects of visual features, which are more versati

  77. Madeleine D. Breshears, Justin Pothoof, Rajiv Giridharagopal, David S. Ginger

    We spatially resolve photocarrier dynamics in halide perovskites using time-resolved electrostatic force microscopy (trEFM) to map surface potential equilibration during photoexcitation. Following treatment with different surface passivation agents, we show that trEFM probes dynamics directly related to surface recombination velocity and carrier lifetimes co

  78. Maryam Hosseini, Reem Yassawi

    For every Toeplitz sequence $x$ with period structure $(q_i)_{i\geq 1}$, one can identify a period structure ${\bf p}=(p_i)_{i\geq 0}$ which leads to a Bratteli-Vershik realization of the associated Toeplitz shift; we refer to this period structure as {\it constructive}. Let $(X,\sigma,x)$ and $(Y,\sigma,y)$ be Toeplitz shifts where $x\in X$ and $y\in Y$ are

  79. Tomas Ortega, Chun-Yin Huang, Xiaoxiao Li, Hamid Jafarkhani

    Distributed learning algorithms, such as the ones employed in Federated Learning (FL), require communication compression to reduce the cost of client uploads. The compression methods used in practice are often biased, making error feedback necessary both to achieve convergence under aggressive compression and to provide theoretical convergence guarantees. Ho

  80. M. C. Smith, A. D. Leu, K. Miyanishi, M. F. Gely

    We report the achievement of single-qubit gates with sub-part-per-million error rates, in a trapped-ion $^{43}$Ca$^{+}$ hyperfine clock qubit. We explore the speed/fidelity trade-off for gate times $4.4\leq t_{g}\leq35~\mu$s, and benchmark a minimum error per Clifford gate of $1.5(4) \times 10^{-7}$. Calibration errors are suppressed to $< 10^{-8}$, leaving

  81. Marcin Briański, James Davies, Jane Tan

    Negami's famous planar cover conjecture is equivalent to the statement that a connected graph can be embedded in the projective plane if and only if it has a projective planar cover. In 1999, Hlin\v{e}n\'y proposed extending this conjecture to higher genus non-orientable surfaces. In this paper, we put forward a natural extension that encompasses orientable

  82. Damien Gagnier, Ondrej Pejcha

    Three-dimensional hydrodynamical simulations of common envelope evolution are often terminated soon after the initial dynamical plunge of the companion transitions into a long-lasting post-dynamical inspiral with slowly varying semi-major axis, $a_\text{b}$. This premature termination is often due to insufficient numerical resolution and challenges associate

  83. Spencer K. Clark, Oliver Watt-Meyer, Anna Kwa, Jeremy McGibbon

    While autoregressive machine-learning-based emulators have been trained to produce stable and accurate rollouts in the climate of the present-day and recent past, none so far have been trained to emulate the sensitivity of climate to substantial changes in CO$_2$ or other greenhouse gases. As an initial step we couple the Ai2 Climate Emulator version 2 to a

  84. Aryasomayajula Ram Bharadwaj

    Chain-of-Thought (CoT) prompting has significantly enhanced the reasoning abilities of large language models. However, recent studies have shown that models can still perform complex reasoning tasks even when the CoT is replaced with filler(hidden) characters (e.g., "..."), leaving open questions about how models internally process and represent reasoning st

  85. Pranab Sahoo, Ashutosh Tripathi, Sriparna Saha, Samrat Mondal

    Federated Learning (FL) marks a transformative approach to distributed model training by combining locally optimized models from various clients into a unified global model. While FL preserves data privacy by eliminating centralized storage, it encounters significant challenges such as performance degradation, slower convergence, and reduced robustness of th

  86. Xiayin Lou, Peng Luo, Liqiu Meng

    Spatial prediction is a fundamental task in geography. In recent years, with advances in geospatial artificial intelligence (GeoAI), numerous models have been developed to improve the accuracy of geographic variable predictions. Beyond achieving higher accuracy, it is equally important to obtain predictions with uncertainty measures to enhance model credibil

  87. Xuying Li, Zhuo Li, Yuji Kosuga, Yasuhiro Yoshida

    AI agents, powered by large language models (LLMs), have transformed human-computer interactions by enabling seamless, natural, and context-aware communication. While these advancements offer immense utility, they also inherit and amplify inherent safety risks such as bias, fairness, hallucinations, privacy breaches, and a lack of transparency. This paper in

  88. Zihan Cheng, Eric Huang, Vedika Khemani, Michael J. Gullans

    Unitary $k$-designs are distributions of unitary gates that match the Haar distribution up to its $k$-th statistical moment. They are a crucial resource for randomized quantum protocols. However, their implementation on encoded logical qubits is nontrivial due to the need for magic gates, which can require a large resource overhead. In this work, we propose

  89. Anshul Thakur, Yichen Huang, Soheila Molaei, Yujiang Wang

    Shared training approaches, such as multi-task learning (MTL) and gradient-based meta-learning, are widely used in various machine learning applications, but they often suffer from negative transfer, leading to performance degradation in specific tasks. While several optimisation techniques have been developed to mitigate this issue for pre-selected task coh

  90. Brandon D. Stone, Lala Rukh, Gabriel M. Colación, Tara E. Drake

    Kerr-microresonator frequency combs in integrated photonics waveguides are promising technologies for next-generation positioning, navigation, and timing applications, with advantages that include platforms that are mass-producible and CMOS-compatible and spectra that are phase-coherent and octave-spanning. Fundamental thermal noise in the resonator material

  91. D. E. Lyubashevsky, A. A. Pisklyukov, S. V. Klyuchnikov, P. V. Kostryukov

    This study proposes a theoretical model for studying the spin characteristics and angular correlations of fission fragments of heavy nuclei. The mechanisms of spin formation, including the influence of transverse vibrations, are considered and the relationship between the anisotropy of the angular distribution and the correlation coefficient is revealed. The

  92. Erik Burman, Mats G. Larson, Karl Larsson, Carl Lundholm

    We consider an inverse problem involving the reconstruction of the solution to a nonlinear partial differential equation (PDE) with unknown boundary conditions. Instead of direct boundary data, we are provided with a large dataset of boundary observations for typical solutions (collective data) and a bulk measurement of a specific realization. To leverage th

  93. Jiayu Mao, Tongxin Yin, Aylin Yener, Mingyan Liu

    Federated Learning (FL) is a distributed machine learning framework that inherently allows edge devices to maintain their local training data, thus providing some level of privacy. However, FL's model updates still pose a risk of privacy leakage, which must be mitigated. Over-the-air FL (OTA-FL) is an adapted FL design for wireless edge networks that leverag

  94. Noémie C. Combe

    Dually flat statistical manifolds provide a rich toolbox for investigations around the learning process. We prove that such manifolds are Monge-Amp\`ere manifolds. Examples of such manifolds include the space of exponential probability distributions on finite sets and the Boltzmann manifolds. Our investigations of Boltzmann manifolds lead us to prove that Mo

  95. Luca Fanelli, Xiaoyan Su, Ying Wang, Junyong Zhang

    The main mathematical manifestation of the Stark effect in quantum mechanics is the shift and the formation of clusters of eigenvalues when a spherical Hamiltonian is perturbed by lower order terms. Understanding this mechanism turned out to be fundamental in the description of the large-time asymptotics of the associated Schr\"odinger groups and can be resp

  96. Paula S. Ferreira, Ribamar R. R. Reis

    We conducted a review of the fundamental aspects of describing and detecting the Baryon Acoustic Oscillation (BAO) feature in galaxy surveys, emphasizing the optimal tools for constraining this probe based on the type of observation. Additionally, we included new results with two spectroscopic datasets to determine the best-fit model for the power spectrum,

  97. Tom Overman, Diego Klabjan

    Automated feature engineering (AutoFE) is used to automatically create new features from original features to improve predictive performance without needing significant human intervention and domain expertise. Many algorithms exist for AutoFE, but very few approaches exist for the federated learning (FL) setting where data is gathered across many clients and

  98. Akshita Bhagia, Jiacheng Liu, Alexander Wettig, David Heineman

    We develop task scaling laws and model ladders to predict the individual task performance of pretrained language models (LMs) in the overtrained setting. Standard power laws for language modeling loss cannot accurately model task performance. Therefore, we leverage a two-step prediction approach: (1) use model and data size to predict an intermediate loss, t

  99. Yunzhe Zheng, Dong E. Liu

    Magic State Distillation (MSD) has been a research focus for fault-tolerant quantum computing due to the need for non-Clifford resource in gaining quantum advantage. Although many of the MSD protocols so far are based on stabilizer codes with transversal $T$ gates, there exists quite several protocols that don't fall into this class. Here we propose a method

  100. Richard J. Smith

    In a paper from 1987, Gruenhage defined a class of topological spaces that now bear his name, and used it to solve a problem of Talagrand on the existence of dense $G_\delta$ metrizable subsets of Gul'ko compact spaces. Gruenhage's paper became highly influential among researchers in renorming theory, a branch of Banach space theory. In this paper we survey