March 2025 arXiv papers — page 136
Showing 13,501–13,600 of 23,633 papers
Michael J. Baker, Joaquim Iguaz Juan, Aidan Symons, Andrea Thamm
Observation of an exploding black hole would provide the first direct evidence of primordial black holes, the first direct evidence of Hawking radiation, and definitive information on the particles present in nature. However, indirect constraints suggest that direct observation of an exploding Schwarzschild black hole is implausible. We introduce a dark-QED
P. H. Kim, M. Hirschel, J. Suranyi, J. P. Davis
We present a novel cooling method that uses the phase separation and evaporative cooling of 3He to reach and continuously maintain sub-kelvin temperatures. While less complex than a dilution refrigerator, the system performs similarly to a continuous 3He cryostat but with a simpler design, a more efficient cooldown process, and a significantly smaller 3He re
Stefano Baiguera, Vijay Balasubramanian, Pawel Caputa, Shira Chapman
Quantum complexity quantifies the difficulty of preparing a state or implementing a unitary transformation with limited resources. Applications range from quantum computation to condensed matter physics and quantum gravity. We seek to bridge the approaches of these fields, which define and study complexity using different frameworks and tools. We describe se
Leonardo Chataignier, Claus Kiefer, Mritunjay Tyagi
We discuss how the classical notions of time and causal structure may emerge together with quantum-mechanical probabilities from a universal quantum state. For this, the process of decoherence between semiclassical branches is important. Our discussion is based on quantum geometrodynamics, a canonical approach to quantum gravity. In this framework, a particu
J. Scholtz, E. Parlanti, S. Carniani, M. Kohandel
We re-analysed ALMA observations of the [OIII]$\lambda$88$\mu$m emission line in JADES-GS-z14.0, so far the most distant spectroscopically confirmed galaxy at z=14.18. Our analysis shows a tentative detection of a velocity gradient of [OIII]$\lambda$88$\mu$m using three independent tests: 1) construction of moment maps; 2) extraction of integrated spectra fr
Basil M. Smitham, Jeronimo G. C. Martinez, Christie S. Chiu, Andrew A. Houck
In superconducting quantum devices, Purcell filters protect qubit information from decaying into external lines by reducing external coupling at qubit frequencies while maintaining it at readout frequencies. Here, we introduce and demonstrate a novel Purcell filter design that places the readout resonator frequencies in a "linewidth plateau" below the filter
Kryštof Kolář, Dacen Waters, Joshua Folk, Matthew Yankowitz
A central feature of many van der Waals (vdW) materials is the ability to precisely control their charge doping, $n$, and electric displacement field, $D$, using top and bottom gates. For devices composed of only a few layers, it is commonly assumed that $D$ causes the layer-by-layer potential to drop linearly across the structure. Here, we show that this as
SIDM Concerto: Compilation and Data Release of Self-interacting Dark Matter Zoom-in Simulations
astro-ph.COEthan O. Nadler, Demao Kong, Daneng Yang, Hai-Bo Yu
We present SIDM Concerto: $14$ cosmological zoom-in simulations in cold dark matter (CDM) and self-interacting dark matter (SIDM) models based on the Symphony and Milky Way-est suites. SIDM Concerto includes one Large Magellanic Cloud- (LMC-) mass system (host mass $\sim 10^{11}~M_{\mathrm{\odot}}$), two Milky Way (MW) analogs ($\sim 10^{12}~M_{\mathrm{\odot
Naixin Liang, Siang Peng Oh
In classical diffusion, particle step-sizes have a Gaussian distribution. However, in superdiffusion, they have power-law tails, with transport dominated by rare, long L\'evy flights. Similarly, if the time interval between scattering events has power-law tails, subdiffusion occurs. Both forms of anomalous diffusion are seen in cosmic ray (CR) particle track
Valerio De Luca, Loris Del Grosso, Francesco Iacovelli, Andrea Maselli
Binary black hole systems are typically assumed to evolve in vacuum. However, the environment surrounding the binary components can influence their properties, such as their tidal deformability, affecting the gravitational waveform produced by the binary and its interpretation in gravitational wave data analysis. In this work we focus on next-generation expe
GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
cs.CVRongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang
Current image generation and editing methods primarily process textual prompts as direct inputs without reasoning about visual composition and explicit operations. We present Generation Chain-of-Thought (GoT), a novel paradigm that enables generation and editing through an explicit language reasoning process before outputting images. This approach transforms
Xiaoming Zhao, Alexander G. Schwing
Classifier-free guidance has become a staple for conditional generation with denoising diffusion models. However, a comprehensive understanding of classifier-free guidance is still missing. In this work, we carry out an empirical study to provide a fresh perspective on classifier-free guidance. Concretely, instead of solely focusing on classifier-free guidan
Rohit Gandikota, David Bau
Distilled diffusion models generate images in far fewer timesteps but suffer from reduced sample diversity when generating multiple outputs from the same prompt. To understand this phenomenon, we first investigate whether distillation damages concept representations by examining if the required diversity is properly learned. Surprisingly, distilled models re
The Curse of Conditions: Analyzing and Improving Optimal Transport for Conditional Flow-Based Generation
cs.LGHo Kei Cheng, Alexander Schwing
Minibatch optimal transport coupling straightens paths in unconditional flow matching. This leads to computationally less demanding inference as fewer integration steps and less complex numerical solvers can be employed when numerically solving an ordinary differential equation at test time. However, in the conditional setting, minibatch optimal transport fa
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
cs.CVZhaoyi Li, Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui
Despite promising performance on open-source large vision-language models (LVLMs), transfer-based targeted attacks often fail against closed-source commercial LVLMs. Analyzing failed adversarial perturbations reveals that the learned perturbations typically originate from a uniform distribution and lack clear semantic details, resulting in unintended respons
Yanming Zhang, Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang
This paper introduces V$^2$Edit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content preservation with editing task fulfillment, our approach employs a progressive strategy that decomposes complex editing tasks into a sequence of simpler subtasks. Each subtask is c
Eliahu Horwitz, Nitzan Kurer, Jonathan Kahana, Liel Amar
Public model repositories now contain millions of models, yet most models remain undocumented and effectively lost. In this position paper, we advocate for charting the world's model population in a unified structure we call the Model Atlas: a graph that captures models, their attributes, and the weight transformations that connect them. The Model Atlas enab
Subhajit Maity, Killian Hitsman, Xin Li, Aritra Dutta
Kolmogorov-Arnold networks (KANs) are a remarkable innovation that consists of learnable activation functions, with the potential to capture more complex relationships from data. Presently, KANs are deployed by replacing multilayer perceptrons (MLPs) in deep networks, including advanced architectures such as vision Transformers (ViTs). This work asks whether
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
cs.CVJiaming Liu, Hao Chen, Pengju An, Zhuoyang Liu
A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic environments. Recent autoregressive vision-language-action (VLA) methods inherit common-sense reasoning capabilities from vision-language models (VLMs) for next action-token prediction.
Hang Yin, Xiuwei Xu, Lingqing Zhao, Ziwei Wang
In this paper, we propose a general framework for universal zero-shot goal-oriented navigation. Existing zero-shot methods build inference framework upon large language models (LLM) for specific tasks, which differs a lot in overall pipeline and fails to generalize across different types of goal. Towards the aim of universal zero-shot navigation, we propose
Tianjiao Yu, Vedant Shah, Muntasir Wahed, Kiet A. Nguyen
Expressing confidence is challenging for embodied agents navigating dynamic multimodal environments, where uncertainty arises from both perception and decision-making processes. We present the first work investigating embodied confidence elicitation in open-ended multimodal environments. We introduce Elicitation Policies, which structure confidence assessmen
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
cs.CVZiyu Guo, Ray Zhang, Hao Chen, Jialin Gao
The rapid advancement of Large Multi-modal Models (LMMs) has enabled their application in scientific problem-solving, yet their fine-grained capabilities remain under-explored. In this paper, we introduce SciVerse, a multi-modal scientific evaluation benchmark to thoroughly assess LMMs across 5,735 test instances in five distinct versions. We aim to investig
Mert Albaba, Chenhao Li, Markos Diomataris, Omid Taheri
Acquiring physically plausible motor skills across diverse and unconventional morphologies-including humanoid robots, quadrupeds, and animals-is essential for advancing character simulation and robotics. Traditional methods, such as reinforcement learning (RL) are task- and body-specific, require extensive reward function engineering, and do not generalize w
Lingteng Qiu, Xiaodong Gu, Peihao Li, Qi Zuo
Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits their generalization ability. Conversely, optimization-base
Boqian Li, Haiwen Feng, Zeyu Cai, Michael J. Black
Fitting a body to a 3D clothed human point cloud is a common yet challenging task. Traditional optimization-based approaches use multi-stage pipelines that are sensitive to pose initialization, while recent learning-based methods often struggle with generalization across diverse poses and garment types. We propose Equivariant Tightness Fitting for Clothed Hu
Jordan Huang, Thomas J. DiNapoli, Gavin Rockwood, Ming Yuan
Circuit quantum electrodynamics (cQED) with superconducting cavities coupled to nonlinear circuits like transmons offers a promising platform for hardware-efficient quantum information processing. We address critical challenges in realizing this architecture by weakening the dispersive coupling while also demonstrating fast, high-fidelity multimode control b
Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun
Normalization layers are ubiquitous in modern neural networks and have long been considered essential. This work demonstrates that Transformers without normalization can achieve the same or better performance using a remarkably simple technique. We introduce Dynamic Tanh (DyT), an element-wise operation $DyT($x$) = \tanh(\alpha $x$)$, as a drop-in replacemen
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
cs.CVAyesha Ishaq, Jean Lahoud, Ketan More, Omkar Thawakar
While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive processing before decisions can be made. In this domain, a
Kshitij Ambilduke, Ben Peters, Sonal Sannigrahi, Anil Keshwani
We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating text input in both language directions. Spire integrates the speech modality into an existing multilingual LM via speech discretization and continued pre-training using only 42.5K hou
Andy Zhou, Ron Arel
We introduce Tempest, a multi-turn adversarial framework that models the gradual erosion of Large Language Model (LLM) safety through a tree search perspective. Unlike single-turn jailbreaks that rely on one meticulously engineered prompt, Tempest expands the conversation at each turn in a breadth-first fashion, branching out multiple adversarial prompts tha
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
cs.CVChen Chen, Rui Qian, Wenze Hu, Tsu-Jui Fu
In this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protocols. We evaluate a range of DiT-based architectures--including PixArt-style and MMDiT variants--and compare them with a standard DiT variant which directly processes concatenated te
Andy Zhou
Adapting large language models to multiple tasks can cause cross-skill interference, where improvements for one skill degrade another. While methods such as LoRA impose orthogonality constraints at the weight level, they do not fully address interference in hidden-state representations. We propose Compositional Subspace Representation Fine-tuning (CS-ReFT),
Ayush Jain, Alexander Swerdlow, Yuzhou Wang, Sergio Arnaud
Progress in 3D vision-language learning has been hindered by the scarcity of large-scale 3D datasets. We introduce UniVLG, a unified architecture for 2D and 3D vision-language understanding that bridges the gap between existing 2D-centric models and the rich 3D sensory data available in embodied systems. Our approach initializes most model weights from pre-t
Jinyang Li, En Yu, Sijia Chen, Wenbing Tao
Open-vocabulary multiple object tracking aims to generalize trackers to unseen categories during training, enabling their application across a variety of real-world scenarios. However, the existing open-vocabulary tracker is constrained by its framework structure, isolated frame-level perception, and insufficient modal interactions, which hinder its performa
Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang
Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reason visual content, resulting in suboptimal performance on com
Bolin Chen, Baoquan Zhao, Haoran Xie, Yi Cai
Style transfer involves transferring the style from a reference image to the content of a target image. Recent advancements in LoRA-based (Low-Rank Adaptation) methods have shown promise in effectively capturing the style of a single image. However, these approaches still face significant challenges such as content inconsistency, style misalignment, and cont
Advait Gupta, NandaKiran Velaga, Dang Nguyen, Tianyi Zhou
Text-to-image models like stable diffusion and DALLE-3 still struggle with multi-turn image editing. We decompose such a task as an agentic workflow (path) of tool use that addresses a sequence of subtasks by AI tools of varying costs. Conventional search algorithms require expensive exploration to find tool paths. While large language models (LLMs) possess
Shane Farnsworth
We solve an open problem in spectral geometry: the construction of finite-dimensional, discrete geometries coordinatized by non-simple, exceptional Jordan algebras. The approach taken is readily generalisable to broad classes of nonassociative geometries, opening the door to the spectral geometric desciption of gauge theories with exceptional symmetries. We
Preserving the minimum principle on the entropy for the compressible Euler Equations with general equations of state
math.NABennett Clayton, Eric J. Tovar
This paper is concerned with constructing an invariant-domain preserving approximation technique for the compressible Euler equations with general equations of state that preserves the minimum principle on the physical entropy. We derive a sufficient wave speed estimate for the Riemann problem under some mild thermodynamic assumptions on the equation of stat
Karen Habermann, Stephen C. Preston, Stefan Sommer
We provide a full characterization of geodesic completeness for spaces of configurations of landmarks with smooth Riemannian metrics that satisfy a rotational and translation invariance and which are induced from metrics on subgroups of the diffeomorphism group for the shape domain. These spaces are widely used for applications in shape analysis, for example
J. A. Acevedo Barroso, B. Clément, F. Courbin, R. Gavazzi
Recent wide-field galaxy surveys have led to an explosion in the number of galaxy-scale strong gravitational lens candidates. However, the vast majority of them feature massive luminous red galaxies as the main deflectors, with late-type galaxies being vastly under-represented. This work presents a dedicated search for lensing by edge-on late-type galaxies i
Knot reconstruction of the scalar primordial power spectrum with Planck, ACT, and SPT CMB data
astro-ph.COAntonio Raffaelli, Mario Ballardini
We investigate a non-parametric Bayesian method for reconstructing the primordial power spectrum (PPS) of scalar perturbations using temperature and polarisation data from the {\em Planck}, ACT, and SPT CMB experiments. This reconstruction method is based on linear splines for the PPS between nodes in $k$-space whose amplitudes and positions are allowed to v
Hierarchical Bayesian inference for uncertainty quantification of thermal grease rheology
cond-mat.softPranay P. Nagrani, Akshay J. Thomas, Amy M. Marconnet, Ivan C. Christov
Rheologically complex soft solids such as thermal greases consist of filler particles within a polymer matrix. These materials find applications in improving the conformity of solid-solid contacts and enhancing heat transfer. Complex soft solids exhibit a transient non-Newtonian rheological response, including thixotropy and viscoelasticity. Previously, stre
Utilizing discrete variable representations for decoherence-accurate numerical simulation of superconducting circuits
quant-phBrittany Richman, C. J. Lobb, Jacob M. Taylor
Given the prevalence of superconducting platforms for uses in quantum computing and quantum sensing, the simulation of quantum superconducting circuits has become increasingly important for identifying system characteristics and modeling their relevant dynamics. Various numerical tools and software packages have been developed with this purpose in mind, typi
Mehrab Momennia, Olivier Sarbach
We present a nontrivial extension of the problem of spherical accretion of a collisionless kinetic gas into the standard Schwarzschild black hole. This extension consists of replacing the Schwarzschild black hole by generic static and spherically symmetric black hole spacetimes with the aim of studying the effects of either modified gravitational theories be
Severin Heidrich, Till Beemelmanns, Alexey Nekrasov, Bastian Leibe
Autonomous driving has the potential to significantly enhance productivity and provide numerous societal benefits. Ensuring robustness in these safety-critical systems is essential, particularly when vehicles must navigate adverse weather conditions and sensor corruptions that may not have been encountered during training. Current methods often overlook unce
Yingshuang Zou, Yikang Ding, Chuanrui Zhang, Jiazhe Guo
Recent breakthroughs in radiance fields have significantly advanced 3D scene reconstruction and novel view synthesis (NVS) in autonomous driving. Nevertheless, critical limitations persist: reconstruction-based methods exhibit substantial performance deterioration under significant viewpoint deviations from training trajectories, while generation-based techn
Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
cs.ROQi Lv, Hao Li, Xiang Deng, Rui Shao
Despite the significant success of imitation learning in robotic manipulation, its application to bimanual tasks remains highly challenging. Existing approaches mainly learn a policy to predict a distant next-best end-effector pose (NBP) and then compute the corresponding joint rotation angles for motion using inverse kinematics. However, they suffer from tw
Keyframe-oriented Vision Token Pruning: Enhancing Efficiency of Large Vision Language Models on Long-Form Video Processing
cs.LGYudong Liu, Jingwei Sun, Yueqian Lin, Jingyang Zhang
Vision language models (VLMs) demonstrate strong capabilities in jointly processing visual and textual data. However, they often incur substantial computational overhead due to redundant visual information, particularly in long-form video scenarios. Existing approaches predominantly focus on either vision token pruning, which may overlook spatio-temporal dep
Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition
cs.CVJun Yu, Lingsi Zhu, Yanjun Chi, Yunxiang Zhang
Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal temporal signals. To address the limitations of existing methods--insufficient exploitation of cross-modal synergies, sensiti
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
cs.CVJinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer
Object Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such as hidden states, encode the "overall truthfulness" of generated responses. However, it remains under-explored how internal states in LVLMs
Berat Yenilen, Arnau Sala, Hendrik Bluhm, Markus Müller
Qubit shuttling promises to advance some quantum computing platforms to the qubit register sizes needed for effective quantum error correction (QEC), but also introduces additional errors whose impact must be evaluated. The established method to investigate the performance of QEC codes in a realistic scenario is to employ a standard noise model known as circ
Ludwig Scheuchenpflug, Sebastian Esser, Robert Gruhl, Max Hirschberger
Emergent electromagnetic induction (EEMI) induced through current-driven spin dynamics was recently predicted and subsequently observed in helical spin magnets, opening a new direction in spintronics and paving the way towards further miniaturization of electronic circuit elements. In contrast to conventional inductors consisting of coil-like structures whos
Cai Dieball, Aljaž Godec
Recently, a thermodynamic bound on correlation times was formulated in [A. Dechant, J. Garnier-Brun, S.-i. Sasa, Phys. Rev. Lett. 131, 167101 (2023)], showing how the decay of correlations in Langevin dynamics is bounded by short-time fluctuations and dissipation. Whereas these original results only address very long observation times in steady-state dynamic
Chandramouli Chowdhury, Arthur Lipstein, Joe Marshall, Jiajie Mei
The basic observables in cosmology are known as in-in correlators. Recent calculations have revealed that in-in correlators in four dimensional de Sitter space exhibit hidden simplicity stemming from a close relation to scattering amplitudes in flat space. In this paper we explain how to make this property manifest by dressing flat space Feynman diagrams wit
Yang Zheng, Menglei Chai, Delio Vicini, Yuxiao Zhou
We present GroomLight, a novel method for relightable hair appearance modeling from multi-view images. Existing hair capture methods struggle to balance photorealistic rendering with relighting capabilities. Analytical material models, while physically grounded, often fail to fully capture appearance details. Conversely, neural rendering approaches excel at
Rui Hu, Lianghui Zhu, Yuxuan Zhang, Tianheng Cheng
Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existing datasets, including limited object categories, ins
Giuseppe De Laurentis, Harald Ita, Ben Page, Vasily Sotnikov
We present compact two-loop QCD corrections in the leading-color approximation for the production of an electroweak vector boson, $V = \{W^{\pm}, Z,\gamma^\star\}$, in association with two light jets ($Vjj$) at hadron colliders. Leptonic decays of the electroweak boson are included at the amplitude level. Working in the analytic-reconstruction approach, we d
Antonia van Betteray, Matthias Rottmann, Karsten Kahl
The structural analogies of ResNets and Multigrid (MG) methods such as common building blocks like convolutions and poolings where already pointed out by He et al.\ in 2016. Multigrid methods are used in the context of scientific computing for solving large sparse linear systems arising from partial differential equations. MG methods particularly rely on two
Sebastian Grieninger, Sergio Morales-Tejera, Pau G. Romeu
We extend previous holographic studies of the Chiral Magnetic Effect (CME) by incorporating a time-dependent magnetic field. Various magnetic field profiles proposed in the literature are implemented, and their impact on the CME signal is analyzed in both static and expanding backgrounds. Interestingly, the integrated chiral magnetic current can exhibit a no
Yubin Hong, Chaofan Li, Jingyi Zhang, Yingxia Shao
Retrieval-Augmented Generation (RAG) enables large language models to provide more precise and pertinent responses by incorporating external knowledge. In the Query-Focused Summarization (QFS) task, GraphRAG-based approaches have notably enhanced the comprehensiveness and diversity of generated responses. However, existing GraphRAG-based approaches predomina
Hao He, Ceyuan Yang, Shanchuan Lin, Yinghao Xu
This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos with large camera movement. We take an approach that progressive
Analysis and sample-size determination for $2^K$ audit experiments with binary response and application to identification of effect of racial discrimination on access to justice
stat.MENicole Pashley, Brian Libgober, Tirthankar Dasgupta
Social scientists have increasingly turned to audit experiments to investigate discrimination in the market for jobs, loans, housing and other opportunities. In a typical audit experiment, researchers assign ``signals'' (the treatment) to subjects at random and compare success rates across treatment conditions. In the recent past there has been increased int
Dmitriy Rumynin, James Taylor
We consider Brauer's 14th Problem in the context of "Real" structures on finite groups and their antilinear representations. The problem is to count the number of characters of each different type using "group theory". While Brauer's original problem deals only with three types (real, complex and quaternionic), here we consider the ten types coming from Dyso
Yuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma
Recent advances in video generation can produce realistic, minute-long single-shot videos with scalable diffusion transformers. However, real-world narrative videos require multi-shot scenes with visual and dynamic consistency across shots. In this work, we introduce Long Context Tuning (LCT), a training paradigm that expands the context window of pre-traine
Ilia V. Zalivako, Andrey Yu. Chernyavskiy, Anastasiia S. Nikolaeva, Alexander S. Borisenko
Factoring integers is considered as a computationally-hard problem for classical methods, whereas there exists polynomial-time Shor's quantum algorithm for solving this task. However, requirements for running the Shor's algorithm for realistic tasks, which are beyond the capabilities of existing and upcoming generations of quantum computing devices, motivate
Omar Costilla-Reyes, Morgan Talbot
Body Dysmorphic Disorder (BDD) is a highly prevalent and frequently underdiagnosed condition characterized by persistent, intrusive preoccupations with perceived defects in physical appearance. In this extended analysis, we employ multiple machine learning approaches to predict treatment outcomes -- specifically treatment response and remission -- with an em
Justin Sahs, Ryan Pyle, Fabio Anselmi, Ankit Patel
Despite classical statistical theory predicting severe overfitting, modern massively overparameterized neural networks still generalize well. This unexpected property is attributed to the network's so-called implicit bias, which describes its propensity to converge to solutions that generalize effectively, among the many possible that correctly label the tra
Chaoqun Wang, Jie Yang, Xiaobin Hong, Ruimao Zhang
Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quality annotated data, which is costly and labor-intensive. To address this issue, we propose unlocking the value of abundant yet unlabeled data to improve the language-driving model i
Once bitten, twice shy: A modeling framework for incorporating heterogeneous mosquito biting into transmission models
q-bio.PEKyle J. -M. Dahlin, Michael A. Robert, Lauren M. Childs
The risk of mosquito-borne disease outbreaks is tightly linked to the frequency at which mosquitoes feed on blood, also known as the biting rate. However, standard models of mosquito-borne disease transmission inherently assume that mosquitoes bite only once per reproductive cycle -- an assumption commonly violated in nature. Drivers of multiple biting also
Holographic study of shear viscosity and butterfly velocity for magnetic field-driven quantum criticality
hep-thJun-Kun Zhao, Li Li
We investigate the shear viscosity and butterfly velocity of a magnetic field-induced quantum phase transition in five dimensional Einstein-Maxwell-Chern-Simons theory, which is holographically dual to a class of strongly coupled quantum field theories with chiral anomalies. Our analysis reveals that the ratio of longitudinal shear viscosity to entropy densi
Piotr Budzyński
Assorted weighted shifts over finite rooted directed trees are studied. Their complex symmetry is characterized.
Yiming Jia, Jiachen Li, Xiang Yue, Bo Li
Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-quality and diverse training data. In this work, we aim to address the scarcity of reasoning-focused multimodal datasets. We propose VisualWebInstruct, a novel approach that leverag
Simulating charging characteristics of lithium iron phosphate by electro-ionic optimization on a quantum annealer
cond-mat.mtrl-sciTobias Binninger, Yin-Ying Ting, Konstantin Köster, Nils Bruch
The rapid evolution of quantum computing hardware opens up new avenues in the simulation of energy materials. Today's quantum annealers are able to tackle complex combinatorial optimization problems. A formidable challenge of this type is posed by materials with site-occupational disorder for which atomic arrangements with a low, or lowest, energy must be fo
Ishaq Aden-Ali
We prove an upper bound on the expected $\ell_p$ injective norm of sums of subgaussian random tensors. Our proof is simple and does not rely on any explicit geometric or chaining arguments. Instead, it follows from a simple application of the PAC-Bayesian lemma, a tool that has proven effective at controlling the suprema of certain ``smooth'' empirical proce
Chaoqun Wang, Xiaobin Hong, Wenzhong Li, Ruimao Zhang
LiDAR-based 3D object detection presents significant challenges due to the inherent sparsity of LiDAR points. A common solution involves long-term temporal LiDAR data to densify the inputs. However, efficiently leveraging spatial-temporal information remains an open problem. In this paper, we propose a novel Semantic-Supervised Spatial-Temporal Fusion (ST-Fu
Benchmarking low-power flopping-mode spin qubit fidelities in Si/SiGe devices with alloy disorder
cond-mat.mes-hallSteve M. Young, Mitchell Brickson, Jason R. Petta, N. Tobias Jacobson
In the "flopping-mode" regime of electron spin resonance, a single electron confined in a double quantum dot is electrically driven in the presence of a magnetic field gradient. The increased dipole moment of the charge in the flopping mode significantly reduces the amount of power required to drive spin rotations. However, the susceptibility of flopping-mod
Félix Cabello Sánchez, Willian Corrêa, Ben-Hur Eidt
We generalize the complex interpolation formula $(X, X')_{\frac{1}{2}} = L^2$ from the context of Banach function spaces to that of directional Banach spaces. We also obtain a formula for the interpolation of matrix weighted spaces of vector valued functions via interpolation functors, and apply our formula to the particular case of interpolation of matr
Nina Vesseron, Louis Béthune, Marco Cuturi
The canonical approach in generative modeling is to split model fitting into two blocks: define first how to sample noise (e.g. Gaussian) and choose next what to do with it (e.g. using a single map or flows). We explore in this work an alternative route that ties sampling and mapping. We find inspiration in moment measures, a result that states that for any
Brune Massoulié, Clément Erignoux, Cristina Toninelli, Werner Krauth
We discuss non-reversible Markov-chain Monte Carlo algorithms that, for particle systems, rigorously sample the positional Boltzmann distribution and that have faster than physical dynamics. These algorithms all feature a non-thermal velocity distribution. They are exemplified by the lifted TASEP (totally asymmetric simple exclusion process), a one-dimension
Naomi Sagan, Amir Dembo, Matthew Ho, Tsachy Weissman
We study a family of processes generated according to sequential probability assignments induced by the LZ78 universal compressor. We characterize entropic and distributional properties such as their entropy and relative entropy rates, finite-state compressibility and log loss of their realizations, and the empirical distributions that they induce. Though no
Afrar Jahin, Arif Hassan Zidan, Wei Zhang, Yu Bao
With the rapid advancement of Artificial Intelligence (AI), Large Language Models (LLMs) have significantly impacted a wide array of domains, including healthcare, engineering, science, education, and mathematical reasoning. Among these, mathematical reasoning remains a particularly challenging capability, often requiring multi-step logic and abstract genera
David Criens, Michael Kupper
The objective of this paper is to investigate the connection between penalty functions from stochastic optimal control, convex semigroups from analysis and convex expectations from probability theory. Our main result provides a one-to-one relation between these objects. As an application, we use the representation via penality functions and duality arguments
Yongchang Hao, Mengyao Zhai, Hossein Hajimirsadeghi, Sepidehsadat Hosseini
Transformer models have demonstrated exceptional performance across a wide range of applications. Though forming the foundation of Transformer models, the dot-product attention does not scale well to long-context data since its time requirement grows quadratically with context length. In this work, we propose Radar, a training-free approach that accelerates
A Simple Description of the Hyperk\"{a}hler Structure of the Cotangent Bundle of Projective Space via Quantization
math.SGJoshua Lackman
Quantization identifies the cotangent bundle of projective space with the (non-Hermitian) rank-$1$ projections of a Hilbert space. We use this identification to study the natural geometric structures of these cotangent bundles and those of Grassmanians. In particular, we show that the quantization map is an isometric and complex embedding $T^*\mathbb{P}\math
Mingzhou Yin, Matthias A. Müller
Low-rank matrix regression is a fundamental problem in data science with various applications in systems and control. Nuclear norm regularization has been widely applied to solve this problem due to its convexity. However, it suffers from high computational complexity and the inability to directly specify the rank. This work introduces a novel framework for
Haopeng Li, Jinyue Yang, Guoqi Li, Huan Wang
We introduce ARPG, a novel visual Autoregressive model that enables Randomized Parallel Generation, addressing the inherent limitations of conventional raster-order approaches, which hinder inference efficiency and zero-shot generalization due to their sequential, predefined token generation order. Our key insight is that effective random-order modeling nece
Nannan Wu, Zengqiang Yan, Nong Sang, Li Yu
Training a model that effectively handles both common and rare data-i.e., achieving performance fairness-is crucial in federated learning (FL). While existing fair FL methods have shown effectiveness, they remain vulnerable to mislabeled data. Ensuring robustness in fair FL is therefore essential. However, fairness and robustness inherently compete, which ca
Egor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova
Despite their remarkable performance, large language models lack elementary safety features, making them susceptible to numerous malicious attacks. In particular, previous work has identified the absence of an intrinsic separation between instructions and data as the root cause of the success of prompt injection attacks. In this work, we propose a new archit
Marek Gazdzicki, Daniel Kikola, Ivan Pidhurskyi, Leonardo Tinti
Teleportation, introduced in science fiction literature, is an instantaneous change of the position of a macroscopic object. Two teleportation-like phenomena have been predicted by quantum mechanics: quantum teleportation and, more recently, quantum particle teleportation. Here, we introduce the third teleportation-like phenomenon - apparent teleportation. I
Controlling the dynamical phase diagram of a spinor BEC using time-dependent potentials
cond-mat.quant-gasQ. Guan, D. Blume, R. J. Lewis-Swan
We theoretically investigate the spin-mixing dynamics of a spinor BEC subject to a time-dependent confining potential. Our study provides a theory framework for the experimental results reported in Phys. Rev. A 109, 043309 (2024). We exploit the disparity of energy scales associated with the spatial and internal (spin) degrees of freedom under typical experi
Discontinuous Galerkin discretization of conservative dynamical low-rank approximation schemes for the Vlasov-Poisson equation
math.NAAndré Uschmajew, Andreas Zeiser
A numerical dynamical low-rank approximation (DLRA) scheme for the solution of the Vlasov-Poisson equation is presented. Based on the formulation of the DLRA equations as Friedrichs' systems in a continuous setting, it combines recently proposed conservative DLRA methods with a discontinuous Galerkin discretization. The resulting scheme is shown to ensure ma
Soham Das, Santiago Paternain, Luiz F. O. Chamon, Ceyhun Eksin
We propose the concept of a Lagrangian game to solve constrained Markov games. Such games model scenarios where agents face cost constraints in addition to their individual rewards, that depend on both agent joint actions and the evolving environment state over time. Constrained Markov games form the formal mechanism behind safe multiagent reinforcement lear
Kirill Solovev, Nicolas Pröllochs
Community-based fact-checking is a promising approach to address misinformation on social media at scale. However, an understanding of what makes community-created fact-checks helpful to users is still in its infancy. In this paper, we analyze the determinants of the helpfulness of community-created fact-checks. For this purpose, we draw upon a unique datase
Georg Jäger, Nils-Jonathan Friedrich, Hauke Petersen, Benjamin Noack
Robot navigation in complex environments necessitates controllers that prioritize safety while remaining performant and adaptable. Traditional controllers like Regulated Pure Pursuit, Dynamic Window Approach, and Model-Predictive Path Integral, while reliable, struggle to adapt to dynamic conditions. Reinforcement Learning offers adaptability but state-wise
Stijn Cambie
We establish the order of the maximum length of an increasing sequence, bounded by $n$, in which the largest prime divisor of the elements form a decreasing sequence.
Benoît Collins, Akihiro Miyagawa
We exhibit several bounds for operator norms of the sum of $\epsilon$-free semicircular random variables introduced in the paper of Speicher and Wysocza\'{n}ski. In particular, using the first and second largest eigenvalues of the adjacency matrix $\epsilon$, we show analogs of the operator-valued Khintchine-type inequality obtained by Haagerup and Pisier.
Margot Niels, Tom Vanackere, Ewoud Vissers, Tingting Zhai
The rapid expansion of cloud computing and artificial intelligence has driven the demand for faster optical components in data centres to unprecedented levels. A key advancement in this field is the integration of multiple photonic components onto a single chip, enhancing the performance of optical transceivers. Here, silicon photonics, benefiting from matur
Short-term AI literacy intervention does not reduce over-reliance on incorrect ChatGPT recommendations
cs.CYBrett Puppart, Jaan Aru
In this study, we examined whether a short-form AI literacy intervention could reduce the adoption of incorrect recommendations from large language models. High school seniors were randomly assigned to either a control or an intervention group, which received an educational text explaining ChatGPT's working mechanism, limitations, and proper use. Participant
Ofir Gorodetsky, Mo Dick Wong
Let $\alpha$ be a Steinhaus random multiplicative function. For a wide class of multiplicative functions $f$ we construct a multiplicative chaos measure arising from the Dirichlet series of $\alpha f$, in the whole $L^1$-regime. Our method does not rely on the thick point approach or Gaussian approximation, and uses a modified second moment method with the h