Skip to content

May 2025 arXiv papers — page 18

Showing 1,7011,800 of 24,552 papers

  1. Haohan Chi, Huan-ang Gao, Ziming Liu, Jianing Liu

    Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced f

  2. Justin Lazarow, Kai Kang, Afshin Dehghan

    We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM)

  3. Emilie Despontin, Stephane Detournay, Sudipta Dutta, Dima Fontaine

    We investigate anisotropic conformal Carroll field theories and their holographic duals. On the field theory side, we focus on the case with scaling exponent $z=0$ in two and three spacetime dimensions. These theories exhibit infinite-dimensional symmetry algebras, including supertranslations and superrotations, and are closely related to, but distinct from,

  4. Ziyin Zhang, Jiahao Xu, Zhiwei He, Tian Liang

    Theorem proving serves as a major testbed for evaluating complex reasoning abilities in large language models (LLMs). However, traditional automated theorem proving (ATP) approaches rely heavily on formal proof systems that poorly align with LLMs' strength derived from informal, natural language knowledge acquired during pre-training. In this work, we propos

  5. Akashah Shabbir, Muhammad Akhtar Munir, Akshay Dudhane, Muhammad Umer Sheikh

    Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing evaluations often focus on general-purpose or multimodal scenarios, leaving a gap in domain-specific benchmarks that assess tool-use capabilities in complex remote sensing use cases. W

  6. Declan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell

    Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invarian

  7. Rodrigo Voivodic

    This work presents a formalism for deriving likelihoods of the cosmological density field directly from first principles within Perturbation Theory (PT). By assuming a perturbative expansion around the Gaussian initial density field and additional stochastic components, we analytically compute two forms of the likelihood. Full marginalization over all underl

  8. Paul Gölz, Nika Haghtalab, Kunhe Yang

    After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have diverse preferences. As a result, it is not even clear that

  9. Shay Sadovsky, Gaoyong Zhang

    This paper establishes two new geometric inequalities in the dual Brunn-Minkowski theory. The first, originally conjectured by Lutwak, is the Brunn-Minkowski inequality for dual quermassintegrals of origin-symmetric convex bodies. The second, generalizing Ball's volume ratio inequality, is a reverse isoperimetric inequality: among all origin-symmetric convex

  10. Hugo Henry, Kelly Cohen

    This study investigates the application of Genetic Fuzzy Systems (GFS) to model the self-noise generated by airfoils, a key issue in aeroaccoustics with significant implications for aerospace, automotive and drone applications. Using the publicly available Airfoil Self Noise dataset, various Fuzzy regression strategies are explored and compared. The paper ev

  11. Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi

    Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and generation. While they excel in zero-shot and transfer learning scenarios, VLMs remain susceptible to misclassification, often yielding confident yet incorrect predictions. This l

  12. Qiang Wang, Xiang Song, Yuhang He, Jizhou Han

    Deep neural networks (DNNs) often underperform in real-world, dynamic settings where data distributions change over time. Domain Incremental Learning (DIL) offers a solution by enabling continual model adaptation, with Parameter-Isolation DIL (PIDIL) emerging as a promising paradigm to reduce knowledge conflicts. However, existing PIDIL methods struggle with

  13. Bowei Chen, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz

    Self-captured full-body videos are popular, but most deployments require mounted cameras, carefully-framed shots, and repeated practice. We propose a more convenient solution that enables full-body video capture using handheld mobile devices. Our approach takes as input two static photos (front and back) of you in a mirror, along with an IMU motion reference

  14. Amber Yijia Zheng, Yu Zhang, Jun Hu, Raymond A. Yeh

    High-quality photography in extreme low-light conditions is challenging but impactful for digital cameras. With advanced computing hardware, traditional camera image signal processor (ISP) algorithms are gradually being replaced by efficient deep networks that enhance noisy raw images more intelligently. However, existing regression-based models often minimi

  15. Yufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang

    We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference subjects, and copy-paste artifacts. To address these issues,

  16. Tudor D. Stanescu, Sumanta Tewari

    In a recent experiment, flux dependent oscillations of the quantum capacitance were observed in a one dimensional spin-orbit coupled semiconductor superconductor heterostructure connected end to end via a quantum dot and threaded by a magnetic flux. In the topological superconducting phase of the heterostructure, the oscillations corresponding to different f

  17. Ronghuan Wu, Wanchao Su, Jing Liao

    Image vectorization is a powerful technique that converts raster images into vector graphics, enabling enhanced flexibility and interactivity. However, popular image vectorization tools struggle with occluded regions, producing incomplete or fragmented shapes that hinder editability. While recent advancements have explored optimization-based and learning-bas

  18. Connar Rowan, Henry Whitehead, Gaia Fabj, Philip Kirkeberg

    The frequency of compact object interactions in AGN discs is naturally tied to the number of objects embedded within it. We investigate the evolution of black holes in the nuclear stellar cluster on inclined orbits to the AGN disc by performing adiabatic hydrodynamical simulations of isolated black hole disc crossings over a range of disc densities and incli

  19. Xiaojuan Wang, Aleksander Holynski, Brian Curless, Ira Kemelmacher

    We present a framework for generating music-synchronized, choreography aware animal dance videos. Our framework introduces choreography patterns -- structured sequences of motion beats that define the long-range structure of a dance -- as a novel high-level control signal for dance video generation. These patterns can be automatically estimated from human da

  20. Carl Ingebretsen, Bryce T. Bolin, Robert Jedicke, Peter Vereš

    Imminent impactors may be detected only a few hours before their impact with Earth, providing a brief opportunity to characterize them before impact. We describe the characterization of imminent impactor 2024 RW$_1$, which was discovered by the Catalina Sky Survey on 2024 September 4 at 05:43 UTC, before it entered the atmosphere near the northern Philippine

  21. Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri

    Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale. Their quadratic memory and time complexity, however, bound their applicability in longer sequences and so has motivated researchers to explore effective alternative architect

  22. Weijie Wang, Donny Y. Chen, Zeyu Zhang, Duochao Shi

    Feed-forward 3D Gaussian Splatting (3DGS) models have recently emerged as a promising solution for novel view synthesis, enabling one-pass inference without the need for per-scene 3DGS optimization. However, their scalability is fundamentally constrained by the limited capacity of their models, leading to degraded performance or excessive memory consumption

  23. Truong Jack Luu, Binny M. Samuel

    The democratization of generative AI introduces new forms of human-AI interaction and raises urgent safety, ethical, and cybersecurity concerns. We develop a socio-technical explanation for how generative AI enables and scales cybercrime. Drawing on affordance theory and technological amplification, we argue that generative AI systems create new action possi

  24. Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso

    Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by na\"ively aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal nature of emotions, hindering inter-emotion understanding and often resulting in a wide modality gap between the audio and text embeddings due

  25. P. J. Pessi, R. Lunnan, J. Sollerman, L. Yan

    AT2022rze is a luminous, ambiguous transient located South-East of the geometric center of its host galaxy at redshift z = 0.08. The host appears to be formed by a merging galaxy system. The observed characteristics of AT2022rze are reminiscent of active galactic nuclei (AGN), tidal disruption events (TDEs), and superluminous supernovae (SLSNe). The transien

  26. Jun-Hsiang Yao, Mingzheng Li, Jiayi Liu, Yuxiao Li

    The Digital Twin Brain (DTB) is an advanced artificial intelligence framework that integrates spiking neurons to simulate complex cognitive functions and collaborative behaviors. For domain experts, visualizing the DTB's simulation outcomes is essential to understanding complex cognitive activities. However, this task poses significant challenges due to DTB

  27. Mohamad Chehade, Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy

    Aligning large language models with humans is challenging due to the inherently multifaceted nature of preference feedback. While existing approaches typically frame this as a multi-objective optimization problem, they often overlook how humans actually make decisions. Research on bounded rationality suggests that human decision making follows satisficing st

  28. Song Wang, Gongfan Fang, Lingdong Kong, Xiangtai Li

    Existing reasoning segmentation approaches typically fine-tune multimodal large language models (MLLMs) using image-text pairs and corresponding mask labels. However, they exhibit limited generalization to out-of-distribution scenarios without an explicit reasoning process. Although recent efforts leverage reinforcement learning through group-relative policy

  29. Darryl Hannan, Timothy Doster, Henry Kvinge, Adam Attarian

    Collecting high quality data for object detection tasks is challenging due to the inherent subjectivity in labeling the boundaries of an object. This makes it difficult to not only collect consistent annotations across a dataset but also to validate them, as no two annotators are likely to label the same object using the exact same coordinates. These challen

  30. Minrui Luo, Fuhang Kuang, Yu Wang, Zirui Liu

    Parameter-Efficient Fine-Tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA), are indispensable for efficiently customizing Large Language Models (LLMs). However, vanilla LoRA suffers from slow convergence speed and knowledge forgetting problems. Recent studies have leveraged the power of designed LoRA initialization, to enhance the fine-tuning ef

  31. Zexi Liu, Jingyi Chai, Xinyu Zhu, Shuo Tang

    The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering. However, the dominant prompt-based paradigm exhibits limitations: smaller models lack the capacity to learn from execution trajectories for generalization, while large proprietary models incur high computational

  32. Fan Bai, Hamid Hassanzadeh, Ardavan Saeedi, Mark Dredze

    In-context learning (ICL) enables large language models (LLMs) to perform new tasks using only a few demonstrations. However, in Named Entity Recognition (NER), existing ICL methods typically rely on task-agnostic semantic similarity for demonstration retrieval, which often yields less relevant examples and leads to inferior results. We introduce DEER, a tra

  33. Sean Current, Ziqi Chen, Daniel Adu-Ampratwum, Xia Ning

    Methods for automatic chemical retrosynthesis have found recent success through the application of models traditionally built for natural language processing, primarily through transformer neural networks. These models have demonstrated significant ability to translate between the SMILES encodings of chemical products and reactants, but are constrained as a

  34. Arun Verma, Indrajit Saha, Makoto Yokoo, Bryan Kian Hsiang Low

    This paper considers a contextual bandit problem involving multiple agents, where a learner sequentially observes the contexts and the agent's reported arms, and then selects the arm that maximizes the system's overall reward. Existing work in contextual bandits assumes that agents truthfully report their arms, which is unrealistic in many real-life applicat

  35. Andreas Auer, Patrick Podest, Daniel Klotz, Sebastian Böck

    In-context learning, the ability of large language models to perform tasks using only examples provided in the prompt, has recently been adapted for time series forecasting. This paradigm enables zero-shot prediction, where past values serve as context for forecasting future values, making powerful forecasting tools accessible to non-experts and increasing t

  36. Saulo Queiroz

    In this work, we present the \emph{twiddless fast Fourier transform (TFFT)}, a novel algorithm for computing the $N$-point discrete Fourier transform (DFT). The TFFT's divide strategy builds on recent results that decimate an $N$-point signal (by a factor of $p$) into an $N/p$-point compressed signal whose DFT readily yields $N/p$ coefficients of the origina

  37. Mengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie

    Large Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature. Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all components when applied to new domains. We introduce Workforce

  38. Olaf Dössel, Axel Loewe

    This review focuses on the computerized modeling of the electrophysiology of the human atria, emphasizing the simulation of common arrhythmias such as atrial flutter (AFlut) and atrial fibrillation (AFib). Which components of the model are necessary to accurately model arrhythmogenic tissue modifications, including remodeling, cardiomyopathy, and fibrosis, t

  39. Tianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang

    Test-Time Training (TTT) models context dependencies by adapting part of the model's weights (referred to as fast weights) during inference. This fast weight, akin to recurrent states in RNNs, stores temporary memories of past tokens in the current sequence. Existing TTT methods struggled to show effectiveness in handling long-context data, due to their inef

  40. Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu

    We introduce AnySplat, a feed forward network for novel view synthesis from uncalibrated image collections. In contrast to traditional neural rendering pipelines that demand known camera poses and per scene optimization, or recent feed forward methods that buckle under the computational weight of dense views, our model predicts everything in one shot. A sing

  41. Jinzhe Li, Gengxu Li, Yi Chang, Yuan Wu

    Large language models (LLMs) have witnessed rapid advancements, demonstrating remarkable capabilities. However, a notable vulnerability persists: LLMs often uncritically accept flawed or contradictory premises, leading to inefficient reasoning and unreliable outputs. This emphasizes the significance of possessing the \textbf{Premise Critique Ability} for LLM

  42. Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo, Matthew J Thompson

    Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the

  43. Roksana Goworek, Harpal Karlcut, Muhammad Shezad, Nijaguna Darshana

    This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and s

  44. Zixiang Xu, Yanbo Wang, Yue Huang, Jiayi Ye

    Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning ability - the capacity to interpret social contexts, infer others' mental states, and assess the truthfulness of presented in

  45. Mark P. Hertzberg, Oleksandr S. Stashko

    We study static, spherically symmetric neutron stars in a class of scalar-tensor theories with non-canonical kinetic terms (K-essence) obeying all causality and hyperbolicity conditions. These models have non-trivial dynamics that lead to a type of anti-screening of the scalar. They lead to small corrections in the solar system due to a small coupling, but c

  46. Vyacheslav R. Misko, Franco Nori, Wim De Malsche

    Selecting active matter based on its motility represents a challenging task, as it requires different approaches than common separation techniques intended for separation based on, e.g., size, shape, density, and flexibility. This motility-based selection is important for, e.g., selecting biological species, such as bacteria or highly motile sperm cells for

  47. Christopher D. Rosin

    Large Language Models (LLMs) with reasoning are trained to iteratively generate and refine their answers before finalizing them, which can help with applications to mathematics and code generation. We apply code generation with reasoning LLMs to a specific task in the mathematical field of combinatorial design. This field studies diverse types of combinatori

  48. Anja Randecker

    Siegel-Veech constants are powerful tools for counting saddle connections on a translation surface. Their computation can be involved, most famously with recursive formulas that use intricate combinatorics or intersection theory. From these formulas, asymptotics of Siegel-Veech constants for growing genus can be extracted. We extend the known asymptotics to

  49. Zeinab Nezami, Syed Danial Ali Shah, Maryam Hafeez, Karim Djemame

    This paper envisions 6G as a self-evolving telecom ecosystem, where AI-driven intelligence enables dynamic adaptation beyond static connectivity. We explore the key enablers of autonomous communication systems, spanning reconfigurable infrastructure, adaptive middleware, and intelligent network functions, alongside multi-agent collaboration for distributed d

  50. Dionysis Christopoulos, Sotiris Spanos, Eirini Baltzi, Valsamis Ntouskos

    We introduce SLIMP (Skin Lesion Image-Metadata Pre-training) for learning rich representations of skin lesions through a novel nested contrastive learning approach that captures complex relationships between images and metadata. Melanoma detection and skin lesion classification based solely on images, pose significant challenges due to large variations in im

  51. Lucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller

    Reinforcement learning (RL) has significantly advanced the control of physics-based and robotic characters that track kinematic reference motion. However, methods typically rely on a weighted sum of conflicting reward functions, requiring extensive tuning to achieve a desired behavior. Due to the computational cost of RL, this iterative process is a tedious,

  52. Shenzhe Zhu, Jiao Sun, Yi Nian, Tobin South

    AI agents are increasingly used in consumer-facing applications to assist with tasks such as product search, negotiation, and transaction execution. In this paper, we explore a future scenario where both consumers and merchants authorize AI agents to fully automate negotiations and transactions. We aim to answer two key questions: (1) Do different LLM agents

  53. José Á. Sánchez Gómez, Weibin Mo, Junlong Zhao, Yufeng Liu

    Graphical models are popular tools for exploring relationships among a set of variables. The Gaussian graphical model (GGM) is an important class of graphical models, where the conditional dependence among variables is represented by nodes and edges in a graph. In many real applications, we are interested in detecting hubs in graphical models, which refer to

  54. Utku Demir, Yalin E. Sagduyu, Tugba Erpek, Hossein Jafari

    In connected and autonomous vehicles, machine learning for safety message classification has become critical for detecting malicious or anomalous behavior. However, conventional approaches that rely on centralized data collection or purely local training face limitations due to the large scale, high mobility, and heterogeneous data distributions inherent in

  55. Danny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu

    Vision-language-action (VLA) models provide a powerful approach to training control policies for physical systems, such as robots, by combining end-to-end learning with transfer of semantic knowledge from web-scale vision-language model (VLM) training. However, the constraints of real-time control are often at odds with the design of VLMs: the most powerful

  56. Mohamad Alansari, Sajid Javed, Iyyakutti Iyappan Ganapathi, Sara Alansari

    VOT remains a fundamental yet challenging task in computer vision due to dynamic appearance changes, occlusions, and background clutter. Traditional trackers, relying primarily on visual cues, often struggle in such complex scenarios. Recent advancements in VLMs have shown promise in semantic understanding for tasks like open-vocabulary detection and image c

  57. Ruida Wang, Yuxin Li, Yi R. Fung, Tong Zhang

    Enhancing the mathematical reasoning capabilities of LLMs has garnered significant attention in both the mathematical and computer science communities. Recent works have made substantial progress in both Natural Language (NL) reasoning and Formal Language (FL) reasoning by leveraging the potential of pure Reinforcement Learning (RL) methods on base models. H

  58. Nathan Lichtlé, Alexi Canesse, Zhe Fu, Hossein Nick Zinat Matin

    We introduce (U)NFV, a modular neural network architecture that generalizes classical finite volume (FV) methods for solving hyperbolic conservation laws. Hyperbolic partial differential equations (PDEs) are challenging to solve, particularly conservation laws whose physically relevant solutions contain shocks and discontinuities. FV methods are widely used

  59. Ziling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao

    Final-answer-based metrics are commonly used for evaluating large language models (LLMs) on math word problems, often taken as proxies for reasoning ability. However, such metrics conflate two distinct sub-skills: abstract formulation (capturing mathematical relationships using expressions) and arithmetic computation (executing the calculations). Through a d

  60. Oleksii Furman, Ulvi Movsum-zada, Patryk Marszalek, Maciej Zięba

    Counterfactual explanations play a pivotal role in explainable artificial intelligence (XAI) by offering intuitive, human-understandable alternatives that elucidate machine learning model decisions. Despite their significance, existing methods for generating counterfactuals often require constant access to the predictive model, involve computationally intens

  61. Ruben Burkard, Benedikt Schneider, Björn Sbierski

    The high-temperature series expansion for quantum spin models is a well-established tool to compute thermodynamic quantities and equal-time spin correlations, in particular for frustrated interactions. We extend the scope of this expansion to the dynamic Matsubara spin-spin correlator and develop an algorithm that yields exact expansion coefficients in the f

  62. Rashmiranjan Bhutia, Stephy Jose, Prasad Perlekar, Kabir Ramola

    Theoretical descriptions of the stepping-stone model, a cornerstone of spatial population genetics, have long overlooked diffusive noise arising from migration dynamics. We derive an exact fluctuating hydrodynamic description of this model from microscopic rules, which we then use to demonstrate that diffusive noise significantly alters early-time genetic de

  63. Ajay Sharma, Sakshi Chaudhary, Aishwarya Sarath, Debanjan Bose

    A comprehensive analysis of quasi-periodic oscillations (QPOs) in the gamma-ray emissions of blazars. Utilizing 15 years of Fermi-LAT observations of seven blazars in our sample, we identify both long-term and transient quasi-periodic oscillations in the gamma-ray light curves, with timescales ranging from a few months to years. These periodicities were dete

  64. Hiroshi Kera, Nico Pelleriti, Yuki Ishihara, Max Zimmer

    Solving systems of polynomial equations, particularly those with finitely many solutions, is a crucial challenge across many scientific fields. Traditional methods like Gröbner and Border bases are fundamental but suffer from high computational costs, which have motivated recent Deep Learning approaches to improve efficiency, albeit at the expense of output

  65. Madelyne Xiao, Palak Jain, Micha Gorelick, Sarah Scheffler

    WhatsApp and many other commonly used communication platforms guarantee end-to-end encryption (E2EE), which requires that service providers lack the cryptographic keys to read communications on their own platforms. WhatsApp's privacy-preserving design makes it difficult to study important phenomena like the spread of misinformation or political messaging, as

  66. Ran Zhang, Mohannad Elhamod

    The rapid advancement of LLMs has led to the creation of diverse agentic systems in data analysis, utilizing LLMs' capabilities to improve insight generation and visualization. In this paper, we present an agentic system that automates the data-to-dashboard pipeline through modular LLM agents capable of domain detection, concept extraction, multi-perspective

  67. Li Ren, Chen Chen, Liqiang Wang, Kien Hua

    Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models by partially fine-tuning learnable tokens while keeping most model parameters frozen. Recent research has explored modifying the connection structures of the prompts. However, the fundamental correlation and distr

  68. Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao

    MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on interpreting AIGC videos rema

  69. Jingyun Yang, Isabella Huang, Brandon Vu, Max Bajracharya

    Learned visuomotor policies are capable of performing increasingly complex manipulation tasks. However, most of these policies are trained on data collected from limited robot positions and camera viewpoints. This leads to poor generalization to novel robot positions, which limits the use of these policies on mobile platforms, especially for precise tasks li

  70. Hao Tian, Shengmin Jin, Reza Zafarani

    The spectral properties of traditional (dyadic) graphs, where an edge connects exactly two vertices, are widely studied in different applications. These spectral properties are closely connected to the structural properties of dyadic graphs. We generalize such connections and characterize higher-order networks by their spectral information. We first split th

  71. Arul Shankar, Ila Varma

    We compute the asymptotic number of octic number fields whose Galois groups over $\mathbb Q$ are isomorphic to $D_4$, the symmetries of a square, when ordering such fields by their absolute discriminants. In particular, we verify the strong form of Malle's conjecture for such octic $D_4$-fields and obtain the constant of proportionality. Our result answers t

  72. Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza

    Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data. However, the ge

  73. James Tanner, Morgan Sonderegger, Jane Stuart-Smith, Jeff Mielke

    Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic in

  74. Hanbo Xu, Xinyang Liu, Lei Wang

    This study presents an optimized hybrid design integrating a distributed Bragg reflector (DBR) and a TiO2 nanocylinder metasurface to enhance light extraction efficiency (LEE) and beam directionality(narrow divergence angle) in light-emitting diodes (LEDs) based on gallium nitride (GaN).Parametric simulations were used to identify an optimal device architect

  75. Caroline Wang, Arrasy Rahman, Jiaxun Cui, Yoonchang Sung

    Learning to collaborate with previously unseen partners is a fundamental generalization challenge in multi-agent learning, known as Ad Hoc Teamwork (AHT). Existing AHT approaches often adopt a two-stage pipeline, where first, a fixed population of teammates is generated with the idea that they should be representative of the teammates that will be seen at de

  76. Raffles Xingqi Zhu, Charlie S. Burlingham, Olivier Mercier, Phillip Guan

    Stereoscopic head-mounted displays (HMDs) render and present binocular images to create an egocentric, 3D percept to the HMD user. Within this render and presentation pipeline there are potential rendering camera and viewing position errors that can induce deviations in the depth and distance that a user perceives compared to the underlying intended geometry

  77. Martin Obaidi, Jakob Droste, Hannah Deters, Marc Herrmann

    As software systems grow increasingly complex, explainability has become a crucial non-functional requirement for transparency, user trust, and regulatory compliance. Eliciting explainability requirements is challenging, as different methods capture varying levels of detail and structure. This study examines the efficiency and effectiveness of three commonly

  78. Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian

    Transformer-based language models have demonstrated impressive capabilities across a range of complex reasoning tasks. Prior theoretical work exploring the expressive power of transformers has shown that they can efficiently perform multi-step reasoning tasks involving parallelizable computations. However, the learnability of such constructions, particularly

  79. Rachel Cummings, Alessandro Epasto, Jieming Mao, Tamalika Mukherjee

    The turnstile continual release model of differential privacy captures scenarios where a privacy-preserving real-time analysis is sought for a dataset evolving through additions and deletions. In typical applications of real-time data analysis, both the length of the stream $T$ and the size of the universe $|U|$ from which data come can be extremely large. T

  80. Bo Zhao, Nima Dehmamy, Robin Walters, Rose Yu

    Neural network minima are often connected by curves along which train and test loss remain nearly constant, a phenomenon known as mode connectivity. While this property has enabled applications such as model merging and fine-tuning, its theoretical explanation remains unclear. We propose a new approach to exploring the connectedness of minima using parameter

  81. Farshad Rostami Ghadi, Kai-Kit Wong, F. Javier Lopez-Martinez, George C. Alexandropoulos

    This letter investigates the performance of emerging wireless communication systems assisted by a fluid reconfigurable intelligent surface (FRIS). Unlike conventional reconfigurable intelligent surfaces (RISs), an FRIS consists of fluid-inspired metamaterials arranged in a densely packed matrix of sub-elements over a surface. It dynamically activates specifi

  82. Ángel Cuevas, Javier Chagoya, C. Ortiz

    In the derivation of the Einstein field equations via Hamilton's principle, the inclusion of a boundary term is essential to render the variational problem well-posed, as it addresses variations that do not vanish at the boundary of the spacetime manifold. Typically, this term is chosen as the Gibbons-Hawking-York boundary term. In this work, we propose an a

  83. R. Bachev, Tushar Tripathi, Alok C. Gupta, A. Kurtenkov

    OT 355 (4FGL J1734.3 + 3858) is a relatively rarely studied but highly variable, moderate-redshift (z = 0.975) flat-spectrum radio quasar (blazar). With this work, we aim to study its optical variability on different timescales, which can help us to better understand the physical processes in relativistic jets operating in blazar-type active galactic nuclei.

  84. Piotr Bartman-Szwarc, Adil M. Bagirov, Anna Ochal

    In this paper, we employ a global aggregate subgradient method for the numerical solution of hemivariational inequality problems arising in contact mechanics. The method integrates a global search procedure to identify starting points for a local minimization algorithm. The algorithm consists of two types of steps: null steps and serious steps. In each null

  85. Moinak Bhattacharya, Judy Huang, Amna F. Sher, Gagandeep Singh

    Accurately predicting immunotherapy response in Non-Small Cell Lung Cancer (NSCLC) remains a critical unmet need. Existing radiomics and deep learning-based predictive models rely primarily on pre-treatment imaging to predict categorical response outcomes, limiting their ability to capture the complex morphological and textural transformations induced by imm

  86. Jianjun Zhao

    Quantum computing has demonstrated the potential to solve computationally intensive problems more efficiently than classical methods. Many software engineering tasks, such as test case selection, static analysis, code clone detection, and defect prediction, involve complex optimization, search, or classification, making them candidates for quantum enhancemen

  87. Aya Kayal, Sattar Vakili, Laura Toni, Da-shan Shiu

    Bayesian optimization (BO) with preference-based feedback has recently garnered significant attention due to its emerging applications. We refer to this problem as Bayesian Optimization from Human Feedback (BOHF), which differs from conventional BO by learning the best actions from a reduced feedback model, where only the preference between two actions is re

  88. Amir Said, Xin Zhao, Marta Karczewicz, Jianle Chen

    Intra-frame prediction in the High Efficiency Video Coding (HEVC) standard can be empirically improved by applying sets of recursive two-dimensional filters to the predicted values. However, this approach does not allow (or complicates significantly) the parallel computation of pixel predictions. In this work we analyze why the recursive filters are effectiv

  89. Manish Shetty, Naman Jain, Jinjian Liu, Vijay Kethanaboyina

    Developing high-performance software is a complex task that requires specialized expertise. We introduce GSO, a benchmark for evaluating language models' capabilities in developing high-performance software. We develop an automated pipeline that generates and executes performance tests to analyze repository commit histories to identify 102 challenging optimi

  90. C. W. J. Beenakker

    We calculate the full counting statistics of charge transfer in a chiral Majorana interferometer - a setup where a Dirac mode (an electron-hole mode) is split into two Majorana modes that encircle a number of h/2e vortices in a topological superconductor. Without any coupling to the environment it is known that the low-energy charge transfer is deterministic

  91. Syeda Abeera Amir, Artur Agaronyan, William Gaillard, Chima Oluigbo

    Accurately localizing the brain regions that triggers seizures and predicting whether a patient will be seizure-free after surgery are vital for surgical planning and patient management in drug-resistant epilepsy. Stereo-electroencephalography (sEEG) delivers high-fidelity intracranial recordings that enable clinicians to precisely locate epileptogenic netwo

  92. Sung Soo Moon, Sebastian E. Ahnert

    Many real-world networks have associated metadata that assigns categorical labels to nodes. Analysis of these annotations can complement the topological analysis of complex networks. Annotated networks have typically been used to evaluate community detection approaches. Here, we introduce an approach that combines the quantitative analysis of annotations and

  93. Lang Cao, Jingxian Xu, Hanbing Liu, Jinyu Wang

    Tables are a fundamental medium for organizing and analyzing data, making table reasoning a critical capability for intelligent systems. Although large language models (LLMs) exhibit strong general reasoning abilities, they still struggle with accurate numerical reasoning over tabular data, particularly in complex table settings beyond simple relational look

  94. Nariman Naderi, Zahra Atf, Peter R Lewis, Aref Mahjoub far

    This paper investigates how prompt engineering techniques impact both accuracy and confidence elicitation in Large Language Models (LLMs) applied to medical contexts. Using a stratified dataset of Persian board exam questions across multiple specialties, we evaluated five LLMs - GPT-4o, o3-mini, Llama-3.3-70b, Llama-3.1-8b, and DeepSeek-v3 - across 156 confi

  95. Jeremy Brazas, Atish Mitra

    The $\pi_n$-wild set $\mathbf{w}_{n}(X)$ of a topological space $X$ is the subspace of $X$ consisting of the points at which there exists a shrinking sequence of essential based maps $S^n\to X$. In this paper, we show that the homotopy type of $\mathbf{w}_{n}(X)$ is a homotopy invariant of $X$ and, in analogy to the known one-dimensional case, we show that f

  96. Ahmed Almheiri

    Bousso and Stanford (BS) argued that the black hole final state proposal leads to acausal effects and ill-defined probabilities for the AMPS experiment. We identify a loophole in their analysis using insights from entanglement wedge reconstruction and replica wormholes. We trace the cause of the BS problems to the misidentification of the physical interior w

  97. Niklas Freymuth, Tobias Würth, Nicolas Schreiber, Balazs Gyenes

    The cost and accuracy of simulating complex physical systems using the Finite Element Method (FEM) scales with the resolution of the underlying mesh. Adaptive meshes improve computational efficiency by refining resolution in critical regions, but typically require task-specific heuristics or cumbersome manual design by a human expert. We propose Adaptive Mes

  98. Beong-woo Kwak, Minju Kim, Dongha Lim, Hyungjoo Chae

    Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capa

  99. Size Wu, Zhonghua Wu, Zerui Gong, Qingyi Tao

    In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (

  100. Ziteng Gao, Mike Zheng Shou

    This paper presents Diffusion via Autoregressive models (D-AR), a new paradigm recasting the image diffusion process as a vanilla autoregressive procedure in the standard next-token-prediction fashion. We start by designing the tokenizer that converts images into sequences of discrete tokens, where tokens in different positions can be decoded into different