May 2025 arXiv papers — page 18
Showing 1,701–1,800 of 24,552 papers
Haohan Chi, Huan-ang Gao, Ziming Liu, Jianing Liu
Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced f
Justin Lazarow, Kai Kang, Afshin Dehghan
We revisit scene-level 3D object detection as the output of an object-centric framework capable of both localization and mapping using 3D oriented boxes as the underlying geometric primitive. While existing 3D object detection approaches operate globally and implicitly rely on the a priori existence of metric camera poses, our method, Rooms from Motion (RfM)
Emilie Despontin, Stephane Detournay, Sudipta Dutta, Dima Fontaine
We investigate anisotropic conformal Carroll field theories and their holographic duals. On the field theory side, we focus on the case with scaling exponent $z=0$ in two and three spacetime dimensions. These theories exhibit infinite-dimensional symmetry algebras, including supertranslations and superrotations, and are closely related to, but distinct from,
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
cs.CLZiyin Zhang, Jiahao Xu, Zhiwei He, Tian Liang
Theorem proving serves as a major testbed for evaluating complex reasoning abilities in large language models (LLMs). However, traditional automated theorem proving (ATP) approaches rely heavily on formal proof systems that poorly align with LLMs' strength derived from informal, natural language knowledge acquired during pre-training. In this work, we propos
Akashah Shabbir, Muhammad Akhtar Munir, Akshay Dudhane, Muhammad Umer Sheikh
Recent progress in large language models (LLMs) has enabled tool-augmented agents capable of solving complex real-world tasks through step-by-step reasoning. However, existing evaluations often focus on general-purpose or multimodal scenarios, leaving a gap in domain-specific benchmarks that assess tool-use capabilities in complex remote sensing use cases. W
Declan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell
Sequence models such as transformers require inputs to be represented as one-dimensional sequences. In vision, this typically involves flattening images using a fixed row-major (raster-scan) order. While full self-attention is permutation-equivariant, modern long-sequence transformers increasingly rely on architectural approximations that break this invarian
Rodrigo Voivodic
This work presents a formalism for deriving likelihoods of the cosmological density field directly from first principles within Perturbation Theory (PT). By assuming a perturbative expansion around the Gaussian initial density field and additional stochastic components, we analytically compute two forms of the likelihood. Full marginalization over all underl
Paul Gölz, Nika Haghtalab, Kunhe Yang
After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings where users have diverse preferences. As a result, it is not even clear that
Shay Sadovsky, Gaoyong Zhang
This paper establishes two new geometric inequalities in the dual Brunn-Minkowski theory. The first, originally conjectured by Lutwak, is the Brunn-Minkowski inequality for dual quermassintegrals of origin-symmetric convex bodies. The second, generalizing Ball's volume ratio inequality, is a reverse isoperimetric inequality: among all origin-symmetric convex
Hugo Henry, Kelly Cohen
This study investigates the application of Genetic Fuzzy Systems (GFS) to model the self-noise generated by airfoils, a key issue in aeroaccoustics with significant implications for aerospace, automotive and drone applications. Using the publicly available Airfoil Self Noise dataset, various Fuzzy regression strategies are explored and compared. The paper ev
Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi
Vision-Language Models (VLMs) have demonstrated strong capabilities in aligning visual and textual modalities, enabling a wide range of applications in multimodal understanding and generation. While they excel in zero-shot and transfer learning scenarios, VLMs remain susceptible to misclassification, often yielding confident yet incorrect predictions. This l
Qiang Wang, Xiang Song, Yuhang He, Jizhou Han
Deep neural networks (DNNs) often underperform in real-world, dynamic settings where data distributions change over time. Domain Incremental Learning (DIL) offers a solution by enabling continual model adaptation, with Parameter-Isolation DIL (PIDIL) emerging as a promising paradigm to reduce knowledge conflicts. However, existing PIDIL methods struggle with
Bowei Chen, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz
Self-captured full-body videos are popular, but most deployments require mounted cameras, carefully-framed shots, and repeated practice. We propose a more convenient solution that enables full-body video capture using handheld mobile devices. Our approach takes as input two static photos (front and back) of you in a mirror, along with an IMU motion reference
Amber Yijia Zheng, Yu Zhang, Jun Hu, Raymond A. Yeh
High-quality photography in extreme low-light conditions is challenging but impactful for digital cameras. With advanced computing hardware, traditional camera image signal processor (ISP) algorithms are gradually being replaced by efficient deep networks that enhance noisy raw images more intelligently. However, existing regression-based models often minimi
Yufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang
We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference subjects, and copy-paste artifacts. To address these issues,
Fermion parity and quantum capacitance oscillation with partially separated Majorana and quasi-Majorana modes
cond-mat.mes-hallTudor D. Stanescu, Sumanta Tewari
In a recent experiment, flux dependent oscillations of the quantum capacitance were observed in a one dimensional spin-orbit coupled semiconductor superconductor heterostructure connected end to end via a quantum dot and threaded by a magnetic flux. In the topological superconducting phase of the heterostructure, the oscillations corresponding to different f
Ronghuan Wu, Wanchao Su, Jing Liao
Image vectorization is a powerful technique that converts raster images into vector graphics, enabling enhanced flexibility and interactivity. However, popular image vectorization tools struggle with occluded regions, producing incomplete or fragmented shapes that hinder editability. While recent advancements have explored optimization-based and learning-bas
Hydrodynamic simulations of black hole evolution in AGN discs I: orbital alignment of highly inclined satellites
astro-ph.HEConnar Rowan, Henry Whitehead, Gaia Fabj, Philip Kirkeberg
The frequency of compact object interactions in AGN discs is naturally tied to the number of objects embedded within it. We investigate the evolution of black holes in the nuclear stellar cluster on inclined orbits to the AGN disc by performing adiabatic hydrodynamical simulations of isolated black hole disc crossings over a range of disc densities and incli
Xiaojuan Wang, Aleksander Holynski, Brian Curless, Ira Kemelmacher
We present a framework for generating music-synchronized, choreography aware animal dance videos. Our framework introduces choreography patterns -- structured sequences of motion beats that define the long-range structure of a dance -- as a novel high-level control signal for dance video generation. These patterns can be automatically estimated from human da
Carl Ingebretsen, Bryce T. Bolin, Robert Jedicke, Peter Vereš
Imminent impactors may be detected only a few hours before their impact with Earth, providing a brief opportunity to characterize them before impact. We describe the characterization of imminent impactor 2024 RW$_1$, which was discovered by the Catalina Sky Survey on 2024 September 4 at 05:43 UTC, before it entered the atmosphere near the northern Philippine
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri
Transformers have been established as the most popular backbones in sequence modeling, mainly due to their effectiveness in in-context retrieval tasks and the ability to learn at scale. Their quadratic memory and time complexity, however, bound their applicability in longer sequences and so has motivated researchers to explore effective alternative architect
Weijie Wang, Donny Y. Chen, Zeyu Zhang, Duochao Shi
Feed-forward 3D Gaussian Splatting (3DGS) models have recently emerged as a promising solution for novel view synthesis, enabling one-pass inference without the need for per-scene 3DGS optimization. However, their scalability is fundamentally constrained by the limited capacity of their models, leading to degraded performance or excessive memory consumption
Truong Jack Luu, Binny M. Samuel
The democratization of generative AI introduces new forms of human-AI interaction and raises urgent safety, ethical, and cybersecurity concerns. We develop a socio-technical explanation for how generative AI enables and scales cybercrime. Drawing on affordance theory and technological amplification, we argue that generative AI systems create new action possi
EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast
cs.LGShreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso
Current emotion-based contrastive language-audio pretraining (CLAP) methods typically learn by na\"ively aligning audio samples with corresponding text prompts. Consequently, this approach fails to capture the ordinal nature of emotions, hindering inter-emotion understanding and often resulting in a wide modality gap between the audio and text embeddings due
The ambiguous AT2022rze: Changing-look AGN mimicking a supernova in a merging galaxy system
astro-ph.HEP. J. Pessi, R. Lunnan, J. Sollerman, L. Yan
AT2022rze is a luminous, ambiguous transient located South-East of the geometric center of its host galaxy at redshift z = 0.08. The host appears to be formed by a merging galaxy system. The observed characteristics of AT2022rze are reminiscent of active galactic nuclei (AGN), tidal disruption events (TDEs), and superluminous supernovae (SLSNe). The transien
Jun-Hsiang Yao, Mingzheng Li, Jiayi Liu, Yuxiao Li
The Digital Twin Brain (DTB) is an advanced artificial intelligence framework that integrates spiking neurons to simulate complex cognitive functions and collaborative behaviors. For domain experts, visualizing the DTB's simulation outcomes is essential to understanding complex cognitive activities. However, this task poses significant challenges due to DTB
Mohamad Chehade, Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy
Aligning large language models with humans is challenging due to the inherently multifaceted nature of preference feedback. While existing approaches typically frame this as a multi-objective optimization problem, they often overlook how humans actually make decisions. Research on bounded rationality suggests that human decision making follows satisficing st
Song Wang, Gongfan Fang, Lingdong Kong, Xiangtai Li
Existing reasoning segmentation approaches typically fine-tune multimodal large language models (MLLMs) using image-text pairs and corresponding mask labels. However, they exhibit limited generalization to out-of-distribution scenarios without an explicit reasoning process. Although recent efforts leverage reinforcement learning through group-relative policy
Darryl Hannan, Timothy Doster, Henry Kvinge, Adam Attarian
Collecting high quality data for object detection tasks is challenging due to the inherent subjectivity in labeling the boundaries of an object. This makes it difficult to not only collect consistent annotations across a dataset but also to validate them, as no two annotators are likely to label the same object using the exact same coordinates. These challen
SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA
cs.LGMinrui Luo, Fuhang Kuang, Yu Wang, Zirui Liu
Parameter-Efficient Fine-Tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA), are indispensable for efficiently customizing Large Language Models (LLMs). However, vanilla LoRA suffers from slow convergence speed and knowledge forgetting problems. Recent studies have leveraged the power of designed LoRA initialization, to enhance the fine-tuning ef
Zexi Liu, Jingyi Chai, Xinyu Zhu, Shuo Tang
The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering. However, the dominant prompt-based paradigm exhibits limitations: smaller models lack the capacity to learn from execution trajectories for generalization, while large proprietary models incur high computational
Fan Bai, Hamid Hassanzadeh, Ardavan Saeedi, Mark Dredze
In-context learning (ICL) enables large language models (LLMs) to perform new tasks using only a few demonstrations. However, in Named Entity Recognition (NER), existing ICL methods typically rely on task-agnostic semantic similarity for demonstration retrieval, which often yields less relevant examples and leads to inferior results. We introduce DEER, a tra
Sean Current, Ziqi Chen, Daniel Adu-Ampratwum, Xia Ning
Methods for automatic chemical retrosynthesis have found recent success through the application of models traditionally built for natural language processing, primarily through transformer neural networks. These models have demonstrated significant ability to translate between the SMILES encodings of chemical products and reactants, but are constrained as a
Arun Verma, Indrajit Saha, Makoto Yokoo, Bryan Kian Hsiang Low
This paper considers a contextual bandit problem involving multiple agents, where a learner sequentially observes the contexts and the agent's reported arms, and then selects the arm that maximizes the system's overall reward. Existing work in contextual bandits assumes that agents truthfully report their arms, which is unrealistic in many real-life applicat
Andreas Auer, Patrick Podest, Daniel Klotz, Sebastian Böck
In-context learning, the ability of large language models to perform tasks using only examples provided in the prompt, has recently been adapted for time series forecasting. This paradigm enables zero-shot prediction, where past values serve as context for forecasting future values, making powerful forecasting tools accessible to non-experts and increasing t
Saulo Queiroz
In this work, we present the \emph{twiddless fast Fourier transform (TFFT)}, a novel algorithm for computing the $N$-point discrete Fourier transform (DFT). The TFFT's divide strategy builds on recent results that decimate an $N$-point signal (by a factor of $p$) into an $N/p$-point compressed signal whose DFT readily yields $N/p$ coefficients of the origina
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
cs.AIMengkang Hu, Yuhang Zhou, Wendong Fan, Yuzhou Nie
Large Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature. Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all components when applied to new domains. We introduce Workforce
Computerized Modeling of Electrophysiology and Pathoelectrophysiology of the Atria -- How Much Detail is Needed?
physics.comp-phOlaf Dössel, Axel Loewe
This review focuses on the computerized modeling of the electrophysiology of the human atria, emphasizing the simulation of common arrhythmias such as atrial flutter (AFlut) and atrial fibrillation (AFib). Which components of the model are necessary to accurately model arrhythmogenic tissue modifications, including remodeling, cardiomyopathy, and fibrosis, t
Tianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang
Test-Time Training (TTT) models context dependencies by adapting part of the model's weights (referred to as fast weights) during inference. This fast weight, akin to recurrent states in RNNs, stores temporary memories of past tokens in the current sequence. Existing TTT methods struggled to show effectiveness in handling long-context data, due to their inef
Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu
We introduce AnySplat, a feed forward network for novel view synthesis from uncalibrated image collections. In contrast to traditional neural rendering pipelines that demand known camera poses and per scene optimization, or recent feed forward methods that buckle under the computational weight of dense views, our model predicts everything in one shot. A sing
Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models
cs.CLJinzhe Li, Gengxu Li, Yi Chang, Yuan Wu
Large language models (LLMs) have witnessed rapid advancements, demonstrating remarkable capabilities. However, a notable vulnerability persists: LLMs often uncritically accept flawed or contradictory premises, leading to inefficient reasoning and unreliable outputs. This emphasizes the significance of possessing the \textbf{Premise Critique Ability} for LLM
Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo, Matthew J Thompson
Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the
Roksana Goworek, Harpal Karlcut, Muhammad Shezad, Nijaguna Darshana
This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and s
Zixiang Xu, Yanbo Wang, Yue Huang, Jiayi Ye
Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning ability - the capacity to interpret social contexts, infer others' mental states, and assess the truthfulness of presented in
Mark P. Hertzberg, Oleksandr S. Stashko
We study static, spherically symmetric neutron stars in a class of scalar-tensor theories with non-canonical kinetic terms (K-essence) obeying all causality and hyperbolicity conditions. These models have non-trivial dynamics that lead to a type of anti-screening of the scalar. They lead to small corrections in the solar system due to a small coupling, but c
Motility-dependent selective transport of active matter in trap arrays: Separation methods based on trapping-detrapping and deterministic lateral displacement
cond-mat.softVyacheslav R. Misko, Franco Nori, Wim De Malsche
Selecting active matter based on its motility represents a challenging task, as it requires different approaches than common separation techniques intended for separation based on, e.g., size, shape, density, and flexibility. This motility-based selection is important for, e.g., selecting biological species, such as bacteria or highly motile sperm cells for
Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems
cs.AIChristopher D. Rosin
Large Language Models (LLMs) with reasoning are trained to iteratively generate and refine their answers before finalizing them, which can help with applications to mathematics and code generation. We apply code generation with reasoning LLMs to a specific task in the mathematical field of combinatorial design. This field studies diverse types of combinatori
Anja Randecker
Siegel-Veech constants are powerful tools for counting saddle connections on a translation surface. Their computation can be involved, most famously with recursive formulas that use intricate combinatorics or intersection theory. From these formulas, asymptotics of Siegel-Veech constants for growing genus can be extracted. We extend the known asymptotics to
Zeinab Nezami, Syed Danial Ali Shah, Maryam Hafeez, Karim Djemame
This paper envisions 6G as a self-evolving telecom ecosystem, where AI-driven intelligence enables dynamic adaptation beyond static connectivity. We explore the key enablers of autonomous communication systems, spanning reconfigurable infrastructure, adaptive middleware, and intelligent network functions, alongside multi-agent collaboration for distributed d
Dionysis Christopoulos, Sotiris Spanos, Eirini Baltzi, Valsamis Ntouskos
We introduce SLIMP (Skin Lesion Image-Metadata Pre-training) for learning rich representations of skin lesions through a novel nested contrastive learning approach that captures complex relationships between images and metadata. Melanoma detection and skin lesion classification based solely on images, pose significant challenges due to large variations in im
Lucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller
Reinforcement learning (RL) has significantly advanced the control of physics-based and robotic characters that track kinematic reference motion. However, methods typically rely on a weighted sum of conflicting reward functions, requiring extensive tuning to achieve a desired behavior. Due to the computational cost of RL, this iterative process is a tedious,
The Automated but Risky Game: Modeling and Benchmarking Agent-to-Agent Negotiations and Transactions in Consumer Markets
cs.AIShenzhe Zhu, Jiao Sun, Yi Nian, Tobin South
AI agents are increasingly used in consumer-facing applications to assist with tasks such as product search, negotiation, and transaction execution. In this paper, we explore a future scenario where both consumers and merchants authorize AI agents to fully automate negotiations and transactions. We aim to answer two key questions: (1) Do different LLM agents
José Á. Sánchez Gómez, Weibin Mo, Junlong Zhao, Yufeng Liu
Graphical models are popular tools for exploring relationships among a set of variables. The Gaussian graphical model (GGM) is an important class of graphical models, where the conditional dependence among variables is represented by nodes and edges in a graph. In many real applications, we are interested in detecting hubs in graphical models, which refer to
Distributed Federated Learning for Vehicular Network Security: Anomaly Detection Benefits and Multi-Domain Attack Threats
cs.NIUtku Demir, Yalin E. Sagduyu, Tugba Erpek, Hossein Jafari
In connected and autonomous vehicles, machine learning for safety message classification has become critical for detecting malicious or anomalous behavior. However, conventional approaches that rely on centralized data collection or purely local training face limitations due to the large scale, high mobility, and heterogeneous data distributions inherent in
Danny Driess, Jost Tobias Springenberg, Brian Ichter, Lili Yu
Vision-language-action (VLA) models provide a powerful approach to training control policies for physical systems, such as robots, by combining end-to-end learning with transfer of semantic knowledge from web-scale vision-language model (VLM) training. However, the constraints of real-time control are often at odds with the design of VLMs: the most powerful
Mohamad Alansari, Sajid Javed, Iyyakutti Iyappan Ganapathi, Sara Alansari
VOT remains a fundamental yet challenging task in computer vision due to dynamic appearance changes, occlusions, and background clutter. Traditional trackers, relying primarily on visual cues, often struggle in such complex scenarios. Recent advancements in VLMs have shown promise in semantic understanding for tasks like open-vocabulary detection and image c
Ruida Wang, Yuxin Li, Yi R. Fung, Tong Zhang
Enhancing the mathematical reasoning capabilities of LLMs has garnered significant attention in both the mathematical and computer science communities. Recent works have made substantial progress in both Natural Language (NL) reasoning and Formal Language (FL) reasoning by leveraging the potential of pure Reinforcement Learning (RL) methods on base models. H
Nathan Lichtlé, Alexi Canesse, Zhe Fu, Hossein Nick Zinat Matin
We introduce (U)NFV, a modular neural network architecture that generalizes classical finite volume (FV) methods for solving hyperbolic conservation laws. Hyperbolic partial differential equations (PDEs) are challenging to solve, particularly conservation laws whose physically relevant solutions contain shocks and discontinuities. FV methods are widely used
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation
cs.CLZiling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao
Final-answer-based metrics are commonly used for evaluating large language models (LLMs) on math word problems, often taken as proxies for reasoning ability. However, such metrics conflate two distinct sub-skills: abstract formulation (capturing mathematical relationships using expressions) and arithmetic computation (executing the calculations). Through a d
Oleksii Furman, Ulvi Movsum-zada, Patryk Marszalek, Maciej Zięba
Counterfactual explanations play a pivotal role in explainable artificial intelligence (XAI) by offering intuitive, human-understandable alternatives that elucidate machine learning model decisions. Despite their significance, existing methods for generating counterfactuals often require constant access to the predictive model, involve computationally intens
Ruben Burkard, Benedikt Schneider, Björn Sbierski
The high-temperature series expansion for quantum spin models is a well-established tool to compute thermodynamic quantities and equal-time spin correlations, in particular for frustrated interactions. We extend the scope of this expansion to the dynamic Matsubara spin-spin correlator and develop an algorithm that yields exact expansion coefficients in the f
Rashmiranjan Bhutia, Stephy Jose, Prasad Perlekar, Kabir Ramola
Theoretical descriptions of the stepping-stone model, a cornerstone of spatial population genetics, have long overlooked diffusive noise arising from migration dynamics. We derive an exact fluctuating hydrodynamic description of this model from microscopic rules, which we then use to demonstrate that diffusive noise significantly alters early-time genetic de
Exploring Year-timescale Gamma-ray Quasi-Periodic Oscillations in Blazars: Evidence for Supermassive Binary Black Holes Scenario
astro-ph.HEAjay Sharma, Sakshi Chaudhary, Aishwarya Sarath, Debanjan Bose
A comprehensive analysis of quasi-periodic oscillations (QPOs) in the gamma-ray emissions of blazars. Utilizing 15 years of Fermi-LAT observations of seven blazars in our sample, we identify both long-term and transient quasi-periodic oscillations in the gamma-ray light curves, with timescales ranging from a few months to years. These periodicities were dete
Hiroshi Kera, Nico Pelleriti, Yuki Ishihara, Max Zimmer
Solving systems of polynomial equations, particularly those with finitely many solutions, is a crucial challenge across many scientific fields. Traditional methods like Gröbner and Border bases are fundamental but suffer from high computational costs, which have motivated recent Deep Learning approaches to improve efficiency, albeit at the expense of output
Madelyne Xiao, Palak Jain, Micha Gorelick, Sarah Scheffler
WhatsApp and many other commonly used communication platforms guarantee end-to-end encryption (E2EE), which requires that service providers lack the cryptographic keys to read communications on their own platforms. WhatsApp's privacy-preserving design makes it difficult to study important phenomena like the spread of misinformation or political messaging, as
Data-to-Dashboard: Multi-Agent LLM Framework for Insightful Visualization in Enterprise Analytics
cs.AIRan Zhang, Mohannad Elhamod
The rapid advancement of LLMs has led to the creation of diverse agentic systems in data analysis, utilizing LLMs' capabilities to improve insight generation and visualization. In this paper, we present an agentic system that automates the data-to-dashboard pipeline through modular LLM agents capable of domain detection, concept extraction, multi-perspective
Li Ren, Chen Chen, Liqiang Wang, Kien Hua
Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models by partially fine-tuning learnable tokens while keeping most model parameters frozen. Recent research has explored modifying the connection structures of the prompts. However, the fundamental correlation and distr
Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao
MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on interpreting AIGC videos rema
Jingyun Yang, Isabella Huang, Brandon Vu, Max Bajracharya
Learned visuomotor policies are capable of performing increasingly complex manipulation tasks. However, most of these policies are trained on data collected from limited robot positions and camera viewpoints. This leads to poor generalization to novel robot positions, which limits the use of these policies on mobile platforms, especially for precise tasks li
Hao Tian, Shengmin Jin, Reza Zafarani
The spectral properties of traditional (dyadic) graphs, where an edge connects exactly two vertices, are widely studied in different applications. These spectral properties are closely connected to the structural properties of dyadic graphs. We generalize such connections and characterize higher-order networks by their spectral information. We first split th
Arul Shankar, Ila Varma
We compute the asymptotic number of octic number fields whose Galois groups over $\mathbb Q$ are isomorphic to $D_4$, the symmetries of a square, when ordering such fields by their absolute discriminants. In particular, we verify the strong form of Malle's conjecture for such octic $D_4$-fields and obtain the constant of proportionality. Our result answers t
Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza
Seminal work by Huebner et al. (2021) showed that language models (LMs) trained on English Child-Directed Language (CDL) can reach similar syntactic abilities as LMs trained on much larger amounts of adult-directed written text, suggesting that CDL could provide more effective LM training material than the commonly used internet-crawled data. However, the ge
James Tanner, Morgan Sonderegger, Jane Stuart-Smith, Jeff Mielke
Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic in
Enhanced Light Extraction and Beam Focusing in GaN LEDs Using Hybrid Metasurface-Distributed Bragg Reflector Structures
physics.opticsHanbo Xu, Xinyang Liu, Lei Wang
This study presents an optimized hybrid design integrating a distributed Bragg reflector (DBR) and a TiO2 nanocylinder metasurface to enhance light extraction efficiency (LEE) and beam directionality(narrow divergence angle) in light-emitting diodes (LEDs) based on gallium nitride (GaN).Parametric simulations were used to identify an optimal device architect
Caroline Wang, Arrasy Rahman, Jiaxun Cui, Yoonchang Sung
Learning to collaborate with previously unseen partners is a fundamental generalization challenge in multi-agent learning, known as Ad Hoc Teamwork (AHT). Existing AHT approaches often adopt a two-stage pipeline, where first, a fixed population of teammates is generated with the idea that they should be representative of the teammates that will be seen at de
Raffles Xingqi Zhu, Charlie S. Burlingham, Olivier Mercier, Phillip Guan
Stereoscopic head-mounted displays (HMDs) render and present binocular images to create an egocentric, 3D percept to the HMD user. Within this render and presentation pipeline there are potential rendering camera and viewing position errors that can induce deviations in the depth and distance that a user perceives compared to the underlying intended geometry
How to Elicit Explainability Requirements? A Comparison of Interviews, Focus Groups, and Surveys
cs.SEMartin Obaidi, Jakob Droste, Hannah Deters, Marc Herrmann
As software systems grow increasingly complex, explainability has become a crucial non-functional requirement for transparency, user trust, and regulatory compliance. Eliciting explainability requirements is challenging, as different methods capture varying levels of detail and structure. This study examines the efficiency and effectiveness of three commonly
Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian
Transformer-based language models have demonstrated impressive capabilities across a range of complex reasoning tasks. Prior theoretical work exploring the expressive power of transformers has shown that they can efficiently perform multi-step reasoning tasks involving parallelizable computations. However, the learnability of such constructions, particularly
Differentially Private Space-Efficient Algorithms for Counting Distinct Elements in the Turnstile Model
cs.DSRachel Cummings, Alessandro Epasto, Jieming Mao, Tamalika Mukherjee
The turnstile continual release model of differential privacy captures scenarios where a privacy-preserving real-time analysis is sought for a dataset evolving through additions and deletions. In typical applications of real-time data analysis, both the length of the stream $T$ and the size of the universe $|U|$ from which data come can be extremely large. T
Bo Zhao, Nima Dehmamy, Robin Walters, Rose Yu
Neural network minima are often connected by curves along which train and test loss remain nearly constant, a phenomenon known as mode connectivity. While this property has enabled applications such as model merging and fine-tuning, its theoretical explanation remains unclear. We propose a new approach to exploring the connectedness of minima using parameter
Performance Analysis of Wireless Communication Systems Assisted by Fluid Reconfigurable Intelligent Surfaces
cs.ITFarshad Rostami Ghadi, Kai-Kit Wong, F. Javier Lopez-Martinez, George C. Alexandropoulos
This letter investigates the performance of emerging wireless communication systems assisted by a fluid reconfigurable intelligent surface (FRIS). Unlike conventional reconfigurable intelligent surfaces (RISs), an FRIS consists of fluid-inspired metamaterials arranged in a densely packed matrix of sub-elements over a surface. It dynamically activates specifi
Ángel Cuevas, Javier Chagoya, C. Ortiz
In the derivation of the Einstein field equations via Hamilton's principle, the inclusion of a boundary term is essential to render the variational problem well-posed, as it addresses variations that do not vanish at the boundary of the spacetime manifold. Typically, this term is chosen as the Gibbons-Hawking-York boundary term. In this work, we propose an a
Optical Photometric Monitoring of the Blazar OT 355 and Local Standard Stars' Calibration
astro-ph.HER. Bachev, Tushar Tripathi, Alok C. Gupta, A. Kurtenkov
OT 355 (4FGL J1734.3 + 3858) is a relatively rarely studied but highly variable, moderate-redshift (z = 0.975) flat-spectrum radio quasar (blazar). With this work, we aim to study its optical variability on different timescales, which can help us to better understand the physical processes in relativistic jets operating in blazar-type active galactic nuclei.
Piotr Bartman-Szwarc, Adil M. Bagirov, Anna Ochal
In this paper, we employ a global aggregate subgradient method for the numerical solution of hemivariational inequality problems arising in contact mechanics. The method integrates a global search procedure to identify starting points for a local minimization algorithm. The algorithm consists of two types of steps: null steps and serious steps. In each null
Moinak Bhattacharya, Judy Huang, Amna F. Sher, Gagandeep Singh
Accurately predicting immunotherapy response in Non-Small Cell Lung Cancer (NSCLC) remains a critical unmet need. Existing radiomics and deep learning-based predictive models rely primarily on pre-treatment imaging to predict categorical response outcomes, limiting their ability to capture the complex morphological and textural transformations induced by imm
Jianjun Zhao
Quantum computing has demonstrated the potential to solve computationally intensive problems more efficiently than classical methods. Many software engineering tasks, such as test case selection, static analysis, code clone detection, and defect prediction, involve complex optimization, search, or classification, making them candidates for quantum enhancemen
Aya Kayal, Sattar Vakili, Laura Toni, Da-shan Shiu
Bayesian optimization (BO) with preference-based feedback has recently garnered significant attention due to its emerging applications. We refer to this problem as Bayesian Optimization from Human Feedback (BOHF), which differs from conventional BO by learning the best actions from a reduced feedback model, where only the preference between two actions is re
Amir Said, Xin Zhao, Marta Karczewicz, Jianle Chen
Intra-frame prediction in the High Efficiency Video Coding (HEVC) standard can be empirically improved by applying sets of recursive two-dimensional filters to the predicted values. However, this approach does not allow (or complicates significantly) the parallel computation of pixel predictions. In this work we analyze why the recursive filters are effectiv
Manish Shetty, Naman Jain, Jinjian Liu, Vijay Kethanaboyina
Developing high-performance software is a complex task that requires specialized expertise. We introduce GSO, a benchmark for evaluating language models' capabilities in developing high-performance software. We develop an automated pipeline that generates and executes performance tests to analyze repository commit histories to identify 102 challenging optimi
C. W. J. Beenakker
We calculate the full counting statistics of charge transfer in a chiral Majorana interferometer - a setup where a Dirac mode (an electron-hole mode) is split into two Majorana modes that encircle a number of h/2e vortices in a topological superconductor. Without any coupling to the environment it is known that the low-energy charge transfer is deterministic
Dual-Task Graph Neural Network for Joint Seizure Onset Zone Localization and Outcome Prediction using Stereo EEG
eess.SPSyeda Abeera Amir, Artur Agaronyan, William Gaillard, Chima Oluigbo
Accurately localizing the brain regions that triggers seizures and predicting whether a patient will be seizure-free after surgery are vital for surgical planning and patient management in drug-resistant epilepsy. Stereo-electroencephalography (sEEG) delivers high-fidelity intracranial recordings that enable clinicians to precisely locate epileptogenic netwo
Sung Soo Moon, Sebastian E. Ahnert
Many real-world networks have associated metadata that assigns categorical labels to nodes. Analysis of these annotations can complement the topological analysis of complex networks. Annotated networks have typically been used to evaluate community detection approaches. Here, we introduce an approach that combines the quantitative analysis of annotations and
Formula-R1: Incentivizing LLM Reasoning over Complex Tables with Numerical Computation via Formula-Driven Reinforcement Learning
cs.AILang Cao, Jingxian Xu, Hanbing Liu, Jinyu Wang
Tables are a fundamental medium for organizing and analyzing data, making table reasoning a critical capability for intelligent systems. Although large language models (LLMs) exhibit strong general reasoning abilities, they still struggle with accurate numerical reasoning over tabular data, particularly in complex table settings beyond simple relational look
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
cs.CYNariman Naderi, Zahra Atf, Peter R Lewis, Aref Mahjoub far
This paper investigates how prompt engineering techniques impact both accuracy and confidence elicitation in Large Language Models (LLMs) applied to medical contexts. Using a stratified dataset of Persian board exam questions across multiple specialties, we evaluated five LLMs - GPT-4o, o3-mini, Llama-3.3-70b, Llama-3.1-8b, and DeepSeek-v3 - across 156 confi
Jeremy Brazas, Atish Mitra
The $\pi_n$-wild set $\mathbf{w}_{n}(X)$ of a topological space $X$ is the subspace of $X$ consisting of the points at which there exists a shrinking sequence of essential based maps $S^n\to X$. In this paper, we show that the homotopy type of $\mathbf{w}_{n}(X)$ is a homotopy invariant of $X$ and, in analogy to the known one-dimensional case, we show that f
Ahmed Almheiri
Bousso and Stanford (BS) argued that the black hole final state proposal leads to acausal effects and ill-defined probabilities for the AMPS experiment. We identify a loophole in their analysis using insights from entanglement wedge reconstruction and replica wormholes. We trace the cause of the BS problems to the misidentification of the physical interior w
Niklas Freymuth, Tobias Würth, Nicolas Schreiber, Balazs Gyenes
The cost and accuracy of simulating complex physical systems using the Finite Element Method (FEM) scales with the resolution of the underlying mesh. Adaptive meshes improve computational efficiency by refining resolution in critical regions, but typically require task-specific heuristics or cumbersome manual design by a human expert. We propose Adaptive Mes
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
cs.CLBeong-woo Kwak, Minju Kim, Dongha Lim, Hyungjoo Chae
Large language models (LLMs) have demonstrated strong capabilities in using external tools to address user inquiries. However, most existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions. To fill this gap, we introduce ToolHaystack, a benchmark for testing the tool use capa
Size Wu, Zhonghua Wu, Zerui Gong, Qingyi Tao
In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (
Ziteng Gao, Mike Zheng Shou
This paper presents Diffusion via Autoregressive models (D-AR), a new paradigm recasting the image diffusion process as a vanilla autoregressive procedure in the standard next-token-prediction fashion. We start by designing the tokenizer that converts images into sequences of discrete tokens, where tokens in different positions can be decoded into different