Skip to content

May 2025 arXiv papers — page 81

Showing 8,0018,100 of 24,552 papers

  1. Rohan Ghuge, Vidya Muthukumar, Sahil Singla

    We study \emph{online multicalibration}, a framework for ensuring calibrated predictions across multiple groups in adversarial settings, across $T$ rounds. Although online calibration is typically studied in the $\ell_1$ norm, prior approaches to online multicalibration have taken the indirect approach of obtaining rates in other norms (such as $\ell_2$ and

  2. Apar Pokhrel, Gia Dao

    Parking space occupancy detection is a critical component in the development of intelligent parking management systems. Traditional object detection approaches, such as YOLOv8, provide fast and accurate vehicle detection across parking lots but can struggle with borderline cases, such as partially visible vehicles, small vehicles (e.g., motorcycles), and poo

  3. Hassan Wasswa, Hussein Abbass, Timothy Lynar

    Due to the exponential rise in IoT-based botnet attacks, researchers have explored various advanced techniques for both dimensionality reduction and attack detection to enhance IoT security. Among these, Variational Autoencoders (VAE), Vision Transformers (ViT), and Graph Neural Networks (GNN), including Graph Convolutional Networks (GCN) and Graph Attention

  4. Zafarullah Mahmood, Soliman Ali, Jiading Zhu, Mohamed Abdelwahab

    The conversational capabilities of Large Language Models (LLMs) suggest that they may be able to perform as automated talk therapists. It is crucial to know if these systems would be effective and adhere to known standards. We present a counsellor chatbot that focuses on motivating tobacco smokers to quit smoking. It uses a state-of-the-art LLM and a widely

  5. Chi-Chun Zhou, Shuai A. Chen, Yu-Zhu Chen, Yao Shen

    Quantum matter in three spatial dimensions is observed to consist exclusively of bosons and fermions. Whether this empirical fact follows from basic consistency requirements of quantum theory itself or must be imposed as an additional principle has for 80 years remained a fundamental conceptual gap. Here we close this gap by establishing a no-go theorem that

  6. Rares-Darius Buhai, Jun-Ting Hsieh, Aayush Jain, Pravesh K. Kothari

    There is a growing body of work on proving hardness results for average-case estimation problems by bounding the low-degree advantage (LDA) - a quantitative estimate of the closeness of low-degree moments - between a null distribution and a related planted distribution. Such hardness results are now ubiquitous not only for foundational average-case problems

  7. Xianzhong Ding, Yunkai Zhang, Binbin Chen, Donghao Ying

    Modern industry-scale data centers need to manage a large number of virtual machines (VMs). Due to the continual creation and release of VMs, many small resource fragments are scattered across physical machines (PMs). To handle these fragments, data centers periodically reschedule some VMs to alternative PMs, a practice commonly referred to as VM reschedulin

  8. Chinmay Talegaonkar, Nikhil Gandudi Suresh, Zachary Novack, Yash Belhe

    Recent monocular metric depth estimation (MMDE) methods have made notable progress towards zero-shot generalization. However, they still exhibit a significant performance drop on out-of-distribution datasets. We address this limitation by injecting defocus blur cues at inference time into Marigold, a \textit{pre-trained} diffusion model for zero-shot, scale-

  9. Hassan Wasswa, Hussein Abbass, Timothy Lynar

    With the rise of IoT-based botnet attacks, researchers have explored various learning models for detection, including traditional machine learning, deep learning, and hybrid approaches. A key advancement involves deploying attention mechanisms to capture long-term dependencies among features, significantly improving detection accuracy. However, most models t

  10. Parsa Moradi, Hanzaleh Akabrinodehi, Mohammad Ali Maddah-Ali

    In this paper, we investigate the adversarial robustness of nonparametric regression, a fundamental problem in machine learning, under the setting where an adversary can arbitrarily corrupt a subset of the input data. While the robustness of parametric regression has been extensively studied, its nonparametric counterpart remains largely unexplored. We chara

  11. Wasif Husain

    In this study, the impact of neutron decay into dark matter and various dark matter self-interaction strengths on neutron star properties have been explored. Using the quark-meson coupling (QMC) model for nucleon-only equations of state (EoSs), the effects of different matter compositions have been compared, including strange matter and self-interacting dark

  12. N. Benjamin Erichson, Vinicius Mikuni, Dongwei Lyu, Yang Gao

    We introduce FLEX (FLow EXpert), a backbone architecture for generative modeling of spatio-temporal physical systems using diffusion models. FLEX operates in the residual space rather than on raw data, a modeling choice that we motivate theoretically, showing that it reduces the variance of the velocity field in the diffusion model, which helps stabilize tra

  13. Yanting Miao, William Loh, Pacal Poupart, Suraj Kothawade

    Recent work uses reinforcement learning (RL) to fine-tune text-to-image diffusion models, improving text-image alignment and sample quality. However, existing approaches introduce unnecessary complexity: they cache the full sampling trajectory, depend on differentiable reward models or large preference datasets, or require specialized guidance techniques. Mo

  14. Maximiliano Cristiá, Alfredo Capozucca, Gianfranco Rossi

    {log} (read 'setlog') was born as a Constraint Logic Programming (CLP) language where sets and binary relations are first-class citizens, thus fostering set programming. Internally, {log} is a constraint satisfiability solver implementing decision procedures for several fragments of set theory. Hence, {log} can be used as a declarative, set, logic programmin

  15. David Porlles, Wei Chen

    The momentum space of conventional superconductors is recently recognized to possess a quantum metric defined from the overlap of filled quasihole states at neighboring momenta. For multiband superconductors with arbitrary intraband and interband s-wave pairing, we elaborate that their superfluid weight in London equations is given by the momentum integratio

  16. Álvaro Cabezas-Clavijo, Pavel Sidorenko-Bautista

    This study analyzes the performance of eight generative artificial intelligence chatbots -- ChatGPT, Claude, Copilot, DeepSeek, Gemini, Grok, Le Chat, and Perplexity -- in their free versions, in the task of generating academic bibliographic references within the university context. A total of 400 references were evaluated across the five major areas of know

  17. Yizhou Xu, Florent Krzakala, Lenka Zdeborová

    The Restricted Boltzmann Machine (RBM) is one of the simplest generative neural networks capable of learning input distributions. Despite its simplicity, the analysis of its performance in learning from the training data is only well understood in cases that essentially reduce to singular value decomposition of the data. Here, we consider the limit of a larg

  18. Gizem Gultekin-Varkonyi

    Legal AI systems are increasingly being adopted by judicial and legal system deployers and providers worldwide to support a range of applications. While they offer potential benefits such as reducing bias, increasing efficiency, and improving accountability, they also pose significant risks, requiring a careful balance between opportunities, and legal and et

  19. Nuno Crokidakis

    The Siege of Syracuse (214 - 212 BC) was a decisive event in the Second Punic War, leading to the city's fall to Rome despite its formidable defenses, including the war machines devised by Archimedes. In this work, we propose a mathematical model to describe the dynamics of the siege, incorporating the depletion of resources, the decline of Syracuse'

  20. Qian Chen, Mohamed Elrefaie, Angela Dai, Faez Ahmed

    Surrogate modeling has emerged as a powerful tool to accelerate Computational Fluid Dynamics (CFD) simulations. Existing 3D geometric learning models based on point clouds, voxels, meshes, or graphs depend on explicit geometric representations that are memory-intensive and resolution-limited. For large-scale simulations with millions of nodes and cells, exis

  21. Deepak Kumar, D. Tripathi, Sunil Hans

    Let $P(z)$ be a polynomial of degree $n$. In $2004$, Aziz and Rather \cite{aziz2004some} investigated the dependence of \[\bigg|P(Rz)-αP(z)+β\biggl\{\biggl(\frac{R+1}{2}\biggr)^n-|α|\biggr\}P(z)\bigg|, \ \text{for} \ z \in B(\mathbb{D}),\] on $\max_{z\in B(\mathbb{D})}|P(z)|$, for every real and complex number $α, β$ satisfying $|α| \leq 1$, $|β| \leq 1$, an

  22. Ruizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao

    The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited

  23. Tomoro Yanase, Shin-ichiro Shima, Seiya Nishizawa, Hirofumi Tomita

    Clouds play a central role in climate physics by interacting with precipitation, radiation, and circulation. Despite being a fundamental issue in convective organization, the self-aggregation of clouds lacks a theoretical explanation due to its complexity. In this study, we introduce an idealized mathematical model where the system's state is represented

  24. Thomas Chen, Patricia Muñoz Ewald

    We analyze geometric aspects of the gradient descent algorithm in Deep Learning (DL), and give a detailed discussion of the circumstance that in underparametrized DL networks, zero loss minimization can generically not be attained. As a consequence, we conclude that the distribution of training inputs must necessarily be non-generic in order to produce zero

  25. Yuheng Wu, Jianwen Xie, Denghui Zhang, Zhaozhuo Xu

    Theory-of-Mind (ToM) tasks pose a unique challenge for large language models (LLMs), which often lack the capability for dynamic logical reasoning. In this work, we propose DEL-ToM, a framework that improves verifiable ToM reasoning through inference-time scaling rather than architectural changes. Our approach decomposes ToM tasks into a sequence of belief u

  26. Leland Russell, Ezekiel A. Rein, Anatalya Piatigorsky, Jennifer T. Heath

    In this work, the force due to radiation pressure is measured with sub-10 pN sensitivity, corresponding to less than 2 mW of optical power. The apparatus adds homemade reflectors to a commercial Cavendish balance, which consists of a torsion pendulum with a built-in capacitance position sensor. When driven by four 5 mW laser diodes, with square-wave modulati

  27. Simon C. Tait, Michael J. Williams, Joseph Bayley, Bryan W. Barr

    Gravitational wave detectors, such as LIGO, are predominantly limited by coating Brownian thermal noise (CTN), arising from mechanical losses in the Bragg mirror coatings used on test-mass optics. Accurately characterizing and minimizing these losses is crucial for enhancing detector sensitivity. This paper introduces a general mathematical and statistical f

  28. Justin D. Norman, Michael U. Rivera, D. Alex Hughes

    Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern, there is little scientific work that attempts to measure the prevalence of language model hallucination in a comprehensive way. In this paper, we argue that language models should

  29. Ninda Nurseha Amalina, Heungjo An

    Unattended scheduled appointments, defined as patient no-shows, adversely affect both healthcare providers and patients' health, disrupting the continuity of care, operational efficiency, and the efficient allocation of medical resources. Accurate predictive modeling is needed to reduce the impact of no-shows. Although machine learning methods, such as logis

  30. Zhewei Yao, Guoheng Sun, Lukasz Borchmann, Gaurav Nuti

    Translating natural language into SQL (Test2SQL) is a longstanding challenge at the intersection of natural language understanding and structured data access. While large language models (LLMs) have significantly improved fluency in SQL generation, producing correct and executable SQL--particularly for complex queries--remains a bottleneck. We present Arctic

  31. Dillon Lohr, Michael J. Proulx, Mehedi Hasan Raju, Oleg V. Komogortsev

    This paper investigates the feasibility of fusing two eye-centric authentication modalities-eye movements and periocular images-within a calibration-free authentication system. While each modality has independently shown promise for user authentication, their combination within a unified gaze-estimation pipeline has not been thoroughly explored at scale. In

  32. Chaoyi Jiang, Sungwoo Kim, Lei Gao, Hossein Entezari Zarch

    Masked autoregressive (MAR) models unify the strengths of masked and autoregressive generation by predicting tokens in a fixed order using bidirectional attention for image generation. While effective, MAR models suffer from significant computational overhead, as they recompute attention and feed-forward representations for all tokens at every decoding step,

  33. Ankita Kushwaha, Kiran Ravish, Preeti Lamba, Pawan Kumar

    Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and deployment of agents. This survey provides a mathematically rigorous overview of SafeRL formulations based on Constrained Markov Decision Processes (CMDPs) and extensions to Multi-Agent Safe RL (SafeMARL). We re

  34. Dibyajyoti Nayak, Somdatta Goswami

    Accurate temporal extrapolation remains a fundamental challenge for neural operators modeling dynamical systems, where predictions must extend far beyond the training horizon. Conventional DeepONet approaches rely on two limited paradigms: fixed-horizon rollouts, which predict full spatiotemporal solutions while ignoring temporal causality, and autoregressiv

  35. Tinghan Ye, Amira Hijazi, Pascal Van Hentenryck

    Accurate estimation of order fulfillment time is critical for e-commerce logistics, yet traditional rule-based approaches often fail to capture the inherent uncertainties in delivery operations. This paper introduces a novel framework for distributional forecasting of order fulfillment time, leveraging Conformal Predictive Systems and Cross Venn-Abers Predic

  36. Philip G. Judge

    This study attempts to establish a basis for understanding how methods used in research in solar physics have evolved since World War II (WWII). The goal is to begin to explore if and how the changing research environment affects the training of young scientists, and the future of solar physics research at our institutions. A strategy based upon a sample of

  37. Muhammad Umar Farooq, Daniel Kaiser

    Large message transmissions in libp2p GossipSub lead to longer than expected network-wide message dissemination times and very high bandwidth utilization. This article identifies key issues responsible for this behavior and proposes modifications to the protocol for transmitting large messages. These modifications preserve the GossipSub resilience and fit we

  38. I. Papuccio-Fernández, A. A. Reynoso, A. E. Bruchhausen, A. S. Kuznetsov

    Phonon lasers, as their photon counterparts, rely on the physics of stimulated emission. Arguably, because light does not require a material substrate to propagate, while sound does, the impact of the two technologies has however been highly contrasting, with "sasers" (for sound amplification by stimulated emission of radiation) mostly remaining as an academ

  39. Tahina Ramananandro, Gabriel Ebner, Guido Martínez, Nikhil Swamy

    Incorrect handling of security-critical data formats, particularly in low-level languages, are the root cause of many security vulnerabilities. Provably correct parsing and serialization tools that target languages like C can help. Towards this end, we present PulseParse, a library of verified parser and serializer combinators for non-malleable binary format

  40. Mihail Cocos

    We establish that any affine manifold $(M,\nabla)$ endowed with a parallel volume form $\omega,$ admits, in any conformal class of Riemannian metrics, a representative $H$ for which $\nabla$ is the Levi-Civita connection. This provides a constructive proof that such manifolds are necessarily complete, generalizing the "if" direction of Markus' conjecture \ci

  41. Stefan van der Jagt, Erik Osinga, Reinout J. van Weeren, George K. Miley

    The radio jets of radio galaxies in galaxy clusters are often bent due to the ram pressure of the intracluster medium. In this paper we start with a well-defined sample of galaxy clusters and subsequently identifying tailed radio sources in these known environments. Our sample consists of 81 galaxy clusters from the Planck ESZ cluster sample. We present a ca

  42. Xin You, Minghui Zhang, Hanxiao Zhang, Jie Yang

    Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist simultaneously. However, in the preoperative data acquisition stage, the slight movement of patients may result in dynamic backgr

  43. Hitesh Laxmichand Patel, Amit Agarwal, Arion Das, Bhargava Kumar

    Enterprise customers are increasingly adopting Large Language Models (LLMs) for critical communication tasks, such as drafting emails, crafting sales pitches, and composing casual messages. Deploying such models across different regions requires them to understand diverse cultural and linguistic contexts and generate safe and respectful responses. For enterp

  44. Maryam Dialameh, Rezaul Karim, Hossein Rajabzadeh, Omar Mohamed Awad

    This paper introduces ECHO-LLaMA, an efficient LLaMA architecture designed to improve both the training speed and inference throughput of LLaMA architectures while maintaining its learning capacity. ECHO-LLaMA transforms LLaMA models into shared KV caching across certain layers, significantly reducing KV computational complexity while maintaining or improvin

  45. Amit Agarwal, Srikant Panda, Kulbhushan Pachauri

    In this work, we propose Few Shot Domain Adapting Graph (FS-DAG), a scalable and efficient model architecture for visually rich document understanding (VRDU) in few-shot settings. FS-DAG leverages domain-specific and language/vision specific backbones within a modular framework to adapt to diverse document types with minimal data. The model is robust to prac

  46. Dylan Kline

    This study bridges cognitive science and neural network design by examining whether artificial models exhibit human-like forgetting curves. Drawing upon Ebbinghaus' seminal work on memory decay and principles of spaced repetition, we propose a quantitative framework to measure information retention in neural networks. Our approach computes the recall probabi

  47. Hossein Adeli, Sun Minni, Nikolaus Kriegeskorte

    A major goal of neuroscience is to understand brain computations during visual processing in naturalistic settings. A dominant approach is to use image-computable deep neural networks trained with different task objectives as a basis for linear encoding models. However, in addition to requiring estimation of a large number of linear encoding parameters, this

  48. Dylan Kline

    Foundational game-image encoders often overfit to game-specific visual styles, undermining performance on downstream tasks when applied to new games. We present a method that combines contrastive learning and domain-adversarial training to learn game-invariant visual features. By simultaneously encouraging similar content to cluster and discouraging game-spe

  49. Soren DeHaan, Yuanze Liu, Johan Bollen, Sa'ul A. Blanco

    The proliferation of Large Language Models (LLMs) in late 2022 has impacted academic writing, threatening credibility, and causing institutional uncertainty. We seek to determine the degree to which LLMs are used to generate critical text as opposed to being used for editing, such as checking for grammar errors or inappropriate phrasing. In our study, we ana

  50. Zackary Rackauckas, Julia Hirschberg

    We introduce VoxRAG, a modular speech-to-speech retrieval-augmented generation system that bypasses transcription to retrieve semantically relevant audio segments directly from spoken queries. VoxRAG employs silence-aware segmentation, speaker diarization, CLAP audio embeddings, and FAISS retrieval using L2-normalized cosine similarity. We construct a 50-que

  51. Daniel Král', Filip Kučerák, Ander Lamaison, Gábor Tardos

    In the 1980s, Erdős and Sós initiated the study of Turán hypergraph problems with a uniformity condition on the distribution of edges, i.e., determining density thresholds for the existence of a hypergraph H in a host hypergraph with edges uniformly distributed. In particular, Erdős and Sós asked to determine the uniform Turán densities of the hypergraphs $K

  52. Weichen Tang, Zhenglu Li, Cheng Chen, Yu He

    We present a first-principles study based on density functional theory (DFT) on the electronic and structural properties of Ta2NiSe5, a layered transition metal chalcogenide that has been considered as a possible candidate for an excitonic insulator. Our systematic DFT results however provide a non-excitonic mechanism for the experimentally observed electron

  53. Kaveen Hiniduma, Dylan Ryan, Suren Byna, Jean Luca Bez

    AI Data Readiness Inspector (AIDRIN) is a framework to evaluate and improve data preparedness for AI applications. It addresses critical data readiness dimensions such as data quality, bias, fairness, and privacy. This paper details enhancements to AIDRIN by focusing on user interface improvements and integration with a privacy-preserving federated learning

  54. Ruaridh Mon-Williams, Max Taylor-Davies, Elizabeth Mieczkowski, Natalia Velez

    Humans are remarkably adept at collaboration, able to infer the strengths and weaknesses of new partners in order to work successfully towards shared goals. To build AI systems with this capability, we must first understand its building blocks: does such flexibility require explicit, dedicated mechanisms for modelling others -- or can it emerge spontaneously

  55. Jiachen Jiang, Yuxin Dong, Jinxin Zhou, Zhihui Zhu

    In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks without weight updates by learning from demonstration sequences. While ICL shows strong empirical performance, its internal representational mechanisms are not yet well understood. In this work, we conduct a statistical geometric analysis of ICL representations to investigat

  56. Tiago Fonseca, Clarisse Sousa, Ricardo Venâncio, Pedro Pires

    The electrification of transportation and the increased adoption of decentralized renewable energy generation have added complexity to managing Renewable Energy Communities (RECs). Integrating Electric Vehicle (EV) charging with building energy systems like heating, ventilation, air conditioning (HVAC), photovoltaic (PV) generation, and battery storage prese

  57. Zackary Rackauckas, Julia Hirschberg

    Synthesizing expressive Japanese character speech poses unique challenges due to pitch-accent sensitivity and stylistic variability. This paper empirically evaluates two open-source text-to-speech models--VITS and Style-BERT-VITS2 JP Extra (SBV2JE)--on in-domain, character-driven Japanese speech. Using three character-specific datasets, we evaluate models ac

  58. Shahriar Aslani

    We prove that a Ma\~n\'e generic real-analytic $D$-Hamiltonian H, subjected to a totally non-holonomic real-analytic distribution $D$, has no non-trivial normal $D$-singular orbits of minimal rank. If $D$ has co-rank 1, this implies that $H + u$, where $u$ is a generic real-analytic potential, does not admit non-trivial normal $D$-singular orbits.

  59. Giovanni Ferrami, Nathan J. Adams, Lewi Westcott, Thomas Harvey

    We present four galaxy scale lenses discovered in two JWST blank-fields: the ~ 54 arcmin^2 of the PEARLS North-Ecliptic-Pole Time-Domain Field (NEP TDF) and in the ~ 90 arcmin^2 of CEERS. We perform the search by visual inspection of NIRCam photometric data, obtaining an initial list of 16 lens candidates. We down-select this list to 5 high-confidence lens c

  60. Alyson East, Elizabeth G. Campolongo, Luke Meyers, S M Rayeed

    1) Biological collections house millions of specimens with digital images increasingly available through open-access platforms. However, most imaging protocols were developed for human interpretation without considering automated analysis requirements. As computer vision applications revolutionize taxonomic identification and trait extraction, a critical gap

  61. Jiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning

    Achieving better alignment between vision embeddings and Large Language Models (LLMs) is crucial for enhancing the abilities of Multimodal LLMs (MLLMs), particularly for recent models that rely on powerful pretrained vision encoders and LLMs. A common approach to connect the pretrained vision encoder and LLM is through a projector applied after the vision en

  62. Nick Fischer, Marvin Künnemann, Mirza Redžić, Julian Stieß

    Is detecting a $k$-clique in $k$-partite regular (hyper-)graphs as hard as in the general case? Intuition suggests yes, but proving this -- especially for hypergraphs -- poses notable challenges. Concretely, we consider a strong notion of regularity in $h$-uniform hypergraphs, where we essentially require that any subset of at most $h-1$ is incident to a uni

  63. Soham Dutta, Arnab Saha

    Optical tweezers can confine position as well as orientation of a Brownian particle by simultaneously exerting restoring force and torque on it. Here we have proposed the theoretical model of a microscopic Stirling engine, using a passive Brownian ellipsoid as its working substance. The position and the orientation degrees of freedom (DoF) of the ellipsoid i

  64. Xiangqi Wang, Yue Huang, Yanbo Wang, Xiaonan Luo

    LLMs often need effective configurations, like temperature and reasoning steps, to handle tasks requiring sophisticated reasoning and problem-solving, ranging from joke generation to mathematical reasoning. Existing prompting approaches usually adopt general-purpose, fixed configurations that work 'well enough' across tasks but seldom achieve task-specific o

  65. Barbara Puccio, Federico Castagna, Allan Tucker, Pierangelo Veltri

    Despite their staggering capabilities as assistant tools, often exceeding human performances, Large Language Models (LLMs) are still prone to jailbreak attempts from malevolent users. Although red teaming practices have already identified and helped to address several such jailbreak techniques, one particular sturdy approach involving role-playing (which we

  66. Harim Kim, Yuhan Wang, Minkyu Ahn, Heeyoul Choi

    Unsupervised anomaly detection (UAD) in medical imaging is crucial for identifying pathological abnormalities without requiring extensive labeled data. However, existing diffusion-based UAD models rely solely on imaging features, limiting their ability to distinguish between normal anatomical variations and pathological anomalies. To address this, we propose

  67. Blessing Airehenbuwa, Touseef Hasan, Souvika Sarkar, Ujjwal Guin

    The proliferation of electronic devices has greatly transformed every aspect of human life, such as communication, healthcare, transportation, and energy. Unfortunately, the global electronics supply chain is vulnerable to various attacks, including piracy of intellectual properties, tampering, counterfeiting, information leakage, side-channel, and fault inj

  68. Nam H. Le, Josh Bongard

    Genetic Programming (GP) has traditionally entangled the evolution of symbolic representations with their performance-based evaluation, often relying solely on raw fitness scores. This tight coupling makes GP solutions more fragile and prone to overfitting, reducing their ability to generalize. In this work, we propose LaSER (Latent Semantic Representation R

  69. Philipp Pilar, Markus Heinonen, Niklas Wahlström

    Physics-informed neural networks (PINNs) have proven an effective tool for solving differential equations, in particular when considering non-standard or ill-posed settings. When inferring solutions and parameters of the differential equation from data, uncertainty estimates are preferable to point estimates, as they give an idea about the accuracy of the so

  70. Pu Yang, J. A. Barria

    This paper presents a Wavelet Probabilistic Recurrent Convolutional Network (WPRCN) for Multivariate Time Series Classification (MTSC), especially effective in handling non-stationary environments, data scarcity and noise perturbations. We introduce a versatile wavelet probabilistic module designed to extract and analyse the probabilistic features, which can

  71. Xinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze

    Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important fo

  72. Anna Ivagnes, Giovanni Stabile, Gianluigi Rozza

    In this paper, we propose an equation-based parametric Reduced Order Model (ROM), whose accuracy is improved with data-driven terms added into the reduced equations. These additions have the aim of reintroducing contributions that in standard reduced-order approaches are not taken into account. In particular, in this work we focus on a Proper Orthogonal Deco

  73. Jianhao Ma, Geyu Liang, Salar Fattahi

    Implicit regularization refers to the tendency of local search algorithms to converge to low-dimensional solutions, even when such structures are not explicitly enforced. Despite its ubiquity, the mechanism underlying this behavior remains poorly understood, particularly in over-parameterized settings. We analyze gradient descent dynamics and identify three

  74. Sousannah Abdalla, Sabur Baidya

    Gesture recognition presents a promising avenue for interfacing with unmanned aerial vehicles (UAVs) due to its intuitive nature and potential for precise interaction. This research conducts a comprehensive comparative analysis of vision-based hand gesture detection methodologies tailored for UAV Control. The existing gesture recognition approaches involving

  75. Marcio Gameiro, Brittany Gelb, Konstantin Mischaikow

    The identification of dynamics from time series data is a problem of general interest. It is well established that dynamics on the level of invariant sets, the primary objects of interest in the classical theory of dynamical systems, is not computable. We recall a coarser characterization of dynamics based on order theory and algebraic topology and prove tha

  76. D. A. Beleño-Molina, L. Olguín, L. F. Miranda, M. E. Contreras

    We present a spectroscopic investigation of 25 objects previously reported as possible Planetary Nebulae (PNe) in recent catalogs to obtain their physical properties and to establish their true nature. We found 11 objects showing intense emission lines, 11 where it was not possible to measure $\mathrm{H{\beta}}$, and three where no lines are present. We have

  77. Selina Carter, Arun K Kuchibhotla

    The construction of confidence intervals and hypothesis tests for functionals is a cornerstone of statistical inference. Traditionally, the most efficient procedures - such as the Wald interval or the Likelihood Ratio Test - require both a point estimator and a consistent estimate of its asymptotic variance. However, when estimators are derived from online o

  78. Harrison Tuckman, Eric Neuscamman

    Modeling charge transfer well can require treating post-excitation orbital relaxations and handling medium to large molecules in realistic environments. By combining a state-specific correlation treatment with such orbital relaxations, Aufbau suppressed coupled cluster has proven accurate for charge transfer, but, like many coupled cluster methods, it strugg

  79. William M Feldman, Zhonggan Huang

    We homogenize the Laplace and heat equations with the Neumann data oscillating in the ``vertical" $u$-variable. These are simplified models for interface motion in heterogeneous media, particularly capillary contact lines. The homogenization limit reveals a pinning effect at zero tangential slope, leading to a novel singularly anisotropic pinned Neumann cond

  80. Mayesha Tasnim, Erman Acar, Sennay Ghebreab

    The design of fair and efficient algorithms for allocating public resources, such as school admissions, housing, or medical residency, has a profound social impact. In one-sided matching problems, where individuals are assigned to items based on ranked preferences, a fundamental trade-off exists between efficiency and strategyproofness. Existing algorithms l

  81. Geraldo Botelho, Ariel Monção

    We give a necessary condition and a sufficient condition on the Banach lattices E and F so that an operator from E to F is DW-compact whenever its adjoint is DW-compact. We do the same, with different conditions, for DW-DP operators. Moreover, we characterize the Banach lattices E and F for which the adjoint of every DW-compact operator from E to F is DW-com

  82. Phat Thanh Dang, Saahil Thoppay, Wang Yang, Qifan Wang

    Large language models suffer issues when operated on long contexts that are larger than their training context length due to the standard position encoding for tokens in the attention layer. Tokens a long distance apart will rarely have an effect on each other and long prompts yield unexpected results. To solve this problem, we propose SELF (Self-Extend the

  83. Maryam Heydari, Hanieh Moghaddasi, Mir Vahid Hosseini, Mehdi Askari

    We theoretically investigate current-induced spin polarization in disordered topological insulator thin films with broken inversion symmetry under an applied in-plane electric field. Utilizing the Kubo formalism within the self-consistent Born approximation and incorporating vertex corrections to account for multiple scattering events, we analyze how disorde

  84. Ting-Wei Li, Ruizhong Qiu, Hanghang Tong

    Graph domain adaptation (GDA) is a fundamental task in graph machine learning, with techniques like shift-robust graph neural networks (GNNs) and specialized training procedures to tackle the distribution shift problem. Although these model-centric approaches show promising results, they often struggle with severe shifts and constrained computational resourc

  85. Dario A. Leon, Mikael Kuisma, Mikkel Ohm Sauer, Jakob K. Svaneborg

    We introduce a computationally efficient method to calculate the quasiparticle (QP) band structure of general van der Waals (vdW) heterostructures. A layer-projected scissors (LAPS) operator, which depends on the one-body density matrix, is added to the density functional theory (DFT) Hamiltonian. The LAPS operator corrects the band edges of the individual l

  86. Petr Kourzanov, Anmol

    In order to truly benefit from RISC-V ISA modularity, the community has to address the issue of compositionality, going beyond modules at the specification level covering larger subsets of the RISC-V development flow including emulation, simulation and verification. In this paper we introduce modular SAIL, an experiment to inject compositionality into the SA

  87. Linus Bleistein, Aurélien Bellet, Julie Josse

    We consider the problem of solving the optimal transport problem between two empirical distributions with missing values. Our main assumption is that the data is missing completely at random (MCAR), but we allow for heterogeneous missingness probabilities across features and across the two distributions. As a first contribution, we show that the Wasserstein

  88. Krti Tallam

    We present a robust neural watermarking framework for scientific data integrity, targeting high-dimensional fields common in climate modeling and fluid simulations. Using a convolutional autoencoder, binary messages are invisibly embedded into structured data such as temperature, vorticity, and geopotential. Our method ensures watermark persistence under los

  89. Samuel L. Foley, Margaret E. Johnson

    Cellular decision-making based on information received from the external environment is frequently initiated by transmembrane receptors. These receptors are known to propagate such information by triggering a series of irreversible, energy-consuming reactions. While this active mechanism ensures switch-like responses, here we show how spontaneous self-assemb

  90. Diego Alexander Castro Guevara

    In this paper we study the problem \[ \begin{cases} -\Delta_d u = \mu_0 &\text{ in } G\\ u = 0 &\text{ on } \partial G \end{cases} \] where, $\Delta_d$ represent the discret Laplacian, and $\mu_0$ it is a measure defined in the vertex of the graph $G=(V,E)$. Here $V$ defined the vertex of the graph, $E$ its edges and $\partial G$ its boundary. We prove that

  91. Seamus Somerstep, Vinod Raman, Unique Subedi, Yuekai Sun

    Using the bit string generation problem as a case study, we theoretically compare two standard methods for adapting large language models to new tasks. The first, referred to as supervised fine-tuning, involves training a new next token predictor on good generations. The second method, Best-of-N, trains a reward model to select good responses from a collecti

  92. Sushrut Kumar, Joshua Romero, Jung-Hee Seo, Massimiliano Fatica

    Immersed boundary methods (IBMs) facilitate the simulation of flows around stationary, moving, and deforming bodies on Cartesian grids. However, extending these simulations to the large grid sizes required for realistic flow problems remains a significant computational challenge. In this work, we present the implementation and acceleration of ViCar3D, a shar

  93. D. Kaledin

    This is essentially an illustration for the general technology of homotopical enhancements developed recently in arxiv:2409.17489. We take the derived category of an abelian category, and we look at the full subcategory spanned by complexes of length 2. This has a natural refinement to a 2-category that we call "the 2-category of extensions". However, just u

  94. Nilanjana Laha, Nilson Chapagain, Victoria Cicherski, Aaron Sonabend-W

    Patients with chronic diseases often receive treatments at multiple time points, or stages. Our goal is to learn the optimal dynamic treatment regime (DTR) from longitudinal patient data. When both the number of stages and the number of treatment levels per stage are arbitrary, estimating the optimal DTR reduces to a sequential, weighted, multiclass classifi

  95. Steffen Schotthöfer, Cory Hauck

    Two benchmark problems for linear radiation transport and derived from the literature are presented in detail and several quantities of interest are defined. High-resolution simulations are computed using standard, robust numerical methods and implemented using HPC resources. The goal of these simulations is to provide reference solutions for new discretizat

  96. Prateek Jaiswal, Esmaeil Keyvanshokooh, Junyu Cao

    Randomized clinical trials often require large patient cohorts before drawing definitive conclusions, yet abundant observational data from parallel studies remains underutilized due to confounding and hidden biases. To bridge this gap, we propose Deconfounded Warm-Start Thompson Sampling (DWTS), a practical approach that leverages a Doubly Debiased LASSO (DD

  97. Diyuan Wu, Aleksandr Shevchenko, Samet Oymak, Marco Mondelli

    Token embeddings play a crucial role in language modeling but, despite this practical relevance, their theoretical understanding remains limited. Our paper addresses the gap by characterizing the structure of embeddings obtained via gradient descent. Specifically, we consider a one-layer softmax attention model with a linear head for binary classification, i

  98. Peilin Wu, Mian Zhang, Xinlu Zhang, Xinya Du

    Agentic Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) by enabling dynamic, multi-step reasoning and information retrieval. However, these systems often exhibit sub-optimal search behaviors like over-search (retrieving redundant information) and under-search (failing to retrieve necessary information), which hinder efficien

  99. Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert

    The biases exhibited by text-to-image (TTI) models are often treated as independent, though in reality, they may be deeply interrelated. Addressing bias along one dimension - such as ethnicity or age - can inadvertently affect another, like gender, either mitigating or exacerbating existing disparities. Understanding these interdependencies is crucial for de

  100. Marc Schouler, Anca Belme, Paola Cinnella

    Aerodynamic shape optimization in industry still faces challenges related to robustness and scalability. This aspect becomes crucial for advanced optimizations that rely on expensive high-fidelity flow solvers, where computational budget constraints only allow a very limited number of simulations within the optimization loop. To address these challenges, we