Skip to content

October 2025 arXiv papers — page 75

Showing 7,4017,500 of 25,213 papers

  1. Alfred Li, Ankit Mahajan, Sandeep Sharma

    Auxiliary-field quantum Monte Carlo (AFQMC) is typically formulated as an open-ended random walk in an overcomplete space of Slater determinants, implemented through a Langevin equation. However, the explicit form of the underlying Fokker-Planck equation governing the walker population distribution has remained unknown. In this paper, we derive the Fokker-Pl

  2. Thomas Cornish, David Alonso, Boris Leistedt, Kevin Wolz

    Recent work has developed a formalism for computing angular power spectra directly from catalogues containing field values at discrete positions on the sky, thereby circumventing the need to create pixelised maps of the fields, as well as avoiding aliasing and finite-resolution effects. We adapt this formalism to incorporate template deprojection for mitigat

  3. Yoo Jung Kim, Michael P. Fitzgerald, Sébastien Vievard, Jonathan Lin

    Resolving fine details of astronomical objects provides critical insights into their underlying physical processes. This drives in part the desire to construct ever-larger telescopes and interferometer arrays and to observe at shorter wavelength to lower the diffraction limit of angular resolution. Alternatively, one can aim to overcome the diffraction limit

  4. Amir Siraj, Christopher F. Chyba, Scott Tremaine

    The orbits of small bodies in the outer solar system are particularly sensitive to gravitational perturbations, including stellar flybys. Stellar clusters, with low velocity dispersions and high number densities, can be the source of strong and frequent flybys. As a result, we can infer what properties of the solar birth environment would be incompatible wit

  5. Luis F. Alday, Elisabetta Armanini, Kelian Häring, Alexander Zhiboedov

    We study scattering on the Coulomb branch of planar ${\mathcal{N}}=4$ SYM at finite 't Hooft coupling. This setup defines a family of classical open-string S-matrices that smoothly interpolates between perturbative parton scattering at weak coupling and flat-space string scattering at strong coupling. We focus on the four-point amplitude, which exhibits a re

  6. Roy J. Zhao, Mark R. Morris, Matthew J. Hankins, Angela S. Cotera

    We present an analysis of high-resolution mid-infrared observations at 25 and 37 $μm$ of the Sagittarius C Complex (Sgr C) in the Central Molecular Zone (CMZ), based on data from the SOFIA/FORCAST Galactic Center Legacy Survey. Enabled by the high bright-source limit of the FORCAST instrument, we perform a map-level dust temperature and optical depth analysi

  7. Cesar A. Gallegos, Rafael M. Magaldi, Andrew Millis, Steven R. White

    We investigate the weak interaction integer quantum Hall (IQH) phase, the intermediate interaction phase identified as a chiral spin liquid (CSL) and the transition between them in the triangular lattice Hofstadter-Hubbard model at a density of one electron per site in an orbital magnetic field corresponding to one-quarter flux per plaquette. Our primary too

  8. Antoine Petitjean, Anja Butter, Kevin Greif, Sofia Palacios Schweitzer

    Unfolding, for example of distortions imparted by detectors, provides suitable and publishable representations of LHC data. Many methods for unbinned and high-dimensional unfolding using machine learning have been proposed, but no generative method scales to the several hundred dimensions necessary to fully characterize LHC collisions. This paper proposes a

  9. Jonathan Mercedes-Feliz, Daniel Anglés-Alcázar, Boon Kiat Oh, Rachel K. Cochrane

    Central starbursts and Active Galactic Nuclei (AGN) are thought to be fueled by either galaxy interactions or secular processes in gravitationally unstable discs. We employ cosmological hydrodynamic simulations from the Feedback in Realistic Environments (FIRE) project to propose a new nuclear fueling scenario based on the transition that galaxies undergo fr

  10. Aleksandra Calovic, Katerina S. Klos, Robert B. Hudson, James E. Dale

    The removal of gas left over from star formation has long been thought to dominate the dynamical evolution, and dissolution of star-forming regions. Feedback from massive stars from their stellar winds, photoionising radiation and supernovae is postulated to expel significant amounts of gas, altering the gravitational potential energy of the star-forming reg

  11. Matthew Blakeney, Luke Corcoran, Marius de Leeuw, Balazs Pozsgay

    We show that every fusion category containing a non-invertible, self-dual object $a$ gives rise to an integrable anyonic chain whose Hamiltonian density satisfies the Temperley-Lieb algebra. This spin chain arises by considering the projection onto the identity channel in the fusion process $a\otimes a$. We relate these models to Pasquier's construction of A

  12. Jeffrey V. Backus

    The recently-developed "scalar-scaffolding" formulation of gluon amplitudes casts the Yang-Mills (YM) amplitude as a well-defined Laurent series expansion in scalar variables, valid for any spacetime dimension and helicity configuration. In this letter, we exploit this new perspective to develop conceptually novel methods of computing YM tree amplitudes. Fir

  13. Soubhik Kumar, Michael Nee

    Extra dimensions are present in many beyond the Standard Model scenarios, most notably in string theory. However, direct signatures of extra dimensions are difficult to observe in many cases. This is the situation, for example, if the energy scales associated with extra dimensions are close to the string or Grand Unification scale. The energetic early univer

  14. Juan A. Carretero, Philippe Grandclément, Carlos Palenzuela, Marcelo Salgado

    Rotating hairy black holes (RHBHs) are axisymmetric equilibrium solutions of the Einstein--Klein--Gordon equations, consisting of a spinning black hole surrounded by a toroidal distribution of complex scalar field. Despite their potential astrophysical relevance, the stability of these configurations -- naturally expected to form through superradiant growth

  15. Luke Staszewski, Asmi Haldar, Pieter W. Claeys, Alexander Wietek

    In isolated quantum many-body systems periodically driven in time, the asymptotic dynamics at late times can exhibit distinct behavior such as thermalization or dynamical freezing. Understanding the properties of and the convergence towards infinite-time (nonequilibrium) steady states however remains a challenging endeavor. We propose a physically motivated

  16. Simone Cavazzoni, Giovanni Ragazzi, Paolo Bordone, Matteo G. A. Paris

    The maximum work that can be extracted from a quantum battery is bounded by the ergotropy of the system, which is determined by the spectral properties of the Hamiltonian. In this paper, we employ the formalism of quantum walks to investigate how the topology of the battery and the chirality of the Hamiltonian influence its performance as an energy storage u

  17. Parth Nayak, Michael Walther, Daniel Gruen

    Deep learning (DL) has been shown to outperform traditional, human-defined summary statistics of the Ly{\alpha} forest in constraining key astrophysical and cosmological parameters owing to its ability to tap into the realm of non-Gaussian information. An understanding of the impact of nuisance effects such as noise on such field-level frameworks, however, s

  18. Atharv Sonwane, Isadora White, Hyunji Lee, Matheus Pereira

    High quality bugs are key to training the next generation of language model based software engineering (SWE) agents. We introduce a novel method for synthetic generation of difficult and diverse bugs. Our method instructs SWE Agents to introduce a feature into the codebase whereby they may unintentionally break tests, resulting in bugs. Prior approaches ofte

  19. Aritra Ghosh, Nilamoni Daloi, M. Bhattacharya

    We theoretically propose a quantum heat engine using a setup consisting of a ring-trapped Bose-Einstein condensate placed in a Fabry-P\'erot cavity where the optical field carries orbital angular momentum. We first show that the cavity-enhanced light-atom coupling leads to the emergence of polaritonic modes whose character can be reversibly switched between

  20. Jackson Hassell, Dan Zhang, Hannah Kim, Tom Mitchell

    We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without parameter updates. While conventional approaches like fine-tuning are often costly, inflexible, and opaque, we propose a memory-augmented framework that leverages LLM-generated critiques grounded in labeled data.

  21. Dominik Kempa, Tomasz Kociumaka

    In this work, we study the limits of compressed data structures, i.e., structures that support various queries on an input text $T\in\Sigma^n$ using space proportional to the size of $T$ in compressed form. Nearly all fundamental queries can currently be efficiently supported in $O(\delta(T)\log^{O(1)}n)$ space, where $\delta(T)$ is the substring complexity,

  22. Ilona Demler, Saumya Chauhan, Georgia Gkioxari

    We introduce ITTO, a challenging new benchmark suite for evaluating and diagnosing the capabilities and limitations of point tracking methods. Our videos are sourced from existing datasets and egocentric real-world recordings, with high-quality human annotations collected through a multi-stage pipeline. ITTO captures the motion complexity, occlusion patterns

  23. Jacob Berg, Chuning Zhu, Yanda Bao, Ishan Durugkar

    Planning with world models offers a powerful paradigm for robotic control. Conventional approaches train a model to predict future frames conditioned on current frames and actions, which can then be used for planning. However, the objective of predicting future pixels is often at odds with the actual planning objective; strong pixel reconstruction does not a

  24. Jake Poznanski, Luca Soldaini, Kyle Lo

    We present olmOCR 2, the latest in our family of powerful OCR systems for converting digitized print documents, like PDFs, into clean, naturally ordered plain text. olmOCR 2 is powered by olmOCR-2-7B-1025, a specialized, 7B vision language model (VLM) trained using reinforcement learning with verifiable rewards (RLVR), where our rewards are a diverse set of

  25. Veronica Giardini, Luca Guariento, Andrea Fantini, Shawn Storm

    We report on the realization of a platform for trapping and manipulating individual $^{88}$Sr atoms in optical tweezers. A first cooling stage based on a blue shielded magneto-optical trap (MOT) operating on the $^1S_0$ -> $^1P_1$ transition at 461 nm enables us to trap approximately $4\times 10^6$ atoms at a temperature of 6.8 mK. Further cooling is achieve

  26. Dominik Kempa, Tomasz Kociumaka

    We study the fundamental question of how efficiently suffix array entries can be accessed when the array cannot be stored explicitly. The suffix array $SA_T[1..n]$ of a text $T$ of length $n$ encodes the lexicographic order of its suffixes and underlies numerous applications in pattern matching, data compression, and bioinformatics. Previous work established

  27. Siyang Wu, Jack Nugent, Willow Yang, Jia Deng

    Monocular depth estimation is an important task with rapid progress, but how to evaluate it is not fully resolved, as evidenced by a lack of standardization in existing literature and a large selection of evaluation metrics whose trade-offs and behaviors are not fully understood. This paper contributes a novel, quantitative analysis of existing metrics in te

  28. Haoming Ning, Brian Nugent

    We extend the notions of higher Du Bois and higher rational singularities to pairs in the sense of the minimal model program. We extend numerous results to these higher pairs, including Bertini type theorems, stability under finite maps and that m-rational pairs are m-Du Bois. We prove these using a generalized Kov\'acs-Schwede-type injectivity theorem for p

  29. Filipe Ferreira de Oliveira, Matheus Becali Rocha, Renato A. Krohling

    In this paper, we propose an approach to support the diagnosis of urinary tract diseases, with a focus on bladder cancer, using SHAP (SHapley Additive exPlanations)-based feature selection to enhance the transparency and effectiveness of predictive models. Six binary classification scenarios were developed to distinguish bladder cancer from other urological

  30. Johnny Tian-Zheng Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Wang

    We present Hubble, a suite of fully open-source large language models (LLMs) for the scientific study of LLM memorization. Hubble models come in standard and perturbed variants: standard models are pretrained on a large English corpus, and perturbed models are trained in the same way but with controlled insertion of text (e.g., book passages, biographies, an

  31. Lukas Kueß, Ernst Paunzen

    The pre-main-sequence evolution of the chemically peculiar (CP) stars on the upper main sequence is still a vast mystery and not well understood. Our analysis of young associations and open clusters aims to find (very) young CP stars to try to put a lower boundary on the age of such objects. Using three catalogues of open clusters and associations, we determ

  32. Gyubeum Lim, Yemo Koo, Vijay Krishna Madisetti

    Understanding long-context visual information remains a fundamental challenge for vision-language models, particularly in agentic tasks such as GUI control and web navigation. While web pages and GUI environments are inherently structured documents, current VLMs typically neglect decision-oriented document understanding in their training objectives. Existing

  33. Virgile Guémard

    In this work, we prove that for any $m>1$, there exists a family of good qudit quantum codes supporting transversal logical $\mathsf{C}^{m-1}\mathsf{Z}$ gates that can address specified logical qudits and be largely executed in parallel. Building on the family of good quantum error-correcting codes presented in He et al. (2025), which support addressable and

  34. Yusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong

    Recent advances in multimodal models have demonstrated remarkable text-guided image editing capabilities, with systems like GPT-4o and Nano-Banana setting new benchmarks. However, the research community's progress remains constrained by the absence of large-scale, high-quality, and openly accessible datasets built from real images. We introduce Pico-Banana-4

  35. Guoyun Zhang

    The integration of Large Language Models (LLMs) with optimization modeling offers a promising avenue for advancing decision-making in operations research (OR). Traditional optimization methods,such as linear programming, mixed integer programming, and simulation depend heavily on domain expertise to translate real-world problems into solvable mathematical mo

  36. Xichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan

    Reinforcement learning from verifiable rewards has emerged as a powerful technique for enhancing the complex reasoning abilities of Large Language Models (LLMs). However, these methods are fundamentally constrained by the ''learning cliff'' phenomenon: when faced with problems far beyond their current capabilities, models consistently fail, yielding a persis

  37. David Mora, Viraat Aryabumi, Wei-Yin Ko, Sara Hooker

    Synthetic data has become a cornerstone for scaling large language models, yet its multilingual use remains bottlenecked by translation-based prompts. This strategy inherits English-centric framing and style and neglects cultural dimensions, ultimately constraining model generalization. We argue that the overlooked prompt space-the very inputs that define tr

  38. Carl-Johan Fauvelle Munck af Rosensch"old, Feras M. Awaysheh, Ahmad Awad

    In-memory key-value datastores have become indispensable building blocks of modern cloud-native infrastructures, yet their evolution faces scalability, compatibility, and sustainability constraints. The current literature lacks an experimental evaluation of state-of-the-art tools in the domain. This study addressed this timely gap by benchmarking Redis alter

  39. Boris Alexeev, Dustin G. Mixon

    We resolve a $1000 Erd\H{o}s prize problem, complete with formal verification generated by a large language model. In over a dozen papers, beginning in 1976 and spanning two decades, Paul Erd\H{o}s repeatedly posed one of his "favourite" conjectures: every finite Sidon set can be extended to a finite perfect difference set. We establish that {1, 2, 4, 8, 13}

  40. Maret Einasto, Peeter Tenjes, Rain Kipper, Pekka Heinämäki

    We study the substructure, connectivity, and galaxy content of galaxy clusters A1656 and 1367 in the Coma supercluster and of A1185 in the Leo supercluster with the aim of understanding the evolution of clusters from turnaround to virialisation. We used data from the SDSS DR10 MAIN galaxy sample and from DESI cluster catalogues. The projected phase space dia

  41. Xiaozhen Qiao, Jingkai Zhao, Yuqiu Jiang, Xianda Guo

    Vision-Language Models (VLMs) demonstrate impressive zero-shot generalization through large-scale image-text pretraining, yet their performance can drop once the deployment distribution diverges from the training distribution. To address this, Test-Time Adaptation (TTA) methods update models using unlabeled target data. However, existing approaches often ign

  42. Sandra Malagon, Monica A. Ulloa Ruiz, Tatiana Elizabeth Sandoval Plaza, Gabriel Rafael Rosario Bolívar

    The rapid escalation of computational requirements for training large-scale language models has reinforced structural asymmetries between high-capacity jurisdictions and countries in the Global South. This paper examines the technical and fiscal feasibility of sovereign-scale language model training in Brazil and Mexico under conditions of constrained hardwa

  43. B. Bale, G. Tautvaisiene, R. Minkeviciute, A. Drazdauskas

    Aims: We carried out a detailed investigation of Lithium and CNO abundances, including carbon isotope ratios, in RS CVn stars to assess the role of magnetic activity in the mixing of stellar atmospheres. Methods: We obtained high-resolution spectra at the Moletai Astronomical Observatory. Lithium abundances were determined by spectral synthesis of the 6707 A

  44. Ji Ma, Albert Casella

    Public and nonprofit organizations often hesitate to adopt AI tools because most models are opaque even though standard approaches typically analyze aggregate patterns rather than offering actionable, case-level guidance. This study tests a practitioner-in-the-loop workflow that pairs transparent decision-tree models with large language models (LLMs) to impr

  45. Vishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang

    Spoken Question-Answering (SQA) is a core capability for useful and interactive artificial intelligence systems. Recently, several speech-language models (SpeechLMs) have been released with a specific focus on improving their SQA performance. However, a lack of controlled ablations of pretraining data processing and curation makes it challenging to understan

  46. C. Murray, R. Kou, J. G. Bartlett

    We explore the observational prospects for detecting gravitational lensing induced by cosmological matter currents, a relativistic correction to the standard density lensing effect arising from the motion of matter. We propose to isolate this contribution by cross-correlating the weak-lensing convergence field with a reconstructed cosmic momentum field infer

  47. Joseph Bak-Coleman, Cailin O'Connor, Carl Bergstrom, Jevin West

    Emerging information technologies like social media, search engines, and AI can have a broad impact on public health, political institutions, social dynamics, and the natural world. It is critical to develop a scientific understanding of these impacts to inform evidence-based technology policy that minimizes harm and maximizes benefits. Unlike most other glo

  48. Roey Magen, Gal Vardi

    Transformers have demonstrated impressive in-context learning (ICL) capabilities, raising the question of whether they can serve as metalearners that adapt to new tasks using only a small number of in-context examples, without any further training. While recent theoretical work has studied transformers' ability to perform ICL, most of these analyses do not a

  49. Rohith Kuditipudi, Jing Huang, Sally Zhu, Diyi Yang

    Suppose Alice trains an open-weight language model and Bob uses a blackbox derivative of Alice's model to produce text. Can Alice prove that Bob is using her model, either by querying Bob's derivative model (query setting) or from the text alone (observational setting)? We formulate this question as an independence testing problem--in which the null hypothes

  50. Jhionathan de Lima, Cristiano Francisco Woellner

    Herein, we conduct a comprehensive investigation of Hexa-graphyne (HXGY), a planar carbon allotrope formed by distorted hexagonal and rectangular rings incorporating sp and sp$^2$-hybridized carbon atoms. First-principles calculations confirm its energetic, dynamical and thermal stability (up to at least 1000 K). Regarding its band structure, this material e

  51. Benjamin Bergougnoux, Vera Chekan, Giannos Stamoulis

    For a graph $G$, the parameter treedepth measures the minimum depth among all forests $F$, called elimination forests, such that $G$ is a subgraph of the ancestor-descendant closure of $F$. We introduce a logic, called neighborhood operator logic with acyclicity, connectivity and clique constraints ($\mathsf{NEO}_2[\mathsf{FRec}]+\mathsf{ACK}$ for short), th

  52. Tomás Dodds, Wang Ngai Yeung, Claudia Mellado, Mathias-Felipe de Lima-Santos

    Using (generative) artificial intelligence tools and systems in journalism is expected to increase journalists' production rates, transform newsrooms' economic models, and further personalize the audience's news consumption practices. Since its release in 2022, OpenAI's ChatGPT and other large language models have raised the alarms inside news organizations,

  53. Saptarshi Sengupta, Zhengyu Zhou, Jun Araki, Xingbo Wang

    Tool calling has become increasingly popular for Large Language Models (LLMs). However, for large tool sets, the resulting tokens would exceed the LLM's context window limit, making it impossible to include every tool. Hence, an external retriever is used to provide LLMs with the most relevant tools for a query. Existing retrieval models rank tools based on

  54. Archana Warrier, Dat Nguyen, Michelangelo Naim, Moksh Jain

    World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as next-frame prediction or task return, and (ii) do not test whether a learned model supports diverse queries about the environment. In contrast, humans build $\textit{general

  55. Manuel Kauers, Isaac Wood

    Continuing recent investigations of bounding the tensor rank of matrix multiplication using flip graphs, we present here improved rank bounds for about thirty matrix formats.

  56. H. McCright, I. G. Abel, I. Haber, P. G. O'Shea

    Dispersive shock waves (DSWs) are expanding nonlinear wave trains that arise when dispersion regularizes a steepening front, a phenomenon observed in fluids, plasmas, optics, and superfluids. Here we report the first experimental observation of DSWs in an intense electron beam, using the University of Maryland Electron Ring (UMER). A localized induction-cell

  57. Nishant Balepur, Dang Nguyen, Dayeon Ki

    Multi-modal large language models (MLMs) are often assessed on static, individual benchmarks -- which cannot jointly assess MLM capabilities in a single task -- or rely on human or model pairwise comparisons -- which is highly subjective, expensive, and allows models to exploit superficial shortcuts (e.g., verbosity) to inflate their win-rates. To overcome t

  58. Mengxiang Zhu, Riccardo Rastelli

    As a core policy tool for China in addressing climate risks, green finance plays a strategically important role in shaping carbon mitigation outcomes. This study investigates the nonlinear and interaction effects of green finance on carbon emission intensity (CEI) using Chinese provincial panel data from 2000 to 2022. The Climate Physical Risk Index (CPRI) i

  59. Shixuan Liu, Yue He, Haotian Wang, Wenjing Yang

    Data-driven methods offer efficient and robust solutions for analyzing complex dynamical systems but rely on the assumption of I.I.D. data, driving the development of generalization techniques for handling environmental differences. These techniques, however, are limited by their dependence on environment labels, which are often unavailable during training d

  60. Miguel Sánchez de la Rosa, Francisco J. andújar, Jesus Escudero-Sahuquillo, José L. Sánchez

    The increase in computation and storage has led to a significant growth in the scale of systems powering applications and services, raising concerns about sustainability and operational costs. In this paper, we explore power-saving techniques in high-performance computing (HPC) and datacenter networks, and their relation with performance degradation. From th

  61. Prashant Kodali, Vaishnavi Shivkumar, Swarang Joshi, Monojit Choudhary

    We study model merging as a practical alternative to conventional adaptation strategies for code-mixed NLP. Starting from a multilingual base model, we: (i) perform continued pre-training (CPT) on unlabeled code-mixed text to obtain an adapted checkpoint, (ii) merge checkpoint with the base model, and (iii) fine-tune (FT) on the downstream task data. We eval

  62. Tomas Valencia Zuluaga, Simon Pang, Jean-Paul Watson

    We propose explicitly incorporating large-scale load siting into a stochastic nodal power system capacity expansion planning model that concurrently co-optimizes generation, transmission and storage expansion. The potential operational flexibility of some of these large loads is also taken into account by considering them as consisting of a set of tranches w

  63. Adam Karczmarz, Wojciech Nadara, Marek Sokołowski

    In this paper, we show new strongly polynomial work-depth tradeoffs for computing single-source shortest paths (SSSP) in non-negatively weighted directed graphs in parallel. Most importantly, we prove that directed SSSP can be solved within $\tilde{O}(m+n^{2-\epsilon})$ work and $\tilde{O}(n^{1-\epsilon})$ depth for some positive $\epsilon>0$. In particular,

  64. Yuezhou Hu, Jiaxin Guo, Xinyu Feng, Tuo Zhao

    Speculative Decoding (SD) accelerates large language model inference by employing a small draft model to generate predictions, which are then verified by a larger target model. The effectiveness of SD hinges on the alignment between these models, which is typically enhanced by Knowledge Distillation (KD). However, conventional KD methods aim to minimize the

  65. Anand Choudhary, Yasser Sulaıman, Lukas Mauch, Ghouthi Boukli Hacene

    Sparse fine-tuning techniques adapt LLMs to downstream tasks by only tuning a sparse subset of model parameters. However, the effectiveness of sparse adaptation depends on optimally selecting the model parameters to be fine-tuned. In this work, we introduce a novel sparse fine-tuning technique named GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Par

  66. Mahdiyar Mousavi-Sadr, Fatemeh S. Tabatabaei, Alexander Wolszczan, Ghassem Gozaliasl

    Radio observations provide a window into a planet's interior and play a crucial role in studying its atmosphere and surface, key factors to find potential habitability. The discovery of thousands of exoplanets, together with advances in radio astronomy through the Square Kilometre Array (SKA), motivates the search for planetary-scale radio emissions. Here, w

  67. Domantas Kuryla, Fabian Berger, Gábor Csányi, Angelos Michaelides

    Training of general-purpose machine learning interatomic potentials (MLIPs) relies on large datasets with properties usually computed with density functional theory (DFT). A pre-requisite for accurate MLIPs is that the DFT data are well converged to minimize numerical errors. A possible symptom of errors in DFT force components is nonzero net force. Here, we

  68. Alessio Zaccone

    Phonon spectra in solids often display anomalies that defy the simple Debye law, most prominently the van Hove singularity in crystals and the boson peak in glasses. Although traditionally regarded as distinct, both features are increasingly recognized as sharing a common physical origin. In a recent work, G. Ding et al. (Nat. Phys. 2025) propose a resonant-

  69. Euodia Dodd, Nataša Krčo, Igor Shilov, Yves-Alexandre de Montjoye

    Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability, the TPR at low FPR, to membershi

  70. André G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi

    Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings. In this work, we provide a comprehensive empirical study analyzing the impact of several multilingual design choices, such as training data composition, encoder selection, and text

  71. Jad Zarzour, Matthew Jablonski

    The integration of Industrial Internet of Things (IIoT) devices into manufacturing environments has accelerated the transition to Industry 4.0, but has also introduced new cybersecurity risks. This paper conducts a comprehensive security analysis of a commercial smart air compressor, revealing critical vulnerabilities including hardcoded credentials, unauthe

  72. Zhengyuan Du, Kangning Liu, Zhe-fei Yu

    We study the (type 0B) $\mathcal{N}=1$ supersymmetric complex Liouville string ($\text{S}\mathbb{C}\text{LS}$), a supersymmetric extension of the bosonic complex Liouville string ($\mathbb{C}\text{LS}$). We compute the sphere three-point amplitudes (including NS-NS-NS and NS-R-R types) and find they share the same form as the sphere three-point amplitude of

  73. Piotr Budzyński

    Weakly centered and spectrally weakly cenetered weighted composition operators in $L^2$-spaces are characterized. Criteria for existence of invariant subspaces are given. Additional results and examples are supplied.

  74. Craig Sanders, Billy Dickson, Sahaj Singh Maini, Robert Nosofsky

    In cognitive science and AI, a longstanding question is whether machines learn representations that align with those of the human mind. While current models show promise, it remains an open question whether this alignment is superficial or reflects a deeper correspondence in the underlying dimensions of representation. Here we introduce a methodology to prob

  75. Xichen Zhang, Sitong Wu, Haoru Tan, Shaozuo Yu

    The long chain-of-thought (LongCoT) capability is central to the recent breakthroughs achieved by large language models in complex reasoning tasks. However, the accompanying issue of ''underthinking'', where models exhibit shallow reasoning by frequently switching thoughts without sufficient exploration, limits both performance and token efficiency. To addre

  76. Nubio Vidal, Naghmeh Moradpoor, Leandros Maglaras

    Operational technology (OT) networks are increasingly coupled with information technology (IT), expanding the attack surface and complicating incident response. Although OT standards emphasise incident reporting and evidence preservation, they do not specify what data to capture during an incident, which hinders coordination across stakeholders. In contrast,

  77. Jan Zelinka, Oliver Kost, Marek Hrúz

    We present a data generation framework designed to simulate spoofing attacks and randomly place attack scenarios worldwide. We apply deep neural network-based models for spoofing detection, utilizing Long Short-Term Memory networks and Transformer-inspired architectures. These models are specifically designed for online detection and are trained using the ge

  78. Hongyu Ding, Xinyue Liang, Yudong Fang, You Wu

    In this paper, we propose SEA, a novel approach for active robot exploration through semantic map prediction and a reinforcement learning-based hierarchical exploration policy. Unlike existing learning-based methods that rely on one-step waypoint prediction, our approach enhances the agent's long-term environmental understanding to facilitate more efficient

  79. Vinay Banakar, Suli Yang, Kan Wu, Andrea C. Arpaci-Dusseau

    Memory tiering in datacenters does not achieve its full potential due to hotness fragmentation -- the intermingling of hot and cold objects within memory pages. This fragmentation prevents page-based reclamation systems from distinguishing truly hot pages from pages containing mostly cold objects, fundamentally limiting memory efficiency despite highly skewe

  80. James C. Knight, Johanna Senk, Thomas Nowotny

    The majority of research in both training Artificial Neural Networks (ANNs) and modeling learning in biological brains focuses on synaptic plasticity, where learning equates to changing the strength of existing connections. However, in biological brains, structural plasticity - where new connections are created and others removed - is also vital, not only fo

  81. Max Dupré la Tour, Manuel Lafond, Ndiamé Ndiaye

    Leaf powers and pairwise compatibility graphs were introduced over twenty years ago as simplified graph models for phylogenetic trees. Despite significant research, several properties of these graph classes remain poorly understood. In this paper, we establish that the recognition problem for both classes is NP-complete. We extend this hardness result to a b

  82. Manuchehr Aminian, Kristin M. Kurianski

    We investigate the application of a framework for sparse model identification of differential equations from timeseries data in the context of compartmental models in epidemiology. Such frameworks often seek a sparse representation from a polynomial basis in the state variables which reproduces the timeseries. Out-of-the-box approaches for the underlying spa

  83. Mohamed ElShehaby, Ashraf Matrawy

    Adversarial attacks pose significant challenges to Machine Learning (ML) systems and especially Deep Neural Networks (DNNs) by subtly manipulating inputs to induce incorrect predictions. This paper investigates whether increasing the layer depth of deep neural networks affects their robustness against adversarial attacks in the Network Intrusion Detection Sy

  84. Shaohang Jia, Zhiyong Huang, Zhi Yu, Mingyang Hou

    Quantization-Aware Training (QAT) is a critical technique for deploying deep neural networks on resource-constrained devices. However, existing methods often face two major challenges: the highly non-uniform distribution of activations and the static, mismatched codebooks used in weight quantization. To address these challenges, we propose Adaptive Distribut

  85. Joseph Casale, Andrew Silverschotz, Joseph DeSimone

    Top-K masking schemes have been proposed as a method to promote sparse representations in Information Retrieval (IR) tasks, as a simple alternative to Floating Point Operations per Second (FLOPS) regularization. Algorithms such as Bilingual Lexical and Document Expansion Model (BLADE), adopt this approach as a post-processing stage. We propose using Top-P Dy

  86. W. -J. Ong, Z. Y. Xu, R. Grzywacz, A. Ravlić

    In an experiment performed at the Facility for Rare Isotope Beams (FRIB) using the FRIB Decay Station initiator (FDSi), 15 new half lives of isotopes near $^{54}$Ca were measured. A new method of extracting lifetimes from experimental data, taking into account the unknown $\beta$-delayed neutron emission branches of very neutron-rich nuclei, was developed to

  87. Georges Habib, Andreas Savas-Halilaj

    We investigate harmonic unit vector fields with totally geodesic integral curves on 3-manifolds. Under mild curvature assumptions, we classify both the vector fields and the manifolds that support them. Our results are inspired by Carriere's classification of Riemannian flows on compact three-manifolds, as well as by the works of Geiges and Belgun on Killing

  88. Jiacheng Liu, Xinyu Wang, Yuqi Lin, Zhikai Wang

    Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existi

  89. Mostafa Ameli, Sulthana Shams, Van Anh Le, Alexander Skabardonis

    The traffic assignment problem is essential for traffic flow analysis, traditionally solved using mathematical programs under the Equilibrium principle. These methods become computationally prohibitive for large-scale networks due to non-linear growth in complexity with the number of OD pairs. This study introduces a novel data-driven approach using deep neu

  90. Aman Bilkhoo, Mehran Hosseini, Milad Kazemi, Nicola Paoletti

    Counterfactual explanations (CFXs) provide human-understandable justifications for model predictions, enabling actionable recourse and enhancing interpretability. To be reliable, CFXs must avoid regions of high predictive uncertainty, where explanations may be misleading or inapplicable. However, existing methods often neglect uncertainty or lack principled

  91. Qilin Ye, Deqing Fu, Robin Jia, Vatsal Sharan

    Transformers often fail to learn generalizable algorithms, instead relying on brittle heuristics. Using graph connectivity as a testbed, we explain this phenomenon both theoretically and empirically. We consider a simplified Transformer architecture, the Disentangled Transformer, and prove that an $L$-layer model can compute connectivity in graphs with diame

  92. Ameesh Shah, William Chen, Adwait Godbole, Federico Mora

    Solving complex real-world control tasks often takes multiple tries: if we fail at first, we reflect on what went wrong, and change our strategy accordingly to avoid making the same mistake. In robotics, Vision-Language-Action models (VLAs) offer a promising path towards solving complex control tasks, but lack the ability to contextually and dynamically read

  93. Robbie King, Robin Kothari, Ryan Babbush, Sergio Boixo

    This note presents a simplified version of the OTOC$^{(2)}$ problem that was recently experimentally implemented by Google Quantum AI and collaborators. We present a formulation of the problem for growing input size and hope this spurs further theoretical work on the problem.

  94. Rajat De, Dominik Kempa

    Compressed indexing is a powerful technique that enables efficient querying over data stored in compressed form, significantly reducing memory usage and often accelerating computation. While extensive progress has been made for one-dimensional strings, many real-world datasets (such as images, maps, and adjacency matrices) are inherently two-dimensional and

  95. Catherine Villeneuve, Benjamin Akera, Mélisande Teng, David Rolnick

    Species distribution models (SDMs), which aim to predict species occurrence based on environmental variables, are widely used to monitor and respond to biodiversity change. Recent deep learning advances for SDMs have been shown to perform well on complex and heterogeneous datasets, but their effectiveness remains limited by spatial biases in the data. In thi

  96. Yi Liu

    The following criterion is proved in this paper. If the Alexander polynomial of a knot $K\subset S^3$ has a zero of odd order on the complex unit circle, then there exists a continuous family of irreducible representations $\pi_1(S^3\setminus K)\to \mathrm{SL}(2,\mathbb{R})$ converging to an abelian representation of noncentral elliptic type. As an applicati

  97. Priyaranjan Pattnayak, Hussain Bohra

    Large Language Models (LLMs) are transforming software creation by enabling zero code development platforms. Our survey reviews recent platforms that let users build applications without writing code, by leveraging LLMs as the brains of the development process. We adopt a broad survey methodology, categorizing platforms based on key dimensions such as interf

  98. Eric Hics, Vinhthuy Phan, Kriangsiri Malasri

    Computer science's increased recognition as a prominent field of study has attracted students with diverse academic backgrounds. This has significantly increased the already high failure rates in introductory courses. To address this challenge, it is essential to identify struggling students early on. Incorporating in-class coding exercises in these courses

  99. Geyang Wang, Alexander Barg, Navin Kashyap

    Recoverable systems provide coarse models of data storage on the two-dimensional square lattice, where each site reconstructs its value from neighboring sites according to a specified local rule. To study the typical behavior of recoverable patterns, this work introduces an interaction potential on the local recovery regions of the lattice, which defines a c

  100. Zhicheng Jin, Xiaotong Sun, Li Zhen, Weihua Gu

    The literature on transportation network companies (TNCs), also known as ride-hailing services, has often characterized these service providers as predominantly substitutive to public transit (PT). However, as TNC markets expand and mature, the complementary and substitutive relationships with PT may shift. To explore whether such a transformation is occurri