Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs
Joseph K. L. Lee, Mehrdad Malekmohammadi, Hong-Sheng Zheng, Shuli Shu, Cheick Doumbia, Kalman Szenes, Mehran Zamani Abnili, Thomas Ainsworth, Matthew Seymour, Thomas Germain, Leonhard Neuhaus, Josh Izaac, Lee J. O'Riordan
Abstract
Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time quantum error correction (QEC) demand performance that traditional interpreted environments cannot provide. FPGAs and ASICs play a central role at these layers, but their specialized programming models make development rigid and time-consuming. CPUs, GPUs, and other accelerators introduce a different challenge: as infrastructure becomes increasingly heterogeneous, programming across different devices and their associated abstractions becomes more complex. Allowing researchers to write workloads in high-level languages that map to low-latency execution across diverse distributed target platforms will enable the development of key infrastructure for utility-scale quantum systems. For this, we introduce Backline, a heterogeneous compilation and runtime framework built within PennyLane and Catalyst. Backline allows us to design and build quantum-classical workloads for high-performance and low-latency devices, with compilation directly from a Python interface through MLIR. We demonstrate the compilation and execution of several quantum workloads with low-latency data movement across a mix of CPUs, GPUs, and FPGAs, for both local and distributed remote hardware targets, all from a vendor-agnostic Python frontend. With an AMD VPK120 FPGA board as the controller, issuing each round from its hardware-handshake engine, we measured median steady-state round-trip latencies over RoCE v2 of 2.305~μs to an AMD Ryzen Threadripper PRO CPU and 4.5~μs to an AMD Instinct MI210 GPU across 106-1 rounds per path, demonstrating microsecond-scale synchronous co-processing.
Create a lesson
Related papers
Low-rank propagation for tridiagonalizable open quantum systems: near-linear scaling with system size
Roman Ovsiannikov, Kurt Jacobs, Andrii G. Sotnikov et al.
Superradiant Mpemba Relaxation in a Dicke Ladder
Matheus G. H. Santos, Hugo Sanchez, Italo M. de Araújo et al.
Thermalization and dephasing in an isolated system of coupled qubits
Jukka P. Pekola, Bayan Karimi
Effective Study of Superconducting Quantum Circuits
Carlos Raul Javier Valdez, Hector Hugo Hernandez Hernandez, Guillermo Chacon-Acosta
A Quantum Phase-based Comparator
Alessandro Berti, Alessandro Poggiali
Exploring Asymmetric QEC Code Concatenation
Sayam Sethi, Maxwell Poster, Aditi Awasthi et al.