FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch
Dhiren Mukesh Khatri
Abstract
Training molecular machine-learning models on ephemeral or memory-constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage. FoldPipe is a lightweight Python orchestration layer for already-sharded PyTorch and PyTorch Geometric data. It retrieves one shard ahead in a background thread while the consumer trains on the current shard, keeping the number of live shard payloads bounded with respect to total dataset size. Asynchronous prefetch and bounded buffering are established systems techniques rather than novel scheduling algorithms. FoldPipe's contribution is a small integration targeted at native .pt molecular shards together with a source-pinned empirical characterization of its operating regime. We evaluate a SchNet energy-and-force workload on MD17 aspirin using 20 paired, order-alternating benchmark passes on a Tesla T4. Each pass processes five pinned shards containing 25,000 structures. FoldPipe records 16.33 s mean I/O-compute overlap, compared with zero by construction for the sequential bounded baseline. Mean pass time is 76.78 s for FoldPipe and 83.37 s for the baseline. However, the geometric mean paired speedup is 1.059× with a 95% bootstrap interval from 0.878× to 1.288×. The experiment therefore verifies the overlap mechanism but is inconclusive about a reliable wall-clock speed advantage under the observed public-network variability.
Create a lesson
Related papers
TOPIQ: Statistical Error Propagation for Quantity-of-Interest Prediction under Lossy Compression
Youyuan Liu, Bo Jiang, Taolue Yang et al.
Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models
Gokulakannan Sakthivel, Jerry Wu, Amogh Rajendra et al.
Direct-Operable SIMD Bit-Slicing: A Framework for Memory-Efficient Predicate Evaluation
Arunkumar Mathiyazhagan
When Structure is Silent: Opportunities for Algorithmic Dispatch in Linear Algebra
Emmanuel Lujan, Alan Edelman
Portable to Efficient: Auto-Tuning Hardware-Agnostic GPU Kernels in Julia
Floris-Jan Willemsen, Evelyne Ringoot, Alan Edelman
Accelerating Performance Inference over Closed Systems by Asymptotic Methods
Giuliano Casale