Performance Evaluation of RED-ONION: A High-Speed Disk-to-Disk Transfer System
Keichi Takahashi, Hiroaki Kataoka, Takeo Hosomi, Ayahiro Takaki, Yasunori Kakizawa, Shuichi Ihara, Nobuaki Hashizume, Susumu Date
Abstract
Modern experimental instruments produce data faster than general-purpose file transfer interfaces can move it, so delivery to the computing infrastructure has become a bottleneck in the research process. At many universities and research institutes, moreover, the instruments that generate research data and the high-performance computing systems that analyze it are separated both geographically and organizationally, because each demands its own expertise and installation environment. Connecting the two seamlessly is a pressing challenge for data-driven science. This article presents RED-ONION, a high-speed disk-to-disk transfer system that connects research facilities, on campus and beyond, to a computing center. The system combines data transfer nodes, a dedicated high-bandwidth network, an all-flash parallel file system, and multi-threaded transfer software that parallelizes network transmission and storage access. The design targets the wire rate both along the entire path, from the read on the sender storage to the write on the receiver storage, and for a single file between one pair of nodes rather than only in aggregate over many files or nodes. We describe the end-to-end optimizations across the transfer software, the operating system, and the storage that this requires. We evaluate a prototype deployed over a 100 Gbps transpacific path between Atlanta and Tokyo with a 150 ms round-trip time, on which a single 1 TB file transfer reached 90 Gbps, delivering a terabyte in approximately 95 s. Moving a dataset of this size therefore becomes a routine step, and the computing center serves an instrument as if the two were co-located.
Create a lesson
Related papers
Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret
Xia Jiang, Lu Liu, Gang Feng
A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems
Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata et al.
Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model
Minchen Yu
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction
Alfonso López-Ruiz, Diego Royo
Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC
Stepan Nassyr, Prateek Chawla, Daniel Seibel et al.