FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval
Long Yang, Yu Mao, Yuchen Shao, Yumiao Zhao, Yaqi Li, Xuan Liu, Xiaolong Shen, Tao Yu, Gezi Li, Jing Wang, Chengcheng Wan, Liang Shi
Abstract
With the rapid expansion of large-scale media-text corpora, multi-modal datasets increasingly require efficient storage and retrieval. Existing formats such as Files, TDP, and FFRecord work adequately for uni-modal data but expose fundamental limitations in multi-modal settings, including storage redundancy, massive small-file overheads, cache-unfriendly layouts, and heavy index structures. These issues jointly inflate storage and memory usage and make I/O the dominant bottleneck in real training workloads. We present FFSlim, a lightweight format for storing and retrieving multi-modal data. FFSlim improves storage efficiency and loading throughput through three components: a unified file format that removes media duplication and avoids small-file proliferation; an adaptive retrieval mechanism that enables low-overhead pair-level access and accelerates repeated media loading; and a redundancy detection and aggregation module that converts existing datasets into the FFSlim layout. The experimental results demonstrate that FFSlim achieves 2.07x and 8.26x higher data loading and write throughput on average than the strongest baseline, with minimal storage and index overhead. Consequently, these underlying I/O accelerations enable FFSlim to reduce end-to-end training time by 5.36%-14.18% across seven diverse multi-modal models.
Create a lesson
Related papers
Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation
Heyuan Yao, Chutong Gao, Yuan Lyu et al.
The Price of Remembering: A Calibrated Energy Law for Computation
Mohamed Amine Bergach
DART: Aiming for Tail-Delay Control in Reconfigurable Networks
Hossein Mohammadalizadeh, Holger Karl
Spectral Analysis for Sparse Matrix Computation: Insights and Potential
Ruifeng Zhang, Xipeng Shen
Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton et al.
Adaptation Fidelity of SPEC CPU2026
Doa'a Al-Otoom, Mahesh Madhav