Presentation: From S3 to GPU in One Copy: Rethinking Data Loading for ML Training

| Source: InfoQ AI/ML

Tags: Vortex, columnar file format, ML training, GPU, data loading, SpiralDB, Linux Foundation, zero-copy

Vortex, an open-source columnar file format now under the Linux Foundation, achieves S3-to-GPU data streaming at up to 60 Gbps via zero-copy pipelines and dynamic column pruning — eliminating the CPU bottleneck that throttles ML training data loading without requiring upfront dataset reformatting.

Details

ML training pipelines routinely bottleneck on data loading: CPUs spend cycles decoding and copying data before it ever reaches the GPU. Vortex, an open-source columnar file format developed at SpiralDB and now part of the Linux Foundation (LF AI & Data), directly attacks this with a zero-copy memory pipeline that streams data from S3 straight to the GPU. The format uses three interlocking techniques: cascading lightweight encodings that compress data while keeping it GPU-decodable, layout-based segment pruning that skips unneeded column bytes at read time, and a unified memory pipeline that avoids the traditional S3 → CPU → system RAM → GPU copy chain. In a QCon London demonstration, Vortex streamed a 4K video (three columns, ~8M pixels per frame at 60fps, ~13 Gbps) from S3 to GPU with real-time visualization — the GPU itself re-encoded the output to avoid downlinking the full stream. Theoretical throughput ceiling is 60 Gbps, which would saturate high-bandwidth interconnects rather than the format itself. Critically, no upfront data reprocessing is required — existing datasets can be read with dynamic pruning applied on the fly. Linux Foundation membership signals a permissive open-source license appropriate for commercial ML production use. The format is actively maintained by Onur Satici (SpiralDB) as core maintainer.