Skip to content

AI/ML Data Pipelines

This guide covers flexFS configuration for machine learning workflows including training data ingest, model checkpointing, and multi-GPU/multi-node distributed training.

WorkloadFile SizesAccess Pattern
Training datasets (ImageNet, etc.)Millions of small files or large archivesRandom read, high IOPS
Large datasets (video, medical imaging)GiB-scale filesSequential read
Model checkpoints1-100 GiBSequential write, occasional read
TFRecord / WebDataset shards100 MiB - 1 GiBSequential read
Logs and metricsSmall, append-onlySequential write

For Sharded Datasets (TFRecord, WebDataset)

Section titled “For Sharded Datasets (TFRecord, WebDataset)”
Terminal window
configure.flexfs create volume \
--name training-data \
--metaStoreID 1 --blockStoreID 1 \
--blockSize 4MiB \
--compression lz4

If your dataset consists of millions of small files (e.g., individual images):

Terminal window
configure.flexfs create volume \
--name image-data \
--metaStoreID 1 --blockStoreID 1 \
--blockSize 256KiB \
--compression lz4

Small files don’t consume a full block in object storage. The smaller block size instead reduces read amplification and per-block cache memory when the access pattern is many small random reads.

Terminal window
configure.flexfs create volume \
--name checkpoints \
--metaStoreID 1 --blockStoreID 1 \
--blockSize 4MiB \
--compression zstd \
--retention 30d

Zstd provides better compression ratios than LZ4 for checkpoint data, which is written infrequently and read occasionally.

Terminal window
mount.flexfs start training-data /mnt/data \
--diskFolder '/local-nvme/cache/<pid>' \
--diskQuota 80% \
--diskMaxBlockSize 0 \
--noAtime

Key settings:

  • --diskFolder on a local NVMe drive absorbs repeated reads of hot data (e.g., the same dataset across multiple epochs). Keep the <pid> component so each mount process gets its own cache folder.
  • --diskQuota 80% allows the cache to use most of the local SSD.
  • --diskMaxBlockSize 0 lets full blocks into the disk cache. The default, 256K, keeps them out. The limit applies to the stored block size after compression and encryption, which can be slightly larger than the volume block size for incompressible data such as images.
  • --noAtime eliminates metadata writes from reads.

Every training node mounts the same volume:

Terminal window
# On each node
mount.flexfs start training-data /mnt/data \
--diskFolder '/local-nvme/cache/<pid>' \
--diskQuota 50% \
--diskMaxBlockSize 0 \
--noAtime

All nodes see the same filesystem. Data loaders on each node read different shards, and the local disk cache captures the working set per node.

PyTorch’s DataLoader with num_workers > 0 spawns multiple reader processes. FlexFS handles concurrent reads from the same mount point:

dataset = ImageFolder('/mnt/data/imagenet/train', transform=transform)
loader = DataLoader(dataset, batch_size=256, num_workers=8, pin_memory=True)

DataLoader worker processes run as the same user as the training process, so they access the mount without any additional FUSE configuration.

Write checkpoints directly to flexFS. For large models, the sequential write throughput is determined by the network path to storage (direct or via proxy):

torch.save(model.state_dict(), '/mnt/checkpoints/epoch_10.pt')

For hybrid cloud deployments with on-prem compute, enable --diskWriteback (with --diskQuota) or use a proxy group to mask write latency.

The Hugging Face datasets library uses memory-mapped Arrow files. To share its caches across nodes, point them at flexFS with the HF_HOME environment variable, which relocates both the datasets cache and the Hub download cache, before starting Python:

Terminal window
export HF_HOME=/mnt/data/hf-cache

HF_DATASETS_CACHE relocates only the Arrow cache and HF_HUB_CACHE only the Hub downloads. To choose the location for a single dataset, pass cache_dir to load_dataset(). See the Hugging Face cache documentation.

  1. Pre-shard your data: Use formats like TFRecord or WebDataset instead of millions of small files. This reduces metadata overhead.
  2. Cache sizing: Size the local disk cache to hold at least one epoch’s worth of data per node if possible.
  3. Proxy groups: For multi-region training, deploy proxy groups near GPU clusters to avoid cross-region reads.
  4. Separate volumes: Use different volumes for immutable training data (read-only tokens, long retention) and mutable checkpoints (write access, shorter retention).
  5. Read-only mounts: Mount training data as read-only (--ro) on training nodes to eliminate any write-path overhead.