Skip to content

Data Flow

FlexFS separates metadata operations from block data operations. Metadata flows through an RPC protocol to the metadata server; block data flows over HTTPS REST to object storage (or through proxy servers). This page traces both paths in detail.

When an application reads a file, the request passes through the kernel’s FUSE layer, into mount.flexfs, and then fans out to the metadata server and block storage.

Read path sequence diagram showing request flow from application through kernel, mount, cache tiers, and storage Read path sequence diagram showing request flow from application through kernel, mount, cache tiers, and storage
  1. FUSE dispatch: The kernel delivers a FUSE_READ request containing the file handle, offset, and length. mount.flexfs translates the offset into a block index: blockIdx = offset / blockSize.

  2. Block key lookup: The mount client looks up the block’s storage key from a local cache of inode-to-block-key mappings. If the mapping is not cached, it issues an RPC to the metadata server to retrieve the current block keys for the inode.

  3. Cache lookup: The block is looked up by its composite key (inode number, block index, storage key) in the L1 memory cache first, then the L2 disk cache. The L1 cache is an in-memory LRU with configurable capacity. The L2 disk cache uses an LRU eviction policy for clean blocks and can optionally operate in writeback mode.

  4. Remote fetch: On a cache miss, the block is fetched from either a proxy server (if proxy groups are configured and reachable) or directly from object storage. Concurrent requests for the same block are coalesced into a single downstream fetch (request coalescing).

  5. Processing pipeline: The raw block from storage is first decrypted (AES-256-GCM, if encryption is enabled), then decompressed (LZ4, Snappy, or zstd, depending on volume configuration). The result is a full-size block matching the volume’s configured block size.

  6. Prefetching: When a sequential read pattern is detected, mount.flexfs prefetches subsequent blocks in the background. Prefetched blocks populate the cache before they are needed, reducing read latency for sequential workloads.

Writes follow a similar but reversed flow, with dirty block management adding an additional layer.

Write path sequence diagram showing request flow from application through dirty cache, background flush, and storage upload Write path sequence diagram showing request flow from application through dirty cache, background flush, and storage upload
  1. FUSE dispatch: The kernel delivers a FUSE_WRITE request. mount.flexfs determines the target block index and copies the data into that block’s dirty buffer. A partial-block write does not read the existing block at write time; the unwritten parts are filled in from the existing block when the dirty block is flushed or read.

  2. Dirty block cache: The modified block is stored in the in-memory dirty block cache. The FUSE write returns immediately, giving the application low-latency write acknowledgment. When the dirty cache fills, the oldest dirty block is handed off for upload; the write waits only until an upload slot is free, and the upload itself runs in the background.

  3. Sync trigger: Dirty blocks are flushed to storage on any of these events:

    • Application calls fsync() or fdatasync()
    • File handle is flushed (close())
    • An open file has had no new writes for about 5 seconds
    • Dirty cache capacity pressure
    • During sequential writes, blocks a few positions behind the current write are flushed early
    • The volume is unmounted
  4. Processing pipeline: Before upload, the block is compressed using the volume’s configured algorithm (LZ4 by default, or Snappy, zstd, or none). If encryption is enabled, the compressed block is then encrypted with AES-256-GCM. The processing order on write is: compress then encrypt. On read, it is reversed: decrypt then decompress.

  5. Block key allocation: Each new version of a block is stored as a new object with its own key. This allows flexFS to retain previous versions of blocks for the volume’s configured retention period, enabling point-in-time recovery.

  6. Upload: The processed block is uploaded to object storage (or through a proxy server). A block that is entirely zeros is not uploaded; it is recorded as a hole. If local disk writeback caching is enabled, the block is written to the disk cache first and the upload happens asynchronously in the background, further reducing write latency.

  7. Metadata update: After the block is stored (uploaded, or written to the local disk cache in writeback mode), mount.flexfs records the new block key for that inode and block index on the metadata server, no later than when the file is synced or closed.

Every block in object storage is identified by three components:

ComponentDescriptionExample
Inode number (ino)The file’s unique inode number42
Block index (idx)The block’s position within the file (0-indexed)3
Storage key (key)Version key, unique to each version of the block1711234567_890123456

These components are combined to form the object key in the storage bucket:

{prefix}/{volumeID}/{ino}/{idx}/{key}

where {prefix} is the block store prefix. For example: flexfs/0b5c3a9e-7f41-4d2a-9e6b-2c8d1f4a7e30/42/3/1711234567_890123456

When a prefix contains the string partition, flexFS replaces it with a hash of the inode and block index to distribute objects across many key prefixes. This improves throughput on storage backends that partition by key prefix. See Partition-based key distribution.

The mount client maintains a persistent WebSocket connection (wss://) to the metadata server, communicating via a custom binary RPC protocol carried as WebSocket binary messages. Each RPC carries a request payload and returns a response with a status code. The protocol supports the full range of filesystem operations:

  • Namespace: Lookup, MkDir, MkNod (also used to create regular files), Symlink, Link, Rename, Unlink, RmDir
  • Attributes: GetAttr, SetAttr
  • Data: block key lookups and updates
  • Directory: ReadDir, ReadDirPlus (paginated directory streams)
  • Locking: Lock, Unlock (both POSIX fcntl and BSD flock semantics)
  • Extended attributes: returned with a file’s attributes and set with SetAttr; RemoveXattr removes one
  • Session: Hello (opens the session), Ping, StatFs
  • Notifications: the metadata server pushes change events to other connected mounts

Permission and ACL checks are made by the kernel or the mount client, not by a separate RPC.

To cut round trips, the mount client caches metadata locally: Lookup results (name → inode resolutions, including negative lookups) in an internal dentry cache, attributes in an attribute cache, and inode-to-block-key mappings — each a bounded, per-mount cache kept coherent across mounts by the invalidation event bus. See Caching Architecture.

TLS is enabled by default on the RPC connection. It can be disabled for testing.