Caching Architecture
FlexFS employs a three-tier caching architecture to minimize latency and reduce the number of requests to object storage. Each tier serves a different access pattern and can be independently configured.
Cache tiers at a glance
Section titled “Cache tiers at a glance”| Tier | Location | Eviction | Writeback | Edition |
|---|---|---|---|---|
| L1 | mount client memory | LRU | No (read cache) | Both |
| L2 | mount client disk | LRU (clean blocks) | Optional | Both |
| L3 | proxy server disk | LRU (clean blocks) | Yes | Enterprise |
L1: In-memory LRU cache
Section titled “L1: In-memory LRU cache”The L1 cache is an in-memory LRU (Least Recently Used) cache that sits at the top of the store pipeline, closest to the application. It caches decoded (decompressed, decrypted) blocks, so cache hits avoid decryption and decompression.
Key characteristics
Section titled “Key characteristics”- Capacity: Sized automatically from system RAM. The internal
--memCapacityflag overrides it with a percentage of RAM (e.g.2%), a size (e.g.4G,512M), or a number of blocks (e.g.2000). The cache cannot be turned off. - Eviction: LRU. When the cache is full, the least recently used block is evicted and its buffer is returned to the pool. Blocks also expire 8 hours after they are cached.
- Coalesced fetches: Concurrent requests for the same block are coalesced into a single downstream fetch (request coalescing). This prevents cache stampedes when many threads read the same file region simultaneously.
- Prefetch integration: Prefetched blocks are inserted into the L1 cache. If a block is already cached, the prefetch is skipped.
- Write-through: A block being written is added to the L1 cache as it is passed downstream.
L2: On-disk cache
Section titled “L2: On-disk cache”The L2 cache persists blocks to local disk, providing much larger capacity than memory. It caches blocks in their processed form (after compression and encryption), meaning cache hits still avoid network I/O but do require decompression and decryption.
Configuration flags
Section titled “Configuration flags”| Flag | Description |
|---|---|
--diskFolder | Directory for cached block files |
--diskMaxBlockSize | Maximum processed block size to cache on disk. Blocks larger than this after processing bypass the disk cache. Set to 0 for no limit. |
--diskQuota | Maximum disk usage for the cache (e.g., 5%, 64M, 10G) |
--diskSync | Fsync each block written in writeback mode to local disk before acknowledging it (see below) |
--diskWriteback | Enable writeback mode (see below) |
See the mount.flexfs CLI reference for types and defaults.
LRU eviction
Section titled “LRU eviction”The disk cache maintains an LRU list for clean blocks (blocks that have been persisted to object storage). When a newer version of a block is cached, for example after the block is rewritten, the version it replaces stays readable but moves to the eviction end of the list. When the cache needs space for new blocks:
- A clean block is evicted: one replaced by a newer version if there is one, otherwise the least recently used one.
- Its disk file is deleted.
- If no clean blocks remain (the cache is full of dirty blocks in writeback mode), the new block bypasses the cache entirely.
Dirty blocks are never evicted. A dirty block replaced by a newer version is still uploaded, and is then the first to be evicted.
Writeback mode
Section titled “Writeback mode”When --diskWriteback is enabled, the disk cache operates as a writeback cache for writes:
- The mount client writes the processed block to the disk cache.
- The write is acknowledged to the application immediately (low latency).
- A pool of background worker threads reads dirty blocks from disk and persists them to object storage asynchronously.
- After a dirty block is successfully persisted, it transitions to a clean state and becomes eligible for LRU eviction.
Writeback mode is especially useful for:
- Latency-sensitive workloads: Write latency is bounded by local disk speed rather than object storage round-trip time.
- Bursty write patterns: The disk cache absorbs write bursts while background workers drain to storage at a sustainable rate.
- On-premises deployments: When object storage is in a remote cloud region, writeback caching masks the network latency.
The number of writeback workers matches the mount’s block-operation concurrency limit. The downstream store retries failures for up to 24 hours with randomized backoff capped at 5 seconds. If an upload still fails, the mount logs it once, keeps serving the block from the disk cache, and retries the upload with backoff of up to one minute until it succeeds. Reading a block back from the disk cache for upload is retried with backoff of up to 10 seconds. A block file that is missing or fails its integrity check is not retried: the mount logs an error, and a damaged file is renamed with a .corrupt suffix. Unmounting waits for pending uploads, but stops waiting if they make no progress for 30 seconds or the mount receives a stop signal; unsynced writes still in memory are then dropped. A block not yet uploaded when the mount process stops stays on local disk and is uploaded by a later read-write mount of the volume on that host that uses the same --diskFolder setting (see Performance tuning). This recovers data from files that were synced or closed before the mount stopped.
L3: Proxy group cache (Enterprise)
Section titled “L3: Proxy group cache (Enterprise)”Proxy groups provide a shared caching layer between mount clients and object storage. They function like a content delivery network (CDN) for block data.
How proxy groups work
Section titled “How proxy groups work”-
Group selection: When a mount client starts, and whenever its volume’s proxy groups change, it sends a health-check request to every server of every proxy group configured for its volume. It selects the group with the lowest round-trip time (RTT), taking a group’s RTT from its slowest server; a group counts only if all of its servers answer within 125 ms. After a proxy request still fails after 15 seconds of retries, the client goes direct for 5 minutes and then probes again, preferring any other group that answers.
-
Block routing: Within the selected group, blocks are distributed across the proxy servers so that:
- The same block always routes to the same proxy server (cache consistency).
- Adding or removing a proxy server only redistributes a fraction of blocks (minimal cache disruption).
-
Proxy-side caching: Each proxy server maintains its own disk cache. On a GET, if the block is cached locally, it is returned immediately. On a cache miss, the proxy fetches the block from object storage, caches it, and returns it to the client.
-
Proxy-side writeback: Proxy servers acknowledge PUT requests as soon as the block is written to the proxy’s local disk cache, then asynchronously flush the block to object storage. The block is flushed to disk with fsync before the acknowledgment only when proxy.flexfs runs with
--sync. -
Graceful fallback: If no proxy group answers the probe, or a proxy request still fails after 15 seconds of retries, mount clients automatically bypass the proxy and communicate directly with object storage for 5 minutes, then probe again in the background and return only to a group that answers. When the volume sets
maxProxied, each file’s blocks beyond that count also go directly to object storage.
Proxy group configuration
Section titled “Proxy group configuration”Proxy groups are created and associated with volumes via configure.flexfs:
- proxy-group: Defines a group with a provider, region, and comma-separated list of proxy server addresses.
- volume-proxy-group: Associates a proxy group with a volume. A volume can have multiple proxy groups; the mount client selects the best one based on RTT.
Max proxied blocks
Section titled “Max proxied blocks”Volumes have a maxProxied setting that limits how many blocks per file (by block index) are routed through proxies. Blocks beyond this index bypass the proxy and go directly to object storage. This is useful for large files where only the first portion benefits from proxy caching. When set to 0, all blocks are eligible for proxying.
Dirty block cache
Section titled “Dirty block cache”In addition to the three cache tiers, mount.flexfs maintains an in-memory dirty block cache for write operations. This cache holds modified blocks that have not yet been synced to storage.
Dirty blocks are flushed when:
- The application calls
fsync(),fdatasync(), orclose() - An open file has had no new writes for about 5 seconds
- The dirty cache reaches capacity (back-pressure flush)
- During sequential writes, blocks a few positions behind the current write are flushed early
- The volume is unmounted
Prefetching
Section titled “Prefetching”Mount.flexfs includes a block prefetcher that proactively loads blocks into the L1 cache before they are requested.
Prefetched blocks share the same request coalescing with regular reads, so a prefetch and a concurrent read for the same block result in a single downstream fetch.
Buffer pool
Section titled “Buffer pool”All block buffers are managed by a reusable buffer pool to minimize memory allocation overhead and garbage collection pressure. The pool capacity is sized automatically.
The pool keeps a set of buffers ready for reuse; it does not cap how many can be in use at once. The mount’s actual buffer usage is driven by demand, and mostly by the in-memory cache capacity, so that is the setting to adjust if the mount is using more memory than you expect. flexfs_mount_buffer_pool_borrowed and flexfs_mount_mem_cache_size report both.
FUSE kernel caching
Section titled “FUSE kernel caching”In addition to the userspace caches described above, mount.flexfs configures the Linux FUSE kernel module’s caching behavior, giving the kernel a TTL for cached file attributes, a TTL for cached directory entries, and a separate TTL for cached negative lookups. The attribute TTL is long by default; the two directory-entry TTLs default to one second.
These kernel-level caches reduce the number of FUSE round trips for repeated stat() and directory lookups. They are separate from and additive to the block caching tiers.
Filesystem statistics (statfs)
Section titled “Filesystem statistics (statfs)”Volume statistics (df, statfs(2)) are served from a short-lived cache in the mount client itself: an answer is reused for a few seconds (default 3) before the metadata server is asked again. Volume-level figures move slowly and have no invalidation channel, so a brief TTL is safe — and it is what absorbs repeated free-space polling, which would otherwise cost one metadata round trip per poll. During a metadata-server outage the last known answer is served regardless of age, so df keeps working.
Attribute coherence (event-driven)
Section titled “Attribute coherence (event-driven)”Cross-mount attribute and data-page coherence is event-driven. When another mount modifies a file, the metadata server notifies every other connected client (never the mount that made the change), and the receiving client invalidates the affected kernel cache entries immediately — typically within milliseconds. The attribute-cache TTL is therefore only a fallback: an upper bound on how long a missed notification could leave attributes stale, not the expected coherence window. flexfs_mount_ledger_size reports the working set of tracked inodes, and out-of-order notifications that would otherwise roll a file’s attributes (its size included) backward are discarded and counted in flexfs_mount_stale_attr_installs_rejected_total — a steady trickle on a busy multi-mount volume is normal, each one a stale update correctly ignored rather than lost work.
If the notification connection is disrupted and events may have been missed, the client re-invalidates every inode the kernel still references, so no stale attribute or data page survives the gap — which is what makes the long attribute-cache TTL safe. Recovery flushes inodes, not names, so cached names remain bounded by the much shorter directory-entry TTL across a disruption, exactly as they are the rest of the time (see below). Health metrics are listed under metrics, with ready-made alert rules under alerting.
Close-to-open freshness
Section titled “Close-to-open freshness”Invalidation notifications are fast, but they are not instant. On their own they would leave a brief window in which a reader that opens a file at the moment a peer’s close() returns could still be served the previous contents.
Opening the file closes that window. Whenever a mount opens a file it checks the file’s current state with the metadata server, and if another mount has changed the data since it last looked, it discards what it had cached before the open() returns. An application therefore never reads stale data from a file it has just opened, however soon after the peer’s write it opens it, and this holds for O_DIRECT readers as well.
Files that have not changed — overwhelmingly the common case — keep their cached copies, so repeat reads are still served from the kernel page cache at full speed.
The check costs one round trip to the metadata server per open(). Opening an already-cached file would otherwise need no round trips at all, so workloads that open very large numbers of files — compiler include scans, interpreter module search, large find traversals — are the ones likely to notice. Opens that cannot be serving stale data skip the check automatically: files this mount already has open, files it has just created, and files being truncated as they are opened.
flexfs_mount_cache_events_total{cache="openrefresh"} reports how many opens paid for the check and how often one found a change, so the cost is measurable on a live mount. The check has no supported off switch: the flag that disables it is internal, because a mount running without it knowingly returns stale data.
Directories are not covered by this. Cross-mount name coherence is bounded by the directory-entry TTL instead, as described next.
Lock coherence
Section titled “Lock coherence”Close-to-open needs a close(), and the idioms that coordinate with file locks do not close — they hold one descriptor open and take a lock around each critical section. Cross-mount locks therefore carry their own cache barriers: releasing one publishes this mount’s pending data and size before the metadata server can pass the lock on, and acquiring one discards this mount’s cached attributes, block keys and kernel pages for the file. An application that locks before it reads and unlocks after it writes sees the same data every mount sees, with no fsync of its own. The semantics are described under File locking.
The acquire barrier drops only the byte range the lock covers, so a byte-range lock does not evict the whole file. Cached attributes are dropped before the lock is granted, and cached pages are dropped with a deadline so that an acquisition never waits indefinitely behind a read already in progress on the same file; flexfs_mount_lock_data_invalidations_deferred_total counts the acquisitions where the page drop finished after the lock was granted.
Directory-entry (name) coherence
Section titled “Directory-entry (name) coherence”Directory-entry (dentry) caching works differently. The kernel caches name → inode lookups for the directory-entry TTL (the internal --entryValid flag) and negative (not-found) results for the negative-lookup TTL (the internal --negEntryValid flag), each one second by default. Cross-mount name coherence is bounded by these TTLs: in the worst case, a create, rename, or delete on one mount becomes visible to a cached name on another mount only after that mount’s TTL expires. (A name a mount has not looked up within the window is unaffected — it is resolved fresh on first access.) A mount’s own mutations keep its own kernel dentry cache correct as part of each syscall reply; the TTL bound applies only to names cached from a peer’s changes.
In practice most peer changes are reflected much sooner than the TTL. When a notification affects a name the kernel may have cached, the mount asks it to drop its cached answer straight away — so a file created, removed, or renamed elsewhere usually appears or disappears within milliseconds, including for names it only saw as part of a directory listing. This is a best-effort improvement rather than a guarantee: there is no way for a client to know every name the kernel has cached, so the directory-entry TTL remains the coherence bound to design against and should be kept short.
For a short period after the notification connection is disrupted, more peer changes fall back to expiring on the directory-entry TTL rather than being reflected immediately, and names are not covered by the recovery flush described above. This is another reason to keep that TTL short.
Mount client name cache. Each mount also caches its own name lookups, including not-found results, so a lookup that reaches the mount client usually needs no metadata round trip. A peer’s create, rename, or delete updates or drops a name this mount already holds as soon as its notification arrives. The cache follows the same TTLs the kernel is given (a TTL of 0 disables that class of record), and it is cleared when the notification connection is disrupted.
If your workload needs strict cross-mount name freshness — for example, one client polls for a file another client will create, or depends on an atomic rename swap being visible on the very next access — set --negEntryValid or --entryValid to 0 for that class of name. The negentryvalid and entryvalid volume flags, when set, override these mount flags. The trade-off is that every such LOOKUP reaches the mount client and, in practice, the metadata server, on every path resolution. Workloads dominated by repeated path resolution — build tools, interpreter module search, shell PATH scans — are best served by the defaults. Attribute and data coherence are unaffected: those remain event-driven regardless of the entry TTLs.
Negative lookups
Section titled “Negative lookups”Negative lookups have their own TTL. When a name does not exist, the kernel caches that fact for --negEntryValid (default one second), so repeated access to a missing path (common with build tools and shell PATH scans) is answered from the kernel without a metadata round trip. Keeping it separate from --entryValid lets a mount cache existing names while still resolving missing ones fresh, which matters when peers create files another mount is waiting for. Negative entries otherwise behave like positive ones (above): a create by this mount clears its own cached negative at once, and a peer’s create over a name this mount had cached as missing usually invalidates the kernel’s negative entry within milliseconds. Because that notification is best-effort, the negative-lookup TTL remains the worst case.
POSIX ACL mode
Section titled “POSIX ACL mode”Extended ACLs (--acl) can be enforced on either side of the FUSE boundary, and which side is in use decides whether the kernel caches at all. mount.flexfs chooses at mount time and logs the result in a line beginning ACL enforcement: kernel or ACL enforcement: client.
Kernel-enforced (Linux 4.9 and later — the normal case). The kernel evaluates ACLs itself, so a cached attribute or name is still checked before it is used. Caching therefore behaves exactly as it does without --acl: the attribute-cache and directory-entry TTL defaults above apply unchanged, and coherence works the same way, because the invalidation the mount client already sends for a peer’s change also drops the kernel’s cached ACL.
Client-enforced (the fallback). The mount client checks every operation itself, using each name lookup as the point where directory-traversal permission is verified. A cached name would let the kernel resolve a path without that check, so the kernel’s attribute and directory-entry caches are turned off (their effective TTLs are 0) and every lookup, including negative ones, reaches the mount client. Cached file data is unaffected. Expect noticeably more metadata round trips — repeated stat() calls on the same file reach the mount client every time instead of being answered by the kernel.
Two situations select the client-enforced path:
- Kernels older than 4.9, which cannot evaluate ACLs for a FUSE filesystem at all. In practice this means RHEL/CentOS 7.
- Root or all squashing, on any kernel, whether set with
--rootSquashor--allSquashor by the volume’s or token’srootsquashorallsquashflag. Squashing cannot be delegated: the kernel grants root every access, and it has no notion of remapping an unprivileged caller at all, so only the mount client can re-evaluate the request as the anonymous user.
The mount client’s own caches are unaffected by any of this and keep working in both modes.