Skip to content

Backup and Recovery

FlexFS stores metadata and block data separately, each with its own durability characteristics and backup strategies.

Data typeStorage locationDurability
Block data (file contents)Cloud object storage (S3, GCS, Azure, OCI)Provided by the cloud provider’s SLA (typically 99.999999999% / 11 nines for S3).
Metadata (inodes, directory entries, attributes, ACLs)Metadata server local databaseDepends on the local disk and your backup strategy.
Configuration (volumes, tokens, accounts, stores)Admin server local databaseDepends on the local disk and your backup strategy.

Block data stored in cloud object storage is inherently durable — you rely on your cloud provider’s durability guarantees. The primary backup concern is the metadata database on the metadata server and the configuration database on the admin server.

The metadata server stores its database in the folder specified by --dbFolder (default ~/.flexfs/meta/data in the home folder of the user running the server). To back up the metadata:

  1. Online checkpoint (recommended): Ask the running server for a consistent point-in-time checkpoint via the POST /v2/backup endpoint, then rsync the staging directory offsite — no downtime required. See Metadata → Maintenance → Online backup for the full workflow and scheduling guidance.

    Terminal window
    curl -fsk -X POST -H "Authorization: Bearer <meta-token>" "https://localhost:8443/v2/backup?blocking=true"
    rsync -a --delete --exclude '.incoming/' ~/.flexfs/meta/data/.backup/ backup-host:/backups/meta/

    blocking=true makes the request wait and return the checkpoint’s status, so curl -f fails loudly on a bad backup; without it you get an immediate 202 and failures go unnoticed.

  2. Filesystem snapshots: If the database resides on a volume that supports snapshots (LVM, ZFS, EBS snapshots), take a snapshot while the server is running. The database is crash-safe and will recover on startup.

  3. Cloud instance snapshots: For cloud-hosted metadata servers, use your provider’s instance or disk snapshot feature (e.g., AWS EBS snapshots, GCP persistent disk snapshots).

  4. Cold copy (fallback): Stop the metadata server, copy the entire database folder, then restart. Only needed if you cannot use the online method. This example is for a server run as a systemd service:

    Terminal window
    sudo manage.flexfs stop meta
    sudo cp -a /root/.flexfs/meta/data /backup/meta-data-$(date +%Y%m%d)
    sudo manage.flexfs start meta

The admin server stores volume definitions, accounts, tokens, block stores, and metadata stores in a SQLite database, the file db in its --dbFolder. By default this is ~/.flexfs/admin/db in the home folder of the user running the server; the examples below use /root/.flexfs/admin/db, the path for a server run as root, such as the systemd service. The database runs in write-ahead log mode: recent changes can sit in the db-wal file beside it, so copying the files while the server is running can produce an unusable backup. Use one of these methods:

  • Online (recommended): Use the SQLite online backup command, which produces a consistent copy while the server keeps running (requires the sqlite3 command-line tool):

    Terminal window
    sudo sqlite3 /root/.flexfs/admin/db ".backup '/backup/admin-db-$(date +%Y%m%d)'"
  • Cold copy: Stop the admin server, copy the database file together with its db-wal and db-shm files if they exist, then restart it. A server that was killed rather than stopped cleanly can leave committed changes in db-wal. This example is for a server run as a systemd service:

    Terminal window
    sudo manage.flexfs stop admin
    sudo mkdir -p /backup/admin-db-$(date +%Y%m%d)
    sudo sh -c 'cp -p /root/.flexfs/admin/db* /backup/admin-db-$(date +%Y%m%d)/'
    sudo manage.flexfs start admin

Filesystem and disk snapshots of the admin server are also safe, as for the metadata server. Also keep copies of the other files in the admin folder that a restore needs: the credentials file creds, the access file access if you use one, and any custom TLS certificate and key.

FlexFS supports time-travel mounting via the --atTime flag, which lets you mount a volume at a specific point in time. This is powered by data retention — both metadata and block data are preserved for the configured retention period.

Terminal window
mount.flexfs start <volume-name> /mnt/recovery --atTime 2026-03-15T10:30:00Z

The --atTime flag accepts an RFC 3339 timestamp. The mount is automatically set to read-only mode.

ScenarioHow time-travel helps
Accidental deletionMount at a time before the deletion, copy the files to the current mount.
Data corruptionMount at a known-good time and compare or restore files.
AuditingView the exact state of the filesystem at a specific date and time.
ComplianceDemonstrate data state at a particular point for regulatory requirements.

Time-travel range depends on the volume’s retention setting, which controls how long deleted and overwritten data is preserved. In Enterprise, set retention when creating or updating a volume with configure.flexfs. In Community, set it with free.flexfs init creds --retention (default 7 days).

While a disk is failing, the metadata server stops accepting changes for any volume whose database it cannot read; see Metadata read errors.

  1. Provision a new server, ideally reachable at the old server’s address (IP or DNS name).
  2. Install flexFS binaries (sudo manage.flexfs download meta --install).
  3. Restore the metadata database from your most recent backup. For an online checkpoint, place the per-volume <volumeID>/ directories into the server’s --dbFolder — each is already a complete, consistent database.
  4. Run meta.flexfs init creds with the same admin server address and token.
  5. Start the metadata server.
  6. If the server’s address changed, point the volumes at it: in Enterprise, run configure.flexfs update meta-store <id> --address <new-addr>; in Community, change metaAddr in the free server’s credentials file and restart the free server.
  7. Mount clients reconnect automatically, picking up a changed address from the admin or free server.
  1. Provision a new server reachable at the old server’s address (IP or DNS name). Metadata servers, mount clients, configure.flexfs, and the CSI driver all connect to that address.
  2. Install flexFS binaries (sudo manage.flexfs download admin --install).
  3. Restore the admin database from backup into the admin server’s --dbFolder: the online backup file as db, or all files of a cold copy.
  4. Restore the other files of the admin folder (~/.flexfs/admin by default): the credentials file creds, or recreate it with admin.flexfs init creds and the same admin token; the access file access, if you use one; and the deploy folder, or repopulate it with manage.flexfs deploy. Restore any custom TLS certificate and key as well.
  5. Start the admin server.
  6. Metadata servers and mount clients reconnect automatically.

Scenario: Block data loss in object storage

Section titled “Scenario: Block data loss in object storage”

Block data loss in a major cloud provider’s object storage is extremely rare. If it occurs:

  1. Files referencing lost blocks will return I/O errors on read.

  2. Use the GET /v2/missing-objects endpoint to enumerate the block keys the metadata server tracks whose backing object is absent from the block store. Each entry reports the affected inode (ino) and block index (idx).

    Terminal window
    curl -fsk -H "Authorization: Bearer <meta-token>" \
    "https://localhost:8443/v2/missing-objects?vid=<volume-name>&live-only=true&pretty=true"

    Omit vid to scan every open volume. live-only=true limits the scan to the blocks backing current file contents; drop it to also check retired blocks, which back time-travel reads within the retention window.

  3. Resolve each inode number to a path with find.flexfs --ino <ino>.

  4. Restore the affected files from external backups if available.

ComponentBackup frequencyRetention
Metadata server databaseDaily (minimum), hourly for critical data. The online checkpoint method needs no downtime, so sub-hourly (e.g. per-minute) schedules are practical.30 days
Admin server databaseDaily30 days
Block dataN/A (cloud provider durability)Per your data lifecycle policy