Skip to content

Persistence & Backup

How cogniverse stores stateful data, what survives which failures, and how to configure the backup destination per environment.

The two-tier model

Cogniverse keeps a clear separation between live data and backups, with each tier in its own failure domain:

Tier What Examples
Primary Where live data lives. Read+write hot path. Vespa document store, the phoenix-postgres database, MinIO bucket contents
Backup Periodic snapshots in a different failure domain. Read on disaster recovery. S3-compatible object storage (in-cluster MinIO for dev, R2 / B2 / AWS S3 for prod)

A backup target in the same failure domain as primary is the OVH SBG2 fire (2021) and GitLab.com 2017 incident mistake. Cogniverse's backup CronWorkflow defaults to the in-cluster MinIO for dev convenience, but operators MUST point it at a separate failure domain for any production workload — see Backup destination below.

Durability matrix

What survives what, per primary-storage configuration:

Operation hostStorage.enabled=true (laptop dev) local-path PVC (single-node prod) Replicated CSI (Longhorn / Rook / cloud SC)
Pod restart ✓ ✓ ✓
StatefulSet recreate ✓ ✓ ✓
Laptop / node reboot ✓ ✓ ✓
helm uninstall cogniverse ✓ (host fs untouched) ✗ (reclaimPolicy: Delete) depends on SC
k3d cluster delete (or equivalent) ✓ (data on host fs) ✗ (Docker volume goes with cluster) ✗
Single node loss in multi-node cluster N/A ✗ (node-pinned) ✓ (replicated)
Disk failure on the only node ✗ ✗ ✗ (need backup tier)

The backup tier is what defends against disk failure, accidental helm uninstall, and full cluster loss. Set up Tier B even when Tier A looks safe.

Per-component primary storage

Each stateful component has its own <name>.persistence block in values.yaml. The chart honours the same shape for all of them:

<component>:
  persistence:
    enabled: true              # off → ephemeral emptyDir (don't do this)
    storageClass: "fast-ssd"   # cloud: gp3 / pd-ssd / managed-csi
    accessMode: ReadWriteOnce
    size: "100Gi"
    annotations: {}

Components

Component values key Default size Notes
Vespa vespa.persistence 100 Gi Document store + config server — holds every schema (video/image/document/audio embeddings, agent_memories, knowledge graph, provenance, tenant/org metadata, adapter registry) and per-tenant config overrides written through ConfigStore (schema config_metadata; scopes backend, gateway_agent, telemetry all land here — ConfigManager methods default to service="backend"). One tar backs up all of it. hostStorage.enabled=true overrides to hostPath bind-mount.
Phoenix phoenix.persistence 50 Gi Working directory only; traces, datasets and annotations live in phoenix-postgres. Backed up in mode: postgres: a pg_dump custom-format archive of that database plus a tar of the read-only mounted working directory, in one <service>-<timestamp>.tar.
MinIO minio.persistence 100 Gi Default backup destination on dev. See MinIO durability below.
HF model cache (per pod) hfCache.persistence 50 Gi each Off by default (enabled: false). When enabled, one PVC per inference svc + runtime + ingestor, pre-warmed via init container.
Redis redis.persistence 10 Gi Ingestion job queue and status streams, /ingestion/start job status, agent registrations made over /agents/register, the annotation review queue, A2A task state, agent-session message queues, conversation turn order and suspended /v1 turns, workflow and ingestion task events with their cancellations and each tenant's active tasks, approval and workflow locks (AOF on). One instance, never Redis Cluster (see Scaling). Lose it = re-ingest in-flight jobs and re-register agents registered over HTTP; in-flight A2A tasks, queued session messages, suspended /v1 turns and the progress events of running workflows are lost; the next annotation cycle re-queues spans still awaiting review.
LLM (builtin) llm.ollama.persistence (llm.engine: ollama) or llm.vllm.persistence (llm.engine: vllm) 100 Gi Model files for the in-cluster LLM. Only the PVC matching the selected llm.engine renders.
Semantic router semanticRouter.router.persistence 10 Gi Classifier bundle cache (~GB) for the vllm-sr sidecar so it isn't re-downloaded on every restart/rollout. emptyDir when disabled.

MinIO durability (load-bearing for dev)

MinIO IS the backup destination on dev clusters. If MinIO data dies with the cluster, the backup strategy is theatre. The chart provides two MinIO storage modes:

Mode 1 — PVC (default, suitable for cloud / multi-node prod)

minio:
  persistence:
    storageClass: "fast-ssd"   # any real CSI provider
    size: "200Gi"

Backed by the configured storageClass. Durability follows whatever that SC provides — replicated CSI (Longhorn, Rook, cloud SC) gives node-loss tolerance; local-path does not.

Mode 2 — hostPath bind-mount (dev / single-laptop)

minio:
  persistence:
    hostPath: /host-data/minio   # k3d node path; bind-mounted from laptop

The chart skips PVC provisioning entirely. The MinIO Deployment mounts /host-data/minio directly. The cogniverse cluster CLI bind-mounts ~/.local/share/cogniverse → /host-data on the k3d node, so MinIO data ends up at ~/.local/share/cogniverse/minio on the laptop fs and survives k3d cluster delete.

Backup destination

The vespa backup CronWorkflow (and any other backup CronWorkflow added later) uploads to an S3-compatible endpoint. One config block switches between dev and cloud — same template, same code path, same workflow.

Enabling backup for local dev

hostStorage.backup.enabled defaults to false — opt in explicitly:

hostStorage:
  backup:
    enabled: true
    bucket: cogniverse-backups
    schedule: "0 3 * * *"
    retainLast: 7
    services:
      - name: vespa
        dataPath: /opt/vespa/var
        podLabel: app.kubernetes.io/component=vespa

No s3 block needed. Backup goes to in-cluster cogniverse-minio. Pair this with minio.persistence.hostPath (above) so the bucket data actually survives cluster destruction.

If services is omitted, the chart's built-in default already backs up both vespa (via kubectl exec + tar) and phoenix nightly — only override services to change which pods are covered or add more. Phoenix's mode follows phoenix.postgres.enabled: postgres while its rows live in the phoenix-postgres database, volume-mount while they live in the SQLite database in the working directory. The postgres dump reads its own archive back and fails the step unless Phoenix's tables and rows are in it, so a wiped, unmigrated or renamed database never publishes a snapshot over the retention window.

values.prod.yaml and values.k3s.yaml both set hostStorage.backup.enabled: true, so each renders cogniverse-backup-vespa and cogniverse-backup-phoenix.

To point at an in-cluster MinIO Deployment whose secret was renamed, set hostStorage.backup.s3.existingSecret. The CronWorkflow uses that secret when configured and otherwise uses the chart's own <release>-minio secret. The top-level hostStorage.backup.existingSecret key is invalid.

The chart's own local-dev overlay, charts/cogniverse/values.k3s.yaml (what cogniverse up deploys with on k3d), already sets hostStorage.enabled: true and hostStorage.backup.enabled: true. It does not set minio.persistence.hostPath, so on that overlay today MinIO falls back to a local-path PVC and the nightly backups do not survive k3d cluster delete — add the minio.persistence.hostPath override from MinIO durability on top of it (e.g. via --set minio.persistence.hostPath=/host-data/minio) to close that gap.

Cloud — Cloudflare R2

hostStorage:
  backup:
    enabled: true
    bucket: my-cogniverse-backups   # bucket pre-created in R2
    schedule: "0 3 * * *"
    retainLast: 30
    services:
      - name: vespa
        dataPath: /opt/vespa/var
        podLabel: app.kubernetes.io/component=vespa
    s3:
      endpoint: "https://<account>.r2.cloudflarestorage.com"
      existingSecret: "cogniverse-r2-creds"
      region: "auto"

Operator pre-creates the credentials secret with both the access key and secret under specific keys (rootUser + rootPassword, regardless of the actual S3 provider — this keeps the chart provider-agnostic):

kubectl -n cogniverse create secret generic cogniverse-r2-creds \
  --from-literal=rootUser=<r2-access-key-id> \
  --from-literal=rootPassword=<r2-secret-access-key>

Cloud — AWS S3

hostStorage:
  backup:
    s3:
      endpoint: "https://s3.us-east-1.amazonaws.com"
      existingSecret: "cogniverse-aws-creds"
      region: "us-east-1"

Same secret shape: rootUser = AWS access key id, rootPassword = AWS secret key. For long-term retention enable bucket-level object lock + versioning on the AWS side; the chart doesn't manage bucket policy.

Cloud — Backblaze B2

hostStorage:
  backup:
    s3:
      endpoint: "https://s3.us-west-002.backblazeb2.com"
      existingSecret: "cogniverse-b2-creds"
      region: "us-west-002"

Recipes by environment

Single-laptop dev (k3d)

Goal: data + backups survive k3d cluster delete.

# charts/cogniverse/values.k3s.yaml already ships hostStorage.enabled=true
# and hostStorage.backup.enabled=true (this is what `cogniverse up` deploys
# on k3d). Layer the minio block below on top — it isn't in that overlay
# yet — to get full k3d-cluster-delete survival:
hostStorage:
  enabled: true                # vespa + phoenix bind-mounted to host
  path: /host-data
  backup:
    enabled: true              # backup CronWorkflow on
    services:
      - name: vespa
        dataPath: /opt/vespa/var
        podLabel: app.kubernetes.io/component=vespa
    # No s3 block → defaults to in-cluster MinIO

minio:
  persistence:
    hostPath: /host-data/minio  # MinIO bucket also on host fs

What survives: - Pod / STS restart: ✓ - Laptop reboot: ✓ - helm uninstall: ✓ (host fs untouched) - k3d cluster delete: ✓ (everything lives in ~/.local/share/cogniverse/) - Laptop nvme failure: ✗ (no offsite copy)

Single-node bare-metal prod

Goal: tolerate node disk failures via offsite backup; node loss = full restore from R2.

# values-singlenode.yaml
hostStorage:
  enabled: false               # use real PVCs, not hostPath
  backup:
    enabled: true
    schedule: "0 */6 * * *"    # every 6h
    retainLast: 30
    services:
      - {name: vespa, dataPath: /opt/vespa/var, podLabel: app.kubernetes.io/component=vespa}
    s3:
      endpoint: "https://<account>.r2.cloudflarestorage.com"
      existingSecret: "cogniverse-r2-creds"
      region: "auto"

vespa: {persistence: {storageClass: "local-path", size: "500Gi"}}
phoenix: {persistence: {storageClass: "local-path", size: "100Gi"}}
minio: {persistence: {storageClass: "local-path", size: "1Ti"}}

Multi-node cloud prod (EKS / GKE / AKS)

Goal: tolerate node loss via replicated CSI; survive AZ loss via offsite backup; survive region loss via cross-region object replication.

# values-cloud.yaml
hostStorage:
  enabled: false
  backup:
    enabled: true
    schedule: "0 */4 * * *"
    retainLast: 90
    services:
      - {name: vespa, dataPath: /opt/vespa/var, podLabel: app.kubernetes.io/component=vespa}
    s3:
      endpoint: "https://s3.us-east-1.amazonaws.com"
      existingSecret: "cogniverse-aws-creds"
      region: "us-east-1"

vespa: {persistence: {storageClass: "gp3", size: "1Ti"}}
phoenix: {persistence: {storageClass: "gp3", size: "200Gi"}}
minio: {persistence: {storageClass: "gp3", size: "5Ti"}}

For prod, additionally wire up Velero at the cluster layer for resource-level backups (Deployments, ConfigMaps, Secrets) — the cogniverse CronWorkflow only covers the application data volumes, not the K8s metadata.

Restore procedure

The dance to recover a stateful component from a MinIO/S3 backup.

Vespa

  1. Take a final pre-restore backup (in case the restore overwrites live state you didn't intend). The backup schedule is a CronWorkflow (cogniverse-backup-vespa) with an inline workflowSpec, not a standalone WorkflowTemplate, so trigger an ad hoc run of it with:

    argo submit --from cronworkflow/cogniverse-backup-vespa -n cogniverse
    

  2. Stop the StatefulSet:

    kubectl -n cogniverse scale sts cogniverse-vespa --replicas=0
    

  3. Mount the volume in a utility pod that has the backup credentials

  4. tar:

    apiVersion: v1
    kind: Pod
    metadata: {name: restore-vespa, namespace: cogniverse}
    spec:
      restartPolicy: Never
      securityContext: {runAsUser: 0}
      containers:
      - name: restore
        image: cogniverse/runtime-rocm:0.1.0-dev
        env:
        - {name: MINIO_ENDPOINT, value: "http://cogniverse-minio:9000"}
        - {name: MINIO_BUCKET, value: "cogniverse-backups"}
        - {name: MINIO_KEY, value: "vespa/vespa-<TIMESTAMP>.tar"}
        - {name: MINIO_ACCESS_KEY, valueFrom: {secretKeyRef: {name: cogniverse-minio, key: rootUser}}}
        - {name: MINIO_SECRET_KEY, valueFrom: {secretKeyRef: {name: cogniverse-minio, key: rootPassword}}}
        command: ["sh","-c"]
        args:
        - |
          set -eu
          python -c "import boto3, os; boto3.client('s3', endpoint_url=os.environ['MINIO_ENDPOINT'], aws_access_key_id=os.environ['MINIO_ACCESS_KEY'], aws_secret_access_key=os.environ['MINIO_SECRET_KEY']).download_file(os.environ['MINIO_BUCKET'], os.environ['MINIO_KEY'], '/tmp/restore.tar')"
          rm -rf /data/* /data/.[!.]* 2>/dev/null || true
          tar -xf /tmp/restore.tar -C /data --strip-components=1
        volumeMounts: [{name: data, mountPath: /data}]
      volumes:
      - name: data
        hostPath: {path: /host-data/vespa, type: DirectoryOrCreate}  # adjust for cloud SC
    
    Apply with kubectl apply -f restore-pod.yaml and wait for Succeeded.

  5. Scale the StatefulSet back up:

    kubectl -n cogniverse scale sts cogniverse-vespa --replicas=1
    kubectl -n cogniverse rollout status sts/cogniverse-vespa
    

  6. Verify by querying the runtime as the rest of the system would:

    kubectl -n cogniverse exec deploy/cogniverse-runtime -c runtime -- \
      curl -s 'http://cogniverse-vespa:8080/search/?yql=select+%2A+from+sources+%2A+where+true&hits=0' \
      | python3 -c 'import json,sys; print(json.load(sys.stdin)["root"]["fields"])'
    

Phoenix

phoenix/phoenix-<TIMESTAMP>.tar holds four members:

Member What
database.dump custom-format pg_dump of the phoenix database — projects, traces, spans, datasets, annotations
database.list that archive's table of contents
restore.env the PGHOST / PGPORT / PGDATABASE / PGUSER it was dumped from
working-assets.tar the PHOENIX_WORKING_DIR tree, members relative to its root
  1. Stop Phoenix so nothing writes while the database is replaced:

    kubectl -n cogniverse scale sts cogniverse-phoenix --replicas=0
    

  2. Unpack the snapshot on the operator's machine:

    mc alias set dest "$MINIO_ENDPOINT" "$MINIO_ACCESS_KEY" "$MINIO_SECRET_KEY"
    mc cp dest/cogniverse-backups/phoenix/phoenix-<TIMESTAMP>.tar .
    tar -xf phoenix-<TIMESTAMP>.tar
    cat restore.env                     # the database this archive came from
    grep ' TABLE DATA public ' database.list | wc -l
    

  3. Restore the database beside the live one, then swap. pg_restore into a populated database leaves the old rows in place, so it goes into a new database that is renamed over the old one:

    kubectl -n cogniverse cp database.dump \
      cogniverse-phoenix-postgres-0:/tmp/database.dump
    kubectl -n cogniverse exec cogniverse-phoenix-postgres-0 -- sh -c '
      set -eu
      psql -U phoenix -d postgres -v ON_ERROR_STOP=1 \
        -c "CREATE DATABASE phoenix_restored"
      pg_restore --exit-on-error --single-transaction --no-owner \
        --no-privileges -U phoenix -d phoenix_restored /tmp/database.dump
      psql -U phoenix -d postgres -v ON_ERROR_STOP=1 \
        -c "ALTER DATABASE phoenix RENAME TO phoenix_prerestore"
      psql -U phoenix -d postgres -v ON_ERROR_STOP=1 \
        -c "ALTER DATABASE phoenix_restored RENAME TO phoenix"'
    
    phoenix_prerestore holds the state the restore replaced; drop it once the verification in step 5 passes.

  4. Restore the working directory into Phoenix's volume with a utility pod shaped like the Vespa one above, mounting Phoenix's storage at /data and unpacking working-assets.tar:

        args:
        - |
          set -eu
          python -c "import boto3, os; boto3.client('s3', endpoint_url=os.environ['MINIO_ENDPOINT'], aws_access_key_id=os.environ['MINIO_ACCESS_KEY'], aws_secret_access_key=os.environ['MINIO_SECRET_KEY']).download_file(os.environ['MINIO_BUCKET'], os.environ['MINIO_KEY'], '/tmp/restore.tar')"
          mkdir -p /tmp/snapshot && tar -xf /tmp/restore.tar -C /tmp/snapshot
          rm -rf /data/* /data/.[!.]* 2>/dev/null || true
          tar -xf /tmp/snapshot/working-assets.tar -C /data
    
    with MINIO_KEY set to phoenix/phoenix-<TIMESTAMP>.tar and the volume hostPath: {path: /host-data/phoenix} on hostStorage clusters or persistentVolumeClaim: {claimName: data-cogniverse-phoenix-0} otherwise.

  5. Start Phoenix and verify the rows came back:

    kubectl -n cogniverse scale sts cogniverse-phoenix --replicas=1
    kubectl -n cogniverse rollout status sts/cogniverse-phoenix
    kubectl -n cogniverse exec cogniverse-phoenix-postgres-0 -- \
      psql -U phoenix -d phoenix -At -c \
      'SELECT (SELECT count(*) FROM projects), (SELECT count(*) FROM spans),
              (SELECT count(*) FROM datasets), (SELECT count(*) FROM span_annotations)'
    

With phoenix.postgres.enabled=false the snapshot is a single tar of the working directory holding Phoenix's SQLite database; restore it with steps 1, 4 and 5, unpacking the snapshot itself rather than an inner working-assets.tar.

Backup verification (the "0" in 3-2-1-1-0)

A backup that's never been restored is a hypothesis. Cogniverse exercises restore through normal cluster lifecycle ops:

  • Every helm upgrade that touches a stateful component re-mounts the same hostPath / PVC and the data must come back intact. If it doesn't, the upgrade fails fast.
  • Every cogniverse up / k3d cluster delete cycle is implicit hostStorage durability proof — pods restart against the same ~/.local/share/cogniverse/{vespa,phoenix,minio} directories.
  • Periodic explicit restore drill (recommended quarterly): follow the Restore procedure above using a recent backup tarball, pointed at a fresh sandbox cluster. Verify the doc count + a known query match production.

For cloud deployments where the backup target is offsite (R2 / B2 / AWS S3), this drill also exercises the network path + credentials — catching expired secrets, IAM mis-grants, region misconfigs that silently rot if never tested.

Why these defaults

Drawing on:

  • 3-2-1-1-0 rule: 3 copies, 2 media, 1 offsite, 1 immutable, 0 verification errors (SNIA cloud-native interpretation)
  • Failure-domain separation: backup must not share power, network, hypervisor, or credentials with primary
  • Vespa data management (docs): Vespa OSS has no built-in snapshot — the chart's kubectl exec tar approach is what they recommend for self-hosted operators (with the caveat that it's not crash-consistent)
  • Velero best practices: backup target must be S3-compatible object storage, never the source cluster's storage
  • MinIO production guide: distributed mode (4+ nodes, erasure coding) is the only production-grade self-hosted MinIO. Single-node MinIO is staging / cache / dev — exactly how cogniverse uses it.

Out of scope (track separately)

  • Velero integration for K8s resource backup (Deployments, ConfigMaps). The CronWorkflow we ship only covers application data volumes; for full cluster recovery use Velero alongside.
  • CSI VolumeSnapshots for crash-consistent snapshots. Requires a real CSI provider (Longhorn, cloud SC). The current tar approach reads the live filesystem while the source process keeps writing.
  • HF model cache backup. Explicitly NOT backed up — it's rebuildable from HF Hub via the populate Job (hfCache.persistence.minio.models).