Skip to content

Kubernetes Deployment Guide


Overview

Cogniverse provides production-ready Helm charts for Kubernetes deployment with:

  • Helm Chart: charts/cogniverse/ (plus the bundled openshell subchart at charts/cogniverse/charts/openshell)

  • StatefulSets (persistent storage): Vespa, Phoenix, and LLM — only when llm.engine: ollama (the chart default); llm.engine: vllm deploys the LLM as a Deployment instead (see LLM serving below)

  • Deployments: Runtime (with an optional in-pod quality-monitor sidecar container), Web client, Ingestor workers (dequeue ingestion jobs from Redis), MinIO, Redis, Semantic Router (Envoy + vllm-sr router), Messaging Gateway (Telegram/Slack, disabled by default), LLM (only when llm.engine: vllm), and one Deployment per entry under inference.* — vllm_colpali (ColQwen3 token-embed), colbert_pylate, code_colbert_pylate, denseon (dense single-vector embeddings), gliner (zero-shot NER), clap_embed (audio embeddings), face_embed, vllm_asr (transcription), vllm_llm_student, vllm_llm_teacher, and video_embed (remote sidecar)

  • LLM serving: llm.engine: ollama is the chart default (deploys Ollama as a StatefulSet). Set llm.engine: vllm to deploy a built-in vLLM Deployment instead (the ROCm GPU overlay, values.rocm.yaml, switches to vllm by default) — or llm.engine: external to point at an existing endpoint with no pod. Configure once via llm.builtin.enabled and llm.engine in your values file.

See models-and-inference.md for the canonical list of every model, image source, and deployment style (custom-built sidecars vs official vLLM/Ollama images, CPU vs ROCm, student vs teacher LLM).

  • Argo Workflows: chart-managed WorkflowTemplates (job-runner, optimization-runner) plus scheduled CronWorkflows for optimization, distillation, synthetic-data generation, monthly reports, and backup/cleanup maintenance — see argo-workflows.md

  • Auto-scaling: HPA for Runtime

  • Ingress: NGINX (Traefik on K3s) with TLS/SSL

  • Init Jobs: Schema deployment, model pulling — plus the optional HF-cache MinIO-populate Job and openshell mTLS-cert Job when their features are enabled

  • Networking: optional cluster-wide NetworkPolicy (networkPolicy.enabled) and per-agent egress allow-listing (networkPolicy.agentEgress.enabled)


Prerequisites

Cluster Requirements

Minimum:

  • Kubernetes 1.24+

  • 3 nodes (1 master, 2 workers)

  • 32GB RAM per node

  • 100GB+ storage per node

Recommended:

  • Kubernetes 1.27+

  • 5+ nodes

  • 64GB RAM per node

  • GPU nodes for vLLM inference sidecars (or for Ollama if used)

  • NVMe/SSD storage

Required Tools

# Helm 3.x
helm version

# kubectl
kubectl version --client

# Optional: K3s (lightweight Kubernetes)
curl -sfL https://get.k3s.io | sh -

Storage Class

# Check available storage classes
kubectl get storageclass

# For K3s, use local-path (default)
# For EKS, use gp3
# For GKE, use standard-rwo

Quick Start

1. Add Helm Repository (if published)

# Add Cogniverse Helm repo
helm repo add cogniverse https://charts.cogniverse.ai
helm repo update

# Or use local charts
cd cogniverse/charts

2. Install with Default Values

# Create namespace
kubectl create namespace cogniverse

# Install chart
helm install cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --create-namespace

# Check status
helm status cogniverse -n cogniverse
kubectl get pods -n cogniverse

3. Access Services

# Port-forward Runtime API
kubectl port-forward -n cogniverse svc/cogniverse-runtime 8000:8000

# Port-forward the web client
kubectl port-forward -n cogniverse svc/cogniverse-web 4000:4000

# Access
open http://localhost:8000/docs
open http://localhost:4000

Production Deployment

Create Production Values

Create values.prod.yaml:

# Production configuration
image:
  tag: "2.0.0"
  pullPolicy: IfNotPresent

global:
  imageRegistry: "your-registry.io"
  imagePullSecrets:
    - name: regcred

# Vespa configuration
vespa:
  replicaCount: 3
  persistence:
    enabled: true
    storageClass: "fast-ssd"
    size: "200Gi"
  resources:
    requests:
      cpu: "4"
      memory: "16Gi"
    limits:
      cpu: "8"
      memory: "32Gi"

# Runtime auto-scaling
runtime:
  replicaCount: 3
  autoscaling:
    enabled: true
    minReplicas: 3
    maxReplicas: 20
    targetCPUUtilizationPercentage: 70

# Ingress with SSL
ingress:
  enabled: true
  className: "nginx"
  annotations:
    cert-manager.io/cluster-issuer: "letsencrypt-prod"
  hosts:
    - host: cogniverse.your-domain.com
      paths:
        - path: /api
          pathType: Prefix
          service: runtime
          port: 8000
        - path: /
          pathType: Prefix
          service: web
          port: 4000
  tls:
    - secretName: cogniverse-tls
      hosts:
        - cogniverse.your-domain.com

# GPU configuration for LLM (nodeSelector, tolerations, and the
# nvidia.com/gpu or amd.com/gpu resource request are all wired
# automatically off llm.device — no manual resource block needed)
llm:
  engine: vllm
  device: cuda
  gpuCount: 1

Deploy to Production

# Deploy with production values
helm install cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --create-namespace \
  --values values.prod.yaml

# Verify deployment
kubectl get all -n cogniverse

Rolling Out a New Version

helm upgrade cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --values values.prod.yaml

Four behaviors to know during a rollout — each is deliberate and easy to misread as a failure:

Preflight: confirm the cluster has no orphan schemas.

curl -sfX POST "$RUNTIME_URL/admin/reconcile-orphans?dry_run=true" | jq .

All five response lists (orphan_schemas, orphan_tenants, unrecovered_schemas, tenant_orphan_schemas, tenant_orphan_tenants) empty means nothing to do. Names under tenant_orphan_schemas are carried by every deploy; clear them with ?dry_run=false&remove_tenant_orphans=true. Any name in unrecovered_schemas must be resolved before the next tenant delete: the schema redeploy refuses (raises rather than drops) while a deployed schema cannot be confirmed as an orphan, so one unattributable schema blocks every tenant delete and reconcile until an operator resolves it. See Orphan reconciliation.

Runtime pods legitimately sit 0/1 Ready for up to ~20 minutes. The FastAPI lifespan waits for Vespa, bootstraps the metadata schemas, and blocks on deploy convergence before binding port 8000 — worst case ~810s on a loaded cluster. The startupProbe budget (30s initial delay + 80 × 15s = 1230s) covers this, and the liveness probe only begins once startup succeeds. Running + 0/1 Ready inside that window is normal; deleting the pod to "unstick" it restarts convergence from zero. Investigate the pod logs only once the budget is exhausted and the probe itself has restarted the pod.

Tenant schemas move to the release's definitions after startup. Once a runtime pod has started, it redeploys every tenant schema registered with a definition other than the one the release ships in configs/schemas/, one application package per tenant, while it already serves. A change Vespa refuses without a validation override stays unapplied and is reported. Once the rollout has replaced every pod, an empty drifted list means every tenant runs the shipped definitions:

curl -s "$RUNTIME_URL/admin/schemas/drift" | jq .

See Schema changes in a release.

Worker restarts spend a job's redelivery budget. Killing an ingestion worker mid-job — which every rollout does — strands the job until the reaper reclaims it (INGEST_REAPER_MIN_IDLE_MS, default 5 min) and redelivers it to a live worker, incrementing a per-job delivery counter that never resets. Past INGEST_REAPER_MAX_DELIVERIES (default 5) the job is presumed poison and moved to the ingest:queue:dead stream with a failed terminal event. One rollout costs one redelivery; repeated restarts while long jobs are in flight can dead-letter healthy jobs. After repeated restarts, check the dead stream:

kubectl exec -n cogniverse deploy/cogniverse-redis -- \
  redis-cli XRANGE ingest:queue:dead - +

Dead-lettering clears the job's in-flight idempotency marker and writes no done marker, so re-submitting the same source through POST /ingestion/upload re-enqueues it as a fresh job. See Troubleshooting → Deployment and Rollout.


K3s Local Deployment

Overview

K3s is a lightweight Kubernetes distribution perfect for local development, testing, and edge deployments. It:

  • Runs on minimal resources (single node with 4GB RAM)

  • Includes built-in storage (local-path provisioner)

  • Has Traefik ingress controller by default

  • Supports full Kubernetes APIs

Use Cases:

  • Local development and testing

  • CI/CD pipelines

  • Edge deployments

  • Learning Kubernetes

Quick Start with K3s

The easiest way to deploy locally is using the CLI:

# Start all services via k3d/Helm
cogniverse up

# Check status
cogniverse status

Argo Workflows on K3s:

Argo Workflows works perfectly on K3s for local testing of batch processing workflows:

# Deploy with Argo Workflows
cogniverse up

# Access Argo UI locally
kubectl port-forward -n cogniverse svc/cogniverse-argo-workflows-server 2746:2746
open http://localhost:2746

# Submit a workflow (e.g., video ingestion)
argo submit workflows/video-ingestion.yaml \
  -n cogniverse \
  --parameter video-dir="/data/videos" \
  --parameter tenant-id="default" \
  --parameter profiles="video_colpali_smol500_mv_frame"

# Watch workflow progress
argo watch <workflow-name> -n cogniverse

# List all workflows
argo list -n cogniverse

This gives you the full Kubernetes + Argo experience locally without needing a cloud cluster!

Manual K3s Installation

If you prefer manual setup:

# Install K3s
curl -sfL https://get.k3s.io | sh -s - \
  --write-kubeconfig-mode 644 \
  --disable traefik  # Optional: disable if using custom ingress

# Wait for K3s to be ready
sudo k3s kubectl get nodes

# Setup kubeconfig
mkdir -p $HOME/.kube
sudo cp /etc/rancher/k3s/k3s.yaml $HOME/.kube/config
sudo chown $(id -u):$(id -g) $HOME/.kube/config

# Verify
kubectl get nodes

K3s-Specific Configuration

The chart already ships charts/cogniverse/values.k3s.yaml, which cogniverse up applies automatically — you don't need to write your own for the dev CLI path. The example below shows the structure to follow if you're deploying to a standalone K3s cluster by hand (helm install without cogniverse up) and want your own values.k3s.yaml overrides:

# K3s-specific Helm values
# Optimized for local development

# Use local-path storage (K3s default)
global:
  storageClass: "local-path"

# Reduced resources for single-node deployment
vespa:
  replicaCount: 1
  persistence:
    enabled: true
    storageClass: "local-path"
    size: "20Gi"
  resources:
    requests:
      cpu: "1"
      memory: "4Gi"
    limits:
      cpu: "2"
      memory: "8Gi"

runtime:
  replicaCount: 1
  autoscaling:
    enabled: false  # Disable for single node
  resources:
    requests:
      cpu: "500m"
      memory: "2Gi"
    limits:
      cpu: "1"
      memory: "4Gi"

web:
  replicaCount: 1
  resources:
    requests:
      cpu: "100m"
      memory: "512Mi"
    limits:
      cpu: "1"
      memory: "512Mi"

phoenix:
  replicaCount: 1
  # The Phoenix UI address a browser reaches (the runtime's PHOENIX_UI_URL);
  # the web client's trace and dataset links point at it.
  uiUrl: "http://localhost:26006"
  persistence:
    enabled: true
    storageClass: "local-path"
    size: "10Gi"
  resources:
    requests:
      cpu: "500m"
      memory: "1Gi"
    limits:
      cpu: "1"
      memory: "2Gi"

llm:
  engine: ollama
  device: cpu
  nodeSelector: {}
  tolerations: []
  ollama:
    persistence:
      enabled: true
      storageClass: "local-path"
      size: "20Gi"
    resources:
      requests:
        cpu: "1"
        memory: "4Gi"
      limits:
        cpu: "2"
        memory: "8Gi"
    models:
      - "gemma3:4b"

# Ingress with Traefik (K3s default)
ingress:
  enabled: true
  className: "traefik"
  hosts:
    - host: cogniverse.local
      paths:
        - path: /api
          pathType: Prefix
          service: runtime
          port: 8000
        - path: /
          pathType: Prefix
          service: web
          port: 4000
  tls: []  # No TLS for local development

# Local development tenant
config:
  tenants:
    - id: "default"
      name: "Default Tenant"
  llmModels:
    - "gemma3:4b"

# Enable init jobs
initJobs:
  schemaDeployment:
    enabled: true
  modelPulling:
    enabled: true

Deploy to K3s

# Create namespace
kubectl create namespace cogniverse

# Deploy with K3s values (CPU host)
helm install cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --values values.k3s.yaml \
  --wait \
  --timeout 10m

# Deploy with K3s + ROCm overlay (AMD GPU host)
helm install cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --values values.k3s.yaml \
  --values values.rocm.yaml \
  --wait \
  --timeout 10m

# Check status
kubectl get pods -n cogniverse -w

cogniverse up (the dev CLI) composes these layers automatically when it detects a GPU host — detect_torch_backend() (libs/cli/cogniverse_cli/images.py) picks values.rocm.yaml or values.cuda.yaml to overlay on values.k3s.yaml, and labels the k3d node accordingly (see the up command in libs/cli/cogniverse_cli/main.py). A CPU-only host gets no overlay — the CPU defaults already baked into values.yaml apply (see also values.cpu.yaml for the standalone CPU-override profile used outside cogniverse up).

GPU passthrough (ROCm + CUDA)

ROCm (AMD): GPU access is via hostPath volume mounts of /dev/kfd and /dev/dri into the pods, not the legacy k8s device plugin. This means:

  • The k3d cluster must be created with the host devices bind-mounted. cogniverse up does this automatically when /dev/kfd is detected on the host:
k3d cluster create cogniverse \
  --volume /dev/kfd:/dev/kfd@server:0 \
  --volume /dev/dri:/dev/dri@server:0 \
  ...
  • The k3d node must carry the label amd.com/gpu.present=true so the chart's nodeSelector schedules vLLM pods. cogniverse up applies this via kubectl label. For manual helm install, apply it once: kubectl label node --all amd.com/gpu.present=true --overwrite.

  • Pods need supplementalGroups for the host's render and video group ids (default 992 and 44 on Debian/Ubuntu). Override per service for distros with different ids:

inference:
  vllm_colpali:
    rocm:
      supplementalGroups: [109, 18]   # Fedora render=109, video=18

The chart resolves supplementalGroups via dig() so an absent inference.<svc>.rocm block falls through to the default [992, 44] rather than nil-derefing.

CUDA (NVIDIA): still uses nvidia.com/gpu resource requests via the NVIDIA device plugin. Apply the nvidia.com/gpu.present=true label on the node and set GPU resource requests on the relevant pods in values.cuda.yaml.

Unified-pool memory budget

On an APU host the GPU pool is carved out of system RAM rather than being separate memory. On the ROCm reference host:

Measurement Source Value
System RAM /proc/meminfo MemTotal 123.46 GiB
GPU pool mem_info_gtt_total 96 GiB
Dedicated VRAM carve-out mem_info_vram_total 2 GiB

Each vLLM service's --gpu-memory-utilization is a fraction of the 96 GiB pool, and every GiB it claims is a GiB the desktop, the CPU-side cluster services and the test containers no longer have. The fractions must sum to well under 1.0. Summing them to ~1.0 leaves the host with no eviction headroom: allocation pressure then drives svm_range_restore churn that starves the compositor.

The budget after reserving 77.5 GiB for the non-GPU workloads — cluster services (31.5 GiB of memory limits), desktop and daemons, the e2e suite's own containers, the pylate pods that allocate from the pool without declaring a fraction, GLiNER, and page-cache slack — leaves 45.96 GiB, so the enabled fractions are capped at 0.47 (45.12 GiB). The deployed configuration, with both chat models on Modal, sums to 0.27.

Size each fraction from need, not from habit: the weights at their served precision plus a KV allowance for the configured --max-model-len x --max-num-seqs. A fraction above that need does not make the model faster, it only denies the memory to everything else. The Tomoro pooling encoder pins its cache in bytes, so its fraction is vLLM's startup free-memory guard, sized for the weights, that cache and one step's activations, which --max-num-batched-tokens bounds.

tests/charts/test_gpu_memory_budget.py renders the overlay, sums the enabled fractions and fails when they exceed the cap. It also fails if a service's rendered form escapes its parser, so a service cannot drop out of the budget silently.

With the student and the distillation teacher both resident the fractions sum to 0.69, over the cap, which is why this host serves them from Modal (values.modal-llm.yaml).

Pods that allocate without a fraction

--gpu-memory-utilization is a vLLM flag. The pylate services (colbert_pylate, code_colbert_pylate) serve through libs/cli/cogniverse_cli/modal_inference/servers/pylate.py, which is plain PyTorch: the chart passes them DEVICE=cuda (torch's namespace, HIP on ROCm), and allocation is whatever the caching allocator grows to. There is no fraction knob, and the served encode call passes no batch_size while the request model accepts an unbounded input list, so peak allocation scales with the caller's payload rather than with a configured ceiling.

Because peak allocation follows the request, the request is bounded. Each pylate pod carries MAX_INPUT_ITEMS, MAX_INPUT_CHARS and ENCODE_BATCH_SIZE, set per service under inference.<svc>.requestLimits and defaulted by the chart:

Bound Default Basis
maxInputItems 256 eight times the 32-text chunk the canonical client sends
maxInputChars 2000000 covers the few-items-but-enormous payload shape
encodeBatchSize 32 sentence-transformers' own default, pinned explicitly

Over-limit requests are rejected with 413 naming the limit and the received size, before the model loads or encodes. They are never truncated to fit: a truncated encode would return embeddings for a subset under a success status, which the caller cannot distinguish from a complete result.

The declared memory bound remains the pod memory limit (4Gi each), reserved in the budget above. Note that a Kubernetes memory limit constrains the container's cgroup; pool allocations made through the amdgpu driver are not necessarily charged there, so treat that limit as a budgeting declaration rather than an enforced ceiling.

A pod reaches the pool exactly when the chart mounts /dev/kfd and /dev/dri into it. The budget test uses that as its marker, so any service rendered with GPU access is picked up automatically, and it fails when such a pod declares neither a fraction nor a memory limit. gliner renders without those mounts and is CPU-only.

First-deploy validation

Nothing above is confirmed by rendering alone. On the first deploy of these values, check in this order and stop at the first failure:

  1. vllm_colpali boots at 0.18. Highest risk: the prior 0.45 was chosen because vLLM's startup free-memory guard rejects this multimodal model when the fraction is too low for its transient startup profile. Failure signal: the pod never reaches ready and its log carries a vLLM memory-profiling error rather than a served-model line. Remedy is to raise this one fraction until it boots, then re-check the budget total.
  2. vllm_llm_student loads at 0.22. Its ~16 GiB weight figure is inferred, not measured. Failure signal: an out-of-memory abort during weight load, or a KV-cache-too-small startup error.
  3. vllm_asr loads at 0.04 and serves a transcription.
  4. Residency stays under the ceiling. Read /sys/class/drm/card1/device/mem_info_gtt_used with all models ready and again under sweep traffic. Expect it below the 51.84 GiB cap; sustained readings above it mean a fraction is under-declared relative to what the model actually takes.
  5. The startup chain completes inside its deadlines. Each gated pod should leave Init when its predecessor turns ready, not when its pacing deadline elapses. Failure signal: a pod's startup-gate log ends with the deadline line rather than the serving line, which means the chain is pacing on timeouts instead of on readiness.
  6. The host stays responsive during a full sweep — no svm_range_restore workers saturating CPU.
  7. The pylate request bounds hold in practice. Rendering proves the pods carry the limits; only a run shows what they peak at. With ingest running, watch each pylate pod against its 4Gi limit and watch mem_info_gtt_used while it encodes. Failure signal: pool usage climbing with ingest batch size, which means the request bounds are not translating into a memory ceiling and the per-process fraction below is needed.
  8. Whether set_per_process_memory_fraction is warranted. Read torch.cuda.get_device_properties(0).total_memory from inside a pylate pod. That value is the denominator any fraction would use, and it decides whether the knob can express a 4Gi ceiling at all.

Model startup pacing

Every inference pod is created at once, so by default every model reserves its share of the GPU pool simultaneously. Where the GPU pool is carved out of system RAM, the enabled models' --gpu-memory-utilization fractions can sum past 1.0 and that simultaneous reservation drives the kernel into an eviction/restore cycle that starves the rest of the host.

inferenceStartup.sequence is an ordered list of inference keys. The pod at position N runs a startup-gate init container that waits for position N-1 to answer /health before the model process starts, so GPU memory is reserved one model at a time:

inferenceStartup:
  sequence:
    - vllm_llm_teacher
    - vllm_colpali
    - vllm_llm_student
  perLinkTimeoutSeconds: 600
  progressAllowanceSeconds: 900
  • The first entry that runs a pod and any key absent from the list load immediately. An empty list disables pacing, which is the right setting for a discrete-GPU host with headroom.
  • Order the list by descending --gpu-memory-utilization so the largest allocation lands against an unfragmented pool.
  • Each pod waits on the nearest earlier entry that runs a pod in the release. Entries that are disabled or served from an externalUrl are skipped, never waited on, and never leave their successor ungated.
  • Waiting happens in an init container, so readiness and liveness — both measured from the main container's start — are unaffected, and a waiting pod never reports unhealthy.
  • The gate polls /health every 5 s with curl from curlimages/curl:8.14.1, the image the chart's init Jobs use, and starts the model on an HTTP 200. No cogniverse image is involved, so a runtime release leaves every gated model pod's template unchanged and its model loaded. Air-gapped clusters mirror the image with the rest of the third-party set.
  • Weight downloads run in the model-warm init container ahead of the gate; they touch network and disk rather than the GPU, so they still run concurrently across pods. model-warm runs from the pod's own server image, except with the MinIO mirror (hfCache.persistence.minio.enabled), where it needs boto3 and runs from the runtime image; only then does a runtime release change, and restart, the inference pods.
  • Position N gives up after N x perLinkTimeoutSeconds. The deadline scales with position because all init containers start together, so a flat deadline would release the whole tail of the chain at once. A gated pod's rollout progress deadline is extended past its own wait budget so a deliberately waiting pod is not reported as a stalled rollout.

Local Access Setup

For domain-based access (optional):

# Add entry to /etc/hosts
echo "127.0.0.1 cogniverse.local" | sudo tee -a /etc/hosts

# Access via domain
open http://cogniverse.local
open http://cogniverse.local/api/docs

Or use port-forwarding:

# Runtime API
kubectl port-forward -n cogniverse svc/cogniverse-runtime 8000:8000
open http://localhost:8000/docs

# Web client
kubectl port-forward -n cogniverse svc/cogniverse-web 4000:4000
open http://localhost:4000

# Phoenix
kubectl port-forward -n cogniverse svc/cogniverse-phoenix 6006:6006
open http://localhost:6006

K3s Operations

View K3s logs:

sudo journalctl -u k3s -f

Restart K3s:

sudo systemctl restart k3s

Stop K3s:

sudo systemctl stop k3s

Uninstall K3s:

# Uninstall Cogniverse first
helm uninstall cogniverse -n cogniverse
kubectl delete namespace cogniverse

# Uninstall K3s
/usr/local/bin/k3s-uninstall.sh

K3s Troubleshooting

Issue: Pods stuck in Pending

# Check node resources
kubectl describe nodes

# K3s typically runs on single node, check if resources exhausted
kubectl top nodes
kubectl top pods -A

# Consider reducing resource requests in values.k3s.yaml

Issue: Storage provisioning fails

# Check local-path provisioner
kubectl get pods -n kube-system | grep local-path

# Check storage class
kubectl get storageclass

# Verify PVC
kubectl get pvc -n cogniverse
kubectl describe pvc <pvc-name> -n cogniverse

# Local-path uses /var/lib/rancher/k3s/storage
sudo ls -la /var/lib/rancher/k3s/storage/

Issue: Cannot connect to cluster

# Check K3s service
sudo systemctl status k3s

# Check kubeconfig
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
sudo kubectl get nodes

# Or copy to ~/.kube/config
sudo cp /etc/rancher/k3s/k3s.yaml ~/.kube/config
sudo chown $(id -u):$(id -g) ~/.kube/config

Issue: Ingress not working

# Check Traefik (K3s default ingress)
kubectl get pods -n kube-system | grep traefik

# View Traefik logs
kubectl logs -n kube-system -l app.kubernetes.io/name=traefik

# Check ingress resources
kubectl get ingress -n cogniverse
kubectl describe ingress cogniverse -n cogniverse

K3s vs Full Kubernetes

Feature K3s Full Kubernetes
Resource Usage 512MB RAM minimum 2GB+ RAM minimum
Installation Single binary Multi-component
Storage local-path (built-in) Requires CSI driver
Ingress Traefik (built-in) Requires installation
Use Case Development, Edge Production, Scale
Node Requirement 1+ nodes 3+ nodes (HA)

When to Use K3s

✅ Good for:

  • Local development and testing

  • CI/CD testing pipelines

  • Learning Kubernetes concepts

  • Edge deployments

  • Small-scale production (single tenant)

❌ Not ideal for:

  • Large-scale production (100+ tenants)

  • High-availability requirements

  • Heavy resource workloads

  • Multi-region deployments


Configuration

Multi-Tenant Setup

Update values.yaml:

config:
  tenants:
    - id: "acme_corp"
      name: "Acme Corporation"
    - id: "globex_inc"
      name: "Globex Inc"

# Enable multi-tenant resource quotas
multiTenant:
  enabled: true
  quotas:
    cpu: "10"
    memory: "20Gi"
    storage: "100Gi"

GPU Configuration

For a GPU-backed LLM pod, set llm.device — the chart wires the matching nodeSelector, tolerations, and nvidia.com/gpu / amd.com/gpu resource request automatically:

llm:
  engine: ollama          # or vllm
  device: cuda            # cpu | cuda | rocm | mps (vllm only)
  gpuCount: 1
  ollama:
    resources:
      limits:
        memory: "16Gi"
      requests:
        cpu: "4"
        memory: "8Gi"

Persistent Storage

See persistence-and-backup.md for the full durability matrix, per-component storage knobs, the dev/cloud config recipes, the backup destination switch (in-cluster MinIO vs Cloudflare R2 / Backblaze B2 / AWS S3), and the restore procedure.

Quick reference:

# Cloud / multi-node prod — replicated CSI for primary, S3 for backup
vespa:    {persistence: {storageClass: "gp3", size: "1Ti"}}
phoenix:  {persistence: {storageClass: "gp3", size: "200Gi"}}
minio:    {persistence: {storageClass: "gp3", size: "5Ti"}}
hfCache:  {persistence: {enabled: true, storageClass: "gp3", size: "100Gi"}}
hostStorage:
  backup:
    enabled: true
    s3:
      endpoint: "https://s3.us-east-1.amazonaws.com"
      existingSecret: "cogniverse-aws-creds"
# Single-laptop dev — hostStorage everywhere, in-cluster MinIO is the
# backup target. Survives ``k3d cluster delete``.
hostStorage:
  enabled: true
  backup: {enabled: true}
minio:
  persistence:
    hostPath: /host-data/minio

Operations

Upgrade Deployment

# Update values
nano values.prod.yaml

# Upgrade with new values
helm upgrade cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --values values.prod.yaml \
  --reuse-values

# Check upgrade status
helm history cogniverse -n cogniverse

Rollback

# List releases
helm history cogniverse -n cogniverse

# Rollback to previous version
helm rollback cogniverse -n cogniverse

# Rollback to specific revision
helm rollback cogniverse 3 -n cogniverse

Scaling

# Manual scaling
kubectl scale deployment cogniverse-runtime \
  --replicas=5 \
  -n cogniverse

# Or update via Helm
helm upgrade cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --set runtime.replicaCount=5 \
  --reuse-values

# uvicorn worker processes per runtime pod; each is a full copy of the
# runtime, so runtime.resources memory must hold all of them
helm upgrade cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --set runtime.workers=2 \
  --reuse-values

Every replica and every worker holds its own in-process state; the runtime module guide's Deployment section lists what a follow-up request on another process does not see.

Every replica and worker shares the one Redis instance the chart deploys (redis.replicaCount: 1). It must stay a single Redis, not Redis Cluster: the runtime's conversation ledger (cogniverse_runtime/session_state.py) updates a context's turn clock, its pending saves and the shared lost-turn record in one Lua script, and Redis Cluster refuses a script whose keys span hash slots.

Each runtime worker reaches its shared and session state (agent registrations, annotations, /ingestion/start jobs, conversation order, /v1 continuations, task events) through one client and connection pool of at most 128 connections, named cogniverse-runtime-state:<pod>:<pid>:<suffix>; the A2A task store and the cluster-events channel hold their own connections. To count a pod's state connections:

kubectl -n cogniverse exec deploy/cogniverse-redis -- redis-cli CLIENT LIST \
  | grep -o 'name=cogniverse-runtime-state:[^ ]*' | sort | uniq -c

Each worker shows one name. Its count is one after startup and grows to the most commands the worker has had in flight at once (pooled connections stay open), never past 128.

Backup & Restore

Backup:

# Backup Helm values
helm get values cogniverse -n cogniverse > backup-values.yaml

# Backup PVCs
kubectl get pvc -n cogniverse -o yaml > backup-pvcs.yaml

# Snapshot volumes (cloud-specific)
# AWS EBS: Use EBS snapshots
# GCP: Use persistent disk snapshots

Restore:

# Restore from backup
helm install cogniverse ./charts/cogniverse \
  --namespace cogniverse \
  --values backup-values.yaml


Monitoring

Check Pod Status

# Get all pods
kubectl get pods -n cogniverse -o wide

# Describe pod
kubectl describe pod cogniverse-runtime-xxx -n cogniverse

# View logs
kubectl logs -f cogniverse-runtime-xxx -n cogniverse

# Previous logs
kubectl logs --previous cogniverse-runtime-xxx -n cogniverse

Resource Usage

# Node resources
kubectl top nodes

# Pod resources
kubectl top pods -n cogniverse

# Detailed pod metrics
kubectl describe node <node-name>

Health Checks

# Check all services
kubectl get svc -n cogniverse

# Test health endpoints
kubectl run curl --image=curlimages/curl -i --rm --restart=Never -- \
  curl http://cogniverse-runtime:8000/health

# Check HPA status
kubectl get hpa -n cogniverse

Troubleshooting

Pods Not Starting

# Check pod events
kubectl describe pod <pod-name> -n cogniverse

# Common issues:
# 1. Image pull errors
kubectl get events -n cogniverse --sort-by='.lastTimestamp'

# 2. Resource constraints
kubectl describe nodes

# 3. Storage issues
kubectl get pvc -n cogniverse
kubectl describe pvc <pvc-name> -n cogniverse

Init Jobs Failing

# Check init job status
kubectl get jobs -n cogniverse

# View init job logs
kubectl logs job/cogniverse-schema-deployment -n cogniverse

# Delete and re-run
kubectl delete job cogniverse-schema-deployment -n cogniverse
helm upgrade cogniverse ./charts/cogniverse -n cogniverse --reuse-values

For each config.tenants entry the schema-deployment Job first registers the tenant (POST /admin/tenants with created_by: "helm:cogniverse", which also deploys the tenant's base schemas), then deploys config.defaultProfiles.video for it when one is selected. A create answering 200 or 201 logs Registered tenant <id>, a 409 logs Tenant <id> is already registered and both proceed; any other status logs Response (<status>): <body> and Tenant registration failed for tenant <id> and fails the attempt. The chart's tenants are therefore registered tenants: the web client can act for them and uploads and queries for them are accepted.

Each runtime call in the schema-deployment Job allows --max-time 580 seconds: one attempt (the deploy lease wait, one Vespa prepareandactivate request, the convergence wait and a margin); a tenant create activates its base schemas in one such deploy. On any deploy status other than success or already_deployed the Job fails and Kubernetes retries it up to backoffLimit; a retry that meets a deploy still holding the lease answers failed within the lease wait, and ingestion deploys a missing schema on first use. activeDeadlineSeconds caps the whole Job at the runtime's startup-probe budget plus (backoffLimit + 1) × tenants × 2 × 580 seconds (a create and a deploy per tenant) plus 300 seconds of retry delay per retry (with restartPolicy: OnFailure a retry restarts the container in place, after the kubelet's crash-loop backoff, capped at five minutes). Helm waits up to its --timeout (the CLI passes 10m) for the hook: longer than one attempt, shorter than the Job's deadline, so a Job still retrying fails helm install/upgrade (and cogniverse up) while it keeps running.

Service Connection Issues

# Test service connectivity
kubectl run test-pod --image=busybox -i --rm --restart=Never -- \
  wget -O- http://cogniverse-vespa:8080/ApplicationStatus

# Check service endpoints
kubectl get endpoints -n cogniverse

# Verify DNS
kubectl run test-dns --image=busybox -i --rm --restart=Never -- \
  nslookup cogniverse-vespa.cogniverse.svc.cluster.local

Best Practices

Resource Management

  1. Set resource requests and limits

    resources:
      requests:
        cpu: "2"
        memory: "4Gi"
      limits:
        cpu: "4"
        memory: "8Gi"
    

  2. Use Pod Disruption Budgets

    apiVersion: policy/v1
    kind: PodDisruptionBudget
    metadata:
      name: runtime-pdb
    spec:
      minAvailable: 1
      selector:
        matchLabels:
          app.kubernetes.io/component: runtime
    

Security

  1. Use Network Policies

    networkPolicy:
      enabled: true
    

  2. Enable RBAC

    rbac:
      create: true
    

  3. Use Secrets for sensitive data

    kubectl create secret generic cogniverse-secrets \
      --from-literal=api-key=xxx \
      -n cogniverse
    

High Availability

  1. Multi-replica deployments
  2. Pod anti-affinity rules
  3. Health checks configured
  4. Auto-scaling enabled
  5. PVC backup strategy

Deployment Scripts

For automated deployment, use the CLI:

# Start all services via k3d/Helm
cogniverse up

# Check status
cogniverse status

See cogniverse --help for full options.