Deployment Guide¶
Overview¶
This guide covers verified, implemented deployment patterns:
-
Local Development: Docker-based setup for development
-
Modal Serverless: GPU-accelerated VLM frame description generation
-
Multi-Tenant: Schema deployment and isolation
Core Services¶
- Vespa: Multi-tenant vector database (ports 8080, 19071)
- Phoenix: Telemetry dashboard + OTLP collector (ports 6006, 4317)
- Ollama: Local LLM inference (port 11434)
Runtime and Agents¶
The deployed system runs a single Runtime process (port 8000) that exposes REST + an in-process A2A JSON-RPC server (/a2a). All 23 agent classes registered in configs/config.json agents.* (14 enabled by default, 9 knowledge/federation agents shipped disabled — see the full roster table below) are DSPy-based classes registered with an in-process AgentRegistry and dispatched directly by the Runtime — they are not separate deployments or A2A services on their own ports. gateway_agent and orchestrator_agent together form the routing entry point that plans and hands off to the specialized agents.
search_agent.py, summarizer_agent.py, detailed_report_agent.py, and text_analysis_agent.py each expose their own standalone FastAPI app with a uvicorn.run(...) call, for running that one agent alone during development (default ports 8002, 8003, 8004, 8005 respectively) — this is not how the chart deploys the system. All other agent classes (orchestrator_agent, gateway_agent, entity_extraction_agent, query_enhancement_agent, profile_selection_agent, image_search_agent, audio_analysis_agent, document_agent, deep_research_agent, coding_agent, and the nine knowledge/federation agents) have no standalone entry point at all; a port value on their constructor or in configs/config.json is agent-config metadata used to build the A2A card, not a way to launch them as their own server.
Service Architecture¶
flowchart TB
Client["<span style='color:#000'>Client / Web client</span>"]
Client --> Runtime["<span style='color:#000'>Runtime<br/>Port 8000<br/>REST + A2A JSON-RPC</span>"]
subgraph Registry["<span style='color:#000'>In-process AgentRegistry dispatch</span>"]
Gateway["<span style='color:#000'>gateway_agent</span>"]
Orchestrator["<span style='color:#000'>orchestrator_agent</span>"]
Search["<span style='color:#000'>search_agent</span>"]
Summarizer["<span style='color:#000'>summarizer_agent</span>"]
DetailedReport["<span style='color:#000'>detailed_report_agent</span>"]
QueryEnh["<span style='color:#000'>query_enhancement_agent</span>"]
EntityExt["<span style='color:#000'>entity_extraction_agent</span>"]
end
Runtime --> Gateway
Runtime --> Orchestrator
Gateway --> Orchestrator
Orchestrator --> Search
Orchestrator --> Summarizer
Orchestrator --> DetailedReport
Orchestrator --> QueryEnh
Orchestrator --> EntityExt
Search --> Vespa["<span style='color:#000'>Vespa<br/>Port 8080, 19071<br/>Vector Database</span>"]
Runtime -.-> Phoenix["<span style='color:#000'>Phoenix<br/>Port 6006, 4317<br/>Telemetry</span>"]
Summarizer --> Ollama["<span style='color:#000'>Ollama<br/>Port 11434<br/>Local LLM</span>"]
Vespa --> VespaData[("<span style='color:#000'>Video Embeddings</span>")]
Phoenix --> PhoenixData[("<span style='color:#000'>Experiments</span>")]
Ollama --> OllamaData[("<span style='color:#000'>Models</span>")]
style Client fill:#90caf9,stroke:#1565c0,color:#000
style Runtime fill:#ffcc80,stroke:#ef6c00,color:#000
style Gateway fill:#ce93d8,stroke:#7b1fa2,color:#000
style Orchestrator fill:#ce93d8,stroke:#7b1fa2,color:#000
style Search fill:#ce93d8,stroke:#7b1fa2,color:#000
style Summarizer fill:#ce93d8,stroke:#7b1fa2,color:#000
style DetailedReport fill:#ce93d8,stroke:#7b1fa2,color:#000
style QueryEnh fill:#ce93d8,stroke:#7b1fa2,color:#000
style EntityExt fill:#ce93d8,stroke:#7b1fa2,color:#000
style Vespa fill:#a5d6a7,stroke:#388e3c,color:#000
style Phoenix fill:#a5d6a7,stroke:#388e3c,color:#000
style Ollama fill:#a5d6a7,stroke:#388e3c,color:#000
style VespaData fill:#b0bec5,stroke:#546e7a,color:#000
style PhoenixData fill:#b0bec5,stroke:#546e7a,color:#000
style OllamaData fill:#b0bec5,stroke:#546e7a,color:#000 The diagram above shows the core dispatch path with a representative subset of agents. All 23 registered agent classes are dispatched the same way, through the in-process AgentRegistry — none of them run as separate deployments.
Complete Agent Roster¶
configs/config.json agents.* registers 23 agents. 14 are enabled by default; the 9 disabled ones are knowledge-graph and multi-tenant federation agents still under active development.
Generation + Routing Agents
| Agent | Enabled | Purpose |
|---|---|---|
gateway_agent | yes | LLM-free A2A entry point that classifies queries via GLiNER and routes simple ones directly, complex ones to the orchestrator. |
orchestrator_agent | yes | Autonomous coordinator that plans a multi-agent workflow with DSPy then executes it by calling sub-agents over A2A HTTP. |
summarizer_agent | yes | A2A summarizer that turns search results into structured summaries with a thinking phase and VLM visual analysis. |
detailed_report_agent | yes | A2A agent that generates comprehensive detailed reports (executive summary, findings, technical + visual analysis, recommendations) with optional RLM synthesis. |
profile_selection_agent | yes | Uses DSPy LLM reasoning to pick the optimal backend search profile for a query, with a heuristic fallback. |
query_enhancement_agent | yes | Expands and rewrites queries with synonyms, context, and RRF variants using DSPy. |
entity_extraction_agent | yes | Tiered NER agent: fast GLiNER + SpaCy path (no LLM) with a DSPy ChainOfThought fallback. |
Search & Analysis Agents
| Agent | Enabled | Purpose |
|---|---|---|
search_agent | yes | Multi-modal retrieval agent that searches Vespa across video/image/text/audio/document modalities with DSPy query rewriting and RRF ensemble fusion. |
image_search_agent | yes | ColPali multi-vector image similarity search against Vespa, with semantic and hybrid (BM25+ColPali) modes and image-to-image lookup. |
document_agent | yes | Dual-strategy document search: ColPali visual (page-as-image), ColBERT/BM25 text, or hybrid, with keyword-based auto strategy selection. |
text_analysis_agent | yes | Runtime-configurable DSPy text analysis agent (sentiment/summary/entities) with per-tenant persisted config and its own FastAPI /analyze endpoint. |
audio_analysis_agent | yes | Whisper transcription + Vespa audio search agent supporting transcript (BM25), acoustic (CLAP nearest-neighbor), and hybrid modes. |
Research + Coding Agents
| Agent | Enabled | Purpose |
|---|---|---|
deep_research_agent | yes | Multi-step research agent that decomposes a query, iteratively gathers evidence via parallel searches, and synthesizes a cited report. |
coding_agent | yes | Iterative coding agent that searches code semantically, plans and generates code with DSPy, and runs it in an OpenShell sandbox, looping on failures. |
Knowledge-Graph & Reasoning Agents (all disabled by default except audit_explanation_agent)
| Agent | Enabled | Purpose |
|---|---|---|
citation_tracing_agent | no | Read-only agent that walks a memory's provenance chain back to its primary sources. |
contradiction_reconciliation_agent | no | Resolves conflict sets by applying a knowledge schema's contradiction policy over member memories. |
multi_document_synthesis_agent | no | Synthesises a coherent answer across N source documents while preserving the citation graph. |
kg_traversal_agent | no | Structurally walks kg_node/entity_fact and kg_edge memories from a seed entity into a node+edge graph view. |
temporal_reasoning_agent | no | Compares a subject's knowledge across explicit time windows using provenance.written_at. |
knowledge_summarization_agent | no | Distills a knowledge subgraph into a structured, citation-aware summary with optional admin-gated promotion to the org trunk. |
audit_explanation_agent | yes | Explains why a given answer memory was produced — its derivation chain, per-source trust, and active contradictions. |
Multi-Tenant + Federation Agents (disabled by default)
| Agent | Enabled | Purpose |
|---|---|---|
cross_tenant_comparison_agent | no | Read-only A2A agent that compares per-tenant views of one subject across all tenants in an org via the federation read path. |
federated_query_agent | no | Read-only A2A agent that answers a free-text query by aggregating federated reads across multiple tenants in the same org, with an optional RLM summariser. |
None of the 23 agents run on their own port in a deployed cluster — they are all dispatched in-process by the Runtime. See Runtime and Agents above for which ones also expose a standalone dev-mode FastAPI server.
Local Development¶
Quick Setup¶
# Clone repository
git clone https://github.com/amit-jain/cogniverse.git
cd cogniverse
# Install uv 0.12.19, then dependencies with the PyTorch extra for this host
curl -LsSf https://astral.sh/uv/0.12.19/install.sh | sh
scripts/install_with_gpu.sh
# Start Vespa
docker run -d --name vespa \
-p 8080:8080 -p 19071:19071 \
-v vespa-data:/opt/vespa/var \
vespaengine/vespa:latest
# Start Phoenix
docker run -d --name phoenix \
-p 6006:6006 -p 4317:4317 \
-v phoenix-data:/data \
-e PHOENIX_WORKING_DIR=/data \
arizephoenix/phoenix:latest
# Start Ollama
docker run -d --name ollama \
-p 11434:11434 \
-v ollama-data:/root/.ollama \
ollama/ollama:latest
# Pull required Ollama models
docker exec ollama ollama pull gemma3:4b
# Verify services
curl http://localhost:8080/ApplicationStatus # Vespa
curl http://localhost:6006/health # Phoenix
curl http://localhost:11434/api/tags # Ollama
Service Ports¶
| Service | Port | Protocol | Purpose |
|---|---|---|---|
| Vespa HTTP | 8080 | HTTP | Document feed & search queries |
| Vespa Config | 19071 | HTTP | Schema deployment |
| Phoenix | 6006 | HTTP | Telemetry & experiments dashboard |
| Phoenix OTLP | 4317 | gRPC | OTLP span collection (served by Phoenix) |
| Ollama | 11434 | HTTP | LLM inference API |
| Runtime | 8000 | HTTP | REST API + in-process A2A JSON-RPC entry point |
None of the 23 agents registered under configs/config.json agents.* (see Complete Agent Roster above) have their own port in a deployed cluster — all are dispatched in-process by the Runtime. Only search_agent.py, summarizer_agent.py, detailed_report_agent.py, and text_analysis_agent.py define their own standalone FastAPI app + uvicorn.run(...) for running that one agent on its own during development (default ports 8002, 8003, 8004, and 8005 respectively). All other agent classes have no standalone entry point — orchestrator_agent's constructor defaults port=8013, gateway_agent's port=8014, entity_extraction_agent's port=8010, profile_selection_agent's port=8011, and query_enhancement_agent's port=8012, but these are only agent-config metadata (used to build the A2A card), not a way to run the agent as its own server. The nine knowledge-graph/federation agent classes work the same way, using their own _DEFAULT_PORT module constants (8019–8027, one port per agent in roster order) purely as A2A card metadata.
Environment Configuration¶
Create a .env file in the workspace root:
# Backend (read by BootstrapConfig — REQUIRED)
BACKEND_URL=http://localhost
BACKEND_PORT=8080
# JAX (for X-CLIP models)
JAX_PLATFORM_NAME=cpu
# Config file path (optional — defaults to configs/config.json)
# COGNIVERSE_CONFIG=/path/to/config.json
# Tenant ID is per-request (not an env var)
# Agents receive tenant_id in each A2A task payload
Deployment Options¶
Cogniverse supports multiple deployment methods depending on your needs:
Unified Deployment (k3d/Helm)¶
Full-stack bring-up on a fresh machine. Two strategies share the same cluster deploy: A — all inference served by local cluster pods; B — selected inference services on Modal, called by the cluster over https://*.modal.run.
Prerequisites¶
docker,kubectl,helm(checked bycogniverse up; it offers to install what is missing), plusk3dwhen no existing Kubernetes cluster is reachable.uvwith the workspace synced (see Setup & Installation);cogniverseis the CLI entry point.
Credentials¶
The secret syncs resolve each value in this order: the environment variable, then ./.env, then ~/.env (each may be a directory holding one <VAR>.env file per secret, or a single file of KEY=value lines), then any tool-specific location.
| Value | Cluster Secret | Needed for |
|---|---|---|
HF_TOKEN / HUGGING_FACE_HUB_TOKEN, else ~/.cache/huggingface/token | hf-token | Gated HuggingFace models (e.g. the Gemma LLM) |
TELEGRAM_BOT_TOKEN | cogniverse-messaging-secrets | --messaging only |
COGNIVERSE_INFERENCE_API_KEY | cogniverse-inference-api-key | Strategy B only |
Strategy A — fully local¶
cogniverse up # flags: --llm auto|builtin|external, --llm-url,
# --image-source, --messaging, --sandbox, --sandbox-endpoint
cogniverse status
cogniverse up runs, in order: environment detection (a running k3d cluster counts as local); prerequisite check; host-LLM probe on :11434 (--llm auto); k3d cluster create with NodePort range 1-65535, workspace / HF-cache / host-storage mounts and AMD GPU device mounts, then CoreDNS upstreams pinned to 1.1.1.1 8.8.8.8 (cluster start fails if pinning fails); image build + k3d image import + prune of the superseded generation; values composition (values.k3s.yaml plus values.<backend>.yaml for the detected torch backend); third-party image pre-pull; secret syncs (hf-token, cogniverse-inference-api-key — each a warn-and-skip when the local value is absent); Argo controller install; sandbox wiring; helm install/upgrade of release cogniverse into namespace cogniverse; workflow-template deploy; kubectl wait for pod readiness (300s); port-forwards and HTTP health checks (Vespa :19071, Runtime :28000, Web client :28400, Phoenix :26006, LLM :11434, Argo :2746).
Re-sync cluster Secrets later with cogniverse secrets sync (--required fails instead of warning).
Strategy B — local cluster + Modal-hosted inference¶
Modal credentials stay on the control plane — they are never synced to the cluster. modal token new writes ~/.modal.toml; MODAL_TOKEN_ID / MODAL_TOKEN_SECRET override it per-process.
- Generate the shared bearer key and store it where the sync reads it:
- Create the Modal-side secrets the apps require:
modal secret create cogniverse-inference-api-key \
COGNIVERSE_INFERENCE_API_KEY=$(cat ~/.env/COGNIVERSE_INFERENCE_API_KEY.env)
modal secret create hf-token HF_TOKEN=<token> # vllm_llm_student and vllm_llm_teacher only
- Deploy and warm the Modal services (canonical names:
vllm_colpali,colbert_pylate,code_colbert_pylate,denseon,gliner,video_embed,vllm_llm_student,vllm_llm_teacher,vllm_asr,clap_embed,face_embed):
cogniverse inference modal deploy denseon colbert_pylate
cogniverse inference modal warm denseon colbert_pylate # prints each base URL
-
Point the chart at the printed endpoints: set
inference.<svc>.externalUrlin the values overlay the deploy applies (values.k3s.yaml+values.<backend>.yamlforcogniverse up;values.prod.yamlfor prod), or--setit on a manualhelm upgrade. The service keeps itsINFERENCE_SERVICE_URLSentry, pointed at Modal, and its local pod, Service, model-cache PVC, and image build are skipped — see Models and Inference.externalUrlis rejected onvllm_llm_student, whose endpoint is overridden viaruntime.primaryLLM.apiBaseinstead;vllm_llm_teacheracceptsexternalUrl. -
cogniverse up. The sync places the key incogniverse-inference-api-key; the runtime and ingestor resolve the Modal URLs fromINFERENCE_SERVICE_URLSand attachAuthorization: Bearer $COGNIVERSE_INFERENCE_API_KEYtohttps://*.modal.runcalls. After rotating the key, re-runcogniverse secrets syncand recreate the Modal secret.
See detailed guides:
-
Kubernetes Deployment Guide - K8s/K3s/Helm deployment
-
Istio Service Mesh Guide - Service mesh with mTLS, Phoenix tracing, DNS-based multi-cluster
-
Argo Workflows Guide - Batch processing workflows
Modal Deployment (Serverless GPU)¶
Heavy inference services run on Modal via the cogniverse inference modal lifecycle (deploy / warm / release / status / qualify / undeploy), authenticated with COGNIVERSE_INFERENCE_API_KEY. Setup, service names, and the per-service inference.<svc>.externalUrl chart override are covered in Unified Deployment — Strategy B and Models and Inference.
Multi-Tenant Schema Deployment¶
Cogniverse supports multi-tenant deployment with per-tenant schema isolation.
Schema Deployment (runtime admin API)¶
All production schema deployment goes through the runtime's admin API so it funnels through SchemaRegistry.deploy_schema and the VespaBackend.deploy_schemas merge path (which discovers deployed document types from the live cluster and refuses to silently drop peer-tenant schemas). The in-cluster init job at charts/cogniverse/templates/init-jobs.yaml wires this up as a post-install step that registers each .Values.config.tenants entry (POST /admin/tenants; a 409 for an already registered tenant proceeds) and deploys .Values.config.defaultProfiles.video (none when it is empty) for it:
RUNTIME_URL="http://$HOST:$PORT"
# Register a tenant (creates tenant metadata records)
curl -sfX POST "$RUNTIME_URL/admin/tenants" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production"}'
# Deploy a profile's content schema for that tenant
curl -sfX POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production", "force": false}'
Tenant-Aware Schema Deployment (programmatic path — what the admin API calls internally):
from cogniverse_foundation.config.utils import create_default_config_manager, get_config
from cogniverse_core.common.tenant_utils import SYSTEM_TENANT_ID
from cogniverse_core.registries.schema_registry import SchemaRegistry
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from cogniverse_vespa.backend import VespaBackend
from pathlib import Path
# Initialize configuration (SYSTEM_TENANT_ID reads cluster-wide base config;
# per-tenant overrides merge in when a specific tenant is resolved later.)
config_manager = create_default_config_manager()
config = get_config(tenant_id=SYSTEM_TENANT_ID, config_manager=config_manager)
# Get the cluster-wide backend config that tenant overrides merge on top of.
backend_config = config_manager.get_backend_config(tenant_id=SYSTEM_TENANT_ID)
# Initialize schema loader
schemas_dir = Path("configs/schemas")
schema_loader = FilesystemSchemaLoader(schemas_dir)
# Initialize backend with all required parameters
# VespaBackend constructor requires:
# backend_config: BackendConfig instance (REQUIRED)
# schema_loader: SchemaLoader instance (REQUIRED)
# config_manager: ConfigManager instance (REQUIRED)
backend = VespaBackend(
backend_config=backend_config,
schema_loader=schema_loader,
config_manager=config_manager
)
# Initialize schema registry
registry = SchemaRegistry(
config_manager=config_manager,
backend=backend,
schema_loader=schema_loader
)
# Deploy schemas for multiple real tenants. There is no "default" tenant
# in cogniverse — every tenant must have been registered via the admin
# API or created here programmatically.
tenants = ["acme_corp", "globex_inc"]
for tenant_id in tenants:
# This example deploys only the video content schemas; the same
# pattern applies to any other profile schema (image_colpali_mv,
# audio_content, document_text, document_visual, lateon_mv,
# code_lateon_mv) by adjusting the filter below.
all_schemas = schema_loader.list_available_schemas()
video_schemas = [
s for s in all_schemas
if s.startswith("video_") and not s.endswith("_metadata")
]
for base_schema in video_schemas:
# Deploy tenant-specific schema (name: {base_schema}_{tenant_id})
tenant_schema_name = registry.deploy_schema(
tenant_id=tenant_id,
base_schema_name=base_schema,
force=False
)
print(f"Deployed: {tenant_schema_name}")
Deploy Schemas¶
# Production: loop tenant × profile against the runtime admin API
# (same pattern as the Helm init job):
RUNTIME_URL="http://localhost:8080"
for tenant in acme:production globex:staging; do
curl -sfX POST "$RUNTIME_URL/admin/tenants" \
-H 'Content-Type: application/json' \
-d "{\"tenant_id\": \"$tenant\"}"
# Add every content profile the tenant needs, e.g.:
# video_colpali_smol500_mv_frame image_colpali_mv document_text_semantic
for profile in video_colpali_smol500_mv_frame; do
curl -sfX POST "$RUNTIME_URL/admin/profiles/$profile/deploy" \
-H 'Content-Type: application/json' \
-d "{\"tenant_id\": \"$tenant\"}"
done
done
# Dev: single-schema JSON deploy from the filesystem
JAX_PLATFORM_NAME=cpu uv run python scripts/deploy_json_schema.py \
configs/schemas/video_colpali_smol500_mv_frame_schema.json
Note: The runtime admin API is the production path and the only route that preserves peer-tenant schemas through redeploys. deploy_json_schema.py is kept for single-schema dev iteration.
Available Schemas¶
The configs/schemas/ directory (21 files) breaks down into three categories:
- Content/profile schemas — one per
configs/config.jsonprofiles.*entry, deployed on demand throughPOST /admin/profiles/{profile}/deployand always tenant-suffixed. - Knowledge/memory schemas —
agent_memories,knowledge_graph,provenance,wiki_pages. Also tenant-suffixed, but deployed automatically on tenant bootstrap (memory_init.py,ingestion_worker/worker.py) rather than through the profile-deploy route. - Cluster-wide metadata schemas —
organization_metadata_schema.json,tenant_metadata_schema.json,config_metadata_schema.json,adapter_registry_schema.json. Never tenant-suffixed; one instance serves the whole cluster.
ranking_strategies.json is not a Vespa schema — it is a schema-name-keyed map of ranking-profile definitions consumed by the backend at query time.
The schema names inside the schema JSON files do not include the _schema.json suffix.
| Schema File | Schema Name (in JSON) | Profile | Embedding Model | Dimensions | Tenant Suffix |
|---|---|---|---|---|---|
video_colpali_smol500_mv_frame_schema.json | video_colpali_smol500_mv_frame | video_colpali_smol500_mv_frame | TomoroAI/tomoro-colqwen3-embed-4b | 320 per patch | _<tenant_id> |
video_colqwen_omni_mv_chunk_30s_schema.json | video_colqwen_omni_mv_chunk_30s | video_colqwen_omni_mv_chunk_30s | TomoroAI/tomoro-colqwen3-embed-4b | 320 per patch | _<tenant_id> |
video_xclip_sv_chunk_6s_schema.json | video_xclip_sv_chunk_6s | video_xclip_sv_chunk_6s | microsoft/xclip-large-patch14 | 768 | _<tenant_id> |
image_colpali_mv_schema.json | image_colpali_mv | image_colpali_mv | TomoroAI/tomoro-colqwen3-embed-4b | 320 per patch | _<tenant_id> |
audio_content_schema.json | audio_content | audio_clap_semantic | laion/clap-htsat-unfused | 128 per token + 512 dense | _<tenant_id> |
document_text_schema.json | document_text | document_text_semantic | lightonai/LateOn | 128 per token | _<tenant_id> |
document_visual_schema.json | document_visual | document_visual_colpali | TomoroAI/tomoro-colqwen3-embed-4b | 320 per patch | _<tenant_id> |
lateon_mv_schema.json | lateon_mv | lateon_mv | lightonai/LateOn | 128 per token | _<tenant_id> |
code_lateon_mv_schema.json | code_lateon_mv | code_lateon_mv | lightonai/LateOn-Code-edge | 48 per token | _<tenant_id> |
Knowledge/memory schemas (tenant-suffixed, bootstrap-deployed — no profiles.* entry):
| Schema File | Schema Name (in JSON) | Deployed by |
|---|---|---|
agent_memories_schema.json | agent_memories | memory_init.py, MemoryManager |
knowledge_graph_schema.json | knowledge_graph | main.py tenant bootstrap, ingestion_worker/worker.py |
provenance_schema.json | provenance | provenance_store.py |
wiki_pages_schema.json | wiki_pages | main.py tenant bootstrap |
Example: For tenant acme_corp, content-schema deployment produces video_colpali_smol500_mv_frame_acme_corp, image_colpali_mv_acme_corp, document_text_acme_corp, and so on for every profile deployed for that tenant; bootstrap deployment produces agent_memories_acme_corp, knowledge_graph_acme_corp, provenance_acme_corp, and wiki_pages_acme_corp.
Each tenant gets completely isolated schemas with their own documents and indexes.
Monitoring & Observability¶
Phoenix Dashboard Access¶
# Access local Phoenix dashboard
open http://localhost:6006
# View tenant-specific traces (each tenant has isolated project)
# Project names follow "cogniverse-{tenant_id}" (TelemetryConfig
# .get_project_name; there is no "default" tenant/project):
# - cogniverse-acme_corp
# - cogniverse-globex_inc
Phoenix Telemetry Integration¶
Phoenix telemetry is automatically enabled with per-tenant project isolation:
Tracked Spans (per tenant):
-
Query processing spans (tenant-specific)
-
Agent routing decisions (tenant-isolated)
-
Vespa search operations (tenant schema-specific)
-
Embedding generation (tenant context)
-
Multi-modal reranking (tenant metrics)
Tenant Isolation:
-
Each tenant has its own Phoenix project
-
No cross-tenant trace visibility
-
Tenant-specific experiment tracking
-
Isolated span datasets per tenant
Example Span Attributes:
{
"tenant_id": "acme_corp",
"schema_name": "video_colpali_smol500_mv_frame_acme_corp",
"phoenix_project": "cogniverse-acme_corp",
"query": "machine learning tutorial",
"agent_type": "search"
}
Troubleshooting¶
Vespa Connection Issues¶
# Check Vespa health
curl http://localhost:8080/state/v1/health
# Restart Vespa
docker restart vespa
# Check logs
docker logs vespa
Phoenix Not Recording Spans¶
# Check Phoenix is running
docker ps | grep phoenix
# Verify OTLP endpoint is reachable (Phoenix serves OTLP on port 4317)
curl -s http://localhost:4317 || echo "Phoenix OTLP not reachable"
# Check Phoenix logs
docker logs phoenix
Ollama Model Issues¶
# List installed models
docker exec ollama ollama list
# Remove and re-pull model
docker exec ollama ollama rm gemma3:4b
docker exec ollama ollama pull gemma3:4b
SDK Package Deployment¶
Building Distribution Packages¶
# Build all 12 SDK packages for distribution
for dir in libs/*/; do
echo "Building $(basename $dir)..."
(cd "$dir" && uv build)
done
# Packages created in dist/ directory (all 12 packages):
# Foundation Layer:
# - cogniverse-sdk-0.1.0-py3-none-any.whl
# - cogniverse-foundation-0.1.0-py3-none-any.whl
# Core Layer:
# - cogniverse-core-0.1.0-py3-none-any.whl
# - cogniverse-evaluation-0.1.0-py3-none-any.whl
# - cogniverse-telemetry-phoenix-0.1.0-py3-none-any.whl
# Implementation Layer:
# - cogniverse-agents-0.1.0-py3-none-any.whl
# - cogniverse-vespa-0.1.0-py3-none-any.whl
# - cogniverse-synthetic-0.1.0-py3-none-any.whl
# - cogniverse-finetuning-0.1.0-py3-none-any.whl
# Application Layer:
# - cogniverse-runtime-0.1.0-py3-none-any.whl
# - cogniverse-cli-0.1.0-py3-none-any.whl
# - cogniverse-messaging-0.1.0-py3-none-any.whl
Installing from Wheels¶
# Install core package only
pip install dist/cogniverse-core-0.1.0-py3-none-any.whl
# Install agents package (includes core dependency)
pip install dist/cogniverse-agents-0.1.0-py3-none-any.whl
# Install all packages
pip install dist/*.whl
Publishing to PyPI (Optional)¶
# Build all packages
for dir in libs/*/; do
(cd "$dir" && uv build)
done
# Publish to PyPI
for dir in libs/*/; do
(cd "$dir" && uv publish)
done
Related Documentation¶
- Setup & Installation - Complete installation guide for UV workspace
- Configuration - Multi-tenant configuration management
- Multi-Tenant Operations - Multi-tenant deployment patterns
- Istio Service Mesh - Production service mesh with mTLS and Phoenix tracing
- SDK Architecture - SDK package structure
- Performance & Monitoring - Performance targets and monitoring
- Modal Deployment - Serverless GPU deployment