Cogniverse Quickstart¶
Get started with Cogniverse in 5 minutes.
Prerequisites¶
- Python 3.12
- uv 0.12.19
- Docker, kubectl, helm, and k3d (
cogniverse upchecks for them and offers to install missing ones) - Deno 2.0+ (required for RLM sandboxed code execution:
curl -fsSL https://deno.land/install.sh | sh)
1. Clone and Install¶
# Clone the repository
git clone https://github.com/amit-jain/cogniverse.git
cd cogniverse
# Install dependencies with the PyTorch extra for this host
scripts/install_with_gpu.sh
source .venv/bin/activate
# Linux + ROCm: stop `uv run` from re-syncing away the ROCm torch wheels
export UV_NO_SYNC=1
2. Start Services¶
# Deploy the full stack (Vespa, Phoenix, LLM, Runtime, Web client) via k3d.
# By default the LLM is vLLM gemma-4-e4b-it on ROCm hosts, Ollama gemma3:4b on
# CPU and CUDA hosts.
cogniverse up
# Wait for services to be ready (~30 seconds). k3d's loadbalancer publishes
# each service's NodePort on localhost — these differ from the in-cluster
# ports (e.g. Runtime listens on 8000 inside the pod, but is reachable at
# 28000 on the host).
curl http://localhost:8080/ApplicationStatus # Vespa
curl http://localhost:26006/health # Phoenix (NodePort 26006)
# Backend connection env vars for any host-side script or CLI that talks to
# Vespa directly (create_default_config_manager() requires these)
export BACKEND_URL=http://localhost
export BACKEND_PORT=8080
3. Create a Tenant & Deploy Its Schemas¶
Schema deployment is always per-tenant in cogniverse — there is no "default" tenant. Create a tenant first; the runtime provisions the tenant's metadata schemas. Then deploy the profile's content schema for that tenant.
RUNTIME_URL=http://localhost:28000
# Create the tenant (auto-provisions org + deploys tenant metadata)
curl -sfX POST "$RUNTIME_URL/admin/tenants" \
-H 'Content-Type: application/json' \
-d '{
"tenant_id": "quickstart",
"created_by": "admin"
}'
# Deploy the profile's content schema for that tenant
curl -sfX POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "quickstart", "force": false}'
Note: The tenant management API is available on the main Runtime (NodePort 28000, in-cluster port 8000) at
/admin/tenants, or as a standalone service on port 9000 if run separately (python -m cogniverse_runtime.admin.tenant_manager, not deployed bycogniverse up).
4. Run Your First Search¶
# quick_search.py
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# Requires: schemas deployed and tenant created (see step 3), and
# BACKEND_URL / BACKEND_PORT exported (see step 2)
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Create agent with dependencies
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# Search by text — tenant_id is required per-request
results = agent.search_by_text(
query="machine learning tutorial",
tenant_id="quickstart",
top_k=10,
)
for result in results:
print(f"- {result.get('id', 'unknown')}: {result.get('score', 0):.2f}")
5. Process a Video¶
The easiest way to ingest videos is via the CLI script:
# Ingest a single directory of videos
uv run python scripts/run_ingestion.py \
--tenant-id quickstart \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa \
--profile video_colpali_smol500_mv_frame
# Ingest with multiple profiles for richer retrieval
uv run python scripts/run_ingestion.py \
--tenant-id quickstart \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
video_xclip_sv_chunk_6s
# Test mode — limited frames for faster iteration
uv run python scripts/run_ingestion.py \
--tenant-id quickstart \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
--test-mode --max-frames 1
6. Query the Runtime API¶
cogniverse up (step 2) already deployed the Runtime FastAPI server — it's reachable on the host at NodePort 28000 (in-cluster port 8000).
# Test the health endpoint
curl http://localhost:28000/health
# Execute a search via API
curl -X POST http://localhost:28000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "machine learning tutorial",
"profile": "video_colpali_smol500_mv_frame",
"top_k": 10,
"tenant_id": "quickstart"
}'
For local development without k3d (hot-reload on code changes), run the server directly against the same Vespa instance instead:
7. Open the Web Client¶
cogniverse up (step 2) also deployed the web client. It is reachable on the host at NodePort 28400.
To run it locally without k3d, see Web Client.
Optional: Enable Messaging Gateway¶
Allow users to interact with Cogniverse via Telegram:
# Start with Telegram bot enabled
cogniverse up --messaging
# Or set in Helm values (production):
# messaging.enabled: true
# messaging.mode: webhook # polling for dev
# Generate an invite token for a user
curl -X POST http://localhost:28000/admin/messaging/invite \
-H "Content-Type: application/json" \
-d '{"tenant_id": "quickstart", "expires_in_hours": 24}'
# Returns: {"token": "abc123...", "tenant_id": "quickstart"}
# User sends: /start abc123... to the bot to register
Optional: Wiki Knowledge Base¶
The wiki knowledge base automatically saves substantial agent interactions as searchable pages in Vespa. It requires no configuration — the wiki_pages schema is deployed automatically, per-tenant, the first time a tenant saves or searches the wiki (no manual schema step needed).
# Save a session manually
curl -X POST http://localhost:28000/wiki/save \
-H "Content-Type: application/json" \
-d '{
"query": "how does ColPali work?",
"response": {"answer": "ColPali uses patch-level embeddings..."},
"entities": ["ColPali", "patch_embeddings"],
"agent_name": "summarizer_agent",
"tenant_id": "quickstart"
}'
# Search the wiki
curl -X POST http://localhost:28000/wiki/search \
-H "Content-Type: application/json" \
-d '{"query": "ColPali embeddings", "tenant_id": "quickstart", "top_k": 5}'
Via Telegram (after connecting the messaging gateway):
/wiki save — Save the current session to the wiki
/wiki search ColPali — Search the wiki knowledge base
/wiki topic ColPali — Look up a topic page by name
/wiki index — Show the wiki index
/wiki lint — Check wiki for orphan, stale, empty, or malformed pages
Auto-filing triggers automatically when an interaction has 3+ extracted entities, comes from detailed_report_agent or deep_research_agent, or spans 4+ conversation turns.
Optional: Enable Quality Monitor¶
The quality monitor runs as its own Deployment, continuously evaluating all agents and triggering optimization when quality degrades:
# Run directly
TELEMETRY_OTLP_ENDPOINT=localhost:4317 \
python -m cogniverse_runtime.quality_monitor_cli \
--tenant-id quickstart \
--runtime-url http://localhost:28000 \
--phoenix-url http://localhost:26006 \
--llm-model qwen3:4b
# Enabled by default in Helm (runtime.qualityMonitor.enabled: true)
What's Next?¶
| Goal | Documentation |
|---|---|
| Understand the architecture | Architecture Overview |
| Use the interactive coding agent | Coding Agent CLI |
| Use the Knowledge Management Layer | User Guide — Knowledge Management |
| Extract a knowledge graph from code / docs | Knowledge Graph |
| Create a custom agent | Creating Agents Tutorial |
| Learn key concepts | Glossary |
| Deep dive into modules | Core Module, Foundation, Agents, Runtime, Vespa Backend |
| Configure profiles | Configuration System |
| Run tests | Testing Guide |
Project Structure¶
cogniverse/
├── libs/ # 12-package workspace
│ ├── sdk/ # Pure interfaces
│ ├── foundation/ # Config + telemetry
│ ├── core/ # Agent base, orchestration, caching
│ ├── agents/ # Agent implementations + strategy learner
│ ├── vespa/ # Vespa backend
│ ├── evaluation/ # Metrics, experiments, quality monitor
│ ├── finetuning/ # LLM fine-tuning (SFT, DPO)
│ ├── telemetry-phoenix/ # Phoenix telemetry provider
│ ├── synthetic/ # Training data generation
│ ├── runtime/ # FastAPI server + quality monitor CLI
│ ├── messaging/ # Telegram messaging gateway
│ └── cli/ # cogniverse CLI (deploy, manage)
├── clients/
│ └── web/ # Web client (the Cogniverse UI)
├── configs/ # Configuration files
│ ├── config.json # Main config
│ └── schemas/ # Schema definitions
├── scripts/ # Utility scripts
├── tests/ # Test suites
└── docs/ # Documentation
Common Commands¶
# Run tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v
# Lint code
make lint-all
# Format code
uv run ruff format .
# Run ingestion
uv run python scripts/run_ingestion.py --tenant-id quickstart --video_dir data/videos --backend vespa --profile video_colpali_smol500_mv_frame
# View Phoenix traces
open http://localhost:26006
Troubleshooting¶
Vespa won't start:
# cogniverse deploys to a k3d/Kubernetes cluster, not raw docker containers
# -- Vespa runs as the cogniverse-vespa StatefulSet
cogniverse status
cogniverse logs vespa -f
# Check port 8080 is free
lsof -i :8080
Import errors:
# Ensure all packages are installed
uv sync
# Check PYTHONPATH
python -c "import cogniverse_core; print('OK')"
Search returns no results:
-
Check tenant_id matches ingested data
-
Verify schema exists:
curl http://localhost:8080/document/v1/ -
Ensure embeddings were generated during ingestion