Skip to content

Cogniverse User Guide

Complete guide for using Cogniverse - the general-purpose multi-agent AI platform for content intelligence and beyond.


Table of Contents

  1. Introduction
  2. Getting Started
  3. Core Features
  4. Basic Operations
  5. Advanced Usage
  6. API Reference
  7. Configuration
  8. Troubleshooting
  9. Best Practices

Introduction

Key Capabilities

For Content Managers:

  • Ingest and index video libraries with multiple embedding strategies
  • Search across videos using natural language queries
  • Get visual relevance scores and frame-level results
  • Monitor performance with built-in dashboards

For Data Scientists:

  • Run experiments with different embedding models and search strategies
  • Evaluate search quality with reference-free and visual LLM metrics
  • Optimize routing agents using synthetic data generation
  • Track all experiments in Phoenix with full observability

For Developers:

  • RESTful API for all operations
  • Multi-tenant support with tenant isolation
  • Configurable embedding profiles and search strategies
  • Plugin architecture for custom components

Architecture at a Glance

See the full Platform Overview for the architecture diagram showing the complete system including agents, content store, memory, telemetry, optimization loop, and training pipeline.


Getting Started

Prerequisites

Before using Cogniverse, ensure you have:

  • Python 3.12+ installed
  • 16GB+ RAM (32GB recommended for large video libraries)
  • Docker for running Vespa, Phoenix, and Ollama
  • GPU (optional but recommended for video processing)
  • uv package manager: pip install uv

Quick Start

Follow these steps to get Cogniverse running in 5 minutes:

1. Installation

# Clone repository
git clone <repository-url>
cd cogniverse

# Install all dependencies
uv sync

# Start infrastructure services
cogniverse up

2. Verify Services

# Check Vespa (should return JSON status)
curl http://localhost:8080/ApplicationStatus

# Check Phoenix (should return "ok")
curl http://localhost:26006/health

# Check Ollama (should return list of models)
curl http://localhost:11434/api/tags

3. Ingest Sample Videos

# Ingest videos with ColPali frame-level embeddings
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/testset/evaluation/sample_videos \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id default

4. Run Your First Query

# Test multi-agent search
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
  --profiles video_colpali_smol500_mv_frame \
  --test-multiple-strategies

Success! You should see search results with ranked videos and relevance scores.

Next Steps

  • View Results: Open the web client at http://localhost:28400 and the Phoenix UI at http://localhost:26006
  • Try API: Use the REST API at http://localhost:28000/docs
  • Configure: Customize profiles in configs/config.json
  • Web UI: Chat with any registered agent and manage the stack in the browser — see the web client
  • Coding Agent: Run cogniverse code to start an interactive coding REPL with streaming — see Coding Agent CLI
  • Knowledge Graph: Run cogniverse index ./path --type code to build a searchable knowledge graph — see Knowledge Graph
  • Learn More: Continue reading this guide

CLI Reference

The cogniverse CLI manages the full stack:

Command Purpose
cogniverse up Deploy all services (Vespa, Phoenix, LLM, Runtime, Web client) via k3d
cogniverse up --messaging Deploy with Telegram gateway enabled
cogniverse down Stop all services
cogniverse down --keep-data Stop services but preserve volumes
cogniverse status Show health of all services
cogniverse logs <service> View logs (runtime, web, vespa, phoenix, llm, argo)
cogniverse logs <service> --follow Stream logs in real-time
cogniverse code Interactive coding agent REPL
cogniverse index <path> --type code Build a knowledge graph from code
cogniverse graph stats Show knowledge graph statistics
cogniverse graph search <query> Search the knowledge graph
cogniverse graph neighbors <node> Find related nodes
cogniverse graph path <src> <dst> Find path between nodes
cogniverse sandbox sync Sync OpenShell gateway mTLS certs into the cluster (after rotation)
cogniverse sandbox status Show OpenShell gateway status and cluster sync state
cogniverse secrets sync Re-sync cluster Secrets (hf-token, cogniverse-messaging-secrets) from env vars, ./.env, or ~/.env
cogniverse admin reconcile-orphans Find (and, with --confirm, drop) Vespa schema orphans not in the schema registry
cogniverse admin merge-article-nodes Report (and, with --apply, perform) merges of KG nodes the_<id> / a_<id> / an_<id> into the tenant's <id> node; --tenant scopes to one tenant; --exclude ID (repeatable) never merges that article id
cogniverse admin invite <tenant_id> Mint a messaging invite token; prints the /start <token> the user sends to the bot

Core Features

Search videos using different modalities:

from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

# Create agent — profile sets the default embedding model
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)

# Search by text (synchronous) — tenant_id is per-request
results = agent.search_by_text(
    query="machine learning tutorial",
    tenant_id="your_org:production",
    top_k=10,
)

for result in results:
    print(f"Video: {result.get('video_id', 'unknown')}")
    print(f"Score: {result.get('score', 0):.2f}")
# Search across different embedding profiles using ensemble mode
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps, SearchInput
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

# Create agent with a default profile
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)

# Single profile search
colpali_results = agent.search_by_text(
    query="cooking tutorial",
    tenant_id="your_org:production",
    top_k=10,
)

# Search with a different profile via SearchInput for ensemble
xclip_deps = SearchAgentDeps(profile="video_xclip_sv_chunk_6s")
xclip_agent = SearchAgent(deps=xclip_deps, config_manager=config_manager, schema_loader=schema_loader)

xclip_results = xclip_agent.search_by_text(
    query="cooking tutorial",
    tenant_id="your_org:production",
    top_k=10,
)
# Search with date filters
results = agent.search_by_text(
    query="machine learning tutorial",
    tenant_id="your_org:production",
    top_k=10,
    start_date="2024-01-01",
    end_date="2024-12-31",
)

2. Intelligent Query Routing

Cogniverse automatically routes queries to the optimal execution agent via the GatewayAgent → OrchestratorAgent pipeline. The orchestrator plans and executes preprocessing agents (QueryEnhancementAgent, EntityExtractionAgent, ProfileSelectionAgent) and threads their enrichment outputs — enhanced_query, entities, relationships, query_variants — directly onto the execution agent's typed input via AgentTask:

import asyncio
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps, OrchestratorInput
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_foundation.config.utils import create_default_config_manager

async def main():
    config_manager = create_default_config_manager()
    registry = AgentRegistry(tenant_id="your_org:production", config_manager=config_manager)
    orchestrator = OrchestratorAgent(deps=OrchestratorDeps(), registry=registry, config_manager=config_manager)

    result = await orchestrator._process_impl(
        OrchestratorInput(
            query="cooking recipes with pasta",
            tenant_id="your_org:production",
        )
    )

    print(result.execution_summary)

asyncio.run(main())

Routing Features:

  • Entity Extraction: Identifies people, places, concepts using GLiNER
  • Relationship Detection: Finds relationships between entities
  • Query Enhancement: Enriches queries with context
  • Modality Detection: Classifies query type (factual, conceptual, visual, etc.)
  • Confidence Scoring: Provides routing confidence for transparency

3. Multiple Embedding Models

Choose the best embedding model for your use case:

Model Type Best For Dimensions
ColPali Frame-level Visual documents, text-rich videos 320 (patch)
X-CLIP Temporal video Text-to-video clip retrieval 768
ColQwen3 Omni Multi-modal Text+visual fusion 320 (patch)

Switching Models:

# Use X-CLIP for global video understanding
uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_xclip_sv_chunk_6s \
  --tenant-id default

4. Hybrid Search Strategies

Combine multiple search methods for better results:

# Ranking strategies are plain rank-profile-name strings (validated against
# the deployed Vespa schema), not an enum:
strategies = [
    "bm25_only",           # Text-only BM25 (fastest for keyword queries)
    "float_float",         # Dense float embeddings (highest visual accuracy)
    "binary_binary",       # Binary embeddings (fastest visual search)
    "float_binary",        # Float query, binary index (speed/accuracy balance)
    "phased",              # Two-phase: binary retrieval, float reranking
    "hybrid_float_bm25",   # Visual + text hybrid (best overall accuracy)
    "hybrid_binary_bm25",  # Fast hybrid (binary visual + text)
    "hybrid_bm25_binary",  # Text matches ranked by binary visual + text
    "hybrid_bm25_float",   # Text matches ranked by float visual + text
    # plus the "_no_description" hybrid variants for ColPali/ColQwen schemas
]

# A strategy string is passed at query time via the `strategy` field of the
# /search/ request (or the query_dict given to VespaSearchBackend.search).
# The SearchAgent handles this automatically based on profile configuration.

Cogniverse remembers user context for personalized results:

from cogniverse_core.memory.manager import Mem0MemoryManager
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# Initialize required dependencies
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

# Get memory manager (singleton per tenant via __new__)
memory = Mem0MemoryManager(tenant_id="your_org:production")

# Initialize with required parameters
memory.initialize(
    backend_host="localhost",
    backend_port=8080,
    llm_model="openai/google/gemma-4-e4b-it",
    embedding_model="lightonai/DenseOn",
    llm_base_url="http://localhost:11434",
    embedder_base_url="http://localhost:29006",
    config_manager=config_manager,
    schema_loader=schema_loader,
)

# Add user preference
memory.add_memory(
    content="User prefers beginner-level Python tutorials",
    tenant_id="your_org:production",
    agent_name="search_agent",
    metadata={"user_id": "user_123"}
)

# Search memories
user_memories = memory.search_memory(
    query="Python tutorial preferences",
    tenant_id="your_org:production",
    agent_name="search_agent",
    top_k=5
)

6. Comprehensive Telemetry

Track everything with Phoenix:

# Phoenix UI: traces, spans and experiments
open http://localhost:26006

# Web client: analytics, evaluation and memory views over the same telemetry
open http://localhost:28400

Web client views over telemetry:

  • Analytics: Traces with latency over time, histograms, outliers and root causes
  • Evaluation: Golden-set search quality per profile and strategy
  • Routing evaluation: Routing decisions, accuracy and confidence calibration
  • Memory: View, search, add and delete stored memories
  • Embedding atlas: Documents placed by their embeddings

Basic Operations

Video Ingestion

Single Profile Ingestion

Ingest videos with one embedding model:

# Ingest with ColPali (frame-based)
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id default

Multi-Profile Ingestion

Ingest with multiple models for best coverage:

# Ingest with ColPali, X-CLIP, and ColQwen
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_colpali_smol500_mv_frame \
            video_xclip_sv_chunk_6s \
            video_colqwen_omni_mv_chunk_30s \
  --tenant-id default

Ingestion Options:

  • --video_dir: Directory containing content files
  • --content-dir: Alias for --video_dir
  • --media-root-uri: Media root URI (e.g. s3://corpus/, pvc://media/) for non-filesystem sources; overrides --video_dir / --content-dir when set
  • --content-type: Content type to ingest (choices: video, image, audio, document; default: video)
  • --profile: Embedding profile(s) - can specify multiple space-separated values
  • --tenant-id: Tenant ID for multi-tenancy (required — no default)
  • --backend: Backend to use (choices: byaldi, vespa; default: vespa)
  • --max-concurrent: Maximum concurrent items to process (default: 3)
  • --output_dir: Output directory for processed data
  • --max-frames: Maximum frames per video / images per batch
  • --test-mode: Use test mode with limited frames
  • --debug: Enable debug mode

Note: Processing options like keyframe extraction, transcription, and frame sampling are configured per-profile in configs/config.json, not via CLI arguments.

Check Ingestion Status

# Confirm documents are searchable via the production /search/ endpoint.
# A non-empty results list means content was indexed for the tenant.
curl -X POST http://localhost:28000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "test",
    "tenant_id": "acme",
    "profile": "video_colpali_smol500_mv_frame",
    "top_k": 1
  }'

# For an async ingestion job, poll its status instead:
# curl http://localhost:28000/ingestion/status/{job_id}

Searching Videos

Use the REST API for production applications:

# Text search
curl -X POST http://localhost:28000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "machine learning tutorial",
    "top_k": 10,
    "strategy": "hybrid_float_bm25",
    "tenant_id": "default"
  }'

Response:

{
  "query": "machine learning tutorial",
  "profile": "video_colpali_smol500_mv_frame",
  "strategy": "hybrid_float_bm25",
  "results_count": 10,
  "results": [
    {
      "document_id": "doc_123",
      "score": 0.95,
      "metadata": {
        "source_id": "video_123",
        "frame_id": "frame_45"
      },
      "highlights": {}
    }
  ],
  "session_id": null
}

Use the Python SDK for scripting:

from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)

# search_by_text is synchronous — no await needed
results = agent.search_by_text(
    query="cooking pasta",
    tenant_id="your_org:production",
    top_k=10,
)

for result in results:
    print(f"Video: {result.get('video_id', 'unknown')}")
    print(f"Score: {result.get('score', 0):.2f}")

Advanced Search Options

# Search with date filters
results = agent.search_by_text(
    query="tutorial",
    tenant_id="your_org:production",
    top_k=10,
    start_date="2024-01-01",
    end_date="2024-12-31",
)

# Search with more results for client-side filtering
results = agent.search_by_text(
    query="Python tutorial",
    tenant_id="your_org:production",
    top_k=50,
)

Running Evaluations

Quick Evaluation

Evaluate search quality on a dataset:

# Run evaluation with Phoenix tracking
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
  --tenant-id acme:acme \
  --dataset-name golden_eval_v1 \
  --profiles video_colpali_smol500_mv_frame \
  --all-strategies \
  --quality-evaluators

This will:

  1. Load evaluation dataset with ground truth
  2. Run queries through all search strategies
  3. Compute quality metrics (MRR, NDCG, Precision@K)
  4. Track results in Phoenix
  5. Generate comparison charts

Custom Evaluation Dataset

Create your own evaluation dataset as a CSV file with query, expected_videos (comma-separated video IDs), and an optional category column:

query,expected_videos,category
machine learning basics,"video_123,video_456",general
Python tutorial for beginners,video_789,general
# Run evaluation on custom dataset
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
  --tenant-id acme:acme \
  --csv-path evaluation_dataset.csv \
  --profiles video_colpali_smol500_mv_frame

Compare Embedding Models

Compare different embedding models:

# Run experiments with multiple profiles
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
  --tenant-id acme:acme \
  --dataset-name golden_eval_v1 \
  --profiles video_colpali_smol500_mv_frame \
             video_xclip_sv_chunk_6s \
             video_colqwen_omni_mv_chunk_30s \
  --all-strategies

Results Table:

Profile                              | MRR@10 | NDCG@10 | Precision@5
-------------------------------------|--------|---------|-------------
video_colpali_smol500_mv_frame       | 0.82   | 0.79    | 0.75
video_xclip_sv_chunk_6s   | 0.78   | 0.74    | 0.70
video_colqwen_omni_mv_chunk_30s      | 0.85   | 0.82    | 0.78


Advanced Usage

Multi-Tenant Configuration

Set up multiple tenants with isolated data:

# Multi-tenancy is handled at the request level, not via separate config objects.
# SystemConfig is global infrastructure config (one per deployment).
# Tenant isolation is achieved by passing tenant_id per request:

from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)

# Each tenant_id produces isolated results:
# - Vespa documents are filtered by a tenant-scoped schema. A bare tenant_id
#   like "acme_corp" is canonicalized to "acme_corp:acme_corp" (org:tenant),
#   giving the schema name video_colpali_smol500_mv_frame_acme_corp_acme_corp
#   (use "org:tenant" form, e.g. "acme_corp:production", to avoid the
#   doubled suffix).
# - Memory: separate Mem0 namespaces
results_acme = agent.search_by_text(query="tutorial", tenant_id="acme_corp", top_k=10)
results_startup = agent.search_by_text(query="tutorial", tenant_id="startup_inc", top_k=10)

Tenant Lifecycle:

# Ingest data for tenant (tenant is implicitly created on first use)
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/new_customer_videos \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id new_customer

# Explicit tenant creation/deletion APIs are also available directly on
# the main runtime:
# curl -X POST http://localhost:28000/admin/tenants \
#   -d '{"tenant_id": "new_customer", "created_by": "admin"}'
# curl -X DELETE http://localhost:28000/admin/tenants/new_customer

Custom Embedding Profiles

Create custom profiles for specific use cases:

{
  "backend": {
    "profiles": {
      "video_custom_highres_frame": {
        "type": "video",
        "description": "Custom high-resolution frame-based profile",
        "schema_name": "video_custom_highres_frame",
        "embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
        "pipeline_config": {
          "extract_keyframes": true,
          "keyframe_strategy": "fps",
          "keyframe_fps": 2.0,
          "transcribe_audio": true,
          "generate_descriptions": true,
          "generate_embeddings": true
        },
        "strategies": {
          "segmentation": {
            "class": "FrameSegmentationStrategy",
            "params": {
              "fps": 2.0,
              "threshold": 0.999,
              "max_frames": 200
            }
          },
          "embedding": {
            "class": "MultiVectorEmbeddingStrategy",
            "params": {}
          }
        },
        "embedding_type": "multi_vector"
      }
    }
  }
}
# Deploy custom profile schema
JAX_PLATFORM_NAME=cpu uv run python scripts/deploy_json_schema.py \
  configs/schemas/video_custom_highres_frame.json

# Ingest with custom profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_custom_highres_frame \
  --tenant-id default

DSPy Optimization

Optimize routing and search agents via the optimization CLI:

# Optimize gateway confidence thresholds
python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id default

# Optimize entity extraction
python -m cogniverse_runtime.optimization_cli --mode entity-extraction --tenant-id default

# Optimize profile performance
python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id default

# Full optimization workflow (all modes)
python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id default

# Triggered optimization (run when quality degrades)
python -m cogniverse_runtime.optimization_cli --mode triggered \
  --tenant-id default --agents search_agent,summarizer_agent \
  --trigger-dataset optimization-trigger-default-20260403_040000

# Cleanup old logs
python -m cogniverse_runtime.optimization_cli --mode cleanup --log-retention-days 7

Available optimization modes:

Mode What It Optimizes
gateway-thresholds GLiNER confidence thresholds
entity-extraction Entity extraction accuracy
profile Profile performance ranking
workflow Full end-to-end optimization pipeline
triggered On-demand when quality monitor fires
simba Query enhancement (SIMBA), from cogniverse.query_enhancement spans
online-routing-eval Scores cogniverse.routing spans (routing_outcome + confidence) without retraining
synthetic Generates synthetic training data for one or more optimizer types (simba, profile, workflow by default)
rollback Restores an agent's active prompts/demos artifact to a previously-snapshotted version
ab-compare Runs an RLM A/B comparison over a Phoenix queries dataset
egress-netpol Emits Kubernetes NetworkPolicy CRDs from per-agent policy YAMLs (no --tenant-id)
monthly-reports Generates the monthly usage + performance report (no --tenant-id)
cleanup Purge old optimization logs and expired memories (no --tenant-id)

Batch Processing

Process large video libraries efficiently using the CLI:

# Process videos in batches using the ingestion script
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/batch1 \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id default

# Or via the API
curl -X POST http://localhost:28000/ingestion/start \
  -H "Content-Type: application/json" \
  -d '{
    "video_dir": "data/batch1",
    "profile": "video_colpali_smol500_mv_frame",
    "tenant_id": "default",
    "batch_size": 10
  }'

# Check status
curl http://localhost:28000/ingestion/status/{job_id}

API Authentication

Authentication is handled via tenant isolation. Each request includes a tenant_id:

# Search with tenant ID
curl -X POST http://localhost:28000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "tutorial",
    "top_k": 10,
    "tenant_id": "acme_corp"
  }'

Note: Tenant isolation provides logical separation of data. API key authentication is planned for future releases.

Telegram Messaging Gateway

Users can interact with Cogniverse via Telegram after receiving an invite token:

Admin: generate an invite token

curl -X POST http://localhost:28000/admin/messaging/invite \
  -H "Content-Type: application/json" \
  -d '{"tenant_id": "acme_corp", "expires_in_hours": 24}'

# Returns: {"token": "abc123def456...", "tenant_id": "acme_corp"}

User: register and use the bot

  1. Send /start abc123def456... to the bot to link your account
  2. Use commands to interact with agents:
Command Agent Example
/search <query> search_agent /search machine learning tutorial
/summarize <query> summarizer_agent /summarize the Python basics video
/report <query> detailed_report_agent /report Q4 content performance
/research <query> deep_research_agent /research best practices for async Python
/code <query> coding_agent /code write a FastAPI health endpoint
/wiki save — Save current session to the wiki
/wiki search <query> — Search the wiki knowledge base
/wiki topic <name> — Look up a topic page by name
/wiki index — Show the full wiki index
/wiki lint — Check wiki for orphan, stale, empty, or malformed pages
/instructions set <text> — Set custom agent instructions for your tenant
/instructions show — Show current tenant instructions
/memories list — List memories (add agent=<name> to filter)
/memories clear strategies — Clear strategy learner memories
/jobs list — List scheduled agent jobs
/jobs create "<cron>" <query> — Create a new scheduled job
/jobs delete <job_id> — Delete a scheduled job
Plain text gateway_agent what videos do you have on transformers?
Photo/video search_agent Send a frame to search for similar content
/help — Show all available commands

Conversation history is maintained via Mem0 across sessions. The gateway runs in polling mode for development (GATEWAY_MODE=polling) and webhook mode for production (GATEWAY_MODE=webhook with TELEGRAM_WEBHOOK_URL set).

Gateway Architecture

flowchart TD
    TG["<span style='color:#000'>Telegram User</span>"]
    BOT["<span style='color:#000'>Telegram Bot API<br/>(webhook / polling)</span>"]
    GW["<span style='color:#000'>MessagingGateway</span>"]
    CR["<span style='color:#000'>command_router<br/>parse_message()</span>"]
    AUTH["<span style='color:#000'>InviteTokenManager<br/>claim_token()</span>"]
    UM["<span style='color:#000'>UserTenantMapper<br/>get_tenant_id()</span>"]
    CM["<span style='color:#000'>ConversationManager<br/>get_history() / store_turn()</span>"]
    RC["<span style='color:#000'>RuntimeClient<br/>POST /agents/{name}/process</span>"]
    FMT["<span style='color:#000'>format_agent_response()<br/>chunk at 4096 chars</span>"]

    TG -->|"sends message"| BOT
    BOT -->|"Update"| GW
    GW --> CR
    CR -->|"ParsedCommand<br/>(agent_name, query)"| GW
    GW --> AUTH
    AUTH -->|"tenant_id"| UM
    UM -->|"tenant confirmed"| GW
    GW --> CM
    CM -->|"conversation history"| RC
    RC -->|"agent response"| FMT
    FMT -->|"chunked messages"| TG

    style TG fill:#81d4fa,stroke:#0288d1,color:#000
    style BOT fill:#90caf9,stroke:#1565c0,color:#000
    style GW fill:#ce93d8,stroke:#7b1fa2,color:#000
    style CR fill:#a5d6a7,stroke:#388e3c,color:#000
    style AUTH fill:#ffcc80,stroke:#ef6c00,color:#000
    style UM fill:#ffcc80,stroke:#ef6c00,color:#000
    style CM fill:#b0bec5,stroke:#546e7a,color:#000
    style RC fill:#64b5f6,stroke:#1565c0,color:#000
    style FMT fill:#a5d6a7,stroke:#388e3c,color:#000

End-to-End User Flow

sequenceDiagram
    participant ADM as Admin
    participant RT as Runtime API
    participant USR as Telegram User
    participant BOT as Telegram Bot API
    participant GW as MessagingGateway
    participant MEM as Mem0 Memory

    ADM->>RT: POST /admin/messaging/invite<br/>{tenant_id, expires_in_hours}
    RT-->>ADM: {token: "abc123..."}
    ADM->>USR: share invite token out-of-band

    USR->>BOT: /start abc123...
    BOT->>GW: Update (start command + token)
    GW->>RT: POST /admin/messaging/register<br/>{platform, external_user_id, token}
    RT->>RT: InviteTokenManager.claim_token()
    RT->>MEM: UserTenantMapper.register_user()
    RT->>RT: InviteTokenManager.mark_token_used()
    RT-->>GW: {tenant_id}
    GW-->>USR: "Registered as acme_corp."

    USR->>BOT: /search machine learning tutorial
    BOT->>GW: Update (search command)
    GW->>GW: parse_message() → search_agent
    GW->>MEM: ConversationManager.get_history(chat_id)
    MEM-->>GW: prior turns
    GW->>RT: POST /agents/search_agent/process
    RT-->>GW: {results: [...], message: "..."}
    GW->>GW: format_agent_response() → chunk at 4096 chars
    GW-->>USR: search results
    GW->>MEM: ConversationManager.store_turn(user + assistant)

    USR->>BOT: what else do you have on this topic?
    BOT->>GW: Update (plain text)
    GW->>GW: parse_message() → gateway_agent
    GW->>MEM: ConversationManager.get_history(chat_id)
    MEM-->>GW: prior turns (multi-turn context)
    GW->>RT: POST /agents/gateway_agent/process<br/>(with conversation_history)
    RT-->>GW: {message: "..."}
    GW-->>USR: routed response
    GW->>MEM: ConversationManager.store_turn()

Wiki Knowledge Base

Cogniverse automatically saves agent interactions as searchable wiki pages. Pages are stored in Vespa using hybrid search (semantic + BM25) and indexed per tenant.

Page Types

Type Description
Topic page Named page that grows over time — new content is appended each time the topic is mentioned. Stable doc_id based on the entity name slug.
Session page Point-in-time capture of a single agent interaction — one page per conversation. Cross-references the topic pages it touched.

A separate wiki_index document is maintained per tenant listing all pages and summaries.

Auto-Filing

After every agent dispatch, the system checks whether the interaction is substantial enough to auto-file as a wiki session. An interaction is filed automatically when any of the following is true:

  • 3 or more entities were extracted from the response
  • The agent is detailed_report_agent or deep_research_agent
  • The conversation has 4 or more turns

Auto-filing is fire-and-forget (non-blocking). Failures are logged but never surfaced to the user.

Auto-Filing Flow

flowchart TD
    AGENT["<span style='color:#000'>Agent Interaction<br/>(any agent dispatch)</span>"]
    CHECK["<span style='color:#000'>_should_auto_file()<br/>entities ≥ 3<br/>agent in AUTO_FILE_AGENTS<br/>turn_count ≥ 4</span>"]
    SKIP["<span style='color:#000'>Skip<br/>(interaction too brief)</span>"]
    WM["<span style='color:#000'>WikiManager<br/>save_session()</span>"]
    TOPIC["<span style='color:#000'>Topic Pages<br/>(upsert per entity)</span>"]
    SESSION["<span style='color:#000'>Session Page<br/>(point-in-time capture)</span>"]
    INDEX["<span style='color:#000'>wiki_index<br/>(rebuilt per tenant)</span>"]
    VESPA["<span style='color:#000'>Vespa wiki_pages schema<br/>hybrid search (semantic + BM25)</span>"]

    AGENT --> CHECK
    CHECK -->|"no"| SKIP
    CHECK -->|"yes"| WM
    WM --> TOPIC
    WM --> SESSION
    WM --> INDEX
    TOPIC --> VESPA
    SESSION --> VESPA
    INDEX --> VESPA

    style AGENT fill:#ce93d8,stroke:#7b1fa2,color:#000
    style CHECK fill:#ffcc80,stroke:#ef6c00,color:#000
    style SKIP fill:#b0bec5,stroke:#546e7a,color:#000
    style WM fill:#ce93d8,stroke:#7b1fa2,color:#000
    style TOPIC fill:#81c784,stroke:#388e3c,color:#000
    style SESSION fill:#81c784,stroke:#388e3c,color:#000
    style INDEX fill:#81c784,stroke:#388e3c,color:#000
    style VESPA fill:#90caf9,stroke:#1565c0,color:#000

Telegram /wiki Commands

Use these commands in Telegram to interact with the wiki directly:

Command Description
/wiki save Save the current session to the wiki
/wiki search <query> Search the wiki knowledge base
/wiki topic <name> Look up a topic page by name
/wiki index Show the full wiki index
/wiki lint Check wiki for orphan, stale, empty, or malformed pages

REST API

Endpoint Method Description
/wiki/save POST Persist an agent interaction as a wiki page
/wiki/search POST Full-text search over wiki pages
/wiki/topic/{slug} GET Retrieve a topic page by slug
/wiki/index GET Return the rendered wiki index
/wiki/lint GET Report orphan, stale, empty, and malformed pages
/wiki/topic/{slug} DELETE Delete a topic page by slug

The lint response includes malformed_pages entries for missing, invalid, or timezone-naive updated_at values. Each entry names the document, field, and stored value, and contributes to issues_found.

# Save a wiki page
curl -X POST http://localhost:28000/wiki/save \
  -H "Content-Type: application/json" \
  -d '{
    "query": "machine learning basics",
    "response": {"answer": "ML is..."},
    "entities": ["machine_learning"],
    "agent_name": "summarizer_agent",
    "tenant_id": "acme_corp"
  }'

# Search wiki pages
curl -X POST http://localhost:28000/wiki/search \
  -H "Content-Type: application/json" \
  -d '{"query": "machine learning", "tenant_id": "acme_corp", "top_k": 5}'

# Get a topic page
curl "http://localhost:28000/wiki/topic/machine_learning?tenant_id=acme_corp"

# Get the wiki index
curl "http://localhost:28000/wiki/index?tenant_id=acme_corp"

# Run lint checks
curl "http://localhost:28000/wiki/lint?tenant_id=acme_corp"

# Delete a topic page
curl -X DELETE \
  "http://localhost:28000/wiki/topic/machine_learning?tenant_id=acme_corp"

RLM (Recursive Language Model)

RLM enables agents to process context that exceeds normal token limits by recursively decomposing inputs using a Python REPL. Available on search, report, code, and research agents.

To activate, set the rlm field on the agent's typed input. Note: the unified runtime's REST shortcut (POST /agents/{name}/process) does not forward an rlm field — activate RLM via the Python SDK, which calls the agent's typed process() entrypoint directly:

import asyncio
from cogniverse_agents.detailed_report_agent import (
    DetailedReportAgent,
    DetailedReportDeps,
    DetailedReportInput,
)
from cogniverse_core.agents.rlm_options import RLMOptions
from cogniverse_foundation.config.utils import create_default_config_manager

async def main():
    config_manager = create_default_config_manager()
    deps = DetailedReportDeps()  # deps are tenant-agnostic; tenant_id is per-request
    agent = DetailedReportAgent(deps=deps, config_manager=config_manager)

    result = await agent.process(
        DetailedReportInput(
            query="Analyze these results",
            tenant_id="default",
            rlm=RLMOptions(enabled=True, max_iterations=5),
        )
    )
    print(result.rlm_synthesis)

asyncio.run(main())

RLM is opt-in and disabled by default. When enabled, telemetry metrics (depth, calls, tokens, latency) are included in the response for A/B testing.


Knowledge Management

Cogniverse ships a full Knowledge Management Layer built on top of Mem0+Vespa. Every memory write is governed by a KnowledgeSchema that controls retention, sensitivity, pin authority, provenance requirement, contradiction policy, and default trust.

Schema-driven retention

Retention Behaviour
PERMANENT Never auto-deleted. Default when the kind is unregistered.
EPHEMERAL_SESSION Cleared when the session ends (via DELETE /admin/tenants/{tenant_id}/sessions/{session_id}).
EPHEMERAL_DAYS(N) Soft-deleted at N days, hard-deleted at 2N days. Restorable inside the soft-delete window.
SCHEMA_DRIVEN Custom cleanup_hook on the schema.

Provenance and citations

Every write can carry a Provenance record describing who wrote the memory, how it was derived (direct_ingest, extraction, synthesis, etc.), and which source memories or external URLs it cites. Use the CitationTracingAgent to walk the chain back to primary sources.

Contradiction detection

When two memories disagree about the same subject, ContradictionDetector groups them into a ConflictSet. The schema's contradiction_policy resolves conflicts at retrieval time: latest_wins, trust_ranked, or preserve_both (all copies surfaced with metadata["disputed"]=True). Use ContradictionReconciliationAgent to surface and resolve open conflict sets.

Trust ranking

Trust is derived from the schema's default_trust and the write's derivation_kind. It ages slowly (≈0.005 pt/day above baseline), and can be boosted by user/admin endorsements. At retrieval, results are ranked by relevance × trust × confidence. Direct human assertions (user_assert) outrank agent inferences by default.

Federation (org trunk + tenant overlays)

FederatedQueryAgent reads from both the caller's tenant and the org's shared trunk in one call, with tenant overlay winning on collision. KnowledgeSummarizationAgent can promote a summary into the org trunk so all tenants in the same org see it.

Pinning

Memories can be pinned (by users, tenant admins, or org admins, each with quota limits) so they survive lifecycle cleanup and trust decay. Pinned memories are never auto-deleted. Use PinService from code or via the admin API.

Knowledge agents

Nine specialized agents operate on the knowledge layer:

Agent What it does
AuditExplanationAgent Explains why an answer was produced (provenance + trust + contradictions)
CitationTracingAgent Walks provenance chains back to primary sources
ContradictionReconciliationAgent Surfaces and resolves conflict sets
FederatedQueryAgent Queries tenant + org-trunk in one call
KnowledgeGraphTraversalAgent Traverses the knowledge graph by entity and relationship
KnowledgeSummarizationAgent Summarizes a knowledge slice with citations
MultiDocumentSynthesisAgent Synthesizes across multiple source documents
TemporalReasoningAgent Answers questions about knowledge change over time
CrossTenantComparisonAgent Compares knowledge views across tenants (org-admin scoped)

For full API details see Core Module — Memory Management and Agents Module — Knowledge Agents.


API Reference

REST API Endpoints

Search Endpoint

POST /search/

Request:

{
  "query": "string",
  "top_k": 10,
  "strategy": "hybrid_float_bm25",
  "profile": "video_colpali_smol500_mv_frame",
  "tenant_id": "default",
  "filters": {}
}

Response:

{
  "query": "string",
  "profile": "video_colpali_smol500_mv_frame",
  "strategy": "hybrid_float_bm25",
  "results_count": 10,
  "results": [
    {
      "document_id": "string",
      "score": 0.95,
      "metadata": {},
      "highlights": {}
    }
  ],
  "session_id": null
}

Ingestion Endpoint

POST /ingestion/start

Request:

{
  "video_dir": "/path/to/videos",
  "profile": "video_colpali_smol500_mv_frame",
  "backend": "vespa",
  "tenant_id": "default",
  "batch_size": 10
}

Response:

{
  "job_id": "abc123",
  "status": "started",
  "message": "Ingestion job started successfully"
}

Check Status:

GET /ingestion/status/{job_id}

Status Response:

{
  "job_id": "abc123",
  "status": "processing",
  "videos_processed": 5,
  "videos_total": 10,
  "errors": []
}

Health Check

GET /health

Response (the chat LLM's endpoint answered 404: nothing is deployed for the model):

{
  "status": "degraded",
  "service": "cogniverse-runtime",
  "backends": {
    "registered": 1,
    "backends": ["vespa"]
  },
  "agents": {
    "registered": 3,
    "agents": ["search_agent", "gateway_agent", "orchestrator_agent"]
  },
  "dependencies": {
    "llm": {
      "status": "not_serving",
      "endpoints": [
        {
          "endpoint": "http://cogniverse-semantic-router-envoy:8801/v1",
          "model": "openai/cogniverse-classification",
          "route": "pro",
          "state": "not_serving",
          "upstream_status": 404,
          "failure": null,
          "reason": "answered HTTP 404: nothing is deployed for this model; calls fail fast until the next recheck",
          "observed_at": "2026-10-01T17:10:07.512301+00:00",
          "recheck_in_s": 21.4
        }
      ]
    }
  }
}

status is healthy, degraded (the chat LLM is not_serving or failing; search still serves) or, with HTTP 503, unhealthy (the search backend is unreachable). dependencies.llm.status is not_called until this worker process has made an LM call.

Python SDK Reference

SearchAgent

from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)

# search_by_text — tenant_id is per-request
results = agent.search_by_text(
    query="machine learning",
    tenant_id="your_org:production",
    top_k=10,
    start_date="2024-01-01",  # Optional
    end_date="2024-12-31",    # Optional
)

OrchestratorAgent

from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps, OrchestratorInput
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_foundation.config.utils import create_default_config_manager

config_manager = create_default_config_manager()
registry = AgentRegistry(tenant_id="your_org:production", config_manager=config_manager)
orchestrator = OrchestratorAgent(deps=OrchestratorDeps(), registry=registry, config_manager=config_manager)

# Orchestrator plans preprocessing + execution agents and returns OrchestratorOutput
result = await orchestrator._process_impl(
    OrchestratorInput(query="machine learning tutorial", tenant_id="your_org:production")
)

# OrchestratorOutput.final_output contains the aggregated execution result.
# Enrichment (enhanced_query, entities, relationships, query_variants) is threaded
# from preprocessing agent outputs onto execution agent inputs via AgentTask fields
# by the module-level _merge_enrichment() helper — not returned as a top-level field.
print(result.execution_summary)

Mem0MemoryManager

from cogniverse_core.memory.manager import Mem0MemoryManager
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# Initialize required dependencies first
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

# Instantiate (per-tenant singleton pattern)
memory = Mem0MemoryManager(tenant_id="your_org:production")

# Initialize with all required parameters
memory.initialize(
    backend_host="localhost",
    backend_port=8080,
    llm_model="openai/google/gemma-4-e4b-it",
    embedding_model="lightonai/DenseOn",
    llm_base_url="http://localhost:11434",
    embedder_base_url="http://localhost:29006",
    config_manager=config_manager,
    schema_loader=schema_loader,
)

# Add memory (requires tenant_id and agent_name)
memory.add_memory(
    content="User prefers Python tutorials",
    tenant_id="your_org:production",
    agent_name="search_agent"
)

# Search memory
relevant_memories = memory.search_memory(
    query="tutorial preferences",
    tenant_id="your_org:production",
    agent_name="search_agent",
    top_k=5
)

# Get all memories
all_memories = memory.get_all_memories(tenant_id="your_org:production", agent_name="search_agent")

Configuration

System Configuration

Configuration is loaded from configs/config.json. The system auto-discovers this file from: 1. COGNIVERSE_CONFIG environment variable (if set) 2. configs/config.json (from current directory) 3. ../configs/config.json (one level up) 4. ../../configs/config.json (two levels up)

See Profile Configuration below for the actual config.json structure.

Profile Configuration

Configure embedding profiles in configs/config.json:

{
  "backend": {
    "profiles": {
      "video_colpali_smol500_mv_frame": {
        "type": "video",
        "description": "Frame-based ColPali profile",
        "schema_name": "video_colpali_smol500_mv_frame",
        "embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
        "pipeline_config": {
          "extract_keyframes": true,
          "keyframe_strategy": "fps",
          "keyframe_fps": 0.5,
          "transcribe_audio": true,
          "generate_descriptions": true,
          "generate_embeddings": true
        },
        "strategies": {
          "segmentation": {
            "class": "FrameSegmentationStrategy",
            "params": {
              "fps": 0.5,
              "threshold": 0.999,
              "max_frames": 3000
            }
          },
          "embedding": {
            "class": "MultiVectorEmbeddingStrategy",
            "params": {}
          }
        },
        "embedding_type": "multi_vector"
      }
    }
  }
}

Environment Variables

The following environment variables are honored by the system:

# Configuration File Discovery
export COGNIVERSE_CONFIG=/path/to/config.json  # Override config file path

# JAX Configuration (required for X-CLIP models)
export JAX_PLATFORM_NAME=cpu  # Required on Apple Silicon or systems without GPU

# HuggingFace (for model downloads)
export HF_TOKEN=your_token_here  # HuggingFace access token for gated models

Note: Most configuration is done via configs/config.json. Environment variables are minimal - primarily JAX_PLATFORM_NAME for X-CLIP compatibility and COGNIVERSE_CONFIG to override the config file location.


Troubleshooting

Common Issues

Issue: "ModuleNotFoundError: No module named 'cogniverse_core'"

Solution:

cd /path/to/cogniverse
uv sync
source .venv/bin/activate  # or .venv\Scripts\activate on Windows

Issue: "Vespa connection refused"

Solution:

# Check if Vespa is running
cogniverse status

# Check logs
cogniverse logs vespa

# Vespa runs as a k3d-managed StatefulSet pod, not a standalone
# container — restart via kubectl, not `docker restart`:
kubectl rollout restart statefulset/cogniverse-vespa -n cogniverse

Issue: "Phoenix not recording spans"

Solution:

# Verify Phoenix endpoint
echo $TELEMETRY_OTLP_ENDPOINT

# Should be: localhost:4317 (gRPC)

# Test connectivity
curl http://localhost:26006/health

Issue: "Out of memory during ingestion"

Solution:

# Reduce concurrent processing
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_colpali_smol500_mv_frame \
  --max-concurrent 1  # Reduce from default 3

# Or use binary embeddings via ranking strategies
# Binary embeddings are configured in schema_config and used via ranking strategies

Issue: "Slow search performance"

Solutions:

  1. Use binary embeddings instead of float
  2. Enable caching in config.json
  3. Use BM25-only for text queries
  4. Reduce top_k to get fewer results
# Binary embeddings are configured per-profile in schema_config.binary_dim
# Use standard profiles - they support both float and binary embeddings
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id default

Debug Mode

Enable debug logging by configuring the logging level in your Python script or using standard Python logging configuration:

# Run ingestion with verbose output
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/videos \
  --profile video_colpali_smol500_mv_frame \
  --tenant-id default

# Check logs
tail -f outputs/logs/*.log

Note: Logging levels are configured programmatically or via configs/config.json, not via environment variables.

Getting Help

  • Documentation: Home
  • GitHub Issues: Report bugs
  • Web Client: http://localhost:28400
  • Phoenix UI: http://localhost:26006 for traces and experiments
  • API Docs: http://localhost:28000/docs for interactive API documentation

Best Practices

For Content Managers

  1. Use Multiple Profiles: Ingest with ColPali (frames), X-CLIP (global), and ColQwen (chunks) for best coverage
  2. Enable Transcription: Always transcribe audio for text search
  3. Monitor Quality: Run evaluations weekly to track search quality
  4. Organize by Tenant: Use separate tenants for different content libraries

For Data Scientists

  1. Track Experiments: Always run evaluations through Phoenix for reproducibility
  2. Use Synthetic Data: Enable synthetic data generation for routing optimization
  3. Compare Strategies: Test multiple search strategies on your dataset
  4. Monitor Drift: Track query distribution changes in Phoenix

For Developers

  1. Use SDK: Prefer Python SDK over direct API calls for better error handling
  2. Handle Errors: Always catch and handle exceptions
  3. Implement Caching: Cache frequent queries at application level
  4. Test Multi-Tenant: Test with multiple tenants to ensure isolation

Performance Tips

  1. Binary Embeddings: Use binary embeddings for 4x faster search with minimal accuracy loss
  2. Batch Ingestion: Process videos in batches of 10-20 for optimal throughput
  3. Enable Caching: Enable LRU cache for repeated queries
  4. Use BM25 First: For pure text queries, use BM25-only strategy
  5. Prewarm Cache: Warm up caches with common queries after ingestion

Next Steps

For Users

Developer Resources

For DevOps