Cogniverse User Guide¶
Complete guide for using Cogniverse - the general-purpose multi-agent AI platform for content intelligence and beyond.
Table of Contents¶
- Introduction
- Getting Started
- Core Features
- Basic Operations
- Advanced Usage
- API Reference
- Configuration
- Troubleshooting
- Best Practices
Introduction¶
Key Capabilities¶
For Content Managers:
- Ingest and index video libraries with multiple embedding strategies
- Search across videos using natural language queries
- Get visual relevance scores and frame-level results
- Monitor performance with built-in dashboards
For Data Scientists:
- Run experiments with different embedding models and search strategies
- Evaluate search quality with reference-free and visual LLM metrics
- Optimize routing agents using synthetic data generation
- Track all experiments in Phoenix with full observability
For Developers:
- RESTful API for all operations
- Multi-tenant support with tenant isolation
- Configurable embedding profiles and search strategies
- Plugin architecture for custom components
Architecture at a Glance¶
See the full Platform Overview for the architecture diagram showing the complete system including agents, content store, memory, telemetry, optimization loop, and training pipeline.
Getting Started¶
Prerequisites¶
Before using Cogniverse, ensure you have:
- Python 3.12+ installed
- 16GB+ RAM (32GB recommended for large video libraries)
- Docker for running Vespa, Phoenix, and Ollama
- GPU (optional but recommended for video processing)
- uv package manager:
pip install uv
Quick Start¶
Follow these steps to get Cogniverse running in 5 minutes:
1. Installation¶
# Clone repository
git clone <repository-url>
cd cogniverse
# Install all dependencies
uv sync
# Start infrastructure services
cogniverse up
2. Verify Services¶
# Check Vespa (should return JSON status)
curl http://localhost:8080/ApplicationStatus
# Check Phoenix (should return "ok")
curl http://localhost:26006/health
# Check Ollama (should return list of models)
curl http://localhost:11434/api/tags
3. Ingest Sample Videos¶
# Ingest videos with ColPali frame-level embeddings
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/testset/evaluation/sample_videos \
--profile video_colpali_smol500_mv_frame \
--tenant-id default
4. Run Your First Query¶
# Test multi-agent search
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
--profiles video_colpali_smol500_mv_frame \
--test-multiple-strategies
Success! You should see search results with ranked videos and relevance scores.
Next Steps¶
- View Results: Open the web client at http://localhost:28400 and the Phoenix UI at http://localhost:26006
- Try API: Use the REST API at http://localhost:28000/docs
- Configure: Customize profiles in
configs/config.json - Web UI: Chat with any registered agent and manage the stack in the browser — see the web client
- Coding Agent: Run
cogniverse codeto start an interactive coding REPL with streaming — see Coding Agent CLI - Knowledge Graph: Run
cogniverse index ./path --type codeto build a searchable knowledge graph — see Knowledge Graph - Learn More: Continue reading this guide
CLI Reference¶
The cogniverse CLI manages the full stack:
| Command | Purpose |
|---|---|
cogniverse up | Deploy all services (Vespa, Phoenix, LLM, Runtime, Web client) via k3d |
cogniverse up --messaging | Deploy with Telegram gateway enabled |
cogniverse down | Stop all services |
cogniverse down --keep-data | Stop services but preserve volumes |
cogniverse status | Show health of all services |
cogniverse logs <service> | View logs (runtime, web, vespa, phoenix, llm, argo) |
cogniverse logs <service> --follow | Stream logs in real-time |
cogniverse code | Interactive coding agent REPL |
cogniverse index <path> --type code | Build a knowledge graph from code |
cogniverse graph stats | Show knowledge graph statistics |
cogniverse graph search <query> | Search the knowledge graph |
cogniverse graph neighbors <node> | Find related nodes |
cogniverse graph path <src> <dst> | Find path between nodes |
cogniverse sandbox sync | Sync OpenShell gateway mTLS certs into the cluster (after rotation) |
cogniverse sandbox status | Show OpenShell gateway status and cluster sync state |
cogniverse secrets sync | Re-sync cluster Secrets (hf-token, cogniverse-messaging-secrets) from env vars, ./.env, or ~/.env |
cogniverse admin reconcile-orphans | Find (and, with --confirm, drop) Vespa schema orphans not in the schema registry |
cogniverse admin merge-article-nodes | Report (and, with --apply, perform) merges of KG nodes the_<id> / a_<id> / an_<id> into the tenant's <id> node; --tenant scopes to one tenant; --exclude ID (repeatable) never merges that article id |
cogniverse admin invite <tenant_id> | Mint a messaging invite token; prints the /start <token> the user sends to the bot |
Core Features¶
1. Multi-Modal Video Search¶
Search videos using different modalities:
Text-to-Video Search¶
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Create agent — profile sets the default embedding model
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# Search by text (synchronous) — tenant_id is per-request
results = agent.search_by_text(
query="machine learning tutorial",
tenant_id="your_org:production",
top_k=10,
)
for result in results:
print(f"Video: {result.get('video_id', 'unknown')}")
print(f"Score: {result.get('score', 0):.2f}")
Multi-Profile Search¶
# Search across different embedding profiles using ensemble mode
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps, SearchInput
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Create agent with a default profile
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# Single profile search
colpali_results = agent.search_by_text(
query="cooking tutorial",
tenant_id="your_org:production",
top_k=10,
)
# Search with a different profile via SearchInput for ensemble
xclip_deps = SearchAgentDeps(profile="video_xclip_sv_chunk_6s")
xclip_agent = SearchAgent(deps=xclip_deps, config_manager=config_manager, schema_loader=schema_loader)
xclip_results = xclip_agent.search_by_text(
query="cooking tutorial",
tenant_id="your_org:production",
top_k=10,
)
Date-Filtered Search¶
# Search with date filters
results = agent.search_by_text(
query="machine learning tutorial",
tenant_id="your_org:production",
top_k=10,
start_date="2024-01-01",
end_date="2024-12-31",
)
2. Intelligent Query Routing¶
Cogniverse automatically routes queries to the optimal execution agent via the GatewayAgent → OrchestratorAgent pipeline. The orchestrator plans and executes preprocessing agents (QueryEnhancementAgent, EntityExtractionAgent, ProfileSelectionAgent) and threads their enrichment outputs — enhanced_query, entities, relationships, query_variants — directly onto the execution agent's typed input via AgentTask:
import asyncio
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps, OrchestratorInput
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_foundation.config.utils import create_default_config_manager
async def main():
config_manager = create_default_config_manager()
registry = AgentRegistry(tenant_id="your_org:production", config_manager=config_manager)
orchestrator = OrchestratorAgent(deps=OrchestratorDeps(), registry=registry, config_manager=config_manager)
result = await orchestrator._process_impl(
OrchestratorInput(
query="cooking recipes with pasta",
tenant_id="your_org:production",
)
)
print(result.execution_summary)
asyncio.run(main())
Routing Features:
- Entity Extraction: Identifies people, places, concepts using GLiNER
- Relationship Detection: Finds relationships between entities
- Query Enhancement: Enriches queries with context
- Modality Detection: Classifies query type (factual, conceptual, visual, etc.)
- Confidence Scoring: Provides routing confidence for transparency
3. Multiple Embedding Models¶
Choose the best embedding model for your use case:
| Model | Type | Best For | Dimensions |
|---|---|---|---|
| ColPali | Frame-level | Visual documents, text-rich videos | 320 (patch) |
| X-CLIP | Temporal video | Text-to-video clip retrieval | 768 |
| ColQwen3 Omni | Multi-modal | Text+visual fusion | 320 (patch) |
Switching Models:
# Use X-CLIP for global video understanding
uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_xclip_sv_chunk_6s \
--tenant-id default
4. Hybrid Search Strategies¶
Combine multiple search methods for better results:
# Ranking strategies are plain rank-profile-name strings (validated against
# the deployed Vespa schema), not an enum:
strategies = [
"bm25_only", # Text-only BM25 (fastest for keyword queries)
"float_float", # Dense float embeddings (highest visual accuracy)
"binary_binary", # Binary embeddings (fastest visual search)
"float_binary", # Float query, binary index (speed/accuracy balance)
"phased", # Two-phase: binary retrieval, float reranking
"hybrid_float_bm25", # Visual + text hybrid (best overall accuracy)
"hybrid_binary_bm25", # Fast hybrid (binary visual + text)
"hybrid_bm25_binary", # Text matches ranked by binary visual + text
"hybrid_bm25_float", # Text matches ranked by float visual + text
# plus the "_no_description" hybrid variants for ColPali/ColQwen schemas
]
# A strategy string is passed at query time via the `strategy` field of the
# /search/ request (or the query_dict given to VespaSearchBackend.search).
# The SearchAgent handles this automatically based on profile configuration.
5. Memory-Aware Search¶
Cogniverse remembers user context for personalized results:
from cogniverse_core.memory.manager import Mem0MemoryManager
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# Initialize required dependencies
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Get memory manager (singleton per tenant via __new__)
memory = Mem0MemoryManager(tenant_id="your_org:production")
# Initialize with required parameters
memory.initialize(
backend_host="localhost",
backend_port=8080,
llm_model="openai/google/gemma-4-e4b-it",
embedding_model="lightonai/DenseOn",
llm_base_url="http://localhost:11434",
embedder_base_url="http://localhost:29006",
config_manager=config_manager,
schema_loader=schema_loader,
)
# Add user preference
memory.add_memory(
content="User prefers beginner-level Python tutorials",
tenant_id="your_org:production",
agent_name="search_agent",
metadata={"user_id": "user_123"}
)
# Search memories
user_memories = memory.search_memory(
query="Python tutorial preferences",
tenant_id="your_org:production",
agent_name="search_agent",
top_k=5
)
6. Comprehensive Telemetry¶
Track everything with Phoenix:
# Phoenix UI: traces, spans and experiments
open http://localhost:26006
# Web client: analytics, evaluation and memory views over the same telemetry
open http://localhost:28400
Web client views over telemetry:
- Analytics: Traces with latency over time, histograms, outliers and root causes
- Evaluation: Golden-set search quality per profile and strategy
- Routing evaluation: Routing decisions, accuracy and confidence calibration
- Memory: View, search, add and delete stored memories
- Embedding atlas: Documents placed by their embeddings
Basic Operations¶
Video Ingestion¶
Single Profile Ingestion¶
Ingest videos with one embedding model:
# Ingest with ColPali (frame-based)
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_colpali_smol500_mv_frame \
--tenant-id default
Multi-Profile Ingestion¶
Ingest with multiple models for best coverage:
# Ingest with ColPali, X-CLIP, and ColQwen
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_colpali_smol500_mv_frame \
video_xclip_sv_chunk_6s \
video_colqwen_omni_mv_chunk_30s \
--tenant-id default
Ingestion Options:
--video_dir: Directory containing content files--content-dir: Alias for--video_dir--media-root-uri: Media root URI (e.g.s3://corpus/,pvc://media/) for non-filesystem sources; overrides--video_dir/--content-dirwhen set--content-type: Content type to ingest (choices: video, image, audio, document; default: video)--profile: Embedding profile(s) - can specify multiple space-separated values--tenant-id: Tenant ID for multi-tenancy (required — no default)--backend: Backend to use (choices: byaldi, vespa; default: vespa)--max-concurrent: Maximum concurrent items to process (default: 3)--output_dir: Output directory for processed data--max-frames: Maximum frames per video / images per batch--test-mode: Use test mode with limited frames--debug: Enable debug mode
Note: Processing options like keyframe extraction, transcription, and frame sampling are configured per-profile in configs/config.json, not via CLI arguments.
Check Ingestion Status¶
# Confirm documents are searchable via the production /search/ endpoint.
# A non-empty results list means content was indexed for the tenant.
curl -X POST http://localhost:28000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "test",
"tenant_id": "acme",
"profile": "video_colpali_smol500_mv_frame",
"top_k": 1
}'
# For an async ingestion job, poll its status instead:
# curl http://localhost:28000/ingestion/status/{job_id}
Searching Videos¶
REST API Search¶
Use the REST API for production applications:
# Text search
curl -X POST http://localhost:28000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "machine learning tutorial",
"top_k": 10,
"strategy": "hybrid_float_bm25",
"tenant_id": "default"
}'
Response:
{
"query": "machine learning tutorial",
"profile": "video_colpali_smol500_mv_frame",
"strategy": "hybrid_float_bm25",
"results_count": 10,
"results": [
{
"document_id": "doc_123",
"score": 0.95,
"metadata": {
"source_id": "video_123",
"frame_id": "frame_45"
},
"highlights": {}
}
],
"session_id": null
}
Python SDK Search¶
Use the Python SDK for scripting:
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# search_by_text is synchronous — no await needed
results = agent.search_by_text(
query="cooking pasta",
tenant_id="your_org:production",
top_k=10,
)
for result in results:
print(f"Video: {result.get('video_id', 'unknown')}")
print(f"Score: {result.get('score', 0):.2f}")
Advanced Search Options¶
# Search with date filters
results = agent.search_by_text(
query="tutorial",
tenant_id="your_org:production",
top_k=10,
start_date="2024-01-01",
end_date="2024-12-31",
)
# Search with more results for client-side filtering
results = agent.search_by_text(
query="Python tutorial",
tenant_id="your_org:production",
top_k=50,
)
Running Evaluations¶
Quick Evaluation¶
Evaluate search quality on a dataset:
# Run evaluation with Phoenix tracking
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--dataset-name golden_eval_v1 \
--profiles video_colpali_smol500_mv_frame \
--all-strategies \
--quality-evaluators
This will:
- Load evaluation dataset with ground truth
- Run queries through all search strategies
- Compute quality metrics (MRR, NDCG, Precision@K)
- Track results in Phoenix
- Generate comparison charts
Custom Evaluation Dataset¶
Create your own evaluation dataset as a CSV file with query, expected_videos (comma-separated video IDs), and an optional category column:
query,expected_videos,category
machine learning basics,"video_123,video_456",general
Python tutorial for beginners,video_789,general
# Run evaluation on custom dataset
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--csv-path evaluation_dataset.csv \
--profiles video_colpali_smol500_mv_frame
Compare Embedding Models¶
Compare different embedding models:
# Run experiments with multiple profiles
JAX_PLATFORM_NAME=cpu uv run python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--dataset-name golden_eval_v1 \
--profiles video_colpali_smol500_mv_frame \
video_xclip_sv_chunk_6s \
video_colqwen_omni_mv_chunk_30s \
--all-strategies
Results Table:
Profile | MRR@10 | NDCG@10 | Precision@5
-------------------------------------|--------|---------|-------------
video_colpali_smol500_mv_frame | 0.82 | 0.79 | 0.75
video_xclip_sv_chunk_6s | 0.78 | 0.74 | 0.70
video_colqwen_omni_mv_chunk_30s | 0.85 | 0.82 | 0.78
Advanced Usage¶
Multi-Tenant Configuration¶
Set up multiple tenants with isolated data:
# Multi-tenancy is handled at the request level, not via separate config objects.
# SystemConfig is global infrastructure config (one per deployment).
# Tenant isolation is achieved by passing tenant_id per request:
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# Each tenant_id produces isolated results:
# - Vespa documents are filtered by a tenant-scoped schema. A bare tenant_id
# like "acme_corp" is canonicalized to "acme_corp:acme_corp" (org:tenant),
# giving the schema name video_colpali_smol500_mv_frame_acme_corp_acme_corp
# (use "org:tenant" form, e.g. "acme_corp:production", to avoid the
# doubled suffix).
# - Memory: separate Mem0 namespaces
results_acme = agent.search_by_text(query="tutorial", tenant_id="acme_corp", top_k=10)
results_startup = agent.search_by_text(query="tutorial", tenant_id="startup_inc", top_k=10)
Tenant Lifecycle:
# Ingest data for tenant (tenant is implicitly created on first use)
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/new_customer_videos \
--profile video_colpali_smol500_mv_frame \
--tenant-id new_customer
# Explicit tenant creation/deletion APIs are also available directly on
# the main runtime:
# curl -X POST http://localhost:28000/admin/tenants \
# -d '{"tenant_id": "new_customer", "created_by": "admin"}'
# curl -X DELETE http://localhost:28000/admin/tenants/new_customer
Custom Embedding Profiles¶
Create custom profiles for specific use cases:
{
"backend": {
"profiles": {
"video_custom_highres_frame": {
"type": "video",
"description": "Custom high-resolution frame-based profile",
"schema_name": "video_custom_highres_frame",
"embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
"pipeline_config": {
"extract_keyframes": true,
"keyframe_strategy": "fps",
"keyframe_fps": 2.0,
"transcribe_audio": true,
"generate_descriptions": true,
"generate_embeddings": true
},
"strategies": {
"segmentation": {
"class": "FrameSegmentationStrategy",
"params": {
"fps": 2.0,
"threshold": 0.999,
"max_frames": 200
}
},
"embedding": {
"class": "MultiVectorEmbeddingStrategy",
"params": {}
}
},
"embedding_type": "multi_vector"
}
}
}
}
# Deploy custom profile schema
JAX_PLATFORM_NAME=cpu uv run python scripts/deploy_json_schema.py \
configs/schemas/video_custom_highres_frame.json
# Ingest with custom profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_custom_highres_frame \
--tenant-id default
DSPy Optimization¶
Optimize routing and search agents via the optimization CLI:
# Optimize gateway confidence thresholds
python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id default
# Optimize entity extraction
python -m cogniverse_runtime.optimization_cli --mode entity-extraction --tenant-id default
# Optimize profile performance
python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id default
# Full optimization workflow (all modes)
python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id default
# Triggered optimization (run when quality degrades)
python -m cogniverse_runtime.optimization_cli --mode triggered \
--tenant-id default --agents search_agent,summarizer_agent \
--trigger-dataset optimization-trigger-default-20260403_040000
# Cleanup old logs
python -m cogniverse_runtime.optimization_cli --mode cleanup --log-retention-days 7
Available optimization modes:
| Mode | What It Optimizes |
|---|---|
gateway-thresholds | GLiNER confidence thresholds |
entity-extraction | Entity extraction accuracy |
profile | Profile performance ranking |
workflow | Full end-to-end optimization pipeline |
triggered | On-demand when quality monitor fires |
simba | Query enhancement (SIMBA), from cogniverse.query_enhancement spans |
online-routing-eval | Scores cogniverse.routing spans (routing_outcome + confidence) without retraining |
synthetic | Generates synthetic training data for one or more optimizer types (simba, profile, workflow by default) |
rollback | Restores an agent's active prompts/demos artifact to a previously-snapshotted version |
ab-compare | Runs an RLM A/B comparison over a Phoenix queries dataset |
egress-netpol | Emits Kubernetes NetworkPolicy CRDs from per-agent policy YAMLs (no --tenant-id) |
monthly-reports | Generates the monthly usage + performance report (no --tenant-id) |
cleanup | Purge old optimization logs and expired memories (no --tenant-id) |
Batch Processing¶
Process large video libraries efficiently using the CLI:
# Process videos in batches using the ingestion script
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/batch1 \
--profile video_colpali_smol500_mv_frame \
--tenant-id default
# Or via the API
curl -X POST http://localhost:28000/ingestion/start \
-H "Content-Type: application/json" \
-d '{
"video_dir": "data/batch1",
"profile": "video_colpali_smol500_mv_frame",
"tenant_id": "default",
"batch_size": 10
}'
# Check status
curl http://localhost:28000/ingestion/status/{job_id}
API Authentication¶
Authentication is handled via tenant isolation. Each request includes a tenant_id:
# Search with tenant ID
curl -X POST http://localhost:28000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "tutorial",
"top_k": 10,
"tenant_id": "acme_corp"
}'
Note: Tenant isolation provides logical separation of data. API key authentication is planned for future releases.
Telegram Messaging Gateway¶
Users can interact with Cogniverse via Telegram after receiving an invite token:
Admin: generate an invite token
curl -X POST http://localhost:28000/admin/messaging/invite \
-H "Content-Type: application/json" \
-d '{"tenant_id": "acme_corp", "expires_in_hours": 24}'
# Returns: {"token": "abc123def456...", "tenant_id": "acme_corp"}
User: register and use the bot
- Send
/start abc123def456...to the bot to link your account - Use commands to interact with agents:
| Command | Agent | Example |
|---|---|---|
/search <query> | search_agent | /search machine learning tutorial |
/summarize <query> | summarizer_agent | /summarize the Python basics video |
/report <query> | detailed_report_agent | /report Q4 content performance |
/research <query> | deep_research_agent | /research best practices for async Python |
/code <query> | coding_agent | /code write a FastAPI health endpoint |
/wiki save | — | Save current session to the wiki |
/wiki search <query> | — | Search the wiki knowledge base |
/wiki topic <name> | — | Look up a topic page by name |
/wiki index | — | Show the full wiki index |
/wiki lint | — | Check wiki for orphan, stale, empty, or malformed pages |
/instructions set <text> | — | Set custom agent instructions for your tenant |
/instructions show | — | Show current tenant instructions |
/memories list | — | List memories (add agent=<name> to filter) |
/memories clear strategies | — | Clear strategy learner memories |
/jobs list | — | List scheduled agent jobs |
/jobs create "<cron>" <query> | — | Create a new scheduled job |
/jobs delete <job_id> | — | Delete a scheduled job |
| Plain text | gateway_agent | what videos do you have on transformers? |
| Photo/video | search_agent | Send a frame to search for similar content |
/help | — | Show all available commands |
Conversation history is maintained via Mem0 across sessions. The gateway runs in polling mode for development (GATEWAY_MODE=polling) and webhook mode for production (GATEWAY_MODE=webhook with TELEGRAM_WEBHOOK_URL set).
Gateway Architecture¶
flowchart TD
TG["<span style='color:#000'>Telegram User</span>"]
BOT["<span style='color:#000'>Telegram Bot API<br/>(webhook / polling)</span>"]
GW["<span style='color:#000'>MessagingGateway</span>"]
CR["<span style='color:#000'>command_router<br/>parse_message()</span>"]
AUTH["<span style='color:#000'>InviteTokenManager<br/>claim_token()</span>"]
UM["<span style='color:#000'>UserTenantMapper<br/>get_tenant_id()</span>"]
CM["<span style='color:#000'>ConversationManager<br/>get_history() / store_turn()</span>"]
RC["<span style='color:#000'>RuntimeClient<br/>POST /agents/{name}/process</span>"]
FMT["<span style='color:#000'>format_agent_response()<br/>chunk at 4096 chars</span>"]
TG -->|"sends message"| BOT
BOT -->|"Update"| GW
GW --> CR
CR -->|"ParsedCommand<br/>(agent_name, query)"| GW
GW --> AUTH
AUTH -->|"tenant_id"| UM
UM -->|"tenant confirmed"| GW
GW --> CM
CM -->|"conversation history"| RC
RC -->|"agent response"| FMT
FMT -->|"chunked messages"| TG
style TG fill:#81d4fa,stroke:#0288d1,color:#000
style BOT fill:#90caf9,stroke:#1565c0,color:#000
style GW fill:#ce93d8,stroke:#7b1fa2,color:#000
style CR fill:#a5d6a7,stroke:#388e3c,color:#000
style AUTH fill:#ffcc80,stroke:#ef6c00,color:#000
style UM fill:#ffcc80,stroke:#ef6c00,color:#000
style CM fill:#b0bec5,stroke:#546e7a,color:#000
style RC fill:#64b5f6,stroke:#1565c0,color:#000
style FMT fill:#a5d6a7,stroke:#388e3c,color:#000 End-to-End User Flow¶
sequenceDiagram
participant ADM as Admin
participant RT as Runtime API
participant USR as Telegram User
participant BOT as Telegram Bot API
participant GW as MessagingGateway
participant MEM as Mem0 Memory
ADM->>RT: POST /admin/messaging/invite<br/>{tenant_id, expires_in_hours}
RT-->>ADM: {token: "abc123..."}
ADM->>USR: share invite token out-of-band
USR->>BOT: /start abc123...
BOT->>GW: Update (start command + token)
GW->>RT: POST /admin/messaging/register<br/>{platform, external_user_id, token}
RT->>RT: InviteTokenManager.claim_token()
RT->>MEM: UserTenantMapper.register_user()
RT->>RT: InviteTokenManager.mark_token_used()
RT-->>GW: {tenant_id}
GW-->>USR: "Registered as acme_corp."
USR->>BOT: /search machine learning tutorial
BOT->>GW: Update (search command)
GW->>GW: parse_message() → search_agent
GW->>MEM: ConversationManager.get_history(chat_id)
MEM-->>GW: prior turns
GW->>RT: POST /agents/search_agent/process
RT-->>GW: {results: [...], message: "..."}
GW->>GW: format_agent_response() → chunk at 4096 chars
GW-->>USR: search results
GW->>MEM: ConversationManager.store_turn(user + assistant)
USR->>BOT: what else do you have on this topic?
BOT->>GW: Update (plain text)
GW->>GW: parse_message() → gateway_agent
GW->>MEM: ConversationManager.get_history(chat_id)
MEM-->>GW: prior turns (multi-turn context)
GW->>RT: POST /agents/gateway_agent/process<br/>(with conversation_history)
RT-->>GW: {message: "..."}
GW-->>USR: routed response
GW->>MEM: ConversationManager.store_turn() Wiki Knowledge Base¶
Cogniverse automatically saves agent interactions as searchable wiki pages. Pages are stored in Vespa using hybrid search (semantic + BM25) and indexed per tenant.
Page Types¶
| Type | Description |
|---|---|
| Topic page | Named page that grows over time — new content is appended each time the topic is mentioned. Stable doc_id based on the entity name slug. |
| Session page | Point-in-time capture of a single agent interaction — one page per conversation. Cross-references the topic pages it touched. |
A separate wiki_index document is maintained per tenant listing all pages and summaries.
Auto-Filing¶
After every agent dispatch, the system checks whether the interaction is substantial enough to auto-file as a wiki session. An interaction is filed automatically when any of the following is true:
- 3 or more entities were extracted from the response
- The agent is
detailed_report_agentordeep_research_agent - The conversation has 4 or more turns
Auto-filing is fire-and-forget (non-blocking). Failures are logged but never surfaced to the user.
Auto-Filing Flow¶
flowchart TD
AGENT["<span style='color:#000'>Agent Interaction<br/>(any agent dispatch)</span>"]
CHECK["<span style='color:#000'>_should_auto_file()<br/>entities ≥ 3<br/>agent in AUTO_FILE_AGENTS<br/>turn_count ≥ 4</span>"]
SKIP["<span style='color:#000'>Skip<br/>(interaction too brief)</span>"]
WM["<span style='color:#000'>WikiManager<br/>save_session()</span>"]
TOPIC["<span style='color:#000'>Topic Pages<br/>(upsert per entity)</span>"]
SESSION["<span style='color:#000'>Session Page<br/>(point-in-time capture)</span>"]
INDEX["<span style='color:#000'>wiki_index<br/>(rebuilt per tenant)</span>"]
VESPA["<span style='color:#000'>Vespa wiki_pages schema<br/>hybrid search (semantic + BM25)</span>"]
AGENT --> CHECK
CHECK -->|"no"| SKIP
CHECK -->|"yes"| WM
WM --> TOPIC
WM --> SESSION
WM --> INDEX
TOPIC --> VESPA
SESSION --> VESPA
INDEX --> VESPA
style AGENT fill:#ce93d8,stroke:#7b1fa2,color:#000
style CHECK fill:#ffcc80,stroke:#ef6c00,color:#000
style SKIP fill:#b0bec5,stroke:#546e7a,color:#000
style WM fill:#ce93d8,stroke:#7b1fa2,color:#000
style TOPIC fill:#81c784,stroke:#388e3c,color:#000
style SESSION fill:#81c784,stroke:#388e3c,color:#000
style INDEX fill:#81c784,stroke:#388e3c,color:#000
style VESPA fill:#90caf9,stroke:#1565c0,color:#000 Telegram /wiki Commands¶
Use these commands in Telegram to interact with the wiki directly:
| Command | Description |
|---|---|
/wiki save | Save the current session to the wiki |
/wiki search <query> | Search the wiki knowledge base |
/wiki topic <name> | Look up a topic page by name |
/wiki index | Show the full wiki index |
/wiki lint | Check wiki for orphan, stale, empty, or malformed pages |
REST API¶
| Endpoint | Method | Description |
|---|---|---|
/wiki/save | POST | Persist an agent interaction as a wiki page |
/wiki/search | POST | Full-text search over wiki pages |
/wiki/topic/{slug} | GET | Retrieve a topic page by slug |
/wiki/index | GET | Return the rendered wiki index |
/wiki/lint | GET | Report orphan, stale, empty, and malformed pages |
/wiki/topic/{slug} | DELETE | Delete a topic page by slug |
The lint response includes malformed_pages entries for missing, invalid, or timezone-naive updated_at values. Each entry names the document, field, and stored value, and contributes to issues_found.
# Save a wiki page
curl -X POST http://localhost:28000/wiki/save \
-H "Content-Type: application/json" \
-d '{
"query": "machine learning basics",
"response": {"answer": "ML is..."},
"entities": ["machine_learning"],
"agent_name": "summarizer_agent",
"tenant_id": "acme_corp"
}'
# Search wiki pages
curl -X POST http://localhost:28000/wiki/search \
-H "Content-Type: application/json" \
-d '{"query": "machine learning", "tenant_id": "acme_corp", "top_k": 5}'
# Get a topic page
curl "http://localhost:28000/wiki/topic/machine_learning?tenant_id=acme_corp"
# Get the wiki index
curl "http://localhost:28000/wiki/index?tenant_id=acme_corp"
# Run lint checks
curl "http://localhost:28000/wiki/lint?tenant_id=acme_corp"
# Delete a topic page
curl -X DELETE \
"http://localhost:28000/wiki/topic/machine_learning?tenant_id=acme_corp"
RLM (Recursive Language Model)¶
RLM enables agents to process context that exceeds normal token limits by recursively decomposing inputs using a Python REPL. Available on search, report, code, and research agents.
To activate, set the rlm field on the agent's typed input. Note: the unified runtime's REST shortcut (POST /agents/{name}/process) does not forward an rlm field — activate RLM via the Python SDK, which calls the agent's typed process() entrypoint directly:
import asyncio
from cogniverse_agents.detailed_report_agent import (
DetailedReportAgent,
DetailedReportDeps,
DetailedReportInput,
)
from cogniverse_core.agents.rlm_options import RLMOptions
from cogniverse_foundation.config.utils import create_default_config_manager
async def main():
config_manager = create_default_config_manager()
deps = DetailedReportDeps() # deps are tenant-agnostic; tenant_id is per-request
agent = DetailedReportAgent(deps=deps, config_manager=config_manager)
result = await agent.process(
DetailedReportInput(
query="Analyze these results",
tenant_id="default",
rlm=RLMOptions(enabled=True, max_iterations=5),
)
)
print(result.rlm_synthesis)
asyncio.run(main())
RLM is opt-in and disabled by default. When enabled, telemetry metrics (depth, calls, tokens, latency) are included in the response for A/B testing.
Knowledge Management¶
Cogniverse ships a full Knowledge Management Layer built on top of Mem0+Vespa. Every memory write is governed by a KnowledgeSchema that controls retention, sensitivity, pin authority, provenance requirement, contradiction policy, and default trust.
Schema-driven retention¶
| Retention | Behaviour |
|---|---|
PERMANENT | Never auto-deleted. Default when the kind is unregistered. |
EPHEMERAL_SESSION | Cleared when the session ends (via DELETE /admin/tenants/{tenant_id}/sessions/{session_id}). |
EPHEMERAL_DAYS(N) | Soft-deleted at N days, hard-deleted at 2N days. Restorable inside the soft-delete window. |
SCHEMA_DRIVEN | Custom cleanup_hook on the schema. |
Provenance and citations¶
Every write can carry a Provenance record describing who wrote the memory, how it was derived (direct_ingest, extraction, synthesis, etc.), and which source memories or external URLs it cites. Use the CitationTracingAgent to walk the chain back to primary sources.
Contradiction detection¶
When two memories disagree about the same subject, ContradictionDetector groups them into a ConflictSet. The schema's contradiction_policy resolves conflicts at retrieval time: latest_wins, trust_ranked, or preserve_both (all copies surfaced with metadata["disputed"]=True). Use ContradictionReconciliationAgent to surface and resolve open conflict sets.
Trust ranking¶
Trust is derived from the schema's default_trust and the write's derivation_kind. It ages slowly (≈0.005 pt/day above baseline), and can be boosted by user/admin endorsements. At retrieval, results are ranked by relevance × trust × confidence. Direct human assertions (user_assert) outrank agent inferences by default.
Federation (org trunk + tenant overlays)¶
FederatedQueryAgent reads from both the caller's tenant and the org's shared trunk in one call, with tenant overlay winning on collision. KnowledgeSummarizationAgent can promote a summary into the org trunk so all tenants in the same org see it.
Pinning¶
Memories can be pinned (by users, tenant admins, or org admins, each with quota limits) so they survive lifecycle cleanup and trust decay. Pinned memories are never auto-deleted. Use PinService from code or via the admin API.
Knowledge agents¶
Nine specialized agents operate on the knowledge layer:
| Agent | What it does |
|---|---|
AuditExplanationAgent | Explains why an answer was produced (provenance + trust + contradictions) |
CitationTracingAgent | Walks provenance chains back to primary sources |
ContradictionReconciliationAgent | Surfaces and resolves conflict sets |
FederatedQueryAgent | Queries tenant + org-trunk in one call |
KnowledgeGraphTraversalAgent | Traverses the knowledge graph by entity and relationship |
KnowledgeSummarizationAgent | Summarizes a knowledge slice with citations |
MultiDocumentSynthesisAgent | Synthesizes across multiple source documents |
TemporalReasoningAgent | Answers questions about knowledge change over time |
CrossTenantComparisonAgent | Compares knowledge views across tenants (org-admin scoped) |
For full API details see Core Module — Memory Management and Agents Module — Knowledge Agents.
API Reference¶
REST API Endpoints¶
Search Endpoint¶
Request:
{
"query": "string",
"top_k": 10,
"strategy": "hybrid_float_bm25",
"profile": "video_colpali_smol500_mv_frame",
"tenant_id": "default",
"filters": {}
}
Response:
{
"query": "string",
"profile": "video_colpali_smol500_mv_frame",
"strategy": "hybrid_float_bm25",
"results_count": 10,
"results": [
{
"document_id": "string",
"score": 0.95,
"metadata": {},
"highlights": {}
}
],
"session_id": null
}
Ingestion Endpoint¶
Request:
{
"video_dir": "/path/to/videos",
"profile": "video_colpali_smol500_mv_frame",
"backend": "vespa",
"tenant_id": "default",
"batch_size": 10
}
Response:
Check Status:
Status Response:
{
"job_id": "abc123",
"status": "processing",
"videos_processed": 5,
"videos_total": 10,
"errors": []
}
Health Check¶
Response (the chat LLM's endpoint answered 404: nothing is deployed for the model):
{
"status": "degraded",
"service": "cogniverse-runtime",
"backends": {
"registered": 1,
"backends": ["vespa"]
},
"agents": {
"registered": 3,
"agents": ["search_agent", "gateway_agent", "orchestrator_agent"]
},
"dependencies": {
"llm": {
"status": "not_serving",
"endpoints": [
{
"endpoint": "http://cogniverse-semantic-router-envoy:8801/v1",
"model": "openai/cogniverse-classification",
"route": "pro",
"state": "not_serving",
"upstream_status": 404,
"failure": null,
"reason": "answered HTTP 404: nothing is deployed for this model; calls fail fast until the next recheck",
"observed_at": "2026-10-01T17:10:07.512301+00:00",
"recheck_in_s": 21.4
}
]
}
}
}
status is healthy, degraded (the chat LLM is not_serving or failing; search still serves) or, with HTTP 503, unhealthy (the search backend is unreachable). dependencies.llm.status is not_called until this worker process has made an LM call.
Python SDK Reference¶
SearchAgent¶
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
deps = SearchAgentDeps(profile="video_colpali_smol500_mv_frame")
agent = SearchAgent(deps=deps, config_manager=config_manager, schema_loader=schema_loader)
# search_by_text — tenant_id is per-request
results = agent.search_by_text(
query="machine learning",
tenant_id="your_org:production",
top_k=10,
start_date="2024-01-01", # Optional
end_date="2024-12-31", # Optional
)
OrchestratorAgent¶
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps, OrchestratorInput
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_foundation.config.utils import create_default_config_manager
config_manager = create_default_config_manager()
registry = AgentRegistry(tenant_id="your_org:production", config_manager=config_manager)
orchestrator = OrchestratorAgent(deps=OrchestratorDeps(), registry=registry, config_manager=config_manager)
# Orchestrator plans preprocessing + execution agents and returns OrchestratorOutput
result = await orchestrator._process_impl(
OrchestratorInput(query="machine learning tutorial", tenant_id="your_org:production")
)
# OrchestratorOutput.final_output contains the aggregated execution result.
# Enrichment (enhanced_query, entities, relationships, query_variants) is threaded
# from preprocessing agent outputs onto execution agent inputs via AgentTask fields
# by the module-level _merge_enrichment() helper — not returned as a top-level field.
print(result.execution_summary)
Mem0MemoryManager¶
from cogniverse_core.memory.manager import Mem0MemoryManager
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# Initialize required dependencies first
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Instantiate (per-tenant singleton pattern)
memory = Mem0MemoryManager(tenant_id="your_org:production")
# Initialize with all required parameters
memory.initialize(
backend_host="localhost",
backend_port=8080,
llm_model="openai/google/gemma-4-e4b-it",
embedding_model="lightonai/DenseOn",
llm_base_url="http://localhost:11434",
embedder_base_url="http://localhost:29006",
config_manager=config_manager,
schema_loader=schema_loader,
)
# Add memory (requires tenant_id and agent_name)
memory.add_memory(
content="User prefers Python tutorials",
tenant_id="your_org:production",
agent_name="search_agent"
)
# Search memory
relevant_memories = memory.search_memory(
query="tutorial preferences",
tenant_id="your_org:production",
agent_name="search_agent",
top_k=5
)
# Get all memories
all_memories = memory.get_all_memories(tenant_id="your_org:production", agent_name="search_agent")
Configuration¶
System Configuration¶
Configuration is loaded from configs/config.json. The system auto-discovers this file from: 1. COGNIVERSE_CONFIG environment variable (if set) 2. configs/config.json (from current directory) 3. ../configs/config.json (one level up) 4. ../../configs/config.json (two levels up)
See Profile Configuration below for the actual config.json structure.
Profile Configuration¶
Configure embedding profiles in configs/config.json:
{
"backend": {
"profiles": {
"video_colpali_smol500_mv_frame": {
"type": "video",
"description": "Frame-based ColPali profile",
"schema_name": "video_colpali_smol500_mv_frame",
"embedding_model": "TomoroAI/tomoro-colqwen3-embed-4b",
"pipeline_config": {
"extract_keyframes": true,
"keyframe_strategy": "fps",
"keyframe_fps": 0.5,
"transcribe_audio": true,
"generate_descriptions": true,
"generate_embeddings": true
},
"strategies": {
"segmentation": {
"class": "FrameSegmentationStrategy",
"params": {
"fps": 0.5,
"threshold": 0.999,
"max_frames": 3000
}
},
"embedding": {
"class": "MultiVectorEmbeddingStrategy",
"params": {}
}
},
"embedding_type": "multi_vector"
}
}
}
}
Environment Variables¶
The following environment variables are honored by the system:
# Configuration File Discovery
export COGNIVERSE_CONFIG=/path/to/config.json # Override config file path
# JAX Configuration (required for X-CLIP models)
export JAX_PLATFORM_NAME=cpu # Required on Apple Silicon or systems without GPU
# HuggingFace (for model downloads)
export HF_TOKEN=your_token_here # HuggingFace access token for gated models
Note: Most configuration is done via configs/config.json. Environment variables are minimal - primarily JAX_PLATFORM_NAME for X-CLIP compatibility and COGNIVERSE_CONFIG to override the config file location.
Troubleshooting¶
Common Issues¶
Issue: "ModuleNotFoundError: No module named 'cogniverse_core'"¶
Solution:
Issue: "Vespa connection refused"¶
Solution:
# Check if Vespa is running
cogniverse status
# Check logs
cogniverse logs vespa
# Vespa runs as a k3d-managed StatefulSet pod, not a standalone
# container — restart via kubectl, not `docker restart`:
kubectl rollout restart statefulset/cogniverse-vespa -n cogniverse
Issue: "Phoenix not recording spans"¶
Solution:
# Verify Phoenix endpoint
echo $TELEMETRY_OTLP_ENDPOINT
# Should be: localhost:4317 (gRPC)
# Test connectivity
curl http://localhost:26006/health
Issue: "Out of memory during ingestion"¶
Solution:
# Reduce concurrent processing
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_colpali_smol500_mv_frame \
--max-concurrent 1 # Reduce from default 3
# Or use binary embeddings via ranking strategies
# Binary embeddings are configured in schema_config and used via ranking strategies
Issue: "Slow search performance"¶
Solutions:
- Use binary embeddings instead of float
- Enable caching in config.json
- Use BM25-only for text queries
- Reduce top_k to get fewer results
# Binary embeddings are configured per-profile in schema_config.binary_dim
# Use standard profiles - they support both float and binary embeddings
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_colpali_smol500_mv_frame \
--tenant-id default
Debug Mode¶
Enable debug logging by configuring the logging level in your Python script or using standard Python logging configuration:
# Run ingestion with verbose output
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/videos \
--profile video_colpali_smol500_mv_frame \
--tenant-id default
# Check logs
tail -f outputs/logs/*.log
Note: Logging levels are configured programmatically or via configs/config.json, not via environment variables.
Getting Help¶
- Documentation: Home
- GitHub Issues: Report bugs
- Web Client: http://localhost:28400
- Phoenix UI: http://localhost:26006 for traces and experiments
- API Docs: http://localhost:28000/docs for interactive API documentation
Best Practices¶
For Content Managers¶
- Use Multiple Profiles: Ingest with ColPali (frames), X-CLIP (global), and ColQwen (chunks) for best coverage
- Enable Transcription: Always transcribe audio for text search
- Monitor Quality: Run evaluations weekly to track search quality
- Organize by Tenant: Use separate tenants for different content libraries
For Data Scientists¶
- Track Experiments: Always run evaluations through Phoenix for reproducibility
- Use Synthetic Data: Enable synthetic data generation for routing optimization
- Compare Strategies: Test multiple search strategies on your dataset
- Monitor Drift: Track query distribution changes in Phoenix
For Developers¶
- Use SDK: Prefer Python SDK over direct API calls for better error handling
- Handle Errors: Always catch and handle exceptions
- Implement Caching: Cache frequent queries at application level
- Test Multi-Tenant: Test with multiple tenants to ensure isolation
Performance Tips¶
- Binary Embeddings: Use binary embeddings for 4x faster search with minimal accuracy loss
- Batch Ingestion: Process videos in batches of 10-20 for optimal throughput
- Enable Caching: Enable LRU cache for repeated queries
- Use BM25 First: For pure text queries, use BM25-only strategy
- Prewarm Cache: Warm up caches with common queries after ingestion
Next Steps¶
For Users¶
- Advanced Features: See Advanced Configuration
- Deployment: See Production Deployment
- Monitoring: See Performance Monitoring
Developer Resources¶
- Developer Guide: See DEVELOPER_GUIDE.md
- Architecture: See architecture/overview.md
- Module Docs: See modules/sdk.md for package-specific documentation
For DevOps¶
- Deployment: See operations/deployment.md (use
cogniverse up) - Kubernetes Deployment: See operations/kubernetes-deployment.md
- Multi-Tenant Operations: See operations/multi-tenant-ops.md