Cogniverse Developer Guide¶
Complete guide for developers contributing to Cogniverse - a general-purpose multi-agent AI platform with self-optimization via DSPy optimizers. Content understanding is a primary use case, but the architecture supports any agent type.
Table of Contents¶
- Getting Started
- System Architecture
- Development Environment
- Code Organization
- Development Workflows
- Testing Strategy
- Contributing Code
- Best Practices
- Troubleshooting
Getting Started¶
For New Developers¶
Welcome to Cogniverse! This guide will help you:
- Understand the system architecture
- Set up your development environment
- Navigate the codebase
- Write and test code
- Submit contributions
Prerequisites¶
Before you begin, ensure you have:
- Python 3.12+ installed
- Git for version control
- Docker for running services
- uv package manager:
pip install uv - IDE (VS Code, PyCharm recommended)
- 16GB+ RAM (32GB recommended)
Quick Start¶
# Clone repository
git clone <repository-url>
cd cogniverse
# Install dependencies
uv sync
# Start services
cogniverse up
# Run tests to verify setup
uv run pytest tests/common/ -v
# Success! You're ready to develop
System Architecture¶
12-Package Layered Architecture¶
Cogniverse uses a UV workspace with a layered architecture:
flowchart TB
subgraph APP["<span style='color:#000'><b>APPLICATION LAYER</b></span>"]
runtime["<span style='color:#000'><b>runtime</b><br/>FastAPI Server · Quality Monitor</span>"]
messaging["<span style='color:#000'><b>messaging</b><br/>Telegram Gateway</span>"]
cogcli["<span style='color:#000'><b>cli</b><br/>cogniverse CLI</span>"]
end
subgraph IMPL["<span style='color:#000'><b>IMPLEMENTATION LAYER</b></span>"]
agents["<span style='color:#000'><b>agents</b><br/>Routing & Search</span>"]
vespa["<span style='color:#000'><b>vespa</b><br/>Backend Integration</span>"]
finetuning["<span style='color:#000'><b>finetuning</b><br/>SFT & DPO</span>"]
telemetry_phoenix["<span style='color:#000'><b>telemetry-phoenix</b><br/>Phoenix Provider</span>"]
end
subgraph CORE["<span style='color:#000'><b>CORE LAYER</b></span>"]
core["<span style='color:#000'><b>core</b><br/>Base Classes & Registries</span>"]
evaluation["<span style='color:#000'><b>evaluation</b><br/>Experiments & Metrics</span>"]
synthetic["<span style='color:#000'><b>synthetic</b><br/>Training Data</span>"]
end
subgraph FOUND["<span style='color:#000'><b>FOUNDATION LAYER</b></span>"]
sdk["<span style='color:#000'><b>sdk</b><br/>Interfaces & Types</span>"]
foundation["<span style='color:#000'><b>foundation</b><br/>Config & Telemetry</span>"]
end
APP --> IMPL
IMPL --> CORE
CORE --> FOUND
%% Styling - Foundation Layer (green)
style sdk fill:#a5d6a7,stroke:#388e3c,color:#000
style foundation fill:#a5d6a7,stroke:#388e3c,color:#000
%% Styling - Core Layer (purple)
style core fill:#ce93d8,stroke:#7b1fa2,color:#000
style evaluation fill:#ce93d8,stroke:#7b1fa2,color:#000
style synthetic fill:#ce93d8,stroke:#7b1fa2,color:#000
%% Styling - Implementation Layer (orange)
style agents fill:#ffcc80,stroke:#ef6c00,color:#000
style vespa fill:#ffcc80,stroke:#ef6c00,color:#000
style finetuning fill:#ffcc80,stroke:#ef6c00,color:#000
style telemetry_phoenix fill:#ffcc80,stroke:#ef6c00,color:#000
%% Styling - Application Layer (blue)
style runtime fill:#90caf9,stroke:#1565c0,color:#000
style messaging fill:#90caf9,stroke:#1565c0,color:#000
style cogcli fill:#90caf9,stroke:#1565c0,color:#000 Key Principles¶
- Layered Dependencies: Dependencies flow downward only (no cycles)
- SDK Foundation: Zero internal dependencies, pure interfaces
- Separation of Concerns: Each package has clear responsibilities
- Plugin Architecture: Telemetry providers via entry points
- Multi-Tenancy: Complete tenant isolation at every layer
Package Responsibilities¶
| Package | Layer | Purpose | Key Modules |
|---|---|---|---|
| sdk | Foundation | Backend interfaces, Document model | interfaces/, document.py |
| foundation | Foundation | Config base, telemetry interfaces | config/, telemetry/ |
| core | Core | Base classes, registries, memory | agents/, common/, registries/ |
| evaluation | Core | Experiments, metrics, datasets, quality monitor | core/, metrics/, evaluators/, quality_monitor.py |
| synthetic | Core | Synthetic data generation | service.py, generators/ |
| agents | Implementation | Routing, search, orchestration, strategy learner, wiki knowledge base | routing/, search/, orchestrator/, tools/, optimizer/, wiki/ |
| vespa | Implementation | Vespa backend, schema management | config/, registry/ |
| finetuning | Implementation | LLM fine-tuning (SFT, DPO) | training/, dataset/, registry/ |
| telemetry-phoenix | Implementation | Phoenix telemetry provider (plugin) | provider.py, evaluation/ |
| runtime | Application | FastAPI server, ingestion, optimization CLI, quality monitor CLI | routers/, ingestion/, admin/, optimization_cli.py, quality_monitor_cli.py |
| messaging | Application | Telegram messaging gateway | gateway.py, auth.py, command_router.py |
| cli | Application | cogniverse CLI (up, down, status, code, index, graph, logs, secrets, admin, sandbox) | main.py, cluster.py, deploy.py, code.py, graph.py |
Quality Monitor Architecture¶
The quality monitor is a continuous evaluation loop in its own Deployment (cogniverse-quality-monitor), reaching the runtime over its Service. It applies two independent scoring strategies:
- Golden set evaluation (every 2h by default): runs a curated set of queries against the runtime API and scores results using IR metrics (MRR, NDCG, Precision@5). Baselines are stored as Phoenix datasets.
- Live traffic evaluation (every 4h by default): samples recent spans from Phoenix, uses an LLM judge to score response quality, and detects drift from the stored baseline.
When either strategy finds quality below threshold, the monitor submits an Argo workflow to trigger the optimization CLI (--mode triggered) with the degraded agents and a scored trigger dataset.
flowchart TB
QM["<span style='color:#000'><b>QualityMonitor</b><br/>cogniverse_evaluation.quality_monitor</span>"]
evalGolden["<span style='color:#000'><b>evaluate_golden_set()</b><br/>→ GoldenEvalResult (MRR, NDCG, Precision@5)</span>"]
evalLive["<span style='color:#000'><b>evaluate_live_traffic()</b><br/>→ LiveEvalResult (per-agent scores)</span>"]
updateBaseline["<span style='color:#000'><b>update_baseline()</b><br/>Stores new baseline in Phoenix dataset</span>"]
growGolden["<span style='color:#000'><b>grow_golden_set()</b><br/>Adds high-scoring live queries to golden set</span>"]
runLoop["<span style='color:#000'><b>run()</b><br/>Async loop: golden every goldenIntervalSeconds,<br/>live every liveIntervalSeconds</span>"]
QM --> evalGolden
QM --> evalLive
QM --> updateBaseline
QM --> growGolden
QM --> runLoop
style QM fill:#ce93d8,stroke:#7b1fa2,color:#000
style evalGolden fill:#ba68c8,stroke:#7b1fa2,color:#000
style evalLive fill:#ba68c8,stroke:#7b1fa2,color:#000
style updateBaseline fill:#ba68c8,stroke:#7b1fa2,color:#000
style growGolden fill:#ba68c8,stroke:#7b1fa2,color:#000
style runLoop fill:#ba68c8,stroke:#7b1fa2,color:#000 CLI entry point: python -m cogniverse_runtime.quality_monitor_cli Helm: runtime.qualityMonitor.enabled: true
Strategy Learner Architecture¶
The strategy learner (cogniverse_agents.optimizer.strategy_learner.StrategyLearner) distills execution traces into reusable agent strategies stored in Vespa memory via Mem0. It runs as part of the --mode triggered optimization flow.
Two distillation paths run in sequence:
- Pattern extraction (no LLM): statistical analysis of scored traces grouped by agent. Identifies keyword categories (temporal, object, action, comparison) that correlate with high or low scores. Produces org-level
Strategyobjects. - LLM contrastive distillation (optional, needs
llm_config): pairs a high-scoring and low-scoring trace per agent, feeds them to a DSPyPredictmodule to identify what made the difference.
Strategies are stored in Vespa memory with type=strategy metadata under a reserved agent name _strategy_store. Deduplication uses Jaccard similarity (threshold: 0.9) to avoid storing redundant strategies.
Two-level scoping:
- Org-level strategies: stored under
org_id(prefix oftenant_idbefore:) - User-level strategies: stored under
tenant_id
Runtime retrieval via MemoryAwareMixin.get_strategies(query, top_k) in cogniverse_agents.memory_aware_mixin:
# Inside any agent that extends MemoryAwareMixin:
strategies = self.get_strategies(query="cooking tutorial", top_k=5)
# Returns formatted Markdown string for prompt injection, or None
RLM (Recursive Language Model) Architecture¶
DetailedReportAgent, CodingAgent, DeepResearchAgent, and WikiManager now support RLM (Recursive Language Model) for near-infinite context processing. Pass rlm=RLMOptions(enabled=True) in the agent input to activate. WikiManager uses RLM automatically when topic content exceeds 50,000 characters during merging.
Development Environment¶
Initial Setup¶
1. Install UV Package Manager¶
# macOS/Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
# Verify installation
uv --version
2. Clone and Setup Repository¶
# Clone
git clone <repository-url>
cd cogniverse
# Install workspace (all 12 packages + dependencies)
uv sync
# Activate virtual environment
source .venv/bin/activate # macOS/Linux
.venv\Scripts\activate # Windows
3. Start Infrastructure Services¶
# Start Vespa, Phoenix, Ollama
cogniverse up
# Verify services
curl http://localhost:8080/ApplicationStatus # Vespa
curl http://localhost:6006/health # Phoenix
curl http://localhost:11434/api/tags # Ollama
4. Verify Installation¶
# Run tests
uv run pytest tests/common/ -v
# Verify all packages installed
uv pip list | grep cogniverse
# Expected: 12 packages (sdk, foundation, core, evaluation, cli, etc.)
Development Workflow: Three Loops¶
cogniverse up deploys in dev mode: the k3d cluster mounts your working tree (libs/, scripts/, configs/schemas/, data/) over the images, so pods run the code in your checkout — the image only provides the base environment. That splits day-to-day work into three loops, fastest first:
1. Inner loop — code changes (seconds). Pods read your tree; a Python process picks up edits when it starts. Edit, then restart just the component you changed:
No build, no tag, no helm. Tests run straight off the tree too (uv run pytest ...). This covers most development.
2. Deploy loop — cogniverse up (minutes, occasional). Required only when something outside the mounts changes: dependencies (pyproject.toml / uv.lock bake into image layers), Dockerfiles, chart templates or values, new sidecars — or when you want to refresh the deployed artifact identity. It builds image tags absent from the host, imports tags absent from k3d, and helm-upgrades.
Each image has a git-derived version (0.1.devN+g<sha>) from the latest commit to its Dockerfile, copied inputs after .dockerignore filtering, or the ignore file itself. A tests-only commit keeps every image tag; a source change re-stamps only the images that consume it. The packaged chart uses the runtime version as its scalar package version while Helm receives each image tag independently. A consequence worth knowing: because of the dev mounts, the image tag on a dev cluster records the base image's provenance, not the running code's — the running code is always your tree as of the last pod restart. (In dev mode this extends to Argo-spawned optimization pods too — they inherit the same source mounts.) Run cogniverse up when you need image content and tree to coincide — for workloads that run from the image alone (the inference sidecars, or any non-dev deployment where devMode is off).
3. Release loop — make release VERSION=x.y.z (rare, deliberate). Cuts the v<version> tag; CI builds and publishes the wheels, images, and OCI chart from that tag. See docs/development/publishing-guide.md for the full release process.
Rule of thumb: code → restart; deps/chart/Dockerfile → cogniverse up; shipping → release tag.
IDE Setup¶
VS Code Configuration¶
Create .vscode/settings.json:
{
"python.defaultInterpreterPath": "${workspaceFolder}/.venv/bin/python",
"python.testing.pytestEnabled": true,
"python.testing.pytestArgs": [
"tests"
],
"python.linting.enabled": true,
"python.linting.ruffEnabled": true,
"python.formatting.provider": "black",
"editor.formatOnSave": true,
"files.exclude": {
"**/__pycache__": true,
"**/*.pyc": true,
".venv": true
}
}
Create .vscode/launch.json:
{
"version": "0.2.0",
"configurations": [
{
"name": "Python: Debug Tests",
"type": "python",
"request": "launch",
"module": "pytest",
"args": ["tests/", "-v"],
"console": "integratedTerminal"
},
{
"name": "Python: FastAPI Server",
"type": "python",
"request": "launch",
"module": "uvicorn",
"args": [
"cogniverse_runtime.main:app",
"--reload"
],
"env": {
"JAX_PLATFORM_NAME": "cpu"
}
}
]
}
PyCharm Configuration¶
- Set Python Interpreter: File → Settings → Project → Python Interpreter → Select
.venv/bin/python - Enable Pytest: Settings → Tools → Python Integrated Tools → Testing → pytest
- Configure Ruff: Settings → Tools → External Tools → Add Ruff
- Mark Directories: Right-click
libs/*/cogniverse_*→ Mark Directory as → Sources Root
Code Organization¶
Workspace Structure¶
flowchart TB
subgraph ROOT["<span style='color:#000'><b>Cogniverse Workspace</b></span>"]
direction TB
subgraph LIBS["<span style='color:#000'><b>libs/</b><br/>All 12 workspace packages</span>"]
sdk["<span style='color:#000'><b>sdk/</b><br/>Foundation: Pure interfaces</span>"]
foundation["<span style='color:#000'><b>foundation/</b><br/>Foundation: Config & telemetry</span>"]
core["<span style='color:#000'><b>core/</b><br/>Core: Base classes & registries</span>"]
evaluation["<span style='color:#000'><b>evaluation/</b><br/>Core: Experiments & metrics</span>"]
synthetic["<span style='color:#000'><b>synthetic/</b><br/>Core: Synthetic data</span>"]
agents["<span style='color:#000'><b>agents/</b><br/>Implementation: Agents</span>"]
vespa["<span style='color:#000'><b>vespa/</b><br/>Implementation: Vespa backend</span>"]
finetuning["<span style='color:#000'><b>finetuning/</b><br/>Implementation: Fine-tuning</span>"]
telemetry_phoenix["<span style='color:#000'><b>telemetry-phoenix/</b><br/>Implementation: Phoenix plugin</span>"]
runtime["<span style='color:#000'><b>runtime/</b><br/>Application: FastAPI server</span>"]
messaging2["<span style='color:#000'><b>messaging/</b><br/>Application: Telegram gateway</span>"]
cogcli2["<span style='color:#000'><b>cli/</b><br/>Application: cogniverse CLI</span>"]
end
subgraph TESTS["<span style='color:#000'><b>tests/</b><br/>Test suite</span>"]
test_common["<span style='color:#000'>common/</span>"]
test_agents["<span style='color:#000'>agents/</span>"]
test_routing["<span style='color:#000'>routing/</span>"]
test_memory["<span style='color:#000'>memory/</span>"]
test_ingestion["<span style='color:#000'>ingestion/</span>"]
test_evaluation["<span style='color:#000'>evaluation/</span>"]
test_telemetry["<span style='color:#000'>telemetry/</span>"]
test_backends["<span style='color:#000'>backends/</span>"]
test_finetuning["<span style='color:#000'>finetuning/</span>"]
test_synthetic["<span style='color:#000'>synthetic/</span>"]
test_admin["<span style='color:#000'>admin/</span>"]
test_events["<span style='color:#000'>events/</span>"]
test_system["<span style='color:#000'>system/</span>"]
test_cli["<span style='color:#000'>cli/</span>"]
test_core["<span style='color:#000'>core/</span>"]
test_e2e["<span style='color:#000'>e2e/</span>"]
test_foundation["<span style='color:#000'>foundation/</span>"]
test_messaging["<span style='color:#000'>messaging/</span>"]
test_runtime["<span style='color:#000'>runtime/</span>"]
test_utils["<span style='color:#000'>utils/</span>"]
test_charts["<span style='color:#000'>charts/</span>"]
end
subgraph SCRIPTS["<span style='color:#000'><b>scripts/</b><br/>Operational scripts</span>"]
script_ingestion["<span style='color:#000'>run_ingestion.py</span>"]
script_schema["<span style='color:#000'>deploy_json_schema.py</span>"]
script_experiments["<span style='color:#000'>run_experiments_with_visualization.py</span>"]
script_phoenix["<span style='color:#000'>start_phoenix.py</span>"]
script_other["<span style='color:#000'>...and more</span>"]
end
subgraph CONFIGS["<span style='color:#000'><b>configs/</b><br/>Configuration files</span>"]
config_json["<span style='color:#000'>config.json</span>"]
config_schemas["<span style='color:#000'>schemas/</span>"]
config_profiles["<span style='color:#000'>profiles/</span>"]
config_examples["<span style='color:#000'>examples/</span>"]
config_policies["<span style='color:#000'>agent_policies/</span>"]
end
subgraph DOCS["<span style='color:#000'><b>docs/</b><br/>Documentation</span>"]
docs_arch["<span style='color:#000'>architecture/</span>"]
docs_dev["<span style='color:#000'>development/</span>"]
docs_diagrams["<span style='color:#000'>diagrams/</span>"]
docs_modules["<span style='color:#000'>modules/</span>"]
docs_ops["<span style='color:#000'>operations/</span>"]
docs_testing["<span style='color:#000'>testing/</span>"]
docs_other["<span style='color:#000'>...and more</span>"]
end
root_pyproject["<span style='color:#000'>pyproject.toml</span>"]
root_uvlock["<span style='color:#000'>uv.lock</span>"]
end
%% Styling - Foundation packages (green)
style sdk fill:#a5d6a7,stroke:#388e3c,color:#000
style foundation fill:#a5d6a7,stroke:#388e3c,color:#000
%% Styling - Core packages (purple)
style core fill:#ce93d8,stroke:#7b1fa2,color:#000
style evaluation fill:#ce93d8,stroke:#7b1fa2,color:#000
style synthetic fill:#ce93d8,stroke:#7b1fa2,color:#000
%% Styling - Implementation packages (orange)
style agents fill:#ffcc80,stroke:#ef6c00,color:#000
style vespa fill:#ffcc80,stroke:#ef6c00,color:#000
style finetuning fill:#ffcc80,stroke:#ef6c00,color:#000
style telemetry_phoenix fill:#ffcc80,stroke:#ef6c00,color:#000
%% Styling - Application packages (blue)
style runtime fill:#90caf9,stroke:#1565c0,color:#000
style messaging2 fill:#90caf9,stroke:#1565c0,color:#000
style cogcli2 fill:#90caf9,stroke:#1565c0,color:#000
%% Styling - Supporting directories (grey)
style TESTS fill:#b0bec5,stroke:#546e7a,color:#000
style SCRIPTS fill:#b0bec5,stroke:#546e7a,color:#000
style CONFIGS fill:#b0bec5,stroke:#546e7a,color:#000
style DOCS fill:#b0bec5,stroke:#546e7a,color:#000 Package Structure Pattern¶
Each package follows this structure:
flowchart TB
subgraph PKG["<span style='color:#000'><b>libs/my_package/</b></span>"]
direction TB
pyproject["<span style='color:#000'><b>pyproject.toml</b><br/>Package configuration</span>"]
readme["<span style='color:#000'><b>README.md</b><br/>Package documentation</span>"]
subgraph SRC["<span style='color:#000'><b>cogniverse_my_package/</b><br/>Python package</span>"]
init["<span style='color:#000'>__init__.py</span>"]
subgraph MOD1["<span style='color:#000'><b>module1/</b><br/>Feature module</span>"]
mod1_init["<span style='color:#000'>__init__.py</span>"]
mod1_impl["<span style='color:#000'>implementation.py</span>"]
end
subgraph MOD2["<span style='color:#000'><b>module2/</b><br/>Feature module</span>"]
mod2_init["<span style='color:#000'>__init__.py</span>"]
mod2_impl["<span style='color:#000'>implementation.py</span>"]
end
subgraph UTILS["<span style='color:#000'><b>utils/</b><br/>Utilities</span>"]
helpers["<span style='color:#000'>helpers.py</span>"]
end
end
tests["<span style='color:#000'><b>tests/</b><br/>Package-specific tests (optional)</span>"]
end
%% Styling
style pyproject fill:#b0bec5,stroke:#546e7a,color:#000
style readme fill:#b0bec5,stroke:#546e7a,color:#000
style SRC fill:#90caf9,stroke:#1565c0,color:#000
style tests fill:#b0bec5,stroke:#546e7a,color:#000 Naming Conventions¶
Packages:
- Installable name:
cogniverse-my-package(hyphens) - Import name:
cogniverse_my_package(underscores)
Modules:
- Lowercase with underscores:
search_agent.py - Avoid abbreviations:
config.pynotcfg.py
Classes:
- PascalCase:
SearchAgent,VespaSchemaManager - Descriptive names:
OrchestratorAgentnotOA
Functions:
- snake_case:
get_tenant_id(),deploy_schema() - Verb + noun:
create_agent(),load_config()
Constants:
- UPPER_SNAKE_CASE:
MAX_BATCH_SIZE,DEFAULT_TIMEOUT
Development Workflows¶
Working on a Single Package¶
# Navigate to package
cd libs/agents
# Install package in editable mode
uv pip install -e .
# Make changes
vim cogniverse_agents/orchestrator_agent.py
# Run package tests
cd ../..
uv run pytest tests/agents/ -v
# Run linting
uv run ruff check libs/agents/
# Format code
uv run ruff format libs/agents/
Working Across Multiple Packages¶
# Make changes in core
vim libs/core/cogniverse_core/agents/base.py
# Make changes in agents (uses core)
vim libs/agents/cogniverse_agents/orchestrator_agent.py
# Changes in core immediately visible to agents (workspace)
uv run pytest tests/agents/ -v
Adding a New Feature¶
Example: Add a new ingestion processor
-
Create feature branch:
-
Implement in appropriate package:
# libs/runtime/cogniverse_runtime/ingestion/processors/custom_processor.py from pathlib import Path from cogniverse_runtime.ingestion.processor_base import BaseProcessor class CustomVideoProcessor(BaseProcessor): """Custom processor for specialized video handling""" PROCESSOR_NAME = "custom_video" def __init__(self, logger, **kwargs): super().__init__(logger, **kwargs) def process(self, video_path: Path, context: dict) -> dict: # Implement custom processing logic frames = self._extract_frames(video_path) embeddings = self._generate_embeddings(frames) return {"embeddings": embeddings, "metadata": context} -
Add tests:
# tests/ingestion/test_custom_processor.py import logging import pytest from pathlib import Path from cogniverse_runtime.ingestion.processors.custom_processor import CustomVideoProcessor def test_custom_processor(): logger = logging.getLogger(__name__) processor = CustomVideoProcessor(logger=logger) result = processor.process(Path("test.mp4"), {}) assert "embeddings" in result -
Run tests:
-
Update documentation:
-
Commit and push:
Adding a New Package Dependency¶
To a specific package:
cd libs/agents
uv add scikit-learn>=1.3.0 # Adds to agents/pyproject.toml
cd ../..
uv sync # Update workspace
To workspace root (shared dependency):
# Navigate to workspace root (if not already there)
cd cogniverse # Or your workspace root directory
uv add numpy>=1.24.0 # Adds to root pyproject.toml
uv sync
Running Scripts¶
# Ingestion
uv run python scripts/run_ingestion.py \
--tenant-id acme:acme \
--video_dir data/videos \
--backend vespa \
--profile video_colpali_smol500_mv_frame
# Experiments
uv run python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--dataset-path data/testset/evaluation/video_search_queries.csv \
--dataset-name golden_eval_v1 \
--profiles video_colpali_smol500_mv_frame
# Deploy schema
uv run python scripts/deploy_json_schema.py \
configs/schemas/video_colpali_smol500_mv_frame_schema.json
# Optimization CLI — per-agent batch modes
python -m cogniverse_runtime.optimization_cli --mode simba --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode entity-extraction --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode cleanup --log-retention-days 7
# Optimization CLI — triggered mode (called by Argo when quality degrades)
# Loads scored examples from a Phoenix trigger dataset, compiles DSPy modules
# for each flagged agent, then runs strategy distillation.
python -m cogniverse_runtime.optimization_cli \
--mode triggered \
--tenant-id default \
--agents search,summary \
--trigger-dataset optimization-trigger-default-20260403_040000
# Optimization CLI — telemetry-first batch modes (A2A architecture)
# Each mode reads spans from Phoenix, optimizes the relevant agent, and saves artifacts.
python -m cogniverse_runtime.optimization_cli --mode simba --tenant-id default --lookback-hours 24
python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id default --lookback-hours 24
python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id default --lookback-hours 24
python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id default --lookback-hours 24
# Optimization CLI — remaining modes
python -m cogniverse_runtime.optimization_cli --mode online-routing-eval --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode synthetic --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode rollback --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode ab-compare --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode egress-netpol --tenant-id default
python -m cogniverse_runtime.optimization_cli --mode monthly-reports --tenant-id default
Full --mode set: cleanup, triggered, simba, workflow, gateway-thresholds, online-routing-eval, profile, entity-extraction, synthetic, rollback, ab-compare, egress-netpol, monthly-reports.
Testing Strategy¶
Test Organization¶
Tests are organized by package/module:
flowchart TB
subgraph TESTS["<span style='color:#000'><b>tests/</b><br/>Test suite organized by package and feature</span>"]
direction TB
subgraph PKG_TESTS["<span style='color:#000'><b>Package Tests</b></span>"]
test_common["<span style='color:#000'><b>common/</b><br/>cogniverse_core tests</span>"]
test_agents["<span style='color:#000'><b>agents/</b><br/>cogniverse_agents tests</span>"]
test_backends["<span style='color:#000'><b>backends/</b><br/>Backend integration tests</span>"]
test_evaluation["<span style='color:#000'><b>evaluation/</b><br/>Evaluation framework tests</span>"]
test_finetuning["<span style='color:#000'><b>finetuning/</b><br/>Fine-tuning tests</span>"]
test_synthetic["<span style='color:#000'><b>synthetic/</b><br/>Synthetic data tests</span>"]
test_telemetry["<span style='color:#000'><b>telemetry/</b><br/>Telemetry provider tests</span>"]
test_core["<span style='color:#000'><b>core/</b><br/>cogniverse_core unit tests</span>"]
test_foundation["<span style='color:#000'><b>foundation/</b><br/>cogniverse_foundation tests</span>"]
end
subgraph FEATURE_TESTS["<span style='color:#000'><b>Feature Tests</b></span>"]
test_routing["<span style='color:#000'><b>routing/</b><br/>Routing integration tests</span>"]
test_memory["<span style='color:#000'><b>memory/</b><br/>Memory system tests</span>"]
test_ingestion["<span style='color:#000'><b>ingestion/</b><br/>Ingestion pipeline tests</span>"]
test_events["<span style='color:#000'><b>events/</b><br/>Event system tests</span>"]
test_admin["<span style='color:#000'><b>admin/</b><br/>Admin API tests</span>"]
test_messaging["<span style='color:#000'><b>messaging/</b><br/>Messaging gateway tests</span>"]
test_runtime["<span style='color:#000'><b>runtime/</b><br/>Runtime integration tests</span>"]
test_cli["<span style='color:#000'><b>cli/</b><br/>CLI tests</span>"]
end
subgraph SUPPORT_TESTS["<span style='color:#000'><b>Support</b></span>"]
test_e2e["<span style='color:#000'><b>e2e/</b><br/>End-to-end tests</span>"]
test_system["<span style='color:#000'><b>system/</b><br/>System-level tests</span>"]
test_utils["<span style='color:#000'><b>utils/</b><br/>Test utilities & helpers</span>"]
test_charts["<span style='color:#000'><b>charts/</b><br/>Helm chart validation tests</span>"]
end
end
%% Styling
style PKG_TESTS fill:#ce93d8,stroke:#7b1fa2,color:#000
style FEATURE_TESTS fill:#90caf9,stroke:#1565c0,color:#000
style SUPPORT_TESTS fill:#b0bec5,stroke:#546e7a,color:#000 Running Tests¶
Full test suite:
Package-specific tests:
# Core package
uv run pytest tests/common/ -v
# Agents package
uv run pytest tests/agents/ -v
# Integration tests
uv run pytest tests/routing/integration/ -v
Single test:
With coverage:
Writing Tests¶
Unit test example:
# tests/agents/unit/test_orchestrator_agent.py
import pytest
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps
from cogniverse_core.registries.agent_registry import AgentRegistry
@pytest.fixture
def orchestrator(config_manager):
registry = AgentRegistry(tenant_id="test:unit", config_manager=config_manager)
return OrchestratorAgent(
deps=OrchestratorDeps(), registry=registry, config_manager=config_manager
)
@pytest.mark.unit
class TestOrchestratorAgent:
@pytest.mark.ci_fast
def test_orchestrator_initialization(self, orchestrator):
"""Test OrchestratorAgent initializes correctly"""
assert orchestrator.deps is not None
assert hasattr(orchestrator, "logger")
Integration test example:
# tests/routing/integration/test_orchestration_pipeline.py
import pytest
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps, OrchestratorInput
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_vespa.search_backend import VespaSearchBackend
@pytest.mark.integration
@pytest.mark.asyncio
async def test_full_orchestration_pipeline(config_manager):
"""Test orchestrator with real Vespa backend"""
registry = AgentRegistry(tenant_id="test:integration", config_manager=config_manager)
orchestrator = OrchestratorAgent(
deps=OrchestratorDeps(), registry=registry, config_manager=config_manager
)
result = await orchestrator._process_impl(
OrchestratorInput(
query="machine learning tutorial",
tenant_id="test:integration",
)
)
assert result is not None
Test Fixtures¶
Workspace-level fixtures (tests/conftest.py):
import pytest
@pytest.fixture
def config_manager(backend_config_env):
"""
Create ConfigManager with backend store for testing.
Requires backend_config_env fixture to set environment variables.
"""
from cogniverse_foundation.config.utils import create_default_config_manager
return create_default_config_manager()
@pytest.fixture
def config_manager_memory():
"""
Create ConfigManager with in-memory store for unit testing.
Does not require any backend infrastructure (Vespa, etc.).
"""
from cogniverse_foundation.config.manager import ConfigManager
from tests.utils.memory_store import InMemoryConfigStore
store = InMemoryConfigStore()
store.initialize()
return ConfigManager(store=store)
@pytest.fixture
def telemetry_manager_without_phoenix():
"""
Standard telemetry manager fixture for tests that don't need real Phoenix.
Sets up telemetry with mock endpoints.
"""
import cogniverse_foundation.telemetry.manager as telemetry_manager_module
from cogniverse_foundation.telemetry.config import BatchExportConfig, TelemetryConfig
from cogniverse_foundation.telemetry.manager import TelemetryManager
from cogniverse_foundation.telemetry.registry import get_telemetry_registry
# Reset TelemetryManager singleton AND clear provider cache
TelemetryManager.reset()
get_telemetry_registry().clear_cache()
config = TelemetryConfig(
otlp_endpoint="http://localhost:24317",
provider_config={
"http_endpoint": "http://localhost:26006",
"grpc_endpoint": "http://localhost:24317",
},
batch_config=BatchExportConfig(use_sync_export=True),
)
# Set as the global singleton
manager = TelemetryManager(config=config)
telemetry_manager_module._telemetry_manager = manager
yield manager
# Cleanup
TelemetryManager.reset()
get_telemetry_registry().clear_cache()
Contributing Code¶
Code Review Process¶
- Create feature branch:
git checkout -b feature/my-feature - Implement feature: Write code + tests
- Run tests:
uv run pytest -v - Run linting:
uv run ruff check . && uv run ruff format . - Commit changes:
git commit -m "Add feature X" - Push branch:
git push origin feature/my-feature - Create PR: Use GitHub PR template
- Address feedback: Make requested changes
- Merge: Once approved, squash and merge
Pull Request Template¶
## Description
Brief description of changes
## Type of Change
- [ ] Bug fix
- [ ] New feature
- [ ] Breaking change
- [ ] Documentation update
## Testing
- [ ] Unit tests added/updated
- [ ] Integration tests added/updated
- [ ] All tests passing
## Documentation
- [ ] README updated
- [ ] Module docs updated
- [ ] API docs updated
## Checklist
- [ ] Code follows style guidelines
- [ ] Self-review completed
- [ ] Comments added for complex code
- [ ] No warnings from linters
Commit Message Guidelines¶
Format:
Rules:
- Start with a verb:
Add,Fix,Update,Refactor,Remove - Keep subject line under 72 characters
- Add body explaining WHY for non-trivial changes
- No meta-commentary (test counts, phase numbers, "all tests pass")
Examples:
Add BM25 rerank search strategy
Enable hybrid search combining semantic and lexical matching
for improved recall on keyword-heavy queries.
Refactor tenant context logic to middleware
Centralizes tenant isolation checks to reduce code duplication
across 12 endpoint handlers.
Pre-Commit Checklist:
- Run
uv run pytest- 100% pass rate required - Run
uv run ruff check- no lint errors - Fix implementation to satisfy tests, never weaken tests
- Update documentation for significant changes
Best Practices¶
Code Quality¶
-
Type Hints: Always use type hints
-
Docstrings: Document all public functions
-
Error Handling: Use specific exceptions
-
Logging: Use structured logging
Performance¶
-
Async/Await: Use async for I/O operations
-
Batch Processing: Process in batches
-
Caching: Cache expensive operations
Security¶
-
Tenant Isolation: Always use tenant_id
-
Input Validation: Validate all inputs
Troubleshooting¶
Common Issues¶
Issue: Import errors
Issue: Tests failing
# Check services are running
docker ps | grep -E "vespa|phoenix|ollama"
# Restart a single component (dev mounts serve your working tree)
kubectl rollout restart deployment/cogniverse-runtime -n cogniverse
# Full rebuild + redeploy (recreates anything missing)
cogniverse up
Issue: Out of memory
# Reduce ingestion concurrency
uv run python scripts/run_ingestion.py --tenant-id acme:acme --video_dir data/videos --max-concurrent 1
# Force JAX (X-CLIP) onto CPU instead of GPU
export JAX_PLATFORM_NAME=cpu
Debug Mode¶
Enable debug logging:
export LOG_LEVEL=DEBUG
export TELEMETRY_HTTP_ENDPOINT=http://localhost:6006
export TELEMETRY_OTLP_ENDPOINT=localhost:4317
uv run pytest tests/agents/ -v -s # -s shows print statements
Next Steps¶
Recommended Reading¶
- Read Architecture Overview
- Read SDK Architecture
- Explore Module Documentation
For Contributors¶
- Read Testing Guide
- Review module documentation in Module Documentation
For Advanced Features¶
- Read Multi-Tenant Architecture
- Read System Flows
- Read Performance Monitoring
- Read Multi-Agent Interactions
Last Updated: 2026-07-05