Skip to content

Comprehensive System Testing Plan

Purpose: Bottom-up manual testing and learning guide for Cogniverse Approach: Start with foundational subsystems, validate each layer, build upward Goal: Thorough understanding and validation of the complete system


Testing Philosophy

Bottom-Up Approach:

  1. Validate foundation before building on it

  2. Each layer depends on layers below it

  3. Fix issues at lowest level first

  4. Document learnings and discoveries

For Each Subsystem:

  • ✅ Purpose: Understand what it does and why

  • ✅ Basic Tests: Core functionality works

  • ✅ Advanced Tests: Edge cases and features

  • ✅ Integration Tests: Works with dependent layers

  • ✅ Learnings: Document key insights


Layer 1: Storage Foundation (Vespa)

Purpose: Vector database for embeddings and metadata storage

1.1 Vespa Service Health

Test Vespa is running:

# Check Vespa service status
curl http://localhost:8080/state/v1/health

# Expected: {"status": {"code": "up"}}

Verify Vespa configuration endpoint:

curl http://localhost:8080/ApplicationStatus

# Should show application status

Learning Points:

  • Vespa is the foundation - everything else needs it

  • Default port: 8080

  • State API provides health checks

1.2 Schema Validation

Check deployed schemas:

# List all schemas
curl "http://localhost:8080/search/?yql=select+*+from+sources+*+where+true+limit+0"

# Check specific tenant schema exists
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_default+where+true+limit+1"

Test schema structure:

# Get document count per schema
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_default+where+true+limit+0" | jq '.root.fields.totalCount'

Learning Points:

  • Schemas are tenant-isolated: {profile}_{tenant_id}

  • Multiple profiles can exist per tenant

  • Default tenant: "default"

1.3 Document Ingestion Test

Ingest a test video:

# Create test video directory
mkdir -p /tmp/test_video

# Download or copy a sample video (MP4)
# For testing: use any short video file

# Run ingestion for single profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
    --video_dir /tmp/test_video \
    --backend vespa \
    --profile video_colpali_smol500_mv_frame \
    --tenant-id test_basic

# Monitor logs
tail -f outputs/logs/ingestion_*.log

Verify ingestion:

# Check document count increased
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+true+limit+10" | jq '.root.fields.totalCount'

# Inspect a document
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+true+limit+1" | jq '.root.children[0].fields'

Learning Points:

  • Ingestion creates embeddings and metadata

  • Documents have video_id, frame_number, embeddings, metadata

  • Each profile has different embedding dimensions

1.4 Basic Search Test

Test vector search:

# Simple text query
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+userQuery()&query=test+video&hits=5"

# Check results structure
curl "http://localhost:8080/search/?yql=select+video_id,video_title,frame_number+from+video_colpali_smol500_mv_frame_test_basic+where+userQuery()&query=test&hits=3" | jq '.root.children[] | {video_id: .fields.video_id, title: .fields.video_title}'

Learning Points:

  • Vespa provides vector search via embeddings

  • Results ranked by relevance

  • Can filter by fields (video_id, timestamp, etc.)

✅ Layer 1 Complete: Vespa storage working, schemas deployed, documents ingested and searchable


Layer 2: Telemetry Foundation (Phoenix)

Purpose: Observability and tracing for all operations

2.1 Phoenix Service Health

Check Phoenix server:

# Verify Phoenix is running
curl http://localhost:6006/healthz

# Expected: 200 OK

Access Phoenix UI:

# Open Phoenix dashboard
open http://localhost:6006

# Verify UI loads

Learning Points:

  • Phoenix: observability platform for LLM apps

  • Port 6006

  • UI provides span visualization

2.2 Project Structure

List Phoenix projects:

# Run in Python REPL
from phoenix.client import Client

client = Client(base_url='http://localhost:6006')

# List projects (returns a list of dict-like phoenix.client v1.Project)
projects = client.projects.list()
for project in projects:
    print(f"Project: {project['name']}")

Expected projects:

  • cogniverse-test:unit (search, routing and orchestration spans of tenant test:unit, the tenant the run below searches as)

  • cogniverse-test:unit-{service} for a management service's spans, for example cogniverse-test:unit-synthetic_data

Learning Points:

  • Projects isolate telemetry by tenant

  • Naming: cogniverse-{tenant_id} for user operations, cogniverse-{tenant_id}-{service} for management operations

2.3 Span Collection Test

Generate telemetry by running a search:

# Run comprehensive test (generates spans)
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
    --profiles video_colpali_smol500_mv_frame \
    --test-multiple-strategies \
    --max-queries 3

# Wait for execution to complete

Query spans from Phoenix:

from phoenix.client import Client
from datetime import datetime, timedelta, timezone

client = Client(base_url='http://localhost:6006')

# Get search spans. `timeout=` is required explicitly — the method's own
# default is 5s (independent of any client-level timeout) and a loaded
# project can easily exceed that, silently reading as "no spans".
spans_df = client.spans.get_spans_dataframe(
    project_name="cogniverse-test:unit",
    start_time=datetime.now(timezone.utc) - timedelta(hours=1),
    timeout=60,
)
search_spans = spans_df[spans_df["name"] == "search_service.search"]

print(f"Search spans collected: {len(search_spans)}")
print(f"Columns: {search_spans.columns.tolist()}")

# Inspect a span
if len(search_spans) > 0:
    print("\nSample span:")
    print(search_spans.iloc[0][['name', 'attributes.query', 'attributes.latency_ms']])

Learning Points:

  • Spans capture operation traces

  • Attributes store metadata (query, profile, strategy, latency, result rows)

  • Can query by time range and project

2.4 Span Attributes Validation

Check span structure:

# Get span with all attributes
span = search_spans.iloc[0]

# Key attributes for search spans
search_attributes = [
    'attributes.query',
    'attributes.profile',
    'attributes.strategy',
    'attributes.top_k',
    'attributes.backend',
    'attributes.result_granularity',
    'attributes.latency_ms',
    'attributes.num_results',
    'attributes.top_score',
    'attributes.output.value',   # JSON list of the result rows
    'attributes.tenant',         # {'id': <tenant_id>}
]

for attr in search_attributes:
    if attr in span.index:
        print(f"{attr}: {span[attr]}")

Learning Points:

  • Search spans include query, profile, strategy, result count and the result rows (output.value)

  • Can filter spans by attributes

  • Used for optimization and analytics

✅ Layer 2 Complete: Phoenix collecting telemetry, spans queryable, attributes validated


Layer 3: Memory Foundation (Mem0)

Purpose: Conversation memory and context for agents

3.1 Memory Service Health

Check Mem0 initialization:

from cogniverse_core.memory.manager import Mem0MemoryManager

# Initialize manager (requires tenant_id)
manager = Mem0MemoryManager(tenant_id="your_org:production")
manager.initialize(
    backend_host="localhost",
    backend_port=8080,
    llm_model="openai/google/gemma-4-e4b-it",
    embedding_model="lightonai/DenseOn",
    llm_base_url="http://localhost:11434",
    embedder_base_url="http://localhost:8002",  # DenseOn sidecar endpoint
    config_manager=config_manager,
    schema_loader=schema_loader,
)

print(f"Memory manager initialized: {manager.memory is not None}")

Learning Points:

  • Mem0 stores agent memories and user preferences

  • Initialized on first use

  • Isolated by tenant_id + agent_name

3.2 Memory Storage Test

Add a memory:

# Add memory for routing agent
manager.add_memory(
    content="User prefers video results for cooking queries",
    tenant_id="your_org:production",
    agent_name="orchestrator_agent",
    metadata={"context": "preference_learning"}
)

print("Memory added successfully")

Retrieve memory:

# Search for memory
results = manager.search_memory(
    query="cooking preferences",
    tenant_id="your_org:production",
    agent_name="orchestrator_agent",
    top_k=5
)

for i, result in enumerate(results, 1):
    print(f"\nMemory {i}:")
    print(f"  Content: {result['memory']}")
    print(f"  Score: {result.get('score', 0.0):.3f}")
    print(f"  Metadata: {result.get('metadata', {})}")

Note: search_memory returns List[Dict[str, Any]] (plain dicts, not objects) — always use result[...]/result.get(...), never attribute access.

Learning Points:

  • Memories are searchable (semantic search)

  • Score indicates relevance

  • Metadata provides context

3.3 Memory Lifecycle

Get all memories:

# Get all memories for agent
all_memories = manager.get_all_memories(
    tenant_id="your_org:production",
    agent_name="orchestrator_agent"
)

print(f"Total memories: {len(all_memories)}")

Delete a memory:

# Get memory ID
memory_id = all_memories[0]['id'] if all_memories else None

if memory_id:
    manager.delete_memory(
        memory_id=memory_id,
        tenant_id="your_org:production",
        agent_name="orchestrator_agent"
    )
    print(f"Deleted memory: {memory_id}")

Learning Points:

  • Memories persist across sessions

  • Can be deleted individually or bulk cleared

  • Memory manager handles CRUD operations

✅ Layer 3 Complete: Mem0 storing and retrieving memories, lifecycle operations working


Layer 4: Core Library (cogniverse_core)

Purpose: Configuration, telemetry, and memory management abstractions

4.1 Configuration Management

Test SystemConfig:

from cogniverse_foundation.config.unified_config import SystemConfig

# SystemConfig is GLOBAL — one per deployment, not per-tenant.
# It has no tenant_id field.
system_config = SystemConfig(
    llm_model="gpt-4",
    backend_url="http://localhost",
    backend_port=8080
)

print(f"LLM Model: {system_config.llm_model}")
print(f"Search Backend: {system_config.search_backend}")
print(f"Backend URL: {system_config.backend_url}")
print(f"Backend Port: {system_config.backend_port}")

Test tenant-specific override (via BackendConfig):

from cogniverse_foundation.config.unified_config import BackendConfig
from cogniverse_foundation.config.utils import create_default_config_manager

# Tenant-specific overrides use BackendConfig (not SystemConfig, which
# is global) and are persisted via ConfigManager
config_manager = create_default_config_manager()
tenant_backend_config = BackendConfig(
    tenant_id="test_tenant",
    url="http://localhost",
    port=8080,
)
config_manager.set_backend_config(tenant_backend_config)

print(f"Tenant backend URL: {tenant_backend_config.url}")

Learning Points:

  • SystemConfig defined in foundation layer; it is GLOBAL (one per deployment), not per-tenant

  • Tenant-specific configurations (backend URL, profiles) use BackendConfig via ConfigManager

  • Default configs in libs/foundation/

4.2 Telemetry Manager

Test TelemetryProvider:

from cogniverse_foundation.telemetry.config import TelemetryConfig
from cogniverse_telemetry_phoenix.provider import PhoenixProvider

# TelemetryConfig is generic (provider-agnostic)
config = TelemetryConfig()

# PhoenixProvider takes no constructor args — it's configured via initialize()
telemetry = PhoenixProvider()
telemetry.initialize({
    "tenant_id": "your_org:production",
    "http_endpoint": "http://localhost:6006",
    "grpc_endpoint": "localhost:4317",
})

# Check Phoenix connection
print(f"Phoenix endpoint: http://localhost:6006")
print(f"Telemetry enabled: {config.enabled}")

Learning Points:

  • PhoenixProvider (telemetry-phoenix package) implements TelemetryProvider interface (foundation)

  • Auto-creates projects per tenant

  • Provides span export utilities via foundation layer interfaces

4.3 Tenant-Aware Components

Test TenantAwareAgentMixin:

from cogniverse_core.agents.tenant_aware_mixin import TenantAwareAgentMixin

class TestComponent(TenantAwareAgentMixin):
    def __init__(self, tenant_id: str):
        super().__init__(tenant_id=tenant_id)

    def get_info(self):
        return f"Tenant: {self.tenant_id}"

# Create instance
component = TestComponent(tenant_id="acme:production")
print(component.get_info())

Learning Points:

  • TenantAwareAgentMixin provides tenant isolation

  • All agents/components inherit this

✅ Layer 4 Complete: Core abstractions working, config/telemetry/memory managers functional


Layer 5: Backend Layer (cogniverse_vespa)

Purpose: Search backend implementation using sdk interfaces

5.1 Backend Implementation

Test backend usage:

from cogniverse_vespa.search_backend import VespaSearchBackend
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# VespaSearchBackend resolves schema/profile at query time
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
backend = VespaSearchBackend(
    config={
        "url": "http://localhost",
        "port": 8080,  # Use your Vespa instance's actual port
        "profiles": {"video_colpali_smol500_mv_frame": {}},
        "default_profiles": {"video": "video_colpali_smol500_mv_frame"},
    },
    config_manager=config_manager,
    schema_loader=schema_loader,
)

print(f"Backend type: {type(backend).__name__}")

Learning Points:

  • Backend implementations (vespa package) use interfaces from sdk layer

  • Provides tenant-aware search operations

  • Handles schema management and routing

5.2 Profile Management

Test profile loading:

# Profiles are managed via config_manager, not directly through backend
from cogniverse_foundation.config.utils import get_config

config = get_config(tenant_id="your_org:production", config_manager=config_manager)

# List available profiles from config
profiles = config.get("backend")["profiles"]
print(f"Available profiles: {list(profiles.keys())}")

# Get profile config
if "video_colpali_smol500_mv_frame" in profiles:
    profile_config = profiles["video_colpali_smol500_mv_frame"]
    print(f"\nProfile config:")
    print(f"  Model loader: {profile_config.get('model_loader')}")
    print(f"  Embedding dim: {profile_config.get('schema_config', {}).get('embedding_dim')}")

Learning Points:

  • Profiles define processing pipelines

  • Each profile: encoder type + embedding dim + processing strategy

  • Managed via foundation config layer, not directly through backend

5.3 Tenant Schema Management

Test schema deployment:

from cogniverse_vespa.json_schema_parser import JsonSchemaParser
from cogniverse_vespa.search_backend import VespaSearchBackend
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# Initialize schema parser
parser = JsonSchemaParser()
schema = parser.load_schema_from_json_file("configs/schemas/video_colpali_smol500_mv_frame_schema.json")

# Initialize the search backend
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
backend = VespaSearchBackend(
    config={
        "url": "http://localhost",
        "port": 8080,  # Use your Vespa instance's actual port
        "profiles": {"video_colpali_smol500_mv_frame": {}},
        "default_profiles": {"video": "video_colpali_smol500_mv_frame"},
    },
    config_manager=config_manager,
    schema_loader=schema_loader,
)

# Schema deployment is handled automatically during ingestion.
# Tenant scoping is applied per query via the query_dict tenant_id;
# tenant-specific schemas follow the pattern {profile}_{tenant_id}.
print("Schema will be deployed for tenant: test_tenant")
print("Schema naming: video_colpali_smol500_mv_frame_test_tenant")

Learning Points:

  • Schema management in vespa package

  • Tenant-specific schema isolation

  • Auto-creates schemas on first use

5.4 Search Execution

Test search via backend:

# Simple search using VespaSearchBackend.search(query_dict)
# strategy is a rank-profile-name string; query_embeddings is optional
# (a numpy array) for visual/hybrid strategies.
results = backend.search({
    "query": "test video",
    "type": "video",
    "profile": "video_colpali_smol500_mv_frame",
    "strategy": "bm25_only",  # or "hybrid_float_bm25", "binary_binary", etc.
    "top_k": 5,
    "tenant_id": "your_org:production",  # REQUIRED
    # "query_embeddings": <numpy array>,  # for visual/hybrid strategies
})

# search() returns List[SearchResult]
print(f"Search results: {len(results)} found")
for i, result in enumerate(results[:3], 1):
    print(f"\n{i}. score={result.score:.4f}")
    print(f"   Source ID: {result.document.metadata['source_id']}")

Learning Points:

  • Backend abstracts Vespa queries

  • Returns normalized result format

  • Handles tenant-specific schema routing

✅ Layer 5 Complete: Backend abstraction working, profile management functional, schema deployment automated


Layer 6: Synthetic Data Layer (cogniverse_synthetic)

Purpose: Generate training data for optimizer training (implementation layer)

6.1 Service Initialization

Test SyntheticDataService:

from cogniverse_synthetic.service import SyntheticDataService
from cogniverse_foundation.config.unified_config import (
    AgentMappingRule,
    BackendConfig,
    BackendProfileConfig,
    OptimizerGenerationConfig,
    SyntheticGeneratorConfig,
)
from cogniverse_vespa import VespaBackend

# Initialize service with backend and config
profile_name = "video_colpali_smol500_mv_frame"
backend_config = BackendConfig(
    tenant_id="your_org:production",
    url="http://localhost",
    port=8080,
    profiles={
        profile_name: BackendProfileConfig(
            profile_name=profile_name,
            type="video",
            schema_name=profile_name,
            embedding_type="multi_vector",
            pipeline_config={"extract_keyframes": True},
        )
    },
)
generator_config = SyntheticGeneratorConfig(
    tenant_id="your_org:production",
    optimizer_configs={
        "modality": OptimizerGenerationConfig(
            optimizer_type="modality",
            agent_mappings=[
                AgentMappingRule(
                    modality="VIDEO",
                    agent_name="search_agent",
                )
            ],
        )
    },
)
backend = VespaBackend(
    backend_config=backend_config,
    schema_loader=schema_loader,
    config_manager=config_manager,
)
backend.initialize({"tenant_id": "your_org:production"})

service = SyntheticDataService(
    backend=backend,
    backend_config=backend_config,
    generator_config=generator_config,
    agents_config=agents_config,
)

print(f"Service initialized")
print(f"Synthetic data generators available")

Learning Points:

  • Part of implementation layer

  • Service coordinates data generation for optimization

  • Integrates with vespa package for data sampling

6.2 Profile Selection

Test profile selector:

from cogniverse_synthetic.schemas import SyntheticDataRequest

# Create request ("routing" is a registered optimizer name — see 6.4)
request = SyntheticDataRequest(
    optimizer="routing",
    count=10,
    vespa_sample_size=50,
    strategy="diverse",
    max_profiles=2,
    tenant_id="your_org:production"
)

# Generate data
response = await service.generate(request)

print(f"Generated {response.count} examples")
print(f"Selected profiles: {response.selected_profiles}")
print(f"Reasoning: {response.profile_selection_reasoning}")

Learning Points:

  • Profile selector chooses best profiles for optimizer

  • Rule-based (heuristic) or LLM-based

  • Returns reasoning for transparency

6.3 Data Generation

Inspect generated data:

# Check first example (RoutingExperienceSchema fields)
if response.data:
    example = response.data[0]
    print(f"\nSample example:")
    print(f"  Query: {example.get('query')}")
    print(f"  Chosen agent: {example.get('chosen_agent')}")
    print(f"  Routing confidence: {example.get('routing_confidence')}")

Test different optimizers:

# Test profile optimizer
request_profile = SyntheticDataRequest(
    optimizer="profile",
    count=5,
    vespa_sample_size=20,
    tenant_id="your_org:production"
)

response_profile = await service.generate(request_profile)
print(f"\nProfile examples: {response_profile.count}")
print(f"Schema: {response_profile.schema_name}")

Learning Points:

  • Each optimizer has dedicated generator

  • Schema defines example structure

  • Data sampled from real Vespa content

6.4 Optimizer Registry

Test optimizer registry:

from cogniverse_synthetic import OPTIMIZER_REGISTRY
from cogniverse_synthetic.registry import get_optimizer_config

# List all optimizers
for name in OPTIMIZER_REGISTRY.keys():
    config = get_optimizer_config(name)
    print(f"\n{name}:")
    print(f"  Description: {config.description}")
    print(f"  Schema: {config.schema_class.__name__}")
    print(f"  Generator: {config.generator_class_name}")

Learning Points:

  • Registry maps optimizer → generator + schema

  • Supports: profile, routing, workflow, unified, cross_modal

  • Extensible for new optimizers

✅ Layer 6 Complete: Synthetic data generation working, profile selection functional, optimizer-specific generators validated


Layer 7: Agent Layer (cogniverse_agents)

Purpose: Agent implementations (implementation layer)

7.1 Individual Agent Testing

Test SearchAgent:

from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# config_manager and schema_loader are REQUIRED for SearchAgent
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))

# Initialize agent (inherits from core layer base classes)
agent = SearchAgent(
    deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
    config_manager=config_manager,
    schema_loader=schema_loader,
)

# Run search (synchronous) — tenant_id is per-request
result = agent.search_by_text(
    query="machine learning tutorials",
    tenant_id="your_org:production",
    top_k=10,
)

print(f"Agent result:")
print(f"  Videos found: {len(result)}")
print(f"  Profile used: video_colpali_smol500_mv_frame")

Test the gateway router:

from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput

# GatewayAgent classifies and routes with GLiNER — deps has no required fields
router = GatewayAgent(deps=GatewayDeps())

# process is async method
decision = await router.process(GatewayInput(query="Find videos about Python", tenant_id="your_org:production"))

print(f"Routing decision: {decision}")

Learning Points:

  • Agents in implementation layer use core layer base classes

  • Integrate with vespa package for backend operations

  • Tenant-aware by default

7.2 Gateway Agent (Fast-Path Routing)

Test GatewayAgent:

from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput

# GatewayAgent classifies and routes with GLiNER — no LLM call, no registry required
router = GatewayAgent(deps=GatewayDeps())

# Test routing decision (async method)
routing_result = await router.process(
    GatewayInput(query="Show me videos about Python programming", tenant_id="your_org:production")
)

print(f"Routing decision:")
print(f"  Routed to: {routing_result.routed_to}")
print(f"  Confidence: {routing_result.confidence}")
print(f"  Reasoning: {routing_result.reasoning}")

Test routing with different query types:

# Video search query
video_routing = await router.process(GatewayInput(query="Find cooking videos", tenant_id="your_org:production"))

# Report query
report_routing = await router.process(GatewayInput(query="Create a detailed analysis of climate change", tenant_id="your_org:production"))

# Comparison queries
compare_routing = await router.process(GatewayInput(query="Compare Python and Java for web development", tenant_id="your_org:production"))

print(f"Video query → {video_routing.routed_to}")
print(f"Report query → {report_routing.routed_to}")
print(f"Compare query → {compare_routing.routed_to}")

Learning Points:

  • GatewayAgent decides which agent handles a query, using GLiNER entity detection (no LLM call)

  • Uses tiered routing: GatewayAgent (GLiNER, fast path) → OrchestratorAgent (DSPy planning, for complex multi-agent workflows)

  • Returns GatewayOutput with routed_to + confidence + reasoning

7.3 Gateway Threshold Optimization (On-Demand)

Trigger gateway-threshold optimization via the admin API:

import requests

# Trigger on-demand gateway threshold optimization
response = requests.post(
    "http://localhost:8000/admin/tenant/your_org:production/optimize",
    json={"mode": "gateway-thresholds"}
)
result = response.json()
print(f"Workflow: {result['workflow_name']}")
print(f"Status URL: {result['status_url']}")

# status_url is a relative path — resolve against the runtime base URL
status = requests.get(f"http://localhost:8000{result['status_url']}").json()
print(f"Phase: {status['phase']}")

Test the CLI function directly:

from cogniverse_runtime.optimization_cli import (
    _compute_gateway_thresholds,
    GATEWAY_DEFAULT_THRESHOLD,
)
import pandas as pd

# Phoenix stores gateway span attributes as a dict in the
# "attributes.gateway" column (complexity + confidence per span)
spans_df = pd.DataFrame({
    "attributes.gateway": [
        {"complexity": "simple", "confidence": 0.9},
        {"complexity": "complex", "confidence": 0.3},
        {"complexity": "simple", "confidence": 0.8},
        {"complexity": "complex", "confidence": 0.2},
        {"complexity": "simple", "confidence": 0.7},
    ],
})
result = _compute_gateway_thresholds(spans_df)
print(f"Computed threshold: {result['thresholds']['fast_path_confidence_threshold']:.3f}")
print(f"Default threshold: {GATEWAY_DEFAULT_THRESHOLD}")

Learning Points:

  • Gateway thresholds are computed from span data, not reward signals

  • On-demand optimization via POST /admin/tenant/{id}/optimize submits an Argo Workflow

  • GATEWAY_DEFAULT_THRESHOLD = 0.4 is the fallback when insufficient span data exists

7.4 Full Agent Roster (23 Agents)

cogniverse_agents ships 23 A2A agents (see configs/config.json → agents). Every dispatched agent is reachable through the unified runtime dispatcher at POST /agents/{agent_name}/process (registered by libs/runtime/cogniverse_runtime/routers/agents.py); the 7 knowledge-graph agents additionally have bespoke routes under /admin/tenants/{tenant_id}/knowledge/... (libs/runtime/cogniverse_runtime/routers/knowledge.py). Agents marked disabled below have "enabled": false in configs/config.json and must be enabled before the dispatcher will route to them.

Generation + Routing Agents:

Agent Port Enabled Purpose
gateway_agent 8000 (mounted) Yes GLiNER-based fast-path classifier/router (see 7.2)
orchestrator_agent 8013 Yes DSPy two-phase planning + multi-agent execution for complex queries
summarizer_agent 8004 Yes Summarizes search/agent results
detailed_report_agent 8005 Yes Generates detailed multi-section reports
profile_selection_agent 8000 (mounted) Yes Picks the best backend profile for a query (modality + complexity + intent)
query_enhancement_agent 8000 (mounted) Yes Rewrites/expands queries before retrieval
entity_extraction_agent 8000 (mounted) Yes Extracts entities/relationships for routing and KG ingestion

Search & Analysis Agents:

Agent Port Enabled Purpose
search_agent 8002 Yes Video/multi-modal search via the backend (see 7.1)
image_search_agent 8006 Yes Image-specific search
document_agent 8008 Yes Document search and retrieval
text_analysis_agent 8003 Yes Text content analysis
audio_analysis_agent 8007 Yes Audio content analysis/transcription search

Research + Coding Agents:

Agent Port Enabled Purpose
deep_research_agent 8009 Yes Multi-step web/document research
coding_agent 8010 Yes Code generation/analysis tasks

Knowledge-Graph & Reasoning Agents:

Agent Port Enabled Purpose
citation_tracing_agent 8019 No Traces claims back to source citations
contradiction_reconciliation_agent 8020 No Reconciles contradictory claims in the knowledge graph
multi_document_synthesis_agent 8021 No Synthesizes findings across multiple documents
kg_traversal_agent 8022 No Multi-hop knowledge-graph traversal
temporal_reasoning_agent 8025 No Reasons about time-ordered facts/events
knowledge_summarization_agent 8026 No Summarizes knowledge-graph subgraphs
audit_explanation_agent 8027 Yes Explains why a decision/answer was produced

Multi-Tenant + Federation Agents:

Agent Port Enabled Purpose
cross_tenant_comparison_agent 8023 No Compares one subject's view across tenants in an org (ACL-gated, no LLM)
federated_query_agent 8024 No Aggregates federated reads across tenants in an org, with optional RLM summarization

Generic per-agent smoke test (dispatcher path):

# List all registered agents
curl http://localhost:8000/agents/

# Get one agent's card (capabilities, health endpoint)
curl http://localhost:8000/agents/search_agent/card

# Dispatch a request to any agent by name
curl -X POST http://localhost:8000/agents/search_agent/process \
  -H "Content-Type: application/json" \
  -d '{"query": "test video", "tenant_id": "default"}'

Learning Points:

  • 23 agents total: 15 enabled by default, 8 disabled (the KG-reasoning and federation agents, except audit_explanation_agent)

  • gateway_agent, profile_selection_agent, query_enhancement_agent, and entity_extraction_agent are mounted inside the main runtime process (port 8000) rather than run as standalone services

  • KG-reasoning agents also expose direct REST routes under /admin/tenants/{tenant_id}/knowledge/* in addition to the generic dispatcher path

✅ Layer 7 Complete: Individual agents working, routing decisions functional, optimizers trainable


Layer 8: Runtime Layer (cogniverse_runtime)

Purpose: FastAPI server (application layer) exposing all functionality via REST API

8.1 Runtime Service Startup

Start runtime server:

# Start server (application layer)
JAX_PLATFORM_NAME=cpu uv run python -m cogniverse_runtime.main

# Verify startup in logs
# Expected: "Uvicorn running on http://0.0.0.0:8000"

Check health endpoint:

curl http://localhost:8000/health

# Expected: {"status": "healthy", "service": "cogniverse-runtime",
#            "backends": {"registered": N, "backends": [...]},
#            "agents": {"registered": N, "agents": [...]}}
# Returns HTTP 503 with {"status": "unhealthy", ...} if the registries
# can't be assembled (e.g. missing BACKEND_URL)

Learning Points:

  • Runtime is in application layer

  • Integrates agents, vespa, synthetic, and evaluation packages

  • Port 8000 (default)

  • Health check validates all services

8.2 Search Endpoints

Test search API:

# Unified search endpoint
curl -X POST http://localhost:8000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "machine learning tutorials",
    "tenant_id": "default",
    "profile": "video_colpali_smol500_mv_frame",
    "strategy": "hybrid",
    "top_k": 5
  }'

# Should return JSON with results

Test streaming search:

curl -X POST http://localhost:8000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "Python programming",
    "tenant_id": "default",
    "profile": "video_colpali_smol500_mv_frame",
    "top_k": 10,
    "stream": true
  }'

# Returns SSE stream for real-time results

Learning Points:

  • Single unified /search/ endpoint for all search operations

  • Set stream: true for server-sent events (SSE) streaming

  • Results include profile info and result count

8.3 Routing Endpoints

Test routing via agent:

# Fast-path routing is handled by GatewayAgent via POST /agents/gateway_agent/process
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput

# GatewayAgent classifies and routes with GLiNER — deps has no required fields
router = GatewayAgent(deps=GatewayDeps())

# Route query (async method)
decision = await router.process(GatewayInput(query="Show me cooking videos", tenant_id="your_org:production"))

print(f"Routing decision:")
print(f"  Routed to: {decision.routed_to}")
print(f"  Confidence: {decision.confidence}")
print(f"  Reasoning: {decision.reasoning}")

Learning Points:

  • Fast-path routing is handled by GatewayAgent via the A2A endpoint

  • Returns GatewayOutput with routed_to, confidence, and reasoning

  • Complex queries are routed to OrchestratorAgent for multi-step planning

8.4 Synthetic Data Endpoints

Test synthetic data generation:

# Reuse the live service initialized in Layer 6.1.
from cogniverse_synthetic.schemas import SyntheticDataRequest

# Create request ("routing" is a registered optimizer name)
request = SyntheticDataRequest(
    optimizer="routing",
    count=10,
    vespa_sample_size=50,
    strategy="diverse",
    max_profiles=2,
    tenant_id="your_org:production"
)

# Generate data
response = await service.generate(request)
print(f"Generated {response.count} examples")
print(f"Selected profiles: {response.selected_profiles}")

List available optimizers:

from cogniverse_synthetic import OPTIMIZER_REGISTRY

for name in OPTIMIZER_REGISTRY.keys():
    print(f"- {name}")

Learning Points:

  • Synthetic data service used programmatically (implementation layer)

  • OPTIMIZER_REGISTRY lists available optimizer types

  • Integrates with backend for content sampling

8.5 Admin Endpoints

cogniverse_runtime.admin.tenant_manager defines the /admin/organizations and /admin/tenants routes. main.py mounts this same router on the unified runtime (port 8000), so these endpoints work there in normal operation. The module can also run standalone (uv run python -m cogniverse_runtime.admin.tenant_manager, default port 9000) for isolated testing without booting the full runtime.

Create organization (via the unified runtime, port 8000):

curl -X POST http://localhost:8000/admin/organizations \
  -H "Content-Type: application/json" \
  -d '{
    "org_id": "acme",
    "org_name": "Acme Corporation",
    "created_by": "admin"
  }'

Create tenant (via the unified runtime, port 8000):

curl -X POST http://localhost:8000/admin/tenants \
  -H "Content-Type: application/json" \
  -d '{
    "tenant_id": "acme:production",
    "created_by": "admin"
  }'

Learning Points:

  • Tenant management endpoints are mounted on the main runtime (port 8000) and are also runnable standalone via cogniverse_runtime.admin.tenant_manager (default port 9000) for isolated testing

  • Organization → Tenant hierarchy

  • Tenant format: {org}:{name}

✅ Layer 8 Complete: Runtime API functional, all endpoints responding, multi-tenant support working


Layer 9: End-to-End Integration Tests

All 12 Packages Working Together

Purpose: Validate complete workflows across all layers

9.1 Complete Search Workflow

Test full search pipeline:

# 1. Ingest test video
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
    --video_dir /tmp/test_video \
    --backend vespa \
    --profile video_colpali_smol500_mv_frame \
    --tenant-id e2e_test

# 2. Run comprehensive search test
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
    --profiles video_colpali_smol500_mv_frame \
    --test-multiple-strategies \
    --max-queries 5

# 3. Verify Phoenix captured spans

Validate in Phoenix:

from phoenix.client import Client

client = Client(base_url='http://localhost:6006')
spans_df = client.spans.get_spans_dataframe(
    project_name="cogniverse-e2e_test-search",
    timeout=60,
)

print(f"Search spans collected: {len(spans_df)}")
print(f"Avg latency: {spans_df['latency_ms'].mean():.2f}ms")

Learning Points:

  • Complete flow: Ingest → Search → Telemetry

  • All layers working together

  • End-to-end latency tracking

9.2 Routing + Search Workflow

Test routing to search:

# Complete routing + search workflow
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path

# 1. Route the query (async method) — GatewayAgent classifies with GLiNER
router = GatewayAgent(deps=GatewayDeps())
decision = await router.process(GatewayInput(query="Find videos about machine learning", tenant_id="your_org:production"))

print(f"Routing: {decision.routed_to}")

# 2. If search_agent, execute search
if decision.routed_to == "search_agent":
    config_manager = create_default_config_manager()
    schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
    agent = SearchAgent(
        deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
        config_manager=config_manager,
        schema_loader=schema_loader,
    )
    results = agent.search_by_text(query="machine learning", tenant_id="your_org:production", top_k=5)
    print(f"Found {len(results)} results")

Learning Points:

  • Routing determines agent

  • Agent executes search

  • Results returned to user

9.3 Optimization Workflow

Test complete optimization cycle:

# 1. Generate synthetic data (if needed)
# This is typically done programmatically, see Layer 6 examples above

# 2. Run profile optimization
JAX_PLATFORM_NAME=cpu uv run python -m cogniverse_runtime.optimization_cli \
  --mode profile --tenant-id default

# 3. Check results in Phoenix — the optimization run creates a new
# experiment with the optimized module's scores vs baseline

# 4. Verify the compiled artifact was saved to the telemetry provider's
# dataset store (loaded by agents automatically on next restart)

Learning Points:

  • Synthetic data → Training → Optimized model

  • Complete optimization pipeline

  • Results include improvement metrics

9.4 Multi-Tenant Workflow

Test tenant isolation:

# 1. Create organization (unified runtime, port 8000 — see 8.5)
curl -X POST http://localhost:8000/admin/organizations \
  -H "Content-Type: application/json" \
  -d '{"org_id": "test_org", "org_name": "Test Org", "created_by": "admin"}'

# 2. Create tenant (unified runtime, port 8000 — see 8.5)
curl -X POST http://localhost:8000/admin/tenants \
  -H "Content-Type: application/json" \
  -d '{"tenant_id": "test_org:dev", "org_id": "test_org", "created_by": "admin"}'

# 3. Ingest video for tenant
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
    --video_dir /tmp/test_video \
    --backend vespa \
    --profile video_colpali_smol500_mv_frame \
    --tenant-id test_org:dev

# 4. Search with tenant (unified /search/ endpoint)
curl -X POST http://localhost:8000/search/ \
  -H "Content-Type: application/json" \
  -d '{
    "query": "test",
    "tenant_id": "test_org:dev",
    "profile": "video_colpali_smol500_mv_frame"
  }'

# 5. Verify data isolation
# Should not see test_org:dev data when searching with "default" tenant

Learning Points:

  • Complete tenant isolation

  • Separate schemas per tenant

  • Config/telemetry/memory isolated

9.5 Argo Workflow Integration

The scheduled optimization CronWorkflows are Helm templates (charts/cogniverse/templates/optimization-workflows.yaml), not a standalone kubectl apply -f manifest — they're deployed as part of the cogniverse chart release.

Test scheduled optimization:

# 1. Deploy/upgrade the chart (installs the CronWorkflows)
helm upgrade --install cogniverse ./charts/cogniverse -n cogniverse

# 2. Check CronWorkflows created (names are {release}-agent-optimization,
#    {release}-daily-cleanup, etc.)
kubectl get cronworkflow -n cogniverse

# 3. Trigger the weekly agent-optimization CronWorkflow manually
argo submit --from cronwf/cogniverse-agent-optimization -n cogniverse

# 4. Monitor workflow
argo list -n cogniverse
argo logs <workflow-name> -n cogniverse --follow

# 5. Check results
argo get <workflow-name> -n cogniverse -o json | \
  jq '.status.outputs.parameters'

Learning Points:

  • Argo CronWorkflows for batch jobs, defined as Helm templates under charts/cogniverse/templates/

  • agentOptimization runs weekly by default ("0 3 * * 0", Sunday 3 AM UTC, per values.yaml); dailyGateway and dailyCleanup run daily

  • Kubernetes integration via the Argo Workflows CRDs

✅ Layer 9 Complete: End-to-end workflows validated, multi-tenant isolation verified, all integrations working


Final Validation Checklist

Infrastructure

  • Vespa running and healthy
  • Phoenix running and collecting spans
  • Mem0 initialized and storing memories
  • Runtime API responding on all endpoints

Data Layer

  • Videos ingested across multiple profiles
  • Documents searchable via Vespa
  • Embeddings generated correctly
  • Search results ranked properly

Telemetry

  • Phoenix projects created per tenant
  • Spans collected for search/routing/orchestration
  • Span attributes populated correctly
  • Web client Analytics view showing metrics

Configuration

  • System config loaded
  • Tenant-specific overrides working
  • Backend auto-discovery functional
  • Profile management operational

Agents

  • Individual agents (search, summarizer, detailed_report, and others from the 23-agent roster in 7.4) working
  • Gateway agent (fast-path) and orchestrator agent (complex-query planning) making correct routing decisions
  • Agent memory integration functional
  • Multi-agent orchestration working

Optimization

  • Synthetic data generation producing quality examples for each registered optimizer (profile, routing, workflow, unified, cross_modal)
  • Profile optimizer improving profile-selection accuracy
  • Cross-modal optimizer learning fusion patterns
  • DSPy-based routing optimizer (SIMBA/MIPROv2/BootstrapFewShot, auto-selected by training-data size) improving routing decisions
  • Argo workflows submitting and executing

Multi-Tenant

  • Tenant isolation validated
  • Separate schemas per tenant
  • Config/telemetry/memory isolated
  • Cross-tenant data leakage prevented

UI/UX

  • Web client loading all views
  • Analytics showing real data
  • Config management CRUD working
  • Optimization workflows submittable from UI

Troubleshooting Guide

Common Issues

Vespa Connection Failed:

  • Check: curl http://localhost:8080/state/v1/health

  • Fix: Start Vespa service

  • Verify: Port 8080 accessible

Phoenix Not Collecting Spans:

  • Check: Phoenix server running on port 6006

  • Fix: Set PHOENIX_ENDPOINT env var

  • Verify: Spans appear in Phoenix UI

Memory Manager Initialization Failed:

  • Check: mem0ai dependency installed (pinned as mem0ai==1.0.11 in pyproject.toml; already present after uv sync)

  • Fix: uv sync (or uv pip install mem0ai==1.0.11 to add it standalone)

  • Verify: Can create Mem0MemoryManager(tenant_id="your_org:production")

Backend Config Not Found:

  • Check: config.json exists in standard locations

  • Fix: Create config.json with backend section

  • Verify: create_default_config_manager().get_backend_config(tenant_id=...) succeeds

Optimization Training Failed:

  • Check: Enough training examples exist for the mode (--mode simba/profile/entity-extraction scale their DSPy teleprompter choice by training-set size; too few examples still runs, but with weaker few-shot demos)

  • Fix: Generate more examples first via --mode synthetic, or the synthetic mode in the web client's Optimization runs view

  • Verify: check the CLI's logged training_examples count in the run output

Argo Workflow Submission Failed:

  • Check: kubectl or argo CLI configured

  • Fix: Install argo CLI, configure kubectl

  • Verify: argo version succeeds


Learning Outcomes

After completing this testing plan, you will understand:

  1. Architecture: How all layers interact bottom-up
  2. Data Flow: From ingestion → storage → search → telemetry
  3. Tenant Isolation: How multi-tenancy works across layers
  4. Configuration: Auto-discovery and override mechanisms
  5. Optimization: Complete optimization lifecycle
  6. Deployment: Argo Workflows for production automation
  7. Observability: Phoenix telemetry and analytics
  8. APIs: REST endpoints for all functionality
  9. UI Integration: Web client for all operations
  10. Troubleshooting: Common issues and solutions

Next Steps

After completing all layers:

  1. Run full test suite: JAX_PLATFORM_NAME=cpu timeout 7200 uv run pytest
  2. Deploy to staging: Test with real workloads
  3. Performance testing: Benchmark search latency, optimization throughput
  4. Security audit: Verify tenant isolation, auth mechanisms
  5. Documentation review: Update any gaps discovered during testing
  6. Production deployment: Deploy with monitoring and alerting

Testing Tips:

  • Document issues as you find them

  • Take notes on architecture insights

  • Save successful commands for future reference

  • Test edge cases (empty queries, large documents, etc.)

  • Validate error handling (network failures, invalid inputs)

Good luck with your comprehensive testing! 🚀