Comprehensive System Testing Plan¶
Purpose: Bottom-up manual testing and learning guide for Cogniverse Approach: Start with foundational subsystems, validate each layer, build upward Goal: Thorough understanding and validation of the complete system
Testing Philosophy¶
Bottom-Up Approach:
-
Validate foundation before building on it
-
Each layer depends on layers below it
-
Fix issues at lowest level first
-
Document learnings and discoveries
For Each Subsystem:
-
✅ Purpose: Understand what it does and why
-
✅ Basic Tests: Core functionality works
-
✅ Advanced Tests: Edge cases and features
-
✅ Integration Tests: Works with dependent layers
-
✅ Learnings: Document key insights
Layer 1: Storage Foundation (Vespa)¶
Purpose: Vector database for embeddings and metadata storage
1.1 Vespa Service Health¶
Test Vespa is running:
# Check Vespa service status
curl http://localhost:8080/state/v1/health
# Expected: {"status": {"code": "up"}}
Verify Vespa configuration endpoint:
Learning Points:
-
Vespa is the foundation - everything else needs it
-
Default port: 8080
-
State API provides health checks
1.2 Schema Validation¶
Check deployed schemas:
# List all schemas
curl "http://localhost:8080/search/?yql=select+*+from+sources+*+where+true+limit+0"
# Check specific tenant schema exists
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_default+where+true+limit+1"
Test schema structure:
# Get document count per schema
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_default+where+true+limit+0" | jq '.root.fields.totalCount'
Learning Points:
-
Schemas are tenant-isolated:
{profile}_{tenant_id} -
Multiple profiles can exist per tenant
-
Default tenant: "default"
1.3 Document Ingestion Test¶
Ingest a test video:
# Create test video directory
mkdir -p /tmp/test_video
# Download or copy a sample video (MP4)
# For testing: use any short video file
# Run ingestion for single profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir /tmp/test_video \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
--tenant-id test_basic
# Monitor logs
tail -f outputs/logs/ingestion_*.log
Verify ingestion:
# Check document count increased
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+true+limit+10" | jq '.root.fields.totalCount'
# Inspect a document
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+true+limit+1" | jq '.root.children[0].fields'
Learning Points:
-
Ingestion creates embeddings and metadata
-
Documents have video_id, frame_number, embeddings, metadata
-
Each profile has different embedding dimensions
1.4 Basic Search Test¶
Test vector search:
# Simple text query
curl "http://localhost:8080/search/?yql=select+*+from+video_colpali_smol500_mv_frame_test_basic+where+userQuery()&query=test+video&hits=5"
# Check results structure
curl "http://localhost:8080/search/?yql=select+video_id,video_title,frame_number+from+video_colpali_smol500_mv_frame_test_basic+where+userQuery()&query=test&hits=3" | jq '.root.children[] | {video_id: .fields.video_id, title: .fields.video_title}'
Learning Points:
-
Vespa provides vector search via embeddings
-
Results ranked by relevance
-
Can filter by fields (video_id, timestamp, etc.)
✅ Layer 1 Complete: Vespa storage working, schemas deployed, documents ingested and searchable
Layer 2: Telemetry Foundation (Phoenix)¶
Purpose: Observability and tracing for all operations
2.1 Phoenix Service Health¶
Check Phoenix server:
Access Phoenix UI:
Learning Points:
-
Phoenix: observability platform for LLM apps
-
Port 6006
-
UI provides span visualization
2.2 Project Structure¶
List Phoenix projects:
# Run in Python REPL
from phoenix.client import Client
client = Client(base_url='http://localhost:6006')
# List projects (returns a list of dict-like phoenix.client v1.Project)
projects = client.projects.list()
for project in projects:
print(f"Project: {project['name']}")
Expected projects:
-
cogniverse-test:unit(search, routing and orchestration spans of tenanttest:unit, the tenant the run below searches as) -
cogniverse-test:unit-{service}for a management service's spans, for examplecogniverse-test:unit-synthetic_data
Learning Points:
-
Projects isolate telemetry by tenant
-
Naming:
cogniverse-{tenant_id}for user operations,cogniverse-{tenant_id}-{service}for management operations
2.3 Span Collection Test¶
Generate telemetry by running a search:
# Run comprehensive test (generates spans)
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
--profiles video_colpali_smol500_mv_frame \
--test-multiple-strategies \
--max-queries 3
# Wait for execution to complete
Query spans from Phoenix:
from phoenix.client import Client
from datetime import datetime, timedelta, timezone
client = Client(base_url='http://localhost:6006')
# Get search spans. `timeout=` is required explicitly — the method's own
# default is 5s (independent of any client-level timeout) and a loaded
# project can easily exceed that, silently reading as "no spans".
spans_df = client.spans.get_spans_dataframe(
project_name="cogniverse-test:unit",
start_time=datetime.now(timezone.utc) - timedelta(hours=1),
timeout=60,
)
search_spans = spans_df[spans_df["name"] == "search_service.search"]
print(f"Search spans collected: {len(search_spans)}")
print(f"Columns: {search_spans.columns.tolist()}")
# Inspect a span
if len(search_spans) > 0:
print("\nSample span:")
print(search_spans.iloc[0][['name', 'attributes.query', 'attributes.latency_ms']])
Learning Points:
-
Spans capture operation traces
-
Attributes store metadata (query, profile, strategy, latency, result rows)
-
Can query by time range and project
2.4 Span Attributes Validation¶
Check span structure:
# Get span with all attributes
span = search_spans.iloc[0]
# Key attributes for search spans
search_attributes = [
'attributes.query',
'attributes.profile',
'attributes.strategy',
'attributes.top_k',
'attributes.backend',
'attributes.result_granularity',
'attributes.latency_ms',
'attributes.num_results',
'attributes.top_score',
'attributes.output.value', # JSON list of the result rows
'attributes.tenant', # {'id': <tenant_id>}
]
for attr in search_attributes:
if attr in span.index:
print(f"{attr}: {span[attr]}")
Learning Points:
-
Search spans include query, profile, strategy, result count and the result rows (
output.value) -
Can filter spans by attributes
-
Used for optimization and analytics
✅ Layer 2 Complete: Phoenix collecting telemetry, spans queryable, attributes validated
Layer 3: Memory Foundation (Mem0)¶
Purpose: Conversation memory and context for agents
3.1 Memory Service Health¶
Check Mem0 initialization:
from cogniverse_core.memory.manager import Mem0MemoryManager
# Initialize manager (requires tenant_id)
manager = Mem0MemoryManager(tenant_id="your_org:production")
manager.initialize(
backend_host="localhost",
backend_port=8080,
llm_model="openai/google/gemma-4-e4b-it",
embedding_model="lightonai/DenseOn",
llm_base_url="http://localhost:11434",
embedder_base_url="http://localhost:8002", # DenseOn sidecar endpoint
config_manager=config_manager,
schema_loader=schema_loader,
)
print(f"Memory manager initialized: {manager.memory is not None}")
Learning Points:
-
Mem0 stores agent memories and user preferences
-
Initialized on first use
-
Isolated by tenant_id + agent_name
3.2 Memory Storage Test¶
Add a memory:
# Add memory for routing agent
manager.add_memory(
content="User prefers video results for cooking queries",
tenant_id="your_org:production",
agent_name="orchestrator_agent",
metadata={"context": "preference_learning"}
)
print("Memory added successfully")
Retrieve memory:
# Search for memory
results = manager.search_memory(
query="cooking preferences",
tenant_id="your_org:production",
agent_name="orchestrator_agent",
top_k=5
)
for i, result in enumerate(results, 1):
print(f"\nMemory {i}:")
print(f" Content: {result['memory']}")
print(f" Score: {result.get('score', 0.0):.3f}")
print(f" Metadata: {result.get('metadata', {})}")
Note: search_memory returns List[Dict[str, Any]] (plain dicts, not objects) — always use result[...]/result.get(...), never attribute access.
Learning Points:
-
Memories are searchable (semantic search)
-
Score indicates relevance
-
Metadata provides context
3.3 Memory Lifecycle¶
Get all memories:
# Get all memories for agent
all_memories = manager.get_all_memories(
tenant_id="your_org:production",
agent_name="orchestrator_agent"
)
print(f"Total memories: {len(all_memories)}")
Delete a memory:
# Get memory ID
memory_id = all_memories[0]['id'] if all_memories else None
if memory_id:
manager.delete_memory(
memory_id=memory_id,
tenant_id="your_org:production",
agent_name="orchestrator_agent"
)
print(f"Deleted memory: {memory_id}")
Learning Points:
-
Memories persist across sessions
-
Can be deleted individually or bulk cleared
-
Memory manager handles CRUD operations
✅ Layer 3 Complete: Mem0 storing and retrieving memories, lifecycle operations working
Layer 4: Core Library (cogniverse_core)¶
Purpose: Configuration, telemetry, and memory management abstractions
4.1 Configuration Management¶
Test SystemConfig:
from cogniverse_foundation.config.unified_config import SystemConfig
# SystemConfig is GLOBAL — one per deployment, not per-tenant.
# It has no tenant_id field.
system_config = SystemConfig(
llm_model="gpt-4",
backend_url="http://localhost",
backend_port=8080
)
print(f"LLM Model: {system_config.llm_model}")
print(f"Search Backend: {system_config.search_backend}")
print(f"Backend URL: {system_config.backend_url}")
print(f"Backend Port: {system_config.backend_port}")
Test tenant-specific override (via BackendConfig):
from cogniverse_foundation.config.unified_config import BackendConfig
from cogniverse_foundation.config.utils import create_default_config_manager
# Tenant-specific overrides use BackendConfig (not SystemConfig, which
# is global) and are persisted via ConfigManager
config_manager = create_default_config_manager()
tenant_backend_config = BackendConfig(
tenant_id="test_tenant",
url="http://localhost",
port=8080,
)
config_manager.set_backend_config(tenant_backend_config)
print(f"Tenant backend URL: {tenant_backend_config.url}")
Learning Points:
-
SystemConfig defined in foundation layer; it is GLOBAL (one per deployment), not per-tenant
-
Tenant-specific configurations (backend URL, profiles) use BackendConfig via ConfigManager
-
Default configs in libs/foundation/
4.2 Telemetry Manager¶
Test TelemetryProvider:
from cogniverse_foundation.telemetry.config import TelemetryConfig
from cogniverse_telemetry_phoenix.provider import PhoenixProvider
# TelemetryConfig is generic (provider-agnostic)
config = TelemetryConfig()
# PhoenixProvider takes no constructor args — it's configured via initialize()
telemetry = PhoenixProvider()
telemetry.initialize({
"tenant_id": "your_org:production",
"http_endpoint": "http://localhost:6006",
"grpc_endpoint": "localhost:4317",
})
# Check Phoenix connection
print(f"Phoenix endpoint: http://localhost:6006")
print(f"Telemetry enabled: {config.enabled}")
Learning Points:
-
PhoenixProvider (telemetry-phoenix package) implements TelemetryProvider interface (foundation)
-
Auto-creates projects per tenant
-
Provides span export utilities via foundation layer interfaces
4.3 Tenant-Aware Components¶
Test TenantAwareAgentMixin:
from cogniverse_core.agents.tenant_aware_mixin import TenantAwareAgentMixin
class TestComponent(TenantAwareAgentMixin):
def __init__(self, tenant_id: str):
super().__init__(tenant_id=tenant_id)
def get_info(self):
return f"Tenant: {self.tenant_id}"
# Create instance
component = TestComponent(tenant_id="acme:production")
print(component.get_info())
Learning Points:
-
TenantAwareAgentMixin provides tenant isolation
-
All agents/components inherit this
✅ Layer 4 Complete: Core abstractions working, config/telemetry/memory managers functional
Layer 5: Backend Layer (cogniverse_vespa)¶
Purpose: Search backend implementation using sdk interfaces
5.1 Backend Implementation¶
Test backend usage:
from cogniverse_vespa.search_backend import VespaSearchBackend
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# VespaSearchBackend resolves schema/profile at query time
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
backend = VespaSearchBackend(
config={
"url": "http://localhost",
"port": 8080, # Use your Vespa instance's actual port
"profiles": {"video_colpali_smol500_mv_frame": {}},
"default_profiles": {"video": "video_colpali_smol500_mv_frame"},
},
config_manager=config_manager,
schema_loader=schema_loader,
)
print(f"Backend type: {type(backend).__name__}")
Learning Points:
-
Backend implementations (vespa package) use interfaces from sdk layer
-
Provides tenant-aware search operations
-
Handles schema management and routing
5.2 Profile Management¶
Test profile loading:
# Profiles are managed via config_manager, not directly through backend
from cogniverse_foundation.config.utils import get_config
config = get_config(tenant_id="your_org:production", config_manager=config_manager)
# List available profiles from config
profiles = config.get("backend")["profiles"]
print(f"Available profiles: {list(profiles.keys())}")
# Get profile config
if "video_colpali_smol500_mv_frame" in profiles:
profile_config = profiles["video_colpali_smol500_mv_frame"]
print(f"\nProfile config:")
print(f" Model loader: {profile_config.get('model_loader')}")
print(f" Embedding dim: {profile_config.get('schema_config', {}).get('embedding_dim')}")
Learning Points:
-
Profiles define processing pipelines
-
Each profile: encoder type + embedding dim + processing strategy
-
Managed via foundation config layer, not directly through backend
5.3 Tenant Schema Management¶
Test schema deployment:
from cogniverse_vespa.json_schema_parser import JsonSchemaParser
from cogniverse_vespa.search_backend import VespaSearchBackend
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# Initialize schema parser
parser = JsonSchemaParser()
schema = parser.load_schema_from_json_file("configs/schemas/video_colpali_smol500_mv_frame_schema.json")
# Initialize the search backend
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
backend = VespaSearchBackend(
config={
"url": "http://localhost",
"port": 8080, # Use your Vespa instance's actual port
"profiles": {"video_colpali_smol500_mv_frame": {}},
"default_profiles": {"video": "video_colpali_smol500_mv_frame"},
},
config_manager=config_manager,
schema_loader=schema_loader,
)
# Schema deployment is handled automatically during ingestion.
# Tenant scoping is applied per query via the query_dict tenant_id;
# tenant-specific schemas follow the pattern {profile}_{tenant_id}.
print("Schema will be deployed for tenant: test_tenant")
print("Schema naming: video_colpali_smol500_mv_frame_test_tenant")
Learning Points:
-
Schema management in vespa package
-
Tenant-specific schema isolation
-
Auto-creates schemas on first use
5.4 Search Execution¶
Test search via backend:
# Simple search using VespaSearchBackend.search(query_dict)
# strategy is a rank-profile-name string; query_embeddings is optional
# (a numpy array) for visual/hybrid strategies.
results = backend.search({
"query": "test video",
"type": "video",
"profile": "video_colpali_smol500_mv_frame",
"strategy": "bm25_only", # or "hybrid_float_bm25", "binary_binary", etc.
"top_k": 5,
"tenant_id": "your_org:production", # REQUIRED
# "query_embeddings": <numpy array>, # for visual/hybrid strategies
})
# search() returns List[SearchResult]
print(f"Search results: {len(results)} found")
for i, result in enumerate(results[:3], 1):
print(f"\n{i}. score={result.score:.4f}")
print(f" Source ID: {result.document.metadata['source_id']}")
Learning Points:
-
Backend abstracts Vespa queries
-
Returns normalized result format
-
Handles tenant-specific schema routing
✅ Layer 5 Complete: Backend abstraction working, profile management functional, schema deployment automated
Layer 6: Synthetic Data Layer (cogniverse_synthetic)¶
Purpose: Generate training data for optimizer training (implementation layer)
6.1 Service Initialization¶
Test SyntheticDataService:
from cogniverse_synthetic.service import SyntheticDataService
from cogniverse_foundation.config.unified_config import (
AgentMappingRule,
BackendConfig,
BackendProfileConfig,
OptimizerGenerationConfig,
SyntheticGeneratorConfig,
)
from cogniverse_vespa import VespaBackend
# Initialize service with backend and config
profile_name = "video_colpali_smol500_mv_frame"
backend_config = BackendConfig(
tenant_id="your_org:production",
url="http://localhost",
port=8080,
profiles={
profile_name: BackendProfileConfig(
profile_name=profile_name,
type="video",
schema_name=profile_name,
embedding_type="multi_vector",
pipeline_config={"extract_keyframes": True},
)
},
)
generator_config = SyntheticGeneratorConfig(
tenant_id="your_org:production",
optimizer_configs={
"modality": OptimizerGenerationConfig(
optimizer_type="modality",
agent_mappings=[
AgentMappingRule(
modality="VIDEO",
agent_name="search_agent",
)
],
)
},
)
backend = VespaBackend(
backend_config=backend_config,
schema_loader=schema_loader,
config_manager=config_manager,
)
backend.initialize({"tenant_id": "your_org:production"})
service = SyntheticDataService(
backend=backend,
backend_config=backend_config,
generator_config=generator_config,
agents_config=agents_config,
)
print(f"Service initialized")
print(f"Synthetic data generators available")
Learning Points:
-
Part of implementation layer
-
Service coordinates data generation for optimization
-
Integrates with vespa package for data sampling
6.2 Profile Selection¶
Test profile selector:
from cogniverse_synthetic.schemas import SyntheticDataRequest
# Create request ("routing" is a registered optimizer name — see 6.4)
request = SyntheticDataRequest(
optimizer="routing",
count=10,
vespa_sample_size=50,
strategy="diverse",
max_profiles=2,
tenant_id="your_org:production"
)
# Generate data
response = await service.generate(request)
print(f"Generated {response.count} examples")
print(f"Selected profiles: {response.selected_profiles}")
print(f"Reasoning: {response.profile_selection_reasoning}")
Learning Points:
-
Profile selector chooses best profiles for optimizer
-
Rule-based (heuristic) or LLM-based
-
Returns reasoning for transparency
6.3 Data Generation¶
Inspect generated data:
# Check first example (RoutingExperienceSchema fields)
if response.data:
example = response.data[0]
print(f"\nSample example:")
print(f" Query: {example.get('query')}")
print(f" Chosen agent: {example.get('chosen_agent')}")
print(f" Routing confidence: {example.get('routing_confidence')}")
Test different optimizers:
# Test profile optimizer
request_profile = SyntheticDataRequest(
optimizer="profile",
count=5,
vespa_sample_size=20,
tenant_id="your_org:production"
)
response_profile = await service.generate(request_profile)
print(f"\nProfile examples: {response_profile.count}")
print(f"Schema: {response_profile.schema_name}")
Learning Points:
-
Each optimizer has dedicated generator
-
Schema defines example structure
-
Data sampled from real Vespa content
6.4 Optimizer Registry¶
Test optimizer registry:
from cogniverse_synthetic import OPTIMIZER_REGISTRY
from cogniverse_synthetic.registry import get_optimizer_config
# List all optimizers
for name in OPTIMIZER_REGISTRY.keys():
config = get_optimizer_config(name)
print(f"\n{name}:")
print(f" Description: {config.description}")
print(f" Schema: {config.schema_class.__name__}")
print(f" Generator: {config.generator_class_name}")
Learning Points:
-
Registry maps optimizer → generator + schema
-
Supports: profile, routing, workflow, unified, cross_modal
-
Extensible for new optimizers
✅ Layer 6 Complete: Synthetic data generation working, profile selection functional, optimizer-specific generators validated
Layer 7: Agent Layer (cogniverse_agents)¶
Purpose: Agent implementations (implementation layer)
7.1 Individual Agent Testing¶
Test SearchAgent:
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# config_manager and schema_loader are REQUIRED for SearchAgent
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
# Initialize agent (inherits from core layer base classes)
agent = SearchAgent(
deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
config_manager=config_manager,
schema_loader=schema_loader,
)
# Run search (synchronous) — tenant_id is per-request
result = agent.search_by_text(
query="machine learning tutorials",
tenant_id="your_org:production",
top_k=10,
)
print(f"Agent result:")
print(f" Videos found: {len(result)}")
print(f" Profile used: video_colpali_smol500_mv_frame")
Test the gateway router:
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput
# GatewayAgent classifies and routes with GLiNER — deps has no required fields
router = GatewayAgent(deps=GatewayDeps())
# process is async method
decision = await router.process(GatewayInput(query="Find videos about Python", tenant_id="your_org:production"))
print(f"Routing decision: {decision}")
Learning Points:
-
Agents in implementation layer use core layer base classes
-
Integrate with vespa package for backend operations
-
Tenant-aware by default
7.2 Gateway Agent (Fast-Path Routing)¶
Test GatewayAgent:
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput
# GatewayAgent classifies and routes with GLiNER — no LLM call, no registry required
router = GatewayAgent(deps=GatewayDeps())
# Test routing decision (async method)
routing_result = await router.process(
GatewayInput(query="Show me videos about Python programming", tenant_id="your_org:production")
)
print(f"Routing decision:")
print(f" Routed to: {routing_result.routed_to}")
print(f" Confidence: {routing_result.confidence}")
print(f" Reasoning: {routing_result.reasoning}")
Test routing with different query types:
# Video search query
video_routing = await router.process(GatewayInput(query="Find cooking videos", tenant_id="your_org:production"))
# Report query
report_routing = await router.process(GatewayInput(query="Create a detailed analysis of climate change", tenant_id="your_org:production"))
# Comparison queries
compare_routing = await router.process(GatewayInput(query="Compare Python and Java for web development", tenant_id="your_org:production"))
print(f"Video query → {video_routing.routed_to}")
print(f"Report query → {report_routing.routed_to}")
print(f"Compare query → {compare_routing.routed_to}")
Learning Points:
-
GatewayAgent decides which agent handles a query, using GLiNER entity detection (no LLM call)
-
Uses tiered routing: GatewayAgent (GLiNER, fast path) → OrchestratorAgent (DSPy planning, for complex multi-agent workflows)
-
Returns GatewayOutput with routed_to + confidence + reasoning
7.3 Gateway Threshold Optimization (On-Demand)¶
Trigger gateway-threshold optimization via the admin API:
import requests
# Trigger on-demand gateway threshold optimization
response = requests.post(
"http://localhost:8000/admin/tenant/your_org:production/optimize",
json={"mode": "gateway-thresholds"}
)
result = response.json()
print(f"Workflow: {result['workflow_name']}")
print(f"Status URL: {result['status_url']}")
# status_url is a relative path — resolve against the runtime base URL
status = requests.get(f"http://localhost:8000{result['status_url']}").json()
print(f"Phase: {status['phase']}")
Test the CLI function directly:
from cogniverse_runtime.optimization_cli import (
_compute_gateway_thresholds,
GATEWAY_DEFAULT_THRESHOLD,
)
import pandas as pd
# Phoenix stores gateway span attributes as a dict in the
# "attributes.gateway" column (complexity + confidence per span)
spans_df = pd.DataFrame({
"attributes.gateway": [
{"complexity": "simple", "confidence": 0.9},
{"complexity": "complex", "confidence": 0.3},
{"complexity": "simple", "confidence": 0.8},
{"complexity": "complex", "confidence": 0.2},
{"complexity": "simple", "confidence": 0.7},
],
})
result = _compute_gateway_thresholds(spans_df)
print(f"Computed threshold: {result['thresholds']['fast_path_confidence_threshold']:.3f}")
print(f"Default threshold: {GATEWAY_DEFAULT_THRESHOLD}")
Learning Points:
-
Gateway thresholds are computed from span data, not reward signals
-
On-demand optimization via
POST /admin/tenant/{id}/optimizesubmits an Argo Workflow -
GATEWAY_DEFAULT_THRESHOLD = 0.4is the fallback when insufficient span data exists
7.4 Full Agent Roster (23 Agents)¶
cogniverse_agents ships 23 A2A agents (see configs/config.json → agents). Every dispatched agent is reachable through the unified runtime dispatcher at POST /agents/{agent_name}/process (registered by libs/runtime/cogniverse_runtime/routers/agents.py); the 7 knowledge-graph agents additionally have bespoke routes under /admin/tenants/{tenant_id}/knowledge/... (libs/runtime/cogniverse_runtime/routers/knowledge.py). Agents marked disabled below have "enabled": false in configs/config.json and must be enabled before the dispatcher will route to them.
Generation + Routing Agents:
| Agent | Port | Enabled | Purpose |
|---|---|---|---|
gateway_agent | 8000 (mounted) | Yes | GLiNER-based fast-path classifier/router (see 7.2) |
orchestrator_agent | 8013 | Yes | DSPy two-phase planning + multi-agent execution for complex queries |
summarizer_agent | 8004 | Yes | Summarizes search/agent results |
detailed_report_agent | 8005 | Yes | Generates detailed multi-section reports |
profile_selection_agent | 8000 (mounted) | Yes | Picks the best backend profile for a query (modality + complexity + intent) |
query_enhancement_agent | 8000 (mounted) | Yes | Rewrites/expands queries before retrieval |
entity_extraction_agent | 8000 (mounted) | Yes | Extracts entities/relationships for routing and KG ingestion |
Search & Analysis Agents:
| Agent | Port | Enabled | Purpose |
|---|---|---|---|
search_agent | 8002 | Yes | Video/multi-modal search via the backend (see 7.1) |
image_search_agent | 8006 | Yes | Image-specific search |
document_agent | 8008 | Yes | Document search and retrieval |
text_analysis_agent | 8003 | Yes | Text content analysis |
audio_analysis_agent | 8007 | Yes | Audio content analysis/transcription search |
Research + Coding Agents:
| Agent | Port | Enabled | Purpose |
|---|---|---|---|
deep_research_agent | 8009 | Yes | Multi-step web/document research |
coding_agent | 8010 | Yes | Code generation/analysis tasks |
Knowledge-Graph & Reasoning Agents:
| Agent | Port | Enabled | Purpose |
|---|---|---|---|
citation_tracing_agent | 8019 | No | Traces claims back to source citations |
contradiction_reconciliation_agent | 8020 | No | Reconciles contradictory claims in the knowledge graph |
multi_document_synthesis_agent | 8021 | No | Synthesizes findings across multiple documents |
kg_traversal_agent | 8022 | No | Multi-hop knowledge-graph traversal |
temporal_reasoning_agent | 8025 | No | Reasons about time-ordered facts/events |
knowledge_summarization_agent | 8026 | No | Summarizes knowledge-graph subgraphs |
audit_explanation_agent | 8027 | Yes | Explains why a decision/answer was produced |
Multi-Tenant + Federation Agents:
| Agent | Port | Enabled | Purpose |
|---|---|---|---|
cross_tenant_comparison_agent | 8023 | No | Compares one subject's view across tenants in an org (ACL-gated, no LLM) |
federated_query_agent | 8024 | No | Aggregates federated reads across tenants in an org, with optional RLM summarization |
Generic per-agent smoke test (dispatcher path):
# List all registered agents
curl http://localhost:8000/agents/
# Get one agent's card (capabilities, health endpoint)
curl http://localhost:8000/agents/search_agent/card
# Dispatch a request to any agent by name
curl -X POST http://localhost:8000/agents/search_agent/process \
-H "Content-Type: application/json" \
-d '{"query": "test video", "tenant_id": "default"}'
Learning Points:
-
23 agents total: 15 enabled by default, 8 disabled (the KG-reasoning and federation agents, except
audit_explanation_agent) -
gateway_agent,profile_selection_agent,query_enhancement_agent, andentity_extraction_agentare mounted inside the main runtime process (port 8000) rather than run as standalone services -
KG-reasoning agents also expose direct REST routes under
/admin/tenants/{tenant_id}/knowledge/*in addition to the generic dispatcher path
✅ Layer 7 Complete: Individual agents working, routing decisions functional, optimizers trainable
Layer 8: Runtime Layer (cogniverse_runtime)¶
Purpose: FastAPI server (application layer) exposing all functionality via REST API
8.1 Runtime Service Startup¶
Start runtime server:
# Start server (application layer)
JAX_PLATFORM_NAME=cpu uv run python -m cogniverse_runtime.main
# Verify startup in logs
# Expected: "Uvicorn running on http://0.0.0.0:8000"
Check health endpoint:
curl http://localhost:8000/health
# Expected: {"status": "healthy", "service": "cogniverse-runtime",
# "backends": {"registered": N, "backends": [...]},
# "agents": {"registered": N, "agents": [...]}}
# Returns HTTP 503 with {"status": "unhealthy", ...} if the registries
# can't be assembled (e.g. missing BACKEND_URL)
Learning Points:
-
Runtime is in application layer
-
Integrates agents, vespa, synthetic, and evaluation packages
-
Port 8000 (default)
-
Health check validates all services
8.2 Search Endpoints¶
Test search API:
# Unified search endpoint
curl -X POST http://localhost:8000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "machine learning tutorials",
"tenant_id": "default",
"profile": "video_colpali_smol500_mv_frame",
"strategy": "hybrid",
"top_k": 5
}'
# Should return JSON with results
Test streaming search:
curl -X POST http://localhost:8000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "Python programming",
"tenant_id": "default",
"profile": "video_colpali_smol500_mv_frame",
"top_k": 10,
"stream": true
}'
# Returns SSE stream for real-time results
Learning Points:
-
Single unified /search/ endpoint for all search operations
-
Set stream: true for server-sent events (SSE) streaming
-
Results include profile info and result count
8.3 Routing Endpoints¶
Test routing via agent:
# Fast-path routing is handled by GatewayAgent via POST /agents/gateway_agent/process
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput
# GatewayAgent classifies and routes with GLiNER — deps has no required fields
router = GatewayAgent(deps=GatewayDeps())
# Route query (async method)
decision = await router.process(GatewayInput(query="Show me cooking videos", tenant_id="your_org:production"))
print(f"Routing decision:")
print(f" Routed to: {decision.routed_to}")
print(f" Confidence: {decision.confidence}")
print(f" Reasoning: {decision.reasoning}")
Learning Points:
-
Fast-path routing is handled by GatewayAgent via the A2A endpoint
-
Returns GatewayOutput with routed_to, confidence, and reasoning
-
Complex queries are routed to OrchestratorAgent for multi-step planning
8.4 Synthetic Data Endpoints¶
Test synthetic data generation:
# Reuse the live service initialized in Layer 6.1.
from cogniverse_synthetic.schemas import SyntheticDataRequest
# Create request ("routing" is a registered optimizer name)
request = SyntheticDataRequest(
optimizer="routing",
count=10,
vespa_sample_size=50,
strategy="diverse",
max_profiles=2,
tenant_id="your_org:production"
)
# Generate data
response = await service.generate(request)
print(f"Generated {response.count} examples")
print(f"Selected profiles: {response.selected_profiles}")
List available optimizers:
from cogniverse_synthetic import OPTIMIZER_REGISTRY
for name in OPTIMIZER_REGISTRY.keys():
print(f"- {name}")
Learning Points:
-
Synthetic data service used programmatically (implementation layer)
-
OPTIMIZER_REGISTRY lists available optimizer types
-
Integrates with backend for content sampling
8.5 Admin Endpoints¶
cogniverse_runtime.admin.tenant_manager defines the /admin/organizations and /admin/tenants routes. main.py mounts this same router on the unified runtime (port 8000), so these endpoints work there in normal operation. The module can also run standalone (uv run python -m cogniverse_runtime.admin.tenant_manager, default port 9000) for isolated testing without booting the full runtime.
Create organization (via the unified runtime, port 8000):
curl -X POST http://localhost:8000/admin/organizations \
-H "Content-Type: application/json" \
-d '{
"org_id": "acme",
"org_name": "Acme Corporation",
"created_by": "admin"
}'
Create tenant (via the unified runtime, port 8000):
curl -X POST http://localhost:8000/admin/tenants \
-H "Content-Type: application/json" \
-d '{
"tenant_id": "acme:production",
"created_by": "admin"
}'
Learning Points:
-
Tenant management endpoints are mounted on the main runtime (port 8000) and are also runnable standalone via
cogniverse_runtime.admin.tenant_manager(default port 9000) for isolated testing -
Organization → Tenant hierarchy
-
Tenant format: {org}:{name}
✅ Layer 8 Complete: Runtime API functional, all endpoints responding, multi-tenant support working
Layer 9: End-to-End Integration Tests¶
All 12 Packages Working Together
Purpose: Validate complete workflows across all layers
9.1 Complete Search Workflow¶
Test full search pipeline:
# 1. Ingest test video
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir /tmp/test_video \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
--tenant-id e2e_test
# 2. Run comprehensive search test
JAX_PLATFORM_NAME=cpu uv run python tests/comprehensive_video_query_test_v2.py \
--profiles video_colpali_smol500_mv_frame \
--test-multiple-strategies \
--max-queries 5
# 3. Verify Phoenix captured spans
Validate in Phoenix:
from phoenix.client import Client
client = Client(base_url='http://localhost:6006')
spans_df = client.spans.get_spans_dataframe(
project_name="cogniverse-e2e_test-search",
timeout=60,
)
print(f"Search spans collected: {len(spans_df)}")
print(f"Avg latency: {spans_df['latency_ms'].mean():.2f}ms")
Learning Points:
-
Complete flow: Ingest → Search → Telemetry
-
All layers working together
-
End-to-end latency tracking
9.2 Routing + Search Workflow¶
Test routing to search:
# Complete routing + search workflow
from cogniverse_agents.gateway_agent import GatewayAgent, GatewayDeps, GatewayInput
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
# 1. Route the query (async method) — GatewayAgent classifies with GLiNER
router = GatewayAgent(deps=GatewayDeps())
decision = await router.process(GatewayInput(query="Find videos about machine learning", tenant_id="your_org:production"))
print(f"Routing: {decision.routed_to}")
# 2. If search_agent, execute search
if decision.routed_to == "search_agent":
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
agent = SearchAgent(
deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
config_manager=config_manager,
schema_loader=schema_loader,
)
results = agent.search_by_text(query="machine learning", tenant_id="your_org:production", top_k=5)
print(f"Found {len(results)} results")
Learning Points:
-
Routing determines agent
-
Agent executes search
-
Results returned to user
9.3 Optimization Workflow¶
Test complete optimization cycle:
# 1. Generate synthetic data (if needed)
# This is typically done programmatically, see Layer 6 examples above
# 2. Run profile optimization
JAX_PLATFORM_NAME=cpu uv run python -m cogniverse_runtime.optimization_cli \
--mode profile --tenant-id default
# 3. Check results in Phoenix — the optimization run creates a new
# experiment with the optimized module's scores vs baseline
# 4. Verify the compiled artifact was saved to the telemetry provider's
# dataset store (loaded by agents automatically on next restart)
Learning Points:
-
Synthetic data → Training → Optimized model
-
Complete optimization pipeline
-
Results include improvement metrics
9.4 Multi-Tenant Workflow¶
Test tenant isolation:
# 1. Create organization (unified runtime, port 8000 — see 8.5)
curl -X POST http://localhost:8000/admin/organizations \
-H "Content-Type: application/json" \
-d '{"org_id": "test_org", "org_name": "Test Org", "created_by": "admin"}'
# 2. Create tenant (unified runtime, port 8000 — see 8.5)
curl -X POST http://localhost:8000/admin/tenants \
-H "Content-Type: application/json" \
-d '{"tenant_id": "test_org:dev", "org_id": "test_org", "created_by": "admin"}'
# 3. Ingest video for tenant
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir /tmp/test_video \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
--tenant-id test_org:dev
# 4. Search with tenant (unified /search/ endpoint)
curl -X POST http://localhost:8000/search/ \
-H "Content-Type: application/json" \
-d '{
"query": "test",
"tenant_id": "test_org:dev",
"profile": "video_colpali_smol500_mv_frame"
}'
# 5. Verify data isolation
# Should not see test_org:dev data when searching with "default" tenant
Learning Points:
-
Complete tenant isolation
-
Separate schemas per tenant
-
Config/telemetry/memory isolated
9.5 Argo Workflow Integration¶
The scheduled optimization CronWorkflows are Helm templates (charts/cogniverse/templates/optimization-workflows.yaml), not a standalone kubectl apply -f manifest — they're deployed as part of the cogniverse chart release.
Test scheduled optimization:
# 1. Deploy/upgrade the chart (installs the CronWorkflows)
helm upgrade --install cogniverse ./charts/cogniverse -n cogniverse
# 2. Check CronWorkflows created (names are {release}-agent-optimization,
# {release}-daily-cleanup, etc.)
kubectl get cronworkflow -n cogniverse
# 3. Trigger the weekly agent-optimization CronWorkflow manually
argo submit --from cronwf/cogniverse-agent-optimization -n cogniverse
# 4. Monitor workflow
argo list -n cogniverse
argo logs <workflow-name> -n cogniverse --follow
# 5. Check results
argo get <workflow-name> -n cogniverse -o json | \
jq '.status.outputs.parameters'
Learning Points:
-
Argo CronWorkflows for batch jobs, defined as Helm templates under
charts/cogniverse/templates/ -
agentOptimizationruns weekly by default ("0 3 * * 0", Sunday 3 AM UTC, pervalues.yaml);dailyGatewayanddailyCleanuprun daily -
Kubernetes integration via the Argo Workflows CRDs
✅ Layer 9 Complete: End-to-end workflows validated, multi-tenant isolation verified, all integrations working
Final Validation Checklist¶
Infrastructure¶
- Vespa running and healthy
- Phoenix running and collecting spans
- Mem0 initialized and storing memories
- Runtime API responding on all endpoints
Data Layer¶
- Videos ingested across multiple profiles
- Documents searchable via Vespa
- Embeddings generated correctly
- Search results ranked properly
Telemetry¶
- Phoenix projects created per tenant
- Spans collected for search/routing/orchestration
- Span attributes populated correctly
- Web client Analytics view showing metrics
Configuration¶
- System config loaded
- Tenant-specific overrides working
- Backend auto-discovery functional
- Profile management operational
Agents¶
- Individual agents (search, summarizer, detailed_report, and others from the 23-agent roster in 7.4) working
- Gateway agent (fast-path) and orchestrator agent (complex-query planning) making correct routing decisions
- Agent memory integration functional
- Multi-agent orchestration working
Optimization¶
- Synthetic data generation producing quality examples for each registered optimizer (profile, routing, workflow, unified, cross_modal)
- Profile optimizer improving profile-selection accuracy
- Cross-modal optimizer learning fusion patterns
- DSPy-based routing optimizer (SIMBA/MIPROv2/BootstrapFewShot, auto-selected by training-data size) improving routing decisions
- Argo workflows submitting and executing
Multi-Tenant¶
- Tenant isolation validated
- Separate schemas per tenant
- Config/telemetry/memory isolated
- Cross-tenant data leakage prevented
UI/UX¶
- Web client loading all views
- Analytics showing real data
- Config management CRUD working
- Optimization workflows submittable from UI
Troubleshooting Guide¶
Common Issues¶
Vespa Connection Failed:
-
Check:
curl http://localhost:8080/state/v1/health -
Fix: Start Vespa service
-
Verify: Port 8080 accessible
Phoenix Not Collecting Spans:
-
Check: Phoenix server running on port 6006
-
Fix: Set PHOENIX_ENDPOINT env var
-
Verify: Spans appear in Phoenix UI
Memory Manager Initialization Failed:
-
Check:
mem0aidependency installed (pinned asmem0ai==1.0.11inpyproject.toml; already present afteruv sync) -
Fix:
uv sync(oruv pip install mem0ai==1.0.11to add it standalone) -
Verify: Can create
Mem0MemoryManager(tenant_id="your_org:production")
Backend Config Not Found:
-
Check: config.json exists in standard locations
-
Fix: Create config.json with backend section
-
Verify:
create_default_config_manager().get_backend_config(tenant_id=...)succeeds
Optimization Training Failed:
-
Check: Enough training examples exist for the mode (
--mode simba/profile/entity-extractionscale their DSPy teleprompter choice by training-set size; too few examples still runs, but with weaker few-shot demos) -
Fix: Generate more examples first via
--mode synthetic, or the synthetic mode in the web client's Optimization runs view -
Verify: check the CLI's logged
training_examplescount in the run output
Argo Workflow Submission Failed:
-
Check: kubectl or argo CLI configured
-
Fix: Install argo CLI, configure kubectl
-
Verify:
argo versionsucceeds
Learning Outcomes¶
After completing this testing plan, you will understand:
- Architecture: How all layers interact bottom-up
- Data Flow: From ingestion → storage → search → telemetry
- Tenant Isolation: How multi-tenancy works across layers
- Configuration: Auto-discovery and override mechanisms
- Optimization: Complete optimization lifecycle
- Deployment: Argo Workflows for production automation
- Observability: Phoenix telemetry and analytics
- APIs: REST endpoints for all functionality
- UI Integration: Web client for all operations
- Troubleshooting: Common issues and solutions
Next Steps¶
After completing all layers:
- Run full test suite:
JAX_PLATFORM_NAME=cpu timeout 7200 uv run pytest - Deploy to staging: Test with real workloads
- Performance testing: Benchmark search latency, optimization throughput
- Security audit: Verify tenant isolation, auth mechanisms
- Documentation review: Update any gaps discovered during testing
- Production deployment: Deploy with monitoring and alerting
Testing Tips:
-
Document issues as you find them
-
Take notes on architecture insights
-
Save successful commands for future reference
-
Test edge cases (empty queries, large documents, etc.)
-
Validate error handling (network failures, invalid inputs)
Good luck with your comprehensive testing! 🚀