Note: All 23 agents in the roster (configs/config.json → agents.*) run as handlers inside one unified runtime deployment (libs/runtime/cogniverse_runtime/), not as separate per-agent containers — see Full Agent Roster below and Container Resources for the actual deployment topology.
With Visual Reranking tier (federated multi-tenant reads)
federated_query_agent (FederatedQueryAgent)
8024
disabled
With Visual Reranking tier (federated multi-tenant reads)
Source: agent roster and enabled/disabled flags from configs/config.json → agents.*; class names from libs/agents/cogniverse_agents/*.py. "disabled" agents ship code and tests but are not started by default (agents.<name>.enabled: false); their latency tiers are targets for when enabled, mapped to the closest existing tier above rather than measured.
Source: configs/config.json → video_colpali_smol500_mv_frame, video_colqwen_omni_mv_chunk_30s, video_xclip_sv_chunk_6s (schema_config.embedding_dim / num_patches). The chunk-based profile's own schema_config.model_name is ColQwen2; both it and the frame-based ColPali profile load TomoroAI/tomoro-colqwen3-embed-4b weights.
Note: GEPA optimizer is registered as OptimizerType.GEPA in libs/foundation/cogniverse_foundation/config/agent_config.py. Optimizer selection is configured per-tenant; DSPyAgentOptimizerPipeline does not auto-select based on dataset size.
Note: Available optimizers include: BootstrapFewShot, LabeledFewShot, BootstrapFewShotWithRandomSearch, COPRO, MIPROv2 (via libs/core/cogniverse_core/common/dspy_module_registry.py), plus GEPA and SIMBA (configured via libs/foundation/cogniverse_foundation/config/agent_config.py).
Query enhancement via QueryEnhancementModule (A2A agent)
Note: SIMBA optimization runs as an Argo batch job (not inline). Real-time enhancement is handled by QueryEnhancementAgent (libs/agents/cogniverse_agents/query_enhancement_agent.py) using QueryEnhancementModule.
Implementation: Memory operations are implemented in libs/core/cogniverse_core/memory/ with support for multiple backends including Vespa-based storage.
Spans per batch (BatchExportConfig.max_export_batch_size default)
Export Interval
500ms
Batch export frequency (BatchExportConfig.schedule_delay_millis default)
Max Queue Size
2048
In-memory span queue before drops (BatchExportConfig.max_queue_size default)
Span Storage
30 days
Retention target (not currently enforced by a Phoenix-side TTL)
Source: libs/foundation/cogniverse_foundation/telemetry/config.py → BatchExportConfig defaults, applied by PhoenixTelemetryProvider.configure_span_export in libs/telemetry-phoenix/cogniverse_telemetry_phoenix/provider.py.
Source: charts/cogniverse/values.yaml → runtime.resources, runtime.replicaCount/runtime.autoscaling, vespa.resources/vespa.replicaCount, phoenix.resources/phoenix.replicaCount. There is no per-agent container — libs/runtime/cogniverse_runtime/ hosts every agent from the roster above as handlers inside one FastAPI process, and Vespa runs as a single node (no separate container/content-cluster split) in this chart. There is no standalone "memory service" container; memory operations run in-process via libs/core/cogniverse_core/memory/.
Source: charts/cogniverse/templates/hpa.yaml (only the runtime deployment has a HorizontalPodAutoscaler, gated on runtime.autoscaling.enabled) and charts/cogniverse/values.yaml → runtime.autoscaling.{minReplicas,maxReplicas,targetCPUUtilizationPercentage, targetMemoryUtilizationPercentage}. vespa.replicaCount and phoenix.replicaCount are fixed values with no HPA in this chart; scaling either requires a manual replicaCount change and Vespa content redistribution.
Load testing suite: tests/routing/integration/ — covers integration scenarios for routing, connectivity, and feature integration. Production-load throughput and latency percentile tests are not yet implemented.
# Video ingestion - use integration test with timingJAX_PLATFORM_NAME=cpuuvrunpytesttests/ingestion/integration/-v-k"ingestion"--durations=0# Query latency - use search tests with timingJAX_PLATFORM_NAME=cpuuvrunpytesttests/agents/integration/-v-k"search"--durations=0