Cogniverse Study Guide: Configuration Management¶
Packages: cogniverse_sdk, cogniverse_foundation Module Path: libs/foundation/cogniverse_foundation/config/ and libs/sdk/cogniverse_sdk/interfaces/
Module Overview¶
Purpose¶
The configuration system provides centralized management for all system configurations with:
-
Multi-tenant isolation: Complete configuration separation per tenant
-
Versioning: Full history tracking with rollback capability
-
Pluggable backends: Vespa or custom implementations
-
Type-safe schemas: Strongly typed configuration dataclasses
-
Caching: The system config (the hot-path read) is held in memory and refreshed off the reading thread, so another process's write is served within 60s
-
DSPy integration: Dynamic optimizer and module configuration
Key Components¶
- ConfigStore (sdk): Interface for storage backends
- ConfigManager (foundation): Centralized configuration management with caching
- SystemConfig (foundation): System-wide configuration (LLM, backend, telemetry)
- AgentConfig (foundation): Per-agent DSPy module and optimizer configuration
- VespaConfigStore (vespa): Vespa-based ConfigStore implementation
Package Structure¶
libs/foundation/cogniverse_foundation/config/
├── unified_config.py # SystemConfig, LLMEndpointConfig, LLMConfig,
│ # SemanticRouterConfig, RoutingConfigUnified,
│ # AgentConfigUnified, BackendConfig, BackendProfileConfig
├── agent_config.py # AgentConfig, ModuleConfig, OptimizerConfig,
│ # DSPyModuleType, OptimizerType
├── manager.py # ConfigManager (central read/write API + caching)
├── utils.py # ConfigUtils, get_config(), create_default_config_manager(),
│ # get_config_manager_singleton()
├── bootstrap.py # BootstrapConfig.from_environment() — resolves backend
│ # connection info (BACKEND_URL/BACKEND_PORT + config.json
│ # backend.type) before the ConfigStore itself can connect
├── llm_factory.py # create_dspy_lm(LLMEndpointConfig) — the single
│ # chokepoint every dspy.LM() construction goes through
├── api_mixin.py # ConfigAPIMixin — REST endpoints (GET/PUT /config) for
│ # runtime agent config with ConfigManager persistence
└── semantic_router.py # resolve_semantic_router_headers(), apply_semantic_routing(),
# create_routed_lm(), routed_lm_context_for() — opt-in
# LLM traffic routing through a semantic router
libs/sdk/cogniverse_sdk/interfaces/
└── config_store.py # ConfigStore ABC, ConfigScope, ConfigEntry
libs/vespa/cogniverse_vespa/config/
└── config_store.py # VespaConfigStore(ConfigStore) — Vespa-backed implementation
Architecture Diagram¶
flowchart TB
Apps["<span style='color:#000'>Applications<br/>Agents, Routing, Telemetry</span>"]
Apps --> Manager["<span style='color:#000'>ConfigManager<br/>Singleton Access</span>"]
Manager --> Store["<span style='color:#000'>ConfigStore Interface<br/>(cogniverse_sdk)</span>"]
Store --> Vespa["<span style='color:#000'>VespaConfigStore<br/>(Production: HA, Replication)</span>"]
Store --> Custom["<span style='color:#000'>Custom Implementations<br/>Redis, PostgreSQL, etc.</span>"]
style Apps fill:#90caf9,stroke:#1565c0,color:#000
style Manager fill:#ffcc80,stroke:#ef6c00,color:#000
style Store fill:#ce93d8,stroke:#7b1fa2,color:#000
style Vespa fill:#a5d6a7,stroke:#388e3c,color:#000
style Custom fill:#b0bec5,stroke:#546e7a,color:#000 Configuration Scopes¶
1. System Configuration¶
Global infrastructure settings per tenant:
from cogniverse_foundation.config.unified_config import SystemConfig
# SystemConfig is global infrastructure state (one per deployment) — it
# has no tenant_id field. Per-tenant settings live in RoutingConfigUnified,
# AgentConfig, and TelemetryConfig instead (see below).
system_config = SystemConfig(
llm_model="gpt-4",
base_url="https://api.openai.com/v1",
backend_url="http://localhost",
backend_port=8080,
telemetry_url="http://localhost:6006",
)
Settings Include:
-
LLM providers and models
-
Backend URLs and ports
-
Phoenix telemetry endpoints
-
Agent service URLs (via
agentssection) -
Environment settings
1a. LLM Endpoint Configuration¶
LLMEndpointConfig and LLMConfig from cogniverse_foundation.config.unified_config wire the LLM for every agent and optimizer.
from cogniverse_foundation.config.unified_config import LLMEndpointConfig, LLMConfig
# Primary (student) LLM — used at runtime by all agents
primary = LLMEndpointConfig(
model="openai/google/gemma-4-e4b-it",
api_base="http://localhost:11434/v1", # in-cluster vLLM or Ollama
temperature=0.1,
max_tokens=1000,
request_timeout=120.0, # fail-fast timeout in seconds (default: 120.0)
num_retries=1, # total attempts per call (default: 1)
)
# Teacher LLM — used only during DSPy optimization (BootstrapFewShot's teacher_settings)
teacher = LLMEndpointConfig(
model="openai/cyankiwi/Qwen3.6-27B-AWQ-INT4",
api_base="http://localhost:29011/v1",
request_timeout=120.0,
num_retries=1,
)
llm_config = LLMConfig(primary=primary, teacher=teacher)
Key fields:
| Field | Default | Purpose |
|---|---|---|
model | (required) | LiteLLM model string with provider prefix (openai/<model>, anthropic/<model>, etc.) |
api_base | None | Endpoint URL. None = LiteLLM default routing. |
temperature | 0.1 | Sampling temperature. |
max_tokens | 1000 | Max completion tokens. |
context_window | None | Tokens the endpoint accepts per request. Read by create_budgeted_dspy_lm() only when the endpoint publishes no max_model_len. |
request_timeout | 120.0 | Seconds before litellm raises a timeout. Set low to fail fast on a down endpoint. |
num_retries | 1 | Total call attempts (1 = no retries). DSPy default is higher; this constrains it for fast-fail behavior. |
seed | None | vLLM sampling seed for deterministic output in tests. |
Every dspy.LM instance in the codebase is built by create_dspy_lm() from cogniverse_foundation.config.llm_factory — the single chokepoint for LLM instantiation:
from cogniverse_foundation.config.llm_factory import create_dspy_lm
lm = create_dspy_lm(primary) # wires api_base/api_key/temperature/seed/extra_headers onto dspy.LM
When SystemConfig.semantic_router (SemanticRouterConfig) is enabled, cogniverse_foundation.config.semantic_router.routed_lm_context_for() rewrites the endpoint to target the router instead of the model backend and attaches per-tenant authz headers before calling create_dspy_lm():
from cogniverse_foundation.config.semantic_router import routed_lm_context_for
with routed_lm_context_for(config_manager, tenant_id="acme_corp", agent_name="search_agent"):
... # DSPy calls inside this block use the resolved (routed or direct) LM
Inside a coroutine, routed_lm_context_for_async() resolves the same context in a worker thread and returns it unentered, so the tenant-tier config read never runs on the serving event loop:
from cogniverse_foundation.config.semantic_router import routed_lm_context_for_async
routed_lm = await routed_lm_context_for_async(config_manager, tenant_id="acme_corp", agent_name="search_agent")
with routed_lm:
...
Routing is opt-in and disabled by default (SemanticRouterConfig.enabled = False); when disabled, the endpoint's own api_base is used unchanged.
1b. Agent Registry Configuration¶
The agents section in config.json defines agent URLs for A2A discovery via AgentRegistry. There are 23 agents in ConfigLoader.AGENT_CLASSES, matching the 23 entries in config.json's agents section. Schema shape (2 representative entries):
{
"agents": {
"search_agent": {
"url": "http://localhost:8002",
"enabled": true,
"capabilities": ["search", "video_search", "retrieval"],
"modalities": ["VIDEO"]
},
"gateway_agent": {
"url": "http://localhost:8000",
"enabled": true,
"capabilities": ["gateway", "classification"],
"timeout": 10
}
}
}
Full roster (all 23 agents, grouped by function):
| Group | Agent key | Port | Enabled | Capabilities |
|---|---|---|---|---|
| Search & Analysis | search_agent | 8002 | true | search, video_search, retrieval |
| Search & Analysis | text_analysis_agent | 8003 | true | text_analysis, sentiment, classification |
| Search & Analysis | image_search_agent | 8006 | true | image_search, visual_analysis |
| Search & Analysis | audio_analysis_agent | 8007 | true | audio_analysis, transcription |
| Search & Analysis | document_agent | 8008 | true | document_analysis, pdf_processing |
| Generation + Routing | gateway_agent | 8000 | true | gateway, classification |
| Generation + Routing | entity_extraction_agent | 8000 | true | entity_extraction, ner |
| Generation + Routing | query_enhancement_agent | 8000 | true | query_enhancement, expansion |
| Generation + Routing | profile_selection_agent | 8000 | true | profile_selection |
| Generation + Routing | orchestrator_agent | 8013 | true | orchestration, planning, multi_agent_coordination |
| Generation + Routing | summarizer_agent | 8004 | true | summarization, text_generation |
| Generation + Routing | detailed_report_agent | 8005 | true | detailed_report, analysis, text_generation |
| Research + Coding | deep_research_agent | 8009 | true | deep_research, analysis |
| Research + Coding | coding_agent | 8010 | true | coding, code_generation, code_search |
| Knowledge-Graph & Reasoning | citation_tracing_agent | 8019 | false | citation_tracing, audit, provenance_consumer |
| Knowledge-Graph & Reasoning | contradiction_reconciliation_agent | 8020 | false | contradiction_reconciliation, audit |
| Knowledge-Graph & Reasoning | multi_document_synthesis_agent | 8021 | false | multi_document_synthesis, citation_preservation |
| Knowledge-Graph & Reasoning | kg_traversal_agent | 8022 | false | knowledge_graph_traversal, audit |
| Knowledge-Graph & Reasoning | temporal_reasoning_agent | 8025 | false | temporal_reasoning, audit |
| Knowledge-Graph & Reasoning | knowledge_summarization_agent | 8026 | false | knowledge_summarization, audit, federation_promoter |
| Knowledge-Graph & Reasoning | audit_explanation_agent | 8027 | true | audit_explanation, audit, provenance_consumer |
| Multi-tenant + Federation | cross_tenant_comparison_agent | 8023 | false | cross_tenant_comparison, audit, federation_consumer |
| Multi-tenant + Federation | federated_query_agent | 8024 | false | federated_query, audit, federation_consumer |
Knowledge-graph, reasoning, and federation agents ship "enabled": false by default — operators opt in per deployment. audit_explanation_agent is the one exception in that group, enabled by default.
Key Points: - Agent keys must match AGENT_CLASSES in config_loader.py (e.g., orchestrator_agent, search_agent) — all 23 keys above have a matching entry - ConfigLoader.load_agents() validates each agent class is importable and registers metadata (capabilities, URL) in the AgentRegistry - Set "enabled": false to disable an agent without removing its config - In unified runtime mode, all agents share the same runtime URL; per-request tenant_id, profile, and session_id arrive in the task payload - gateway_agent, entity_extraction_agent, query_enhancement_agent, and profile_selection_agent all default to port 8000 because they run in-process inside the unified gateway rather than as standalone services
2. Agent Configuration¶
Per-agent DSPy module and optimizer settings (per tenant):
from cogniverse_foundation.config.agent_config import AgentConfig, ModuleConfig, DSPyModuleType, OptimizerConfig, OptimizerType
agent_config = AgentConfig(
agent_name="video_search_agent",
agent_version="1.0.0",
agent_description="Video search agent with ReAct module",
agent_url="http://localhost:8002",
capabilities=["video_search", "rerank"],
skills=[{"name": "vespa_search"}, {"name": "rerank"}],
module_config=ModuleConfig(
module_type=DSPyModuleType.REACT,
signature="Question -> Answer",
max_retries=5,
temperature=0.7
),
optimizer_config=OptimizerConfig(
optimizer_type=OptimizerType.GEPA,
max_bootstrapped_demos=4,
max_labeled_demos=16,
num_trials=10
),
llm_model="gpt-4",
llm_temperature=0.7
)
3. Routing Configuration¶
Per-tenant routing settings (RoutingConfigUnified), reflecting the two real routing decisions in the codebase — GatewayAgent's simple/complex classification, and OrchestratorAgent's GLiNER-gated query analysis path. There is no pluggable strategy hierarchy or GRPO reinforcement-learning loop; see docs/modules/routing.md for the full routing architecture.
flowchart LR
Query["<span style='color:#000'>Query</span>"] --> Gateway["<span style='color:#000'>GatewayAgent._is_complex<br/>GLiNER + rule-based heuristics</span>"]
Gateway -->|simple| SimpleRoute["<span style='color:#000'>SIMPLE_ROUTE_MAP<br/>static (modality, generation_type) → agent</span>"]
Gateway -->|complex| Orchestrator["<span style='color:#000'>OrchestratorAgent</span>"]
Orchestrator --> Composable["<span style='color:#000'>ComposableQueryAnalysisModule</span>"]
Composable -->|GLiNER confidence ≥ threshold| PathA["<span style='color:#000'>Path A<br/>GLiNER entities + LLM reformulation</span>"]
Composable -->|GLiNER confidence < threshold| PathB["<span style='color:#000'>Path B<br/>Single unified LLM call</span>"]
style Query fill:#90caf9,stroke:#1565c0,color:#000
style Gateway fill:#ffcc80,stroke:#ef6c00,color:#000
style Orchestrator fill:#ffcc80,stroke:#ef6c00,color:#000
style Composable fill:#ffcc80,stroke:#ef6c00,color:#000
style SimpleRoute fill:#a5d6a7,stroke:#388e3c,color:#000
style PathA fill:#a5d6a7,stroke:#388e3c,color:#000
style PathB fill:#a5d6a7,stroke:#388e3c,color:#000 Configuration (RoutingConfigUnified fields):
-
Tier enable flags:
enable_fast_path,enable_slow_path,enable_fallback(all defaultTrue) -
Confidence thresholds:
fast_path_confidence_threshold(0.7),slow_path_confidence_threshold(0.6) -
GLiNER settings:
gliner_model,gliner_threshold,gliner_device,gliner_labels -
DSPy optimizer knobs:
dspy_enabled,dspy_max_bootstrapped_demos,dspy_max_labeled_demos—BootstrapFewShotis the only DSPy teleprompter actually instantiated at runtime today;SIMBA/GEPA/MIPROv2are validOptimizerTypevalues but are not wired into an active optimization loop -
Caching:
enable_caching,cache_ttl_seconds,max_cache_size -
Tenant-isolated routing config and version history (via
ConfigStore)
4. Telemetry Configuration¶
TelemetryConfig (cogniverse_foundation.telemetry.config) is backend-agnostic — it has zero knowledge of Phoenix/LangSmith specifics; provider-specific settings go in provider_config:
-
Per-tenant project naming:
tenant_project_template(default"cogniverse-{tenant_id}") andtenant_service_templateresolve the project/service nameTelemetryManager.get_provider(tenant_id=...)creates or looks up -
LRU-cached per-tenant providers:
max_cached_tenants(default 100),tenant_cache_ttl_seconds(default 3600) -
Span export:
BatchExportConfig.use_sync_export(defaultFalse) — synchronous export in tests for deterministic assertions, async batching in production -
Instrumentation level:
TelemetryConfig.level(DISABLED/BASIC/DETAILED/VERBOSE) gates which components emit spans viashould_instrument_component() -
Provider selection:
provider("phoenix"/"langsmith"/Nonefor auto-detect)
Storage Backends¶
ConfigStore Interface¶
Configuration storage uses the ConfigStore interface from cogniverse_sdk. Common implementations include VespaConfigStore for Vespa backend storage, which stores configs in the config_metadata schema.
from cogniverse_foundation.config.utils import create_default_config_manager
# Uses the configured backend (e.g., VespaConfigStore)
manager = create_default_config_manager()
# Get the global system configuration (not per-tenant)
system_config = manager.get_system_config()
print(f"LLM: {system_config.llm_model}")
print(f"Backend: {system_config.backend_url}:{system_config.backend_port}")
Features:
-
HA with Vespa replication
-
Consistent with application data
-
Versioned configuration with history
-
Tenant isolation via document IDs
Custom ConfigStore¶
Implement the ConfigStore interface for alternative storage backends.
from cogniverse_vespa.config.config_store import VespaConfigStore
from cogniverse_foundation.config.manager import ConfigManager
# Initialize Vespa store with URL and port
vespa_store = VespaConfigStore(
backend_url="http://localhost",
backend_port=8080
)
# Use with ConfigManager (tenant isolation handled via document IDs)
manager = ConfigManager(store=vespa_store)
Schema Deployment:
# Configuration schema is defined in configs/schemas/config_metadata_schema.json
# Deploy via Vespa application package deployment
# Verify schema is accessible
curl http://localhost:8080/ApplicationStatus | jq '.schemas'
Features:
-
Leverages Vespa's HA/replication
-
Unified storage management
-
Real-time configuration sync
-
Scales with Vespa cluster
Custom Backend Implementation¶
Create custom storage backends by implementing the ConfigStore interface from the sdk layer. ConfigStore is an ABC with 11 abstract methods (initialize, set_config, get_config, get_config_history, list_configs, list_all_configs, delete_config, export_configs, import_configs, get_stats, health_check) — the sketch below shows the pattern for two of them; a real implementation must provide all of them before it can be instantiated:
from cogniverse_sdk.interfaces.config_store import ConfigStore, ConfigScope, ConfigEntry
from typing import Dict, Any, Optional, List
from datetime import datetime
import redis
class RedisConfigStore(ConfigStore):
"""Redis-based configuration storage - implements sdk interface"""
def __init__(self, redis_url: str):
self.redis_client = redis.from_url(redis_url)
self.initialize()
def initialize(self) -> None:
"""Setup Redis indexes and structures"""
# Create sorted sets for versioning
# Setup key-value structures
pass
def set_config(
self,
tenant_id: str,
scope: ConfigScope,
service: str,
config_key: str,
config_value: Dict[str, Any],
) -> ConfigEntry:
"""Store configuration with versioning"""
# Generate version number
# Store in Redis with TTL
# Publish update event
pass
def get_config(
self,
tenant_id: str,
scope: ConfigScope,
service: str,
config_key: str,
version: Optional[int] = None,
) -> Optional[ConfigEntry]:
"""Retrieve configuration by version"""
# Get from Redis cache
# Deserialize JSON
# Return ConfigEntry
pass
Multi-Tenant Configuration¶
Tenant Isolation¶
Each tenant has completely isolated configuration:
sequenceDiagram
participant App as Application
participant Manager as ConfigManager
participant Store as ConfigStore
App->>Manager: get_config(tenant_id="tenant_a", config_manager)
Manager->>Store: get_config(tenant_id="tenant_a", scope=ConfigScope.ROUTING, ...)
Store-->>Manager: ConfigEntry for tenant_a
Manager-->>App: ConfigUtils(tenant_a)
App->>Manager: get_config(tenant_id="tenant_b", config_manager)
Manager->>Store: get_config(tenant_id="tenant_b", scope=ConfigScope.ROUTING, ...)
Store-->>Manager: ConfigEntry for tenant_b
Manager-->>App: ConfigUtils(tenant_b)
Note over App,Store: Per-tenant configurations are completely isolated Example:
from cogniverse_foundation.config.utils import create_default_config_manager, get_config
from cogniverse_foundation.config.unified_config import RoutingConfigUnified
manager = create_default_config_manager()
# Configure per-tenant routing for Tenant A (enterprise customer)
# routing_mode is "tiered" for both tenants below — it's the only mode
# implemented end-to-end today; other values are accepted by the schema
# for forward compatibility but produce no behavior change at dispatch time.
tenant_a_routing = RoutingConfigUnified(
tenant_id="acme_corp",
routing_mode="tiered",
fast_path_confidence_threshold=0.7,
cache_ttl_seconds=300,
optimizer_floors={
"profile_selection": {
"min_samples_for_optimization": 20,
"min_unique_queries": 6,
},
"entity_extraction": {
"min_samples_for_optimization": 58,
"min_unique_queries": 15,
},
},
)
manager.set_routing_config(tenant_a_routing, tenant_id="acme_corp")
# Configure per-tenant routing for Tenant B (different customer)
tenant_b_routing = RoutingConfigUnified(
tenant_id="globex_inc",
routing_mode="tiered",
fast_path_confidence_threshold=0.8,
cache_ttl_seconds=600,
)
manager.set_routing_config(tenant_b_routing, tenant_id="globex_inc")
# Per-tenant configurations are completely isolated
config_a = get_config(tenant_id="acme_corp", config_manager=manager)
config_b = get_config(tenant_id="globex_inc", config_manager=manager)
assert config_a["cache_ttl_seconds"] != config_b["cache_ttl_seconds"]
# Each tenant gets isolated Vespa schemas
# acme_corp → video_colpali_mv_frame_acme_corp
# globex_inc → video_colpali_mv_frame_globex_inc
Tenant Configuration Management¶
from cogniverse_foundation.config.unified_config import RoutingConfigUnified
from cogniverse_foundation.config.utils import get_config
# Create per-tenant routing configuration for a new tenant
new_tenant_routing = RoutingConfigUnified(
tenant_id="new_tenant",
routing_mode="tiered",
fast_path_confidence_threshold=0.7,
)
manager.set_routing_config(new_tenant_routing, tenant_id="new_tenant")
# Copy per-tenant config from one tenant to another
source_config = get_config(tenant_id="tenant_a", config_manager=manager)
staging_routing = RoutingConfigUnified(
tenant_id="tenant_a_staging",
routing_mode=source_config.get("routing_mode", "tiered"),
)
manager.set_routing_config(staging_routing, tenant_id="tenant_a_staging")
# Delete tenant configuration (removes all versions)
from cogniverse_sdk.interfaces.config_store import ConfigScope
manager.store.delete_config(
tenant_id="old_tenant",
scope=ConfigScope.ROUTING,
service="gateway_agent",
config_key="routing_config"
)
DSPy Integration¶
Dynamic Module Configuration¶
from cogniverse_foundation.config.agent_config import AgentConfig, ModuleConfig, DSPyModuleType, OptimizerConfig, OptimizerType
from cogniverse_foundation.config.utils import create_default_config_manager
manager = create_default_config_manager()
# Configure Video Search Agent with ReAct (per tenant)
video_agent_config = AgentConfig(
agent_name="video_search_agent",
agent_version="1.0.0",
agent_description="Video search agent with ReAct module",
agent_url="http://localhost:8002",
capabilities=["video_search", "rerank", "summarize"],
skills=[{"name": "vespa_search"}, {"name": "rerank"}],
module_config=ModuleConfig(
module_type=DSPyModuleType.REACT,
signature="Question -> Answer",
max_retries=5,
temperature=0.7
),
optimizer_config=OptimizerConfig(
optimizer_type=OptimizerType.GEPA,
max_bootstrapped_demos=4,
max_labeled_demos=16,
num_trials=10
),
llm_model="gpt-4",
llm_temperature=0.7
)
manager.set_agent_config(
tenant_id="acme_corp",
agent_name="video_search_agent",
agent_config=video_agent_config
)
Configuration Versioning¶
Version Tracking¶
Every configuration change creates a new version:
from cogniverse_sdk.interfaces.config_store import ConfigScope
# Get configuration history (SystemConfig stored under "_system" sentinel tenant)
history = manager.store.get_config_history(
tenant_id="_system",
scope=ConfigScope.SYSTEM,
service="system",
config_key="system_config",
limit=10
)
for entry in history:
print(f"Version {entry.version}:")
print(f" Updated: {entry.updated_at}")
print(f" Changes: {entry.config_value}")
Rollback Capability¶
from cogniverse_sdk.interfaces.config_store import ConfigScope
from cogniverse_foundation.config.unified_config import SystemConfig
# Get current global system config (no tenant_id argument)
current = manager.get_system_config()
print(f"Current LLM: {current.llm_model}")
# Rollback to specific version by retrieving and re-applying old config
# SystemConfig is stored under the "_system" sentinel tenant
old_entry = manager.store.get_config(
tenant_id="_system",
scope=ConfigScope.SYSTEM,
service="system",
config_key="system_config",
version=5 # Specific version to rollback to
)
if old_entry:
# Re-apply old configuration
old_config = SystemConfig.from_dict(old_entry.config_value)
manager.set_system_config(old_config)
# Verify rollback
rolled_back = manager.get_system_config()
print(f"Rolled back LLM: {rolled_back.llm_model}")
Export/Import¶
Backup Configuration¶
import json
from datetime import datetime
# Export all configurations
export_data = manager.store.export_configs(
tenant_id="production",
include_history=True
)
# Save with timestamp
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
with open(f"config_backup_{timestamp}.json", "w") as f:
json.dump(export_data, f, indent=2, default=str)
print(f"Exported {len(export_data['configs'])} configurations")
Restore Configuration¶
# Load backup
with open("config_backup_20250104_120000.json", "r") as f:
backup_data = json.load(f)
# Import to new environment
imported_count = manager.store.import_configs(
tenant_id="staging",
configs=backup_data
)
print(f"Imported {imported_count} configurations")
Schema-scope rows record the schema registry's deployments for the exported tenant's own Vespa schemas, so an export omits them and import_configs refuses a payload that carries one (ValueError, nothing written). The destination tenant's schemas are deployed through its profiles.
Monitoring and Health¶
Configuration Health Checks¶
# Check storage backend health
if manager.store.health_check():
print("✓ Configuration storage healthy")
else:
print("✗ Configuration storage unavailable")
# Get storage statistics
stats = manager.store.get_stats()
print(f"Total configurations: {stats['total_configs']}")
print(f"Total tenants: {stats['total_tenants']}")
print(f"Total versions: {stats['total_versions']}")
print(f"Configs by scope: {stats['configs_per_scope']}")
print(f"Storage backend: {stats['storage_backend']}")
Best Practices¶
1. Understand System vs Tenant Config¶
# ✅ Good: SystemConfig is global — call with no arguments
system_config = manager.get_system_config()
print(f"Backend: {system_config.backend_url}:{system_config.backend_port}")
# ✅ Good: Per-tenant config uses get_config() with explicit tenant_id
from cogniverse_foundation.config.utils import get_config
tenant_config = get_config(tenant_id="acme_corp", config_manager=manager)
# ❌ Bad: get_system_config does not accept a tenant_id argument
# config = manager.get_system_config(tenant_id="production") # WRONG
2. Version Critical Changes¶
# Before major changes, export current configuration
backup = manager.store.export_configs(
tenant_id="_system",
include_history=True
)
# Save backup to file
import json
from datetime import datetime
timestamp = datetime.now().strftime("%Y%m%d_%H%M%S")
with open(f"config_backup_{timestamp}.json", "w") as f:
json.dump(backup, f, indent=2, default=str)
# Make changes (a new version is created automatically on write)
from cogniverse_foundation.config.unified_config import SystemConfig
current = manager.get_system_config()
updated = SystemConfig.from_dict({**current.to_dict(), "llm_model": "gpt-4-turbo"})
manager.set_system_config(updated)
3. Use Type-Safe Configurations¶
# ✅ Good: Type-safe dataclass (SystemConfig has no tenant_id field)
from cogniverse_foundation.config.unified_config import SystemConfig
config = SystemConfig(
llm_model="gpt-4",
backend_url="http://localhost",
backend_port=8080
)
# ❌ Bad: Raw dictionaries
config = {"llm_model": "gpt-4"} # No validation
4. Use Configuration Templates¶
# Define reusable templates for different deployment environments
# (SystemConfig is global — no tenant_id)
TEMPLATES = {
"development": SystemConfig(
llm_model="gpt-3.5-turbo",
base_url="http://localhost:11434",
backend_url="http://localhost",
backend_port=8080,
telemetry_url="http://localhost:6006"
),
"production": SystemConfig(
llm_model="gpt-4",
base_url="https://api.openai.com/v1",
backend_url="http://production-vespa",
backend_port=8080,
telemetry_url="http://production-phoenix:6006"
)
}
# Apply a template to set the global system config
template = TEMPLATES["production"]
new_config = SystemConfig(
llm_model="claude-3-opus", # Override
base_url=template.base_url,
backend_url=template.backend_url,
backend_port=template.backend_port,
telemetry_url=template.telemetry_url
)
manager.set_system_config(new_config)
Troubleshooting¶
Configuration Not Found¶
from cogniverse_sdk.interfaces.config_store import ConfigScope
from cogniverse_foundation.config.unified_config import SystemConfig
# Check if global system configuration exists
configs = manager.store.list_configs(
tenant_id="_system",
scope=ConfigScope.SYSTEM
)
print(f"Available configs: {configs}")
# Initialize missing system configuration (no tenant_id — SystemConfig is global)
existing_config = manager.store.get_config(
tenant_id="_system",
scope=ConfigScope.SYSTEM,
service="system",
config_key="system_config"
)
if not existing_config:
manager.set_system_config(SystemConfig())
Concurrent Updates¶
from cogniverse_foundation.config.unified_config import SystemConfig
# Handle concurrent updates by checking version history
current_entry = manager.store.get_config(
tenant_id="_system",
scope=ConfigScope.SYSTEM,
service="system",
config_key="system_config"
)
if current_entry is None:
raise RuntimeError("system_config not initialized — call set_system_config first")
# Make your changes
updated_config = SystemConfig.from_dict(current_entry.config_value)
updated_config.llm_model = "gpt-4-turbo"
# Apply update (creates new version automatically)
manager.set_system_config(updated_config)
# Verify version incremented
new_entry = manager.store.get_config(
tenant_id="_system",
scope=ConfigScope.SYSTEM,
service="system",
config_key="system_config"
)
assert new_entry.version == current_entry.version + 1
Storage Backend Issues¶
# create_default_config_manager() reads BACKEND_URL/BACKEND_PORT (env) and
# backend.type (configs/config.json) via BootstrapConfig — it raises
# ValueError if BACKEND_URL is unset or backend.type is anything other
# than "vespa" (the only supported backend). There is no silent fallback.
from cogniverse_foundation.config.utils import create_default_config_manager
manager = create_default_config_manager()
Testing¶
Unit Tests¶
# Run all configuration tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/common/unit/test_agent_config.py -v
JAX_PLATFORM_NAME=cpu uv run pytest tests/common/unit/test_config_api_mixin.py -v
# Test Vespa backend (real Vespa instance)
JAX_PLATFORM_NAME=cpu uv run pytest tests/backends/integration/test_config_store.py -v
JAX_PLATFORM_NAME=cpu uv run pytest tests/backends/unit/test_config_store_yql_escape.py -v
Integration Tests¶
# Test with real backends
cogniverse up # Starts all services including Vespa
# Run integration tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/common/integration/ -v
Related Guides:
-
../architecture/sdk-architecture.md - SDK structure
-
../architecture/multi-tenant.md - Multi-tenant architecture
-
../modules/common.md - Common utilities
-
setup-installation.md - Installation
-
deployment.md - Deployment
Next: deployment.md