Foundation Module¶
Package: cogniverse_foundation Location: libs/foundation/cogniverse_foundation/
Table of Contents¶
- Overview
- Package Structure
- Configuration System
- ConfigManager
- Configuration Types
- Configuration Scopes
- Configuration Inheritance
- ConfigAPIMixin
- Telemetry System
- TelemetryManager
- Span Context
- Session Tracking
- Project Registration
- Span Context Helpers
- Registry System
- Tenant-Scoped Caching
- DSPy Extensions
- Confidence Parsing
- Usage Examples
- Architecture Position
- Testing
Overview¶
The Foundation module provides infrastructure services that all other modules depend on:
- Configuration Management: Multi-tenant, versioned configuration with pluggable
ConfigStorepersistence (default: Vespa) - LLM Wiring:
create_dspy_lmfactory, opt-in semantic-router request routing, and DSPy adapter/model-name helpers - Telemetry Infrastructure: OpenTelemetry-based tracing with tenant isolation
- Provider Abstraction: Pluggable backends for telemetry (Phoenix, etc.) via a generic entry-point plugin registry
- Shared Utilities: Tenant-scoped LRU caching and robust LM-output confidence parsing used across core, agents, evaluation, and finetuning
All configuration and telemetry operations are tenant-aware - tenant_id is required for all operations.
Package Structure¶
flowchart TB
subgraph Foundation["<span style='color:#000'><b>cogniverse_foundation/</b></span>"]
ConfigDir["<span style='color:#000'><b>config/</b><br/>Configuration system</span>"]
TelemetryDir["<span style='color:#000'><b>telemetry/</b><br/>Telemetry system</span>"]
RegistryDir["<span style='color:#000'><b>registry/</b><br/>Generic entry-point plugin registry</span>"]
CachingDir["<span style='color:#000'><b>caching/</b><br/>Tenant-scoped LRU and refreshing caches</span>"]
DspyDir["<span style='color:#000'><b>dspy/</b><br/>DSPy adapters & model-format helpers</span>"]
CommonDir["<span style='color:#000'><b>common/</b><br/>Tenant identity, DSPy registry & Argo client helpers</span>"]
ConfidencePy["<span style='color:#000'>confidence.py<br/>parse_confidence()</span>"]
InferenceSpecs["<span style='color:#000'>inference_specs.py<br/>InferenceServiceSpec, get_inference_service_spec()</span>"]
InitPy["<span style='color:#000'>__init__.py</span>"]
end
subgraph ConfigFiles["<span style='color:#000'><b>config/ files</b></span>"]
Manager["<span style='color:#000'>manager.py<br/>ConfigManager - central API</span>"]
UnifiedConfigFile["<span style='color:#000'>unified_config.py<br/>SystemConfig, BackendConfig, LLMConfig</span>"]
AgentConfigFile["<span style='color:#000'>agent_config.py<br/>AgentConfig, DSPy settings</span>"]
LlmFactory["<span style='color:#000'>llm_factory.py<br/>create_dspy_lm()</span>"]
InferenceAuth["<span style='color:#000'>inference_auth.py<br/>inference_headers</span>"]
SemanticRouterFile["<span style='color:#000'>semantic_router.py<br/>apply_semantic_routing, create_routed_lm</span>"]
LmResponseCache["<span style='color:#000'>lm_response_cache.py<br/>TenantScopedLMCache</span>"]
LmEndpointAvailability["<span style='color:#000'>lm_endpoint_availability.py<br/>LMEndpointAvailability, LMEndpointNotServing</span>"]
Utils["<span style='color:#000'>utils.py<br/>ConfigUtils, create_default_config_manager</span>"]
ApiMixin["<span style='color:#000'>api_mixin.py<br/>ConfigAPIMixin (FastAPI endpoints)</span>"]
Bootstrap["<span style='color:#000'>bootstrap.py<br/>BootstrapConfig</span>"]
VespaStore["<span style='color:#000'>VespaConfigStore<br/>(in cogniverse_vespa)</span>"]
end
subgraph TelemetryFiles["<span style='color:#000'><b>telemetry/ files</b></span>"]
TelManager["<span style='color:#000'>manager.py<br/>TelemetryManager - central API</span>"]
TelConfig["<span style='color:#000'>config.py<br/>TelemetryConfig, TelemetryLevel</span>"]
Registry["<span style='color:#000'>registry.py<br/>TelemetryRegistry</span>"]
Context["<span style='color:#000'>context.py<br/>search_span, encode_span helpers</span>"]
ProvidersDir["<span style='color:#000'><b>providers/</b><br/>base.py - TraceStore, AnnotationStore, etc.</span>"]
end
subgraph CachingFiles["<span style='color:#000'><b>caching/ files</b></span>"]
TenantLru["<span style='color:#000'>tenant_lru.py<br/>TenantLRUCache</span>"]
RefreshingCacheFile["<span style='color:#000'>refreshing_cache.py<br/>RefreshingCache</span>"]
end
subgraph DspyFiles["<span style='color:#000'><b>dspy/ files</b></span>"]
LenientAdapter["<span style='color:#000'>lenient_json_adapter.py<br/>LenientJSONAdapter</span>"]
StructuredAdapter["<span style='color:#000'>structured_json_adapter.py<br/>StructuredJSONAdapter, signature_response_format</span>"]
ModelFormat["<span style='color:#000'>model_format.py<br/>bare_model_name, ensure_provider_prefix</span>"]
end
subgraph CommonFiles["<span style='color:#000'><b>common/ files</b></span>"]
TenantUtilsFile["<span style='color:#000'>tenant_utils.py<br/>SYSTEM_TENANT_ID, require_tenant_id, etc.</span>"]
DspyModuleRegistryFile["<span style='color:#000'>dspy_module_registry.py<br/>DSPyModuleRegistry, DSPyOptimizerRegistry</span>"]
ArgoClientFile["<span style='color:#000'>argo_client.py<br/>build_argo_async_client</span>"]
end
subgraph SeparatePackages["<span style='color:#000'><b>Separate Packages</b></span>"]
SDKInterface["<span style='color:#000'><b>cogniverse_sdk/interfaces/</b><br/>config_store.py<br/>ConfigStore ABC, ConfigScope</span>"]
VespaImpl["<span style='color:#000'><b>cogniverse_vespa/config/</b><br/>config_store.py<br/>VespaConfigStore implementation</span>"]
end
Foundation --> ConfigDir
Foundation --> TelemetryDir
Foundation --> RegistryDir
Foundation --> CachingDir
Foundation --> DspyDir
Foundation --> CommonDir
Foundation --> ConfidencePy
Foundation --> InitPy
ConfigDir --> ConfigFiles
TelemetryDir --> TelemetryFiles
CachingDir --> CachingFiles
DspyDir --> DspyFiles
CommonDir --> CommonFiles
RegistryDir -.-> Registry
style Foundation fill:#a5d6a7,stroke:#388e3c,color:#000
style ConfigDir fill:#ffcc80,stroke:#ef6c00,color:#000
style TelemetryDir fill:#ce93d8,stroke:#7b1fa2,color:#000
style RegistryDir fill:#b0bec5,stroke:#546e7a,color:#000
style CachingDir fill:#81d4fa,stroke:#0288d1,color:#000
style DspyDir fill:#81d4fa,stroke:#0288d1,color:#000
style CommonDir fill:#81d4fa,stroke:#0288d1,color:#000
style ConfidencePy fill:#90caf9,stroke:#1565c0,color:#000
style InitPy fill:#90caf9,stroke:#1565c0,color:#000
style ConfigFiles fill:#ffcc80,stroke:#ef6c00,color:#000
style TelemetryFiles fill:#ce93d8,stroke:#7b1fa2,color:#000
style CachingFiles fill:#81d4fa,stroke:#0288d1,color:#000
style DspyFiles fill:#81d4fa,stroke:#0288d1,color:#000
style CommonFiles fill:#81d4fa,stroke:#0288d1,color:#000
style SeparatePackages fill:#90caf9,stroke:#1565c0,color:#000
style Manager fill:#ffb74d,stroke:#ef6c00,color:#000
style UnifiedConfigFile fill:#ffb74d,stroke:#ef6c00,color:#000
style AgentConfigFile fill:#ffb74d,stroke:#ef6c00,color:#000
style LlmFactory fill:#ffb74d,stroke:#ef6c00,color:#000
style InferenceAuth fill:#ffb74d,stroke:#ef6c00,color:#000
style ArgoClientFile fill:#ffb74d,stroke:#ef6c00,color:#000
style SemanticRouterFile fill:#ffb74d,stroke:#ef6c00,color:#000
style Utils fill:#ffb74d,stroke:#ef6c00,color:#000
style ApiMixin fill:#ffb74d,stroke:#ef6c00,color:#000
style Bootstrap fill:#ffb74d,stroke:#ef6c00,color:#000
style VespaStore fill:#ffb74d,stroke:#ef6c00,color:#000
style TelManager fill:#ba68c8,stroke:#7b1fa2,color:#000
style TelConfig fill:#ba68c8,stroke:#7b1fa2,color:#000
style Registry fill:#ba68c8,stroke:#7b1fa2,color:#000
style Context fill:#ba68c8,stroke:#7b1fa2,color:#000
style ProvidersDir fill:#ba68c8,stroke:#7b1fa2,color:#000
style TenantLru fill:#64b5f6,stroke:#1565c0,color:#000
style LenientAdapter fill:#64b5f6,stroke:#1565c0,color:#000
style ModelFormat fill:#64b5f6,stroke:#1565c0,color:#000
style TenantUtilsFile fill:#64b5f6,stroke:#1565c0,color:#000
style DspyModuleRegistryFile fill:#64b5f6,stroke:#1565c0,color:#000
style SDKInterface fill:#64b5f6,stroke:#1565c0,color:#000
style VespaImpl fill:#64b5f6,stroke:#1565c0,color:#000 Configuration System¶
ConfigManager¶
ConfigManager is the central configuration API. All configuration operations go through this class.
Key Features:
- Multi-tenant configuration with tenant isolation
- Version history tracking for all configurations
- In-process caching: the system config is held in a
RefreshingCachewithsystem_config_refresh_s(default 5s) andsystem_config_max_staleness_s(default 60s), on the same terms as the scoped configs below:set_system_configholds what it wrote at once, another process's write (the runtime storing its deployment overrides) is served within 60s, and a store outage past that bound raises. Pinned inference URLs are laid over every value served. Per-tenant scoped configs (routing, telemetry, backend, agent, durable execution, tenant instructions) are held in aRefreshingCache(see Tenant-Scoped Caching below). A value older thanscoped_config_refresh_s(default 5s) is still served while one background thread re-reads it, so a request never waits on that store read. A value that reachesscoped_config_max_staleness_s(default 60s) is never served: the caller reads the store itself, and a failed read raises rather than answering with defaults. A failed background refresh is logged and leaves the last value in place up to that bound. Setters on the same manager invalidate immediately. A write made by another process is served within 60s, and within about 5s plus one store read for a key read continuously. Concurrent reads of one key share one store read, and a completed setter cannot be overwritten by an older in-flight read. Profile add, update and delete read the stored backend config, not the held copy, so they never write back over another process's profile change. ConfigUtilsparses a discoveredconfig.jsononce per file modification across concurrent callers and returns an isolated copy to each instance. Invalid JSON and file-access failures propagate with the file path instead of being treated as an empty configuration.- Tenant backend profiles inherit the shipped catalog and system-tenant stored profiles absent from that catalog. Non-empty tenant fields override the base; nested fields merge and the
vlm_endpointfield remains system-owned. This merge also applies when no JSON backend section exists. Store read failures propagate, and callers receive isolated profile values.ConfigUtils(tenant_id, config_manager).backend_profiles()returns that merged catalog, reading only the backend configs. - Pluggable backend persistence via
ConfigStoreinterface (VespaConfigStore)
from cogniverse_foundation.config.utils import create_default_config_manager
# Initialize config manager
config_manager = create_default_config_manager()
# Get global system configuration (no tenant_id argument — SystemConfig is deployment-wide)
system_config = config_manager.get_system_config()
# Set agent configuration (agent_config built as shown under Configuration Types → AgentConfig)
config_manager.set_agent_config(
tenant_id="acme",
agent_name="orchestrator_agent",
agent_config=agent_config
)
Inference endpoint credentials¶
cogniverse_foundation.config.inference_auth.inference_headers(base_url) returns immutable headers for a canonical HTTP(S) root URL. *.modal.run endpoints require HTTPS and COGNIVERSE_INFERENCE_API_KEY; other URLs get no headers. endpoint_root(url) reduces any endpoint URL (.../v1, .../v1/chat/completions) to that root.
The runtime API and ingestion worker parse INFERENCE_SERVICE_URLS with cogniverse_runtime.inference_services.parse_inference_service_urls. The runtime API persists the parsed endpoints into SystemConfig; the ingestion worker pins them on its ConfigManager with pin_inference_service_urls, so its reads carry them without writing to the store.
API Reference:
| Method | Description |
|---|---|
get_system_config() | Get global system configuration (deployment-wide, not per-tenant) |
set_system_config(system_config) | Set global system configuration |
pin_inference_service_urls(service_urls) | Serve service_urls as SystemConfig.inference_service_urls on this manager's reads, without persisting them |
get_agent_config(tenant_id, agent_name) | Get agent configuration |
set_agent_config(tenant_id, agent_name, agent_config) | Set agent configuration |
get_agent_config_history(tenant_id, agent_name, limit=10) | Get version history after canonicalizing the tenant identifier |
get_routing_config(tenant_id="your_org:production", service="gateway_agent") | Get routing configuration |
set_routing_config(routing_config, tenant_id=None, service="gateway_agent") | Set routing configuration |
get_durable_execution_config(tenant_id, service="optimization") | Get durable-execution enablement (default off) |
set_durable_execution_config(durable_config, tenant_id=None, service="optimization") | Set durable-execution enablement |
get_telemetry_config(tenant_id="your_org:production", service="telemetry") | Get telemetry configuration |
set_telemetry_config(telemetry_config, tenant_id=None, service="telemetry") | Set telemetry configuration |
get_backend_config(tenant_id="your_org:production", service="backend") | Get backend configuration |
get_stored_backend_config(tenant_id, service="backend") | Backend configuration as the store holds it now, past the held copy another process's write may not have reached yet; for writes that decide on what the tenant has stored |
set_backend_config(backend_config, tenant_id=None, service="backend") | Set backend configuration |
get_tenant_instructions_config(tenant_id) | Get raw tenant instructions value (TTL-cached; {"text": ..., "updated_at": ...} or None) |
get_backend_profile(profile_name, tenant_id="your_org:production", service="backend") | Get specific backend profile |
add_backend_profile(profile, tenant_id="your_org:production", service="backend", *, replace=True) | Add/update a backend profile; returns BackendProfileWrite(profile, version), the version of the tenant's backend config this write produced (for an identical re-add that writes nothing, the version already holding it). replace=False raises BackendProfileExistsError when the name is already stored, checked in the same compare-and-set as the write |
update_backend_profile(profile_name, overrides, base_tenant_id=SYSTEM_TENANT_ID, target_tenant_id=None, service="backend") | Partial profile update; inherits from base_tenant_id, saves to target_tenant_id, and returns BackendProfileWrite with the merged profile and the target config version the update produced. Raises BackendProfileNotFoundError when the profile is not stored for base_tenant_id, checked in the same compare-and-set as the write |
list_backend_profiles(tenant_id="your_org:production", service="backend") | List all backend profiles |
delete_backend_profile(profile_name, tenant_id="your_org:production", service="backend") | Delete backend profile; False when none was stored |
The three profile writes rewrite the tenant's whole backend config, and every runtime process and replica writes it. Each is a compare-and-set read-modify-write through ConfigStore.update_config: the change is applied to the config as stored, and re-applied to the newer one whenever another writer lands first, so no process's profile change is overwritten. A change that leaves the stored config as it was writes no new version (startup's reaffirm_system_profiles on every worker writes once). Every profile write, one that wrote nothing included, drops the manager's held copy of that tenant's backend config, so its next read serves what it found stored. A write that loses every attempt raises ConfigWriteConflictError, and a store failure raises; either way nothing is written.
Searches read profile writes from the store: the shared search backend resolves the querying tenant's profiles per request through get_config, so a write is searchable on the writing process at once and on every other process and replica within scoped_config_max_staleness_s (60 s by default, about scoped_config_refresh_s for a tenant searched continuously). Nothing pushes profile changes into backend instances. | get_config_value(tenant_id, scope, service, config_key, default=None) | Get arbitrary config value by scope | | set_config_value(tenant_id, scope, service, config_key, config_value) | Set arbitrary config value by scope | | get_all_configs(tenant_id, scope=None) | Get all configs for a tenant, optionally filtered by scope | | export_configs(tenant_id, output_path) | Export all configs to JSON with a timezone-aware UTC exported_at | | get_stats() | Get configuration statistics |
Configuration Types¶
SystemConfig - Infrastructure settings:
from cogniverse_foundation.config.unified_config import SystemConfig
system_config = SystemConfig(
search_backend="vespa",
backend_url="http://localhost",
backend_port=8080
)
AgentConfig - Agent-specific settings:
from cogniverse_foundation.config.agent_config import (
AgentConfig, ModuleConfig, OptimizerConfig,
DSPyModuleType, OptimizerType
)
agent_config = AgentConfig(
agent_name="orchestrator_agent",
agent_version="1.0.0",
agent_description="Routes queries to appropriate agents",
agent_url="http://localhost:8001",
capabilities=["routing", "query_analysis", "entity_extraction", "conversation_memory"],
skills=[{"name": "route_query", "description": "Route to best agent"}],
module_config=ModuleConfig(
module_type=DSPyModuleType.CHAIN_OF_THOUGHT,
signature="query -> routing_decision"
),
optimizer_config=OptimizerConfig(
optimizer_type=OptimizerType.MIPRO_V2,
num_trials=50
),
llm_model="gpt-4",
llm_temperature=0.7,
llm_max_tokens=2000
)
BackendConfig - Backend and profile settings:
from cogniverse_foundation.config.unified_config import BackendConfig, BackendProfileConfig
backend_config = BackendConfig(
tenant_id="acme",
backend_type="vespa",
url="http://localhost",
port=8080,
profiles={
"video_colpali_mv_frame": BackendProfileConfig(
profile_name="video_colpali_mv_frame",
type="video",
embedding_model="colpali",
schema_name="video_colpali_smol500_mv_frame",
pipeline_config={"chunk_strategy": "frame"},
strategies={"default_top_k": 10}
)
}
)
model_loader names the loader ingestion embeds with and process_type (one of PROCESS_TYPES: direct_video, frame_based, video_chunks) the processing type ingestion would otherwise infer. extra_config holds every other profile key (inference_services, model_config, result_granularity, semantic_model, ...); to_dict() writes it beside the named fields and from_dict() collects any unnamed key into it.
Config sections (config/sections.py) - the configs an operator edits as forms:
from cogniverse_foundation.config.sections import CONFIG_SECTIONS, ConfigValueError
routing = CONFIG_SECTIONS["routing"]
schema = routing.schema() # the form's JSON schema
current = routing.default("acme:prod", "gateway_agent")
edited = routing.from_form({"routing_mode": "direct"}, current)
stored = routing.dump(edited, "acme:prod") # what ConfigManager stores
Each ConfigSection is a config dataclass (model) and where it is stored (scope, config_key, a fixed service or one per entry, tenant or system), with load/dump between the stored dict and the dataclass. schema() is the dataclass's JSON schema without the fixed_fields the location sets, with secret_fields marked writeOnly; form_value() dumps a config for the form with secrets null. from_form(value, current) applies the submitted fields to current: a field left out keeps its value, a null secret keeps it and "" clears it, and a key the schema does not know or a value the dataclass refuses raises ConfigValueError naming each, as does a changed value outside a field's choices (search_backend, environment, routing_mode, the telemetry provider from the registered providers) or ranges (backend_port, 1 to 65535); schema() carries them as enum, minimum and maximum, and a stored value outside them is kept when a save leaves it unchanged. Sections: system (SystemConfig), routing (RoutingConfigUnified), telemetry (TelemetryConfig), agent (AgentConfig, one per agent) and durable_execution (DurableExecutionConfig); section_for(scope, config_key) finds the section of a stored entry.
ConfigManager.compare_and_set_entry(tenant_id, scope, service, config_key, value, expected_version=...) stores a value as the next version when the stored one is expected_version (0: none) and returns None when another write landed first; forget_held_configs(tenant_id) drops what the manager holds for a tenant (the system config for _system). The module function forget_held_backend_configs(tenant_id) drops the tenant's backend config from every ConfigManager in the process; the runtime runs it on every runtime and ingestion worker when a profile is written. forget_held_tenant_configs(tenant_id) drops everything every ConfigManager in the process holds for a tenant (the system config for _system); the runtime runs it on every runtime and ingestion worker when a config is saved, restored or imported. SystemConfig.to_dict() shows llm_api_key as "***"; to_dict(redact=False) is the stored form.
Profile servability - whether a tenant can be served a profile:
from cogniverse_foundation.config.unified_config import (
PROFILE_EMBEDDING_SERVICE_UNCONFIGURED,
PROFILE_SCHEMA_NOT_DEPLOYED,
PROFILE_SERVABLE,
profile_base_schema_name,
profile_is_servable,
profile_servability,
)
state = profile_servability(
"document_text_semantic",
profile.to_dict(),
system_config.inference_service_urls,
deployed_schemas={"document_text"},
)
assert state == PROFILE_SERVABLE
profile_servability names why a profile can or cannot be served: its embedding service must resolve to a URL in inference_service_urls (PROFILE_EMBEDDING_SERVICE_UNCONFIGURED otherwise) and its base schema — profile_base_schema_name, the profile's schema_name or its own name — must be in deployed_schemas (PROFILE_SCHEMA_NOT_DEPLOYED otherwise). profile_is_servable is the boolean form. The deployed set is per tenant and comes from cogniverse_core.registries.schema_registry.tenant_deployed_schema_names; callers in cogniverse_agents.profile_selection_agent compose the two, so this module keeps no dependency on core.
tenant_profile_servability(config_manager, tenant_id) there lists the tenant's profiles with their state: the profiles it stored and, for each schema it has deployed that none of those reads, the catalog's search profiles (video, document, image, audio, code, wiki) declaring that schema. A registered tenant stores no profile but has the built-in video_colpali_smol500_mv_frame schema deployed, so that profile is servable for it, alongside any profile it adds. Each row carries the catalog definition from ConfigUtils.backend_profiles(). servable_tenant_profiles and tenant_usable_profile_names (which raises when none is servable) filter it to the servable rows; GET /search/profiles, answer grounding, profile selection, synthetic profile data and the optimizer read the tenant's profiles through them.
FieldMappingConfig - Canonical content fields used by synthetic generation:
from cogniverse_foundation.config.unified_config import FieldMappingConfig
field_mappings = FieldMappingConfig()
assert field_mappings.topic_fields == [
"video_title",
"audio_title",
"image_title",
"document_title",
"chunk_name",
"title",
]
assert field_mappings.description_fields == [
"segment_description",
"image_description",
"full_text",
"source_code",
"content",
"description",
]
assert field_mappings.transcript_fields == ["audio_transcript", "transcript"]
assert field_mappings.entity_fields == ["video_title", "segment_description"]
assert field_mappings.temporal_fields == {
"start": "start_time",
"end": "end_time",
}
assert field_mappings.metadata_fields == {}
Override these mappings only when a deployed schema uses different canonical fields. The defaults preserve titles, descriptive text, transcripts, entity sources, temporal bounds, and metadata for the configured video, audio, image, document, code, and wiki schemas. FieldMappingConfig.from_dict({}) hydrates these same defaults; unknown keys raise instead of being ignored.
SyntheticGeneratorConfig - Synthetic generation timeout and floor:
from cogniverse_foundation.config.unified_config import SyntheticGeneratorConfig
generator_config = SyntheticGeneratorConfig(
tenant_id="acme",
synthetic_generation_timeout_seconds=300.0,
synthetic_generation_floor_count=1,
)
The timeout bounds production callback work. The floor bounds how far a synthetic run may fall back when the candidate pool is exhausted. Both fields validate as positive values and are required in the deployable synthetic config.
RoutingConfigUnified - Routing agent settings:
from cogniverse_foundation.config.unified_config import RoutingConfigUnified
routing_config = RoutingConfigUnified(
tenant_id="acme",
routing_mode="tiered", # "tiered", "ensemble", "hybrid"
enable_fast_path=True,
fast_path_confidence_threshold=0.4,
)
# Per-optimizer floors can override the global 100/3 defaults when a tenant
# needs calibrated promotion thresholds for specific optimization modes.
# routing_config.optimizer_floors = {
# "profile_selection": {
# "min_samples_for_optimization": 20,
# "min_unique_queries": 6,
# },
# "entity_extraction": {
# "min_samples_for_optimization": 58,
# "min_unique_queries": 15,
# },
# }
DurableExecutionConfig - Durable-execution enablement for long-running optimization/eval workflows:
from cogniverse_foundation.config.unified_config import DurableExecutionConfig
durable_config = DurableExecutionConfig(
tenant_id="acme",
enabled=True, # checkpoint + resume the triggered-optimization job (default False)
)
TelemetryConfig - Telemetry settings:
from cogniverse_foundation.telemetry.config import TelemetryConfig, TelemetryLevel
telemetry_config = TelemetryConfig(
enabled=True,
level=TelemetryLevel.DETAILED,
otlp_enabled=True,
otlp_endpoint="localhost:4317"
)
LLMEndpointConfig - Single LLM endpoint wiring:
from cogniverse_foundation.config.unified_config import LLMEndpointConfig
endpoint = LLMEndpointConfig(
model="openai/gpt-4o", # always provider-prefixed (DSPy/LiteLLM convention)
api_base="http://localhost:8101/v1",
api_key=None, # None for local OAI-compat servers
temperature=0.1,
max_tokens=1000,
request_timeout=120.0, # seconds before giving up on a slow endpoint
num_retries=1,
seed=42 # optional; enables bit-stable output on vLLM
)
| Field | Default | Description |
|---|---|---|
model | required | Provider-prefixed model string, e.g. "openai/gpt-4o" |
api_base | None | Override endpoint URL |
api_key | None | None for keyless local servers |
temperature | 0.1 | Sampling temperature |
max_tokens | 1000 | Max output tokens |
context_window | None | Tokens the endpoint accepts per request; used only when the endpoint publishes no max_model_len |
request_timeout | 120.0 | Per-request timeout in seconds |
num_retries | 1 | Retry count on transient errors |
seed | None | vLLM sampling seed for reproducibility |
adapter_path | None | LoRA/fine-tuned artifact path (bookkeeping only; not read by create_dspy_lm) |
extra_body | None | Provider-specific request params |
extra_headers | None | Static HTTP headers sent with every request (e.g. semantic-router tenant/tier headers) |
LLMConfig - Multi-role LLM configuration:
from cogniverse_foundation.config.unified_config import LLMConfig, LLMEndpointConfig
llm_config = LLMConfig(
primary=LLMEndpointConfig(model="openai/gpt-4o", api_base="http://localhost:8101/v1"),
teacher=LLMEndpointConfig(model="openai/gpt-4o-mini", api_base="http://localhost:8101/v1"),
overrides={
# Per-component partial overrides merged onto primary at resolve time
"summarizer_agent": {"max_tokens": 2000},
}
)
# Resolve the effective config for a component
resolved = llm_config.resolve("summarizer_agent")
primary is the global default for all DSPy modules and also the student model during optimization. teacher is optional for non-optimization processes, but every teacher-dependent optimization must configure it explicitly: resolve_teacher() returns an isolated copy for BootstrapFewShot(teacher_settings={"lm": ...}) and raises when the role is absent instead of falling back to primary. overrides holds per-component partial dicts — only differing fields need to be specified; resolve(component) merges them field-by-field onto a copy of primary (never through to_dict(), which masks api_key — the resolved endpoint keeps the real key).
is_modal_inference_url(base_url) classifies canonical inference roots: it returns True only for root HTTPS *.modal.run URLs, and inference_headers(base_url) uses that predicate to return the shared Modal bearer only for those roots. Non-Modal roots remain keyless.
create_dspy_lm(config: LLMEndpointConfig) -> dspy.LM (cogniverse_foundation.config.llm_factory) is the single chokepoint every dspy.LM() construction in the codebase goes through. It wires api_base/api_key/temperature/max_tokens/timeout/num_retries onto the LM, merges seed into extra_body, forwards extra_headers, and resolves api_key through resolve_inference_api_key(api_base, api_key). Raises ValueError if config.model is empty.
create_dspy_lm(config, tenant_id=...) binds responses to a canonical tenant in TenantScopedLMCache (config/lm_response_cache.py). Its key includes the complete request, including model, messages, response format, tools, sampling and routing headers; only the API key and cache-control flag are excluded. DSPy caching is disabled, and an unbound LM does not cache. BodyBoundedLM.for_tenant(tenant_id) binds the ambient LM per request.
The process cache loads semantic_router.response_cache_ttl_seconds and response_cache_max_entries from the config file ConfigUtils._discover_config_file() resolves, on first use; no config file raises. Entries expire on a monotonic clock and evict in LRU order. Threads and event loops share one upstream call per key. Hits return independent copies with cache_hit=True and zero additional usage. Failed calls retain the provider exception with tenant, model and key context, store nothing, and allow the next request to retry. Invalid cache configuration raises with its file path.
BodyBoundedLM holds every call's max_tokens within the completion budget bound by cogniverse_foundation.config.lm_output_budget.bound_output_token_budget(budget): the smaller of that budget and the endpoint's configured max_tokens is sent, and the response cache keys on the value sent. output_token_budget_from(context) reads context["max_output_tokens"], returns None when absent and raises ValueError for anything but a positive integer.
create_budgeted_dspy_lm(config: LLMEndpointConfig) -> dspy.LM builds the same LM as a BudgetedLM (cogniverse_foundation.config.budgeted_lm), which fits every request inside the window its endpoint serves. On first call it resolves the window through resolve_context_window(api_base, declared=config.context_window, model=...) and holds it as a TokenBudget(model, context_window, reserved_output=config.max_tokens); input_budget is context_window - reserved_output. What the endpoint publishes as max_model_len at {api_base}/models wins (ResolvedContextWindow(tokens, source="served")), because the window belongs to the deployment rather than to the model; a listing that carries none falls back to config.context_window (source="declared"). Requests over that allowance shed whole few-shot demonstrations, oldest first, and log how many were dropped. A prompt that still overflows with none left raises PromptBudgetExceededError naming the window, the reservation, the measured input and the number dropped — as does a provider-side context-window rejection, so a caller never sees a bare litellm error. An endpoint that publishes no max_model_len and declares no context_window, or one with no api_base or no max_tokens, raises ContextWindowUnavailableError instead of being budgeted by guess. Used wherever a prompt can approach the window: the DSPy bootstrap teacher, the entity-extraction optimizer, the ingestion LM context (ingest_lm_context_for) and the ingestion worker's default LM. resolve_context_window memoizes per (api_base, model) for the process (32 entries, oldest evicted; clear_context_window_memo() drops them), so ingestion paths that build one LM per segment read the listing once. fitting_demonstrations(demos, budget=, render=, count_tokens=, calls=) sizes a compiled program before it is sent at all: the largest number of demos, kept from the front, whose request render(demos, call) fits budget.input_budget for every one of calls (zero when a call overflows with none). Calls are taken longest request first; each is counted once at the longest prefix every earlier call fit, and one that overflows it bisects below it, so a counter that pays a round-trip per count is called about once per call. served_message_counter(api_base, model) is such a counter: it counts messages through the endpoint's POST {root}/tokenize (vLLM) with the bare served model name, so the count is the served tokenizer and chat template's, the prompt_tokens the completion reports. litellm counts a model it has no tokenizer for with OpenAI's cl100k_base, which undercounts the gemma student by about a quarter on transcript text. A count is idempotent and a fitting pass asks hundreds, so an unreachable endpoint or a 5xx answer is asked again up to retries times (1 s, 2 s, 4 s ... apart; the optimizer passes the endpoint's num_retries, as its LM calls are retried). An endpoint that still cannot answer, or answers 4xx or without an integer count, raises TokenCountUnavailableError.
resolve_inference_api_key(api_base, api_key) is the one key-resolution rule for every OpenAI-compatible client, dspy.LM or not (Mem0's LLM provider, litellm.completion/rerank, the LLM/visual judges' chat-completions POSTs): an explicit key wins; otherwise the COGNIVERSE_INFERENCE_API_KEY bearer is sent whenever it is set (an in-cluster router may forward to Modal, and a self-hosted server ignores it); a *.modal.run api_base with no bearer raises naming the variable; a self-hosted api_base with nothing configured gets "not-required"; with no api_base the configured key passes through untouched.
from cogniverse_foundation.config.llm_factory import create_dspy_lm
from cogniverse_foundation.config.unified_config import LLMEndpointConfig
lm = create_dspy_lm(LLMEndpointConfig(model="openai/gpt-4o", api_base="http://localhost:8101/v1"))
SemanticRouterConfig - Opt-in routing of LLM calls through a vLLM Semantic Router:
from cogniverse_foundation.config.unified_config import SemanticRouterConfig
router_config = SemanticRouterConfig(
enabled=True,
semantic_router_url="http://semantic-router:8801/v1",
routed_model="openai/auto",
classification_model="openai/cogniverse-classification",
vision_model="openai/cogniverse-vision",
short_reasoning_model="openai/cogniverse-short-reasoning",
)
When enabled, cogniverse_foundation.config.semantic_router rewrites an LLMEndpointConfig to target semantic_router_url instead of the model backend, sets model to the router entry the call site takes (the router resolves models by its own catalog, aliases and entrypoints, not raw provider ids), and attaches two authz headers per request: tenant identity (user_id_header, default x-authz-user-id) and tenant tier (tier_header, default x-authz-user-groups, the caller's resolved RouterTier). When disabled, the endpoint passes through unchanged. The tier is the tenant's stored attribute: cogniverse_foundation.config.tenant_tiers holds it in the config store under scope ROUTING / service semantic_router / key tenant_tier, resolve_tenant_tier(config_accessor, tenant_id) reads it through a per-tenant RefreshingCache (refreshed off the request thread after TENANT_TIER_REFRESH_S, 15 s; never served past TENANT_TIER_MAX_STALENESS_S, 30 s) invalidated by every write in the process, an unset tenant is DEFAULT_ROUTER_TIER, and a store failure routes as DEFAULT_ROUTER_TIER with a WARNING. Claim extraction during ingestion keeps the direct primary endpoint.
Each call site names its router entry. CLASSIFICATION_CALL_SITES (entity extraction, gateway, orchestrator, profile selection, query enhancement, search) produce a bounded output and send classification_model, which names the chart router's cogniverse-classification entrypoint: its recipe chooses the decision from the tenant tier alone, so no domain classifier runs; every tier's bounded call is served by basic-chat (it never crosses to the teacher), and the decision's exact response cache still applies. SHORT_REASONING_CALL_SITES (deep-research decomposition and evidence evaluation) send short_reasoning_model, which names the cogniverse-short-reasoning entrypoint: its recipe also tests the tier alone and serves pro-reasoning with reasoning on for pro and basic-chat with reasoning off for free and base, with the same exact response cache. FREE_FORM_CALL_SITES send routed_model (the auto alias), where the router classifies the content and may promote the call to the reasoning model; a call site in neither set takes auto. Every decision in every routing profile caches with mode: exact: on the router's embedding model the closest different-content pair in the evaluation corpus scores higher than the weakest equivalent pair, so no similarity threshold is admissible.
When the router is enabled but inference.vllm_llm_teacher is neither enabled nor external, the chart serves pro-reasoning from the student's Envoy cluster and provider model. Helm NOTES and the rendered manifest name pro_model_unavailable; each pro decision, including short-reasoning-pro, carries tier_degraded: pro_model_unavailable. A served teacher uses its own cluster and has neither warning nor degradation marker.
| Function | Description |
|---|---|
resolve_semantic_router_headers(config, tenant_id) | Resolve the two authz headers, or None when disabled |
apply_semantic_routing(endpoint, config, tenant_id, tier, call_site) | Return a routed copy of endpoint, or the original when disabled |
routed_model_for(config, call_site) | classification_model for a call site in CLASSIFICATION_CALL_SITES, short_reasoning_model for one in SHORT_REASONING_CALL_SITES, routed_model otherwise |
create_routed_lm(endpoint, config, tenant_id, tier, call_site) | apply_semantic_routing + the shared LM construction; returns a RoutedLM (cogniverse_foundation.config.routed_lm), which sends a call carrying image parts on vision_model (the chart's cogniverse-vision entry, serving the multimodal student for every tier) and records the completion's model on the current span as llm.served_model (LLM_SERVED_MODEL_ATTRIBUTE) |
record_served_model(response) | Stamp a completion's model on the current span; a no-op outside any span |
ingest_lm_context_for(endpoint) | Return a direct dspy.context for ingestion-time LM calls (claim extraction); never routed |
routed_lm_context_for(config_manager, tenant_id, agent_name, endpoint=None, call_site=None) | Return a dspy.context binding the routed (or direct) LM for query-time agents, tenant-bound either way — the entry point agents use; call_site (default agent_name) names the router entry |
routed_lm_context_for_async(config_manager, tenant_id, agent_name, endpoint=None, call_site=None) | Await the same context off the event loop — the entry point every coroutine uses, since the tier resolution behind it is a config-store read; the returned context manager is unentered so the caller binds it on its own task |
resolve_semantic_router_config(config_accessor) | Read SemanticRouterConfig off an object exposing get_semantic_router() |
A failed routed completion raises a RoutedLMCallFailed subclass from cogniverse_foundation.config.routed_lm, chained from the litellm error and carrying status, router_code (the provider's error code), tenant_id, tier and routed_model. The router passes the upstream status through, so the class follows it:
| Status | Raised | Retried |
|---|---|---|
| 401, 403 | UpstreamAuthRejected | never |
| 429 | UpstreamRateLimited | up to num_retries |
| 5xx, reset, timeout | UpstreamUnavailable | up to num_retries |
| 404 | UpstreamNotServing (an UpstreamUnavailable) | never |
| any other 4xx | RouterDecodeFailed | never |
RoutedLM spends the endpoint's num_retries itself, only on UpstreamRateLimited and UpstreamUnavailable other than UpstreamNotServing; errors that are not provider or transport failures propagate unchanged. A timeout or refused connection carries the status litellm stamps on it (408 / 500) and is an UpstreamUnavailable regardless.
A 404 means the upstream has nothing deployed for the routed model: an undeployed Modal app answers modal-http: invalid function call, and the router passes the status through (x-vsr-response-path: upstream). UpstreamNotServing carries failed_fast and recheck_in_s (below).
LM endpoint availability¶
cogniverse_foundation.config.lm_endpoint_availability records, per process, the last outcome of every LM endpoint called — keyed by LMEndpoint(api_base, model, route), where route is the router tier on the routed path. BodyBoundedLM, which every LM (RoutedLM, BudgetedLM, create_dspy_lm) sends through, consults it on each call that misses the response cache:
| Last outcome | State | Next call |
|---|---|---|
| a completion | serving | sent |
| HTTP 404 | not_serving | refused without being sent for NOT_SERVING_RECHECK_S (30s), raising LMEndpointNotServing(failed_fast=True); then one call is sent as the recheck while the others keep failing fast, and its outcome replaces the state |
| a 5xx, a timeout, a refused connection | failing | sent — a scaled-to-zero endpoint answers after its cold start, so only a 404 fails fast |
A recheck still unanswered after PROBE_VERDICT_S (5s, five times the slowest measured 404 from an undeployed Modal app) is a cold start, and the calls behind it are sent too. A recheck that ends without an outcome (cancelled, past its caller's deadline) hands the recheck to the next call. The direct path raises LMEndpointNotServing (chained from the litellm NotFoundError on the call that got the 404); RoutedLM raises UpstreamNotServing chained from it. not_serving_cause(exc) finds either through __cause__, and refusal(endpoint) answers the fast failure a call would get without admitting one, for a caller deciding whether to prepare the call at all. Each not-serving call stamps llm.endpoint.state, llm.endpoint.failed_fast and llm.endpoint.recheck_in_s on its span (LLM_ENDPOINT_*_ATTRIBUTE). lm_endpoint_availability().snapshot() lists every endpoint's state, upstream_status, failure, reason, observed_at and recheck_in_s (credentials in the address stripped); the runtime's /health serves it. Nothing is probed in the background: any request to a deployed Modal endpoint boots a GPU container. At most MAX_TRACKED_ENDPOINTS (64) endpoints are tracked, least recently observed evicted first.
A call made under a caller's deadline never outlives it. The caller binds an LMCallDeadline (cogniverse_foundation.config.lm_deadline, bound_lm_call_deadline); BodyBoundedLM sends nothing once the deadline has passed or the caller abandoned it and raises LMCallDeadlineExceeded, a TimeoutError naming the endpoint, the model and the deadline. An OpenAI-compatible request goes through deadline_bound_openai_client(api_base, api_key), whose connects, writes and reads each wait at most the time left at that moment, and a caller joining an identical in-flight call waits no longer than its own deadline. RoutedLM starts no retry or student attempt past the deadline, and a failure under a deadline adds endpoint= and deadline_s= to its message. A request the client hangs up on because the deadline passed or was abandoned raises LMCallDeadlineExceeded chained from the provider's timeout: it does not mark the endpoint failing, and the response cache names it at INFO, since the caller reports the overrun. A timeout or refused connection with no deadline behind it is the provider's own error, recorded as failing and, for a tenant-bound LM, logged at ERROR by the response cache; the "LM rejected the request" error is logged only for a status the endpoint actually answered.
A pro free-form or short-reasoning call whose teacher raises UpstreamUnavailable with status 502, 503, 504, a timeout (408), a transport failure (500 or no status), or UpstreamNotServing (nothing deployed) gets one attempt on classification_model, which serves the student. The LM completion and current span carry tier_degraded: pro_model_unavailable, upstream_status, and upstream_exception_type. The exception type and its status determine eligibility; message text never does. A permanent refusal propagates, and default, bounded-output, and vision calls carry no degradation marker. If the student fails too, its typed failure propagates without another attempt. tier_degradation_context() collects these fields for one dispatch, including calls on worker threads, and restores the enclosing scope on exit.
Configuration Scopes¶
Configurations are organized by scope for isolation:
| Scope | Description | Example Keys |
|---|---|---|
SYSTEM | Infrastructure settings | backend_url, backend_port, summarizer_agent_url |
AGENT | Per-agent settings | module_config, llm_model, llm_temperature |
ROUTING | Routing agent settings | routing_mode, enable_fast_path, gliner_threshold |
TELEMETRY | Telemetry settings | otlp_endpoint, otlp_enabled, level |
SCHEMA | Deployed Vespa schema tracking (used by cogniverse_core.registries.schema_registry) | schema_name, deployment status |
BACKEND | Backend profiles | embedding_model, schema_name, pipeline_config |
VespaConfigStore.compare_and_set_config(tenant_id, scope, service, config_key, config_value, expected_version=...) writes the version after expected_version when that is still the latest. The required keyword expected_version must be nonnegative: zero requires an absent key, and negative values raise ValueError. The version is reserved on the key's version counter only while no other writer holds an unwritten reservation; one older than 30 seconds is treated as abandoned and reserved past. It returns the committed ConfigEntry after a strong read confirms it is the latest version, or None on a version mismatch, a live reservation, or a conditional-write conflict. Every successful version write applies the same keep_versions history retention as set_config, including a stale write that loses the final read; pruning is best-effort. Read/write failures raise. Schema deployment journals and registration completion use this operation to fence stale writers and deletion tombstones.
from cogniverse_sdk.interfaces.config_store import ConfigScope
# Get arbitrary config value by scope
value = config_manager.get_config_value(
tenant_id="acme",
scope=ConfigScope.AGENT,
service="orchestrator_agent",
config_key="optimizer_config"
)
Configuration Inheritance¶
The configuration system uses a layered inheritance model where tenant-specific settings override system defaults:
flowchart TB
subgraph Sources["<span style='color:#000'>Configuration Sources</span>"]
EnvVars["<span style='color:#000'>Environment Variables<br/>COGNIVERSE_CONFIG, etc.</span>"]
ConfigFile["<span style='color:#000'>config.json<br/>Auto-discovered</span>"]
VespaStore["<span style='color:#000'>Vespa Store<br/>Persisted configs</span>"]
end
subgraph Layers["<span style='color:#000'>Configuration Layers</span>"]
direction TB
SystemDefaults["<span style='color:#000'>System Defaults<br/>Hardcoded fallbacks</span>"]
GlobalConfig["<span style='color:#000'>Global Configuration<br/>config.json profiles</span>"]
TenantOverlay["<span style='color:#000'>Tenant Overlay<br/>Per-tenant overrides</span>"]
RuntimeOverride["<span style='color:#000'>Runtime Override<br/>API/query-time params</span>"]
end
subgraph Resolution["<span style='color:#000'>Resolution Order (Bottom Wins)</span>"]
Final["<span style='color:#000'>Final Configuration<br/>Merged result</span>"]
end
EnvVars --> GlobalConfig
ConfigFile --> GlobalConfig
VespaStore --> TenantOverlay
SystemDefaults --> Final
GlobalConfig --> Final
TenantOverlay --> Final
RuntimeOverride --> Final
style Sources fill:#90caf9,stroke:#1565c0,color:#000
style Layers fill:#ffcc80,stroke:#ef6c00,color:#000
style Resolution fill:#a5d6a7,stroke:#388e3c,color:#000
style EnvVars fill:#90caf9,stroke:#1565c0,color:#000
style ConfigFile fill:#90caf9,stroke:#1565c0,color:#000
style VespaStore fill:#90caf9,stroke:#1565c0,color:#000
style SystemDefaults fill:#ffcc80,stroke:#ef6c00,color:#000
style GlobalConfig fill:#ffcc80,stroke:#ef6c00,color:#000
style TenantOverlay fill:#ffcc80,stroke:#ef6c00,color:#000
style RuntimeOverride fill:#ffcc80,stroke:#ef6c00,color:#000
style Final fill:#a5d6a7,stroke:#388e3c,color:#000 Configuration Resolution Example:
from cogniverse_foundation.config.unified_config import BackendConfig, BackendProfileConfig
# System default (hardcoded)
max_frames = 50
# Global config (config.json) - overrides default
config_json = {
"profiles": {
"video_colpali_mv_frame": {
"max_frames": 100
}
}
}
# Tenant overlay (Vespa) - overrides global
config_manager.add_backend_profile(
tenant_id="premium_tenant",
profile=BackendProfileConfig(
profile_name="video_colpali_mv_frame",
type="video",
pipeline_config={"max_frames": 200},
),
)
async def search(query: str, max_frames: int):
... # Runtime override (query param) - overrides all
async def main():
result = await search(query="cats", max_frames=300)
# Final: premium_tenant gets max_frames=300 for this query
return result
Resolution Priority (highest to lowest):
| Priority | Source | Scope | Example |
|---|---|---|---|
| 1 (highest) | Runtime Override | Per-request | Query params, API args |
| 2 | Tenant Overlay | Per-tenant | ConfigManager.set_*_config() (persisted to Vespa) |
| 3 | Global Config | All tenants | config.json profiles |
| 4 (lowest) | System Defaults | Fallback | Hardcoded in classes |
ConfigAPIMixin¶
cogniverse_foundation.config.api_mixin.ConfigAPIMixin adds runtime, persisted configuration REST endpoints to an agent's FastAPI app. It expects the host class to expose self.agent_config (an AgentConfig) plus update_module_config() / update_optimizer_config() methods (provided by DynamicDSPyMixin in cogniverse_core).
class MyAgent(DynamicDSPyMixin, ConfigAPIMixin):
def __init__(self, tenant_id, config_manager):
self.initialize_dynamic_dspy(agent_config)
app = FastAPI()
self.setup_config_endpoints(app, config_manager, tenant_id=tenant_id)
setup_config_endpoints(app, config_manager, tenant_id=None) registers:
| Route | Description |
|---|---|
GET /config | Current AgentConfig as a dict |
GET /config/module | Current DSPy module info |
POST /config/module | Update module config (ModuleConfigUpdate); persists via ConfigManager.set_agent_config |
GET /config/optimizer | Current optimizer info |
POST /config/optimizer | Update optimizer config (OptimizerConfigUpdate); persists |
POST /config/llm | Update LLM fields (LLMConfigUpdate) and reconfigure the DSPy LM; persists |
GET /config/modules/available | List registered DSPy module types |
GET /config/optimizers/available | List registered DSPy optimizer types |
Every mutating endpoint persists through the injected ConfigManager, so changes survive a restart and are versioned like any other config write. The tenant is fixed when the routes are registered; request query parameters cannot redirect a write to another tenant. Mutations are serialized, synchronous store calls run in a worker thread, and a persistence failure restores the prior in-memory configuration before the endpoint returns an error.
Telemetry System¶
TelemetryManager¶
TelemetryManager is a singleton that manages OpenTelemetry tracing with multi-tenant isolation.
Key Features:
- Tenant-isolated tracer providers
- LRU caching of tracers
- Graceful degradation when telemetry unavailable
- Session tracking for multi-turn conversations
- Phoenix integration for trace visualization
from cogniverse_foundation.telemetry.manager import TelemetryManager, get_telemetry_manager
# Get global singleton
telemetry = get_telemetry_manager()
# Create span with tenant isolation
with telemetry.span("search.execute", tenant_id="acme") as span:
span.set_attribute("query", "find videos about cats")
# ... search logic ...
API Reference:
| Method | Description |
|---|---|
span(name, tenant_id, project_name=None, attributes=None, component="agents") | Create tenant-isolated span. component gates emission against TelemetryConfig.level (one of search_service/agents/backend/pipeline/encoder) |
session_span(name, tenant_id, session_id, project_name=None, attributes=None, component="search_service") | Span within a session context; all nested span() calls inherit session_id |
get_tracer(tenant_id, project_name=None) | Get tracer (prefer span()) |
get_provider(tenant_id, project_name=None) | Get telemetry provider for queries (spans/annotations/datasets) |
register_project(tenant_id, project_name, **kwargs) | Register project with config |
force_flush(timeout_millis=10000) | Flush all pending spans |
shutdown() | Graceful shutdown |
get_stats() | Get telemetry statistics |
There is no standalone session() context manager — session tracking is done by calling session_span() with the same session_id for each operation that should share a session, or by nesting span() calls inside one session_span().
Span Context¶
Creating spans with tenant isolation:
from cogniverse_foundation.telemetry.manager import get_telemetry_manager
telemetry = get_telemetry_manager()
async def process(query: str):
# Basic span
with telemetry.span("agent.process", tenant_id="acme") as span:
span.set_attribute("agent.name", "orchestrator_agent")
span.set_attribute("query.length", len(query))
result = await process_query(query)
# Span with project isolation (for management operations)
with telemetry.span(
"experiment.run",
tenant_id="acme",
project_name="experiments", # Separate Phoenix project
attributes={"experiment.name": "optimizer_v2"}
) as span:
await run_experiment()
return result
Session Tracking¶
Track multi-turn conversations across requests:
async def handle_request(query: str, session_id: str):
# At API entry point - establish session context
with telemetry.session_span(
"api.search.request",
tenant_id="acme",
session_id=session_id,
attributes={"query": query, "turn": 3}
) as span:
# All child spans inherit session_id
result = await search_service.search(query)
return result
# Alternative: nest plain spans inside one session_span - all inherit session_id
with telemetry.session_span("session.start", tenant_id="acme", session_id="session-xyz"):
with telemetry.span("operation1", tenant_id="acme") as span1:
pass
with telemetry.span("operation2", tenant_id="acme") as span2:
pass
# Both spans share session_id
Project Registration¶
Register projects with custom endpoints (useful for tests):
# Register with default config
telemetry.register_project(
tenant_id="acme",
project_name="search"
)
# Register with custom endpoints (for tests)
telemetry.register_project(
tenant_id="test-tenant",
project_name="synthetic_data",
otlp_endpoint="http://localhost:24317",
http_endpoint="http://localhost:26006",
use_sync_export=True # Sync export for tests
)
Span Context Helpers¶
cogniverse_foundation.telemetry.context provides pre-shaped span helpers that standardize the attribute set for the three most common span kinds, plus two helpers for enriching an existing span:
| Function | Description |
|---|---|
search_span(tenant_id, query, top_k=10, ranking_strategy="default", profile="unknown", backend="vespa") | Context manager for a search_service.search span (component="search_service") |
encode_span(tenant_id, encoder_type, query_length=0, query="") | Context manager for an encoder.<type>.encode span (component="encoder") |
backend_search_span(tenant_id, backend_type="vespa", schema_name="unknown", ranking_strategy="default", top_k=10, has_embeddings=False, query_text="") | Context manager for a search.execute span (component="backend") |
add_search_results_to_span(span, results, output_value=None) | Set num_results/output.value/top_score attributes and a search_results event; pass output_value to reuse one serialization across both spans |
serialize_search_results(results) | Serialize result rows to the canonical output.value JSON once (feed the result to both add_search_results_to_span calls of a search) |
add_embedding_details_to_span(span, embeddings) | Set embedding_shape/embedding_dtype/norm-mean/norm-std attributes from an embeddings array |
Each context manager wraps TelemetryManager.span(), so it still goes through the tenant-required check and the component-based TelemetryConfig.level gating.
Registry System¶
The cogniverse_foundation/registry/ subpackage provides EntryPointRegistry[T] — a generic plugin registry over importlib.metadata entry points. Subclasses declare an entry-point group and a label; the base handles discovery, manual registration, conflict detection, tenant-scoped or config-keyed caching, and lifecycle-style initialization. Four registries in the codebase are thin subclasses of this base:
| Registry | Subpackage | Entry-point group | Tenant-scoped |
|---|---|---|---|
TelemetryRegistry | cogniverse_foundation.telemetry.registry | cogniverse.telemetry.providers | yes |
EvaluationRegistry | cogniverse_evaluation.providers.registry | cogniverse.evaluation.providers | yes |
WorkflowStoreRegistry | cogniverse_core.registries.workflow_store_registry | cogniverse.workflow.stores | no |
AdapterStoreRegistry | cogniverse_core.registries.adapter_store_registry | cogniverse.adapter.stores | no |
Defining a new registry¶
from cogniverse_foundation.registry import EntryPointRegistry
from my_pkg.interfaces import MyStore
class MyStoreRegistry(EntryPointRegistry[MyStore]):
_entry_point_group = "myapp.stores"
_label = "my store"
# _tenant_scoped = False (default): klass(**config), cached by backend_url/port
# _tenant_scoped = True: klass() + .initialize({**config, tenant_id}), cached by tenant
Implementations register via their pyproject.toml:
Callers fetch instances with MyStoreRegistry.get(name="default", config={...}). Conflict detection is always on — if two installed packages both register name="default" under the same group, discover() raises ValueError rather than silently picking one. Discovery runs once per registry even when threads race the first lookup. For a tenant-scoped registry, get() canonicalizes its tenant_id argument and places that canonical value in the initialization mapping after caller config is merged, so a config payload cannot override the cache's tenant identity. TelemetryRegistry further keys providers by project plus HTTP and gRPC endpoints, so reconfiguring an existing project creates a provider bound to the new destinations.
Tenant-Scoped Caching¶
cogniverse_foundation.caching.TenantLRUCache is a thread-safe, bounded LRU cache keyed by tenant_id. Long-running processes keep per-tenant state (Mem0 memory managers, telemetry providers, backend clients, compiled DSPy modules) in memory; without a bound, a multi-tenant server accumulates one instance per tenant indefinitely. EntryPointRegistry itself uses a TenantLRUCache for its per-instance plugin cache.
from cogniverse_foundation.caching import TenantLRUCache
def close_client(tenant_id: str, client) -> None:
client.close()
cache = TenantLRUCache[MyClient](capacity=16, on_evict=close_client)
# Atomically get-or-build under a lock (the primary entry point)
client = cache.get_or_set("acme", lambda: MyClient(tenant_id="acme"))
| Method | Description |
|---|---|
__init__(capacity, on_evict=None) | capacity must be >= 1; on_evict(key, value) is called (best-effort) on eviction or clear() |
get(key) | Return cached value or None, marking it most-recently-used |
set(key, value) | Insert/replace, evicting the least-recently-used entry if over capacity; a replaced value (different object, same key) also gets on_evict so held resources are released |
get_or_set(key, factory) | Return cached value, or build + cache one atomically (lock held through the factory) |
set_if_absent(key, value) | Insert unless the key is cached; return the winner. For expensive instances built outside the lock: concurrent builders converge on one shared instance, and the loser stays the caller's to release (no on_evict) |
acquire(key) / release(key) | Check a value out against eviction, and give the checkout back. acquire returns None and takes no checkout for an uncached key |
lease(key) / lease_value(value) | Context managers over acquire/release, by key or by the held object; the checkout ends on every exit path including exceptions |
lease_count(key) | Outstanding checkouts on a key |
key_of(value) | The key holding a value by identity, or None |
pop(key, default=None) | Remove and return a value without triggering on_evict |
clear() | Evict every entry (triggers on_evict for each; a checked-out entry's on_evict is deferred to its release) |
keys() / values() | Snapshot of current keys / values (LRU order) |
copy() | Shallow copy preserving capacity, on_evict, and LRU order |
len(cache) / key in cache | Current size / membership check |
COGNIVERSE_TENANT_CACHE_CAPACITY (env var, default 16) sizes the per-registry instance cache inside EntryPointRegistry.
Checkouts¶
A value in use is checked out for the length of that use. Capacity eviction skips checked-out entries and takes the least-recently-used free entry instead, so a cached value serving a request is never closed under it. When every entry is checked out the cache holds more than its capacity and logs one warning naming the overflow and the number of checkouts blocking it; eviction resumes as checkouts end. A close aimed at a checked-out entry — an overwriting set, a clear — is deferred to the release rather than dropped, so its on_evict still runs exactly once.
with cache.lease("acme") as client: # None when "acme" is not cached
client.query(...) # "acme" cannot be evicted here
Tenant-delete eviction¶
register_tenant_cache(cache) and evict_tenant_from_registered_caches(tenant_id) (cogniverse_foundation.caching) let independently-owned TenantLRUCache instances release a deleted tenant's state immediately instead of waiting for LRU pressure. Registration holds only a weak reference (a weakref.WeakSet), so a cache owned by a discarded consumer drops out on garbage collection rather than leaking through the registry.
from cogniverse_foundation.caching import (
TenantLRUCache,
register_tenant_cache,
evict_tenant_from_registered_caches,
)
cache = register_tenant_cache(TenantLRUCache[MyClient](capacity=64))
# On tenant delete:
evicted = evict_tenant_from_registered_caches("acme:production")
evict_tenant_from_registered_caches pops both the given key and its canonical org:tenant form from every registered cache, so simple-form entries are covered too; on_evict does not fire — deletion is a hard drop. It returns the number of entries dropped. The runtime registers six such caches, each of capacity 64 — the per-tenant GatewayAgent cache, the generic-A2A-agent cache (keyed (tenant, agent_name)) and the orchestrator agent cache in agent_dispatcher.py, the per-tenant GraphManager and WikiManager caches in main.py, and the per-tenant ArtifactManager cache in routers/agents.py — and calls evict_tenant_from_registered_caches from release_deleted_tenant, the handler every runtime worker process runs for a tenant delete, so a deleted tenant's cached state is released as part of the delete, not left to linger.
Refreshing values off the request thread¶
cogniverse_foundation.caching.RefreshingCache holds values read from a slow backing store and keeps the read off the caller's thread once a value is held. ConfigManager keeps its system config and its per-tenant scoped configs in one each, TenantRouterTiers its per-tenant router tiers, and DeployedSchemaNames (core module) its per-tenant deployed schema names.
from cogniverse_foundation.caching import RefreshingCache
cache = RefreshingCache[tuple, dict](
name="scoped-config",
refresh_after_s=5.0,
max_staleness_s=60.0,
max_entries=512,
)
value = cache.get(key, lambda: store.read(key))
cache.invalidate(lambda k: k[1] == tenant_id) # after this process writes
| Age of the held value | get(key, read) |
|---|---|
| none held | runs read on the caller's thread; concurrent callers share it; a failure raises to each and caches nothing |
below refresh_after_s | returns it, no read |
refresh_after_s to below max_staleness_s | returns it and starts one background read (daemon thread <name>-refresh); concurrent callers share that read and none waits for it |
max_staleness_s or more | runs read on the caller's thread, or waits on the read already in flight; a failure raises |
Age counts from when the producing read began, so no value is returned max_staleness_s or more after its read started. A failed background read is logged at ERROR with the key and the age of the value still being served. The entry stays, and the key is not refreshed again for refresh_after_s. At most max_background_reads (default 4) background reads run at once. A stale entry found while all are busy is returned, and its refresh starts on a later call. invalidate(matches) drops matching entries and detaches their reads in flight, whose results are never cached. put(key, value) holds a value this process has just written as if read now, detaching the key's read in flight. max_entries bounds the cache, evicting least recently used first. Setting both bounds to 0 reads on every call. The clock argument (default time.monotonic) sets the time source.
get(key, read, accept=...) lets one call refuse a held value: when accept(held) is False the call reads as if nothing were held. It runs under the cache lock and must not call back into the cache. The constructor's keep(value) decides which read results are held; a result it rejects is returned to its callers and drops the key's entry. keys() and items() snapshot the held entries, least recently used first.
DSPy Extensions¶
cogniverse_foundation.dspy hosts DSPy adapter and model-name helpers shared across agents, evaluation, and finetuning.
LenientJSONAdapter (cogniverse_foundation.dspy.lenient_json_adapter) — a dspy.adapters.json_adapter.JSONAdapter subclass that renames common LM field-name variants (e.g. reason/rationale/thought → reasoning, answer/response/output → summary, sub_question → sub_questions) before the parent's strict field-key equality check. A required output the LM left unfilled — absent, null, or a blank string — raises LMOutputIncomplete (an AdapterParseError subclass) whose missing_fields names every such output and whose parsed_result carries the ones the LM did fill. An empty list or dict is a value and passes. An output declared with a default (dspy.OutputField(default="")) is one whose blank answer is valid: a null or blank value takes the default, and only an absent field is incomplete.
from cogniverse_foundation.dspy.lenient_json_adapter import LenientJSONAdapter
import dspy
dspy.configure(adapter=LenientJSONAdapter())
StructuredJSONAdapter / signature_response_format(signature) (cogniverse_foundation/dspy/structured_json_adapter.py) — a JSONAdapter subclass that sends the signature's output fields to the server as an OpenAI response_format of type json_schema: every output field required, additionalProperties: false, strict: true. The schema is derived from the signature, so guided decoding on an OpenAI-compatible engine (vLLM) can only return an object carrying every field. Stock JSONAdapter asks for a schema only when litellm's registry claims the model supports one and otherwise falls back to {"type": "json_object"}, which a bare {} satisfies; this adapter always sends the schema and raises AdapterParseError rather than reprompting under a second adapter when a server ignores it. A field value that fails its declared type — a list item outside a Literal, a missing or extra key on a nested model — raises AdapterParseError too, carrying the response, so every answer outside the schema is one error type.
The schema's name is the signature's class name, except for the placeholder names dspy gives a signature built from a string, rebuilt by with_instructions() or wrapped by ChainOfThought (Signature, StringSignature). Those fall back to the signature's own output field names plus a digest of its field declaration, so two different declarations never share a schema name and rewriting instructions does not rename the schema.
from cogniverse_foundation.dspy import StructuredJSONAdapter
import dspy
with dspy.context(adapter=StructuredJSONAdapter()):
prediction = module(query="...")
bare_model_name(model) / ensure_provider_prefix(model, default_provider="openai") (cogniverse_foundation.dspy.model_format) — strip or add a litellm provider prefix (ollama, ollama_chat, hosted_vllm, openai) on a model id, for sites that talk to an OpenAI-compatible HTTP API directly and need a bare model name (Mem0's embedder/LLM wiring) versus sites that need litellm's required provider/model form.
from cogniverse_foundation.dspy.model_format import bare_model_name, ensure_provider_prefix
bare_model_name("hosted_vllm/Qwen/Qwen2.5-7B-Instruct") # "Qwen/Qwen2.5-7B-Instruct"
ensure_provider_prefix("gemma3:4b") # "openai/gemma3:4b"
Confidence Parsing¶
cogniverse_foundation.confidence.parse_confidence(raw, default=0.0) -> float maps a DSPy module's LM-produced confidence/relevance output — a float, a percent string ("85%"), a label ("high"/"medium"/"low"), a sentence-embedded number ("0.9 (very confident)", "confidence: 0.9", "0.9/1.0"), or an empty string — to a clamped [0.0, 1.0] float, falling back to default on any input it cannot interpret. Lives in foundation so core, agents, evaluation, and finetuning share one implementation instead of each doing a raw float(result.confidence) that crashes on a non-numeric shape.
The sentence-embedded fallback only fires when the string isn't a bare number or exact label match. It first looks for an x/y ratio and accepts it only if x/y lands in [0, 1]; otherwise it takes the first free-standing float token that lands in [0, 1] (digits embedded in a larger token — hex reprs, versions like "1.2.3", identifiers — are skipped). An out-of-range or absent number falls back to default rather than guessing, e.g. "7 (very confident)" stays uninterpretable.
from cogniverse_foundation.confidence import parse_confidence
parse_confidence("85%") # 0.85
parse_confidence("high") # 0.9
parse_confidence(1.5) # 1.0 (clamped)
parse_confidence("") # 0.0 (default)
parse_confidence("0.9 (very confident)") # 0.9 (numeric-extraction fallback)
parse_confidence("confidence: 0.9") # 0.9 (numeric-extraction fallback)
parse_confidence("0.9/1.0") # 0.9 (x/y ratio fallback)
parse_confidence("7 (very confident)") # 0.0 (default; 7 is out of [0, 1])
Usage Examples¶
Complete Configuration Setup¶
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_foundation.config.unified_config import (
SystemConfig,
BackendConfig,
BackendProfileConfig,
)
from cogniverse_foundation.config.agent_config import AgentConfig
# Initialize config manager
config_manager = create_default_config_manager()
# Set global system config (not per-tenant)
system_config = SystemConfig(
search_backend="vespa",
backend_url="http://localhost",
backend_port=8080
)
config_manager.set_system_config(system_config)
# Add backend profile for tenant
profile = BackendProfileConfig(
profile_name="custom_colpali",
type="video",
embedding_model="colpali-v2",
schema_name="video_colpali_custom",
pipeline_config={"chunk_strategy": "frame", "top_k": 20}
)
config_manager.add_backend_profile(profile, tenant_id="acme")
# Set agent config (requires all fields)
from cogniverse_foundation.config.agent_config import (
AgentConfig, ModuleConfig, DSPyModuleType
)
agent_config = AgentConfig(
agent_name="orchestrator_agent",
agent_version="1.0.0",
agent_description="Routes queries",
agent_url="http://localhost:8001",
capabilities=["routing"],
skills=[],
module_config=ModuleConfig(
module_type=DSPyModuleType.CHAIN_OF_THOUGHT,
signature="query -> decision"
),
llm_model="gpt-4",
llm_temperature=0.7
)
config_manager.set_agent_config(
tenant_id="acme",
agent_name="orchestrator_agent",
agent_config=agent_config
)
Argo Workflows HTTP Client¶
build_argo_async_client() returns an httpx.AsyncClient with verify=False and DEFAULT_ARGO_TIMEOUT.
from cogniverse_foundation.common.argo_client import build_argo_async_client
client = build_argo_async_client()
client = build_argo_async_client(30.0)
The tenant job-scheduling route and optimization submissions use it for Argo calls.
Tenant-Specific Profile Overrides¶
from cogniverse_foundation.common.tenant_utils import SYSTEM_TENANT_ID
# Start with system profile, customize for tenant
config_manager.update_backend_profile(
profile_name="video_colpali_mv_frame",
overrides={
"embedding_model": "colpali-custom",
"top_k": 25
},
base_tenant_id=SYSTEM_TENANT_ID, # Inherit from the cluster-wide system profile (the default)
target_tenant_id="acme" # Save to tenant
)
Telemetry with Phoenix¶
from cogniverse_foundation.telemetry.manager import get_telemetry_manager
from cogniverse_foundation.telemetry.config import TelemetryConfig
# Initialize with config
config = TelemetryConfig(
enabled=True,
otlp_endpoint="http://localhost:4317",
service_name="cogniverse",
environment="production"
)
telemetry = TelemetryManager(config=config)
# Use in agent processing
async def process_query(query: str, tenant_id: str):
with telemetry.span(
"agent.routing",
tenant_id=tenant_id,
attributes={"query": query}
) as span:
# Route query
route = await route_query(query)
span.set_attribute("route.agent", route.agent_name)
span.set_attribute("route.confidence", route.confidence)
# Execute with nested span
with telemetry.span(
f"agent.{route.agent_name}",
tenant_id=tenant_id
) as child_span:
result = await execute_agent(route.agent_name, query)
child_span.set_attribute("result.count", len(result.items))
return result
Querying Telemetry Data¶
from datetime import datetime, timezone
async def query_and_annotate():
# Get provider for querying spans
provider = telemetry.get_provider(tenant_id="acme")
# Query spans from Phoenix
spans_df = await provider.traces.get_spans(
project="cogniverse-acme",
start_time=datetime(2025, 1, 1, tzinfo=timezone.utc),
limit=1000
)
# Stream page-sized frames when the result set is large.
async for page in provider.traces.iter_spans(
project="cogniverse-acme",
page_size=512,
columns=("start_time", "end_time", "status_code"),
):
process(page)
# Add annotation (project is required)
await provider.annotations.add_annotation(
span_id="abc123",
name="human_review",
label="approved",
score=1.0,
metadata={"reviewer": "alice"},
project="cogniverse-acme"
)
Architecture Position¶
flowchart TB
subgraph CoreLayer["<span style='color:#000'>Core Layer</span>"]
Core["<span style='color:#000'>cogniverse-core (agents)</span>"]
Evaluation["<span style='color:#000'>cogniverse-evaluation</span>"]
end
subgraph FoundationLayer["<span style='color:#000'>Foundation Layer</span>"]
Foundation["<span style='color:#000'>cogniverse-foundation ◄─ YOU ARE HERE<br/>ConfigManager, TelemetryManager, Provider Registry</span>"]
SDK["<span style='color:#000'>cogniverse-sdk (interfaces)</span>"]
end
CoreLayer --> FoundationLayer
style CoreLayer fill:#ce93d8,stroke:#7b1fa2,color:#000
style Core fill:#ce93d8,stroke:#7b1fa2,color:#000
style Evaluation fill:#ce93d8,stroke:#7b1fa2,color:#000
style FoundationLayer fill:#a5d6a7,stroke:#388e3c,color:#000
style Foundation fill:#a5d6a7,stroke:#388e3c,color:#000
style SDK fill:#a5d6a7,stroke:#388e3c,color:#000 Dependencies (declared in pyproject.toml):
cogniverse-sdk: Pure interfaces (ConfigStore, Backend, etc.)dspy-ai:LenientJSONAdaptersubclassesdspy.adapters.json_adapter.JSONAdapter;create_dspy_lmbuildsdspy.LMopentelemetry-api/sdk: Telemetry infrastructurepydantic: Configuration validationpandas: DataFrame return type for theTraceStore/AnnotationStore/DatasetStoreinterfaces
Not a declared dependency, but imported at runtime:
cogniverse-vespa:VespaConfigStoreis imported lazily insidecreate_default_config_manager()— only required whenbackend.type == "vespa"
Package Structure — cogniverse_foundation/common/:
The shared, dependency-free helpers that the whole stack uses live here so foundation code (config/manager.py, config/utils.py, config/unified_config.py, config/api_mixin.py, registry/entry_point_registry.py) can use them without importing upward into cogniverse-core:
common/tenant_utils.py— tenant identity helpers (SYSTEM_TENANT_ID,require_tenant_id,canonical_tenant_id,parse_tenant_id,validate_tenant_id,get_tenant_storage_path,sanitize_k8s_label_value).cogniverse_core.common.tenant_utilsre-exports these, so existingfrom cogniverse_core.common.tenant_utils import ...call sites keep working; the runtime-coupledassert_tenant_existsstays in core.common/dspy_module_registry.py—DSPyModuleRegistry/DSPyOptimizerRegistry, re-exported fromcogniverse_core.common.dspy_module_registry.
This keeps the Core → Foundation dependency direction one-way: cogniverse-foundation no longer imports cogniverse-core and is usable standalone.
Dependents:
cogniverse-core: Uses ConfigManager, TelemetryManagercogniverse-agents: Uses configuration and telemetrycogniverse-telemetry-phoenix: Implements telemetry provider
Testing¶
# Run foundation tests (confidence, caching, dspy adapters, semantic router, telemetry, config)
JAX_PLATFORM_NAME=cpu uv run pytest tests/foundation/ tests/telemetry/ tests/common/unit/ tests/common/integration/ -v
# Test confidence parsing, TenantLRUCache, dspy model-format, semantic router, entry-point registry
uv run pytest tests/foundation/unit/ -v
# Semantic router end-to-end (spins up a stub upstream)
uv run pytest tests/foundation/integration/test_semantic_router_e2e.py -v
# Test configuration
uv run pytest tests/common/unit/ -v -k "config"
# Test telemetry
uv run pytest tests/telemetry/ -v
# Test with coverage
uv run pytest tests/foundation/ tests/telemetry/ tests/common/unit/ tests/common/integration/ --cov=cogniverse_foundation --cov-report=html
Test Categories:
tests/foundation/unit/-parse_confidence,TenantLRUCache, DSPy model-format helpers,create_dspy_lm, semantic router unit tests,EntryPointRegistrytests/foundation/integration/- Semantic router end-to-end against a stub upstream (_sr_stack/stub_upstream.py)tests/telemetry/- Telemetry manager and provider teststests/common/unit/- Unit tests forAgentConfig,ConfigAPIMixin, and related agent-facing configurationtests/common/integration/- Integration tests with backend (e.g., Vespa) config persistence
Related Documentation¶
- Core Module - Agent base classes that use configuration and telemetry
- Configuration System Guide - Detailed configuration guide
- Telemetry Module - Phoenix provider implementation
- Multi-Tenant Architecture - Tenant isolation patterns
Summary: The Foundation module provides the infrastructure layer for Cogniverse. ConfigManager handles multi-tenant, versioned configuration with pluggable backend persistence (VespaConfigStore). TelemetryManager provides OpenTelemetry tracing with tenant isolation and Phoenix integration. All operations require tenant_id to ensure proper multi-tenant isolation.