Skip to content

Foundation Module

Package: cogniverse_foundation Location: libs/foundation/cogniverse_foundation/


Table of Contents

  1. Overview
  2. Package Structure
  3. Configuration System
  4. ConfigManager
  5. Configuration Types
  6. Configuration Scopes
  7. Configuration Inheritance
  8. ConfigAPIMixin
  9. Telemetry System
  10. TelemetryManager
  11. Span Context
  12. Session Tracking
  13. Project Registration
  14. Span Context Helpers
  15. Registry System
  16. Tenant-Scoped Caching
  17. DSPy Extensions
  18. Confidence Parsing
  19. Usage Examples
  20. Architecture Position
  21. Testing

Overview

The Foundation module provides infrastructure services that all other modules depend on:

  • Configuration Management: Multi-tenant, versioned configuration with pluggable ConfigStore persistence (default: Vespa)
  • LLM Wiring: create_dspy_lm factory, opt-in semantic-router request routing, and DSPy adapter/model-name helpers
  • Telemetry Infrastructure: OpenTelemetry-based tracing with tenant isolation
  • Provider Abstraction: Pluggable backends for telemetry (Phoenix, etc.) via a generic entry-point plugin registry
  • Shared Utilities: Tenant-scoped LRU caching and robust LM-output confidence parsing used across core, agents, evaluation, and finetuning

All configuration and telemetry operations are tenant-aware - tenant_id is required for all operations.


Package Structure

flowchart TB
    subgraph Foundation["<span style='color:#000'><b>cogniverse_foundation/</b></span>"]
        ConfigDir["<span style='color:#000'><b>config/</b><br/>Configuration system</span>"]
        TelemetryDir["<span style='color:#000'><b>telemetry/</b><br/>Telemetry system</span>"]
        RegistryDir["<span style='color:#000'><b>registry/</b><br/>Generic entry-point plugin registry</span>"]
        CachingDir["<span style='color:#000'><b>caching/</b><br/>Tenant-scoped LRU and refreshing caches</span>"]
        DspyDir["<span style='color:#000'><b>dspy/</b><br/>DSPy adapters &amp; model-format helpers</span>"]
        CommonDir["<span style='color:#000'><b>common/</b><br/>Tenant identity, DSPy registry &amp; Argo client helpers</span>"]
        ConfidencePy["<span style='color:#000'>confidence.py<br/>parse_confidence()</span>"]
        InferenceSpecs["<span style='color:#000'>inference_specs.py<br/>InferenceServiceSpec, get_inference_service_spec()</span>"]
        InitPy["<span style='color:#000'>__init__.py</span>"]
    end

    subgraph ConfigFiles["<span style='color:#000'><b>config/ files</b></span>"]
        Manager["<span style='color:#000'>manager.py<br/>ConfigManager - central API</span>"]
        UnifiedConfigFile["<span style='color:#000'>unified_config.py<br/>SystemConfig, BackendConfig, LLMConfig</span>"]
        AgentConfigFile["<span style='color:#000'>agent_config.py<br/>AgentConfig, DSPy settings</span>"]
        LlmFactory["<span style='color:#000'>llm_factory.py<br/>create_dspy_lm()</span>"]
        InferenceAuth["<span style='color:#000'>inference_auth.py<br/>inference_headers</span>"]
        SemanticRouterFile["<span style='color:#000'>semantic_router.py<br/>apply_semantic_routing, create_routed_lm</span>"]
        LmResponseCache["<span style='color:#000'>lm_response_cache.py<br/>TenantScopedLMCache</span>"]
        LmEndpointAvailability["<span style='color:#000'>lm_endpoint_availability.py<br/>LMEndpointAvailability, LMEndpointNotServing</span>"]
        Utils["<span style='color:#000'>utils.py<br/>ConfigUtils, create_default_config_manager</span>"]
        ApiMixin["<span style='color:#000'>api_mixin.py<br/>ConfigAPIMixin (FastAPI endpoints)</span>"]
        Bootstrap["<span style='color:#000'>bootstrap.py<br/>BootstrapConfig</span>"]
        VespaStore["<span style='color:#000'>VespaConfigStore<br/>(in cogniverse_vespa)</span>"]
    end

    subgraph TelemetryFiles["<span style='color:#000'><b>telemetry/ files</b></span>"]
        TelManager["<span style='color:#000'>manager.py<br/>TelemetryManager - central API</span>"]
        TelConfig["<span style='color:#000'>config.py<br/>TelemetryConfig, TelemetryLevel</span>"]
        Registry["<span style='color:#000'>registry.py<br/>TelemetryRegistry</span>"]
        Context["<span style='color:#000'>context.py<br/>search_span, encode_span helpers</span>"]
        ProvidersDir["<span style='color:#000'><b>providers/</b><br/>base.py - TraceStore, AnnotationStore, etc.</span>"]
    end

    subgraph CachingFiles["<span style='color:#000'><b>caching/ files</b></span>"]
        TenantLru["<span style='color:#000'>tenant_lru.py<br/>TenantLRUCache</span>"]
        RefreshingCacheFile["<span style='color:#000'>refreshing_cache.py<br/>RefreshingCache</span>"]
    end

    subgraph DspyFiles["<span style='color:#000'><b>dspy/ files</b></span>"]
        LenientAdapter["<span style='color:#000'>lenient_json_adapter.py<br/>LenientJSONAdapter</span>"]
        StructuredAdapter["<span style='color:#000'>structured_json_adapter.py<br/>StructuredJSONAdapter, signature_response_format</span>"]
        ModelFormat["<span style='color:#000'>model_format.py<br/>bare_model_name, ensure_provider_prefix</span>"]
    end

    subgraph CommonFiles["<span style='color:#000'><b>common/ files</b></span>"]
        TenantUtilsFile["<span style='color:#000'>tenant_utils.py<br/>SYSTEM_TENANT_ID, require_tenant_id, etc.</span>"]
        DspyModuleRegistryFile["<span style='color:#000'>dspy_module_registry.py<br/>DSPyModuleRegistry, DSPyOptimizerRegistry</span>"]
        ArgoClientFile["<span style='color:#000'>argo_client.py<br/>build_argo_async_client</span>"]
    end

    subgraph SeparatePackages["<span style='color:#000'><b>Separate Packages</b></span>"]
        SDKInterface["<span style='color:#000'><b>cogniverse_sdk/interfaces/</b><br/>config_store.py<br/>ConfigStore ABC, ConfigScope</span>"]
        VespaImpl["<span style='color:#000'><b>cogniverse_vespa/config/</b><br/>config_store.py<br/>VespaConfigStore implementation</span>"]
    end

    Foundation --> ConfigDir
    Foundation --> TelemetryDir
    Foundation --> RegistryDir
    Foundation --> CachingDir
    Foundation --> DspyDir
    Foundation --> CommonDir
    Foundation --> ConfidencePy
    Foundation --> InitPy

    ConfigDir --> ConfigFiles
    TelemetryDir --> TelemetryFiles
    CachingDir --> CachingFiles
    DspyDir --> DspyFiles
    CommonDir --> CommonFiles
    RegistryDir -.-> Registry

    style Foundation fill:#a5d6a7,stroke:#388e3c,color:#000
    style ConfigDir fill:#ffcc80,stroke:#ef6c00,color:#000
    style TelemetryDir fill:#ce93d8,stroke:#7b1fa2,color:#000
    style RegistryDir fill:#b0bec5,stroke:#546e7a,color:#000
    style CachingDir fill:#81d4fa,stroke:#0288d1,color:#000
    style DspyDir fill:#81d4fa,stroke:#0288d1,color:#000
    style CommonDir fill:#81d4fa,stroke:#0288d1,color:#000
    style ConfidencePy fill:#90caf9,stroke:#1565c0,color:#000
    style InitPy fill:#90caf9,stroke:#1565c0,color:#000
    style ConfigFiles fill:#ffcc80,stroke:#ef6c00,color:#000
    style TelemetryFiles fill:#ce93d8,stroke:#7b1fa2,color:#000
    style CachingFiles fill:#81d4fa,stroke:#0288d1,color:#000
    style DspyFiles fill:#81d4fa,stroke:#0288d1,color:#000
    style CommonFiles fill:#81d4fa,stroke:#0288d1,color:#000
    style SeparatePackages fill:#90caf9,stroke:#1565c0,color:#000
    style Manager fill:#ffb74d,stroke:#ef6c00,color:#000
    style UnifiedConfigFile fill:#ffb74d,stroke:#ef6c00,color:#000
    style AgentConfigFile fill:#ffb74d,stroke:#ef6c00,color:#000
    style LlmFactory fill:#ffb74d,stroke:#ef6c00,color:#000
    style InferenceAuth fill:#ffb74d,stroke:#ef6c00,color:#000
    style ArgoClientFile fill:#ffb74d,stroke:#ef6c00,color:#000
    style SemanticRouterFile fill:#ffb74d,stroke:#ef6c00,color:#000
    style Utils fill:#ffb74d,stroke:#ef6c00,color:#000
    style ApiMixin fill:#ffb74d,stroke:#ef6c00,color:#000
    style Bootstrap fill:#ffb74d,stroke:#ef6c00,color:#000
    style VespaStore fill:#ffb74d,stroke:#ef6c00,color:#000
    style TelManager fill:#ba68c8,stroke:#7b1fa2,color:#000
    style TelConfig fill:#ba68c8,stroke:#7b1fa2,color:#000
    style Registry fill:#ba68c8,stroke:#7b1fa2,color:#000
    style Context fill:#ba68c8,stroke:#7b1fa2,color:#000
    style ProvidersDir fill:#ba68c8,stroke:#7b1fa2,color:#000
    style TenantLru fill:#64b5f6,stroke:#1565c0,color:#000
    style LenientAdapter fill:#64b5f6,stroke:#1565c0,color:#000
    style ModelFormat fill:#64b5f6,stroke:#1565c0,color:#000
    style TenantUtilsFile fill:#64b5f6,stroke:#1565c0,color:#000
    style DspyModuleRegistryFile fill:#64b5f6,stroke:#1565c0,color:#000
    style SDKInterface fill:#64b5f6,stroke:#1565c0,color:#000
    style VespaImpl fill:#64b5f6,stroke:#1565c0,color:#000

Configuration System

ConfigManager

ConfigManager is the central configuration API. All configuration operations go through this class.

Key Features:

  • Multi-tenant configuration with tenant isolation
  • Version history tracking for all configurations
  • In-process caching: the system config is held in a RefreshingCache with system_config_refresh_s (default 5s) and system_config_max_staleness_s (default 60s), on the same terms as the scoped configs below: set_system_config holds what it wrote at once, another process's write (the runtime storing its deployment overrides) is served within 60s, and a store outage past that bound raises. Pinned inference URLs are laid over every value served. Per-tenant scoped configs (routing, telemetry, backend, agent, durable execution, tenant instructions) are held in a RefreshingCache (see Tenant-Scoped Caching below). A value older than scoped_config_refresh_s (default 5s) is still served while one background thread re-reads it, so a request never waits on that store read. A value that reaches scoped_config_max_staleness_s (default 60s) is never served: the caller reads the store itself, and a failed read raises rather than answering with defaults. A failed background refresh is logged and leaves the last value in place up to that bound. Setters on the same manager invalidate immediately. A write made by another process is served within 60s, and within about 5s plus one store read for a key read continuously. Concurrent reads of one key share one store read, and a completed setter cannot be overwritten by an older in-flight read. Profile add, update and delete read the stored backend config, not the held copy, so they never write back over another process's profile change.
  • ConfigUtils parses a discovered config.json once per file modification across concurrent callers and returns an isolated copy to each instance. Invalid JSON and file-access failures propagate with the file path instead of being treated as an empty configuration.
  • Tenant backend profiles inherit the shipped catalog and system-tenant stored profiles absent from that catalog. Non-empty tenant fields override the base; nested fields merge and the vlm_endpoint field remains system-owned. This merge also applies when no JSON backend section exists. Store read failures propagate, and callers receive isolated profile values. ConfigUtils(tenant_id, config_manager).backend_profiles() returns that merged catalog, reading only the backend configs.
  • Pluggable backend persistence via ConfigStore interface (VespaConfigStore)
from cogniverse_foundation.config.utils import create_default_config_manager

# Initialize config manager
config_manager = create_default_config_manager()

# Get global system configuration (no tenant_id argument — SystemConfig is deployment-wide)
system_config = config_manager.get_system_config()

# Set agent configuration (agent_config built as shown under Configuration Types → AgentConfig)
config_manager.set_agent_config(
    tenant_id="acme",
    agent_name="orchestrator_agent",
    agent_config=agent_config
)

Inference endpoint credentials

cogniverse_foundation.config.inference_auth.inference_headers(base_url) returns immutable headers for a canonical HTTP(S) root URL. *.modal.run endpoints require HTTPS and COGNIVERSE_INFERENCE_API_KEY; other URLs get no headers. endpoint_root(url) reduces any endpoint URL (.../v1, .../v1/chat/completions) to that root.

The runtime API and ingestion worker parse INFERENCE_SERVICE_URLS with cogniverse_runtime.inference_services.parse_inference_service_urls. The runtime API persists the parsed endpoints into SystemConfig; the ingestion worker pins them on its ConfigManager with pin_inference_service_urls, so its reads carry them without writing to the store.

API Reference:

Method Description
get_system_config() Get global system configuration (deployment-wide, not per-tenant)
set_system_config(system_config) Set global system configuration
pin_inference_service_urls(service_urls) Serve service_urls as SystemConfig.inference_service_urls on this manager's reads, without persisting them
get_agent_config(tenant_id, agent_name) Get agent configuration
set_agent_config(tenant_id, agent_name, agent_config) Set agent configuration
get_agent_config_history(tenant_id, agent_name, limit=10) Get version history after canonicalizing the tenant identifier
get_routing_config(tenant_id="your_org:production", service="gateway_agent") Get routing configuration
set_routing_config(routing_config, tenant_id=None, service="gateway_agent") Set routing configuration
get_durable_execution_config(tenant_id, service="optimization") Get durable-execution enablement (default off)
set_durable_execution_config(durable_config, tenant_id=None, service="optimization") Set durable-execution enablement
get_telemetry_config(tenant_id="your_org:production", service="telemetry") Get telemetry configuration
set_telemetry_config(telemetry_config, tenant_id=None, service="telemetry") Set telemetry configuration
get_backend_config(tenant_id="your_org:production", service="backend") Get backend configuration
get_stored_backend_config(tenant_id, service="backend") Backend configuration as the store holds it now, past the held copy another process's write may not have reached yet; for writes that decide on what the tenant has stored
set_backend_config(backend_config, tenant_id=None, service="backend") Set backend configuration
get_tenant_instructions_config(tenant_id) Get raw tenant instructions value (TTL-cached; {"text": ..., "updated_at": ...} or None)
get_backend_profile(profile_name, tenant_id="your_org:production", service="backend") Get specific backend profile
add_backend_profile(profile, tenant_id="your_org:production", service="backend", *, replace=True) Add/update a backend profile; returns BackendProfileWrite(profile, version), the version of the tenant's backend config this write produced (for an identical re-add that writes nothing, the version already holding it). replace=False raises BackendProfileExistsError when the name is already stored, checked in the same compare-and-set as the write
update_backend_profile(profile_name, overrides, base_tenant_id=SYSTEM_TENANT_ID, target_tenant_id=None, service="backend") Partial profile update; inherits from base_tenant_id, saves to target_tenant_id, and returns BackendProfileWrite with the merged profile and the target config version the update produced. Raises BackendProfileNotFoundError when the profile is not stored for base_tenant_id, checked in the same compare-and-set as the write
list_backend_profiles(tenant_id="your_org:production", service="backend") List all backend profiles
delete_backend_profile(profile_name, tenant_id="your_org:production", service="backend") Delete backend profile; False when none was stored

The three profile writes rewrite the tenant's whole backend config, and every runtime process and replica writes it. Each is a compare-and-set read-modify-write through ConfigStore.update_config: the change is applied to the config as stored, and re-applied to the newer one whenever another writer lands first, so no process's profile change is overwritten. A change that leaves the stored config as it was writes no new version (startup's reaffirm_system_profiles on every worker writes once). Every profile write, one that wrote nothing included, drops the manager's held copy of that tenant's backend config, so its next read serves what it found stored. A write that loses every attempt raises ConfigWriteConflictError, and a store failure raises; either way nothing is written.

Searches read profile writes from the store: the shared search backend resolves the querying tenant's profiles per request through get_config, so a write is searchable on the writing process at once and on every other process and replica within scoped_config_max_staleness_s (60 s by default, about scoped_config_refresh_s for a tenant searched continuously). Nothing pushes profile changes into backend instances. | get_config_value(tenant_id, scope, service, config_key, default=None) | Get arbitrary config value by scope | | set_config_value(tenant_id, scope, service, config_key, config_value) | Set arbitrary config value by scope | | get_all_configs(tenant_id, scope=None) | Get all configs for a tenant, optionally filtered by scope | | export_configs(tenant_id, output_path) | Export all configs to JSON with a timezone-aware UTC exported_at | | get_stats() | Get configuration statistics |

Configuration Types

SystemConfig - Infrastructure settings:

from cogniverse_foundation.config.unified_config import SystemConfig

system_config = SystemConfig(
    search_backend="vespa",
    backend_url="http://localhost",
    backend_port=8080
)

AgentConfig - Agent-specific settings:

from cogniverse_foundation.config.agent_config import (
    AgentConfig, ModuleConfig, OptimizerConfig,
    DSPyModuleType, OptimizerType
)

agent_config = AgentConfig(
    agent_name="orchestrator_agent",
    agent_version="1.0.0",
    agent_description="Routes queries to appropriate agents",
    agent_url="http://localhost:8001",
    capabilities=["routing", "query_analysis", "entity_extraction", "conversation_memory"],
    skills=[{"name": "route_query", "description": "Route to best agent"}],
    module_config=ModuleConfig(
        module_type=DSPyModuleType.CHAIN_OF_THOUGHT,
        signature="query -> routing_decision"
    ),
    optimizer_config=OptimizerConfig(
        optimizer_type=OptimizerType.MIPRO_V2,
        num_trials=50
    ),
    llm_model="gpt-4",
    llm_temperature=0.7,
    llm_max_tokens=2000
)

BackendConfig - Backend and profile settings:

from cogniverse_foundation.config.unified_config import BackendConfig, BackendProfileConfig

backend_config = BackendConfig(
    tenant_id="acme",
    backend_type="vespa",
    url="http://localhost",
    port=8080,
    profiles={
        "video_colpali_mv_frame": BackendProfileConfig(
            profile_name="video_colpali_mv_frame",
            type="video",
            embedding_model="colpali",
            schema_name="video_colpali_smol500_mv_frame",
            pipeline_config={"chunk_strategy": "frame"},
            strategies={"default_top_k": 10}
        )
    }
)

model_loader names the loader ingestion embeds with and process_type (one of PROCESS_TYPES: direct_video, frame_based, video_chunks) the processing type ingestion would otherwise infer. extra_config holds every other profile key (inference_services, model_config, result_granularity, semantic_model, ...); to_dict() writes it beside the named fields and from_dict() collects any unnamed key into it.

Config sections (config/sections.py) - the configs an operator edits as forms:

from cogniverse_foundation.config.sections import CONFIG_SECTIONS, ConfigValueError

routing = CONFIG_SECTIONS["routing"]
schema = routing.schema()            # the form's JSON schema
current = routing.default("acme:prod", "gateway_agent")
edited = routing.from_form({"routing_mode": "direct"}, current)
stored = routing.dump(edited, "acme:prod")   # what ConfigManager stores

Each ConfigSection is a config dataclass (model) and where it is stored (scope, config_key, a fixed service or one per entry, tenant or system), with load/dump between the stored dict and the dataclass. schema() is the dataclass's JSON schema without the fixed_fields the location sets, with secret_fields marked writeOnly; form_value() dumps a config for the form with secrets null. from_form(value, current) applies the submitted fields to current: a field left out keeps its value, a null secret keeps it and "" clears it, and a key the schema does not know or a value the dataclass refuses raises ConfigValueError naming each, as does a changed value outside a field's choices (search_backend, environment, routing_mode, the telemetry provider from the registered providers) or ranges (backend_port, 1 to 65535); schema() carries them as enum, minimum and maximum, and a stored value outside them is kept when a save leaves it unchanged. Sections: system (SystemConfig), routing (RoutingConfigUnified), telemetry (TelemetryConfig), agent (AgentConfig, one per agent) and durable_execution (DurableExecutionConfig); section_for(scope, config_key) finds the section of a stored entry.

ConfigManager.compare_and_set_entry(tenant_id, scope, service, config_key, value, expected_version=...) stores a value as the next version when the stored one is expected_version (0: none) and returns None when another write landed first; forget_held_configs(tenant_id) drops what the manager holds for a tenant (the system config for _system). The module function forget_held_backend_configs(tenant_id) drops the tenant's backend config from every ConfigManager in the process; the runtime runs it on every runtime and ingestion worker when a profile is written. forget_held_tenant_configs(tenant_id) drops everything every ConfigManager in the process holds for a tenant (the system config for _system); the runtime runs it on every runtime and ingestion worker when a config is saved, restored or imported. SystemConfig.to_dict() shows llm_api_key as "***"; to_dict(redact=False) is the stored form.

Profile servability - whether a tenant can be served a profile:

from cogniverse_foundation.config.unified_config import (
    PROFILE_EMBEDDING_SERVICE_UNCONFIGURED,
    PROFILE_SCHEMA_NOT_DEPLOYED,
    PROFILE_SERVABLE,
    profile_base_schema_name,
    profile_is_servable,
    profile_servability,
)

state = profile_servability(
    "document_text_semantic",
    profile.to_dict(),
    system_config.inference_service_urls,
    deployed_schemas={"document_text"},
)
assert state == PROFILE_SERVABLE

profile_servability names why a profile can or cannot be served: its embedding service must resolve to a URL in inference_service_urls (PROFILE_EMBEDDING_SERVICE_UNCONFIGURED otherwise) and its base schema — profile_base_schema_name, the profile's schema_name or its own name — must be in deployed_schemas (PROFILE_SCHEMA_NOT_DEPLOYED otherwise). profile_is_servable is the boolean form. The deployed set is per tenant and comes from cogniverse_core.registries.schema_registry.tenant_deployed_schema_names; callers in cogniverse_agents.profile_selection_agent compose the two, so this module keeps no dependency on core.

tenant_profile_servability(config_manager, tenant_id) there lists the tenant's profiles with their state: the profiles it stored and, for each schema it has deployed that none of those reads, the catalog's search profiles (video, document, image, audio, code, wiki) declaring that schema. A registered tenant stores no profile but has the built-in video_colpali_smol500_mv_frame schema deployed, so that profile is servable for it, alongside any profile it adds. Each row carries the catalog definition from ConfigUtils.backend_profiles(). servable_tenant_profiles and tenant_usable_profile_names (which raises when none is servable) filter it to the servable rows; GET /search/profiles, answer grounding, profile selection, synthetic profile data and the optimizer read the tenant's profiles through them.

FieldMappingConfig - Canonical content fields used by synthetic generation:

from cogniverse_foundation.config.unified_config import FieldMappingConfig

field_mappings = FieldMappingConfig()
assert field_mappings.topic_fields == [
    "video_title",
    "audio_title",
    "image_title",
    "document_title",
    "chunk_name",
    "title",
]
assert field_mappings.description_fields == [
    "segment_description",
    "image_description",
    "full_text",
    "source_code",
    "content",
    "description",
]
assert field_mappings.transcript_fields == ["audio_transcript", "transcript"]
assert field_mappings.entity_fields == ["video_title", "segment_description"]
assert field_mappings.temporal_fields == {
    "start": "start_time",
    "end": "end_time",
}
assert field_mappings.metadata_fields == {}

Override these mappings only when a deployed schema uses different canonical fields. The defaults preserve titles, descriptive text, transcripts, entity sources, temporal bounds, and metadata for the configured video, audio, image, document, code, and wiki schemas. FieldMappingConfig.from_dict({}) hydrates these same defaults; unknown keys raise instead of being ignored.

SyntheticGeneratorConfig - Synthetic generation timeout and floor:

from cogniverse_foundation.config.unified_config import SyntheticGeneratorConfig

generator_config = SyntheticGeneratorConfig(
    tenant_id="acme",
    synthetic_generation_timeout_seconds=300.0,
    synthetic_generation_floor_count=1,
)

The timeout bounds production callback work. The floor bounds how far a synthetic run may fall back when the candidate pool is exhausted. Both fields validate as positive values and are required in the deployable synthetic config.

RoutingConfigUnified - Routing agent settings:

from cogniverse_foundation.config.unified_config import RoutingConfigUnified

routing_config = RoutingConfigUnified(
    tenant_id="acme",
    routing_mode="tiered",  # "tiered", "ensemble", "hybrid"
    enable_fast_path=True,
    fast_path_confidence_threshold=0.4,
)

# Per-optimizer floors can override the global 100/3 defaults when a tenant
# needs calibrated promotion thresholds for specific optimization modes.
# routing_config.optimizer_floors = {
#     "profile_selection": {
#         "min_samples_for_optimization": 20,
#         "min_unique_queries": 6,
#     },
#     "entity_extraction": {
#         "min_samples_for_optimization": 58,
#         "min_unique_queries": 15,
#     },
# }

DurableExecutionConfig - Durable-execution enablement for long-running optimization/eval workflows:

from cogniverse_foundation.config.unified_config import DurableExecutionConfig

durable_config = DurableExecutionConfig(
    tenant_id="acme",
    enabled=True,  # checkpoint + resume the triggered-optimization job (default False)
)

TelemetryConfig - Telemetry settings:

from cogniverse_foundation.telemetry.config import TelemetryConfig, TelemetryLevel

telemetry_config = TelemetryConfig(
    enabled=True,
    level=TelemetryLevel.DETAILED,
    otlp_enabled=True,
    otlp_endpoint="localhost:4317"
)

LLMEndpointConfig - Single LLM endpoint wiring:

from cogniverse_foundation.config.unified_config import LLMEndpointConfig

endpoint = LLMEndpointConfig(
    model="openai/gpt-4o",          # always provider-prefixed (DSPy/LiteLLM convention)
    api_base="http://localhost:8101/v1",
    api_key=None,                    # None for local OAI-compat servers
    temperature=0.1,
    max_tokens=1000,
    request_timeout=120.0,           # seconds before giving up on a slow endpoint
    num_retries=1,
    seed=42                          # optional; enables bit-stable output on vLLM
)

Field Default Description
model required Provider-prefixed model string, e.g. "openai/gpt-4o"
api_base None Override endpoint URL
api_key None None for keyless local servers
temperature 0.1 Sampling temperature
max_tokens 1000 Max output tokens
context_window None Tokens the endpoint accepts per request; used only when the endpoint publishes no max_model_len
request_timeout 120.0 Per-request timeout in seconds
num_retries 1 Retry count on transient errors
seed None vLLM sampling seed for reproducibility
adapter_path None LoRA/fine-tuned artifact path (bookkeeping only; not read by create_dspy_lm)
extra_body None Provider-specific request params
extra_headers None Static HTTP headers sent with every request (e.g. semantic-router tenant/tier headers)

LLMConfig - Multi-role LLM configuration:

from cogniverse_foundation.config.unified_config import LLMConfig, LLMEndpointConfig

llm_config = LLMConfig(
    primary=LLMEndpointConfig(model="openai/gpt-4o", api_base="http://localhost:8101/v1"),
    teacher=LLMEndpointConfig(model="openai/gpt-4o-mini", api_base="http://localhost:8101/v1"),
    overrides={
        # Per-component partial overrides merged onto primary at resolve time
        "summarizer_agent": {"max_tokens": 2000},
    }
)

# Resolve the effective config for a component
resolved = llm_config.resolve("summarizer_agent")

primary is the global default for all DSPy modules and also the student model during optimization. teacher is optional for non-optimization processes, but every teacher-dependent optimization must configure it explicitly: resolve_teacher() returns an isolated copy for BootstrapFewShot(teacher_settings={"lm": ...}) and raises when the role is absent instead of falling back to primary. overrides holds per-component partial dicts — only differing fields need to be specified; resolve(component) merges them field-by-field onto a copy of primary (never through to_dict(), which masks api_key — the resolved endpoint keeps the real key).

is_modal_inference_url(base_url) classifies canonical inference roots: it returns True only for root HTTPS *.modal.run URLs, and inference_headers(base_url) uses that predicate to return the shared Modal bearer only for those roots. Non-Modal roots remain keyless.

create_dspy_lm(config: LLMEndpointConfig) -> dspy.LM (cogniverse_foundation.config.llm_factory) is the single chokepoint every dspy.LM() construction in the codebase goes through. It wires api_base/api_key/temperature/max_tokens/timeout/num_retries onto the LM, merges seed into extra_body, forwards extra_headers, and resolves api_key through resolve_inference_api_key(api_base, api_key). Raises ValueError if config.model is empty.

create_dspy_lm(config, tenant_id=...) binds responses to a canonical tenant in TenantScopedLMCache (config/lm_response_cache.py). Its key includes the complete request, including model, messages, response format, tools, sampling and routing headers; only the API key and cache-control flag are excluded. DSPy caching is disabled, and an unbound LM does not cache. BodyBoundedLM.for_tenant(tenant_id) binds the ambient LM per request.

The process cache loads semantic_router.response_cache_ttl_seconds and response_cache_max_entries from the config file ConfigUtils._discover_config_file() resolves, on first use; no config file raises. Entries expire on a monotonic clock and evict in LRU order. Threads and event loops share one upstream call per key. Hits return independent copies with cache_hit=True and zero additional usage. Failed calls retain the provider exception with tenant, model and key context, store nothing, and allow the next request to retry. Invalid cache configuration raises with its file path.

BodyBoundedLM holds every call's max_tokens within the completion budget bound by cogniverse_foundation.config.lm_output_budget.bound_output_token_budget(budget): the smaller of that budget and the endpoint's configured max_tokens is sent, and the response cache keys on the value sent. output_token_budget_from(context) reads context["max_output_tokens"], returns None when absent and raises ValueError for anything but a positive integer.

create_budgeted_dspy_lm(config: LLMEndpointConfig) -> dspy.LM builds the same LM as a BudgetedLM (cogniverse_foundation.config.budgeted_lm), which fits every request inside the window its endpoint serves. On first call it resolves the window through resolve_context_window(api_base, declared=config.context_window, model=...) and holds it as a TokenBudget(model, context_window, reserved_output=config.max_tokens); input_budget is context_window - reserved_output. What the endpoint publishes as max_model_len at {api_base}/models wins (ResolvedContextWindow(tokens, source="served")), because the window belongs to the deployment rather than to the model; a listing that carries none falls back to config.context_window (source="declared"). Requests over that allowance shed whole few-shot demonstrations, oldest first, and log how many were dropped. A prompt that still overflows with none left raises PromptBudgetExceededError naming the window, the reservation, the measured input and the number dropped — as does a provider-side context-window rejection, so a caller never sees a bare litellm error. An endpoint that publishes no max_model_len and declares no context_window, or one with no api_base or no max_tokens, raises ContextWindowUnavailableError instead of being budgeted by guess. Used wherever a prompt can approach the window: the DSPy bootstrap teacher, the entity-extraction optimizer, the ingestion LM context (ingest_lm_context_for) and the ingestion worker's default LM. resolve_context_window memoizes per (api_base, model) for the process (32 entries, oldest evicted; clear_context_window_memo() drops them), so ingestion paths that build one LM per segment read the listing once. fitting_demonstrations(demos, budget=, render=, count_tokens=, calls=) sizes a compiled program before it is sent at all: the largest number of demos, kept from the front, whose request render(demos, call) fits budget.input_budget for every one of calls (zero when a call overflows with none). Calls are taken longest request first; each is counted once at the longest prefix every earlier call fit, and one that overflows it bisects below it, so a counter that pays a round-trip per count is called about once per call. served_message_counter(api_base, model) is such a counter: it counts messages through the endpoint's POST {root}/tokenize (vLLM) with the bare served model name, so the count is the served tokenizer and chat template's, the prompt_tokens the completion reports. litellm counts a model it has no tokenizer for with OpenAI's cl100k_base, which undercounts the gemma student by about a quarter on transcript text. A count is idempotent and a fitting pass asks hundreds, so an unreachable endpoint or a 5xx answer is asked again up to retries times (1 s, 2 s, 4 s ... apart; the optimizer passes the endpoint's num_retries, as its LM calls are retried). An endpoint that still cannot answer, or answers 4xx or without an integer count, raises TokenCountUnavailableError.

resolve_inference_api_key(api_base, api_key) is the one key-resolution rule for every OpenAI-compatible client, dspy.LM or not (Mem0's LLM provider, litellm.completion/rerank, the LLM/visual judges' chat-completions POSTs): an explicit key wins; otherwise the COGNIVERSE_INFERENCE_API_KEY bearer is sent whenever it is set (an in-cluster router may forward to Modal, and a self-hosted server ignores it); a *.modal.run api_base with no bearer raises naming the variable; a self-hosted api_base with nothing configured gets "not-required"; with no api_base the configured key passes through untouched.

from cogniverse_foundation.config.llm_factory import create_dspy_lm
from cogniverse_foundation.config.unified_config import LLMEndpointConfig

lm = create_dspy_lm(LLMEndpointConfig(model="openai/gpt-4o", api_base="http://localhost:8101/v1"))

SemanticRouterConfig - Opt-in routing of LLM calls through a vLLM Semantic Router:

from cogniverse_foundation.config.unified_config import SemanticRouterConfig

router_config = SemanticRouterConfig(
    enabled=True,
    semantic_router_url="http://semantic-router:8801/v1",
    routed_model="openai/auto",
    classification_model="openai/cogniverse-classification",
    vision_model="openai/cogniverse-vision",
    short_reasoning_model="openai/cogniverse-short-reasoning",
)

When enabled, cogniverse_foundation.config.semantic_router rewrites an LLMEndpointConfig to target semantic_router_url instead of the model backend, sets model to the router entry the call site takes (the router resolves models by its own catalog, aliases and entrypoints, not raw provider ids), and attaches two authz headers per request: tenant identity (user_id_header, default x-authz-user-id) and tenant tier (tier_header, default x-authz-user-groups, the caller's resolved RouterTier). When disabled, the endpoint passes through unchanged. The tier is the tenant's stored attribute: cogniverse_foundation.config.tenant_tiers holds it in the config store under scope ROUTING / service semantic_router / key tenant_tier, resolve_tenant_tier(config_accessor, tenant_id) reads it through a per-tenant RefreshingCache (refreshed off the request thread after TENANT_TIER_REFRESH_S, 15 s; never served past TENANT_TIER_MAX_STALENESS_S, 30 s) invalidated by every write in the process, an unset tenant is DEFAULT_ROUTER_TIER, and a store failure routes as DEFAULT_ROUTER_TIER with a WARNING. Claim extraction during ingestion keeps the direct primary endpoint.

Each call site names its router entry. CLASSIFICATION_CALL_SITES (entity extraction, gateway, orchestrator, profile selection, query enhancement, search) produce a bounded output and send classification_model, which names the chart router's cogniverse-classification entrypoint: its recipe chooses the decision from the tenant tier alone, so no domain classifier runs; every tier's bounded call is served by basic-chat (it never crosses to the teacher), and the decision's exact response cache still applies. SHORT_REASONING_CALL_SITES (deep-research decomposition and evidence evaluation) send short_reasoning_model, which names the cogniverse-short-reasoning entrypoint: its recipe also tests the tier alone and serves pro-reasoning with reasoning on for pro and basic-chat with reasoning off for free and base, with the same exact response cache. FREE_FORM_CALL_SITES send routed_model (the auto alias), where the router classifies the content and may promote the call to the reasoning model; a call site in neither set takes auto. Every decision in every routing profile caches with mode: exact: on the router's embedding model the closest different-content pair in the evaluation corpus scores higher than the weakest equivalent pair, so no similarity threshold is admissible.

When the router is enabled but inference.vllm_llm_teacher is neither enabled nor external, the chart serves pro-reasoning from the student's Envoy cluster and provider model. Helm NOTES and the rendered manifest name pro_model_unavailable; each pro decision, including short-reasoning-pro, carries tier_degraded: pro_model_unavailable. A served teacher uses its own cluster and has neither warning nor degradation marker.

Function Description
resolve_semantic_router_headers(config, tenant_id) Resolve the two authz headers, or None when disabled
apply_semantic_routing(endpoint, config, tenant_id, tier, call_site) Return a routed copy of endpoint, or the original when disabled
routed_model_for(config, call_site) classification_model for a call site in CLASSIFICATION_CALL_SITES, short_reasoning_model for one in SHORT_REASONING_CALL_SITES, routed_model otherwise
create_routed_lm(endpoint, config, tenant_id, tier, call_site) apply_semantic_routing + the shared LM construction; returns a RoutedLM (cogniverse_foundation.config.routed_lm), which sends a call carrying image parts on vision_model (the chart's cogniverse-vision entry, serving the multimodal student for every tier) and records the completion's model on the current span as llm.served_model (LLM_SERVED_MODEL_ATTRIBUTE)
record_served_model(response) Stamp a completion's model on the current span; a no-op outside any span
ingest_lm_context_for(endpoint) Return a direct dspy.context for ingestion-time LM calls (claim extraction); never routed
routed_lm_context_for(config_manager, tenant_id, agent_name, endpoint=None, call_site=None) Return a dspy.context binding the routed (or direct) LM for query-time agents, tenant-bound either way — the entry point agents use; call_site (default agent_name) names the router entry
routed_lm_context_for_async(config_manager, tenant_id, agent_name, endpoint=None, call_site=None) Await the same context off the event loop — the entry point every coroutine uses, since the tier resolution behind it is a config-store read; the returned context manager is unentered so the caller binds it on its own task
resolve_semantic_router_config(config_accessor) Read SemanticRouterConfig off an object exposing get_semantic_router()

A failed routed completion raises a RoutedLMCallFailed subclass from cogniverse_foundation.config.routed_lm, chained from the litellm error and carrying status, router_code (the provider's error code), tenant_id, tier and routed_model. The router passes the upstream status through, so the class follows it:

Status Raised Retried
401, 403 UpstreamAuthRejected never
429 UpstreamRateLimited up to num_retries
5xx, reset, timeout UpstreamUnavailable up to num_retries
404 UpstreamNotServing (an UpstreamUnavailable) never
any other 4xx RouterDecodeFailed never

RoutedLM spends the endpoint's num_retries itself, only on UpstreamRateLimited and UpstreamUnavailable other than UpstreamNotServing; errors that are not provider or transport failures propagate unchanged. A timeout or refused connection carries the status litellm stamps on it (408 / 500) and is an UpstreamUnavailable regardless.

A 404 means the upstream has nothing deployed for the routed model: an undeployed Modal app answers modal-http: invalid function call, and the router passes the status through (x-vsr-response-path: upstream). UpstreamNotServing carries failed_fast and recheck_in_s (below).

LM endpoint availability

cogniverse_foundation.config.lm_endpoint_availability records, per process, the last outcome of every LM endpoint called — keyed by LMEndpoint(api_base, model, route), where route is the router tier on the routed path. BodyBoundedLM, which every LM (RoutedLM, BudgetedLM, create_dspy_lm) sends through, consults it on each call that misses the response cache:

Last outcome State Next call
a completion serving sent
HTTP 404 not_serving refused without being sent for NOT_SERVING_RECHECK_S (30s), raising LMEndpointNotServing(failed_fast=True); then one call is sent as the recheck while the others keep failing fast, and its outcome replaces the state
a 5xx, a timeout, a refused connection failing sent — a scaled-to-zero endpoint answers after its cold start, so only a 404 fails fast

A recheck still unanswered after PROBE_VERDICT_S (5s, five times the slowest measured 404 from an undeployed Modal app) is a cold start, and the calls behind it are sent too. A recheck that ends without an outcome (cancelled, past its caller's deadline) hands the recheck to the next call. The direct path raises LMEndpointNotServing (chained from the litellm NotFoundError on the call that got the 404); RoutedLM raises UpstreamNotServing chained from it. not_serving_cause(exc) finds either through __cause__, and refusal(endpoint) answers the fast failure a call would get without admitting one, for a caller deciding whether to prepare the call at all. Each not-serving call stamps llm.endpoint.state, llm.endpoint.failed_fast and llm.endpoint.recheck_in_s on its span (LLM_ENDPOINT_*_ATTRIBUTE). lm_endpoint_availability().snapshot() lists every endpoint's state, upstream_status, failure, reason, observed_at and recheck_in_s (credentials in the address stripped); the runtime's /health serves it. Nothing is probed in the background: any request to a deployed Modal endpoint boots a GPU container. At most MAX_TRACKED_ENDPOINTS (64) endpoints are tracked, least recently observed evicted first.

A call made under a caller's deadline never outlives it. The caller binds an LMCallDeadline (cogniverse_foundation.config.lm_deadline, bound_lm_call_deadline); BodyBoundedLM sends nothing once the deadline has passed or the caller abandoned it and raises LMCallDeadlineExceeded, a TimeoutError naming the endpoint, the model and the deadline. An OpenAI-compatible request goes through deadline_bound_openai_client(api_base, api_key), whose connects, writes and reads each wait at most the time left at that moment, and a caller joining an identical in-flight call waits no longer than its own deadline. RoutedLM starts no retry or student attempt past the deadline, and a failure under a deadline adds endpoint= and deadline_s= to its message. A request the client hangs up on because the deadline passed or was abandoned raises LMCallDeadlineExceeded chained from the provider's timeout: it does not mark the endpoint failing, and the response cache names it at INFO, since the caller reports the overrun. A timeout or refused connection with no deadline behind it is the provider's own error, recorded as failing and, for a tenant-bound LM, logged at ERROR by the response cache; the "LM rejected the request" error is logged only for a status the endpoint actually answered.

A pro free-form or short-reasoning call whose teacher raises UpstreamUnavailable with status 502, 503, 504, a timeout (408), a transport failure (500 or no status), or UpstreamNotServing (nothing deployed) gets one attempt on classification_model, which serves the student. The LM completion and current span carry tier_degraded: pro_model_unavailable, upstream_status, and upstream_exception_type. The exception type and its status determine eligibility; message text never does. A permanent refusal propagates, and default, bounded-output, and vision calls carry no degradation marker. If the student fails too, its typed failure propagates without another attempt. tier_degradation_context() collects these fields for one dispatch, including calls on worker threads, and restores the enclosing scope on exit.

Configuration Scopes

Configurations are organized by scope for isolation:

Scope Description Example Keys
SYSTEM Infrastructure settings backend_url, backend_port, summarizer_agent_url
AGENT Per-agent settings module_config, llm_model, llm_temperature
ROUTING Routing agent settings routing_mode, enable_fast_path, gliner_threshold
TELEMETRY Telemetry settings otlp_endpoint, otlp_enabled, level
SCHEMA Deployed Vespa schema tracking (used by cogniverse_core.registries.schema_registry) schema_name, deployment status
BACKEND Backend profiles embedding_model, schema_name, pipeline_config

VespaConfigStore.compare_and_set_config(tenant_id, scope, service, config_key, config_value, expected_version=...) writes the version after expected_version when that is still the latest. The required keyword expected_version must be nonnegative: zero requires an absent key, and negative values raise ValueError. The version is reserved on the key's version counter only while no other writer holds an unwritten reservation; one older than 30 seconds is treated as abandoned and reserved past. It returns the committed ConfigEntry after a strong read confirms it is the latest version, or None on a version mismatch, a live reservation, or a conditional-write conflict. Every successful version write applies the same keep_versions history retention as set_config, including a stale write that loses the final read; pruning is best-effort. Read/write failures raise. Schema deployment journals and registration completion use this operation to fence stale writers and deletion tombstones.

from cogniverse_sdk.interfaces.config_store import ConfigScope

# Get arbitrary config value by scope
value = config_manager.get_config_value(
    tenant_id="acme",
    scope=ConfigScope.AGENT,
    service="orchestrator_agent",
    config_key="optimizer_config"
)

Configuration Inheritance

The configuration system uses a layered inheritance model where tenant-specific settings override system defaults:

flowchart TB
    subgraph Sources["<span style='color:#000'>Configuration Sources</span>"]
        EnvVars["<span style='color:#000'>Environment Variables<br/>COGNIVERSE_CONFIG, etc.</span>"]
        ConfigFile["<span style='color:#000'>config.json<br/>Auto-discovered</span>"]
        VespaStore["<span style='color:#000'>Vespa Store<br/>Persisted configs</span>"]
    end

    subgraph Layers["<span style='color:#000'>Configuration Layers</span>"]
        direction TB
        SystemDefaults["<span style='color:#000'>System Defaults<br/>Hardcoded fallbacks</span>"]
        GlobalConfig["<span style='color:#000'>Global Configuration<br/>config.json profiles</span>"]
        TenantOverlay["<span style='color:#000'>Tenant Overlay<br/>Per-tenant overrides</span>"]
        RuntimeOverride["<span style='color:#000'>Runtime Override<br/>API/query-time params</span>"]
    end

    subgraph Resolution["<span style='color:#000'>Resolution Order (Bottom Wins)</span>"]
        Final["<span style='color:#000'>Final Configuration<br/>Merged result</span>"]
    end

    EnvVars --> GlobalConfig
    ConfigFile --> GlobalConfig
    VespaStore --> TenantOverlay

    SystemDefaults --> Final
    GlobalConfig --> Final
    TenantOverlay --> Final
    RuntimeOverride --> Final

    style Sources fill:#90caf9,stroke:#1565c0,color:#000
    style Layers fill:#ffcc80,stroke:#ef6c00,color:#000
    style Resolution fill:#a5d6a7,stroke:#388e3c,color:#000
    style EnvVars fill:#90caf9,stroke:#1565c0,color:#000
    style ConfigFile fill:#90caf9,stroke:#1565c0,color:#000
    style VespaStore fill:#90caf9,stroke:#1565c0,color:#000
    style SystemDefaults fill:#ffcc80,stroke:#ef6c00,color:#000
    style GlobalConfig fill:#ffcc80,stroke:#ef6c00,color:#000
    style TenantOverlay fill:#ffcc80,stroke:#ef6c00,color:#000
    style RuntimeOverride fill:#ffcc80,stroke:#ef6c00,color:#000
    style Final fill:#a5d6a7,stroke:#388e3c,color:#000

Configuration Resolution Example:

from cogniverse_foundation.config.unified_config import BackendConfig, BackendProfileConfig

# System default (hardcoded)
max_frames = 50

# Global config (config.json) - overrides default
config_json = {
    "profiles": {
        "video_colpali_mv_frame": {
            "max_frames": 100
        }
    }
}

# Tenant overlay (Vespa) - overrides global
config_manager.add_backend_profile(
    tenant_id="premium_tenant",
    profile=BackendProfileConfig(
        profile_name="video_colpali_mv_frame",
        type="video",
        pipeline_config={"max_frames": 200},
    ),
)

async def search(query: str, max_frames: int):
    ...  # Runtime override (query param) - overrides all

async def main():
    result = await search(query="cats", max_frames=300)
    # Final: premium_tenant gets max_frames=300 for this query
    return result

Resolution Priority (highest to lowest):

Priority Source Scope Example
1 (highest) Runtime Override Per-request Query params, API args
2 Tenant Overlay Per-tenant ConfigManager.set_*_config() (persisted to Vespa)
3 Global Config All tenants config.json profiles
4 (lowest) System Defaults Fallback Hardcoded in classes

ConfigAPIMixin

cogniverse_foundation.config.api_mixin.ConfigAPIMixin adds runtime, persisted configuration REST endpoints to an agent's FastAPI app. It expects the host class to expose self.agent_config (an AgentConfig) plus update_module_config() / update_optimizer_config() methods (provided by DynamicDSPyMixin in cogniverse_core).

class MyAgent(DynamicDSPyMixin, ConfigAPIMixin):
    def __init__(self, tenant_id, config_manager):
        self.initialize_dynamic_dspy(agent_config)
        app = FastAPI()
        self.setup_config_endpoints(app, config_manager, tenant_id=tenant_id)

setup_config_endpoints(app, config_manager, tenant_id=None) registers:

Route Description
GET /config Current AgentConfig as a dict
GET /config/module Current DSPy module info
POST /config/module Update module config (ModuleConfigUpdate); persists via ConfigManager.set_agent_config
GET /config/optimizer Current optimizer info
POST /config/optimizer Update optimizer config (OptimizerConfigUpdate); persists
POST /config/llm Update LLM fields (LLMConfigUpdate) and reconfigure the DSPy LM; persists
GET /config/modules/available List registered DSPy module types
GET /config/optimizers/available List registered DSPy optimizer types

Every mutating endpoint persists through the injected ConfigManager, so changes survive a restart and are versioned like any other config write. The tenant is fixed when the routes are registered; request query parameters cannot redirect a write to another tenant. Mutations are serialized, synchronous store calls run in a worker thread, and a persistence failure restores the prior in-memory configuration before the endpoint returns an error.


Telemetry System

TelemetryManager

TelemetryManager is a singleton that manages OpenTelemetry tracing with multi-tenant isolation.

Key Features:

  • Tenant-isolated tracer providers
  • LRU caching of tracers
  • Graceful degradation when telemetry unavailable
  • Session tracking for multi-turn conversations
  • Phoenix integration for trace visualization
from cogniverse_foundation.telemetry.manager import TelemetryManager, get_telemetry_manager

# Get global singleton
telemetry = get_telemetry_manager()

# Create span with tenant isolation
with telemetry.span("search.execute", tenant_id="acme") as span:
    span.set_attribute("query", "find videos about cats")
    # ... search logic ...

API Reference:

Method Description
span(name, tenant_id, project_name=None, attributes=None, component="agents") Create tenant-isolated span. component gates emission against TelemetryConfig.level (one of search_service/agents/backend/pipeline/encoder)
session_span(name, tenant_id, session_id, project_name=None, attributes=None, component="search_service") Span within a session context; all nested span() calls inherit session_id
get_tracer(tenant_id, project_name=None) Get tracer (prefer span())
get_provider(tenant_id, project_name=None) Get telemetry provider for queries (spans/annotations/datasets)
register_project(tenant_id, project_name, **kwargs) Register project with config
force_flush(timeout_millis=10000) Flush all pending spans
shutdown() Graceful shutdown
get_stats() Get telemetry statistics

There is no standalone session() context manager — session tracking is done by calling session_span() with the same session_id for each operation that should share a session, or by nesting span() calls inside one session_span().

Span Context

Creating spans with tenant isolation:

from cogniverse_foundation.telemetry.manager import get_telemetry_manager

telemetry = get_telemetry_manager()

async def process(query: str):
    # Basic span
    with telemetry.span("agent.process", tenant_id="acme") as span:
        span.set_attribute("agent.name", "orchestrator_agent")
        span.set_attribute("query.length", len(query))
        result = await process_query(query)

    # Span with project isolation (for management operations)
    with telemetry.span(
        "experiment.run",
        tenant_id="acme",
        project_name="experiments",  # Separate Phoenix project
        attributes={"experiment.name": "optimizer_v2"}
    ) as span:
        await run_experiment()

    return result

Session Tracking

Track multi-turn conversations across requests:

async def handle_request(query: str, session_id: str):
    # At API entry point - establish session context
    with telemetry.session_span(
        "api.search.request",
        tenant_id="acme",
        session_id=session_id,
        attributes={"query": query, "turn": 3}
    ) as span:
        # All child spans inherit session_id
        result = await search_service.search(query)
    return result

# Alternative: nest plain spans inside one session_span - all inherit session_id
with telemetry.session_span("session.start", tenant_id="acme", session_id="session-xyz"):
    with telemetry.span("operation1", tenant_id="acme") as span1:
        pass
    with telemetry.span("operation2", tenant_id="acme") as span2:
        pass
    # Both spans share session_id

Project Registration

Register projects with custom endpoints (useful for tests):

# Register with default config
telemetry.register_project(
    tenant_id="acme",
    project_name="search"
)

# Register with custom endpoints (for tests)
telemetry.register_project(
    tenant_id="test-tenant",
    project_name="synthetic_data",
    otlp_endpoint="http://localhost:24317",
    http_endpoint="http://localhost:26006",
    use_sync_export=True  # Sync export for tests
)

Span Context Helpers

cogniverse_foundation.telemetry.context provides pre-shaped span helpers that standardize the attribute set for the three most common span kinds, plus two helpers for enriching an existing span:

Function Description
search_span(tenant_id, query, top_k=10, ranking_strategy="default", profile="unknown", backend="vespa") Context manager for a search_service.search span (component="search_service")
encode_span(tenant_id, encoder_type, query_length=0, query="") Context manager for an encoder.<type>.encode span (component="encoder")
backend_search_span(tenant_id, backend_type="vespa", schema_name="unknown", ranking_strategy="default", top_k=10, has_embeddings=False, query_text="") Context manager for a search.execute span (component="backend")
add_search_results_to_span(span, results, output_value=None) Set num_results/output.value/top_score attributes and a search_results event; pass output_value to reuse one serialization across both spans
serialize_search_results(results) Serialize result rows to the canonical output.value JSON once (feed the result to both add_search_results_to_span calls of a search)
add_embedding_details_to_span(span, embeddings) Set embedding_shape/embedding_dtype/norm-mean/norm-std attributes from an embeddings array

Each context manager wraps TelemetryManager.span(), so it still goes through the tenant-required check and the component-based TelemetryConfig.level gating.


Registry System

The cogniverse_foundation/registry/ subpackage provides EntryPointRegistry[T] — a generic plugin registry over importlib.metadata entry points. Subclasses declare an entry-point group and a label; the base handles discovery, manual registration, conflict detection, tenant-scoped or config-keyed caching, and lifecycle-style initialization. Four registries in the codebase are thin subclasses of this base:

Registry Subpackage Entry-point group Tenant-scoped
TelemetryRegistry cogniverse_foundation.telemetry.registry cogniverse.telemetry.providers yes
EvaluationRegistry cogniverse_evaluation.providers.registry cogniverse.evaluation.providers yes
WorkflowStoreRegistry cogniverse_core.registries.workflow_store_registry cogniverse.workflow.stores no
AdapterStoreRegistry cogniverse_core.registries.adapter_store_registry cogniverse.adapter.stores no

Defining a new registry

from cogniverse_foundation.registry import EntryPointRegistry
from my_pkg.interfaces import MyStore


class MyStoreRegistry(EntryPointRegistry[MyStore]):
    _entry_point_group = "myapp.stores"
    _label = "my store"
    # _tenant_scoped = False (default): klass(**config), cached by backend_url/port
    # _tenant_scoped = True: klass() + .initialize({**config, tenant_id}), cached by tenant

Implementations register via their pyproject.toml:

[project.entry-points."myapp.stores"]
default = "my_pkg.stores.default_impl:DefaultStore"

Callers fetch instances with MyStoreRegistry.get(name="default", config={...}). Conflict detection is always on — if two installed packages both register name="default" under the same group, discover() raises ValueError rather than silently picking one. Discovery runs once per registry even when threads race the first lookup. For a tenant-scoped registry, get() canonicalizes its tenant_id argument and places that canonical value in the initialization mapping after caller config is merged, so a config payload cannot override the cache's tenant identity. TelemetryRegistry further keys providers by project plus HTTP and gRPC endpoints, so reconfiguring an existing project creates a provider bound to the new destinations.


Tenant-Scoped Caching

cogniverse_foundation.caching.TenantLRUCache is a thread-safe, bounded LRU cache keyed by tenant_id. Long-running processes keep per-tenant state (Mem0 memory managers, telemetry providers, backend clients, compiled DSPy modules) in memory; without a bound, a multi-tenant server accumulates one instance per tenant indefinitely. EntryPointRegistry itself uses a TenantLRUCache for its per-instance plugin cache.

from cogniverse_foundation.caching import TenantLRUCache

def close_client(tenant_id: str, client) -> None:
    client.close()

cache = TenantLRUCache[MyClient](capacity=16, on_evict=close_client)

# Atomically get-or-build under a lock (the primary entry point)
client = cache.get_or_set("acme", lambda: MyClient(tenant_id="acme"))
Method Description
__init__(capacity, on_evict=None) capacity must be >= 1; on_evict(key, value) is called (best-effort) on eviction or clear()
get(key) Return cached value or None, marking it most-recently-used
set(key, value) Insert/replace, evicting the least-recently-used entry if over capacity; a replaced value (different object, same key) also gets on_evict so held resources are released
get_or_set(key, factory) Return cached value, or build + cache one atomically (lock held through the factory)
set_if_absent(key, value) Insert unless the key is cached; return the winner. For expensive instances built outside the lock: concurrent builders converge on one shared instance, and the loser stays the caller's to release (no on_evict)
acquire(key) / release(key) Check a value out against eviction, and give the checkout back. acquire returns None and takes no checkout for an uncached key
lease(key) / lease_value(value) Context managers over acquire/release, by key or by the held object; the checkout ends on every exit path including exceptions
lease_count(key) Outstanding checkouts on a key
key_of(value) The key holding a value by identity, or None
pop(key, default=None) Remove and return a value without triggering on_evict
clear() Evict every entry (triggers on_evict for each; a checked-out entry's on_evict is deferred to its release)
keys() / values() Snapshot of current keys / values (LRU order)
copy() Shallow copy preserving capacity, on_evict, and LRU order
len(cache) / key in cache Current size / membership check

COGNIVERSE_TENANT_CACHE_CAPACITY (env var, default 16) sizes the per-registry instance cache inside EntryPointRegistry.

Checkouts

A value in use is checked out for the length of that use. Capacity eviction skips checked-out entries and takes the least-recently-used free entry instead, so a cached value serving a request is never closed under it. When every entry is checked out the cache holds more than its capacity and logs one warning naming the overflow and the number of checkouts blocking it; eviction resumes as checkouts end. A close aimed at a checked-out entry — an overwriting set, a clear — is deferred to the release rather than dropped, so its on_evict still runs exactly once.

with cache.lease("acme") as client:   # None when "acme" is not cached
    client.query(...)                 # "acme" cannot be evicted here

Tenant-delete eviction

register_tenant_cache(cache) and evict_tenant_from_registered_caches(tenant_id) (cogniverse_foundation.caching) let independently-owned TenantLRUCache instances release a deleted tenant's state immediately instead of waiting for LRU pressure. Registration holds only a weak reference (a weakref.WeakSet), so a cache owned by a discarded consumer drops out on garbage collection rather than leaking through the registry.

from cogniverse_foundation.caching import (
    TenantLRUCache,
    register_tenant_cache,
    evict_tenant_from_registered_caches,
)

cache = register_tenant_cache(TenantLRUCache[MyClient](capacity=64))

# On tenant delete:
evicted = evict_tenant_from_registered_caches("acme:production")

evict_tenant_from_registered_caches pops both the given key and its canonical org:tenant form from every registered cache, so simple-form entries are covered too; on_evict does not fire — deletion is a hard drop. It returns the number of entries dropped. The runtime registers six such caches, each of capacity 64 — the per-tenant GatewayAgent cache, the generic-A2A-agent cache (keyed (tenant, agent_name)) and the orchestrator agent cache in agent_dispatcher.py, the per-tenant GraphManager and WikiManager caches in main.py, and the per-tenant ArtifactManager cache in routers/agents.py — and calls evict_tenant_from_registered_caches from release_deleted_tenant, the handler every runtime worker process runs for a tenant delete, so a deleted tenant's cached state is released as part of the delete, not left to linger.

Refreshing values off the request thread

cogniverse_foundation.caching.RefreshingCache holds values read from a slow backing store and keeps the read off the caller's thread once a value is held. ConfigManager keeps its system config and its per-tenant scoped configs in one each, TenantRouterTiers its per-tenant router tiers, and DeployedSchemaNames (core module) its per-tenant deployed schema names.

from cogniverse_foundation.caching import RefreshingCache

cache = RefreshingCache[tuple, dict](
    name="scoped-config",
    refresh_after_s=5.0,
    max_staleness_s=60.0,
    max_entries=512,
)
value = cache.get(key, lambda: store.read(key))
cache.invalidate(lambda k: k[1] == tenant_id)   # after this process writes
Age of the held value get(key, read)
none held runs read on the caller's thread; concurrent callers share it; a failure raises to each and caches nothing
below refresh_after_s returns it, no read
refresh_after_s to below max_staleness_s returns it and starts one background read (daemon thread <name>-refresh); concurrent callers share that read and none waits for it
max_staleness_s or more runs read on the caller's thread, or waits on the read already in flight; a failure raises

Age counts from when the producing read began, so no value is returned max_staleness_s or more after its read started. A failed background read is logged at ERROR with the key and the age of the value still being served. The entry stays, and the key is not refreshed again for refresh_after_s. At most max_background_reads (default 4) background reads run at once. A stale entry found while all are busy is returned, and its refresh starts on a later call. invalidate(matches) drops matching entries and detaches their reads in flight, whose results are never cached. put(key, value) holds a value this process has just written as if read now, detaching the key's read in flight. max_entries bounds the cache, evicting least recently used first. Setting both bounds to 0 reads on every call. The clock argument (default time.monotonic) sets the time source.

get(key, read, accept=...) lets one call refuse a held value: when accept(held) is False the call reads as if nothing were held. It runs under the cache lock and must not call back into the cache. The constructor's keep(value) decides which read results are held; a result it rejects is returned to its callers and drops the key's entry. keys() and items() snapshot the held entries, least recently used first.


DSPy Extensions

cogniverse_foundation.dspy hosts DSPy adapter and model-name helpers shared across agents, evaluation, and finetuning.

LenientJSONAdapter (cogniverse_foundation.dspy.lenient_json_adapter) — a dspy.adapters.json_adapter.JSONAdapter subclass that renames common LM field-name variants (e.g. reason/rationale/thought → reasoning, answer/response/output → summary, sub_question → sub_questions) before the parent's strict field-key equality check. A required output the LM left unfilled — absent, null, or a blank string — raises LMOutputIncomplete (an AdapterParseError subclass) whose missing_fields names every such output and whose parsed_result carries the ones the LM did fill. An empty list or dict is a value and passes. An output declared with a default (dspy.OutputField(default="")) is one whose blank answer is valid: a null or blank value takes the default, and only an absent field is incomplete.

from cogniverse_foundation.dspy.lenient_json_adapter import LenientJSONAdapter
import dspy

dspy.configure(adapter=LenientJSONAdapter())

StructuredJSONAdapter / signature_response_format(signature) (cogniverse_foundation/dspy/structured_json_adapter.py) — a JSONAdapter subclass that sends the signature's output fields to the server as an OpenAI response_format of type json_schema: every output field required, additionalProperties: false, strict: true. The schema is derived from the signature, so guided decoding on an OpenAI-compatible engine (vLLM) can only return an object carrying every field. Stock JSONAdapter asks for a schema only when litellm's registry claims the model supports one and otherwise falls back to {"type": "json_object"}, which a bare {} satisfies; this adapter always sends the schema and raises AdapterParseError rather than reprompting under a second adapter when a server ignores it. A field value that fails its declared type — a list item outside a Literal, a missing or extra key on a nested model — raises AdapterParseError too, carrying the response, so every answer outside the schema is one error type.

The schema's name is the signature's class name, except for the placeholder names dspy gives a signature built from a string, rebuilt by with_instructions() or wrapped by ChainOfThought (Signature, StringSignature). Those fall back to the signature's own output field names plus a digest of its field declaration, so two different declarations never share a schema name and rewriting instructions does not rename the schema.

from cogniverse_foundation.dspy import StructuredJSONAdapter
import dspy

with dspy.context(adapter=StructuredJSONAdapter()):
    prediction = module(query="...")

bare_model_name(model) / ensure_provider_prefix(model, default_provider="openai") (cogniverse_foundation.dspy.model_format) — strip or add a litellm provider prefix (ollama, ollama_chat, hosted_vllm, openai) on a model id, for sites that talk to an OpenAI-compatible HTTP API directly and need a bare model name (Mem0's embedder/LLM wiring) versus sites that need litellm's required provider/model form.

from cogniverse_foundation.dspy.model_format import bare_model_name, ensure_provider_prefix

bare_model_name("hosted_vllm/Qwen/Qwen2.5-7B-Instruct")  # "Qwen/Qwen2.5-7B-Instruct"
ensure_provider_prefix("gemma3:4b")                       # "openai/gemma3:4b"

Confidence Parsing

cogniverse_foundation.confidence.parse_confidence(raw, default=0.0) -> float maps a DSPy module's LM-produced confidence/relevance output — a float, a percent string ("85%"), a label ("high"/"medium"/"low"), a sentence-embedded number ("0.9 (very confident)", "confidence: 0.9", "0.9/1.0"), or an empty string — to a clamped [0.0, 1.0] float, falling back to default on any input it cannot interpret. Lives in foundation so core, agents, evaluation, and finetuning share one implementation instead of each doing a raw float(result.confidence) that crashes on a non-numeric shape.

The sentence-embedded fallback only fires when the string isn't a bare number or exact label match. It first looks for an x/y ratio and accepts it only if x/y lands in [0, 1]; otherwise it takes the first free-standing float token that lands in [0, 1] (digits embedded in a larger token — hex reprs, versions like "1.2.3", identifiers — are skipped). An out-of-range or absent number falls back to default rather than guessing, e.g. "7 (very confident)" stays uninterpretable.

from cogniverse_foundation.confidence import parse_confidence

parse_confidence("85%")                    # 0.85
parse_confidence("high")                   # 0.9
parse_confidence(1.5)                      # 1.0 (clamped)
parse_confidence("")                       # 0.0 (default)
parse_confidence("0.9 (very confident)")   # 0.9 (numeric-extraction fallback)
parse_confidence("confidence: 0.9")        # 0.9 (numeric-extraction fallback)
parse_confidence("0.9/1.0")                # 0.9 (x/y ratio fallback)
parse_confidence("7 (very confident)")     # 0.0 (default; 7 is out of [0, 1])

Usage Examples

Complete Configuration Setup

from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_foundation.config.unified_config import (
    SystemConfig,
    BackendConfig,
    BackendProfileConfig,
)
from cogniverse_foundation.config.agent_config import AgentConfig

# Initialize config manager
config_manager = create_default_config_manager()

# Set global system config (not per-tenant)
system_config = SystemConfig(
    search_backend="vespa",
    backend_url="http://localhost",
    backend_port=8080
)
config_manager.set_system_config(system_config)

# Add backend profile for tenant
profile = BackendProfileConfig(
    profile_name="custom_colpali",
    type="video",
    embedding_model="colpali-v2",
    schema_name="video_colpali_custom",
    pipeline_config={"chunk_strategy": "frame", "top_k": 20}
)
config_manager.add_backend_profile(profile, tenant_id="acme")

# Set agent config (requires all fields)
from cogniverse_foundation.config.agent_config import (
    AgentConfig, ModuleConfig, DSPyModuleType
)
agent_config = AgentConfig(
    agent_name="orchestrator_agent",
    agent_version="1.0.0",
    agent_description="Routes queries",
    agent_url="http://localhost:8001",
    capabilities=["routing"],
    skills=[],
    module_config=ModuleConfig(
        module_type=DSPyModuleType.CHAIN_OF_THOUGHT,
        signature="query -> decision"
    ),
    llm_model="gpt-4",
    llm_temperature=0.7
)
config_manager.set_agent_config(
    tenant_id="acme",
    agent_name="orchestrator_agent",
    agent_config=agent_config
)

Argo Workflows HTTP Client

build_argo_async_client() returns an httpx.AsyncClient with verify=False and DEFAULT_ARGO_TIMEOUT.

from cogniverse_foundation.common.argo_client import build_argo_async_client

client = build_argo_async_client()
client = build_argo_async_client(30.0)

The tenant job-scheduling route and optimization submissions use it for Argo calls.

Tenant-Specific Profile Overrides

from cogniverse_foundation.common.tenant_utils import SYSTEM_TENANT_ID

# Start with system profile, customize for tenant
config_manager.update_backend_profile(
    profile_name="video_colpali_mv_frame",
    overrides={
        "embedding_model": "colpali-custom",
        "top_k": 25
    },
    base_tenant_id=SYSTEM_TENANT_ID,  # Inherit from the cluster-wide system profile (the default)
    target_tenant_id="acme"           # Save to tenant
)

Telemetry with Phoenix

from cogniverse_foundation.telemetry.manager import get_telemetry_manager
from cogniverse_foundation.telemetry.config import TelemetryConfig

# Initialize with config
config = TelemetryConfig(
    enabled=True,
    otlp_endpoint="http://localhost:4317",
    service_name="cogniverse",
    environment="production"
)
telemetry = TelemetryManager(config=config)

# Use in agent processing
async def process_query(query: str, tenant_id: str):
    with telemetry.span(
        "agent.routing",
        tenant_id=tenant_id,
        attributes={"query": query}
    ) as span:
        # Route query
        route = await route_query(query)
        span.set_attribute("route.agent", route.agent_name)
        span.set_attribute("route.confidence", route.confidence)

        # Execute with nested span
        with telemetry.span(
            f"agent.{route.agent_name}",
            tenant_id=tenant_id
        ) as child_span:
            result = await execute_agent(route.agent_name, query)
            child_span.set_attribute("result.count", len(result.items))

        return result

Querying Telemetry Data

from datetime import datetime, timezone

async def query_and_annotate():
    # Get provider for querying spans
    provider = telemetry.get_provider(tenant_id="acme")

    # Query spans from Phoenix
    spans_df = await provider.traces.get_spans(
        project="cogniverse-acme",
        start_time=datetime(2025, 1, 1, tzinfo=timezone.utc),
        limit=1000
    )

    # Stream page-sized frames when the result set is large.
    async for page in provider.traces.iter_spans(
        project="cogniverse-acme",
        page_size=512,
        columns=("start_time", "end_time", "status_code"),
    ):
        process(page)

    # Add annotation (project is required)
    await provider.annotations.add_annotation(
        span_id="abc123",
        name="human_review",
        label="approved",
        score=1.0,
        metadata={"reviewer": "alice"},
        project="cogniverse-acme"
    )

Architecture Position

flowchart TB
    subgraph CoreLayer["<span style='color:#000'>Core Layer</span>"]
        Core["<span style='color:#000'>cogniverse-core (agents)</span>"]
        Evaluation["<span style='color:#000'>cogniverse-evaluation</span>"]
    end

    subgraph FoundationLayer["<span style='color:#000'>Foundation Layer</span>"]
        Foundation["<span style='color:#000'>cogniverse-foundation ◄─ YOU ARE HERE<br/>ConfigManager, TelemetryManager, Provider Registry</span>"]
        SDK["<span style='color:#000'>cogniverse-sdk (interfaces)</span>"]
    end

    CoreLayer --> FoundationLayer

    style CoreLayer fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Core fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Evaluation fill:#ce93d8,stroke:#7b1fa2,color:#000
    style FoundationLayer fill:#a5d6a7,stroke:#388e3c,color:#000
    style Foundation fill:#a5d6a7,stroke:#388e3c,color:#000
    style SDK fill:#a5d6a7,stroke:#388e3c,color:#000

Dependencies (declared in pyproject.toml):

  • cogniverse-sdk: Pure interfaces (ConfigStore, Backend, etc.)
  • dspy-ai: LenientJSONAdapter subclasses dspy.adapters.json_adapter.JSONAdapter; create_dspy_lm builds dspy.LM
  • opentelemetry-api/sdk: Telemetry infrastructure
  • pydantic: Configuration validation
  • pandas: DataFrame return type for the TraceStore/AnnotationStore/DatasetStore interfaces

Not a declared dependency, but imported at runtime:

  • cogniverse-vespa: VespaConfigStore is imported lazily inside create_default_config_manager() — only required when backend.type == "vespa"

Package Structure — cogniverse_foundation/common/:

The shared, dependency-free helpers that the whole stack uses live here so foundation code (config/manager.py, config/utils.py, config/unified_config.py, config/api_mixin.py, registry/entry_point_registry.py) can use them without importing upward into cogniverse-core:

  • common/tenant_utils.py — tenant identity helpers (SYSTEM_TENANT_ID, require_tenant_id, canonical_tenant_id, parse_tenant_id, validate_tenant_id, get_tenant_storage_path, sanitize_k8s_label_value). cogniverse_core.common.tenant_utils re-exports these, so existing from cogniverse_core.common.tenant_utils import ... call sites keep working; the runtime-coupled assert_tenant_exists stays in core.
  • common/dspy_module_registry.py — DSPyModuleRegistry / DSPyOptimizerRegistry, re-exported from cogniverse_core.common.dspy_module_registry.

This keeps the Core → Foundation dependency direction one-way: cogniverse-foundation no longer imports cogniverse-core and is usable standalone.

Dependents:

  • cogniverse-core: Uses ConfigManager, TelemetryManager
  • cogniverse-agents: Uses configuration and telemetry
  • cogniverse-telemetry-phoenix: Implements telemetry provider

Testing

# Run foundation tests (confidence, caching, dspy adapters, semantic router, telemetry, config)
JAX_PLATFORM_NAME=cpu uv run pytest tests/foundation/ tests/telemetry/ tests/common/unit/ tests/common/integration/ -v

# Test confidence parsing, TenantLRUCache, dspy model-format, semantic router, entry-point registry
uv run pytest tests/foundation/unit/ -v

# Semantic router end-to-end (spins up a stub upstream)
uv run pytest tests/foundation/integration/test_semantic_router_e2e.py -v

# Test configuration
uv run pytest tests/common/unit/ -v -k "config"

# Test telemetry
uv run pytest tests/telemetry/ -v

# Test with coverage
uv run pytest tests/foundation/ tests/telemetry/ tests/common/unit/ tests/common/integration/ --cov=cogniverse_foundation --cov-report=html

Test Categories:

  • tests/foundation/unit/ - parse_confidence, TenantLRUCache, DSPy model-format helpers, create_dspy_lm, semantic router unit tests, EntryPointRegistry
  • tests/foundation/integration/ - Semantic router end-to-end against a stub upstream (_sr_stack/stub_upstream.py)
  • tests/telemetry/ - Telemetry manager and provider tests
  • tests/common/unit/ - Unit tests for AgentConfig, ConfigAPIMixin, and related agent-facing configuration
  • tests/common/integration/ - Integration tests with backend (e.g., Vespa) config persistence


Summary: The Foundation module provides the infrastructure layer for Cogniverse. ConfigManager handles multi-tenant, versioned configuration with pluggable backend persistence (VespaConfigStore). TelemetryManager provides OpenTelemetry tracing with tenant isolation and Phoenix integration. All operations require tenant_id to ensure proper multi-tenant isolation.