Skip to content

Cogniverse Study Guide: Scripts & Operations Module

Module Path: scripts/ SDK Packages: Uses all 12 packages (foundation → core → implementation → application)


Table of Contents

  1. Module Overview
  2. Architecture
  3. Core Scripts
  4. Data Flow
  5. Usage Examples
  6. Production Considerations

Module Overview

Purpose

The Scripts & Operations module provides command-line tools for:

  • Video Ingestion: Processing and indexing video content

  • Schema Deployment: Managing Vespa search schemas

  • System Setup: Initializing the system environment

  • Optimization: Running DSPy optimization workflows

  • Experimentation: Conducting Phoenix experiments with visualization

  • Dataset Management: Managing evaluation datasets

Key Features

  • Builder Pattern Ingestion: Fluent API for configurable pipeline construction
  • Multi-Profile Support: Process videos with multiple embedding strategies simultaneously
  • Tenant-Aware Processing: Per-tenant schema isolation and configuration
  • Async Processing: Concurrent video processing with configurable limits
  • Phoenix Integration: Per-tenant experiment tracking with visual analytics
  • Schema Management: JSON-based tenant-specific schema deployment
  • UV Workspace: All scripts use uv run for SDK package management

Script Categories

scripts/
├── Ingestion & Processing
│   ├── run_ingestion.py              # Main video ingestion pipeline
│   ├── test_ingestion.py             # Test ingestion with validation
│   └── backfill_source_url.py        # Backfill source URLs on existing documents
│
├── Deployment & Setup
│   ├── deploy_json_schema.py         # Deploy single JSON schema
│   ├── provision_tenant.py           # Entry point for cogniverse_runtime.provision_tenant
│   ├── setup_ollama.py               # Ollama model setup
│   └── setup_gliner.py               # GLiNER setup
│   (bulk schema deploy flows through the runtime admin API:
│    POST /admin/profiles/{profile}/deploy — see charts/cogniverse/
│    templates/init-jobs.yaml for the in-cluster init job.)
│
├── Optimization & Experiments
│   ├── run_experiments_with_visualization.py  # Phoenix experiments
│   └── auto_optimization_trigger.py  # Automated optimization trigger
│
├── Dataset Management
│   ├── manage_datasets.py            # Create/list/inspect datasets (tenant-scoped)
│   ├── manage_golden_datasets.py     # Golden dataset management
│   ├── create_golden_dataset_from_traces.py  # Golden dataset from traces
│   └── seed_bright_corpus.py         # Seed corpus data
│
├── Reporting & Analysis
│   ├── generate_integrated_evaluation_report.py  # Integrated evaluation report
│   ├── generate_tabbed_html_report.py    # Tabbed HTML report generation
│   ├── generate_langextract_training_data.py  # Training data for LangExtract
│   ├── view_integrated_results.py    # View integrated results
│   └── vlm_caption_bakeoff.py        # VLM caption comparison
│
├── Utilities & Operations
│   ├── discover_tenants.py           # Tenant discovery
│   ├── export_backend_embeddings.py  # Backend embedding export (tenant-aware)
│   ├── manage_phoenix_data.py        # Phoenix data management
│   ├── prune_config_metadata.py      # Config metadata pruning
│   ├── release_manifest.py           # Stage pinned release sources, validate builds, write dist/BUILD_MANIFEST.json, check publishability
│   ├── start_phoenix.py              # Start Phoenix service
│   └── version_bump.py               # Version management
│
└── Shell Scripts (build, local dev, test infrastructure)
    ├── build_packages.sh             # Build the release package set in dependency order
    ├── publish_packages.sh           # Publish SDK packages to (Test)PyPI
    ├── install_with_gpu.sh           # uv sync with the PyTorch backend matching local hardware
    ├── download_test_data.sh         # Download the evaluation dataset
    ├── start_vespa.sh / stop_vespa.sh        # Vespa container lifecycle (persistent volumes)
    ├── start_phoenix.sh              # Start/stop/restart Phoenix via Docker
    ├── setup_evaluation.sh           # Set up the Cogniverse evaluation framework
    ├── setup_local_tests.sh          # Probe every endpoint the local integration tests depend on
    ├── run_local_tests.sh            # Run comprehensive local-only test coverage (skipped in CI)
    ├── run_integration_per_package.sh # Run integration tests one package at a time (isolated processes)
    ├── run_e2e_batched.sh            # Run the e2e suite in two batches with a runtime restart between them
    ├── record_test_goldens.sh        # Re-record integration-test goldens against the live k3d cluster
    ├── analyze_vespa_embeddings.sh   # Export Vespa embeddings and run an integrity check
    ├── memsnap_sampler.sh            # Sample /admin/debug/memsnap at fixed intervals
    └── sweep_state_snapshot.sh       # Snapshot cluster + Vespa + Mem0 registry state during a sweep

Architecture

1. Ingestion Pipeline Architecture

flowchart TB
    Entry["<span style='color:#000'>run_ingestion.py<br/>Command Line Entry Point</span>"]

    TestMode["<span style='color:#000'>Test Mode<br/>build_test_pipeline</span>"]
    SimpleMode["<span style='color:#000'>Simple Mode<br/>build_simple_pipeline</span>"]
    AdvMode["<span style='color:#000'>Advanced Mode<br/>create_pipeline</span>"]

    Pipeline["<span style='color:#000'>VideoIngestionPipeline<br/>• Video Processing<br/>• Embedding Generation<br/>• Vespa Upload</span>"]

    ColPali["<span style='color:#000'>ColPali Profile<br/>Frame-based</span>"]
    X-CLIP["<span style='color:#000'>X-CLIP Profile<br/>Global embeddings</span>"]
    ColQwen["<span style='color:#000'>ColQwen Profile<br/>Chunk-based</span>"]

    Concurrent["<span style='color:#000'>Concurrent Async<br/>Video Processing<br/>max_concurrent=3</span>"]

    Vespa["<span style='color:#000'>Vespa Backend<br/>Bulk Upload</span>"]

    Entry --> TestMode
    Entry --> SimpleMode
    Entry --> AdvMode

    TestMode --> Pipeline
    SimpleMode --> Pipeline
    AdvMode --> Pipeline

    Pipeline --> ColPali
    Pipeline --> X-CLIP
    Pipeline --> ColQwen

    ColPali --> Concurrent
    X-CLIP --> Concurrent
    ColQwen --> Concurrent

    Concurrent --> Vespa

    style Entry fill:#90caf9,stroke:#1565c0,color:#000
    style TestMode fill:#b0bec5,stroke:#546e7a,color:#000
    style SimpleMode fill:#b0bec5,stroke:#546e7a,color:#000
    style AdvMode fill:#b0bec5,stroke:#546e7a,color:#000
    style Pipeline fill:#ffcc80,stroke:#ef6c00,color:#000
    style ColPali fill:#ce93d8,stroke:#7b1fa2,color:#000
    style X-CLIP fill:#ce93d8,stroke:#7b1fa2,color:#000
    style ColQwen fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Concurrent fill:#ffcc80,stroke:#ef6c00,color:#000
    style Vespa fill:#a5d6a7,stroke:#388e3c,color:#000

2. Optimization Workflow Architecture

flowchart TB
    Start["<span style='color:#000'>optimization_cli module<br/>Per-Agent Optimization CLI</span>"]

    Step1["<span style='color:#000'>Step 1: Run Agent Optimization<br/>• Execute per-agent optimizer<br/>• Modes: simba/gateway-thresholds/entity-extraction/etc.</span>"]

    Step2["<span style='color:#000'>Step 2: Persist Artifact<br/>• ArtifactManager.save_blob (per tenant)<br/>• Returns artifact_id</span>"]

    Step3["<span style='color:#000'>Step 3: Agents Load at Startup<br/>• Runtime agents read latest tenant artifact<br/>• No redeploy step</span>"]

    Start --> Step1
    Step1 --> Step2
    Step2 --> Step3

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Step1 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Step2 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Step3 fill:#a5d6a7,stroke:#388e3c,color:#000

3. Experiment Workflow Architecture

flowchart TB
    Start["<span style='color:#000'>run_experiments_with_visualization.py<br/>Phoenix Experiment Runner</span>"]

    Runner["<span style='color:#000'>ExperimentTracker<br/>• From cogniverse_evaluation SDK<br/>• Experiment project isolation<br/>• Quality evaluators optional<br/>• LLM evaluators optional</span>"]

    Dataset["<span style='color:#000'>Dataset Preparation<br/>• Load or create dataset<br/>• CSV parsing<br/>• Phoenix dataset registration</span>"]

    Loop["<span style='color:#000'>Multi-Profile Multi-Strategy Loop<br/>FOR each profile:<br/>  FOR each strategy:<br/>    • Run experiment<br/>    • Track spans<br/>    • Evaluate results<br/>    • Store metrics</span>"]

    Viz["<span style='color:#000'>Visualization Generation<br/>• Profile summary table<br/>• Strategy comparison<br/>• Detailed results<br/>• HTML report optional</span>"]

    Export["<span style='color:#000'>Results Export<br/>• CSV summary<br/>• JSON detailed results<br/>• Phoenix UI links</span>"]

    Start --> Runner
    Runner --> Dataset
    Dataset --> Loop
    Loop --> Viz
    Viz --> Export

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Runner fill:#ffcc80,stroke:#ef6c00,color:#000
    style Dataset fill:#ffcc80,stroke:#ef6c00,color:#000
    style Loop fill:#ffcc80,stroke:#ef6c00,color:#000
    style Viz fill:#ffcc80,stroke:#ef6c00,color:#000
    style Export fill:#a5d6a7,stroke:#388e3c,color:#000

4. Schema Deployment Architecture

flowchart TB
    subgraph Single["<span style='color:#000'>Single Schema Deployment</span>"]
        Single1["<span style='color:#000'>deploy_json_schema.py<br/>Single Schema Deployment</span>"]
        Parser["<span style='color:#000'>JsonSchemaParser<br/>• Load JSON schema file<br/>• Parse to Vespa Schema object<br/>• Validate structure</span>"]
        Package1["<span style='color:#000'>ApplicationPackage<br/>• Create package<br/>• Add schema<br/>• Generate ZIP</span>"]
        Deploy1["<span style='color:#000'>HTTP Deployment<br/>• POST to config server<br/>• Port: 19071<br/>• Endpoint: prepareandactivate</span>"]
        Verify["<span style='color:#000'>Verification<br/>• Check ApplicationStatus<br/>• Verify Vespa responding</span>"]

        Single1 --> Parser
        Parser --> Package1
        Package1 --> Deploy1
        Deploy1 --> Verify
    end

    subgraph Multi["<span style='color:#000'>Multi-Tenant Deployment</span>"]
        Multi1["<span style='color:#000'>Runtime admin API<br/>POST /admin/profiles/{profile}/deploy</span>"]
        Registry["<span style='color:#000'>SchemaRegistry.deploy_schema<br/>• Per-tenant scoping<br/>• Merge with live cluster via<br/>  list_deployed_document_types()</span>"]
        Package2["<span style='color:#000'>ApplicationPackage<br/>• Merged schemas (registry + Vespa)<br/>• allow_schema_removal=False<br/>  (fail-loud on unresolved drops)</span>"]

        Multi1 --> Registry
        Registry --> Package2
    end

    style Single1 fill:#90caf9,stroke:#1565c0,color:#000
    style Multi1 fill:#90caf9,stroke:#1565c0,color:#000
    style Parser fill:#ffcc80,stroke:#ef6c00,color:#000
    style Package1 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Registry fill:#ffcc80,stroke:#ef6c00,color:#000
    style Package2 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Deploy1 fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Verify fill:#a5d6a7,stroke:#388e3c,color:#000

Core Scripts

1. run_ingestion.py

Purpose: Main entry point for video ingestion pipeline with builder pattern configuration

Location: scripts/run_ingestion.py (286 lines)

Command Line Arguments:

--tenant-id TENANT       # Tenant ID for schema isolation (required — no default)
--video_dir PATH         # Directory containing content files
--content-dir PATH       # Directory containing content files (alias for --video_dir)
--media-root-uri URI     # Non-filesystem source (e.g. s3://, pvc://); overrides --video_dir
--output_dir PATH        # Output directory for processed data
--backend {byaldi,vespa} # Search backend (default: vespa)
--profile PROFILES       # Processing profiles (space-separated)
--content-type {video,image,audio,document}  # Content type (default: video)
--max-concurrent INT     # Max concurrent items to process (default: 3)
--max-frames INT         # Maximum frames per video / images per batch
--test-mode             # Use test mode with limited frames
--debug                 # Enable debug mode

Usage Modes:

Test Mode (for quick validation):

# Use test pipeline builder (tenant_id is required — no default)
pipeline = build_test_pipeline(
    tenant_id="acme_corp",
    video_dir=Path("data/testset/evaluation/sample_videos"),
    schema="video_colpali_smol500_mv_frame",
    max_frames=10
)

Simple Mode (for standard usage):

# Use simple pipeline builder (tenant_id is required — no default)
pipeline = build_simple_pipeline(
    tenant_id="acme_corp",
    video_dir=Path("data/testset/evaluation/sample_videos"),
    schema="video_colpali_smol500_mv_frame",
    backend="vespa",
    debug=False
)

Advanced Mode (for custom configuration):

# Use fluent builder with custom config
config = (create_config()
    .video_dir(Path("data/videos"))
    .backend("vespa")
    .output_dir(Path("custom/output"))
    .max_frames_per_video(100)
    .build())

from cogniverse_foundation.config.utils import create_default_config_manager
config_manager = create_default_config_manager()

# with_tenant_id() and with_config_manager() are both required before
# build() — the builder raises ValueError if either is missing.
pipeline = (create_pipeline()
    .with_tenant_id("acme_corp")
    .with_config_manager(config_manager)
    .with_config(config)
    .with_schema("video_colpali_smol500_mv_frame")
    .with_debug(True)
    .with_concurrency(5)
    .build())

Multi-Profile Processing (Tenant-Aware):

# Process with multiple profiles simultaneously for specific tenant
# Uses builder pattern from runtime package
from cogniverse_runtime.ingestion.pipeline_builder import build_simple_pipeline

for profile in ["video_colpali_smol500_mv_frame",
                "video_xclip_sv_chunk_6s"]:
    pipeline = build_simple_pipeline(
        tenant_id="acme_corp",  # Tenant-specific processing
        video_dir=Path("data/videos"),
        schema=profile,
        backend="vespa"
    )
    results = await pipeline.process_videos_concurrent(
        video_files,
        max_concurrent=3
    )

Output:

  • Success/failure status per video

  • Documents fed to Vespa

  • Processing time and throughput

  • Per-profile summary statistics


2. deploy_json_schema.py

Purpose: Deploy individual JSON schema files to Vespa

Location: scripts/deploy_json_schema.py (196 lines)

Command Line Arguments:

schema_file              # Path to JSON schema file (required)
--config-host HOST       # Vespa config server host (default: localhost)
--config-port PORT       # Config server port (default: 19071)
--data-host HOST         # Vespa data endpoint host (default: localhost)
--data-port PORT         # Data endpoint port (default: 8080)

Deployment Process:

def deploy_json_schema(schema_file, vespa_host, config_port, data_port):
    # Imports from SDK packages
    from cogniverse_vespa.json_schema_parser import JsonSchemaParser
    from vespa.package import ApplicationPackage

    # 1. Load JSON schema
    with open(schema_file, 'r') as f:
        schema_config = json.load(f)

    # 2. Parse schema using JsonSchemaParser
    parser = JsonSchemaParser()
    schema = parser.parse_schema(schema_config)

    # 3. Create application package
    app_package = ApplicationPackage(name=schema.name.replace('_', ''))
    app_package.add_schema(schema)

    # 4. Deploy via HTTP
    deploy_url = f"http://{vespa_host}:{config_port}/application/v2/tenant/default/prepareandactivate"
    app_zip = app_package.to_zip()
    response = requests.post(deploy_url, headers={"Content-Type": "application/zip"},
                            data=app_zip, timeout=60)

    # 5. Verify deployment
    verify_deployment(schema.name, vespa_host, data_port)

Verification:

def verify_deployment(schema_name, vespa_host, data_port):
    # Check application status endpoint
    response = requests.get(
        f"http://{vespa_host}:{data_port}/ApplicationStatus",
        timeout=5
    )
    return response.status_code == 200

Example Usage:

# Deploy agent memories schema
python scripts/deploy_json_schema.py \
  configs/schemas/agent_memories_schema.json

# Deploy video schema
python scripts/deploy_json_schema.py \
  configs/schemas/video_colpali_smol500_mv_frame_schema.json

# Deploy to remote Vespa instance
python scripts/deploy_json_schema.py \
  configs/schemas/config_metadata_schema.json \
  --config-host vespa.example.com \
  --config-port 19071


3. Bulk per-tenant schema deployment (runtime admin API)

Schema deployment is always per-tenant and always flows through the runtime admin API so it goes through SchemaRegistry.deploy_schema and the VespaBackend.deploy_schemas merge path.

In-cluster path: charts/cogniverse/templates/init-jobs.yaml registers each .Values.config.tenants entry (200/201 created and 409 already registered both proceed; any other status fails the Job attempt) and deploys .Values.config.defaultProfiles.video (none when it is empty) for it by calling the runtime:

curl -sS -w ' HTTP_STATUS=%{http_code}' -X POST "$RUNTIME_URL/admin/tenants" \
  -H "Content-Type: application/json" \
  -d '{"tenant_id": "{{ $tenant.id }}", "created_by": "helm:{{ $.Chart.Name }}"}'

curl -X POST "$RUNTIME_URL/admin/profiles/{{ $profile }}/deploy" \
  -H "Content-Type: application/json" \
  -d '{"tenant_id": "{{ $tenant.id }}", "force": false}'

Local path:

RUNTIME_URL=http://localhost:8080
# Register tenant (deploys tenant_metadata etc.)
curl -X POST "$RUNTIME_URL/admin/tenants" \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id": "acme:production"}'

# Deploy a profile's content schema for that tenant
curl -X POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id": "acme:production", "force": false}'

For a single-schema JSON deploy from the filesystem (dev workflow), scripts/deploy_json_schema.py still works.


3a. cogniverse admin reconcile-orphans

DELETE /admin/tenants/{id} refuses if a peer tenant has an unreconstructable Vespa-only orphan (the redeploy would silently drop the peer's schema). After power loss, SIGKILL between Vespa deploy and registry write, or any other interrupted-deploy path, an operator needs an out-of-band recovery tool — that's cogniverse admin reconcile-orphans.

# Dry-run (default): list orphan schemas + implied tenants, no changes
cogniverse admin reconcile-orphans

# Confirm: drop every orphan tenant in ONE atomic Vespa redeploy
cogniverse admin reconcile-orphans --confirm

# Point at a non-default runtime
cogniverse admin reconcile-orphans --runtime-url http://runtime.cogniverse.svc:28000

The CLI is a thin client over POST /admin/reconcile-orphans?dry_run=…, which internally calls VespaSchemaManager.delete_tenant_schemas_bulk. The endpoint reports:

  • orphan_schemas — Vespa-deployed names with no registry record
  • orphan_tenants — implied tenant ids (recovered by stripping known base prefixes)
  • unrecovered_schemas — orphans whose base prefix isn't in KNOWN_BASES; need operator review

See operations/multi-tenant-ops.md#orphan-reconciliation for the full operator workflow and JSON shapes.


4. provision_tenant.py

Purpose: Cold-bootstrap a tenant's backend resources (Vespa schemas, Mem0 memory schema, Phoenix telemetry project, semantic-router tier) without a live runtime. The tenant-provisioning WorkflowTemplate runs it inside the runtime image.

Location: libs/runtime/cogniverse_runtime/provision_tenant.py, with scripts/provision_tenant.py as the checkout-side entry point.

Command Line Arguments:

--tenant-id TENANT   # Tenant identifier (required)
--step STEP          # Provisioning step: schemas|verify|memory|telemetry|tier (required)
--tier TIER          # Router tier to store (required by --step tier)
--profiles LIST      # Comma-separated backend profiles (required by --step schemas|verify)

Environment: every step needs BACKEND_URL / BACKEND_PORT (the Vespa data endpoint) because every step builds a ConfigManager over the tenant store; VESPA_CONFIG_PORT names the config server the schema deploy posts to. A failing step prints one line naming the cause and exits 1.

--profiles entries resolve through the tenant's merged backend catalog — the cluster profiles in configs/config.json with the tenant's stored overrides on top — so a tenant with no backend rows yet still provisions.

Steps:

  • schemas — Deploys the profiles' schemas through SchemaRegistry.deploy_schemas as one application package, under their tenant-scoped names; a profile that cannot be loaded leaves none of them registered
  • verify — Confirms each profile's tenant schema carries a registry row and answers a YQL query
  • memory — Creates the tenant's Mem0 memory schema via lazy_init_memory
  • telemetry — Exports a required probe span via TelemetryManager.required_span(...) to create the Phoenix project; an unreachable collector fails the step
  • tier — Stores the tenant's semantic-router tier via set_tenant_tier; a tier outside ROUTER_TIERS is refused before any store is built

Usage:

# Initialize memory schema for tenant
uv run python -m cogniverse_runtime.provision_tenant \
  --tenant-id acme:production \
  --step memory

# Initialize Phoenix telemetry project for tenant
uv run python -m cogniverse_runtime.provision_tenant \
  --tenant-id acme:production \
  --step telemetry

# Store the router tier for tenant
uv run python -m cogniverse_runtime.provision_tenant \
  --tenant-id acme:production \
  --step tier --tier pro

Implementation:

def deploy_schemas(tenant_id: str, profiles: List[str]) -> List[str]:
    tenant = canonical_tenant_id(tenant_id)
    _, backend, schema_names = _resolve(tenant, profiles)
    return backend.schema_registry.deploy_schemas(
        tenant_id=tenant, base_schema_names=schema_names
    )

def init_memory(tenant_id: str) -> None:
    config_manager = _config_manager(tenant_id)
    mgr = Mem0MemoryManager(tenant_id)
    lazy_init_memory(mgr, tenant_id, config_manager, auto_create_schema=True)

def init_telemetry(tenant_id: str) -> None:
    tenant = canonical_tenant_id(tenant_id)
    manager = get_telemetry_manager(
        _config_manager(tenant),
        otlp_endpoint=resolve_library_env_defaults()["telemetry_otlp_endpoint"],
    )
    async with manager.required_span("provision.probe", tenant_id=tenant):
        pass


5. cogniverse_runtime.optimization_cli

Purpose: Per-agent optimization CLI, cluster maintenance sweep, and reporting jobs, all driven by DSPy BootstrapFewShot over Phoenix spans

Location: libs/runtime/cogniverse_runtime/optimization_cli.py (3598 lines)

What Gets Optimized (--mode):

  • simba - Compiles the QueryEnhancementAgent's DSPy module from cogniverse.query_enhancement spans (original_query → enhanced_query pairs)

  • gateway-thresholds - Gateway confidence threshold tuning

  • entity-extraction - Entity extraction optimization

  • workflow - Multi-agent workflow orchestration (runs after the per-agent optimizers so it sees their artifacts)

  • profile - Search profile selection optimization

  • cleanup - Daily maintenance sweep: per-tenant Mem0 TTL cleanup (schema-driven), log rotation under LOG_DIR, temp-file cleanup under TEMP_DIR, and config_metadata version vacuum

  • triggered - On-demand optimization for a comma-separated --agents list, driven by a Phoenix --trigger-dataset

  • online-routing-eval - Scores routing spans against ground truth without retraining anything

  • synthetic - Generates pending review batches for one or more optimizer types (--agents, default query_enhancement,profile,routing,entity_extraction)

  • rollback - Restores a single agent's prompts/demonstrations to a prior version (--agent, --prompts-version/--demos-version)

  • ab-compare - Runs an RLM A/B comparison over a Phoenix queries dataset (--queries-dataset, optional --judge-substring)

  • egress-netpol - Generates Kubernetes NetworkPolicy manifests from configs/agent_policies/*.yaml (no tenant, no Phoenix — pure codegen)

  • monthly-reports - Aggregates usage/performance data into JSON reports under --reports-output-dir

How Optimization Works:

Every DSPy-based mode (simba, workflow, gateway-thresholds, profile, entity-extraction) builds a training set from Phoenix spans and compiles the target module with dspy.teleprompt.BootstrapFewShot, scaled by trainset size (_create_teleprompter):

  • < 50 examples → max_bootstrapped_demos=4, max_labeled_demos=8, max_rounds=1
  • >= 50 examples → max_bootstrapped_demos=8, max_labeled_demos=16, max_rounds=2

There is no automatic switch between multiple DSPy optimizer algorithms (no GEPA/MIPRO/SIMBA-the-algorithm selection) — BootstrapFewShot is the only optimizer this CLI compiles with; simba here is the mode name for query-enhancement optimization, not the DSPy SIMBA optimizer.

Command Line Usage:

# Optimize query enhancement (mode: simba)
uv run python -m cogniverse_runtime.optimization_cli \
  --mode simba \
  --tenant-id default

# Optimize gateway thresholds
uv run python -m cogniverse_runtime.optimization_cli \
  --mode gateway-thresholds \
  --tenant-id acme_corp

# Clean up old logs, memories, and config versions (global sweep)
uv run python -m cogniverse_runtime.optimization_cli \
  --mode cleanup \
  --log-retention-days 7

Command Line Options:

--mode CHOICE                # cleanup|triggered|simba|workflow|gateway-thresholds|
                             # online-routing-eval|profile|entity-extraction|synthetic|
                             # rollback|ab-compare|egress-netpol|monthly-reports (required)
--tenant-id ID               # Tenant identifier (required for every mode except cleanup,
                             # egress-netpol, and monthly-reports; if omitted under
                             # --mode cleanup the sweep runs globally across every tenant
                             # in every org — the path the daily-cleanup CronWorkflow takes)
--log-retention-days DAYS    # Days to retain logs (cleanup mode, default: 7)
--memory-retention-days DAYS # Days to retain expired memories (cleanup mode, default: 30)
--lookback-hours HOURS       # Span history window for DSPy-based modes (default: 24.0)

Integration with Argo Workflows:

optimization_cli is invoked directly (not via a separate workflows/ YAML) from CronWorkflow templates rendered by charts/cogniverse/templates/optimization-workflows.yaml, e.g. the weekly agent-optimization workflow's per-mode step:

# charts/cogniverse/templates/optimization-workflows.yaml (run-optimizer step)
- name: gateway-thresholds
  templateRef:
    name: {{ $fullName }}-optimization-runner
    template: run-optimizer
  arguments:
    parameters:
    - name: mode
      value: gateway-thresholds
    - name: tenant-id
      value: "{{workflow.parameters.tenant-id}}"
    - name: lookback-hours
      value: "48"

The web client's Optimization runs view can also submit a one-off run of the same CLI on demand via POST /admin/tenant/{tenant_id}/optimize, which creates an Argo Workflow from the same template.

Scheduled Execution:

CronWorkflow schedules are set in charts/cogniverse/values.yaml under argo.optimization / argo.maintenance and rendered by charts/cogniverse/templates/optimization-workflows.yaml:

  • Daily cleanup (0 4 * * *, 4 AM UTC): --mode cleanup
  • Daily gateway tuning (0 4 * * *, 4 AM UTC): --mode gateway-thresholds (48h lookback), then restarts the runtime deployment
  • Weekly agent optimization (0 3 * * 0, Sunday 3 AM UTC): gateway-thresholds, entity-extraction, simba, profile in parallel (48h lookback, 168h for simba), then workflow, then restarts the runtime deployment
  • Saturday synthetic generation (0 1 * * 6, 1 AM UTC): --mode synthetic --agents query_enhancement,profile,routing,entity_extraction
  • Daily scheduled distillation (0 5 * * *, 5 AM UTC): runs cogniverse_runtime.quality_monitor_cli --once, a separate CLI, not optimization_cli
  • Monthly reports (0 5 1 * *, 1st of month 5 AM UTC): --mode monthly-reports, uploaded to MinIO by a follow-up step

6. run_experiments_with_visualization.py

Purpose: Run Phoenix experiments with comprehensive visualization and quality evaluators

Location: scripts/run_experiments_with_visualization.py (155 lines)

Architecture: This script is a thin CLI wrapper that delegates to ExperimentTracker from the cogniverse_evaluation SDK package.

Experiment Execution:

def main():
    # Parse CLI arguments (tenant-id, dataset-name, profiles, strategies, evaluators, etc.)
    args = parser.parse_args()

    # Initialize ExperimentTracker from SDK
    from cogniverse_evaluation.core.experiment_tracker import ExperimentTracker

    # tenant_id is required — ExperimentTracker has no default tenant
    tracker = ExperimentTracker(
        tenant_id=args.tenant_id,
        experiment_project_name="experiments",
        enable_quality_evaluators=args.quality_evaluators,
        enable_llm_evaluators=args.llm_evaluators,
        evaluator_name=args.evaluator,
        llm_model=args.llm_model,
        llm_base_url=args.llm_base_url,
    )

    # Get configurations (profiles x strategies matrix)
    tracker.get_experiment_configurations(
        profiles=args.profiles,
        strategies=args.strategies,
        all_strategies=args.all_strategies,
    )

    # Create or get dataset
    dataset_name = tracker.create_or_get_dataset(
        dataset_name=args.dataset_name,
        csv_path=args.csv_path,
        force_new=args.force_new,
    )

    # Run all experiments
    experiments = tracker.run_all_experiments(dataset_name)

    # Create and print visualization tables
    tables = tracker.create_visualization_tables()
    tracker.print_visualization(tables)

    # Save results (CSV + JSON) and generate HTML report
    tracker.save_results(tables, experiments)
    tracker.generate_html_report()

Output:

  • Profile summary table

  • Strategy comparison by profile

  • Detailed experiment results

  • CSV summary file

  • JSON detailed results

  • HTML integrated report (if quantitative tests exist)

Command Line Arguments:

# --tenant-id is required; --dataset-path must be a CSV (query, expected_videos,
# category columns) — create_or_get_dataset loads it with pandas.read_csv.
python scripts/run_experiments_with_visualization.py \
  --tenant-id acme:acme \
  --dataset-name golden_eval_v1 \
  --dataset-path data/testset/evaluation/video_search_queries.csv \
  --profiles frame_based_colpali \
  --quality-evaluators \
  --llm-evaluators \
  --evaluator visual_judge \
  --llm-model google/gemma-4-e4b-it


7. manage_datasets.py

Purpose: CLI tool for managing evaluation datasets

Location: scripts/manage_datasets.py (62 lines)

Operations:

List Datasets:

python scripts/manage_datasets.py --tenant-id acme:acme --list

# Output:
# Registered datasets:
#
# Name: golden_eval_v1
#   Dataset ID: ds_abc123
#   Created: 2025-10-07 10:30:00
#   Examples: 50
#   Description: Golden evaluation dataset

Create Dataset:

python scripts/manage_datasets.py \
  --tenant-id acme:acme \
  --create my_dataset \
  --csv data/queries.csv

# Output:
# Dataset 'my_dataset' created with ID: ds_xyz789

Get Dataset Info:

python scripts/manage_datasets.py --tenant-id acme:acme --info golden_eval_v1

# Output:
# Dataset: golden_eval_v1
#   dataset_id: ds_abc123
#   created_at: 2025-10-07 10:30:00
#   num_examples: 50
#   description: Golden evaluation dataset

Implementation:

def main():
    from cogniverse_evaluation.data import DatasetManager

    dm = DatasetManager(tenant_id=args.tenant_id)

    if args.list:
        dataset_names = dm.list_datasets()  # Returns List[str]
        if dataset_names:
            print("\nRegistered datasets:")
            for ds_name in dataset_names:
                info = dm.get_dataset(ds_name)
                print(f"\nName: {ds_name}")
                if info:
                    for key, value in info.items():
                        print(f"  {key}: {value}")
        else:
            print("\nNo datasets registered yet")

    elif args.create and args.csv:
        dataset_id = dm.create_from_csv(
            csv_path=args.csv,
            dataset_name=args.create,
            description=f"Created from {args.csv}",
        )

    elif args.info:
        info = dm.get_dataset(args.info)
        if info:
            print(f"\nDataset: {args.info}")
            for key, value in info.items():
                print(f"  {key}: {value}")
        else:
            print(f"\nDataset '{args.info}' not found")


Data Flow

1. Video Ingestion Flow

flowchart TB
    User(["<span style='color:#000'>User Command</span>"])
    Entry["<span style='color:#000'>run_ingestion.py<br/>Parse args: video_dir, profiles, backend</span>"]
    GetProfiles["<span style='color:#000'>Resolve profiles<br/>(--profile or config default:<br/>video_colpali_smol500_mv_frame)</span>"]
    Build["<span style='color:#000'>Build Pipeline<br/>Test → build_test_pipeline()<br/>Simple → build_simple_pipeline()<br/>Advanced → create_pipeline().with_*().build()</span>"]
    Discover["<span style='color:#000'>Discover content files<br/>video_dir.glob('*.mp4') etc.</span>"]
    Concurrent["<span style='color:#000'>process_videos_concurrent<br/>max_concurrent=3 (default), async</span>"]
    PerVideo["<span style='color:#000'>Per video:<br/>Extract frames/chunks →<br/>Generate embeddings →<br/>Build documents →<br/>Upload to Vespa</span>"]
    Collect["<span style='color:#000'>Collect Results<br/>successful/failed, docs fed,<br/>time, throughput</span>"]
    Loop{"<span style='color:#000'>More profiles?</span>"}
    Summary["<span style='color:#000'>Print Summary<br/>Per-profile stats, overall success rate,<br/>total documents processed</span>"]

    User --> Entry --> GetProfiles --> Build --> Discover --> Concurrent --> PerVideo --> Collect --> Loop
    Loop -->|yes, next profile| Build
    Loop -->|no| Summary

    style User fill:#90caf9,stroke:#1565c0,color:#000
    style Entry fill:#90caf9,stroke:#1565c0,color:#000
    style GetProfiles fill:#b0bec5,stroke:#546e7a,color:#000
    style Build fill:#b0bec5,stroke:#546e7a,color:#000
    style Discover fill:#b0bec5,stroke:#546e7a,color:#000
    style Concurrent fill:#ffcc80,stroke:#ef6c00,color:#000
    style PerVideo fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Collect fill:#a5d6a7,stroke:#388e3c,color:#000
    style Loop fill:#b0bec5,stroke:#546e7a,color:#000
    style Summary fill:#a5d6a7,stroke:#388e3c,color:#000

2. Schema Deployment Flow

Per-tenant deployment flows through the runtime admin API rather than a bulk script; single-schema JSON deploys continue to work via deploy_json_schema.py for dev iteration.

flowchart TB
    Start(["<span style='color:#000'>User / init-job</span>"])
    Post["<span style='color:#000'>POST $RUNTIME_URL/admin/profiles/{profile}/deploy<br/>body: {tenant_id, force: false}</span>"]
    Deploy["<span style='color:#000'>SchemaRegistry.deploy_schema<br/>(tenant_id, base_schema_name)</span>"]
    Transform["<span style='color:#000'>Transform base schema →<br/>tenant-scoped schema name</span>"]
    Merge["<span style='color:#000'>Merge registry schemas + live Vespa<br/>document types<br/>(VespaBackend.deploy_schemas)</span>"]
    Package["<span style='color:#000'>ApplicationPackage<br/>allow_schema_removal=False<br/>(refuses to drop peer-tenant schemas)</span>"]
    VespaDeploy["<span style='color:#000'>_deploy_package → Vespa config server<br/>POST /application/v2/tenant/default/prepareandactivate</span>"]
    Converge["<span style='color:#000'>_wait_for_schema_convergence<br/>(serviceconverge generation wait + per-new-schema feed round-trip; raises on timeout)</span>"]
    DevDeploy["<span style='color:#000'>Single-schema dev deploy:<br/>deploy_json_schema.py → ApplicationPackage → Vespa</span>"]

    Start --> Post --> Deploy --> Transform --> Merge --> Package --> VespaDeploy --> Converge
    Start -.->|dev iteration| DevDeploy

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Post fill:#b0bec5,stroke:#546e7a,color:#000
    style Deploy fill:#ffcc80,stroke:#ef6c00,color:#000
    style Transform fill:#ffcc80,stroke:#ef6c00,color:#000
    style Merge fill:#ffcc80,stroke:#ef6c00,color:#000
    style Package fill:#ce93d8,stroke:#7b1fa2,color:#000
    style VespaDeploy fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Converge fill:#a5d6a7,stroke:#388e3c,color:#000
    style DevDeploy fill:#b0bec5,stroke:#546e7a,color:#000

3. Experiment Workflow Flow

flowchart TB
    Start(["<span style='color:#000'>User Command</span>"])
    Entry["<span style='color:#000'>run_experiments_with_visualization.py</span>"]
    Init["<span style='color:#000'>Initialize ExperimentTracker(tenant_id, ...)<br/>cogniverse_evaluation.core.experiment_tracker<br/>Separate 'experiments' project</span>"]
    Dataset["<span style='color:#000'>create_or_get_dataset()<br/>Load or create dataset from CSV,<br/>register with Phoenix</span>"]
    Config["<span style='color:#000'>get_experiment_configurations()<br/>Filter profiles/strategies,<br/>build profile × strategy matrix</span>"]

    subgraph PerCombo["<span style='color:#000'>FOR each profile × strategy</span>"]
        Create["<span style='color:#000'>Create Experiment<br/>Name: '{profile} - {strategy}', attach to dataset</span>"]
        RunQueries["<span style='color:#000'>Run Search Queries<br/>FOR each query: execute search,<br/>record spans, collect results</span>"]
        Evaluate["<span style='color:#000'>Evaluate Results<br/>quality: relevance/diversity/distribution/temporal<br/>llm: reference-free/reference-based</span>"]
        Store["<span style='color:#000'>Store Results<br/>experiment id, status, scores, span traces</span>"]
        Create --> RunQueries --> Evaluate --> Store
    end

    Viz["<span style='color:#000'>Generate Visualizations<br/>Profile summary, strategy comparison,<br/>detailed results (tabulate)</span>"]
    Export["<span style='color:#000'>Export Results<br/>CSV/JSON under outputs/experiment_results/,<br/>HTML report if quantitative tests exist</span>"]
    Links["<span style='color:#000'>Print Phoenix UI Links<br/>dataset, experiments project, default project</span>"]
    Exit(["<span style='color:#000'>Exit with Summary Statistics</span>"])

    Start --> Entry --> Init --> Dataset --> Config --> PerCombo
    PerCombo --> Viz --> Export --> Links --> Exit

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Entry fill:#90caf9,stroke:#1565c0,color:#000
    style Init fill:#ffcc80,stroke:#ef6c00,color:#000
    style Dataset fill:#ffcc80,stroke:#ef6c00,color:#000
    style Config fill:#ffcc80,stroke:#ef6c00,color:#000
    style Create fill:#ce93d8,stroke:#7b1fa2,color:#000
    style RunQueries fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Evaluate fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Store fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Viz fill:#ffcc80,stroke:#ef6c00,color:#000
    style Export fill:#a5d6a7,stroke:#388e3c,color:#000
    style Links fill:#a5d6a7,stroke:#388e3c,color:#000
    style Exit fill:#a5d6a7,stroke:#388e3c,color:#000

4. Optimization & Deployment Flow

flowchart TB
    Start(["<span style='color:#000'>User Command</span>"])
    Entry["<span style='color:#000'>python -m cogniverse_runtime.optimization_cli<br/>--mode &lt;MODE&gt; --tenant-id &lt;TENANT&gt;</span>"]
    Step1["<span style='color:#000'>Step 1: Run Per-Agent Optimizer<br/>Mode: simba | gateway-thresholds |<br/>entity-extraction | workflow | profile<br/>Collects Phoenix spans; compiles the target<br/>DSPy module with BootstrapFewShot,<br/>scaled by trainset size (&lt;50 vs &gt;=50 examples)</span>"]
    Step2["<span style='color:#000'>Step 2: Persist Artifact<br/>ArtifactManager.save_blob<br/>Stores optimized module / threshold config<br/>per tenant, returns an artifact_id</span>"]
    Step3["<span style='color:#000'>Step 3: Agents Load at Startup<br/>No redeploy: runtime agents read the<br/>latest tenant artifact via ArtifactManager</span>"]
    Summary["<span style='color:#000'>Print Summary<br/>Mode, tenant, artifact_id, status</span>"]
    Exit(["<span style='color:#000'>Exit with Status Code</span>"])

    Start --> Entry --> Step1 --> Step2 --> Step3 --> Summary --> Exit

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Entry fill:#90caf9,stroke:#1565c0,color:#000
    style Step1 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Step2 fill:#ffcc80,stroke:#ef6c00,color:#000
    style Step3 fill:#a5d6a7,stroke:#388e3c,color:#000
    style Summary fill:#a5d6a7,stroke:#388e3c,color:#000
    style Exit fill:#a5d6a7,stroke:#388e3c,color:#000

Usage Examples

Example 1: Basic Video Ingestion

# Process videos with default profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/testset/evaluation/sample_videos \
  --backend vespa

# Output:
# ============================================================
# 🎯 Processing with profile: video_colpali_smol500_mv_frame
# ============================================================
# 🎬 Starting Video Processing Pipeline
# 📁 Video directory: data/testset/evaluation/sample_videos
# 📂 Output directory: outputs/ingestion/video_colpali_smol500_mv_frame
# 🔧 Backend: vespa
# 📹 Found 3 videos to process
#
# ✅ Profile video_colpali_smol500_mv_frame completed!
#    Time: 45.32 seconds
#    Videos: 3/3 successful
#    Documents fed: 180
#    Throughput: 4.0 docs/sec
#    Avg per video: 15.11 seconds

Example 2: Multi-Profile Ingestion

# Process with multiple profiles simultaneously
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
  --video_dir data/testset/evaluation/sample_videos \
  --backend vespa \
  --profile video_colpali_smol500_mv_frame \
           video_xclip_sv_chunk_6s \
           video_colqwen_omni_mv_chunk_30s

# Output shows processing for each profile:
# ============================================================
# 🎯 Processing with profile: video_colpali_smol500_mv_frame
# ============================================================
# ...
# ✅ Profile video_colpali_smol500_mv_frame completed!
#
# ============================================================
# 🎯 Processing with profile: video_xclip_sv_chunk_6s
# ============================================================
# ...
# ✅ Profile video_xclip_sv_chunk_6s completed!
#
# ============================================================
# 📊 Overall Summary
# ============================================================
# Processed 3 profiles
# ✅ video_colpali_smol500_mv_frame: 3/3 videos succeeded, 180 docs in 45.3s
# ✅ video_xclip_sv_chunk_6s: 3/3 videos succeeded, 90 docs in 38.1s
# ⚠️ video_colqwen_omni_mv_chunk_30s: 2/3 videos succeeded, 120 docs in 52.7s

Example 3: Test Mode Ingestion

# Quick validation with limited frames
uv run python scripts/run_ingestion.py \
  --video_dir data/testset/evaluation/sample_videos \
  --backend vespa \
  --test-mode \
  --max-frames 10

# Output:
# 🧪 Using test pipeline builder...
# 🖼️ Max frames: 10
# ✅ Profile video_colpali_smol500_mv_frame completed!
#    Time: 12.45 seconds
#    Videos: 3/3 successful
#    Documents fed: 30  # Only 10 frames per video

Example 4: Schema Deployment

# Deploy single schema
python scripts/deploy_json_schema.py \
  configs/schemas/video_colpali_smol500_mv_frame_schema.json

# Output:
# ============================================================
# Vespa JSON Schema Deployment
# ============================================================
# Schema file: configs/schemas/video_colpali_smol500_mv_frame_schema.json
# Config server: localhost:19071
# Data endpoint: localhost:8080
#
# 📄 Loading schema from video_colpali_smol500_mv_frame_schema.json
# 📦 Processing schema: video_colpali_smol500_mv_frame
# 🚀 Deploying to http://localhost:19071/application/v2/tenant/default/prepareandactivate...
# ✅ Schema 'video_colpali_smol500_mv_frame' deployed successfully!
#
# ⏳ Waiting for deployment to propagate...
#
# 🔍 Verifying 'video_colpali_smol500_mv_frame' deployment...
# ✅ Vespa is running and responding
#
# ============================================================
# Deployment complete!
# ============================================================

# Deploy tenant schemas via the runtime admin API (single profile)
RUNTIME_URL=http://localhost:8080
curl -sfX POST "$RUNTIME_URL/admin/tenants" \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id": "acme:production"}'
curl -sfX POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id": "acme:production", "force": false}'

Example 5: Phoenix Experiments

# Run experiments with quality evaluators
uv run python scripts/run_experiments_with_visualization.py \
  --tenant-id acme:acme \
  --dataset-name golden_eval_v1 \
  --dataset-path data/testset/evaluation/video_search_queries.csv \
  --profiles frame_based_colpali \
  --quality-evaluators

# Output:
# ================================================================================
# PHOENIX EXPERIMENTS WITH VISUALIZATION
# ================================================================================
#
# Timestamp: 2025-10-07 14:30:00
# Experiment Project: experiments (separate from default traces)
# Quality Evaluators: ✅ ENABLED (relevance, diversity, distribution, temporal coverage)
# LLM Evaluators: ❌ DISABLED
#
# Preparing experiment dataset...
# ✅ Dataset ready: http://localhost:6006/datasets/ds_abc123
#
# ============================================================
# Profile: frame_based_colpali
# ============================================================
#
# [1/6] Frame Based Colpali - Binary
#   Strategy: binary_binary
#   ✅ Success
#
# [2/6] Frame Based Colpali - Float
#   Strategy: float_float
#   ✅ Success
#
# [3/6] Frame Based Colpali - Phased
#   Strategy: phased
#   ✅ Success
#
# ...
#
# ================================================================================
# EXPERIMENT RESULTS VISUALIZATION
# ================================================================================
#
# 📊 PROFILE SUMMARY
# ------------------------------------------------------------
# +---------------------+-------+---------+--------+---------------+
# | Profile             | Total | Success | Failed | Success Rate  |
# +=====================+=======+=========+========+===============+
# | frame_based_colpali |     6 |       6 |      0 | 100.0%        |
# +---------------------+-------+---------+--------+---------------+
#
# 🔍 STRATEGY COMPARISON BY PROFILE
# ------------------------------------------------------------
#
# frame_based_colpali:
# Strategy              Description         Status
# --------------------  ------------------  -------------
# binary_binary         Binary              ✅ Success
# float_float           Float               ✅ Success
# float_binary          Float-Binary        ✅ Success
# phased                Phased              ✅ Success
# hybrid_binary_bm25    Hybrid + Desc       ✅ Success
# bm25_only             Text Only           ✅ Success
#
# ================================================================================
# SUMMARY STATISTICS
# ================================================================================
#
# Total Experiments Attempted: 6
# Successful: 6 (100.0%)
# Failed: 0 (0.0%)
#
# ================================================================================
# VIEW IN PHOENIX UI
# ================================================================================
#
# 🔗 Dataset: http://localhost:6006/datasets/ds_abc123
# 🔗 Experiments Project: http://localhost:6006/projects/experiments
# 🔗 Default Project (spans): http://localhost:6006/projects/default
#
# ℹ️  Notes:
#   - Experiments are in separate 'experiments' project
#   - Each experiment has its own traces with detailed spans
#   - Use Phoenix UI to compare experiments side-by-side
#   - Evaluation scores are attached to each experiment
#
# 💾 Results saved to: outputs/experiment_results/experiment_summary_20251007_143000.csv
# 💾 Detailed results saved to: outputs/experiment_results/experiment_details_20251007_143000.json
#
# ✅ All experiments completed!

Example 6: Per-Agent Optimization

# Optimize each agent mode for a tenant
uv run python -m cogniverse_runtime.optimization_cli --mode simba --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode entity-extraction --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id acme_corp

# Clean up old logs afterward
uv run python -m cogniverse_runtime.optimization_cli --mode cleanup --log-retention-days 7

Example 7: Dataset Management

# List all datasets
python scripts/manage_datasets.py --tenant-id acme:acme --list

# Output:
# Registered datasets:
#
# Name: golden_eval_v1
#   Dataset ID: ds_abc123
#   Created: 2025-10-05 10:30:00
#   Examples: 50
#   Description: Golden evaluation dataset v1
#
# Name: video_search_test
#   Dataset ID: ds_def456
#   Created: 2025-10-06 14:15:00
#   Examples: 25
#   Description: Test queries for video search

# Create new dataset
python scripts/manage_datasets.py \
  --tenant-id acme:acme \
  --create my_queries \
  --csv data/my_queries.csv

# Output:
# Dataset 'my_queries' created with ID: ds_ghi789

# Get dataset info
python scripts/manage_datasets.py --tenant-id acme:acme --info golden_eval_v1

# Output:
# Dataset: golden_eval_v1
#   dataset_id: ds_abc123
#   created_at: 2025-10-05 10:30:00
#   num_examples: 50
#   description: Golden evaluation dataset v1

Example 8: Web Client

# Deployed by `cogniverse up`
open http://localhost:28400

# The sidebar lists every agent (chat) and the Operations views:
# Tenants, Backend profiles, Configuration, Ingestion, Optimization runs,
# Memory, Approvals, Annotation queue, Workflow reviews, Profile metrics,
# RLM A/B, Analytics, Evaluation, Embedding atlas, Routing evaluation

See Web Client to run it locally.


Production Considerations

1. Performance Optimization

Ingestion Throughput:

# Adjust concurrency based on available resources
--max-concurrent 5  # More concurrent videos (CPU/memory intensive)
--max-concurrent 1  # Serial processing (safer for limited resources)

# Throughput metrics:
# - ColPali frame-based: ~4-5 docs/sec (GPU)
# - X-CLIP global: ~2-3 docs/sec (GPU)
# - ColQwen chunk-based: ~3-4 docs/sec (GPU)

Async Processing:

# Use async methods for I/O-bound operations
results = await pipeline.process_videos_concurrent(
    video_files,
    max_concurrent=3
)

# Benefits:
# - Non-blocking video processing
# - Better CPU utilization
# - Faster overall throughput

Schema Deployment:

# Production: init-job loops tenant × profile against the runtime admin
# API — see charts/cogniverse/templates/init-jobs.yaml. The runtime
# merges existing Vespa document types into every redeploy so peer
# tenants never get dropped.

# Dev: iterate a single schema file through deploy_json_schema.py.

2. Error Handling

Ingestion Errors:

# Pipeline continues on individual video failures
result = {
    'status': 'failed',
    'error': 'Embedding generation failed',
    'video_path': str(video_path)
}

# Overall statistics still reported:
# ⚠️ Profile frame_based_colpali partially completed!
#    Time: 52.7 seconds
#    Videos: 2/3 successful
#    Failed: 1 videos

Schema Deployment Errors:

# Common errors:
# - "Connection refused" → Vespa not running
# - "HTTP 400" → Schema validation failed
# - "Timeout" → Vespa busy or unresponsive

# deploy_json_schema() failures (non-200 from prepareandactivate, or an
# exception during zip/POST) DO exit non-zero:
if deploy_json_schema(args.schema_file, args.config_host, args.config_port, args.data_port):
    ...
else:
    print("\n❌ Deployment failed")
    sys.exit(1)

# verify_deployment() is called afterward as an informational health
# check (prints a warning on non-200) but its return value is NOT
# checked — a failed verification does not change the exit code.

Experiment Errors:

# Experiments continue on individual strategy failures
if result["status"] == "failed":
    error = result.get("error", "Unknown error")
    if "Text encoder not available" in error:
        print(f"  ⚠️  Skipped: Encoder not available")
    else:
        print(f"  ❌ Failed: {error[:50]}...")

# Final summary shows partial success:
# Total: 10, Successful: 8 (80.0%), Failed: 2 (20.0%)

3. Monitoring and Logging

Ingestion Logs:

# Logs written to outputs/logs/
# - ingestion_pipeline.log  # Main pipeline logs
# - video_processing.log    # Per-video processing
# - embedding_generation.log  # Embedding logs
# - vespa_upload.log        # Upload logs

# View logs:
tail -f outputs/logs/ingestion_pipeline.log

Experiment Logs:

# Experiment results saved to:
# - outputs/experiment_results/experiment_summary_*.csv
# - outputs/experiment_results/experiment_details_*.json

# Phoenix spans capture:
# - Query execution time
# - Search latency
# - Embedding generation time
# - Evaluation scores

# View in Phoenix UI: http://localhost:6006

4. Resource Management

Memory Management:

# Video processing is memory-intensive
# Estimated memory per video:
# - Frame extraction: ~200-500 MB
# - Embedding generation: ~1-2 GB (GPU)
# - Document building: ~50-100 MB

# Total concurrent memory:
# max_concurrent=3 → ~4-6 GB peak memory

# Recommendations:
# - 16 GB RAM minimum for production
# - 32 GB RAM recommended for max_concurrent > 5

GPU Utilization:

# GPU required for:
# - ColPali embedding generation
# - X-CLIP encoding
# - ColQwen embedding generation

# GPU memory requirements:
# - ColPali Smol 500M: ~2 GB VRAM
# - X-CLIP: ~4 GB VRAM
# - ColQwen Omni: ~6 GB VRAM

# For multiple profiles:
# - Process sequentially (GPU memory limits)
# - Or use multiple GPUs with CUDA_VISIBLE_DEVICES

Disk Space:

# Storage requirements per video:
# - Original video: ~50-500 MB
# - Extracted frames: ~10-50 MB
# - Transcripts: ~10-50 KB
# - Embeddings (binary): ~100-500 KB
# - Embeddings (float): ~1-5 MB

# Total for 1000 videos: ~100-500 GB

5. Scaling Strategies

Horizontal Scaling:

# Ingestion:
# - Run multiple ingestion processes
# - Partition videos by directory
# - Each process handles different profiles

# Example:
# Process 1: --profile video_colpali_smol500_mv_frame
# Process 2: --profile video_xclip_sv_chunk_6s
# Process 3: --profile video_colqwen_omni_mv_chunk_30s

Vertical Scaling:

# Increase concurrency with more resources:
--max-concurrent 10  # Requires 32+ GB RAM, 16+ CPU cores

# Benefits:
# - Faster overall throughput
# - Better resource utilization
# - Reduced total processing time

Batch Processing:

# Process large video collections in batches:
videos_per_batch = 100
for i in range(0, len(videos), videos_per_batch):
    batch = videos[i:i+videos_per_batch]
    run_ingestion(batch)

6. Best Practices

Schema Management:

1. Deploy per-tenant via the runtime admin API; the Helm init job
   (charts/cogniverse/templates/init-jobs.yaml) runs tenant × profile
   on install. Never bulk-deploy with allow_schema_removal overrides
   — that path silently drops peer-tenant schemas.
2. Version-control all schema JSON files in configs/schemas/.
3. Test schema changes in staging before production.
4. Ranking strategies are extracted at search time (no separate step).
5. Validate schemas before deployment with deploy_json_schema.py
   against a local Vespa container.

Ingestion Pipeline:

# 1. Test with --test-mode first (max_frames=10)
# 2. Process small batches before full ingestion
# 3. Monitor logs for errors during processing
# 4. Verify documents in Vespa after ingestion
# 5. Use appropriate backend (vespa for production)

Experiments:

# 1. Use separate "experiments" project in Phoenix
# 2. Enable quality evaluators for comprehensive metrics
# 3. Save results to CSV/JSON for analysis
# 4. Compare experiments side-by-side in Phoenix UI
# 5. Export results before deleting experiments

7. Common Issues and Solutions

Issue: "Video processing failed: CUDA out of memory"

# Solution: Reduce concurrency or batch size
--max-concurrent 1  # Process one video at a time
--max-frames 50     # Reduce frames per video

Issue: "Schema deployment failed: Connection refused"

# Solution: Ensure Vespa is running
docker ps | grep vespa  # Check Vespa container
cogniverse up  # Start Vespa if not running


Summary

The Scripts & Operations module provides comprehensive tooling for:

  1. Video Ingestion: Builder pattern pipeline with multi-profile support
  2. Schema Management: JSON-based deployment with validation
  3. Optimization: Complete DSPy optimization and deployment workflow
  4. Experimentation: Phoenix experiments with quality evaluators
  5. Dataset Management: CRUD operations for evaluation datasets
  6. System Setup: Environment initialization and dependency checking

Key Design Patterns:

  • Builder pattern for flexible pipeline configuration

  • Async processing for concurrent video handling

  • Command pattern for CLI tool design

  • Context manager for experiment runner lifecycle

  • Factory pattern for strategy resolution

Production Features:

  • Concurrent async video processing

  • Multi-profile simultaneous ingestion

  • Phoenix experiment tracking with visualization

  • Comprehensive error handling and logging

  • Resource-aware concurrency limits

This module serves as the operational backbone of the Cogniverse system, providing production-grade tools for deployment, ingestion, optimization, and monitoring.


Related Documentation: