Cogniverse Study Guide: Scripts & Operations Module¶
Module Path: scripts/ SDK Packages: Uses all 12 packages (foundation → core → implementation → application)
Table of Contents¶
Module Overview¶
Purpose¶
The Scripts & Operations module provides command-line tools for:
-
Video Ingestion: Processing and indexing video content
-
Schema Deployment: Managing Vespa search schemas
-
System Setup: Initializing the system environment
-
Optimization: Running DSPy optimization workflows
-
Experimentation: Conducting Phoenix experiments with visualization
-
Dataset Management: Managing evaluation datasets
Key Features¶
- Builder Pattern Ingestion: Fluent API for configurable pipeline construction
- Multi-Profile Support: Process videos with multiple embedding strategies simultaneously
- Tenant-Aware Processing: Per-tenant schema isolation and configuration
- Async Processing: Concurrent video processing with configurable limits
- Phoenix Integration: Per-tenant experiment tracking with visual analytics
- Schema Management: JSON-based tenant-specific schema deployment
- UV Workspace: All scripts use
uv runfor SDK package management
Script Categories¶
scripts/
├── Ingestion & Processing
│ ├── run_ingestion.py # Main video ingestion pipeline
│ ├── test_ingestion.py # Test ingestion with validation
│ └── backfill_source_url.py # Backfill source URLs on existing documents
│
├── Deployment & Setup
│ ├── deploy_json_schema.py # Deploy single JSON schema
│ ├── provision_tenant.py # Entry point for cogniverse_runtime.provision_tenant
│ ├── setup_ollama.py # Ollama model setup
│ └── setup_gliner.py # GLiNER setup
│ (bulk schema deploy flows through the runtime admin API:
│ POST /admin/profiles/{profile}/deploy — see charts/cogniverse/
│ templates/init-jobs.yaml for the in-cluster init job.)
│
├── Optimization & Experiments
│ ├── run_experiments_with_visualization.py # Phoenix experiments
│ └── auto_optimization_trigger.py # Automated optimization trigger
│
├── Dataset Management
│ ├── manage_datasets.py # Create/list/inspect datasets (tenant-scoped)
│ ├── manage_golden_datasets.py # Golden dataset management
│ ├── create_golden_dataset_from_traces.py # Golden dataset from traces
│ └── seed_bright_corpus.py # Seed corpus data
│
├── Reporting & Analysis
│ ├── generate_integrated_evaluation_report.py # Integrated evaluation report
│ ├── generate_tabbed_html_report.py # Tabbed HTML report generation
│ ├── generate_langextract_training_data.py # Training data for LangExtract
│ ├── view_integrated_results.py # View integrated results
│ └── vlm_caption_bakeoff.py # VLM caption comparison
│
├── Utilities & Operations
│ ├── discover_tenants.py # Tenant discovery
│ ├── export_backend_embeddings.py # Backend embedding export (tenant-aware)
│ ├── manage_phoenix_data.py # Phoenix data management
│ ├── prune_config_metadata.py # Config metadata pruning
│ ├── release_manifest.py # Stage pinned release sources, validate builds, write dist/BUILD_MANIFEST.json, check publishability
│ ├── start_phoenix.py # Start Phoenix service
│ └── version_bump.py # Version management
│
└── Shell Scripts (build, local dev, test infrastructure)
├── build_packages.sh # Build the release package set in dependency order
├── publish_packages.sh # Publish SDK packages to (Test)PyPI
├── install_with_gpu.sh # uv sync with the PyTorch backend matching local hardware
├── download_test_data.sh # Download the evaluation dataset
├── start_vespa.sh / stop_vespa.sh # Vespa container lifecycle (persistent volumes)
├── start_phoenix.sh # Start/stop/restart Phoenix via Docker
├── setup_evaluation.sh # Set up the Cogniverse evaluation framework
├── setup_local_tests.sh # Probe every endpoint the local integration tests depend on
├── run_local_tests.sh # Run comprehensive local-only test coverage (skipped in CI)
├── run_integration_per_package.sh # Run integration tests one package at a time (isolated processes)
├── run_e2e_batched.sh # Run the e2e suite in two batches with a runtime restart between them
├── record_test_goldens.sh # Re-record integration-test goldens against the live k3d cluster
├── analyze_vespa_embeddings.sh # Export Vespa embeddings and run an integrity check
├── memsnap_sampler.sh # Sample /admin/debug/memsnap at fixed intervals
└── sweep_state_snapshot.sh # Snapshot cluster + Vespa + Mem0 registry state during a sweep
Architecture¶
1. Ingestion Pipeline Architecture¶
flowchart TB
Entry["<span style='color:#000'>run_ingestion.py<br/>Command Line Entry Point</span>"]
TestMode["<span style='color:#000'>Test Mode<br/>build_test_pipeline</span>"]
SimpleMode["<span style='color:#000'>Simple Mode<br/>build_simple_pipeline</span>"]
AdvMode["<span style='color:#000'>Advanced Mode<br/>create_pipeline</span>"]
Pipeline["<span style='color:#000'>VideoIngestionPipeline<br/>• Video Processing<br/>• Embedding Generation<br/>• Vespa Upload</span>"]
ColPali["<span style='color:#000'>ColPali Profile<br/>Frame-based</span>"]
X-CLIP["<span style='color:#000'>X-CLIP Profile<br/>Global embeddings</span>"]
ColQwen["<span style='color:#000'>ColQwen Profile<br/>Chunk-based</span>"]
Concurrent["<span style='color:#000'>Concurrent Async<br/>Video Processing<br/>max_concurrent=3</span>"]
Vespa["<span style='color:#000'>Vespa Backend<br/>Bulk Upload</span>"]
Entry --> TestMode
Entry --> SimpleMode
Entry --> AdvMode
TestMode --> Pipeline
SimpleMode --> Pipeline
AdvMode --> Pipeline
Pipeline --> ColPali
Pipeline --> X-CLIP
Pipeline --> ColQwen
ColPali --> Concurrent
X-CLIP --> Concurrent
ColQwen --> Concurrent
Concurrent --> Vespa
style Entry fill:#90caf9,stroke:#1565c0,color:#000
style TestMode fill:#b0bec5,stroke:#546e7a,color:#000
style SimpleMode fill:#b0bec5,stroke:#546e7a,color:#000
style AdvMode fill:#b0bec5,stroke:#546e7a,color:#000
style Pipeline fill:#ffcc80,stroke:#ef6c00,color:#000
style ColPali fill:#ce93d8,stroke:#7b1fa2,color:#000
style X-CLIP fill:#ce93d8,stroke:#7b1fa2,color:#000
style ColQwen fill:#ce93d8,stroke:#7b1fa2,color:#000
style Concurrent fill:#ffcc80,stroke:#ef6c00,color:#000
style Vespa fill:#a5d6a7,stroke:#388e3c,color:#000 2. Optimization Workflow Architecture¶
flowchart TB
Start["<span style='color:#000'>optimization_cli module<br/>Per-Agent Optimization CLI</span>"]
Step1["<span style='color:#000'>Step 1: Run Agent Optimization<br/>• Execute per-agent optimizer<br/>• Modes: simba/gateway-thresholds/entity-extraction/etc.</span>"]
Step2["<span style='color:#000'>Step 2: Persist Artifact<br/>• ArtifactManager.save_blob (per tenant)<br/>• Returns artifact_id</span>"]
Step3["<span style='color:#000'>Step 3: Agents Load at Startup<br/>• Runtime agents read latest tenant artifact<br/>• No redeploy step</span>"]
Start --> Step1
Step1 --> Step2
Step2 --> Step3
style Start fill:#90caf9,stroke:#1565c0,color:#000
style Step1 fill:#ffcc80,stroke:#ef6c00,color:#000
style Step2 fill:#ffcc80,stroke:#ef6c00,color:#000
style Step3 fill:#a5d6a7,stroke:#388e3c,color:#000 3. Experiment Workflow Architecture¶
flowchart TB
Start["<span style='color:#000'>run_experiments_with_visualization.py<br/>Phoenix Experiment Runner</span>"]
Runner["<span style='color:#000'>ExperimentTracker<br/>• From cogniverse_evaluation SDK<br/>• Experiment project isolation<br/>• Quality evaluators optional<br/>• LLM evaluators optional</span>"]
Dataset["<span style='color:#000'>Dataset Preparation<br/>• Load or create dataset<br/>• CSV parsing<br/>• Phoenix dataset registration</span>"]
Loop["<span style='color:#000'>Multi-Profile Multi-Strategy Loop<br/>FOR each profile:<br/> FOR each strategy:<br/> • Run experiment<br/> • Track spans<br/> • Evaluate results<br/> • Store metrics</span>"]
Viz["<span style='color:#000'>Visualization Generation<br/>• Profile summary table<br/>• Strategy comparison<br/>• Detailed results<br/>• HTML report optional</span>"]
Export["<span style='color:#000'>Results Export<br/>• CSV summary<br/>• JSON detailed results<br/>• Phoenix UI links</span>"]
Start --> Runner
Runner --> Dataset
Dataset --> Loop
Loop --> Viz
Viz --> Export
style Start fill:#90caf9,stroke:#1565c0,color:#000
style Runner fill:#ffcc80,stroke:#ef6c00,color:#000
style Dataset fill:#ffcc80,stroke:#ef6c00,color:#000
style Loop fill:#ffcc80,stroke:#ef6c00,color:#000
style Viz fill:#ffcc80,stroke:#ef6c00,color:#000
style Export fill:#a5d6a7,stroke:#388e3c,color:#000 4. Schema Deployment Architecture¶
flowchart TB
subgraph Single["<span style='color:#000'>Single Schema Deployment</span>"]
Single1["<span style='color:#000'>deploy_json_schema.py<br/>Single Schema Deployment</span>"]
Parser["<span style='color:#000'>JsonSchemaParser<br/>• Load JSON schema file<br/>• Parse to Vespa Schema object<br/>• Validate structure</span>"]
Package1["<span style='color:#000'>ApplicationPackage<br/>• Create package<br/>• Add schema<br/>• Generate ZIP</span>"]
Deploy1["<span style='color:#000'>HTTP Deployment<br/>• POST to config server<br/>• Port: 19071<br/>• Endpoint: prepareandactivate</span>"]
Verify["<span style='color:#000'>Verification<br/>• Check ApplicationStatus<br/>• Verify Vespa responding</span>"]
Single1 --> Parser
Parser --> Package1
Package1 --> Deploy1
Deploy1 --> Verify
end
subgraph Multi["<span style='color:#000'>Multi-Tenant Deployment</span>"]
Multi1["<span style='color:#000'>Runtime admin API<br/>POST /admin/profiles/{profile}/deploy</span>"]
Registry["<span style='color:#000'>SchemaRegistry.deploy_schema<br/>• Per-tenant scoping<br/>• Merge with live cluster via<br/> list_deployed_document_types()</span>"]
Package2["<span style='color:#000'>ApplicationPackage<br/>• Merged schemas (registry + Vespa)<br/>• allow_schema_removal=False<br/> (fail-loud on unresolved drops)</span>"]
Multi1 --> Registry
Registry --> Package2
end
style Single1 fill:#90caf9,stroke:#1565c0,color:#000
style Multi1 fill:#90caf9,stroke:#1565c0,color:#000
style Parser fill:#ffcc80,stroke:#ef6c00,color:#000
style Package1 fill:#ffcc80,stroke:#ef6c00,color:#000
style Registry fill:#ffcc80,stroke:#ef6c00,color:#000
style Package2 fill:#ffcc80,stroke:#ef6c00,color:#000
style Deploy1 fill:#ce93d8,stroke:#7b1fa2,color:#000
style Verify fill:#a5d6a7,stroke:#388e3c,color:#000 Core Scripts¶
1. run_ingestion.py¶
Purpose: Main entry point for video ingestion pipeline with builder pattern configuration
Location: scripts/run_ingestion.py (286 lines)
Command Line Arguments:
--tenant-id TENANT # Tenant ID for schema isolation (required — no default)
--video_dir PATH # Directory containing content files
--content-dir PATH # Directory containing content files (alias for --video_dir)
--media-root-uri URI # Non-filesystem source (e.g. s3://, pvc://); overrides --video_dir
--output_dir PATH # Output directory for processed data
--backend {byaldi,vespa} # Search backend (default: vespa)
--profile PROFILES # Processing profiles (space-separated)
--content-type {video,image,audio,document} # Content type (default: video)
--max-concurrent INT # Max concurrent items to process (default: 3)
--max-frames INT # Maximum frames per video / images per batch
--test-mode # Use test mode with limited frames
--debug # Enable debug mode
Usage Modes:
Test Mode (for quick validation):
# Use test pipeline builder (tenant_id is required — no default)
pipeline = build_test_pipeline(
tenant_id="acme_corp",
video_dir=Path("data/testset/evaluation/sample_videos"),
schema="video_colpali_smol500_mv_frame",
max_frames=10
)
Simple Mode (for standard usage):
# Use simple pipeline builder (tenant_id is required — no default)
pipeline = build_simple_pipeline(
tenant_id="acme_corp",
video_dir=Path("data/testset/evaluation/sample_videos"),
schema="video_colpali_smol500_mv_frame",
backend="vespa",
debug=False
)
Advanced Mode (for custom configuration):
# Use fluent builder with custom config
config = (create_config()
.video_dir(Path("data/videos"))
.backend("vespa")
.output_dir(Path("custom/output"))
.max_frames_per_video(100)
.build())
from cogniverse_foundation.config.utils import create_default_config_manager
config_manager = create_default_config_manager()
# with_tenant_id() and with_config_manager() are both required before
# build() — the builder raises ValueError if either is missing.
pipeline = (create_pipeline()
.with_tenant_id("acme_corp")
.with_config_manager(config_manager)
.with_config(config)
.with_schema("video_colpali_smol500_mv_frame")
.with_debug(True)
.with_concurrency(5)
.build())
Multi-Profile Processing (Tenant-Aware):
# Process with multiple profiles simultaneously for specific tenant
# Uses builder pattern from runtime package
from cogniverse_runtime.ingestion.pipeline_builder import build_simple_pipeline
for profile in ["video_colpali_smol500_mv_frame",
"video_xclip_sv_chunk_6s"]:
pipeline = build_simple_pipeline(
tenant_id="acme_corp", # Tenant-specific processing
video_dir=Path("data/videos"),
schema=profile,
backend="vespa"
)
results = await pipeline.process_videos_concurrent(
video_files,
max_concurrent=3
)
Output:
-
Success/failure status per video
-
Documents fed to Vespa
-
Processing time and throughput
-
Per-profile summary statistics
2. deploy_json_schema.py¶
Purpose: Deploy individual JSON schema files to Vespa
Location: scripts/deploy_json_schema.py (196 lines)
Command Line Arguments:
schema_file # Path to JSON schema file (required)
--config-host HOST # Vespa config server host (default: localhost)
--config-port PORT # Config server port (default: 19071)
--data-host HOST # Vespa data endpoint host (default: localhost)
--data-port PORT # Data endpoint port (default: 8080)
Deployment Process:
def deploy_json_schema(schema_file, vespa_host, config_port, data_port):
# Imports from SDK packages
from cogniverse_vespa.json_schema_parser import JsonSchemaParser
from vespa.package import ApplicationPackage
# 1. Load JSON schema
with open(schema_file, 'r') as f:
schema_config = json.load(f)
# 2. Parse schema using JsonSchemaParser
parser = JsonSchemaParser()
schema = parser.parse_schema(schema_config)
# 3. Create application package
app_package = ApplicationPackage(name=schema.name.replace('_', ''))
app_package.add_schema(schema)
# 4. Deploy via HTTP
deploy_url = f"http://{vespa_host}:{config_port}/application/v2/tenant/default/prepareandactivate"
app_zip = app_package.to_zip()
response = requests.post(deploy_url, headers={"Content-Type": "application/zip"},
data=app_zip, timeout=60)
# 5. Verify deployment
verify_deployment(schema.name, vespa_host, data_port)
Verification:
def verify_deployment(schema_name, vespa_host, data_port):
# Check application status endpoint
response = requests.get(
f"http://{vespa_host}:{data_port}/ApplicationStatus",
timeout=5
)
return response.status_code == 200
Example Usage:
# Deploy agent memories schema
python scripts/deploy_json_schema.py \
configs/schemas/agent_memories_schema.json
# Deploy video schema
python scripts/deploy_json_schema.py \
configs/schemas/video_colpali_smol500_mv_frame_schema.json
# Deploy to remote Vespa instance
python scripts/deploy_json_schema.py \
configs/schemas/config_metadata_schema.json \
--config-host vespa.example.com \
--config-port 19071
3. Bulk per-tenant schema deployment (runtime admin API)¶
Schema deployment is always per-tenant and always flows through the runtime admin API so it goes through SchemaRegistry.deploy_schema and the VespaBackend.deploy_schemas merge path.
In-cluster path: charts/cogniverse/templates/init-jobs.yaml registers each .Values.config.tenants entry (200/201 created and 409 already registered both proceed; any other status fails the Job attempt) and deploys .Values.config.defaultProfiles.video (none when it is empty) for it by calling the runtime:
curl -sS -w ' HTTP_STATUS=%{http_code}' -X POST "$RUNTIME_URL/admin/tenants" \
-H "Content-Type: application/json" \
-d '{"tenant_id": "{{ $tenant.id }}", "created_by": "helm:{{ $.Chart.Name }}"}'
curl -X POST "$RUNTIME_URL/admin/profiles/{{ $profile }}/deploy" \
-H "Content-Type: application/json" \
-d '{"tenant_id": "{{ $tenant.id }}", "force": false}'
Local path:
RUNTIME_URL=http://localhost:8080
# Register tenant (deploys tenant_metadata etc.)
curl -X POST "$RUNTIME_URL/admin/tenants" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production"}'
# Deploy a profile's content schema for that tenant
curl -X POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production", "force": false}'
For a single-schema JSON deploy from the filesystem (dev workflow), scripts/deploy_json_schema.py still works.
3a. cogniverse admin reconcile-orphans¶
DELETE /admin/tenants/{id} refuses if a peer tenant has an unreconstructable Vespa-only orphan (the redeploy would silently drop the peer's schema). After power loss, SIGKILL between Vespa deploy and registry write, or any other interrupted-deploy path, an operator needs an out-of-band recovery tool — that's cogniverse admin reconcile-orphans.
# Dry-run (default): list orphan schemas + implied tenants, no changes
cogniverse admin reconcile-orphans
# Confirm: drop every orphan tenant in ONE atomic Vespa redeploy
cogniverse admin reconcile-orphans --confirm
# Point at a non-default runtime
cogniverse admin reconcile-orphans --runtime-url http://runtime.cogniverse.svc:28000
The CLI is a thin client over POST /admin/reconcile-orphans?dry_run=…, which internally calls VespaSchemaManager.delete_tenant_schemas_bulk. The endpoint reports:
orphan_schemas— Vespa-deployed names with no registry recordorphan_tenants— implied tenant ids (recovered by stripping known base prefixes)unrecovered_schemas— orphans whose base prefix isn't inKNOWN_BASES; need operator review
See operations/multi-tenant-ops.md#orphan-reconciliation for the full operator workflow and JSON shapes.
4. provision_tenant.py¶
Purpose: Cold-bootstrap a tenant's backend resources (Vespa schemas, Mem0 memory schema, Phoenix telemetry project, semantic-router tier) without a live runtime. The tenant-provisioning WorkflowTemplate runs it inside the runtime image.
Location: libs/runtime/cogniverse_runtime/provision_tenant.py, with scripts/provision_tenant.py as the checkout-side entry point.
Command Line Arguments:
--tenant-id TENANT # Tenant identifier (required)
--step STEP # Provisioning step: schemas|verify|memory|telemetry|tier (required)
--tier TIER # Router tier to store (required by --step tier)
--profiles LIST # Comma-separated backend profiles (required by --step schemas|verify)
Environment: every step needs BACKEND_URL / BACKEND_PORT (the Vespa data endpoint) because every step builds a ConfigManager over the tenant store; VESPA_CONFIG_PORT names the config server the schema deploy posts to. A failing step prints one line naming the cause and exits 1.
--profiles entries resolve through the tenant's merged backend catalog — the cluster profiles in configs/config.json with the tenant's stored overrides on top — so a tenant with no backend rows yet still provisions.
Steps:
schemas— Deploys the profiles' schemas throughSchemaRegistry.deploy_schemasas one application package, under their tenant-scoped names; a profile that cannot be loaded leaves none of them registeredverify— Confirms each profile's tenant schema carries a registry row and answers a YQL querymemory— Creates the tenant's Mem0 memory schema vialazy_init_memorytelemetry— Exports a required probe span viaTelemetryManager.required_span(...)to create the Phoenix project; an unreachable collector fails the steptier— Stores the tenant's semantic-router tier viaset_tenant_tier; a tier outsideROUTER_TIERSis refused before any store is built
Usage:
# Initialize memory schema for tenant
uv run python -m cogniverse_runtime.provision_tenant \
--tenant-id acme:production \
--step memory
# Initialize Phoenix telemetry project for tenant
uv run python -m cogniverse_runtime.provision_tenant \
--tenant-id acme:production \
--step telemetry
# Store the router tier for tenant
uv run python -m cogniverse_runtime.provision_tenant \
--tenant-id acme:production \
--step tier --tier pro
Implementation:
def deploy_schemas(tenant_id: str, profiles: List[str]) -> List[str]:
tenant = canonical_tenant_id(tenant_id)
_, backend, schema_names = _resolve(tenant, profiles)
return backend.schema_registry.deploy_schemas(
tenant_id=tenant, base_schema_names=schema_names
)
def init_memory(tenant_id: str) -> None:
config_manager = _config_manager(tenant_id)
mgr = Mem0MemoryManager(tenant_id)
lazy_init_memory(mgr, tenant_id, config_manager, auto_create_schema=True)
def init_telemetry(tenant_id: str) -> None:
tenant = canonical_tenant_id(tenant_id)
manager = get_telemetry_manager(
_config_manager(tenant),
otlp_endpoint=resolve_library_env_defaults()["telemetry_otlp_endpoint"],
)
async with manager.required_span("provision.probe", tenant_id=tenant):
pass
5. cogniverse_runtime.optimization_cli¶
Purpose: Per-agent optimization CLI, cluster maintenance sweep, and reporting jobs, all driven by DSPy BootstrapFewShot over Phoenix spans
Location: libs/runtime/cogniverse_runtime/optimization_cli.py (3598 lines)
What Gets Optimized (--mode):
-
simba- Compiles theQueryEnhancementAgent's DSPy module fromcogniverse.query_enhancementspans (original_query → enhanced_query pairs) -
gateway-thresholds- Gateway confidence threshold tuning -
entity-extraction- Entity extraction optimization -
workflow- Multi-agent workflow orchestration (runs after the per-agent optimizers so it sees their artifacts) -
profile- Search profile selection optimization -
cleanup- Daily maintenance sweep: per-tenant Mem0 TTL cleanup (schema-driven), log rotation underLOG_DIR, temp-file cleanup underTEMP_DIR, andconfig_metadataversion vacuum -
triggered- On-demand optimization for a comma-separated--agentslist, driven by a Phoenix--trigger-dataset -
online-routing-eval- Scores routing spans against ground truth without retraining anything -
synthetic- Generates pending review batches for one or more optimizer types (--agents, defaultquery_enhancement,profile,routing,entity_extraction) -
rollback- Restores a single agent's prompts/demonstrations to a prior version (--agent,--prompts-version/--demos-version) -
ab-compare- Runs an RLM A/B comparison over a Phoenix queries dataset (--queries-dataset, optional--judge-substring) -
egress-netpol- Generates KubernetesNetworkPolicymanifests fromconfigs/agent_policies/*.yaml(no tenant, no Phoenix — pure codegen) -
monthly-reports- Aggregates usage/performance data into JSON reports under--reports-output-dir
How Optimization Works:
Every DSPy-based mode (simba, workflow, gateway-thresholds, profile, entity-extraction) builds a training set from Phoenix spans and compiles the target module with dspy.teleprompt.BootstrapFewShot, scaled by trainset size (_create_teleprompter):
- < 50 examples →
max_bootstrapped_demos=4, max_labeled_demos=8, max_rounds=1 - >= 50 examples →
max_bootstrapped_demos=8, max_labeled_demos=16, max_rounds=2
There is no automatic switch between multiple DSPy optimizer algorithms (no GEPA/MIPRO/SIMBA-the-algorithm selection) — BootstrapFewShot is the only optimizer this CLI compiles with; simba here is the mode name for query-enhancement optimization, not the DSPy SIMBA optimizer.
Command Line Usage:
# Optimize query enhancement (mode: simba)
uv run python -m cogniverse_runtime.optimization_cli \
--mode simba \
--tenant-id default
# Optimize gateway thresholds
uv run python -m cogniverse_runtime.optimization_cli \
--mode gateway-thresholds \
--tenant-id acme_corp
# Clean up old logs, memories, and config versions (global sweep)
uv run python -m cogniverse_runtime.optimization_cli \
--mode cleanup \
--log-retention-days 7
Command Line Options:
--mode CHOICE # cleanup|triggered|simba|workflow|gateway-thresholds|
# online-routing-eval|profile|entity-extraction|synthetic|
# rollback|ab-compare|egress-netpol|monthly-reports (required)
--tenant-id ID # Tenant identifier (required for every mode except cleanup,
# egress-netpol, and monthly-reports; if omitted under
# --mode cleanup the sweep runs globally across every tenant
# in every org — the path the daily-cleanup CronWorkflow takes)
--log-retention-days DAYS # Days to retain logs (cleanup mode, default: 7)
--memory-retention-days DAYS # Days to retain expired memories (cleanup mode, default: 30)
--lookback-hours HOURS # Span history window for DSPy-based modes (default: 24.0)
Integration with Argo Workflows:
optimization_cli is invoked directly (not via a separate workflows/ YAML) from CronWorkflow templates rendered by charts/cogniverse/templates/optimization-workflows.yaml, e.g. the weekly agent-optimization workflow's per-mode step:
# charts/cogniverse/templates/optimization-workflows.yaml (run-optimizer step)
- name: gateway-thresholds
templateRef:
name: {{ $fullName }}-optimization-runner
template: run-optimizer
arguments:
parameters:
- name: mode
value: gateway-thresholds
- name: tenant-id
value: "{{workflow.parameters.tenant-id}}"
- name: lookback-hours
value: "48"
The web client's Optimization runs view can also submit a one-off run of the same CLI on demand via POST /admin/tenant/{tenant_id}/optimize, which creates an Argo Workflow from the same template.
Scheduled Execution:
CronWorkflow schedules are set in charts/cogniverse/values.yaml under argo.optimization / argo.maintenance and rendered by charts/cogniverse/templates/optimization-workflows.yaml:
- Daily cleanup (
0 4 * * *, 4 AM UTC):--mode cleanup - Daily gateway tuning (
0 4 * * *, 4 AM UTC):--mode gateway-thresholds(48h lookback), then restarts the runtime deployment - Weekly agent optimization (
0 3 * * 0, Sunday 3 AM UTC):gateway-thresholds,entity-extraction,simba,profilein parallel (48h lookback, 168h forsimba), thenworkflow, then restarts the runtime deployment - Saturday synthetic generation (
0 1 * * 6, 1 AM UTC):--mode synthetic --agents query_enhancement,profile,routing,entity_extraction - Daily scheduled distillation (
0 5 * * *, 5 AM UTC): runscogniverse_runtime.quality_monitor_cli --once, a separate CLI, notoptimization_cli - Monthly reports (
0 5 1 * *, 1st of month 5 AM UTC):--mode monthly-reports, uploaded to MinIO by a follow-up step
6. run_experiments_with_visualization.py¶
Purpose: Run Phoenix experiments with comprehensive visualization and quality evaluators
Location: scripts/run_experiments_with_visualization.py (155 lines)
Architecture: This script is a thin CLI wrapper that delegates to ExperimentTracker from the cogniverse_evaluation SDK package.
Experiment Execution:
def main():
# Parse CLI arguments (tenant-id, dataset-name, profiles, strategies, evaluators, etc.)
args = parser.parse_args()
# Initialize ExperimentTracker from SDK
from cogniverse_evaluation.core.experiment_tracker import ExperimentTracker
# tenant_id is required — ExperimentTracker has no default tenant
tracker = ExperimentTracker(
tenant_id=args.tenant_id,
experiment_project_name="experiments",
enable_quality_evaluators=args.quality_evaluators,
enable_llm_evaluators=args.llm_evaluators,
evaluator_name=args.evaluator,
llm_model=args.llm_model,
llm_base_url=args.llm_base_url,
)
# Get configurations (profiles x strategies matrix)
tracker.get_experiment_configurations(
profiles=args.profiles,
strategies=args.strategies,
all_strategies=args.all_strategies,
)
# Create or get dataset
dataset_name = tracker.create_or_get_dataset(
dataset_name=args.dataset_name,
csv_path=args.csv_path,
force_new=args.force_new,
)
# Run all experiments
experiments = tracker.run_all_experiments(dataset_name)
# Create and print visualization tables
tables = tracker.create_visualization_tables()
tracker.print_visualization(tables)
# Save results (CSV + JSON) and generate HTML report
tracker.save_results(tables, experiments)
tracker.generate_html_report()
Output:
-
Profile summary table
-
Strategy comparison by profile
-
Detailed experiment results
-
CSV summary file
-
JSON detailed results
-
HTML integrated report (if quantitative tests exist)
Command Line Arguments:
# --tenant-id is required; --dataset-path must be a CSV (query, expected_videos,
# category columns) — create_or_get_dataset loads it with pandas.read_csv.
python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--dataset-name golden_eval_v1 \
--dataset-path data/testset/evaluation/video_search_queries.csv \
--profiles frame_based_colpali \
--quality-evaluators \
--llm-evaluators \
--evaluator visual_judge \
--llm-model google/gemma-4-e4b-it
7. manage_datasets.py¶
Purpose: CLI tool for managing evaluation datasets
Location: scripts/manage_datasets.py (62 lines)
Operations:
List Datasets:
python scripts/manage_datasets.py --tenant-id acme:acme --list
# Output:
# Registered datasets:
#
# Name: golden_eval_v1
# Dataset ID: ds_abc123
# Created: 2025-10-07 10:30:00
# Examples: 50
# Description: Golden evaluation dataset
Create Dataset:
python scripts/manage_datasets.py \
--tenant-id acme:acme \
--create my_dataset \
--csv data/queries.csv
# Output:
# Dataset 'my_dataset' created with ID: ds_xyz789
Get Dataset Info:
python scripts/manage_datasets.py --tenant-id acme:acme --info golden_eval_v1
# Output:
# Dataset: golden_eval_v1
# dataset_id: ds_abc123
# created_at: 2025-10-07 10:30:00
# num_examples: 50
# description: Golden evaluation dataset
Implementation:
def main():
from cogniverse_evaluation.data import DatasetManager
dm = DatasetManager(tenant_id=args.tenant_id)
if args.list:
dataset_names = dm.list_datasets() # Returns List[str]
if dataset_names:
print("\nRegistered datasets:")
for ds_name in dataset_names:
info = dm.get_dataset(ds_name)
print(f"\nName: {ds_name}")
if info:
for key, value in info.items():
print(f" {key}: {value}")
else:
print("\nNo datasets registered yet")
elif args.create and args.csv:
dataset_id = dm.create_from_csv(
csv_path=args.csv,
dataset_name=args.create,
description=f"Created from {args.csv}",
)
elif args.info:
info = dm.get_dataset(args.info)
if info:
print(f"\nDataset: {args.info}")
for key, value in info.items():
print(f" {key}: {value}")
else:
print(f"\nDataset '{args.info}' not found")
Data Flow¶
1. Video Ingestion Flow¶
flowchart TB
User(["<span style='color:#000'>User Command</span>"])
Entry["<span style='color:#000'>run_ingestion.py<br/>Parse args: video_dir, profiles, backend</span>"]
GetProfiles["<span style='color:#000'>Resolve profiles<br/>(--profile or config default:<br/>video_colpali_smol500_mv_frame)</span>"]
Build["<span style='color:#000'>Build Pipeline<br/>Test → build_test_pipeline()<br/>Simple → build_simple_pipeline()<br/>Advanced → create_pipeline().with_*().build()</span>"]
Discover["<span style='color:#000'>Discover content files<br/>video_dir.glob('*.mp4') etc.</span>"]
Concurrent["<span style='color:#000'>process_videos_concurrent<br/>max_concurrent=3 (default), async</span>"]
PerVideo["<span style='color:#000'>Per video:<br/>Extract frames/chunks →<br/>Generate embeddings →<br/>Build documents →<br/>Upload to Vespa</span>"]
Collect["<span style='color:#000'>Collect Results<br/>successful/failed, docs fed,<br/>time, throughput</span>"]
Loop{"<span style='color:#000'>More profiles?</span>"}
Summary["<span style='color:#000'>Print Summary<br/>Per-profile stats, overall success rate,<br/>total documents processed</span>"]
User --> Entry --> GetProfiles --> Build --> Discover --> Concurrent --> PerVideo --> Collect --> Loop
Loop -->|yes, next profile| Build
Loop -->|no| Summary
style User fill:#90caf9,stroke:#1565c0,color:#000
style Entry fill:#90caf9,stroke:#1565c0,color:#000
style GetProfiles fill:#b0bec5,stroke:#546e7a,color:#000
style Build fill:#b0bec5,stroke:#546e7a,color:#000
style Discover fill:#b0bec5,stroke:#546e7a,color:#000
style Concurrent fill:#ffcc80,stroke:#ef6c00,color:#000
style PerVideo fill:#ce93d8,stroke:#7b1fa2,color:#000
style Collect fill:#a5d6a7,stroke:#388e3c,color:#000
style Loop fill:#b0bec5,stroke:#546e7a,color:#000
style Summary fill:#a5d6a7,stroke:#388e3c,color:#000 2. Schema Deployment Flow¶
Per-tenant deployment flows through the runtime admin API rather than a bulk script; single-schema JSON deploys continue to work via deploy_json_schema.py for dev iteration.
flowchart TB
Start(["<span style='color:#000'>User / init-job</span>"])
Post["<span style='color:#000'>POST $RUNTIME_URL/admin/profiles/{profile}/deploy<br/>body: {tenant_id, force: false}</span>"]
Deploy["<span style='color:#000'>SchemaRegistry.deploy_schema<br/>(tenant_id, base_schema_name)</span>"]
Transform["<span style='color:#000'>Transform base schema →<br/>tenant-scoped schema name</span>"]
Merge["<span style='color:#000'>Merge registry schemas + live Vespa<br/>document types<br/>(VespaBackend.deploy_schemas)</span>"]
Package["<span style='color:#000'>ApplicationPackage<br/>allow_schema_removal=False<br/>(refuses to drop peer-tenant schemas)</span>"]
VespaDeploy["<span style='color:#000'>_deploy_package → Vespa config server<br/>POST /application/v2/tenant/default/prepareandactivate</span>"]
Converge["<span style='color:#000'>_wait_for_schema_convergence<br/>(serviceconverge generation wait + per-new-schema feed round-trip; raises on timeout)</span>"]
DevDeploy["<span style='color:#000'>Single-schema dev deploy:<br/>deploy_json_schema.py → ApplicationPackage → Vespa</span>"]
Start --> Post --> Deploy --> Transform --> Merge --> Package --> VespaDeploy --> Converge
Start -.->|dev iteration| DevDeploy
style Start fill:#90caf9,stroke:#1565c0,color:#000
style Post fill:#b0bec5,stroke:#546e7a,color:#000
style Deploy fill:#ffcc80,stroke:#ef6c00,color:#000
style Transform fill:#ffcc80,stroke:#ef6c00,color:#000
style Merge fill:#ffcc80,stroke:#ef6c00,color:#000
style Package fill:#ce93d8,stroke:#7b1fa2,color:#000
style VespaDeploy fill:#ce93d8,stroke:#7b1fa2,color:#000
style Converge fill:#a5d6a7,stroke:#388e3c,color:#000
style DevDeploy fill:#b0bec5,stroke:#546e7a,color:#000 3. Experiment Workflow Flow¶
flowchart TB
Start(["<span style='color:#000'>User Command</span>"])
Entry["<span style='color:#000'>run_experiments_with_visualization.py</span>"]
Init["<span style='color:#000'>Initialize ExperimentTracker(tenant_id, ...)<br/>cogniverse_evaluation.core.experiment_tracker<br/>Separate 'experiments' project</span>"]
Dataset["<span style='color:#000'>create_or_get_dataset()<br/>Load or create dataset from CSV,<br/>register with Phoenix</span>"]
Config["<span style='color:#000'>get_experiment_configurations()<br/>Filter profiles/strategies,<br/>build profile × strategy matrix</span>"]
subgraph PerCombo["<span style='color:#000'>FOR each profile × strategy</span>"]
Create["<span style='color:#000'>Create Experiment<br/>Name: '{profile} - {strategy}', attach to dataset</span>"]
RunQueries["<span style='color:#000'>Run Search Queries<br/>FOR each query: execute search,<br/>record spans, collect results</span>"]
Evaluate["<span style='color:#000'>Evaluate Results<br/>quality: relevance/diversity/distribution/temporal<br/>llm: reference-free/reference-based</span>"]
Store["<span style='color:#000'>Store Results<br/>experiment id, status, scores, span traces</span>"]
Create --> RunQueries --> Evaluate --> Store
end
Viz["<span style='color:#000'>Generate Visualizations<br/>Profile summary, strategy comparison,<br/>detailed results (tabulate)</span>"]
Export["<span style='color:#000'>Export Results<br/>CSV/JSON under outputs/experiment_results/,<br/>HTML report if quantitative tests exist</span>"]
Links["<span style='color:#000'>Print Phoenix UI Links<br/>dataset, experiments project, default project</span>"]
Exit(["<span style='color:#000'>Exit with Summary Statistics</span>"])
Start --> Entry --> Init --> Dataset --> Config --> PerCombo
PerCombo --> Viz --> Export --> Links --> Exit
style Start fill:#90caf9,stroke:#1565c0,color:#000
style Entry fill:#90caf9,stroke:#1565c0,color:#000
style Init fill:#ffcc80,stroke:#ef6c00,color:#000
style Dataset fill:#ffcc80,stroke:#ef6c00,color:#000
style Config fill:#ffcc80,stroke:#ef6c00,color:#000
style Create fill:#ce93d8,stroke:#7b1fa2,color:#000
style RunQueries fill:#ce93d8,stroke:#7b1fa2,color:#000
style Evaluate fill:#ce93d8,stroke:#7b1fa2,color:#000
style Store fill:#ce93d8,stroke:#7b1fa2,color:#000
style Viz fill:#ffcc80,stroke:#ef6c00,color:#000
style Export fill:#a5d6a7,stroke:#388e3c,color:#000
style Links fill:#a5d6a7,stroke:#388e3c,color:#000
style Exit fill:#a5d6a7,stroke:#388e3c,color:#000 4. Optimization & Deployment Flow¶
flowchart TB
Start(["<span style='color:#000'>User Command</span>"])
Entry["<span style='color:#000'>python -m cogniverse_runtime.optimization_cli<br/>--mode <MODE> --tenant-id <TENANT></span>"]
Step1["<span style='color:#000'>Step 1: Run Per-Agent Optimizer<br/>Mode: simba | gateway-thresholds |<br/>entity-extraction | workflow | profile<br/>Collects Phoenix spans; compiles the target<br/>DSPy module with BootstrapFewShot,<br/>scaled by trainset size (<50 vs >=50 examples)</span>"]
Step2["<span style='color:#000'>Step 2: Persist Artifact<br/>ArtifactManager.save_blob<br/>Stores optimized module / threshold config<br/>per tenant, returns an artifact_id</span>"]
Step3["<span style='color:#000'>Step 3: Agents Load at Startup<br/>No redeploy: runtime agents read the<br/>latest tenant artifact via ArtifactManager</span>"]
Summary["<span style='color:#000'>Print Summary<br/>Mode, tenant, artifact_id, status</span>"]
Exit(["<span style='color:#000'>Exit with Status Code</span>"])
Start --> Entry --> Step1 --> Step2 --> Step3 --> Summary --> Exit
style Start fill:#90caf9,stroke:#1565c0,color:#000
style Entry fill:#90caf9,stroke:#1565c0,color:#000
style Step1 fill:#ffcc80,stroke:#ef6c00,color:#000
style Step2 fill:#ffcc80,stroke:#ef6c00,color:#000
style Step3 fill:#a5d6a7,stroke:#388e3c,color:#000
style Summary fill:#a5d6a7,stroke:#388e3c,color:#000
style Exit fill:#a5d6a7,stroke:#388e3c,color:#000 Usage Examples¶
Example 1: Basic Video Ingestion¶
# Process videos with default profile
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa
# Output:
# ============================================================
# 🎯 Processing with profile: video_colpali_smol500_mv_frame
# ============================================================
# 🎬 Starting Video Processing Pipeline
# 📁 Video directory: data/testset/evaluation/sample_videos
# 📂 Output directory: outputs/ingestion/video_colpali_smol500_mv_frame
# 🔧 Backend: vespa
# 📹 Found 3 videos to process
#
# ✅ Profile video_colpali_smol500_mv_frame completed!
# Time: 45.32 seconds
# Videos: 3/3 successful
# Documents fed: 180
# Throughput: 4.0 docs/sec
# Avg per video: 15.11 seconds
Example 2: Multi-Profile Ingestion¶
# Process with multiple profiles simultaneously
JAX_PLATFORM_NAME=cpu uv run python scripts/run_ingestion.py \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa \
--profile video_colpali_smol500_mv_frame \
video_xclip_sv_chunk_6s \
video_colqwen_omni_mv_chunk_30s
# Output shows processing for each profile:
# ============================================================
# 🎯 Processing with profile: video_colpali_smol500_mv_frame
# ============================================================
# ...
# ✅ Profile video_colpali_smol500_mv_frame completed!
#
# ============================================================
# 🎯 Processing with profile: video_xclip_sv_chunk_6s
# ============================================================
# ...
# ✅ Profile video_xclip_sv_chunk_6s completed!
#
# ============================================================
# 📊 Overall Summary
# ============================================================
# Processed 3 profiles
# ✅ video_colpali_smol500_mv_frame: 3/3 videos succeeded, 180 docs in 45.3s
# ✅ video_xclip_sv_chunk_6s: 3/3 videos succeeded, 90 docs in 38.1s
# ⚠️ video_colqwen_omni_mv_chunk_30s: 2/3 videos succeeded, 120 docs in 52.7s
Example 3: Test Mode Ingestion¶
# Quick validation with limited frames
uv run python scripts/run_ingestion.py \
--video_dir data/testset/evaluation/sample_videos \
--backend vespa \
--test-mode \
--max-frames 10
# Output:
# 🧪 Using test pipeline builder...
# 🖼️ Max frames: 10
# ✅ Profile video_colpali_smol500_mv_frame completed!
# Time: 12.45 seconds
# Videos: 3/3 successful
# Documents fed: 30 # Only 10 frames per video
Example 4: Schema Deployment¶
# Deploy single schema
python scripts/deploy_json_schema.py \
configs/schemas/video_colpali_smol500_mv_frame_schema.json
# Output:
# ============================================================
# Vespa JSON Schema Deployment
# ============================================================
# Schema file: configs/schemas/video_colpali_smol500_mv_frame_schema.json
# Config server: localhost:19071
# Data endpoint: localhost:8080
#
# 📄 Loading schema from video_colpali_smol500_mv_frame_schema.json
# 📦 Processing schema: video_colpali_smol500_mv_frame
# 🚀 Deploying to http://localhost:19071/application/v2/tenant/default/prepareandactivate...
# ✅ Schema 'video_colpali_smol500_mv_frame' deployed successfully!
#
# ⏳ Waiting for deployment to propagate...
#
# 🔍 Verifying 'video_colpali_smol500_mv_frame' deployment...
# ✅ Vespa is running and responding
#
# ============================================================
# Deployment complete!
# ============================================================
# Deploy tenant schemas via the runtime admin API (single profile)
RUNTIME_URL=http://localhost:8080
curl -sfX POST "$RUNTIME_URL/admin/tenants" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production"}'
curl -sfX POST "$RUNTIME_URL/admin/profiles/video_colpali_smol500_mv_frame/deploy" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "acme:production", "force": false}'
Example 5: Phoenix Experiments¶
# Run experiments with quality evaluators
uv run python scripts/run_experiments_with_visualization.py \
--tenant-id acme:acme \
--dataset-name golden_eval_v1 \
--dataset-path data/testset/evaluation/video_search_queries.csv \
--profiles frame_based_colpali \
--quality-evaluators
# Output:
# ================================================================================
# PHOENIX EXPERIMENTS WITH VISUALIZATION
# ================================================================================
#
# Timestamp: 2025-10-07 14:30:00
# Experiment Project: experiments (separate from default traces)
# Quality Evaluators: ✅ ENABLED (relevance, diversity, distribution, temporal coverage)
# LLM Evaluators: ❌ DISABLED
#
# Preparing experiment dataset...
# ✅ Dataset ready: http://localhost:6006/datasets/ds_abc123
#
# ============================================================
# Profile: frame_based_colpali
# ============================================================
#
# [1/6] Frame Based Colpali - Binary
# Strategy: binary_binary
# ✅ Success
#
# [2/6] Frame Based Colpali - Float
# Strategy: float_float
# ✅ Success
#
# [3/6] Frame Based Colpali - Phased
# Strategy: phased
# ✅ Success
#
# ...
#
# ================================================================================
# EXPERIMENT RESULTS VISUALIZATION
# ================================================================================
#
# 📊 PROFILE SUMMARY
# ------------------------------------------------------------
# +---------------------+-------+---------+--------+---------------+
# | Profile | Total | Success | Failed | Success Rate |
# +=====================+=======+=========+========+===============+
# | frame_based_colpali | 6 | 6 | 0 | 100.0% |
# +---------------------+-------+---------+--------+---------------+
#
# 🔍 STRATEGY COMPARISON BY PROFILE
# ------------------------------------------------------------
#
# frame_based_colpali:
# Strategy Description Status
# -------------------- ------------------ -------------
# binary_binary Binary ✅ Success
# float_float Float ✅ Success
# float_binary Float-Binary ✅ Success
# phased Phased ✅ Success
# hybrid_binary_bm25 Hybrid + Desc ✅ Success
# bm25_only Text Only ✅ Success
#
# ================================================================================
# SUMMARY STATISTICS
# ================================================================================
#
# Total Experiments Attempted: 6
# Successful: 6 (100.0%)
# Failed: 0 (0.0%)
#
# ================================================================================
# VIEW IN PHOENIX UI
# ================================================================================
#
# 🔗 Dataset: http://localhost:6006/datasets/ds_abc123
# 🔗 Experiments Project: http://localhost:6006/projects/experiments
# 🔗 Default Project (spans): http://localhost:6006/projects/default
#
# ℹ️ Notes:
# - Experiments are in separate 'experiments' project
# - Each experiment has its own traces with detailed spans
# - Use Phoenix UI to compare experiments side-by-side
# - Evaluation scores are attached to each experiment
#
# 💾 Results saved to: outputs/experiment_results/experiment_summary_20251007_143000.csv
# 💾 Detailed results saved to: outputs/experiment_results/experiment_details_20251007_143000.json
#
# ✅ All experiments completed!
Example 6: Per-Agent Optimization¶
# Optimize each agent mode for a tenant
uv run python -m cogniverse_runtime.optimization_cli --mode simba --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode gateway-thresholds --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode entity-extraction --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode workflow --tenant-id acme_corp
uv run python -m cogniverse_runtime.optimization_cli --mode profile --tenant-id acme_corp
# Clean up old logs afterward
uv run python -m cogniverse_runtime.optimization_cli --mode cleanup --log-retention-days 7
Example 7: Dataset Management¶
# List all datasets
python scripts/manage_datasets.py --tenant-id acme:acme --list
# Output:
# Registered datasets:
#
# Name: golden_eval_v1
# Dataset ID: ds_abc123
# Created: 2025-10-05 10:30:00
# Examples: 50
# Description: Golden evaluation dataset v1
#
# Name: video_search_test
# Dataset ID: ds_def456
# Created: 2025-10-06 14:15:00
# Examples: 25
# Description: Test queries for video search
# Create new dataset
python scripts/manage_datasets.py \
--tenant-id acme:acme \
--create my_queries \
--csv data/my_queries.csv
# Output:
# Dataset 'my_queries' created with ID: ds_ghi789
# Get dataset info
python scripts/manage_datasets.py --tenant-id acme:acme --info golden_eval_v1
# Output:
# Dataset: golden_eval_v1
# dataset_id: ds_abc123
# created_at: 2025-10-05 10:30:00
# num_examples: 50
# description: Golden evaluation dataset v1
Example 8: Web Client¶
# Deployed by `cogniverse up`
open http://localhost:28400
# The sidebar lists every agent (chat) and the Operations views:
# Tenants, Backend profiles, Configuration, Ingestion, Optimization runs,
# Memory, Approvals, Annotation queue, Workflow reviews, Profile metrics,
# RLM A/B, Analytics, Evaluation, Embedding atlas, Routing evaluation
See Web Client to run it locally.
Production Considerations¶
1. Performance Optimization¶
Ingestion Throughput:
# Adjust concurrency based on available resources
--max-concurrent 5 # More concurrent videos (CPU/memory intensive)
--max-concurrent 1 # Serial processing (safer for limited resources)
# Throughput metrics:
# - ColPali frame-based: ~4-5 docs/sec (GPU)
# - X-CLIP global: ~2-3 docs/sec (GPU)
# - ColQwen chunk-based: ~3-4 docs/sec (GPU)
Async Processing:
# Use async methods for I/O-bound operations
results = await pipeline.process_videos_concurrent(
video_files,
max_concurrent=3
)
# Benefits:
# - Non-blocking video processing
# - Better CPU utilization
# - Faster overall throughput
Schema Deployment:
# Production: init-job loops tenant × profile against the runtime admin
# API — see charts/cogniverse/templates/init-jobs.yaml. The runtime
# merges existing Vespa document types into every redeploy so peer
# tenants never get dropped.
# Dev: iterate a single schema file through deploy_json_schema.py.
2. Error Handling¶
Ingestion Errors:
# Pipeline continues on individual video failures
result = {
'status': 'failed',
'error': 'Embedding generation failed',
'video_path': str(video_path)
}
# Overall statistics still reported:
# ⚠️ Profile frame_based_colpali partially completed!
# Time: 52.7 seconds
# Videos: 2/3 successful
# Failed: 1 videos
Schema Deployment Errors:
# Common errors:
# - "Connection refused" → Vespa not running
# - "HTTP 400" → Schema validation failed
# - "Timeout" → Vespa busy or unresponsive
# deploy_json_schema() failures (non-200 from prepareandactivate, or an
# exception during zip/POST) DO exit non-zero:
if deploy_json_schema(args.schema_file, args.config_host, args.config_port, args.data_port):
...
else:
print("\n❌ Deployment failed")
sys.exit(1)
# verify_deployment() is called afterward as an informational health
# check (prints a warning on non-200) but its return value is NOT
# checked — a failed verification does not change the exit code.
Experiment Errors:
# Experiments continue on individual strategy failures
if result["status"] == "failed":
error = result.get("error", "Unknown error")
if "Text encoder not available" in error:
print(f" ⚠️ Skipped: Encoder not available")
else:
print(f" ❌ Failed: {error[:50]}...")
# Final summary shows partial success:
# Total: 10, Successful: 8 (80.0%), Failed: 2 (20.0%)
3. Monitoring and Logging¶
Ingestion Logs:
# Logs written to outputs/logs/
# - ingestion_pipeline.log # Main pipeline logs
# - video_processing.log # Per-video processing
# - embedding_generation.log # Embedding logs
# - vespa_upload.log # Upload logs
# View logs:
tail -f outputs/logs/ingestion_pipeline.log
Experiment Logs:
# Experiment results saved to:
# - outputs/experiment_results/experiment_summary_*.csv
# - outputs/experiment_results/experiment_details_*.json
# Phoenix spans capture:
# - Query execution time
# - Search latency
# - Embedding generation time
# - Evaluation scores
# View in Phoenix UI: http://localhost:6006
4. Resource Management¶
Memory Management:
# Video processing is memory-intensive
# Estimated memory per video:
# - Frame extraction: ~200-500 MB
# - Embedding generation: ~1-2 GB (GPU)
# - Document building: ~50-100 MB
# Total concurrent memory:
# max_concurrent=3 → ~4-6 GB peak memory
# Recommendations:
# - 16 GB RAM minimum for production
# - 32 GB RAM recommended for max_concurrent > 5
GPU Utilization:
# GPU required for:
# - ColPali embedding generation
# - X-CLIP encoding
# - ColQwen embedding generation
# GPU memory requirements:
# - ColPali Smol 500M: ~2 GB VRAM
# - X-CLIP: ~4 GB VRAM
# - ColQwen Omni: ~6 GB VRAM
# For multiple profiles:
# - Process sequentially (GPU memory limits)
# - Or use multiple GPUs with CUDA_VISIBLE_DEVICES
Disk Space:
# Storage requirements per video:
# - Original video: ~50-500 MB
# - Extracted frames: ~10-50 MB
# - Transcripts: ~10-50 KB
# - Embeddings (binary): ~100-500 KB
# - Embeddings (float): ~1-5 MB
# Total for 1000 videos: ~100-500 GB
5. Scaling Strategies¶
Horizontal Scaling:
# Ingestion:
# - Run multiple ingestion processes
# - Partition videos by directory
# - Each process handles different profiles
# Example:
# Process 1: --profile video_colpali_smol500_mv_frame
# Process 2: --profile video_xclip_sv_chunk_6s
# Process 3: --profile video_colqwen_omni_mv_chunk_30s
Vertical Scaling:
# Increase concurrency with more resources:
--max-concurrent 10 # Requires 32+ GB RAM, 16+ CPU cores
# Benefits:
# - Faster overall throughput
# - Better resource utilization
# - Reduced total processing time
Batch Processing:
# Process large video collections in batches:
videos_per_batch = 100
for i in range(0, len(videos), videos_per_batch):
batch = videos[i:i+videos_per_batch]
run_ingestion(batch)
6. Best Practices¶
Schema Management:
1. Deploy per-tenant via the runtime admin API; the Helm init job
(charts/cogniverse/templates/init-jobs.yaml) runs tenant × profile
on install. Never bulk-deploy with allow_schema_removal overrides
— that path silently drops peer-tenant schemas.
2. Version-control all schema JSON files in configs/schemas/.
3. Test schema changes in staging before production.
4. Ranking strategies are extracted at search time (no separate step).
5. Validate schemas before deployment with deploy_json_schema.py
against a local Vespa container.
Ingestion Pipeline:
# 1. Test with --test-mode first (max_frames=10)
# 2. Process small batches before full ingestion
# 3. Monitor logs for errors during processing
# 4. Verify documents in Vespa after ingestion
# 5. Use appropriate backend (vespa for production)
Experiments:
# 1. Use separate "experiments" project in Phoenix
# 2. Enable quality evaluators for comprehensive metrics
# 3. Save results to CSV/JSON for analysis
# 4. Compare experiments side-by-side in Phoenix UI
# 5. Export results before deleting experiments
7. Common Issues and Solutions¶
Issue: "Video processing failed: CUDA out of memory"
# Solution: Reduce concurrency or batch size
--max-concurrent 1 # Process one video at a time
--max-frames 50 # Reduce frames per video
Issue: "Schema deployment failed: Connection refused"
# Solution: Ensure Vespa is running
docker ps | grep vespa # Check Vespa container
cogniverse up # Start Vespa if not running
Summary¶
The Scripts & Operations module provides comprehensive tooling for:
- Video Ingestion: Builder pattern pipeline with multi-profile support
- Schema Management: JSON-based deployment with validation
- Optimization: Complete DSPy optimization and deployment workflow
- Experimentation: Phoenix experiments with quality evaluators
- Dataset Management: CRUD operations for evaluation datasets
- System Setup: Environment initialization and dependency checking
Key Design Patterns:
-
Builder pattern for flexible pipeline configuration
-
Async processing for concurrent video handling
-
Command pattern for CLI tool design
-
Context manager for experiment runner lifecycle
-
Factory pattern for strategy resolution
Production Features:
-
Concurrent async video processing
-
Multi-profile simultaneous ingestion
-
Phoenix experiment tracking with visualization
-
Comprehensive error handling and logging
-
Resource-aware concurrency limits
This module serves as the operational backbone of the Cogniverse system, providing production-grade tools for deployment, ingestion, optimization, and monitoring.
Related Documentation:
docs/development/instrumentation.md: Phoenix telemetry and observability