Skip to content

Evaluation Module Study Guide

Package: cogniverse_evaluation (Core Layer) Module Location: libs/evaluation/cogniverse_evaluation/


Package Structure

libs/evaluation/cogniverse_evaluation/
├── __init__.py                          # Package initialization
├── cli.py                               # CLI for evaluation tasks
├── online_evaluator.py                  # Online evaluation pipeline
├── quality_monitor.py                   # Quality monitoring
├── recorded_searches.py                 # Recorded searches scored against a tenant's golden set
├── span_evaluator.py                    # SpanEvaluator for retrospective evaluation
├── core/                                # Core evaluation framework
│   ├── __init__.py
│   ├── experiment_tracker.py            # ExperimentTracker main class
│   ├── solvers.py                       # Inspect AI solvers
│   ├── task.py                          # Evaluation task definitions
│   ├── ground_truth.py                  # Ground truth extraction
│   ├── schema_analyzer.py               # Schema analysis framework
│   ├── inspect_scorers.py               # Inspect AI scorer helpers
│   ├── inspect_model.py                 # Inspect AI "cogniverse" model provider (litellm)
│   ├── reranking.py                     # Reranking logic
│   └── solver_output.py                 # Solver output formatting
├── evaluators/                          # Evaluator implementations
│   ├── agent_evaluators.py              # Per-agent-type evaluator registry (judge prompts + structural evals)
│   ├── routing_evaluator.py             # Routing decision evaluator
│   ├── reference_free.py                # Reference-free evaluators
│   ├── golden_dataset.py                # Golden dataset evaluator
│   ├── llm_judge.py                     # LLM-based evaluators
│   ├── sync_reference_free.py           # Synchronous reference-free evaluators
│   ├── configurable_visual_judge.py     # Visual judge (provider/model from config)
│   ├── _media_helpers.py                # source_url resolution + frame extraction
│   └── base.py                          # Base evaluator classes
├── metrics/                             # Metric definitions
│   └── custom.py                        # Custom metrics
├── data/                                # Data loaders and datasets
│   ├── datasets.py                      # Dataset management
│   ├── storage.py                       # Storage utilities
│   └── traces.py                        # Trace data handling
├── providers/                           # Evaluation provider system
│   ├── base.py                          # Provider interfaces
│   └── registry.py                      # Provider registry
├── plugins/                             # Plugin system
│   ├── video_analyzer.py                # Video schema analyzer
│   ├── document_analyzer.py             # Document schema analyzer
│   └── visual_evaluator.py              # Visual evaluator plugin
└── analysis/                            # Analysis utilities
    └── root_cause_analysis.py           # Root cause analysis

Table of Contents

  1. Module Overview
  2. Architecture Diagrams
  3. Core Components
  4. Usage Examples
  5. Production Considerations
  6. Inspect AI Integration
  7. Plugin System
  8. Testing

Module Overview

Purpose and Responsibilities

The Evaluation Module provides comprehensive experiment tracking and performance evaluation with:

  • Experiment Management: Phoenix-based experiment tracking with visualization
  • Routing Evaluation: Separate evaluation of routing decisions vs search quality
  • Span Analysis: Retrospective evaluation of Phoenix traces
  • Performance Analytics: Statistical analysis and visualization of traces
  • Golden Datasets: Reference datasets for quality benchmarking
  • Multi-Evaluator Support: Quality metrics, LLM judges, visual evaluators

Key Features

  1. ExperimentTracker
  2. Inspect AI-based evaluation framework
  3. Compatible with run_experiments_with_visualization.py
  4. Phoenix integration for trace visualization
  5. Quality and LLM evaluator plugins
  6. Dataset management and golden datasets

  7. RoutingEvaluator

  8. Routing-specific metrics (separate from search quality)
  9. Accuracy, confidence calibration, per-agent precision/recall
  10. Phoenix span analysis for routing decisions
  11. Outcome classification (success, failure, ambiguous)

  12. PhoenixAnalytics

  13. Statistical analysis of traces (latency, throughput, errors)
  14. Outlier detection using IQR and z-score methods
  15. Interactive visualizations (time series, distributions, heatmaps)
  16. Comparison analysis across profiles/strategies

  17. SpanEvaluator

  18. Retrospective evaluation of existing Phoenix spans
  19. Reference-free and golden dataset evaluators
  20. Automatic upload of evaluation results to Phoenix
  21. Batch processing of historical traces

  22. Multi-Turn Evaluation

  23. LLM judges support conversation history context
  24. Session-level evaluation with outcome (success/partial/failure)
  25. Session quality scoring (0-1)
  26. Trajectory-level evaluation for fine-tuning data collection

  27. Session-Based Evaluation

  28. Unified evaluation UI for single and multi-turn conversations
  29. Per-result relevance annotation for individual results
  30. Session-level outcome classification (Success/Partial/Failure)
  31. Integration with Phoenix session tracking

  32. Per-Agent Evaluator Registry (evaluators/agent_evaluators.py)

  33. One AgentEvaluator entry per agent type: search, summary, report, gateway, routing, query_enhancement, entity_extraction, profile_selection
  34. Each entry carries the agent's LLM-judge prompt builder, its judge-payload extraction (search-result list vs summary/report string vs domain dict), and named no-LLM structural evaluators (e.g. routing_outcome, confidence_calibration, enhancement_effect, extraction_yield, profile_confidence_calibration)
  35. QualityMonitor and OnlineEvaluator both dispatch through the registry — adding an agent type to the optimization loop is one registry entry plus an AgentType member
  36. Structural evaluators read the canonical output.value span JSON (read_span_io), with legacy attribute fallbacks

  37. OnlineEvaluator

  38. Real-time scoring of spans as they are produced, via the per-agent registry (agent_type selects the entry; defaults to routing)
  39. Configurable sampling rate to bound evaluation overhead
  40. Dispatches the configured structural evaluators (routing: routing_outcome and confidence_calibration)
  41. Persists scores as online_eval.<evaluator> telemetry annotations for drift detection

  42. QualityMonitor

  43. Continuous, scheduled quality monitoring across all agent types (search, summary, report, gateway, routing, query_enhancement, entity_extraction, profile_selection)
  44. Dual strategy: golden-set evaluation (MRR/nDCG/P@5) + live-traffic LLM-judge sampling; the judge prompt per agent type comes from the evaluator registry
  45. The answer agents are scored on their <ClassName>.process spans; routing/query_enhancement/entity_extraction/profile_selection on their cogniverse.* domain spans
  46. Threshold-based verdicts (SKIP / OPTIMIZE / FULL) that trigger Argo optimization workflows
  47. Composes SpanEvaluator, GoldenDatasetEvaluator, LLMJudgeCore, and PhoenixDatasetStore rather than reimplementing them

  48. Evaluation Provider System

    • Provider-agnostic abstraction (EvaluationProvider, AnalyticsProvider) for experiment tracking, dataset management, and analytics
    • Entry-point based discovery (cogniverse.evaluation.providers) with tenant-scoped caching, mirroring the telemetry provider registry
    • Phoenix is the only concrete implementation shipped today (PhoenixEvaluationProvider)
  49. CLI

    • cogniverse-eval unified command group (evaluate, create-dataset, list-traces, test) for running evaluations and managing datasets from the shell

Dependencies

Internal (from pyproject.toml):

  • cogniverse-foundation: Core configuration and telemetry interfaces

  • cogniverse-sdk: SDK utilities

External:

  • inspect-ai==0.3.272: Evaluation framework

  • pandas==2.3.3: Data analysis

  • numpy==2.4.4: Numerical computations

  • scikit-learn==1.8.0: Statistical methods

  • pillow==12.2.0: Image processing

Note: Phoenix, Plotly, and Tabulate are optional dependencies provided by other packages in the workspace (e.g. cogniverse-telemetry-phoenix supplies the PhoenixEvaluationProvider and PhoenixAnalytics).


Architecture Diagrams

1. Experiment Tracking Architecture

flowchart TB
    Start["<span style='color:#000'>Start Experiment</span>"] --> Config["<span style='color:#000'>Configuration Phase</span>"]

    Config --> GetConfigs["<span style='color:#000'>Get Experiment Configurations</span>"]
    GetConfigs --> Profiles["<span style='color:#000'>Profiles: frame_based, chunk_based</span>"]
    GetConfigs --> Strategies["<span style='color:#000'>Strategies: binary, float, hybrid</span>"]
    GetConfigs --> Registry["<span style='color:#000'>Strategy Registry Integration</span>"]

    Profiles --> CreateDataset["<span style='color:#000'>Create/Get Dataset</span>"]
    Strategies --> CreateDataset
    Registry --> CreateDataset

    CreateDataset --> DataSource{"<span style='color:#000'>Dataset Source</span>"}
    DataSource -->|CSV| LoadCSV["<span style='color:#000'>Load from CSV</span>"]
    DataSource -->|Existing| LoadExisting["<span style='color:#000'>Load from Phoenix</span>"]
    DataSource -->|Golden| LoadGolden["<span style='color:#000'>Load Golden Dataset</span>"]

    LoadCSV --> Execution["<span style='color:#000'>Execution Phase</span>"]
    LoadExisting --> Execution
    LoadGolden --> Execution

    Execution --> RunExperiments["<span style='color:#000'>Run Experiments</span>"]
    RunExperiments --> CreateTask["<span style='color:#000'>Create Inspect AI Task</span>"]
    CreateTask --> TaskConfig["<span style='color:#000'>Configure Task</span>"]
    TaskConfig --> Samples["<span style='color:#000'>Dataset Samples</span>"]
    TaskConfig --> Solver["<span style='color:#000'>Solver Configuration</span>"]
    TaskConfig --> Evaluators["<span style='color:#000'>Evaluator Plugins</span>"]

    Samples --> Execute["<span style='color:#000'>Execute with Inspect AI</span>"]
    Solver --> Execute
    Evaluators --> Execute

    Execute --> RunSolver["<span style='color:#000'>Run Solver on Each Sample</span>"]
    RunSolver --> Collect["<span style='color:#000'>Collect Results</span>"]
    Collect --> RunEvals["<span style='color:#000'>Run Evaluators</span>"]

    RunEvals --> ExtractMetrics["<span style='color:#000'>Extract Metrics</span>"]
    ExtractMetrics --> MRR["<span style='color:#000'>MRR, Recall@10</span>"]
    ExtractMetrics --> Relevance["<span style='color:#000'>Relevance, Diversity</span>"]
    ExtractMetrics --> LLMScores["<span style='color:#000'>LLM Judge Scores</span>"]

    MRR --> LogPhoenix["<span style='color:#000'>Log to Phoenix</span>"]
    Relevance --> LogPhoenix
    LLMScores --> LogPhoenix

    LogPhoenix --> Visualization["<span style='color:#000'>Visualization Phase</span>"]

    Visualization --> CreateViz["<span style='color:#000'>Create Visualization Tables</span>"]
    CreateViz --> ProfileSummary["<span style='color:#000'>Profile Summary</span>"]
    CreateViz --> StrategyComp["<span style='color:#000'>Strategy Comparison</span>"]
    CreateViz --> DetailedResults["<span style='color:#000'>Detailed Results with Metrics</span>"]

    ProfileSummary --> Print["<span style='color:#000'>Print Visualization</span>"]
    StrategyComp --> Print
    DetailedResults --> Print

    Print --> SaveResults["<span style='color:#000'>Save Results</span>"]
    SaveResults --> CSV["<span style='color:#000'>CSV Summary</span>"]
    SaveResults --> JSON["<span style='color:#000'>JSON Detailed Results</span>"]
    SaveResults --> HTML["<span style='color:#000'>HTML Report Optional</span>"]

    CSV --> Phoenix["<span style='color:#000'>Phoenix Backend</span>"]
    JSON --> Phoenix
    HTML --> Phoenix

    Phoenix --> UI["<span style='color:#000'>Phoenix UI</span>"]
    UI --> ExpUI["<span style='color:#000'>Experiments UI</span>"]
    UI --> TraceViewer["<span style='color:#000'>Trace Viewer</span>"]
    UI --> MetricsCharts["<span style='color:#000'>Metrics Charts</span>"]

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style Config fill:#ffcc80,stroke:#ef6c00,color:#000
    style GetConfigs fill:#b0bec5,stroke:#546e7a,color:#000
    style Profiles fill:#b0bec5,stroke:#546e7a,color:#000
    style Strategies fill:#b0bec5,stroke:#546e7a,color:#000
    style Registry fill:#b0bec5,stroke:#546e7a,color:#000
    style CreateDataset fill:#ffcc80,stroke:#ef6c00,color:#000
    style DataSource fill:#b0bec5,stroke:#546e7a,color:#000
    style LoadCSV fill:#a5d6a7,stroke:#388e3c,color:#000
    style LoadExisting fill:#a5d6a7,stroke:#388e3c,color:#000
    style LoadGolden fill:#a5d6a7,stroke:#388e3c,color:#000
    style Execution fill:#ffcc80,stroke:#ef6c00,color:#000
    style RunExperiments fill:#ffcc80,stroke:#ef6c00,color:#000
    style CreateTask fill:#ffcc80,stroke:#ef6c00,color:#000
    style TaskConfig fill:#b0bec5,stroke:#546e7a,color:#000
    style Samples fill:#b0bec5,stroke:#546e7a,color:#000
    style Solver fill:#b0bec5,stroke:#546e7a,color:#000
    style Evaluators fill:#b0bec5,stroke:#546e7a,color:#000
    style Execute fill:#ffcc80,stroke:#ef6c00,color:#000
    style RunSolver fill:#ffcc80,stroke:#ef6c00,color:#000
    style Collect fill:#ffcc80,stroke:#ef6c00,color:#000
    style RunEvals fill:#ffcc80,stroke:#ef6c00,color:#000
    style ExtractMetrics fill:#ffcc80,stroke:#ef6c00,color:#000
    style MRR fill:#a5d6a7,stroke:#388e3c,color:#000
    style Relevance fill:#a5d6a7,stroke:#388e3c,color:#000
    style LLMScores fill:#a5d6a7,stroke:#388e3c,color:#000
    style LogPhoenix fill:#ce93d8,stroke:#7b1fa2,color:#000
    style Visualization fill:#ffcc80,stroke:#ef6c00,color:#000
    style CreateViz fill:#ffcc80,stroke:#ef6c00,color:#000
    style ProfileSummary fill:#b0bec5,stroke:#546e7a,color:#000
    style StrategyComp fill:#b0bec5,stroke:#546e7a,color:#000
    style DetailedResults fill:#b0bec5,stroke:#546e7a,color:#000
    style Print fill:#ffcc80,stroke:#ef6c00,color:#000
    style SaveResults fill:#ffcc80,stroke:#ef6c00,color:#000
    style CSV fill:#a5d6a7,stroke:#388e3c,color:#000
    style JSON fill:#a5d6a7,stroke:#388e3c,color:#000
    style HTML fill:#a5d6a7,stroke:#388e3c,color:#000
    style Phoenix fill:#ce93d8,stroke:#7b1fa2,color:#000
    style UI fill:#a5d6a7,stroke:#388e3c,color:#000
    style ExpUI fill:#a5d6a7,stroke:#388e3c,color:#000
    style TraceViewer fill:#a5d6a7,stroke:#388e3c,color:#000
    style MetricsCharts fill:#a5d6a7,stroke:#388e3c,color:#000

Key Points:

  • Inspect AI framework for evaluation execution

  • Plugin system for extensible evaluators

  • Phoenix integration for visualization

  • Compatible with run_experiments_with_visualization.py output format


2. Routing Evaluator Architecture

sequenceDiagram
    participant Evaluator as Routing Evaluator
    participant Phoenix as Phoenix Backend
    participant Classifier as Outcome Classifier
    participant Calculator as Metrics Calculator

    Note over Evaluator: Step 1: Query Phoenix for Routing Spans
    Evaluator->>Phoenix: query_routing_spans(start_time, end_time, limit)
    Note over Phoenix: Project: cogniverse-{tenant_id}<br/>Filter: name == "cogniverse.routing"<br/>Time range: last N hours<br/>Sort: most recent first
    Phoenix-->>Evaluator: routing_spans[]

    Note over Evaluator: Step 2: Evaluate Each Routing Decision
    loop For each span
        Evaluator->>Evaluator: Extract span attributes
        Note over Evaluator: read_span_io(row)["output"] dict:<br/>chosen_agent, confidence, processing_time

        Evaluator->>Evaluator: _classify_routing_outcome(span_data)
        activate Evaluator
        Note over Evaluator: Check parent span exists?<br/>Check status code == OK?<br/>Check downstream agent executed?<br/>Check error events present?

        alt All checks pass
            Note over Evaluator: SUCCESS
        else Errors or timeouts
            Note over Evaluator: FAILURE
        else Unclear outcome
            Note over Evaluator: AMBIGUOUS
        end
        deactivate Evaluator
    end

    Note over Evaluator: Step 3: Calculate Aggregate Metrics
    Evaluator->>Calculator: calculate_metrics(routing_spans)
    activate Calculator

    Calculator->>Calculator: Routing Accuracy<br/>= successful / total_decisions
    Calculator->>Calculator: Confidence Calibration<br/>= correlation(confidence, success)<br/>(Pearson coefficient)
    Calculator->>Calculator: Average Routing Latency<br/>= mean(latency_ms) over timed decisions
    Calculator->>Calculator: Per-Agent Metrics<br/>Precision = TP / (TP + FP)<br/>Recall = TP / (TP + FN)<br/>F1 = 2 * (P * R) / (P + R)

    Calculator-->>Evaluator: RoutingMetrics{<br/>routing_accuracy: 0.85,<br/>confidence_calibration: 0.72,<br/>avg_routing_latency: 150.5,<br/>per_agent_precision: {...},<br/>per_agent_recall: {...},<br/>per_agent_f1: {...},<br/>total_decisions: 100,<br/>ambiguous_count: 5<br/>}
    deactivate Calculator

    Evaluator-->>Evaluator: Return metrics

Routing vs Search Quality:

  • Routing evaluation focuses on decision quality (right agent chosen?)

  • Search evaluation focuses on result quality (relevant results returned?)

  • Separate metrics enable independent optimization


3. Phoenix Analytics Flow

flowchart TB
    Start["<span style='color:#000'>Analytics Request</span>"] --> FetchTraces["<span style='color:#000'>Step 1: Fetch Traces</span>"]

    FetchTraces --> PhoenixQuery["<span style='color:#000'>Phoenix Client Query</span>"]
    PhoenixQuery --> GetSpans["<span style='color:#000'>get_spans_dataframe</span>"]
    GetSpans --> TimeFilter["<span style='color:#000'>Filter by time range</span>"]
    TimeFilter --> OpFilter["<span style='color:#000'>Filter by operation regex</span>"]
    OpFilter --> ExtractRoot["<span style='color:#000'>Extract root spans traces</span>"]

    ExtractRoot --> ExtractMetrics["<span style='color:#000'>Extract TraceMetrics</span>"]
    ExtractMetrics --> TraceID["<span style='color:#000'>trace_id</span>"]
    ExtractMetrics --> Timestamp["<span style='color:#000'>timestamp</span>"]
    ExtractMetrics --> Duration["<span style='color:#000'>duration_ms</span>"]
    ExtractMetrics --> Operation["<span style='color:#000'>operation</span>"]
    ExtractMetrics --> Status["<span style='color:#000'>status success/error</span>"]
    ExtractMetrics --> Profile["<span style='color:#000'>profile, strategy</span>"]
    ExtractMetrics --> Metadata["<span style='color:#000'>metadata</span>"]

    TraceID --> CalcStats["<span style='color:#000'>Step 2: Calculate Statistics</span>"]
    Timestamp --> CalcStats
    Duration --> CalcStats
    Operation --> CalcStats
    Status --> CalcStats
    Profile --> CalcStats
    Metadata --> CalcStats

    CalcStats --> OverallStats["<span style='color:#000'>Overall Stats</span>"]
    OverallStats --> TotalReq["<span style='color:#000'>Total requests</span>"]
    OverallStats --> TimeRange["<span style='color:#000'>Time range</span>"]
    OverallStats --> ResponseTime["<span style='color:#000'>Response time<br/>mean, median, P50/P75/P90/P95/P99</span>"]
    OverallStats --> SuccessRate["<span style='color:#000'>Success/error rates</span>"]
    OverallStats --> OutlierDetect["<span style='color:#000'>Outlier detection IQR method</span>"]

    CalcStats --> GroupedStats{"<span style='color:#000'>Group By?</span>"}
    GroupedStats -->|Yes| PerProfile["<span style='color:#000'>Per profile/strategy/operation</span>"]
    GroupedStats -->|No| Temporal["<span style='color:#000'>Temporal Patterns</span>"]

    PerProfile --> Count["<span style='color:#000'>Count</span>"]
    PerProfile --> MeanMedian["<span style='color:#000'>Mean/median/P95 duration</span>"]
    PerProfile --> ErrorRate["<span style='color:#000'>Error rate</span>"]

    TotalReq --> Temporal
    TimeRange --> Temporal
    ResponseTime --> Temporal
    SuccessRate --> Temporal
    OutlierDetect --> Temporal
    Count --> Temporal
    MeanMedian --> Temporal
    ErrorRate --> Temporal

    Temporal --> ReqByHour["<span style='color:#000'>Requests by hour</span>"]
    Temporal --> DurationByHour["<span style='color:#000'>Avg duration by hour</span>"]

    ReqByHour --> CreateViz["<span style='color:#000'>Step 3: Create Visualizations</span>"]
    DurationByHour --> CreateViz

    CreateViz --> TimeSeries["<span style='color:#000'>1. Time Series Plot</span>"]
    TimeSeries --> TSMean["<span style='color:#000'>Mean/median/max over time</span>"]
    TimeSeries --> TSBands["<span style='color:#000'>P50/P95 bands</span>"]
    TimeSeries --> TSCount["<span style='color:#000'>Request count</span>"]

    CreateViz --> Distribution["<span style='color:#000'>2. Distribution Plot 4 subplots</span>"]
    Distribution --> Histogram["<span style='color:#000'>Histogram</span>"]
    Distribution --> BoxPlot["<span style='color:#000'>Box plot quartiles, outliers</span>"]
    Distribution --> ViolinPlot["<span style='color:#000'>Violin plot distribution shape</span>"]
    Distribution --> ECDF["<span style='color:#000'>ECDF with percentile lines</span>"]

    CreateViz --> Heatmap["<span style='color:#000'>3. Heatmap</span>"]
    Heatmap --> HourDay["<span style='color:#000'>Hour x Day of week</span>"]
    Heatmap --> ProfileStrategy["<span style='color:#000'>Profile x Strategy</span>"]
    Heatmap --> Aggregation["<span style='color:#000'>Aggregation: count, mean, max</span>"]

    CreateViz --> OutlierPlot["<span style='color:#000'>4. Outlier Plot</span>"]
    OutlierPlot --> ScatterPlot["<span style='color:#000'>Scatter: normal vs outlier points</span>"]
    OutlierPlot --> IQRLine["<span style='color:#000'>IQR threshold line</span>"]
    OutlierPlot --> RefLines["<span style='color:#000'>P50/P95/P99 reference lines</span>"]

    CreateViz --> ComparisonPlot["<span style='color:#000'>5. Comparison Plot 4 subplots</span>"]
    ComparisonPlot --> MeanByGroup["<span style='color:#000'>Mean by group</span>"]
    ComparisonPlot --> MedianByGroup["<span style='color:#000'>Median by group</span>"]
    ComparisonPlot --> P95ByGroup["<span style='color:#000'>P95 by group</span>"]
    ComparisonPlot --> CountByGroup["<span style='color:#000'>Request count by group</span>"]

    TSMean --> GenerateReport["<span style='color:#000'>Step 4: Generate Report</span>"]
    TSBands --> GenerateReport
    TSCount --> GenerateReport
    Histogram --> GenerateReport
    BoxPlot --> GenerateReport
    ViolinPlot --> GenerateReport
    ECDF --> GenerateReport
    HourDay --> GenerateReport
    ProfileStrategy --> GenerateReport
    Aggregation --> GenerateReport
    ScatterPlot --> GenerateReport
    IQRLine --> GenerateReport
    RefLines --> GenerateReport
    MeanByGroup --> GenerateReport
    MedianByGroup --> GenerateReport
    P95ByGroup --> GenerateReport
    CountByGroup --> GenerateReport

    GenerateReport --> ReportJSON["<span style='color:#000'>Comprehensive Report JSON</span>"]
    ReportJSON --> Summary["<span style='color:#000'>summary</span>"]
    ReportJSON --> Statistics["<span style='color:#000'>statistics</span>"]
    ReportJSON --> StatsByProfile["<span style='color:#000'>statistics_by_profile</span>"]
    ReportJSON --> StatsByOp["<span style='color:#000'>statistics_by_operation</span>"]
    ReportJSON --> Visualizations["<span style='color:#000'>visualizations plotly_json</span>"]

    Summary --> SaveFile["<span style='color:#000'>Save to JSON file optional</span>"]
    Statistics --> SaveFile
    StatsByProfile --> SaveFile
    StatsByOp --> SaveFile
    Visualizations --> SaveFile

    style Start fill:#90caf9,stroke:#1565c0,color:#000
    style FetchTraces fill:#ffcc80,stroke:#ef6c00,color:#000
    style PhoenixQuery fill:#ce93d8,stroke:#7b1fa2,color:#000
    style GetSpans fill:#ce93d8,stroke:#7b1fa2,color:#000
    style TimeFilter fill:#b0bec5,stroke:#546e7a,color:#000
    style OpFilter fill:#b0bec5,stroke:#546e7a,color:#000
    style ExtractRoot fill:#b0bec5,stroke:#546e7a,color:#000
    style ExtractMetrics fill:#ffcc80,stroke:#ef6c00,color:#000
    style TraceID fill:#b0bec5,stroke:#546e7a,color:#000
    style Timestamp fill:#b0bec5,stroke:#546e7a,color:#000
    style Duration fill:#b0bec5,stroke:#546e7a,color:#000
    style Operation fill:#b0bec5,stroke:#546e7a,color:#000
    style Status fill:#b0bec5,stroke:#546e7a,color:#000
    style Profile fill:#b0bec5,stroke:#546e7a,color:#000
    style Metadata fill:#b0bec5,stroke:#546e7a,color:#000
    style CalcStats fill:#ffcc80,stroke:#ef6c00,color:#000
    style OverallStats fill:#ffcc80,stroke:#ef6c00,color:#000
    style TotalReq fill:#b0bec5,stroke:#546e7a,color:#000
    style TimeRange fill:#b0bec5,stroke:#546e7a,color:#000
    style ResponseTime fill:#b0bec5,stroke:#546e7a,color:#000
    style SuccessRate fill:#b0bec5,stroke:#546e7a,color:#000
    style OutlierDetect fill:#b0bec5,stroke:#546e7a,color:#000
    style GroupedStats fill:#b0bec5,stroke:#546e7a,color:#000
    style PerProfile fill:#b0bec5,stroke:#546e7a,color:#000
    style Temporal fill:#b0bec5,stroke:#546e7a,color:#000
    style Count fill:#b0bec5,stroke:#546e7a,color:#000
    style MeanMedian fill:#b0bec5,stroke:#546e7a,color:#000
    style ErrorRate fill:#b0bec5,stroke:#546e7a,color:#000
    style ReqByHour fill:#b0bec5,stroke:#546e7a,color:#000
    style DurationByHour fill:#b0bec5,stroke:#546e7a,color:#000
    style CreateViz fill:#ffcc80,stroke:#ef6c00,color:#000
    style TimeSeries fill:#a5d6a7,stroke:#388e3c,color:#000
    style TSMean fill:#b0bec5,stroke:#546e7a,color:#000
    style TSBands fill:#b0bec5,stroke:#546e7a,color:#000
    style TSCount fill:#b0bec5,stroke:#546e7a,color:#000
    style Distribution fill:#a5d6a7,stroke:#388e3c,color:#000
    style Histogram fill:#b0bec5,stroke:#546e7a,color:#000
    style BoxPlot fill:#b0bec5,stroke:#546e7a,color:#000
    style ViolinPlot fill:#b0bec5,stroke:#546e7a,color:#000
    style ECDF fill:#b0bec5,stroke:#546e7a,color:#000
    style Heatmap fill:#a5d6a7,stroke:#388e3c,color:#000
    style HourDay fill:#b0bec5,stroke:#546e7a,color:#000
    style ProfileStrategy fill:#b0bec5,stroke:#546e7a,color:#000
    style Aggregation fill:#b0bec5,stroke:#546e7a,color:#000
    style OutlierPlot fill:#a5d6a7,stroke:#388e3c,color:#000
    style ScatterPlot fill:#b0bec5,stroke:#546e7a,color:#000
    style IQRLine fill:#b0bec5,stroke:#546e7a,color:#000
    style RefLines fill:#b0bec5,stroke:#546e7a,color:#000
    style ComparisonPlot fill:#a5d6a7,stroke:#388e3c,color:#000
    style MeanByGroup fill:#b0bec5,stroke:#546e7a,color:#000
    style MedianByGroup fill:#b0bec5,stroke:#546e7a,color:#000
    style P95ByGroup fill:#b0bec5,stroke:#546e7a,color:#000
    style CountByGroup fill:#b0bec5,stroke:#546e7a,color:#000
    style GenerateReport fill:#ffcc80,stroke:#ef6c00,color:#000
    style ReportJSON fill:#a5d6a7,stroke:#388e3c,color:#000
    style Summary fill:#b0bec5,stroke:#546e7a,color:#000
    style Statistics fill:#b0bec5,stroke:#546e7a,color:#000
    style StatsByProfile fill:#b0bec5,stroke:#546e7a,color:#000
    style StatsByOp fill:#b0bec5,stroke:#546e7a,color:#000
    style Visualizations fill:#b0bec5,stroke:#546e7a,color:#000
    style SaveFile fill:#a5d6a7,stroke:#388e3c,color:#000

Analytics Capabilities:

  • Statistical analysis with percentiles

  • Outlier detection (IQR, z-score)

  • Interactive Plotly visualizations

  • Group-by analysis (profile, strategy, operation)

  • Export to JSON for further analysis


4. Span Evaluator Pipeline

sequenceDiagram
    participant User
    participant SpanEval as Span Evaluator
    participant Phoenix as Telemetry Provider
    participant Eval as Evaluator (per name)

    User->>SpanEval: run_evaluation_pipeline(hours=6, incremental=True)

    Note over SpanEval: Step 1: Fetch Recent Spans
    SpanEval->>Phoenix: get_recent_spans(hours, operation_name, limit)
    activate Phoenix
    Phoenix-->>SpanEval: spans_dataframe
    deactivate Phoenix

    Note over SpanEval: Step 2: Incremental Gate (if incremental=True)
    SpanEval->>Phoenix: _already_evaluated_span_ids(evaluator_names)
    activate Phoenix
    Phoenix->>Phoenix: get_spans + annotations.get_annotations per evaluator
    Phoenix-->>SpanEval: skip_span_ids: {evaluator_name: {span_id,...}}
    deactivate Phoenix

    Note over SpanEval: Step 3: evaluate_spans(spans_df, evaluator_names, skip_span_ids)
    loop For each evaluator_name in evaluator_names
        SpanEval->>SpanEval: Resolve evaluator (golden_evaluator or reference_free_evaluators[name])
        loop For each span not in skip_span_ids[evaluator_name]
            SpanEval->>Eval: evaluate(input=query, output=results, metadata=attributes)
            activate Eval
            Eval-->>SpanEval: EvaluationResult(score, label, explanation)
            deactivate Eval
        end
        SpanEval->>SpanEval: Collect results into evaluator_name -> DataFrame
    end

    Note over SpanEval: Step 4: Upload Evaluations (if upload_evaluations=True)
    SpanEval->>Phoenix: upload_evaluations(eval_results)
    activate Phoenix
    loop For each evaluator's results DataFrame
        Phoenix->>Phoenix: Create + upload annotation per span
    end
    Phoenix-->>SpanEval: Upload complete
    deactivate Phoenix

    Note over SpanEval: Step 5: Generate Summary
    SpanEval->>SpanEval: Aggregate mean_score / score_distribution per evaluator

    SpanEval-->>User: {<br/>num_spans_retrieved: 500,<br/>num_skipped: N,<br/>incremental: true,<br/>evaluators_run: [...],<br/>results: {<br/>  relevance: {num_evaluated, num_skipped, mean_score, score_distribution},<br/>  diversity: {...},<br/>  golden_dataset: {...}<br/>}<br/>}

    Note over User: View results in Phoenix UI

Span Evaluator Features:

  • Retrospective evaluation of existing spans

  • Multiple evaluator support (reference-free, golden dataset) — evaluated sequentially per evaluator name, not in parallel

  • Incremental mode skips (span, evaluator) pairs that already carry that evaluator's annotation

  • Automatic upload to Phoenix for visualization

  • Batch processing of historical traces

  • Summary statistics and distribution analysis


Core Components

1. ExperimentTracker

File: libs/evaluation/cogniverse_evaluation/core/experiment_tracker.py

Purpose: Track and visualize experiments using Inspect AI evaluation framework with Phoenix integration.

Key Attributes:

experiment_project_name: str          # Project name for experiments
output_dir: Path                       # Results directory
enable_quality_evaluators: bool        # Adds the visual quality scorer to the Inspect scorer set (get_configured_scorers via get_visual_scorers)
enable_llm_evaluators: bool           # Adds the visual_judge scorer to the Inspect scorer set
evaluator_name: str                    # Evaluator to use
llm_model: str | None                  # LLM model for evaluators (required when enable_llm_evaluators=True)
llm_base_url: str | None              # Base URL for LLM API
provider: EvaluationProvider           # Evaluation provider
tenant_id: str                         # Tenant identifier
experiments: list[dict]                # Experiment results
configurations: list[dict]             # Experiment configurations
dataset_url: str | None                # Dataset URL

Main Methods:

get_experiment_configurations(profiles: list[str] | None = None, strategies: list[str] | None = None, all_strategies: bool = False) -> list[dict]

Get experiment configurations from strategy registry.

Parameters:

  • profiles: List of profiles to test (None = all)

  • strategies: List of strategies to test (None = common strategies)

  • all_strategies: Test all available strategies

Returns: List of configuration dicts with {profile, strategies: [(name, description)]}

Example:

tracker = ExperimentTracker(tenant_id="your_org:production")

# Get configurations for specific profiles
configs = tracker.get_experiment_configurations(
    profiles=["frame_based_colpali", "chunk_based_xclip"],
    strategies=["binary_binary", "hybrid_float_bm25"]
)

# configs = [
#     {
#         "profile": "frame_based_colpali",
#         "strategies": [
#             ("binary_binary", "Binary"),
#             ("hybrid_float_bm25", "Hybrid Float + Text")
#         ]
#     },
#     ...
# ]


run_experiment(profile: str, strategy: str, dataset_name: str, description: str) -> dict

Run a single experiment using the Inspect AI framework (synchronous).

Parameters:

  • profile: Vespa profile name

  • strategy: Ranking strategy

  • dataset_name: Dataset to evaluate against

  • description: Human-readable experiment description

Returns: Experiment result dict with status, metrics, timestamp

Workflow:

  1. Log experiment start to Phoenix

  2. Create Inspect AI evaluation task

  3. Execute evaluation synchronously (Inspect AI manages its own event loop)

  4. Extract metrics from result

  5. Log completion to Phoenix

  6. Return result dictionary

Example:

tracker = ExperimentTracker(tenant_id="your_org:production")

# Synchronous — do NOT use await or asyncio.run()
result = tracker.run_experiment(
    profile="frame_based_colpali",
    strategy="binary_binary",
    dataset_name="golden_eval_v1",
    description="Frame Based ColPali - Binary"
)

# result = {
#     "status": "success",
#     "profile": "frame_based_colpali",
#     "strategy": "binary_binary",
#     "description": "Frame Based ColPali - Binary",
#     "experiment_name": "frame_based_colpali_binary_binary_20251007_143022",
#     "metrics": {
#         "mrr": 0.85,
#         "recall": 0.92,
#         "relevance": 0.88
#     },
#     "timestamp": "2025-10-07T14:30:22"
# }


create_or_get_dataset(dataset_name: str | None = None, csv_path: str | None = None, force_new: bool = False) -> str

Create or retrieve a dataset for experiments.

Parameters:

  • dataset_name: Name of existing dataset

  • csv_path: Path to CSV file for new dataset

  • force_new: Force creation of new dataset

Returns: Dataset name

Example:

tracker = ExperimentTracker(tenant_id="your_org:production")

# Create from CSV
dataset_name = tracker.create_or_get_dataset(
    dataset_name="my_eval_dataset",
    csv_path="data/testset/evaluation/video_search_queries.csv"
)

# Use existing
dataset_name = tracker.create_or_get_dataset(
    dataset_name="golden_eval_v1"
)


run_all_experiments(dataset_name: str) -> list[dict]

Run all configured experiments.

Parameters:

  • dataset_name: Dataset to evaluate against

Returns: List of experiment result dicts

Output: Prints progress table with success/failure status and metrics

Example:

tracker = ExperimentTracker(tenant_id="your_org:production")

# Configure experiments
tracker.get_experiment_configurations(profiles=["frame_based_colpali"])

# Create dataset
dataset_name = tracker.create_or_get_dataset(csv_path="data/queries.csv")

# Run all experiments
results = tracker.run_all_experiments(dataset_name)

# Output:
# ================================================================
# PHOENIX EXPERIMENTS WITH VISUALIZATION
# ================================================================
# Timestamp: 2025-10-07 14:30:00
# ...
# [1/5] Frame Based ColPali - Binary
#   Strategy: binary_binary
#   ✅ Success
#      mrr: 0.850
#      relevance: 0.880


create_visualization_tables(experiments: list[dict] | None = None, include_quality_metrics: bool = True) -> dict[str, pd.DataFrame]

Create visualization tables from experiment results.

Returns:

{
    "profile_summary": DataFrame,      # Summary by profile
    "detailed_results": DataFrame,     # All experiments with metrics
    "strategy_comparison": DataFrame   # Strategy comparison
}

Example:

tables = tracker.create_visualization_tables()

print(tables["profile_summary"])
# | Profile              | Total | Success | Failed | Success Rate |
# |----------------------|-------|---------|--------|--------------|
# | frame_based_colpali  |   5   |    5    |   0    |   100.0%     |


2. RoutingEvaluator

File: libs/evaluation/cogniverse_evaluation/evaluators/routing_evaluator.py

Purpose: Evaluate routing decisions separately from search quality.

Key Attributes:

provider: TelemetryProvider        # Telemetry provider for querying spans
project_name: str                  # Project name for routing optimization

RoutingOutcome Enum:

SUCCESS = "success"       # Agent completed task successfully
FAILURE = "failure"       # Agent failed, timed out, or returned empty
AMBIGUOUS = "ambiguous"   # Needs human annotation

RoutingMetrics Dataclass:

routing_accuracy: float                    # % successful decisions
confidence_calibration: float              # Correlation(confidence, success)
avg_routing_latency: Optional[float]      # Mean decision time (ms) of the timed decisions; None when none is timed
per_agent_precision: Dict[str, float]     # Precision per agent
per_agent_recall: Dict[str, float]        # Recall per agent
per_agent_f1: Dict[str, float]            # F1 per agent
total_decisions: int                       # Total evaluated
ambiguous_count: int                       # Unclear outcomes

Main Methods:

evaluate_routing_decision(span_data: Dict[str, Any]) -> Tuple[RoutingOutcome, Dict[str, Any]]

Extract and evaluate a single routing decision; delegates to the module function evaluate_routing_span(span_data), which needs no provider.

A decision whose span ended in ERROR is a FAILURE. One that chose an agent, ran inside a request (its parent_id is a non-empty string) and did not end in ERROR is a SUCCESS, whether its status is OK or UNSET (the gateway sets none). A root decision with no request around it is AMBIGUOUS (no_parent_span). A span with no chosen agent or confidence raises ValueError.

Parameters:

  • span_data: Span dict from Phoenix with routing attributes

Returns: (outcome, metrics) tuple

Extracted Metrics:

{
    "chosen_agent": str,        # Agent selected by routing
    "confidence": float,        # Routing confidence score, in [0, 1]
    "latency_ms": Optional[float],  # processing_time when recorded, else the span's duration; None without either
    "success": bool,            # Whether routing succeeded
    "downstream_status": str    # Status description
}

confidence is coerced through cogniverse_foundation.confidence.parse_confidence, which accepts whatever shape a router emits — a float already in [0, 1], a label ("high"/"medium"/"low"), or a percent string ("85%" or "85") — and always returns a clamped [0, 1] float.

Example:

from cogniverse_foundation.telemetry.registry import TelemetryRegistry

# Get telemetry provider
provider = TelemetryRegistry.get(name="phoenix", tenant_id="your_org:production")

tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

evaluator = RoutingEvaluator(provider=provider, project_name=project_name)

span_data = {
    "name": "cogniverse.routing",
    # RoutingEvaluator reads the decision from output.value via read_span_io:
    # read_span_io(span_data)["output"] -> {"chosen_agent": ..., "confidence": ...}
    "attributes.output.value": json.dumps({
        "chosen_agent": "video_search_agent",
        "confidence": 0.92
    }),
    "status_code": "OK"
}

outcome, metrics = evaluator.evaluate_routing_decision(span_data)
# outcome = RoutingOutcome.SUCCESS
# metrics = {
#     "chosen_agent": "video_search_agent",
#     "confidence": 0.92,
#     "latency_ms": 150.5,
#     "success": True,
#     "downstream_status": "completed_successfully"
# }


calculate_metrics(routing_spans: List[Dict[str, Any]]) -> RoutingMetrics

Calculate comprehensive routing metrics from spans.

Parameters:

  • routing_spans: List of routing span dicts

Returns: RoutingMetrics with all calculated metrics

Example:

import asyncio
from cogniverse_foundation.telemetry.registry import TelemetryRegistry

async def calculate_routing_metrics():
    # Get telemetry provider
    tenant_id = "your_org:production"
    provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)
    project_name = f"cogniverse-{tenant_id}"

    evaluator = RoutingEvaluator(provider=provider, project_name=project_name)

    # Get routing spans from Phoenix
    spans = await evaluator.query_routing_spans(limit=100)

    # Calculate metrics
    metrics = evaluator.calculate_metrics(spans)

    print(f"Routing Accuracy: {metrics.routing_accuracy:.2%}")
    print(f"Confidence Calibration: {metrics.confidence_calibration:.3f}")
    print(f"Avg Latency: {metrics.avg_routing_latency:.0f}ms")
    print(f"Video Agent Precision: {metrics.per_agent_precision['video_search_agent']:.2%}")

asyncio.run(calculate_routing_metrics())


async query_routing_spans(start_time: Optional[datetime] = None, end_time: Optional[datetime] = None, limit: int = 100) -> List[Dict[str, Any]]

Query Phoenix for routing spans.

Parameters:

  • start_time: Start of time range

  • end_time: End of time range

  • limit: Max spans to return

Returns: List of routing span dicts

Example:

import asyncio
from datetime import datetime, timedelta
from cogniverse_foundation.telemetry.registry import TelemetryRegistry

async def get_routing_spans():
    # Get telemetry provider
    tenant_id = "your_org:production"
    provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)
    project_name = f"cogniverse-{tenant_id}"

    evaluator = RoutingEvaluator(provider=provider, project_name=project_name)

    # Get last 6 hours of routing decisions
    end_time = datetime.now()
    start_time = end_time - timedelta(hours=6)

    spans = await evaluator.query_routing_spans(
        start_time=start_time,
        end_time=end_time,
        limit=500
    )

    print(f"Retrieved {len(spans)} routing decisions")
    return spans

asyncio.run(get_routing_spans())


3. PhoenixAnalytics

File: libs/telemetry-phoenix/cogniverse_telemetry_phoenix/evaluation/analytics.py

Purpose: Analytics and visualization for Phoenix traces.

Key Attributes:

telemetry_url: str                  # Phoenix endpoint
client: _PhoenixSyncClient          # Phoenix sync client

TraceMetrics Dataclass:

trace_id: str
timestamp: datetime
duration_ms: float
operation: str
status: str
profile: str | None
strategy: str | None
error: str | None
metadata: dict[str, Any]

Main Methods:

get_traces(start_time: datetime | None = None, end_time: datetime | None = None, operation_filter: str | None = None, limit: int = 10000, *, project_name: str) -> list[TraceMetrics]

Fetch traces from Phoenix with filters.

Parameters:

  • start_time: Start of time range

  • end_time: End of time range

  • operation_filter: Regex filter for operation name

  • limit: Max traces to fetch

  • project_name: Phoenix project name (e.g. cogniverse-<tenant_id>)

Returns: List of TraceMetrics objects

Example:

tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

analytics = PhoenixAnalytics()

# Get search operations from last hour
traces = analytics.get_traces(
    start_time=datetime.now() - timedelta(hours=1),
    operation_filter="search_service\\..*",
    project_name=project_name,
)

print(f"Fetched {len(traces)} search traces")


calculate_statistics(traces: list[TraceMetrics], group_by: str | None = None) -> dict[str, Any]

Calculate comprehensive statistics from traces.

Parameters:

  • traces: List of trace metrics

  • group_by: Optional field to group by ("operation", "profile", "strategy")

Returns: Statistics dictionary with:

  • total_requests: Total count

  • time_range: Start/end timestamps

  • response_time: mean, median, min, max, std, P50/P75/P90/P95/P99

  • status: counts, success_rate, error_rate

  • by_{group_by}: Grouped statistics (if group_by specified)

  • temporal: requests/duration by hour

  • outliers: count, percentage, values

Example:

tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

analytics = PhoenixAnalytics()
traces = analytics.get_traces(project_name=project_name)

# Overall stats
stats = analytics.calculate_statistics(traces)
print(f"P95 Latency: {stats['response_time']['p95']:.0f}ms")
print(f"Success Rate: {stats['status']['success_rate']:.2%}")

# Grouped by profile
stats_by_profile = analytics.calculate_statistics(traces, group_by="profile")
for profile, profile_stats in stats_by_profile["by_profile"].items():
    print(f"{profile}: {profile_stats['p95_duration']:.0f}ms P95")


create_time_series_plot(...) -> go.Figure

create_distribution_plot(...) -> go.Figure

create_heatmap(...) -> go.Figure

create_outlier_plot(...) -> go.Figure

create_comparison_plot(...) -> go.Figure

Create interactive Plotly visualizations.

Example:

tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

analytics = PhoenixAnalytics()
traces = analytics.get_traces(project_name=project_name)

# Time series with P50/P95 bands
fig_time = analytics.create_time_series_plot(
    traces,
    metric="duration_ms",
    aggregation="mean",
    time_window="5min"
)
fig_time.show()

# Distribution analysis (4 subplots)
fig_dist = analytics.create_distribution_plot(
    traces,
    metric="duration_ms",
    group_by="profile"
)
fig_dist.show()

# Hour x Day heatmap
fig_heat = analytics.create_heatmap(
    traces,
    x_field="hour",
    y_field="day",
    metric="duration_ms",
    aggregation="mean"
)
fig_heat.show()


generate_report(start_time: datetime | None = None, end_time: datetime | None = None, output_file: str | None = None, *, project_name: str) -> dict[str, Any]

Generate comprehensive analytics report.

Returns: Report dictionary with summary, statistics, and visualizations (as JSON)

Example:

tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

analytics = PhoenixAnalytics()

# Generate last 24h report
report = analytics.generate_report(
    start_time=datetime.now() - timedelta(days=1),
    output_file="outputs/analytics_report.json",
    project_name=project_name,
)

print(f"Analyzed {report['summary']['total_requests']} requests")
print(f"P95 Latency: {report['summary']['p95_response_time']:.0f}ms")
print(f"Outliers: {report['summary']['outlier_percentage']:.1f}%")


4. SpanEvaluator

File: libs/evaluation/cogniverse_evaluation/span_evaluator.py

Purpose: Evaluate existing spans in Phoenix using various evaluators.

Key Attributes:

provider: EvaluationProvider                   # Evaluation provider
project_name: str                              # Project name for telemetry
reference_free_evaluators: dict                # Reference-free evaluators
golden_evaluator: GoldenDatasetEvaluator       # Golden dataset evaluator

Note: SpanEvaluator does not store the tenant_id constructor argument as an attribute — it is only used to resolve the default EvaluationProvider when provider is not passed in.

Main Methods:

async get_recent_spans(hours: int = 6, operation_name: str | None = "search_service.search", limit: int = 1000, require_search_shape: bool = True) -> pd.DataFrame

Retrieve recent spans from Phoenix. An empty frame means genuinely no traffic in the window; a telemetry outage raises (never flattened into an empty frame).

Parameters:

  • hours: Hours to look back

  • operation_name: Filter by operation name (matched server-side)

  • limit: Max spans

  • require_search_shape: When True (default), only spans whose output is a search-result list survive — what the golden/search evaluators consume. QualityMonitor.evaluate_live_traffic passes False so summary/report strings and gateway/routing dicts are returned too, under outputs["value"], instead of being dropped.

Returns: DataFrame with span information

Example:

tenant_id = "your_org:production"
evaluator = SpanEvaluator(
    tenant_id=tenant_id,
    project_name=f"cogniverse-{tenant_id}",
)

# Get last 6 hours of search spans
spans_df = await evaluator.get_recent_spans(
    hours=6,
    operation_name="search_service.search"
)

print(f"Retrieved {len(spans_df)} search spans")


async evaluate_spans(spans_df: pd.DataFrame, evaluator_names: list[str] | None = None, skip_span_ids: dict[str, set[str]] | None = None) -> dict[str, pd.DataFrame]

Evaluate spans using specified evaluators.

Parameters:

  • spans_df: DataFrame of spans to evaluate

  • evaluator_names: List of evaluator names (None = all — reference-free evaluators plus golden_dataset)

  • skip_span_ids: Optional {evaluator_name: {span_id, ...}} of (span, evaluator) pairs to skip (used by run_evaluation_pipeline's incremental mode)

Returns: Dict mapping evaluator name to results DataFrame

Available Evaluators (from create_reference_free_evaluators() plus golden_dataset):

  • relevance: QueryResultRelevanceEvaluator — heuristic relevance from result scores

  • diversity: ResultDiversityEvaluator — unique-video ratio

  • temporal_coverage: TemporalCoverageEvaluator — unique time-segment coverage for video results

  • composite: CompositeEvaluator — weighted combination of relevance + diversity + temporal_coverage

  • golden_dataset: GoldenDatasetEvaluator — comparison against a golden dataset

Example:

tenant_id = "your_org:production"
evaluator = SpanEvaluator(
    tenant_id=tenant_id,
    project_name=f"cogniverse-{tenant_id}",
)

# Get spans
spans_df = await evaluator.get_recent_spans(hours=24)

# Evaluate
eval_results = await evaluator.evaluate_spans(
    spans_df,
    evaluator_names=["relevance", "diversity", "golden_dataset"]
)

# Check results
for eval_name, results_df in eval_results.items():
    mean_score = results_df["score"].mean()
    print(f"{eval_name}: {mean_score:.3f} avg score")


async upload_evaluations(evaluations: dict[str, pd.DataFrame])

Upload evaluation results as annotations.

Example:

tenant_id = "your_org:production"
evaluator = SpanEvaluator(
    tenant_id=tenant_id,
    project_name=f"cogniverse-{tenant_id}",
)

# Evaluate spans
spans_df = await evaluator.get_recent_spans()
eval_results = await evaluator.evaluate_spans(spans_df)

# Upload to Phoenix
await evaluator.upload_evaluations(eval_results)
# Results now visible in Phoenix UI


async run_evaluation_pipeline(hours: int = 6, operation_name: str | None = "search_service.search", evaluator_names: list[str] | None = None, upload_evaluations: bool = True, incremental: bool = True) -> dict[str, Any]

Run complete evaluation pipeline on recent spans.

When incremental=True (default), (span, evaluator) pairs that already carry that evaluator's annotation are skipped — so re-running over the same window only evaluates new spans / new evaluators instead of re-annotating everything. The skip set comes from querying the span annotations already in the telemetry backend (see SpanEvaluator._already_evaluated_span_ids). Pass incremental=False to re-evaluate every retrieved span.

Returns: Summary with num_spans_retrieved, num_skipped, incremental, evaluators_run, and per-evaluator results (each with num_evaluated, num_skipped, mean_score, score_distribution).

Example:

evaluator = SpanEvaluator(tenant_id="acme", project_name="cogniverse-acme")

# Run full pipeline (incremental by default)
summary = await evaluator.run_evaluation_pipeline(
    hours=24,
    evaluator_names=["relevance", "diversity", "golden_dataset"],
    upload_evaluations=True,
)

print(f"Retrieved {summary['num_spans_retrieved']} spans")
print(f"Skipped {summary['num_skipped']} already-evaluated (span, evaluator) pairs")
for eval_name, stats in summary["results"].items():
    print(f"{eval_name}: evaluated={stats['num_evaluated']} skipped={stats['num_skipped']}")


5. Multi-Turn Session Evaluation

Purpose: Evaluate multi-turn conversations at the session level, considering conversation history context.

Session-Level Evaluation Components

PhoenixEvaluationProvider (in libs/telemetry-phoenix/cogniverse_telemetry_phoenix/evaluation/evaluation_provider.py)

Provides session-level evaluation logging:

def log_session_evaluation(
    self,
    session_id: str,
    evaluation_name: str,
    session_score: float,
    session_outcome: str,
    turn_scores: Optional[List[float]] = None,
    explanation: Optional[str] = None,
    metadata: Optional[Dict[str, Any]] = None
) -> None

Parameters:

  • session_id: Unique session identifier

  • evaluation_name: Name of evaluation (e.g., "conversation_quality")

  • session_score: Overall session score (0.0-1.0)

  • session_outcome: Session outcome ("success", "partial", "failure")

  • turn_scores: Optional per-turn scores

  • explanation: Optional explanation

  • metadata: Optional additional metadata

Example:

from cogniverse_evaluation.providers import get_evaluation_provider

provider = get_evaluation_provider(tenant_id="your_org:production")

provider.log_session_evaluation(
    session_id="sess_abc123",
    evaluation_name="conversation_quality",
    session_score=0.9,
    session_outcome="success",
    turn_scores=[0.85, 0.90, 0.95],
    explanation="User successfully found relevant videos",
    metadata={"turns": 3, "topic": "cooking"}
)


LLM-as-Judge Evaluators

The evaluation module provides LLM-based evaluators for video retrieval quality:

1. LLMJudgeCore

Scores query-result relevance via an OAI-compatible LLM endpoint. Used by the quality monitor for live-traffic relevance scoring; _extract_score_from_response parses an X/10 or 0.x rating (or None for an unscored/failed reply), clamped to [0, 1] — an LM reply like "12/10" or "100/10" is capped rather than skewing persisted quality means and the 0.8/0.5 example-classification gates.

from cogniverse_evaluation.evaluators.llm_judge import LLMJudgeCore

judge = LLMJudgeCore(
    model_name="google/gemma-4-e4b-it",
    base_url="http://localhost:11434",
)
score, explanation = judge._extract_score_from_response("Score: 8/10. Relevant.")
# score == 0.8

2. QueryResultRelevanceEvaluator

Evaluates relevance without an LLM, using a heuristic over each result's relevance_score/score field (no embeddings are computed). evaluate is async and takes input/output, matching the shared Evaluator interface.

import asyncio
from cogniverse_evaluation.evaluators.reference_free import QueryResultRelevanceEvaluator

evaluator = QueryResultRelevanceEvaluator(min_score_threshold=0.5)

async def score():
    return await evaluator.evaluate(
        input="machine learning tutorial",
        output=[
            {"video_id": "vid_001", "title": "ML Basics", "score": 0.95},
            {"video_id": "vid_002", "title": "Deep Learning", "score": 0.85},
        ],
    )

result = asyncio.run(score())
# result.score == 0.90, result.label == "highly_relevant"

Sibling reference-free evaluators (cogniverse_evaluation.evaluators.reference_free), all returned together by create_reference_free_evaluators() under the keys relevance / diversity / temporal_coverage / composite:

  • ResultDiversityEvaluator — unique-video ratio (high_diversity ≥0.8, moderate_diversity ≥0.5, else low_diversity); needs ≥2 results.
  • TemporalCoverageEvaluator — counts unique (start_time, end_time) segments across results, normalized to 10 segments (good_coverage ≥0.7, moderate_coverage ≥0.3).
  • CompositeEvaluator(evaluators, weights=None) — runs its component evaluators concurrently via asyncio.gather, combines scores with (normalized) weights, and reports per-component scores/labels in metadata.
  • RetrievalContext — a @dataclass(query, results, metadata=None) context object shared across these evaluators.

Session Tracking Integration

Session evaluation works with the telemetry module's session tracking:

from cogniverse_foundation.telemetry import get_telemetry_manager
from cogniverse_evaluation.providers import get_evaluation_provider

tm = get_telemetry_manager()
provider = get_evaluation_provider(tenant_id="tenant1")

# Start a session
session_id = "user_session_12345"

# Track multiple turns within the session
with tm.session_span("turn_1", tenant_id="tenant1", session_id=session_id):
    # First query-response
    pass

with tm.session_span("turn_2", tenant_id="tenant1", session_id=session_id):
    # Second query-response
    pass

# Evaluate the entire session
provider.log_session_evaluation(
    session_id=session_id,
    evaluation_name="conversation_quality",
    session_score=0.85,
    session_outcome="success",
    turn_scores=[0.80, 0.90],
    explanation="User successfully completed task"
)

Web Client Integration

The web client's search conversation provides unified session evaluation: each result can be rated for relevance, and "Evaluate this conversation" stores an outcome (success, partial, failure) and a 0-1 quality on each search span of the conversation. See Web Client.

Inspect AI Model Provider

File: libs/evaluation/cogniverse_evaluation/core/inspect_model.py

inspect_model(endpoint) gives Inspect AI a model on the cogniverse provider, which sends each call with litellm using the LLMEndpointConfig: the same api_base, key resolution, extra body, headers, sampling and timeout as create_dspy_lm. Inspect AI owns retries; litellm makes one attempt per call. Text and image messages are supported, tool calls are not.

from inspect_ai import eval as inspect_eval

from cogniverse_evaluation.core.inspect_model import inspect_model
from cogniverse_foundation.config.unified_config import LLMEndpointConfig

model = inspect_model(
    LLMEndpointConfig(model="openai/google/gemma-4-e4b-it", api_base=api_base)
)
logs = inspect_eval(task, model=model)

The provider is also registered as the inspect_ai entry point cogniverse, so cogniverse/<litellm model> names it wherever Inspect AI takes a model string.


6. Evaluation Provider System

Files: libs/evaluation/cogniverse_evaluation/providers/base.py, libs/evaluation/cogniverse_evaluation/providers/registry.py

Purpose: Provider-agnostic abstraction for experiment tracking, dataset management, analytics, and monitoring — mirrors the telemetry provider pattern so a non-Phoenix backend (e.g. Langsmith) can be added without touching evaluator code.

Abstract Interfaces (providers/base.py):

class EvaluatorFramework(ABC):
    def get_evaluator_base_class(self) -> type: ...
    def get_evaluation_result_type(self) -> type: ...
    def create_evaluation_result(self, score, label, explanation, metadata=None) -> Any: ...

class EvaluationProvider(ABC):
    def initialize(self, config: dict) -> None: ...
    def create_experiment(self, name, description=None, metadata=None) -> Any: ...
    def create_dataset(self, name, data, description=None, metadata=None) -> Any: ...
    def log_evaluation(self, experiment_id, evaluation_name, score, label=None, explanation=None, metadata=None) -> None: ...
    def create_evaluation_result(self, score, label=None, explanation=None, metadata=None) -> Any: ...
    def get_experiment_url(self, experiment_id: str) -> str: ...
    def get_dataset_url(self, dataset_id: str) -> str: ...

class AnalyticsProvider(ABC):
    async def get_traces(self, start_time=None, end_time=None, operation_filter=None, limit=10000, *, project_name: str) -> list[TraceMetrics]: ...
    def calculate_statistics(self, traces: list[TraceMetrics]) -> dict[str, Any]: ...
    def create_time_series_plot(self, traces, metric="duration") -> Any: ...
    def create_distribution_plot(self, traces, metric="duration") -> Any: ...
    def generate_report(self, traces, format="markdown") -> str: ...

EvaluationProvider is the only interface with a concrete implementation today (PhoenixEvaluationProvider in cogniverse_telemetry_phoenix). AnalyticsProvider documents the contract that PhoenixAnalytics follows but is not literally subclassed by it.

PhoenixEvaluationProvider.create_experiment(name, ...) registers a durable experiment-{name} dataset holding the creation record (event, experiment, description, created_at, metadata) and returns {"id": dataset_name, "name", "description", "metadata", "created_at"}. log_evaluation(experiment_id, ...) appends an evaluation row to that same dataset — so an experiment's full history (creation + every logged evaluation) is readable back from Phoenix — and raises DatasetNotFoundError (a ValueError) if experiment_id was never registered via create_experiment, rather than silently dropping the evaluation. Both are sync facades over the async telemetry.datasets store; calling them from inside a running event loop raises RuntimeError telling the caller to use telemetry.datasets directly instead.

Registry (providers/registry.py):

EvaluationRegistry subclasses cogniverse_foundation.registry.EntryPointRegistry, adding a tenant-scoped default-provider singleton on top of entry-point discovery. Implementations register via the cogniverse.evaluation.providers entry-point group:

[project.entry-points."cogniverse.evaluation.providers"]
phoenix = "cogniverse_telemetry_phoenix.evaluation:PhoenixEvaluationProvider"

Module-level helpers:

from cogniverse_evaluation.providers import (
    get_evaluation_provider,      # get_evaluation_provider(tenant_id=None, name=None, config=None)
    set_evaluation_provider,      # pin a pre-initialized default provider
    register_evaluation_provider, # register_evaluation_provider(name, provider_class) — for testing
    reset_evaluation_provider,    # clear cache + pinned default
)

# tenant_id defaults to SYSTEM_TENANT_ID when omitted
provider = get_evaluation_provider(tenant_id="your_org:production")

7. OnlineEvaluator

File: libs/evaluation/cogniverse_evaluation/online_evaluator.py

Purpose: Score individual cogniverse.routing spans in real time (as opposed to SpanEvaluator's retrospective batch evaluation) and persist the scores as telemetry annotations for drift detection.

Constructor:

OnlineEvaluator(
    provider: TelemetryProvider,
    project_name: str,
    config: OnlineEvaluationConfig | None = None,  # from cogniverse_agents.routing.config
)

When config is omitted, defaults are: enabled=True, sampling_rate=1.0, evaluator_names=["routing_outcome", "confidence_calibration"], persist_scores=True, annotation_name="online_eval".

OnlineEvalResult dataclass:

span_id: str
evaluator_name: str
score: float
label: str
explanation: str
timestamp: datetime

Main Methods:

  • async evaluate_span(span_data: dict) -> list[OnlineEvalResult] — dispatches to _eval_routing_outcome and/or _eval_confidence_calibration per configured evaluator name, respecting sampling_rate; empty list if disabled or not sampled. When persist_scores=True and any result fails to persist as a telemetry annotation, raises RuntimeError naming every failed (span_id, evaluator_name) pair — a silently dropped write would vanish from the drift-detection signal while the caller counted it as persisted.
  • get_statistics() -> dict — returns total_evaluated, total_skipped, sampling_rate, effective_rate, evaluators.

Example:

import asyncio
from cogniverse_evaluation.online_evaluator import OnlineEvaluator
from cogniverse_foundation.telemetry.registry import TelemetryRegistry

async def score_live_span(span_data: dict):
    tenant_id = "your_org:production"
    provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)
    project_name = f"cogniverse-{tenant_id}"
    online_eval = OnlineEvaluator(
        provider=provider,
        project_name=project_name,
    )
    results = await online_eval.evaluate_span(span_data)
    for r in results:
        print(f"{r.evaluator_name}: {r.score:.2f} ({r.label})")

asyncio.run(score_live_span({"context.span_id": "abc123", "status_code": "OK"}))

routing_outcome scores 1.0/0.0/0.5 for SUCCESS/FAILURE/AMBIGUOUS (reusing RoutingEvaluator._classify_routing_outcome); confidence_calibration scores how well the routing span's stated confidence predicted its actual success (well_calibrated / moderately_calibrated / poorly_calibrated).


8. QualityMonitor

File: libs/evaluation/cogniverse_evaluation/quality_monitor.py

Purpose: Continuous, scheduled quality monitor across all agents. Runs two evaluation strategies and decides whether to trigger an Argo optimization workflow — it composes SpanEvaluator, LLMJudgeCore, and the telemetry provider's datasets store rather than reimplementing them. Golden queries come from the tenant's versioned config/golden_set_ground_truth blob via ArtifactManager.

AgentType enum: SEARCH, SUMMARY, REPORT, GATEWAY, ROUTING, QUERY_ENHANCEMENT, ENTITY_EXTRACTION, PROFILE_SELECTION — mapped to span names via SPAN_NAME_BY_AGENT (e.g. "SearchAgent.process"), matching the f"{ClassName}.process" convention emitted by AgentBase.process_span().

Verdict enum: SKIP = 0, OPTIMIZE = 1, FULL = 2.

Constructor:

QualityMonitor(
    tenant_id: str,
    runtime_url: str,
    phoenix_http_endpoint: str,
    llm_base_url: str,
    llm_model: str,
    golden_dataset_path: str,
    argo_api_url: str | None = None,
    argo_namespace: str = "cogniverse",
    golden_eval_interval_seconds: int = 7200,
    live_eval_interval_seconds: int = 14400,
    live_sample_count: int = 20,
    thresholds: QualityThresholds | None = None,
    telemetry_provider=None,
    search_profile: str = "video_colpali_smol500_mv_frame",
    workflow_template: str | None = None,
)

tenant_id is canonicalized to the org:tenant form at construction — every derived name (the Phoenix project live eval reads, quality-baseline-* / optimization-trigger-* dataset names, Argo workflow parameters) uses the canonical form, matching what production span writers emit. Live-traffic evaluation reads the tenant-only user-ops project (cogniverse-{org:tenant}) — the project agents actually write spans to. golden_dataset_path is the one-shot seed source consumed by quality_monitor_cli; the monitor itself reads the tenant blob, not the file. The quality-baseline-* dataset stores one JSON payload column for golden summaries and live per-agent baselines, and the monitor decodes that payload on read so Phoenix's dataframe round-trip cannot stringify metric values. Live evaluation datasets (quality-live-*) use agent as input and retain score, baseline_score, degradation_pct, and sample_count as outputs.

QualityThresholds dataclass (defaults):

golden_mrr_drop_pct: float = 0.10
golden_ndcg_drop_pct: float = 0.10
live_score_floor: float = 0.5
error_rate_ceiling: float = 0.05
latency_p95_ceiling_ms: float = 1000.0
min_samples_for_verdict: int = 10

Golden matching: a golden set names each expected video by its original filename stem (v_-uJnucdW6DY). evaluate_golden_set keys every /search result by result_source_title_key (cogniverse_sdk.document): the source_title the search backend stamps from the schema's document_mapping.title field, which ingestion fills with the original upload basename, with its extension stripped. A tenant whose source_id is a content hash (multipart uploads land in MinIO as {tenant}/{sha256}.{ext}) and one whose source_id is the filename stem key identically. A query that returns a result with no source_title counts as failed.

Result dataclasses: AgentEvalResult (per-agent score/baseline/degradation), GoldenEvalResult (mean_mrr, mean_ndcg, mean_precision_at_5, per-query scores, baseline_mrr captured before the new result is stored, plus failed_query_count/failed_queries naming the /search calls that failed — a partial run is persisted for the record but never used as a comparison baseline), LiveEvalResult (per-AgentType AgentEvalResult map), OptimizationTrigger (payload submitted to the Argo optimization workflow).

Fault contracts: baseline reads (_read_baseline_metric, _get_agent_baseline) return None only for a genuinely absent baseline (first run / only partial runs) and raise on a telemetry outage; the golden blob loader reports {"status": "skipped", "reason": "golden_set_missing"} only when the tenant blob is absent, and raises with {"status": "failed"} on store outages and on corrupt payloads; the golden baseline write raises on failure (a silently lost write would freeze the baseline); the continuous loop treats only the typed missing-blob result as optional, continues live evaluation on its independent cadence, and tries the golden branch again at its next scheduled interval so a later upload is visible; _store_trigger_dataset returns the stored dataset name (or None when there were no example records) and submit_optimization(trigger, trigger_dataset) references exactly that name, so a workflow is never submitted pointing at a dataset that was not created; a span-read outage during evaluate_live_traffic propagates instead of reading as "no traffic".

Spawned workflow pod: workflow_template is the name of the chart's shared optimization WorkflowTemplate. submit_optimization submits a Workflow whose spec is a workflowTemplateRef at it plus the five arguments mode, tenant-id, lookback-hours, agents, trigger-dataset (OPTIMIZATION_WORKFLOW_PARAMETER_NAMES); the template owns the container spec, env, resources, config mount and the per-tenant mutex, so this path and POST /admin/tenant/{id}/optimize spawn the same pod. submit_argo_optimization_workflow raises ValueError when the name is empty rather than submitting a Workflow with no pod spec. See Spawned-Workflow Pod Wiring.

Key methods: check_thresholds(...) decides the Verdict from golden/live results against QualityThresholds, then (when telemetry_provider was passed to the constructor) consults cogniverse_agents.routing.xgboost_meta_models.TrainingDecisionModel.should_train(...) per agent to confirm or override the naive threshold verdict — logged as an override, never silent — falling back to the naive verdicts if the model can't be built or scored. Until a TrainingDecisionModel is trained, the threshold verdicts stand unchanged (logged once per check): its untrained heuristic needs 50 live samples, more than the live_sample_count window holds, so whether there is enough data to train is left to each optimizer's population floor (lookback spans plus approved synthetic data); _build_trigger(...) assembles an OptimizationTrigger when optimization is warranted.

Example:

from cogniverse_evaluation.quality_monitor import QualityMonitor, QualityThresholds

monitor = QualityMonitor(
    tenant_id="your_org:production",
    runtime_url="http://localhost:8000",
    phoenix_http_endpoint="http://localhost:6006",
    llm_base_url="http://localhost:11434",
    llm_model="google/gemma-4-e4b-it",
    golden_dataset_path="data/testset/evaluation/video_search_queries.csv",
    thresholds=QualityThresholds(live_score_floor=0.6),
)

Summarizing routing decisions

File: libs/evaluation/cogniverse_evaluation/evaluators/routing_evaluator.py

summarize_routing_decisions(spans) reads a frame of cogniverse.routing spans with evaluate_routing_span and returns decisions (newest first: span_id, trace_id, start_time in UTC ISO, query, chosen_agent, confidence, outcome, reason, latency_ms (the span's duration) and entity_extraction_failed); total, successes, failures, ambiguous and unreadable (spans it could not read); accuracy (the share that succeeded, None without decisions); confidence_calibration (the Pearson correlation of confidence with success, None when either side is constant or there are fewer than two decisions); latency_ms (mean, p50, p95); and per_agent (agent, decisions, successes, failures, ambiguous, success_rate, mean_confidence, mean_latency_ms, precision, recall, f1; most decisions first, then by agent). GET /admin/tenant/{tenant_id}/routing-decisions serves it.

per_agent_precision_recall_f1(decisions) scores (chosen_agent, succeeded) pairs per agent: a success is a true positive and any other decision a false positive. With no ground truth for where a decision should have gone there are no false negatives, so recall is 1.0 for an agent with a success and 0.0 otherwise. RoutingEvaluator.calculate_metrics uses the same scores.

Scoring recorded searches

File: libs/evaluation/cogniverse_evaluation/recorded_searches.py

score_recorded_searches(spans, golden_rows) scores a tenant's search_service.search spans (SEARCH_SPAN_NAME, recorded by SearchService.search and by every SearchAgent text search) against its canonical golden rows (query, list of expected_videos), without running a search. A span whose stripped query is a golden query is scored under its profile and strategy; the latest successful search per profile, strategy and query counts. Result rows name their source by result_source_title_key, and a source counts once, at its best rank. Each query gets mrr, ndcg (at 10), recall_at_1, recall_at_5 and precision_at_5 from calculate_metrics_suite.

dataset_golden_rows(examples) turns an evaluation dataset's examples (Phoenix's input/output columns, as DatasetManager writes them: input.query, comma-joined output.expected_videos) into the same canonical golden rows; examples without a query or an expected source are left out, and a repeated query keeps its first position and its last expectation.

It returns golden_queries; strategies (per profile and strategy, sorted: queries, the mean of each metric, and success_rate, the share of queries whose first result is expected); queries (per profile and strategy, in golden order: query, expected, retrieved (first 10), searched_at, trace_id and the metrics); unsearched_queries (golden order); failed_searches (searches with an ERROR status); and unscored_searches (searches with a result that has no source title, or no result rows). The runtime serves it as GET /admin/tenant/{tenant_id}/evaluation/golden.


9. Data Layer

Files: libs/evaluation/cogniverse_evaluation/data/{datasets,storage,traces}.py

DatasetManager (data/datasets.py) — sync facade over the telemetry provider's async DatasetStore, used by the CLI's create-dataset command, the runtime's optimization framework router (routers/optimization_framework.py), and scripts/manage_datasets.py:

DatasetManager(tenant_id: str, dataset_store: DatasetStore | None = None)
# tenant_id is canonicalized; the store defaults to the tenant's evaluation
# provider's telemetry.datasets

manager.create_from_csv(csv_path: str, dataset_name: str, description: str | None = None) -> str
manager.create_from_queries(queries: list[dict], dataset_name: str, description: str | None = None) -> str
manager.create_from_json(...) -> str
manager.get_dataset(dataset_name: str) -> dict | None   # None = not found; outage raises
manager.list_datasets() -> list[str]                    # local cache only
manager.update_dataset(dataset_name: str, new_queries: list[dict]) -> bool  # raises ValueError if missing
manager.delete_dataset(dataset_name: str) -> bool
manager.export_dataset(dataset_name: str, output_path: str) -> bool  # raises if missing
manager.create_test_dataset() -> str

A dataset the manager creates is owned by its tenant (the store records tenant_id), which scopes the runtime's dataset evaluation (GET /admin/tenant/{tenant_id}/evaluation/datasets) to it.

expected_videos lists are persisted comma-joined ("v1,v2") — the form that core.ground_truth._resolve_expected_items and core.task split back into item lists. The store raises DatasetNotFoundError (cogniverse_foundation.telemetry.providers.base, a ValueError subclass) for a missing dataset; get_dataset catches it and returns None, while update_dataset/export_dataset surface it (or their own ValueError) to the caller. A backend outage raises the underlying transport error instead of masquerading as no-data.

TelemetryStorage (data/storage.py) — connection/health-check management for the telemetry backend, plus MonitoredSpanExporter (an OTel SpanExporter wrapping export success/failure metrics via ExportMetrics):

TelemetryStorage(config: ConnectionConfig | None = None)

storage.log_experiment_results(...)
storage.get_traces_for_evaluation(trace_ids: list[str] | None = None, start_time: datetime | None = None, limit: int = 1000, *, project: str) -> pd.DataFrame
storage.get_metrics() -> dict
storage.shutdown()
# Context-manager: `with TelemetryStorage() as storage: ...`

get_traces_for_evaluation takes no filter_condition — the underlying provider's get_spans doesn't accept one, so trace_ids filtering happens by fetching the window and matching trace_id/context.trace_id client-side. When the storage isn't connected it raises ConnectionError (it used to degrade to an empty DataFrame, which read as "no traces" indistinguishable from a real outage); once connected, a fetch failure raises RuntimeError.

ConnectionConfig (defaults: http_endpoint="http://localhost:6006", otlp_endpoint="localhost:4317", max_retries=3, connection_timeout_seconds=10.0 (bounds the initial connection probe against the telemetry backend), max_batch_size=512, export_timeout_millis=30000, enable_metrics=True (gates whether MonitoredSpanExporter records success/failure into ExportMetrics at all)) tracks connection health via the ConnectionState enum (DISCONNECTED / CONNECTING / CONNECTED / FAILED).

TraceManager (data/traces.py) — used by the CLI's list-traces command:

tenant_id is canonicalized at construction and the span project is resolved from the loaded telemetry config (get_telemetry_manager().config.get_project_name) — the same tenant_project_template the span writers use — so a template override reads the project production agents actually write to. Batch/live solvers resolve the same way via core/solvers._resolve_project. When storage is not injected, TraceManager also constructs it from that config's required provider_config.http_endpoint and otlp_endpoint; it never falls back to an unrelated localhost backend. A missing query endpoint raises ValueError, and a connection failure names the configured HTTP endpoint.

TraceManager(tenant_id: str, storage: TelemetryStorage | None = None)

manager.get_recent_traces(hours_back: int = 1, limit: int = 100) -> pd.DataFrame
manager.get_traces_by_ids(trace_ids: list[str]) -> pd.DataFrame
manager.extract_trace_data(trace_df: pd.DataFrame) -> list[dict]
manager.get_traces_by_experiment(profile: str, strategy: str, hours_back: int = 24) -> pd.DataFrame
manager.get_trace_statistics(hours_back: int = 24) -> dict
manager.export_traces(output_path: str, hours_back: int = 24) -> bool

extract_trace_data (via the module-level trace_dict_from_span_row) reads the columns the Phoenix span frame actually carries — context.trace_id for identity, start_time/end_time for the derived duration_ms (None when a bound is missing), attributes.output.value parsed from its JSON string — and reconstructs metadata from the flattened attributes.metadata.* columns. get_trace_statistics averages the same start/end-derived durations, excluding rows with a missing bound.

get_traces_by_experiment fetches the window unfiltered, then matches profile/strategy against the flattened attributes.metadata.profile / attributes.metadata.strategy columns client-side — the storage layer takes no filter expression, so values are matched literally (no query-injection surface from profile/strategy names).


10. Ground Truth Strategies

Files: libs/evaluation/cogniverse_evaluation/core/ground_truth.py, core/schema_analyzer.py

Purpose: Extract ground-truth item lists for a query without hardcoding domain assumptions — used by core/solvers.py's batch/live solvers when scoring traces that lack an explicit target.

from cogniverse_evaluation.core.ground_truth import get_ground_truth_strategy

strategy = get_ground_truth_strategy({"ground_truth_strategy": "schema_aware"})
result = await strategy.extract_ground_truth(trace_data, backend=None)
# result = {"expected_items": [...], "confidence": 0.0-1.0, "source": str, "metadata": {...}}

get_ground_truth_strategy(config) maps config["ground_truth_strategy"] to one of four GroundTruthStrategy (ABC) implementations, all with the same async extract_ground_truth(trace_data, backend=None) -> dict contract:

ground_truth_strategy value Class
"schema_aware" (default) SchemaAwareGroundTruthStrategy
"dataset" DatasetGroundTruthStrategy
"backend" BackendGroundTruthStrategy
"hybrid" HybridGroundTruthStrategy

SchemaAwareGroundTruthStrategy delegates schema-specific parsing to a SchemaAnalyzer (core/schema_analyzer.py) resolved via get_schema_analyzer(schema_name, schema_fields); DefaultSchemaAnalyzer is the generic fallback, and SchemaAnalyzerRegistry / register_analyzer(...) is how the video/document/image plugins (see Plugin System below) plug in their own analyzers. Exceptions raised during extraction subclass GroundTruthError (SchemaDiscoveryError, BackendError).

SchemaAwareGroundTruthStrategy caches discovered schema fields per schema name on the strategy instance (self.schema_cache) — discovery costs several backend round-trips (get_schema_info / get_field_mappings / a sample search / list_fields) and the fields are stable for the strategy's lifetime, so only the first trace against a given schema pays that cost. In the get_schema_info-provided path (_parse_schema_info), id-field categorization matches id as a whole token (field == "id", or _id/id_ as a prefix/suffix) rather than a bare substring, so a field like paid_amount or raid_count (both contain "id" as a substring) is never misclassified as an id field the way it would be under plain "id" in field_name.

BackendGroundTruthStrategy delegates to a fresh SchemaAwareGroundTruthStrategy with metadata["high_precision"] = True set on a copy of the trace data (never mutating the caller's dict, so sibling strategies run on the same trace inside HybridGroundTruthStrategy are unaffected). That flag halves the analyzer's max_results search budget (max(1, max_results // 2)), trading recall for precision on the harvested ground truth.


11. Metrics Suite

File: libs/evaluation/cogniverse_evaluation/metrics/custom.py

Pure functions operating on ranked ID lists — used by scorers and QualityMonitor's golden-set evaluation. Every golden consumer (QualityMonitor, the precision_scorer/recall_scorer Inspect scorers over retrieval-solver and trace results, GoldenDatasetEvaluator, and scripts/create_golden_dataset_from_traces.py) builds the ranked ID list from each result's result_source_title_key, never its source_id. The retrieval solver carries each /search result's source_title into its packed results for that. GoldenDatasetEvaluator labels a span whose results carry no source_title not_evaluable; the scorers raise on one:

from cogniverse_evaluation.metrics.custom import (
    calculate_mrr,             # calculate_mrr(results: list[str], expected: list[str]) -> float
    calculate_ndcg,            # calculate_ndcg(results, expected, k: int = 10) -> float
    calculate_precision_at_k,  # calculate_precision_at_k(results, expected, k: int = 5) -> float
    calculate_recall_at_k,     # calculate_recall_at_k(results, expected, k: int = 5) -> float
    calculate_f1_at_k,         # calculate_f1_at_k(results, expected, k: int = 5) -> float
    calculate_map,             # calculate_map(results_list: list[list[str]], expected_list: list[list[str]]) -> float
    calculate_metrics_suite,   # calculate_metrics_suite(results, expected, k_values: list[int] | None = None) -> dict[str, float]
)

metrics = calculate_metrics_suite(
    results=["vid_003", "vid_001", "vid_007"],
    expected=["vid_001", "vid_002"],
    k_values=[1, 5, 10],
)
# metrics = {"mrr": 0.5, "ndcg": ..., "precision@1": 0.0, "recall@1": 0.0, "f1@1": 0.0, ...}

calculate_metrics_suite defaults k_values to [1, 5, 10] and always includes mrr and ndcg (unparameterized) plus precision@k/recall@k/f1@k for each k. Ranked identifiers receive relevance credit at most once, so duplicates still consume rank but cannot inflate a metric above 1.0. Cutoffs must be non-negative.


12. Evaluator Base Classes

File: libs/evaluation/cogniverse_evaluation/evaluators/base.py

Purpose: Provider-delegating base class so evaluator subclasses (e.g. GoldenDatasetEvaluator, QueryResultRelevanceEvaluator, ConfigurableVisualJudge) work against whichever EvaluationProvider is active, without importing a concrete provider type.

class Evaluator:
    """Subclasses implement `evaluate(...)`. `__init_subclass__` injects the
    active provider's evaluator base class into the MRO at subclass-definition
    time (degrades to a debug log, not a hard failure, if the provider isn't
    available yet)."""

class EvaluationResult:
    """`EvaluationResult(score, label, explanation, metadata=None)` — a thin
    `__new__` shim that delegates to `create_evaluation_result`, so callers
    get the provider-specific result type without importing it directly."""

Module-level helpers: get_evaluator_base_class(), get_evaluation_result_type(), and create_evaluation_result(score, label, explanation, metadata=None) — all resolve the current provider via get_evaluation_provider() and delegate to provider.framework.

SyncQueryResultRelevanceEvaluator / SyncResultDiversityEvaluator (evaluators/sync_reference_free.py) provide synchronous (non-async) equivalents of QueryResultRelevanceEvaluator/ResultDiversityEvaluator for call sites that can't await (e.g. Inspect AI scorers); create_sync_evaluators() returns both as a list.


13. CLI

File: libs/evaluation/cogniverse_evaluation/cli.py

Purpose: click-based command group (cogniverse-eval) wrapping evaluation_task (experiment/batch/live modes), dataset creation, and trace listing.

Every command reads the deployment's Phoenix from TELEMETRY_HTTP_ENDPOINT and TELEMETRY_OTLP_ENDPOINT (the chart sets both on the runtime pod) through configure_telemetry_endpoints. A dataset that cannot be loaded fails the run with Loading dataset '<name>' from Phoenix at <endpoint> failed: <error>.

# Run an experiment-mode evaluation
cogniverse-eval evaluate --mode experiment --dataset test_dataset \
    -p frame_based_colpali -s binary_binary

# Evaluate existing traces (batch mode)
cogniverse-eval evaluate --mode batch --dataset test_dataset \
    --tenant-id acme:acme -t trace_id_1 -t trace_id_2

# Live evaluation
cogniverse-eval evaluate --mode live --dataset test_dataset \
    --tenant-id acme:acme

# Create a dataset from CSV
cogniverse-eval create-dataset --name my_dataset --tenant-id acme:acme --csv queries.csv

# List recent traces
cogniverse-eval list-traces --tenant-id acme:acme --hours 2 --limit 50

# Quick smoke test (optionally --tenant-id, defaults to the system tenant)
cogniverse-eval test

inspect_ai.eval returns EvalLogs (a list[EvalLog]). evaluate and test check every log's terminal status: anything other than success prints Inspect evaluation <eval_id>: status=<status> through the failure path, exits 1, and writes no --output file. On success --output receives one row per scored sample: eval_id, sample_id, epoch, input, target, trace_ids and each scorer's value/explanation.

evaluate --mode experiment requires both --profiles/-p and --strategies/-s; evaluate --mode batch/--mode live require --tenant-id (or project_name/tenant_id in the --config file) to resolve the span project to read; create-dataset requires --tenant-id and either --csv or --queries-json; list-traces requires --tenant-id unconditionally (no config-file fallback); test accepts an optional --tenant-id (defaults to SYSTEM_TENANT_ID) so the smoke test runs with zero flags. All commands accept -v/--verbose for debug logging.


14. RootCauseAnalyzer

File: libs/evaluation/cogniverse_evaluation/analysis/root_cause_analysis.py

Purpose: Automated root-cause analysis over a batch of trace objects — separates failed/successful/performance-degraded traces, mines failure patterns, and generates ranked hypotheses with suggested actions.

RootCauseAnalyzer()  # no constructor args; loads a built-in known-issues table
    # (timeout, memory, connection, rate_limit, model_error, data_format —
    # each a regex pattern + category + suggested_action)

analyzer.analyze_failures(
    traces: list[Any],
    include_performance: bool = True,
    performance_threshold_percentile: int = 95,
) -> dict[str, Any]

analyzer.generate_rca_report(analysis: dict, format: str = "markdown") -> str

analyze_failures expects each trace object to expose .status, .error, and .duration_ms attributes; it partitions traces into failed / successful / performance-degraded (successful but above the given duration percentile), then mines FailurePatterns and produces RootCauseHypothesis objects:

The failure, temporal, and performance comparison passes key off an object-identity set built once per batch, so correlation stays linear on large trace batches. Duplicate trace_id values are allowed because the analyzer is operating on span objects, not on trace-id uniqueness; the same trace_id can appear on multiple span records without changing the membership check. trace_id is still used for reporting and links, but not for failure membership.

@dataclass
class FailurePattern:
    pattern_type: str        # 'operation' | 'profile' | 'strategy' | 'time' | 'parameter'
    pattern_value: Any
    failure_rate: float
    occurrence_count: int
    confidence: float
    examples: list[str]
    correlation_strength: float = 0.0

@dataclass
class RootCauseHypothesis:
    hypothesis: str
    confidence: float
    evidence: list[str]
    affected_traces: list[str]
    suggested_action: str
    category: str            # 'configuration' | 'resource' | 'timeout' | 'data' | 'model'
    patterns: list[FailurePattern]

generate_rca_report(analysis, format="markdown") renders the analyze_failures output as a Markdown or HTML report (any other format value falls back to str(analysis)).


Usage Examples

Example 1: Run Experiment Suite with Visualization

"""
Complete experiment workflow with Phoenix visualization.
"""
from cogniverse_evaluation.core.experiment_tracker import ExperimentTracker

# Initialize tracker
tracker = ExperimentTracker(
    tenant_id="your_org:production",
    experiment_project_name="my_experiments",
    enable_quality_evaluators=True,
    enable_llm_evaluators=False
)

# Get experiment configurations
configs = tracker.get_experiment_configurations(
    profiles=["frame_based_colpali", "chunk_based_xclip"],
    strategies=["binary_binary", "hybrid_float_bm25"]
)

print(f"Configured {len(configs)} profiles with {sum(len(c['strategies']) for c in configs)} experiments")

# Create dataset
dataset_name = tracker.create_or_get_dataset(
    dataset_name="golden_eval_v1",
    csv_path="data/testset/evaluation/video_search_queries.csv"
)

# Run all experiments
results = tracker.run_all_experiments(dataset_name)

# Create visualization tables
tables = tracker.create_visualization_tables()

# Print results
tracker.print_visualization(tables)

# Save results
csv_path, json_path = tracker.save_results()
print(f"\nResults saved:")
print(f"  CSV: {csv_path}")
print(f"  JSON: {json_path}")

# Phoenix UI links printed automatically

Output:

================================================================
PHOENIX EXPERIMENTS WITH VISUALIZATION
================================================================
Timestamp: 2025-10-07 14:30:00
Experiment Project: my_experiments
Dataset: golden_eval_v1

Quality Evaluators: ✅ ENABLED
LLM Evaluators: ❌ DISABLED

============================================================
Profile: frame_based_colpali
============================================================

[1/4] Frame Based ColPali - Binary
  Strategy: binary_binary
  ✅ Success
     mrr: 0.850
     relevance: 0.880
...

🔗 Dataset: http://localhost:6006/datasets/golden_eval_v1
🔗 Experiments Project: http://localhost:6006/projects/my_experiments


Example 2: Evaluate Routing Decisions

"""
Analyze routing decision quality from Phoenix spans.
"""
import asyncio
from cogniverse_evaluation.evaluators.routing_evaluator import RoutingEvaluator
from cogniverse_foundation.telemetry.registry import TelemetryRegistry
from datetime import datetime, timedelta

async def evaluate_routing_decisions():
    """Evaluate routing decision quality."""
    # Get telemetry provider
    tenant_id = "your_org:production"
    provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)

    # Initialize evaluator for routing project
    evaluator = RoutingEvaluator(
        provider=provider,
        project_name=f"cogniverse-{tenant_id}"
    )

    # Get routing spans from last 24 hours
    end_time = datetime.now()
    start_time = end_time - timedelta(hours=24)

    routing_spans = await evaluator.query_routing_spans(
        start_time=start_time,
        end_time=end_time,
        limit=500
    )

    print(f"Retrieved {len(routing_spans)} routing decisions from Phoenix")

    # Calculate metrics
    metrics = evaluator.calculate_metrics(routing_spans)

    # Print overall metrics
    print(f"\n{'='*60}")
    print("ROUTING EVALUATION RESULTS")
    print(f"{'='*60}")
    print(f"\nOverall Metrics:")
    print(f"  Total Decisions: {metrics.total_decisions}")
    print(f"  Routing Accuracy: {metrics.routing_accuracy:.2%}")
    print(f"  Confidence Calibration: {metrics.confidence_calibration:.3f}")
    print(f"  Avg Routing Latency: {metrics.avg_routing_latency:.0f}ms")
    print(f"  Ambiguous Decisions: {metrics.ambiguous_count} ({metrics.ambiguous_count/metrics.total_decisions:.1%})")

    # Print per-agent metrics
    print(f"\nPer-Agent Metrics:")
    for agent in metrics.per_agent_precision.keys():
        precision = metrics.per_agent_precision[agent]
        recall = metrics.per_agent_recall[agent]
        f1 = metrics.per_agent_f1[agent]
        print(f"\n  {agent}:")
        print(f"    Precision: {precision:.2%}")
        print(f"    Recall: {recall:.2%}")
        print(f"    F1 Score: {f1:.3f}")

# Run async evaluation
asyncio.run(evaluate_routing_decisions())

Output:

Retrieved 478 routing decisions from Phoenix

============================================================
ROUTING EVALUATION RESULTS
============================================================

Overall Metrics:
  Total Decisions: 478
  Routing Accuracy: 87.45%
  Confidence Calibration: 0.723
  Avg Routing Latency: 152ms
  Ambiguous Decisions: 12 (2.5%)

Per-Agent Metrics:

  video_search_agent:
    Precision: 92.00%
    Recall: 88.50%
    F1 Score: 0.902

  text_agent:
    Precision: 85.00%
    Recall: 79.20%
    F1 Score: 0.820


Example 3: Phoenix Analytics and Visualization

"""
Generate analytics reports with visualizations.
"""
from cogniverse_telemetry_phoenix.evaluation.analytics import PhoenixAnalytics
from datetime import datetime, timedelta

analytics = PhoenixAnalytics(telemetry_url="http://localhost:6006")

# Define analysis period
end_time = datetime.now()
start_time = end_time - timedelta(days=7)
tenant_id = "your_org:production"
project_name = f"cogniverse-{tenant_id}"

# Fetch traces
traces = analytics.get_traces(
    start_time=start_time,
    end_time=end_time,
    operation_filter="search_service\\..*",
    limit=10000,
    project_name=project_name,
)

print(f"Analyzing {len(traces)} search traces from last 7 days")

# Calculate overall statistics
stats = analytics.calculate_statistics(traces)

print(f"\nOverall Statistics:")
print(f"  Total Requests: {stats['total_requests']}")
print(f"  Mean Latency: {stats['response_time']['mean']:.0f}ms")
print(f"  P95 Latency: {stats['response_time']['p95']:.0f}ms")
print(f"  P99 Latency: {stats['response_time']['p99']:.0f}ms")
print(f"  Success Rate: {stats['status']['success_rate']:.2%}")
print(f"  Outliers: {stats['outliers']['percentage']:.1f}%")

# Calculate grouped statistics
stats_by_profile = analytics.calculate_statistics(traces, group_by="profile")

print(f"\nPer-Profile Statistics:")
for profile, profile_stats in stats_by_profile["by_profile"].items():
    print(f"\n  {profile}:")
    print(f"    Count: {profile_stats['count']}")
    print(f"    Mean: {profile_stats['mean_duration']:.0f}ms")
    print(f"    P95: {profile_stats['p95_duration']:.0f}ms")
    print(f"    Error Rate: {profile_stats['error_rate']:.2%}")

# Create visualizations
print("\nGenerating visualizations...")

# Time series with percentile bands
fig_time = analytics.create_time_series_plot(
    traces,
    metric="duration_ms",
    aggregation="mean",
    time_window="1h"
)
fig_time.write_html("outputs/time_series.html")

# Distribution analysis (4 subplots)
fig_dist = analytics.create_distribution_plot(
    traces,
    metric="duration_ms",
    group_by="profile"
)
fig_dist.write_html("outputs/distribution.html")

# Heatmap
fig_heat = analytics.create_heatmap(
    traces,
    x_field="hour",
    y_field="day",
    metric="duration_ms"
)
fig_heat.write_html("outputs/heatmap.html")

# Outlier detection
fig_outlier = analytics.create_outlier_plot(traces)
fig_outlier.write_html("outputs/outliers.html")

# Comparison across profiles
fig_compare = analytics.create_comparison_plot(
    traces,
    compare_field="profile",
    metric="duration_ms"
)
fig_compare.write_html("outputs/comparison.html")

print("✅ Visualizations saved to outputs/")

# Generate comprehensive report
report = analytics.generate_report(
    start_time=start_time,
    end_time=end_time,
    output_file="outputs/analytics_report.json",
    project_name=project_name,
)

print(f"\n✅ Full report saved to outputs/analytics_report.json")

Example 4: Retrospective Span Evaluation

"""
Evaluate existing Phoenix spans and upload results.
"""
import asyncio
from cogniverse_evaluation.span_evaluator import SpanEvaluator

async def evaluate_historical_spans():
    """Evaluate spans from the past week."""
    tenant_id = "your_org:production"
    evaluator = SpanEvaluator(
        tenant_id=tenant_id,
        project_name=f"cogniverse-{tenant_id}",
    )

    # Get spans from last week
    spans_df = await evaluator.get_recent_spans(
        hours=24 * 7,  # 7 days
        operation_name="search_service.search",
        limit=5000
    )

    print(f"Retrieved {len(spans_df)} search spans from last week")

    # Run evaluations
    print("\nRunning evaluations...")
    eval_results = await evaluator.evaluate_spans(
        spans_df,
        evaluator_names=["relevance", "diversity", "golden_dataset"]
    )

    # Print results
    print(f"\nEvaluation Results:")
    for eval_name, results_df in eval_results.items():
        mean_score = results_df["score"].mean()
        distribution = results_df["label"].value_counts()

        print(f"\n  {eval_name}:")
        print(f"    Evaluated: {len(results_df)} spans")
        print(f"    Mean Score: {mean_score:.3f}")
        print(f"    Distribution:")
        for label, count in distribution.items():
            print(f"      {label}: {count} ({count/len(results_df):.1%})")

    # Upload to Phoenix
    print("\nUploading evaluations...")
    await evaluator.upload_evaluations(eval_results)

    print("✅ Evaluations uploaded to Phoenix UI")
    print("   View at: http://localhost:6006/projects/default")

# Run async evaluation
asyncio.run(evaluate_historical_spans())

Output:

Retrieved 3,245 search spans from last week

Running evaluations...

Evaluation Results:

  relevance:
    Evaluated: 3245 spans
    Mean Score: 0.782
    Distribution:
      relevant: 2534 (78.1%)
      not_relevant: 711 (21.9%)

  diversity:
    Evaluated: 3245 spans
    Mean Score: 0.845
    Distribution:
      high_diversity: 2107 (64.9%)
      medium_diversity: 892 (27.5%)
      low_diversity: 246 (7.6%)

  golden_dataset:
    Evaluated: 127 spans
    Mean Score: 0.912
    Distribution:
      exact_match: 89 (70.1%)
      partial_match: 32 (25.2%)
      no_match: 6 (4.7%)

Uploading evaluations...
✅ Evaluations uploaded to Phoenix UI
   View at: http://localhost:6006/projects/default


Example 5: Production Monitoring Pipeline

"""
Production monitoring with routing + analytics + span evaluation.
"""
import asyncio
from datetime import datetime, timedelta
from cogniverse_evaluation.evaluators.routing_evaluator import RoutingEvaluator
from cogniverse_telemetry_phoenix.evaluation.analytics import PhoenixAnalytics
from cogniverse_evaluation.span_evaluator import SpanEvaluator

async def production_monitoring_pipeline():
    """Complete monitoring pipeline for production system."""

    # Time range: last 6 hours
    end_time = datetime.now()
    start_time = end_time - timedelta(hours=6)

    print("="*70)
    print("PRODUCTION MONITORING PIPELINE")
    print("="*70)
    print(f"Time Range: {start_time.strftime('%Y-%m-%d %H:%M')} → {end_time.strftime('%Y-%m-%d %H:%M')}")

    # 1. Routing evaluation
    print("\n[1/3] Evaluating Routing Decisions...")

    # Get telemetry provider
    from cogniverse_foundation.telemetry.registry import TelemetryRegistry
    tenant_id = "your_org:production"
    provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)
    project_name = f"cogniverse-{tenant_id}"

    routing_eval = RoutingEvaluator(
        provider=provider,
        project_name=project_name,
    )
    routing_spans = await routing_eval.query_routing_spans(
        start_time=start_time,
        end_time=end_time,
        limit=1000
    )
    routing_metrics = routing_eval.calculate_metrics(routing_spans)

    print(f"   Routing Accuracy: {routing_metrics.routing_accuracy:.2%}")
    print(f"   Avg Latency: {routing_metrics.avg_routing_latency:.0f}ms")
    print(f"   Confidence Calibration: {routing_metrics.confidence_calibration:.3f}")

    # Alert if routing accuracy drops
    if routing_metrics.routing_accuracy < 0.80:
        print("   ⚠️  WARNING: Routing accuracy below 80%!")

    # 2. Analytics and outlier detection
    print("\n[2/3] Analyzing Search Performance...")
    tenant_id = "your_org:production"
    project_name = f"cogniverse-{tenant_id}"
    analytics = PhoenixAnalytics()
    traces = analytics.get_traces(
        start_time=start_time,
        end_time=end_time,
        operation_filter="search_service\\..*",
        project_name=project_name,
    )
    stats = analytics.calculate_statistics(traces)

    print(f"   Total Requests: {stats['total_requests']}")
    print(f"   P95 Latency: {stats['response_time']['p95']:.0f}ms")
    print(f"   Success Rate: {stats['status']['success_rate']:.2%}")
    print(f"   Outliers: {stats['outliers']['percentage']:.1f}%")

    # Alert if P95 latency is high
    if stats['response_time']['p95'] > 1000:
        print("   ⚠️  WARNING: P95 latency above 1000ms!")

    # Alert if outlier percentage is high
    if stats['outliers']['percentage'] > 5:
        print(f"   ⚠️  WARNING: High outlier percentage ({stats['outliers']['percentage']:.1f}%)!")

    # 3. Span quality evaluation
    print("\n[3/3] Evaluating Search Quality...")
    span_eval = SpanEvaluator(
        tenant_id=tenant_id,
        project_name=project_name,
    )
    spans_df = await span_eval.get_recent_spans(hours=6, limit=500)
    eval_results = await span_eval.evaluate_spans(
        spans_df,
        evaluator_names=["relevance", "diversity"]
    )

    for eval_name, results_df in eval_results.items():
        mean_score = results_df["score"].mean()
        print(f"   {eval_name.capitalize()}: {mean_score:.3f}")

        # Alert if quality drops
        if mean_score < 0.70:
            print(f"   ⚠️  WARNING: {eval_name} score below 0.70!")

    # Upload evaluations
    await span_eval.upload_evaluations(eval_results)

    print("\n" + "="*70)
    print("MONITORING COMPLETE")
    print("="*70)
    print(f"View detailed metrics: http://localhost:6006/projects/default")

# Run monitoring pipeline
asyncio.run(production_monitoring_pipeline())

Output:

======================================================================
PRODUCTION MONITORING PIPELINE
======================================================================
Time Range: 2025-10-07 08:30 → 2025-10-07 14:30

[1/3] Evaluating Routing Decisions...
   Routing Accuracy: 88.50%
   Avg Latency: 145ms
   Confidence Calibration: 0.745

[2/3] Analyzing Search Performance...
   Total Requests: 1,247
   P95 Latency: 782ms
   Success Rate: 96.80%
   Outliers: 3.2%

[3/3] Evaluating Search Quality...
   Relevance: 0.812
   Diversity: 0.867

======================================================================
MONITORING COMPLETE
======================================================================
View detailed metrics: http://localhost:6006/projects/default


Production Considerations

1. Experiment Management

Dataset Versioning:

# Use versioned dataset names
dataset_name = tracker.create_or_get_dataset(
    dataset_name=f"golden_eval_v{version}",
    csv_path="data/golden_dataset.csv"
)

# Track dataset metadata
metadata = {
    "version": "v3",
    "created": datetime.now().isoformat(),
    "num_queries": 500,
    "source": "production_logs"
}

Experiment Reproducibility:

  • Save experiment configurations to JSON

  • Version control evaluation code

  • Record model versions, strategy parameters

  • Store Phoenix project URLs for trace lookup

Cost Management:

  • Limit LLM evaluator usage (expensive)

  • Use quality evaluators first (cheap, fast)

  • Sample large datasets for quick validation

  • Cache evaluation results


2. Routing Evaluation Best Practices

Confidence Calibration Monitoring:

# Good calibration: high correlation (>0.7)
# Poor calibration: low correlation (<0.3)

if routing_metrics.confidence_calibration < 0.5:
    logger.warning(
        "Poor confidence calibration - routing confidence scores "
        "don't predict success well. Consider retraining routing model."
    )

Per-Agent Precision Tracking:

# Identify underperforming agents
for agent, precision in routing_metrics.per_agent_precision.items():
    if precision < 0.75:
        logger.warning(
            f"Agent {agent} has low precision ({precision:.2%}). "
            f"Review agent capabilities or routing logic."
        )

Ambiguous Decision Handling:

# High ambiguous count indicates need for better outcome detection
ambiguous_rate = routing_metrics.ambiguous_count / routing_metrics.total_decisions

if ambiguous_rate > 0.10:
    logger.warning(
        f"High ambiguous decision rate ({ambiguous_rate:.1%}). "
        f"Improve outcome classification or add ground truth labels."
    )


3. Analytics and Alerting

Automated Alerting:

def check_performance_alerts(stats: dict):
    """Check for performance degradation."""
    alerts = []

    # P95 latency alert
    if stats['response_time']['p95'] > 1000:
        alerts.append({
            "level": "warning",
            "metric": "p95_latency",
            "value": stats['response_time']['p95'],
            "threshold": 1000,
            "message": "P95 latency exceeds 1000ms"
        })

    # Error rate alert
    if stats['status']['error_rate'] > 0.05:
        alerts.append({
            "level": "critical",
            "metric": "error_rate",
            "value": stats['status']['error_rate'],
            "threshold": 0.05,
            "message": "Error rate exceeds 5%"
        })

    # Outlier percentage alert
    if stats['outliers']['percentage'] > 10:
        alerts.append({
            "level": "warning",
            "metric": "outlier_percentage",
            "value": stats['outliers']['percentage'],
            "threshold": 10,
            "message": "Outlier percentage exceeds 10%"
        })

    return alerts

Trend Analysis:

# Compare current vs historical performance
def detect_performance_regression(current_stats, historical_baseline):
    """Detect if performance has degraded."""

    # P95 latency regression
    current_p95 = current_stats['response_time']['p95']
    baseline_p95 = historical_baseline['response_time']['p95']

    if current_p95 > baseline_p95 * 1.2:  # 20% degradation
        return {
            "regression_detected": True,
            "metric": "p95_latency",
            "current": current_p95,
            "baseline": baseline_p95,
            "degradation_pct": (current_p95 - baseline_p95) / baseline_p95
        }

    return {"regression_detected": False}


4. Span Evaluation at Scale

Batch Processing:

async def evaluate_spans_in_batches(span_ids: list[str], batch_size: int = 100):
    """Evaluate large number of spans in batches."""
    results = []

    for i in range(0, len(span_ids), batch_size):
        batch = span_ids[i:i+batch_size]
        batch_results = await evaluator.evaluate_spans(batch)
        results.extend(batch_results)

        # Rate limiting
        if i + batch_size < len(span_ids):
            await asyncio.sleep(1)

    return results

Sampling for Large Datasets:

# Sample spans for quick evaluation
sampled_spans = spans_df.sample(n=min(1000, len(spans_df)))
eval_results = await evaluator.evaluate_spans(sampled_spans)

Caching Evaluation Results:

# Cache evaluations to avoid re-evaluation
evaluation_cache = {}

def get_or_evaluate_span(span_id, evaluator):
    if span_id in evaluation_cache:
        return evaluation_cache[span_id]

    result = evaluator.evaluate(span_id)
    evaluation_cache[span_id] = result
    return result


Inspect AI Integration

The evaluation module integrates with Inspect AI for structured evaluation tasks.

Scorers (core/inspect_scorers.py)

Scorers that unpack the structured solver output and score the search results. Precision/recall score against the sample's ground-truth target.

Available Scorers:

Scorer Description
relevance_scorer() Keyword-based relevance (schema-agnostic)
diversity_scorer() Result diversity using video_id deduplication
result_count_scorer() Normalized count of returned results
precision_scorer() Precision@k vs the sample's ground-truth target
recall_scorer() Recall@k vs the sample's ground-truth target

Configuration:

from cogniverse_evaluation.core.inspect_scorers import get_configured_scorers

scorers = get_configured_scorers({
    "use_relevance": True,
    "use_diversity": True,
    "use_result_count": True,
    "use_precision_recall": True,  # precision@k / recall@k vs target
})

Solvers (core/solvers.py)

Solvers execute searches and collect results for scorer evaluation.

Available Solvers:

from cogniverse_evaluation.core.solvers import (
    create_retrieval_solver,
    create_batch_solver,
    create_live_solver
)

# New search solver - runs actual searches
retrieval_solver = create_retrieval_solver(
    profiles=["video_colpali_smol500_mv_frame"],
    strategies=["hybrid_float_bm25", "binary_binary"],
    config={"top_k": 10}
)

# Batch solver - loads existing Phoenix traces with ground truth extraction and
# scores each dataset sample against the traces that answered ITS query.
# ground_truth_strategy must be one of "schema_aware" (default), "dataset",
# "backend", or "hybrid" — any other value silently falls back to schema_aware
# (see core/ground_truth.py::get_ground_truth_strategy).
# The span project comes from "project_name" or is derived from "tenant_id"
# (canonicalized, same derivation the span writers use); the solver raises if
# neither is configured. Trace dicts are built from the columns the Phoenix
# span frame actually carries: context.trace_id for identity, and
# start_time/end_time for the derived duration_ms. A span whose output.value
# does not parse raises rather than reading as a trace that retrieved nothing.
#
# Both the batch and live solvers set state.output to the packed solver-output
# contract the scorers read: one search config per matching trace, keyed by
# trace id, plus state.metadata["trace_ids"]. No trace for the sample's query,
# and no spans at all, are both evaluation errors.
batch_solver = create_batch_solver(
    trace_ids=None,  # None for recent traces
    config={
        "tenant_id": "acme:acme",
        "hours_back": 24,
        "limit": 100,
        "ground_truth_strategy": "hybrid"
    }
)

# Live solver - monitors and evaluates live traces
live_solver = create_live_solver(
    config={
        "tenant_id": "acme:acme",
        "poll_interval": 10,
        "max_iterations": 10
    }
)

Solver output serialization (core/solver_output.py): solvers pass rich result data through Inspect AI's string-only solver→scorer interface via a pack_solver_output(query, search_results, phoenix_trace_id=None, metadata=None) -> str / unpack_solver_output(output_str: str) -> EvaluationOutput pair. EvaluationOutput is a @dataclass(query, search_configs, phoenix_trace_id=None, metadata=None) with to_json()/from_json(). from_json raises ValueError on anything that is not a complete, successful result — unparseable JSON, an empty query, no search configs, a config that did not succeed, or results that are not objects — so a broken payload surfaces as an evaluation error instead of a fabricated 0.0 score. The default scorers propagate that error rather than catching it.

Plugin System

Schema analyzers provide domain-specific evaluation capabilities.

VideoSchemaAnalyzer (plugins/video_analyzer.py)

Analyzes video-specific schemas and queries.

Detection Logic:

def can_handle(self, schema_name: str, schema_fields: dict) -> bool:
    # Checks for "video", "frame", "clip" in schema name
    # Or video-specific fields: video_id, frame_id, audio_transcript, etc.

Query Analysis:

analyzer = VideoSchemaAnalyzer()
constraints = analyzer.analyze_query(
    query="first 30 seconds with cars driving",
    schema_fields={"temporal_fields": ["start_time", "end_time"]}
)
# Returns:
# {
#     "query_type": "video_temporal",
#     "temporal_constraints": {"first_n_seconds": ("30",)},
#     "visual_descriptors": {"motions": ["driving"]},
#     "audio_constraints": {},
#     "frame_constraints": {}
# }

Supported Patterns:

Pattern Constraint Type
first N seconds first_n_seconds
at MM:SS at_timestamp
between MM:SS and MM:SS time_range
frame N frame_number
Colors (red, blue, etc.) visual_descriptors.colors
Motion words (running, driving) visual_descriptors.motions
Scene types (indoor, outdoor) visual_descriptors.scenes
"quoted speech" audio_constraints.exact_speech

VideoTemporalAnalyzer (plugins/video_analyzer.py): a more specialized SchemaAnalyzer that only claims schemas which are both video-like and carry start_time/end_time in their temporal_fields (VideoSchemaAnalyzer handles the video-but-no-temporal-fields case). analyze_query delegates to VideoSchemaAnalyzer first, then layers on advanced patterns: N seconds before/after <event>, during <event>, and throughout the video.

DocumentSchemaAnalyzer (plugins/document_analyzer.py)

Analyzes document/text search schemas.

Detection Logic:

def can_handle(self, schema_name: str, schema_fields: dict) -> bool:
    # Checks for "document", "text", "article", "page" in schema name
    # Or document-specific fields: document_id, title, author, content, etc.

Query Analysis:

analyzer = DocumentSchemaAnalyzer()
constraints = analyzer.analyze_query(
    query='author:"smith" title:"machine learning" after:2024-01-01',
    schema_fields={}
)
# Returns:
# {
#     "query_type": "document_author",
#     "author_constraints": {"author": "smith"},
#     "field_constraints": {"title": "machine learning"},
#     "date_constraints": {"after_date": "2024-01-01"}
# }

ImageSchemaAnalyzer (plugins/document_analyzer.py)

Analyzes image search schemas with color, size, and composition detection.

Supported Patterns:

Pattern Constraint Type
Colors (red, blue) color_constraints
portrait, landscape, square composition.orientation
close-up, wide, aerial composition.shot_type
indoor, outdoor scene.type

VisualEvaluator Plugin (plugins/visual_evaluator.py)

VisualEvaluatorPlugin provides static factory methods that build Inspect AI scorer functions; get_visual_scorers(config) assembles the list based on config flags (not a model name):

from cogniverse_evaluation.plugins.visual_evaluator import get_visual_scorers

scorers = get_visual_scorers({
    "enable_llm_evaluators": True,
    "enable_quality_evaluators": True,
    "evaluator_name": "visual_judge",
})
# Returns: [VisualEvaluatorPlugin.create_visual_judge_scorer("visual_judge"),
#           VisualEvaluatorPlugin.create_quality_scorer()]

create_visual_judge_scorer(evaluator_name="visual_judge") scores each sample through ConfigurableVisualJudge; create_quality_scorer() scores each sample through the synchronous evaluators from evaluators.sync_reference_free.create_sync_evaluators().

Judge failures from raised vision calls are excluded from the mean and reported under metadata["failed_evaluations"]; genuine zeros (no results, no video id) still score 0.0. When every attempted judge call fails the scorer raises instead of reporting a score, so a judge outage cannot masquerade as a uniform quality collapse.

ConfigurableVisualJudge:

Resolves each result's source_url through :class:MediaLocator, extracts frames via cv2, and asks the configured LLM (provider, model, endpoint all sourced from the evaluator config — never constructor defaults) whether they match the query. Each vision call is bounded by the evaluator config's request_timeout_s (default 120) so an unresponsive endpoint fails the evaluation instead of hanging the run. Backend failures raise RuntimeError with the provider, model, and endpoint in the message; the judge never returns a partial result for an outage.

from cogniverse_core.common.media import MediaConfig, MediaLocator
from cogniverse_core.common.tenant_utils import SYSTEM_TENANT_ID
from cogniverse_evaluation.evaluators.configurable_visual_judge import (
    ConfigurableVisualJudge,
)

locator = MediaLocator(tenant_id=SYSTEM_TENANT_ID, config=MediaConfig())
evaluator = ConfigurableVisualJudge(locator=locator, evaluator_name="visual_judge")
result = evaluator.evaluate(
    input={"query": "robots playing soccer"},
    output={"results": [{"video_id": "v1", "source_url": "s3://corpus/v1.mp4"}]},
)
# Returns: score (0-1), label (excellent_match/good_match/partial_match/poor_match), explanation

Testing

Key Test Files

Unit Tests (tests/evaluation/unit/):

  • test_experiment_tracker.py — experiment configuration, dataset management, result formatting, main(args)'s required-args signature
  • test_routing_evaluator.py — routing outcome classification, metric calculation, parse_confidence label/percent coercion
  • test_agent_evaluators.py — per-agent-type registry (evaluators/agent_evaluators.py): every agent type has an entry, judge prompts are byte-identical, routing confidence calibration reads the canonical output.value JSON (with the legacy attribute still honored)
  • test_span_evaluator.py — span evaluation dispatch, evaluator resolution
  • test_quality_monitor.py — threshold checks, verdict decisions, baseline read/write fault contracts (_read_baseline_metric/_get_agent_baseline raise on outage, None only when genuinely absent), _store_trigger_dataset/submit_optimization dataset-name wiring
  • test_online_evaluator_confidence.py — confidence calibration scoring
  • test_data_managers.py — DatasetManager / TraceManager
  • test_storage.py — TelemetryStorage connection/health-check handling, get_traces_for_evaluation raising ConnectionError when disconnected
  • test_ground_truth.py — ground truth strategy dispatch, per-schema discovery caching, BackendGroundTruthStrategy's tighter top_k
  • test_golden_dataset_from_traces.py — scripts/create_golden_dataset_from_traces.py's GoldenDatasetGenerator reads the flattened Phoenix span frame (attributes.input.value/attributes.output.value), not bare input/output, and mines expected_videos as source title keys
  • test_metrics.py — MRR/nDCG/precision/recall/F1/MAP
  • test_evaluators.py — Evaluator base class, reference-free evaluators
  • test_inspect_scorers.py, test_solvers.py, test_task.py, test_reranking.py — Inspect AI integration
  • test_plugin_analyzers.py, test_visual_plugin.py, test_schema_agnostic.py — schema analyzer / visual plugin registration across document, image, and video schemas
  • test_visual_judge_frame_cleanup.py — ConfigurableVisualJudge.evaluate unlinks its extracted temp frame files on both success and failure
  • test_media_helpers.py — _media_helpers source/frame resolution
  • test_multi_turn_llm_judge.py — session-level LLM judging
  • test_cli.py, test_cli_simple.py — CLI command behavior
  • test_root_cause_analysis.py — analysis/root_cause_analysis.py
  • test_query_window_tz.py, test_traces_filter_escape.py — timezone/query-window and YQL-escaping edge cases
  • test_evaluation_provider_session.py — PhoenixEvaluationProvider.log_session_evaluation passes project through to add_annotation, and log_experiment_event emits an OpenTelemetry span named after the event with the event data as span attributes

tests/evaluation/fakes.py provides InMemoryDatasetStore/FailingDatasetStore — real DatasetStore subclasses (not MagicMocks) backed by an in-process dict, shared by test_quality_monitor.py and others to exercise the not-found/outage contract through the same method signatures production code calls, without a live Phoenix.

Integration Tests (tests/evaluation/integration/):

  • test_end_to_end.py — full evaluation pipeline against a real backend
  • test_incremental_span_eval.py — SpanEvaluator incremental skip-set behavior
  • test_golden_baseline_capture.py — golden-set baseline MRR capture for QualityMonitor
  • test_golden_source_title_matching.py — a content-hash tenant and a filename-id tenant served by the real /search route over real Vespa score the same golden set identically through QualityMonitor, the retrieval solver's scorers and GoldenDatasetEvaluator; an untitled result fails its queries without storing a baseline
  • test_xgboost_quality_monitor.py — quality-monitor training-decision modeling
  • test_provider_resolution.py — EvaluationRegistry provider discovery/resolution
  • test_schema_driven_pipeline.py — schema-aware ground truth extraction
  • test_storage_integration.py — TelemetryStorage against a real telemetry backend
  • test_task_real.py — evaluation_task against a real Phoenix dataset
  • test_visual_judge_e2e.py — ConfigurableVisualJudge end-to-end
  • test_dataset_manager_roundtrip.py — DatasetManager create/get/update/delete/export against a real Phoenix dataset store, concurrent creates, and the not-found/outage fault contract
  • test_provider_experiment_roundtrip.py — PhoenixEvaluationProvider.create_experiment/log_evaluation round-trip through a real Phoenix dataset store: the registry dataset and evaluation rows are readable back with exact values, and logging into a never-created experiment raises

Test Scenarios:

  1. Experiment Tracking:

    def test_experiment_configurations():
        """Verify experiment configuration retrieval."""
        tracker = ExperimentTracker(tenant_id="test_org:test")
        configs = tracker.get_experiment_configurations(
            profiles=["test_profile"]
        )
        assert len(configs) > 0
        assert all("strategies" in c for c in configs)
    

  2. Routing Evaluation:

    def test_routing_metrics_calculation():
        """Verify routing metrics calculation."""
        from cogniverse_foundation.telemetry.registry import TelemetryRegistry
    
        # Get telemetry provider
        tenant_id = "your_org:production"
        provider = TelemetryRegistry.get(name="phoenix", tenant_id=tenant_id)
        project_name = f"cogniverse-{tenant_id}"
    
        evaluator = RoutingEvaluator(provider=provider, project_name=project_name)
    
        # Mock spans with known outcomes
        spans = create_mock_routing_spans(
            success_count=80,
            failure_count=20
        )
    
        metrics = evaluator.calculate_metrics(spans)
    
        assert metrics.routing_accuracy == 0.80
        assert metrics.total_decisions == 100
    

  3. Analytics:

    def test_outlier_detection():
        """Verify outlier detection logic."""
        analytics = PhoenixAnalytics()
    
        # Create data with known outliers
        data = np.array([100, 110, 105, 95, 1000, 102])  # 1000 is outlier
        outliers = analytics._detect_outliers(data, method="iqr")
    
        assert 1000 in outliers
        assert len(outliers) == 1
    


Test Coverage:

  • Experiment configuration: ✅

  • Dataset management: ✅

  • Routing evaluation: ✅

  • Analytics calculations: ✅

  • Visualization generation: ✅

  • Phoenix integration: ✅


Summary

The Evaluation Module provides comprehensive experiment tracking and performance analysis with:

Core Features:

  • ✅ Inspect AI-based experiment framework

  • ✅ Routing-specific evaluation (separate from search)

  • ✅ Phoenix analytics with visualizations

  • ✅ Retrospective span evaluation (SpanEvaluator, batch/incremental)

  • ✅ Real-time span evaluation (OnlineEvaluator, sampled)

  • ✅ Continuous cross-agent quality monitoring with optimization triggers (QualityMonitor)

  • ✅ Provider-agnostic evaluation backend abstraction (EvaluationProvider/AnalyticsProvider)

  • ✅ Multi-evaluator support (quality, LLM, golden, reference-free)

  • ✅ CLI for experiment/batch/live evaluation and dataset management

Production Strengths:

  • Experiment reproducibility with versioned datasets

  • Routing confidence calibration monitoring

  • Automated performance alerting

  • Statistical analysis with outlier detection

  • Interactive Plotly visualizations

  • Incremental span evaluation avoids re-annotating already-scored spans

Integration Points:

  • Phoenix for trace storage and visualization (via PhoenixEvaluationProvider)

  • Inspect AI for evaluation execution

  • Quality evaluators for automated assessment

  • Dataset management for golden datasets

  • Argo workflows for triggered re-optimization (QualityMonitor)


For detailed examples and production configurations, see:

  • Architecture Overview: docs/architecture/overview.md

  • Routing Module: docs/modules/routing.md

  • Telemetry Module: docs/modules/telemetry.md

Source Files:

  • ExperimentTracker: libs/evaluation/cogniverse_evaluation/core/experiment_tracker.py

  • RoutingEvaluator: libs/evaluation/cogniverse_evaluation/evaluators/routing_evaluator.py

  • PhoenixAnalytics: libs/telemetry-phoenix/cogniverse_telemetry_phoenix/evaluation/analytics.py

  • SpanEvaluator: libs/evaluation/cogniverse_evaluation/span_evaluator.py

  • OnlineEvaluator: libs/evaluation/cogniverse_evaluation/online_evaluator.py

  • QualityMonitor: libs/evaluation/cogniverse_evaluation/quality_monitor.py

  • Evaluation Provider System: libs/evaluation/cogniverse_evaluation/providers/base.py, providers/registry.py

  • Data Layer: libs/evaluation/cogniverse_evaluation/data/{datasets,storage,traces}.py

  • Ground Truth: libs/evaluation/cogniverse_evaluation/core/ground_truth.py, core/schema_analyzer.py

  • Metrics: libs/evaluation/cogniverse_evaluation/metrics/custom.py

  • Evaluator Base: libs/evaluation/cogniverse_evaluation/evaluators/base.py, evaluators/sync_reference_free.py

  • CLI: libs/evaluation/cogniverse_evaluation/cli.py

  • RootCauseAnalyzer: libs/evaluation/cogniverse_evaluation/analysis/root_cause_analysis.py