Skip to content

Testing Guide

Comprehensive guide to testing practices in Cogniverse.


Table of Contents

  1. Overview
  2. Test Organization
  3. Running Tests
  4. Writing Tests
  5. Fixtures and Mocking
  6. Integration Tests
  7. Test Coverage
  8. CI/CD Testing
  9. Troubleshooting
  10. Related Documentation

Overview

Testing Philosophy

  1. Fix implementation to satisfy tests - Never weaken tests
  2. 100% pass rate required before any commit
  3. No shortcuts - No mocking failures, no disabling tests
  4. Test behavior, not implementation - Tests should survive refactoring

Test Stack

  • pytest: Test framework
  • pytest-asyncio: Async test support
  • pytest-cov: Coverage reporting
  • pytest-mock: Mocking support

uv sync installs everything the tests import, including the reference libraries some integration tests load directly: pylate (dev dependency group), xgboost, boto3 and sentence-transformers. Tests import them plainly; a missing one is an import error, not a skip.


Test Organization

Directory Structure

tests/
├── agents/
│   ├── unit/
│   │   ├── test_agent_startup_no_env_vars.py
│   │   └── ...
│   ├── integration/
│   │   └── ...
│   └── e2e/                         # Standalone-agent e2e config (test_config.py)
├── core/
│   ├── unit/
│   └── integration/
├── runtime/
│   ├── unit/
│   └── integration/
├── foundation/
│   ├── unit/
│   └── integration/
├── messaging/
│   ├── unit/
│   └── integration/
├── cli/
│   └── unit/
├── charts/                          # Helm chart unit tests (test_*_chart.py)
├── system/
│   ├── test_ensemble_comprehensive.py
│   ├── test_ensemble_search_e2e.py
│   ├── vespa_test_manager.py
│   ├── minio_test_manager.py        # MinIOTestManager
│   └── conftest.py
├── ingestion/
│   ├── unit/
│   ├── integration/
│   ├── fixtures/
│   └── pytest.ini                   # Ingestion-scoped markers/addopts
├── evaluation/
│   ├── unit/
│   ├── integration/
│   ├── fixtures/
│   └── conftest.py
├── routing/
│   ├── unit/
│   ├── integration/
│   └── pytest.ini                   # Routing-scoped markers/addopts
├── telemetry/
│   ├── unit/
│   └── integration/
├── finetuning/
│   ├── test_*.py                    # Unit tests live at this level (no unit/ subdir)
│   ├── integration/
│   └── conftest.py
├── memory/
│   ├── unit/
│   ├── integration/
│   └── conftest.py
├── backends/
│   ├── unit/
│   └── integration/
├── events/
│   ├── unit/
│   └── integration/
├── admin/
│   ├── test_profile_api.py          # Integration tests (run against shared_vespa)
│   ├── test_profile_concurrent.py
│   ├── test_profile_multi_tenant.py
│   ├── test_tenant_manager.py
│   ├── unit/
│   └── conftest.py
├── common/
│   ├── unit/
│   └── integration/
├── ui/                              # Currently empty
├── synthetic/
│   ├── unit/
│   └── integration/
├── e2e/                             # Cross-package e2e (A2A gateway, canary, CLI, ...)
│   └── deployment/
├── fixtures/                        # Shared model-endpoint fixtures (llm.py, inference.py, sidecars.py)
├── utils/                           # Shared test helpers (vespa_test_helpers, tenant_helpers,
│                                     # docker_utils, vllm_sidecar, markers, memory_store, ...)
└── conftest.py                      # Shared fixtures (shared_vespa, phoenix_container, etc.)

Naming Conventions

  • Files: test_<module_name>.py
  • Classes: Test<ClassName>
  • Functions: test_<behavior_description>
# test_search_agent.py

class TestSearchAgent:
    """Tests for SearchAgent."""

    def test_search_by_text_returns_results(self):
        """Test that search_by_text returns results for valid query."""
        ...

    def test_search_with_invalid_query(self):
        """Test search with invalid query parameters."""
        ...

    @pytest.mark.asyncio
    async def test_handles_backend_timeout(self):
        """Test graceful handling of backend timeout."""
        ...

Running Tests

Basic Commands

# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v

# Run specific package tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/agents/ -v

# Run specific test file
uv run pytest tests/agents/unit/test_search_agent.py -v

# Run specific test
uv run pytest tests/agents/unit/test_search_agent.py::TestSearchAgent::test_search_by_text -v

E2E helpers outside the e2e conftest

Importing tests/e2e/conftest.py publishes the e2e cluster's Vespa endpoint as TEST_BACKEND_URL/TEST_BACKEND_PORT for the rest of the session, so only modules under tests/e2e import it. Unit and integration tests import e2e helpers from the plain modules beside it: cluster.py (cluster name, kube context, host ports, seeded tenant), tenants.py, sample_corpus.py, inference.py, report.py and the others in tests/e2e. tests/common/unit/test_e2e_conftest_import_guard.py fails any module outside tests/e2e that reaches the conftest, directly or through another module.

Shared e2e cluster lifecycle

Tests that need the shared cogniverse-e2e cluster create it when absent and reuse it across pytest processes and working sessions. They intentionally leave it running between commands because a complete verification campaign can span multiple packages and several days.

If the cluster exists but is stopped, the fixture starts it through the Cogniverse cluster lifecycle and inspects its health and deployed-content fingerprint again before reuse. A healthy, current cluster is reused as-is. A stale or unhealthy cluster fails the test session without deleting or replacing shared state; repair it, or perform the explicit reset below.

k3d fixes the loadbalancer's published ports when the cluster is created, so a cluster created before a host-port mapping was added is unhealthy until it publishes it. The failure names the command that adds the missing mappings, for example k3d cluster edit cogniverse-e2e --port-add 33912:29012@loadbalancer; it recreates only the loadbalancer.

Do not stop the cluster after an individual pytest command. Once every integration and end-to-end check in the campaign has finished, stop it explicitly to release its CPU, RAM, and accelerator allocations:

cogniverse stop --name cogniverse-e2e

The next test requiring the cluster starts it again. Deleting the cluster is a separate explicit reset for stale or damaged state:

k3d cluster delete cogniverse-e2e

E2E_FRESH=1 does not authorize replacement: the cluster must already be absent. A fresh session deletes the cluster at teardown only when that same session created it; ordinary sessions always leave the shared cluster warm.

The fixture also provisions the coding sandbox in host mode: it starts (or reuses) the host OpenShell gateway, syncs its mTLS certs and metadata into the cluster before Helm installs the runtime, and deploys with runtime.sandbox.enabled=true, runtime.sandbox.gatewayEndpoint set to the gateway's own port and runtime.sandbox.hostGatewayIP set to the k3d network gateway (so host.docker.internal resolves in the runtime pod). Those values are part of the deploy identity. The gateway is the e2e stack's own, by name: openshell on port 28080 (OPENSHELL_GATEWAY_HOST_PORT). The deploy and the cert sync fail, naming the active gateway, when that gateway is unregistered, on another port, or not the host's active gateway, so no other gateway is ever deployed into the cluster. Integration tests that start their own gateway (OpenShellTestGateway in tests/agents/integration/conftest.py) register it under a private XDG_CONFIG_HOME, remove only their own container, and assert that the host's ~/.config/openshell stays byte-identical. On reuse the sync runs again and a changed secret rolls the runtime deployment, because the pod mounts the files with subPath.

There is no idle reaper: only the person or automation that knows the whole campaign is complete can safely decide when to stop the shared cluster.

The shared fixture seeds synthetic-generation checks with two real modalities: the tracked tests/system/resources/videos/v_-6dz6tBH77I.mp4 bytes and a JPEG decoded from that video's first frame. It deploys the canonical video and image profiles, then verifies the exact SHA-256 content ID, tenant-scoped MinIO URI, profile, positive documents_fed/chunk counts, and the matching documents persisted in Vespa. An unrelated approximate-nearest-neighbor hit never counts as an existing fixture.

Synthetic generation requests use the singular strategy field. The endpoint checks require exact nonzero samples from those persisted fixtures; profile and workflow examples must remain grounded in their content IDs, routing must use registered canonical agent IDs, and cross-modal examples must combine the real video and image modalities. Run that boundary directly with full output:

uv run pytest tests/e2e/test_api_e2e.py::TestSyntheticDataAPI \
  --tb=long -v > /tmp/cogniverse_synthetic_e2e.log 2>&1

Leave the cluster warm when more campaign checks remain. Use the explicit cogniverse stop --name cogniverse-e2e command above only after the complete integration and end-to-end run has finished.

Test Markers

Static markers are registered in the root pytest.ini (--strict-markers rejects any unregistered marker) plus a per-package override in tests/ingestion/pytest.ini. Exact inference markers are registered by tests/fixtures/inference.py. The commonly used ones:

  • unit, integration, e2e, system — test tier. unit and integration are also applied automatically from the test's location (tests/<pkg>/unit/ or tests/<pkg>/integration/) by tests/fixtures/markers.py::apply_location_markers, wired into pytest_collection_modifyitems in tests/conftest.py and the nested-root conftests (tests/ingestion/, tests/routing/) — so a file that forgets the marker cannot silently fall out of its directory's CI -m selection. A file marked local_only opts out of the location marker (a declared, visible CI exclusion). tests/runtime/unit/test_marker_coverage.py verifies every file against the actual workflow selections.
  • ci_fast — small, essential subset run on every CI push
  • ci_safe — safe to run in CI (no local-only dependencies)
  • local_only — should only run locally, not in CI/CD
  • slow, benchmark — long-running or performance tests
  • e2e_heavy — full CronWorkflow-driven e2e (DSPy optimization, synthetic data generation, distillation); opt-in only, 5-10+ minutes
  • browser — Playwright-driven browser e2e tests
  • telemetry — requires the telemetry/Phoenix provider
  • no_shared_vespa — runtime integration module owns another real boundary; its module fixture leaves the dead backend sentinel in place and does not start the shared Vespa container
  • requires_vespa, requires_docker, requires_gpu, requires_ollama, requires_gliner, requires_models, requires_cv2, requires_ffmpeg — infrastructure/model dependencies
  • requires_inference("vllm_colpali") — exact ColPali/ColQwen HTTP embedding service; requires_inference("video_embed") — exact X-CLIP service. Collection also requests every service a shipped profile using the named embedding service resolves at pipeline init, derived from configs/config.json (vllm_asr for vllm_colpali, video_embed and colbert_pylate). Cluster discovery reads --revision from rendered workload args: a pinned workload is tagged identity_evidence=DEPLOYMENT, and an unpinned one stays ENDPOINT and must report the exact revision from /v1/models.
  • requires_modal_inference("vllm_llm_student") — explicitly opt an exact service into paid Modal provisioning; ordinary requires_inference tests resolve the cluster's endpoint (see "Model endpoints" below).
  • requires_whisper — Whisper model dependency outside the exact inference fixture
  • requires_teacher_model — scales up the vllm-llm-teacher pod (off by default)
  • requires_optimizer_data — exercises non-router optimizers end-to-end against the live cluster (off by default)
  • phoenix, inspect, ragas — evaluation-framework dependencies
  • timeout — tests with custom timeout values

local_only is an explicit run-on-demand policy, not a CI pass. In particular, the real image-search encoder test and real ColPali image-encoding test (both need the cluster's ColPali service) are intentionally excluded from normal CI and also carry slow plus their precise resource markers. Run them locally when changing their boundaries:

uv run pytest -m local_only \
  tests/agents/integration/test_image_search_real_encoder_e2e.py \
  tests/core/integration/test_colpali_encode_image_real.py
# Run only the fast CI subset
JAX_PLATFORM_NAME=cpu uv run pytest tests/agents/integration/ -m ci_fast -v

# Run integration tests but skip anything needing Ollama
uv run pytest tests/agents/integration/ -m "integration and not requires_ollama" -v

With Timeout

# 30-minute timeout for long tests (using system timeout command)
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/ -v

# For per-test timeout, install pytest-timeout first:
# uv pip install pytest-timeout
# Then use:
# uv run pytest tests/ -v --timeout=300

Parallel Execution

Note: pytest-xdist is not currently installed. For parallel execution, install it first:

# Install pytest-xdist
uv pip install pytest-xdist

# Run with 4 workers
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v -n 4

# Auto-detect workers
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v -n auto

With Output

# Show stdout/stderr
uv run pytest tests/ -v -s

# Log to file
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v 2>&1 | tee /tmp/test_output.log

Finding Affected Tests

Before committing, find all tests affected by your changes:

# Find tests for a module
grep -r "def test_" --include="*.py" tests/ | grep -i "search_agent"

# Find all tests in a directory
find tests/agents -name "test_*.py" -exec grep -l "def test_" {} \;

Writing Tests

Basic Test Structure

import pytest
from pathlib import Path
from cogniverse_agents.search_agent import SearchAgent, SearchInput, SearchOutput, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader

class TestSearchAgent:
    """Tests for SearchAgent."""

    @pytest.fixture
    def config_manager(self):
        """Create config manager for agent."""
        return create_default_config_manager()

    @pytest.fixture
    def schema_loader(self):
        """Create schema loader for agent."""
        # FilesystemSchemaLoader requires path to schema directory
        # Use relative path from project root (tests run from project root)
        from pathlib import Path
        schema_path = Path("configs/schemas")
        return FilesystemSchemaLoader(base_path=schema_path)

    @pytest.fixture
    def agent(self, config_manager, schema_loader):
        """Create test agent instance."""
        # Both config_manager and schema_loader are required for SearchAgent
        deps = SearchAgentDeps(
            backend_url="http://localhost",
            backend_port=8080
        )
        return SearchAgent(
            deps=deps,
            schema_loader=schema_loader,
            config_manager=config_manager,
            port=8002  # A2A server port (optional, defaults to 8002)
        )

    @pytest.fixture
    def valid_input(self):
        """Create valid test input."""
        return SearchInput(
            query="machine learning tutorial",
            tenant_id="test-tenant",
            modality="video",
            top_k=10
        )

    def test_search_by_text_returns_results(self, agent, valid_input):
        """Test that search_by_text returns results."""
        results = agent.search_by_text(
            query=valid_input.query,
            tenant_id="test-tenant",
            modality=valid_input.modality,
            top_k=valid_input.top_k
        )

        assert isinstance(results, list)
        # Results is a list of dicts, each with id, score, video_id, frame_id, etc.
        if len(results) > 0:
            assert "video_id" in results[0] or "id" in results[0]

Async Tests

Use @pytest.mark.asyncio for async tests. Note that SearchAgent.search_by_text() is synchronous:

import pytest
from concurrent.futures import ThreadPoolExecutor
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.manager import ConfigManager
from cogniverse_sdk.interfaces.config_store import ConfigStore
from unittest.mock import MagicMock

class TestSearchAgent:

    @pytest.fixture
    def mock_config_store(self):
        """Create mock ConfigStore."""
        store = MagicMock(spec=ConfigStore)
        store.get_config.return_value = None
        return store

    @pytest.fixture
    def config_manager(self, mock_config_store):
        """Create ConfigManager with mock store."""
        return ConfigManager(store=mock_config_store)

    @pytest.fixture
    def schema_loader(self):
        """Create schema loader for agent."""
        from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
        from pathlib import Path
        return FilesystemSchemaLoader(Path("configs/schemas"))

    @pytest.fixture
    def agent(self, config_manager, schema_loader):
        """Create SearchAgent with config manager and schema loader."""
        return SearchAgent(
            deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
            config_manager=config_manager,
            schema_loader=schema_loader,
        )

    def test_search(self, agent):
        """Test search operation (synchronous)."""
        result = agent.search_by_text(query="test query", tenant_id="test", top_k=5)
        assert result is not None

    def test_concurrent_searches(self, agent):
        """Test concurrent search operations using threads."""
        with ThreadPoolExecutor(max_workers=3) as executor:
            futures = [
                executor.submit(agent.search_by_text, query="query1", tenant_id="test", top_k=5),
                executor.submit(agent.search_by_text, query="query2", tenant_id="test", top_k=5),
                executor.submit(agent.search_by_text, query="query3", tenant_id="test", top_k=5),
            ]
            results = [f.result() for f in futures]

        assert len(results) == 3

Parametrized Tests

import pytest

class TestValidation:

    @pytest.mark.parametrize("query,expected_validation", [
        ("", "empty"),
        ("   ", "empty"),
        ("a" * 10001, "too long"),
    ])
    def test_invalid_queries(self, agent, query, expected_validation):
        """Test validation of invalid queries."""
        result = agent.search_by_text(query=query, tenant_id="test-tenant", modality="video", top_k=10)
        # SearchAgent returns a list (may be empty for invalid queries)
        assert isinstance(result, list)

Testing Exceptions

import pytest
from cogniverse_foundation.config.manager import ConfigManager

class TestConfigManager:

    def test_missing_store_raises_error(self):
        """ConfigManager requires a store — passing None raises ValueError.

        Note: get_system_config() itself does NOT raise when no system
        config is stored — it logs a warning and returns SystemConfig()
        defaults. Test the actual required-argument validation instead.
        """
        with pytest.raises(ValueError):
            ConfigManager(store=None)

Fixtures and Mocking

Shared Fixtures (conftest.py)

# tests/conftest.py
import pytest

@pytest.fixture
def config_manager(backend_config_env):
    """Create ConfigManager with backend store for testing.
    Requires backend_config_env fixture to set environment variables.
    """
    from cogniverse_foundation.config.utils import create_default_config_manager
    return create_default_config_manager()

@pytest.fixture
def config_manager_memory():
    """Create ConfigManager with in-memory store for unit testing.
    Does not require any backend infrastructure (Vespa, etc.).
    """
    from cogniverse_foundation.config.manager import ConfigManager
    from tests.utils.memory_store import InMemoryConfigStore
    store = InMemoryConfigStore()
    store.initialize()
    return ConfigManager(store=store)

@pytest.fixture
def workflow_store(telemetry_manager_with_phoenix):
    """Resolve a workflow store via the registry — same path production uses."""
    from cogniverse_core.registries import WorkflowStoreRegistry

    provider = telemetry_manager_with_phoenix.get_provider(
        tenant_id="workflow-store-test"
    )
    # Evict any instance cached under a stale provider so each test resolves clean.
    WorkflowStoreRegistry.clear_cache()
    store = WorkflowStoreRegistry.get(
        name="telemetry",
        config={"telemetry_provider": provider},
    )
    store.initialize()
    return store

Mocking External Services

from unittest.mock import AsyncMock, MagicMock, patch

class TestSearchAgentWithMocks:

    @pytest.fixture
    def mock_backend(self):
        """Create mock backend."""
        backend = MagicMock()
        backend.search = AsyncMock(return_value=[
            {"id": "1", "score": 0.9, "title": "Result 1"},
            {"id": "2", "score": 0.8, "title": "Result 2"},
        ])
        return backend

    def test_search_calls_backend(self, mock_backend, config_manager, schema_loader):
        """Test that search calls backend with correct params."""
        deps = SearchAgentDeps(
            backend_url="http://localhost",
            backend_port=8080
        )
        agent = SearchAgent(
            deps=deps,
            schema_loader=schema_loader,
            config_manager=config_manager,
            port=8002
        )

        result = agent.search_by_text(query="test", tenant_id="test-tenant", modality="video", top_k=5)

        # Verify backend was called (if using mock backend in deps)
        # mock_backend.search.assert_called()

    def test_handles_backend_error(self, mock_backend, config_manager, schema_loader):
        """Test graceful handling of backend error."""
        mock_backend.search.side_effect = ConnectionError("Backend down")

        deps = SearchAgentDeps(
            backend_url="http://localhost",
            backend_port=8080
        )
        agent = SearchAgent(
            deps=deps,
            schema_loader=schema_loader,
            config_manager=config_manager,
            port=8002
        )

        result = agent.search_by_text(query="test", tenant_id="test-tenant", modality="video", top_k=10)

        # Result is a list (may be empty or contain error info depending on implementation)
        assert isinstance(result, list)

Patching

from unittest.mock import patch, MagicMock

class TestWithPatching:

    @patch("cogniverse_agents.search_agent.get_backend_registry")
    def test_with_patched_backend(self, mock_registry, config_manager, schema_loader):
        """Test with patched backend registry."""
        mock_backend = MagicMock()
        mock_registry.return_value.get_search_backend.return_value = mock_backend

        deps = SearchAgentDeps(
            backend_url="http://localhost",
            backend_port=8080
        )
        agent = SearchAgent(
            deps=deps,
            schema_loader=schema_loader,
            config_manager=config_manager
        )

        # Verify backend registry was accessed
        assert agent is not None

Integration Tests

With Real Services

import pytest

@pytest.mark.integration
class TestVespaIntegration:
    """Integration tests requiring Vespa."""

    @pytest.fixture
    def search_backend(self, config_manager, schema_loader):
        """Create real VespaSearchBackend."""
        import os
        from cogniverse_vespa.search_backend import VespaSearchBackend
        backend_url = os.getenv("BACKEND_URL", "http://localhost")
        backend_port = int(os.getenv("BACKEND_PORT", "8080"))
        return VespaSearchBackend(
            config={
                "url": backend_url,
                "port": backend_port,
                "profiles": {"video_colpali_smol500_mv_frame": {}},
                "default_profiles": {"video": "video_colpali_smol500_mv_frame"},
            },
            config_manager=config_manager,
            schema_loader=schema_loader,
        )

    def test_real_search(self, search_backend):
        """Test search against real Vespa."""
        # VespaSearchBackend.search() takes a query_dict and returns
        # List[SearchResult]. strategy is a rank-profile-name string.
        results = search_backend.search({
            "query": "test",
            "type": "video",
            "profile": "video_colpali_smol500_mv_frame",
            "strategy": "bm25_only",
            "top_k": 5,
            "tenant_id": "test:unit",
        })
        assert isinstance(results, list)

Skipping Without Services

import pytest
import os

requires_vespa = pytest.mark.skipif(
    os.getenv("BACKEND_URL") is None,
    reason="BACKEND_URL not set"
)

@requires_vespa
class TestVespaRequired:
    """Tests that require Vespa."""

    def test_vespa_operation(self):
        ...

LLM-driven integration tests (llm_endpoint fixture)

Tests that need a real LLM (visual judges, Inspect-AI eval tasks) consume the provider-agnostic llm_endpoint fixture defined in tests/evaluation/integration/conftest.py. The fixture resolves the endpoint from, in order:

  1. Env vars COGNIVERSE_TEST_LLM_PROVIDER_URI (required) and COGNIVERSE_TEST_LLM_BASE_URL (optional). When a base URL is set, it is propagated to the provider's own env var — only openai/vllm (→ OPENAI_BASE_URL) and anthropic (→ ANTHROPIC_BASE_URL) prefixes are mapped today; a matching *_API_KEY of "not-required" is also injected so Inspect AI's provider client doesn't refuse to initialize against a keyless local endpoint. TEST_LLM_MODEL / TEST_LLM_API_BASE (as set by the ensure_host_ollama fixture in tests/conftest.py) are also accepted as a same-priority alternative.
  2. JSON file at tests/evaluation/integration/resources/test_llm.json (gitignored; a test_llm.example.json is checked in).
  3. llm_config.primary of the session config ensure_host_ollama writes, which points at the Modal Gemma (see "Model endpoints" below).
  4. Otherwise, the fixture fails the test with a message explaining how to configure an endpoint.

Hermetic session configs publish only roles whose exact model endpoint was resolved and verified. A primary-only session points the teacher entry at a dead port instead of exposing an endpoint the fixture did not check. Tests marked requires_teacher_model resolve and verify the distinct teacher before adding that role. Teacher-dependent code must call resolve_teacher(), which fails explicitly when the role was not configured; it never falls back to the primary. Real optimizer integration tests that reach resolve_teacher() carry this marker so collection records the teacher role before the session config is materialized.

Model endpoints

No test starts a model on this host: no model container, no vLLM, Ollama or model-server process. Parallel test runs share the host's memory, and a local model copy holds gigabytes of it. Every model a test uses is served remotely:

  • Chat LLMs (the Gemma student, the Qwen teacher) are served on Modal. ensure_llm (tests/utils/hermetic_llm.py), which ensure_host_ollama, gemma_inference_endpoint and requires_inference("vllm_llm_student") use, reads the role's deployment through ModalInferenceLifecycle.status and accepts it only when its authenticated /v1/models names the exact model and revision. It does not depend on any cluster. An INFERENCE_SERVICE_URLS entry for vllm_llm_student or vllm_llm_teacher replaces the Modal lookup. A missing COGNIVERSE_INFERENCE_API_KEY, an undeployed app or a wrong model raises ModalLlmNotDeployedError or LlmEndpointMismatchError naming the deploy command; the test fails, never skips or falls back.
  • Every other model (ColPali/Tomoro, DenseOn, PyLate LateOn and the code encoder, GLiNER, ASR, CLAP, video-embed, face-embed) is served by the cogniverse-e2e cluster on GPU. remote_inference (tests/fixtures/inference.py) and the requires_inference marker resolve a service from an explicit INFERENCE_SERVICE_URLS entry, else the workload the cogniverse-e2e cluster (then the cogniverse dev cluster) publishes through its load balancer, and validate its model identity. With no endpoint it raises RemoteServiceUnavailable naming the service.

Containers that serve no model (Vespa, Redis, MinIO, Phoenix, the semantic router stack) are still started by the tests' own fixtures.

Before a run that needs models:

# Modal credentials and the inference key live in .env/ (loaded by tests/conftest.py)
ls .env/MODAL_TOKEN_ID.env .env/MODAL_TOKEN_SECRET.env .env/COGNIVERSE_INFERENCE_API_KEY.env

# The chat models must be deployed on Modal
uv run cogniverse inference modal status vllm_llm_student vllm_llm_teacher
uv run cogniverse inference modal deploy vllm_llm_student   # if status fails

# The other models come from the cogniverse-e2e cluster
kubectl --context k3d-cogniverse-e2e -n cogniverse get deploy,svc

When the deploy serves the chat models from Modal, the e2e session runs cogniverse inference modal warm vllm_llm_student before its first test and release at teardown: the query rewrite gives the student 2.8 s, which a runner starting from zero cannot meet. A failed warm fails the session. The deploy itself refuses to build when the host disk holding docker's data is at or above 75%, Vespa's feed-block limit; free space (docker builder prune -f) and rerun.

The cluster is deployed through the e2e path (tests/e2e/deployment), not a plain cogniverse up. tests/common/unit/test_no_local_model_servers.py fails any test code that runs a model image's container, launches a vLLM/Ollama server, or loads a model in-process (SentenceTransformer, PyLate ColBERT, from_pretrained on a model class, Whisper, faster-whisper, FaceAnalysis).

No test loads model weights into the pytest process either: the session plugin tests/fixtures/no_local_models.py replaces the weight-loading entry points of transformers, sentence-transformers (and so PyLate), GLiNER, the Hugging Face hub mixin, Whisper, faster-whisper, InsightFace and open_clip with a refusal (LocalModelLoadForbidden), however the call is reached. Product code that degrades when a model is missing keeps degrading; a test whose result needs the model fails on its assertions. Every refused load is listed, by test, in the session's refused in-process model loads summary. A test that needs a model's output uses the cluster service that serves it (remote_inference, gliner_url, served_semantic_embedder).

Parity tests compare the served model against reference outputs recorded once, on CPU, from the model's own library at the pinned revision: tests/fixtures/model_references/{lateon,denseon}.json carry the model, revision, library versions, device and inputs beside the outputs. Tests never run the recorder; re-record after a pinned revision changes:

uv run python scripts/record_model_references.py

Every pytest session prints a test sidecars section in its terminal summary, captured output or not: each model resolution (resolved-remote with its endpoint and provider, or refused with the reason) and the dead-owner containers the session reaped when it started.

Examples:

# OpenAI for inspect_eval-driven retrieval tests; OPENAI_API_KEY picked up
# automatically by the provider SDK.
COGNIVERSE_TEST_LLM_PROVIDER_URI="openai/gpt-4o-mini" \
uv run pytest tests/evaluation/integration/test_end_to_end.py

# vLLM behind an OpenAI-compatible proxy:
COGNIVERSE_TEST_LLM_PROVIDER_URI="vllm/Qwen/Qwen2.5-7B-Instruct" \
COGNIVERSE_TEST_LLM_BASE_URL="http://vllm.internal:8000/v1" \
uv run pytest tests/evaluation/integration/

# No env vars set: the fixture uses the Modal Gemma ensure_host_ollama resolves.
uv run pytest tests/evaluation/integration/test_visual_judge_e2e.py

Test classes never reference any specific provider, model, or container manager — they only consume llm_endpoint["provider_uri"] and llm_endpoint["base_url"]. API keys are resolved by the provider SDK from its own conventional env var (OPENAI_API_KEY, ANTHROPIC_API_KEY, ...) — the test infrastructure does not re-implement that lookup.

MediaLocator-driven integration tests (media_root_uri)

Tests that exercise the unified-MediaLocator rollout (ingestion, audio transcribe, visual judge) read media via MediaLocator instead of bare filesystem paths. The pipeline accepts a media_root_uri config (e.g. s3://corpus/ or file:///abs/path/) that the locator joins with each video's relative path to produce the canonical source_url written into Vespa. See tests/ingestion/integration/test_pipeline_minio_round_trip.py for the end-to-end pattern (MinIO docker fixture managed by MinIOTestManager); the schema-level field is documented in docs/modules/common.md under "MediaLocator".


Test Coverage

Running with Coverage

# Basic coverage
uv run pytest tests/ --cov=cogniverse_core --cov-report=term

# HTML report
uv run pytest tests/ --cov=cogniverse_core --cov-report=html

# Multiple packages
uv run pytest tests/ \
    --cov=cogniverse_core \
    --cov=cogniverse_agents \
    --cov=cogniverse_foundation \
    --cov-report=html

Coverage Configuration

Coverage is not configured in pyproject.toml or the root pytest.ini — it's specified per invocation via --cov=<path> flags, either on the command line or in a workflow's test step. The per-package tests/ingestion/pytest.ini explicitly notes this in its addopts comment ("no coverage here — handled by CI/Make targets").

Coverage Targets

Most CI workflows collect and report coverage (--cov-report=term-missing, --cov-report=xml) without enforcing a minimum. Only two currently gate the build on a coverage floor via --cov-fail-under:

Workflow Package Enforced Minimum
evaluation-tests.yml cogniverse-evaluation 50%
routing-tests.yml cogniverse-agents (routing) 15%

All other workflows (agents, core, finetuning, ingestion, synthetic, telemetry, vespa, runtime) report coverage but do not fail the build below any specific percentage.


CI/CD Testing

GitHub Actions Workflows

The project has 19 GitHub workflow files: 14 per-module test workflows plus chart-validation.yml, test-integrity.yml, docs.yml, publish-packages.yml, and two manual/release workflows not tied to a single module:

Workflow Module Tests Docker Services
agents-tests.yml cogniverse-agents unit + integration Vespa (ci_fast subset)
chart-validation.yml Helm chart (charts/cogniverse) lint + template + kubeconform None
cli-tests.yml cogniverse-cli unit + integration None
core-tests.yml cogniverse-core (incl. tests/core/*, tests/memory/* and the ci_fast files under tests/utils/; the rest of memory integration is local-tier — it needs the cluster's DenseOn service) unit + integration Vespa
evaluation-tests.yml cogniverse-evaluation unit + integration Phoenix
finetuning-tests.yml cogniverse-finetuning unit + integration Vespa
ingestion-tests.yml cogniverse-runtime (ingestion) unit + integration Vespa
messaging-tests.yml cogniverse-messaging unit + integration None
routing-tests.yml cogniverse-agents (routing) unit + integration Vespa
runtime-tests.yml cogniverse-runtime, cogniverse-foundation, cogniverse-cli, events, cogniverse-messaging (unit for all; integration for runtime + the small events/foundation/messaging suites; the web client's browser suites, test_web_*.py and test_ag_ui_threads.py, run in five parallel web-ops-integration-tests-* jobs and the Optimization framework suite in web-ops-optimization-framework-tests) unit + integration Vespa
synthetic-tests.yml cogniverse-synthetic unit + integration Phoenix
telemetry-tests.yml cogniverse-telemetry-phoenix unit + integration Phoenix
test-integrity.yml Whole-repo test guards (no paths filter) assertion strength + CI coverage None
vespa-tests.yml cogniverse-vespa unit + integration Vespa
docs.yml Documentation mkdocs build + deploy to GitHub Pages None
publish-packages.yml Package publishing (PyPI/TestPyPI) N/A None
mirror-third-party.yml Air-gapped support (workflow_dispatch only) Mirrors third-party images (Vespa, Phoenix, Ollama, …) to a target registry None
release-images.yml Release (v* tag push or workflow_dispatch) Builds/pushes every first-party image + Helm chart OCI artifact None

Workflow Structure

Each test workflow typically has these jobs:

  1. unit-tests - Unit tests; unit-dir jobs select -m unit (satisfied by every test via the location-derived marker) so no file can silently fall out of the selection
  2. integration-tests - Integration tests (with ci_fast subset for quick feedback)
  3. lint - Code linting with ruff/black
  4. test-cli or test-imports - Verify package imports work
  5. coverage-report - Combined coverage (if applicable)

Note: Workflows don't have separate "fast-integration-tests" jobs. Instead, integration-tests jobs use -m ci_fast to run essential tests quickly.

Guards on CI coverage itself

A test CI never runs reports as absence, which reads like success. Three guards close that:

  • tests/runtime/unit/test_marker_coverage.py — every file under a selected path carries markers its selection's -m expression keeps.
  • tests/common/unit/test_ci_coverage_guard.py — every ci_fast test that is not local_only is named by some commit-gating selection whose expression keeps it; every selection's own paths are inside its workflow's paths filter; and every test that walks a tree outside its own package runs in a workflow that fires on changes to that tree.
  • tests/common/unit/test_ci_fast_excludes_model_spawners.py — no ci_fast selection reaches a fixture that needs a model server.

All three read the selections from tests/fixtures/ci_workflows.py, which parses .github/workflows/*.yml; a tag-only workflow gates no commit and its selections do not count. test-integrity.yml runs the whole-tree guards with no paths filter, since any filter would skip them on the commits they exist to catch.

Assertion-strength guard and its waivers

tests/common/unit/test_assertion_strength_guard.py fails a change that leaves any tests/ file with fewer assertions than before (net of assertions moved verbatim into a file the change creates), or that adds a skip, an xfail or an unbounded assertion form. CI compares against the pull request's base or the push's previous commit (ASSERTION_GUARD_BASE); locally it defaults to HEAD~1:

ASSERTION_GUARD_BASE=$(git merge-base HEAD main) \
  uv run pytest tests/common/unit/test_assertion_strength_guard.py

The one accepted loss is of assertions that tested code the change deletes. Each is declared in tests/common/assertion_waivers.toml:

[[waiver]]
file = "tests/utils/test_vllm_sidecar.py"   # the test file that lost assertions
max_net_loss = 173                          # its largest accepted net loss
removed_symbols = ["tests/utils/vllm_sidecar.py:VllmSidecarFactory"]  # path:Name
reason = "Tests of the local vLLM sidecar launch, removed with it."

The guard checks every waiver against the compared range: each named top-level symbol must be gone at HEAD and must have existed in HEAD's history, the file's net loss must not exceed max_net_loss, and a waiver for a file that lost nothing fails as stale. A waiver applies only when at least one of its symbols existed at the base; one whose symbols were all gone before the range covered an earlier change and waives nothing.

CI Fast Integration Tests

Workflows run integration tests with the ci_fast marker on every push to provide quick feedback:

integration-tests:
  runs-on: ubuntu-latest
  timeout-minutes: 60
  needs: unit-tests

  steps:
  - uses: actions/checkout@v4

  - name: Set up Python 3.12
    uses: actions/setup-python@v4
    with:
      python-version: '3.12'

  - name: Install uv
    uses: astral-sh/setup-uv@v4
    with:
      version: "latest"

  - name: Install dependencies
    run: |
      uv sync --all-packages --all-extras
      uv pip install pytest-cov

  - name: Free up disk space for Vespa (if needed)
    run: |
      # Vespa requires <75% disk usage
      sudo rm -rf /usr/share/dotnet
      sudo rm -rf /usr/local/lib/android
      sudo rm -rf /opt/ghc
      sudo rm -rf /opt/hostedtoolcache/CodeQL
      sudo docker image prune -af

  - name: Run integration tests (CI Fast subset)
    run: |
      JAX_PLATFORM_NAME=cpu uv run python -m pytest \
        tests/module/integration/ \
        -m ci_fast \
        -v --tb=long \
        --cov=libs/module/cogniverse_module

Disk Cleanup for Vespa

Vespa requires disk usage below 75%. GitHub runners have ~14GB disk, so cleanup is required:

# Remove ~30GB of unused packages
sudo rm -rf /usr/share/dotnet           # .NET SDK (~6GB)
sudo rm -rf /usr/local/lib/android      # Android SDK (~10GB)
sudo rm -rf /opt/ghc                    # Haskell compiler (~2GB)
sudo rm -rf /opt/hostedtoolcache/CodeQL # CodeQL (~5GB)
sudo docker image prune -af             # Unused Docker images

Docker Service Management

Tests share a single session-scoped Vespa container and manage their own Phoenix containers via fixtures.

Vespa (shared_vespa):

All Vespa integration tests use the shared_vespa session-scoped fixture defined in tests/conftest.py. Each per-package conftest.py re-exports it and provides a compatibility shim under whatever fixture name its tests expect:

# e.g. tests/backends/integration/conftest.py
from tests.conftest import shared_vespa  # noqa: F401

@pytest.fixture(scope="module")
def vespa_instance(shared_vespa):
    """Shim: yields the dict shape tests expect, backed by shared_vespa."""
    yield {
        "http_port": shared_vespa["http_port"],
        "config_port": shared_vespa["config_port"],
        "base_url": shared_vespa["base_url"],
        "container_name": shared_vespa["container_name"],
    }
    # No teardown — shared_vespa owns the container lifecycle.

For tests that need to deploy their own schemas, use the helpers in tests/utils/vespa_test_helpers.py:

from tests.utils.vespa_test_helpers import deploy_tenant_schema, make_config_manager
from tests.utils.tenant_helpers import tenant_id_for_test

@pytest.fixture
def my_schema(shared_vespa, request):
    tenant_id = tenant_id_for_test(request)
    config_manager = make_config_manager(shared_vespa)
    deploy_tenant_schema(
        shared_vespa,
        tenant_id=tenant_id,
        base_schema_name="video_colpali_smol500_mv_frame",
        config_manager=config_manager,
    )
    yield tenant_id

Phoenix:

The phoenix_container fixture is defined in tests/conftest.py (module-scoped). It allocates non-default, per-process ports to avoid conflicts both with local Phoenix instances and with other concurrent pytest sweeps:

  • HTTP: 16006 + port_offset (instead of 6006)
  • gRPC: 14317 + port_offset (instead of 4317)

where port_offset = (os.getpid() % 1000) * 10, giving each process a distinct 10-port-spaced slot in a ~1000-process range.

Image: arizephoenix/phoenix:20.16.0 pinned by digest (never :latest).

Containers are named phoenix_test_pid<pid>_<timestamp> and tagged with the owning pid; on startup the fixture only kills leftover containers matching phoenix_test_pid<its own pid>_* from a prior crashed run of that same process — never another concurrent session's container.

Pre-commit Checklist

Before every commit:

# 1. Find affected tests
grep -r "def test_" --include="*.py" tests/ | grep -i "<module>"

# 2. Run tests
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/ -v 2>&1 | tee /tmp/test_output.log

# 3. Verify 100% pass
grep -E "passed|failed" /tmp/test_output.log

# 4. Lint
uv run make lint-all

Troubleshooting

JAX Platform Errors

# Always set JAX platform for CPU
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v

# Or in conftest.py
import os
os.environ["JAX_PLATFORM_NAME"] = "cpu"

Async Test Not Running

# Ensure pytest-asyncio is installed and configured

# In pytest.ini (already configured):
# asyncio_mode = auto
# asyncio_default_fixture_loop_scope = function
# asyncio_default_test_loop_scope = function  # pinned so a leaked coroutine
#   in one async test can't fail every subsequent async test in the same
#   module/session scope (pytest-asyncio 1.3+ default)

# Use decorator for async tests
@pytest.mark.asyncio
async def test_async_function():
    ...

Import Errors

# Ensure packages are installed
uv sync

# Check import path
uv run python -c "import cogniverse_core; print(cogniverse_core.__file__)"

# Run from project root
cd /path/to/cogniverse
uv run pytest tests/ -v

Flaky Tests

For retries or timeout decorators, install the required packages first:

# Install pytest-timeout for timeout decorator
uv pip install pytest-timeout

# Install pytest-rerunfailures for flaky test retries
uv pip install pytest-rerunfailures

Then use:

# Use retries for network-dependent tests (requires pytest-rerunfailures)
@pytest.mark.flaky(reruns=3)
def test_network_operation():
    ...

# Or add explicit timeout (requires pytest-timeout)
@pytest.mark.timeout(30)
def test_slow_operation():
    ...

Test Isolation

# Each test should be independent
# Use fixtures with appropriate scope

@pytest.fixture  # Default: function scope - fresh for each test
def config():
    return Config()

@pytest.fixture(scope="class")  # Shared within class
def expensive_resource():
    return create_expensive_resource()

@pytest.fixture(scope="session")  # Shared for entire session
def database():
    return setup_database()

Several tests/ subpackages carry their own more detailed testing guide; this file covers cross-cutting practices, these cover package-specific scenarios and fixtures:

  • tests/README.md / tests/CONSOLIDATED_TEST_README.md — full test suite overview
  • tests/agents/README.md — multi-agent routing, A2A protocol, DSPy/GEPA testing
  • tests/agents/e2e/README.md — real-LLM end-to-end tests (Ollama, Vespa, Phoenix)
  • tests/ingestion/README.md — ingestion pipeline test suite
  • tests/routing/README.md — query routing and classification test suite

Summary

  1. Organize tests by package (agents, system, ingestion, evaluation, etc.)
  2. Use fixtures for common setup
  3. Run with JAX_PLATFORM_NAME=cpu to avoid GPU issues
  4. 100% pass rate required before commit
  5. Find affected tests with grep before committing
  6. Use mocks for external services in unit tests
  7. Mark integration tests that require real services