Testing Guide¶
Comprehensive guide to testing practices in Cogniverse.
Table of Contents¶
- Overview
- Test Organization
- Running Tests
- Writing Tests
- Fixtures and Mocking
- Integration Tests
- Test Coverage
- CI/CD Testing
- Troubleshooting
- Related Documentation
Overview¶
Testing Philosophy¶
- Fix implementation to satisfy tests - Never weaken tests
- 100% pass rate required before any commit
- No shortcuts - No mocking failures, no disabling tests
- Test behavior, not implementation - Tests should survive refactoring
Test Stack¶
- pytest: Test framework
- pytest-asyncio: Async test support
- pytest-cov: Coverage reporting
- pytest-mock: Mocking support
uv sync installs everything the tests import, including the reference libraries some integration tests load directly: pylate (dev dependency group), xgboost, boto3 and sentence-transformers. Tests import them plainly; a missing one is an import error, not a skip.
Test Organization¶
Directory Structure¶
tests/
├── agents/
│ ├── unit/
│ │ ├── test_agent_startup_no_env_vars.py
│ │ └── ...
│ ├── integration/
│ │ └── ...
│ └── e2e/ # Standalone-agent e2e config (test_config.py)
├── core/
│ ├── unit/
│ └── integration/
├── runtime/
│ ├── unit/
│ └── integration/
├── foundation/
│ ├── unit/
│ └── integration/
├── messaging/
│ ├── unit/
│ └── integration/
├── cli/
│ └── unit/
├── charts/ # Helm chart unit tests (test_*_chart.py)
├── system/
│ ├── test_ensemble_comprehensive.py
│ ├── test_ensemble_search_e2e.py
│ ├── vespa_test_manager.py
│ ├── minio_test_manager.py # MinIOTestManager
│ └── conftest.py
├── ingestion/
│ ├── unit/
│ ├── integration/
│ ├── fixtures/
│ └── pytest.ini # Ingestion-scoped markers/addopts
├── evaluation/
│ ├── unit/
│ ├── integration/
│ ├── fixtures/
│ └── conftest.py
├── routing/
│ ├── unit/
│ ├── integration/
│ └── pytest.ini # Routing-scoped markers/addopts
├── telemetry/
│ ├── unit/
│ └── integration/
├── finetuning/
│ ├── test_*.py # Unit tests live at this level (no unit/ subdir)
│ ├── integration/
│ └── conftest.py
├── memory/
│ ├── unit/
│ ├── integration/
│ └── conftest.py
├── backends/
│ ├── unit/
│ └── integration/
├── events/
│ ├── unit/
│ └── integration/
├── admin/
│ ├── test_profile_api.py # Integration tests (run against shared_vespa)
│ ├── test_profile_concurrent.py
│ ├── test_profile_multi_tenant.py
│ ├── test_tenant_manager.py
│ ├── unit/
│ └── conftest.py
├── common/
│ ├── unit/
│ └── integration/
├── ui/ # Currently empty
├── synthetic/
│ ├── unit/
│ └── integration/
├── e2e/ # Cross-package e2e (A2A gateway, canary, CLI, ...)
│ └── deployment/
├── fixtures/ # Shared model-endpoint fixtures (llm.py, inference.py, sidecars.py)
├── utils/ # Shared test helpers (vespa_test_helpers, tenant_helpers,
│ # docker_utils, vllm_sidecar, markers, memory_store, ...)
└── conftest.py # Shared fixtures (shared_vespa, phoenix_container, etc.)
Naming Conventions¶
- Files:
test_<module_name>.py - Classes:
Test<ClassName> - Functions:
test_<behavior_description>
# test_search_agent.py
class TestSearchAgent:
"""Tests for SearchAgent."""
def test_search_by_text_returns_results(self):
"""Test that search_by_text returns results for valid query."""
...
def test_search_with_invalid_query(self):
"""Test search with invalid query parameters."""
...
@pytest.mark.asyncio
async def test_handles_backend_timeout(self):
"""Test graceful handling of backend timeout."""
...
Running Tests¶
Basic Commands¶
# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v
# Run specific package tests
JAX_PLATFORM_NAME=cpu uv run pytest tests/agents/ -v
# Run specific test file
uv run pytest tests/agents/unit/test_search_agent.py -v
# Run specific test
uv run pytest tests/agents/unit/test_search_agent.py::TestSearchAgent::test_search_by_text -v
E2E helpers outside the e2e conftest¶
Importing tests/e2e/conftest.py publishes the e2e cluster's Vespa endpoint as TEST_BACKEND_URL/TEST_BACKEND_PORT for the rest of the session, so only modules under tests/e2e import it. Unit and integration tests import e2e helpers from the plain modules beside it: cluster.py (cluster name, kube context, host ports, seeded tenant), tenants.py, sample_corpus.py, inference.py, report.py and the others in tests/e2e. tests/common/unit/test_e2e_conftest_import_guard.py fails any module outside tests/e2e that reaches the conftest, directly or through another module.
Shared e2e cluster lifecycle¶
Tests that need the shared cogniverse-e2e cluster create it when absent and reuse it across pytest processes and working sessions. They intentionally leave it running between commands because a complete verification campaign can span multiple packages and several days.
If the cluster exists but is stopped, the fixture starts it through the Cogniverse cluster lifecycle and inspects its health and deployed-content fingerprint again before reuse. A healthy, current cluster is reused as-is. A stale or unhealthy cluster fails the test session without deleting or replacing shared state; repair it, or perform the explicit reset below.
k3d fixes the loadbalancer's published ports when the cluster is created, so a cluster created before a host-port mapping was added is unhealthy until it publishes it. The failure names the command that adds the missing mappings, for example k3d cluster edit cogniverse-e2e --port-add 33912:29012@loadbalancer; it recreates only the loadbalancer.
Do not stop the cluster after an individual pytest command. Once every integration and end-to-end check in the campaign has finished, stop it explicitly to release its CPU, RAM, and accelerator allocations:
The next test requiring the cluster starts it again. Deleting the cluster is a separate explicit reset for stale or damaged state:
E2E_FRESH=1 does not authorize replacement: the cluster must already be absent. A fresh session deletes the cluster at teardown only when that same session created it; ordinary sessions always leave the shared cluster warm.
The fixture also provisions the coding sandbox in host mode: it starts (or reuses) the host OpenShell gateway, syncs its mTLS certs and metadata into the cluster before Helm installs the runtime, and deploys with runtime.sandbox.enabled=true, runtime.sandbox.gatewayEndpoint set to the gateway's own port and runtime.sandbox.hostGatewayIP set to the k3d network gateway (so host.docker.internal resolves in the runtime pod). Those values are part of the deploy identity. The gateway is the e2e stack's own, by name: openshell on port 28080 (OPENSHELL_GATEWAY_HOST_PORT). The deploy and the cert sync fail, naming the active gateway, when that gateway is unregistered, on another port, or not the host's active gateway, so no other gateway is ever deployed into the cluster. Integration tests that start their own gateway (OpenShellTestGateway in tests/agents/integration/conftest.py) register it under a private XDG_CONFIG_HOME, remove only their own container, and assert that the host's ~/.config/openshell stays byte-identical. On reuse the sync runs again and a changed secret rolls the runtime deployment, because the pod mounts the files with subPath.
There is no idle reaper: only the person or automation that knows the whole campaign is complete can safely decide when to stop the shared cluster.
The shared fixture seeds synthetic-generation checks with two real modalities: the tracked tests/system/resources/videos/v_-6dz6tBH77I.mp4 bytes and a JPEG decoded from that video's first frame. It deploys the canonical video and image profiles, then verifies the exact SHA-256 content ID, tenant-scoped MinIO URI, profile, positive documents_fed/chunk counts, and the matching documents persisted in Vespa. An unrelated approximate-nearest-neighbor hit never counts as an existing fixture.
Synthetic generation requests use the singular strategy field. The endpoint checks require exact nonzero samples from those persisted fixtures; profile and workflow examples must remain grounded in their content IDs, routing must use registered canonical agent IDs, and cross-modal examples must combine the real video and image modalities. Run that boundary directly with full output:
uv run pytest tests/e2e/test_api_e2e.py::TestSyntheticDataAPI \
--tb=long -v > /tmp/cogniverse_synthetic_e2e.log 2>&1
Leave the cluster warm when more campaign checks remain. Use the explicit cogniverse stop --name cogniverse-e2e command above only after the complete integration and end-to-end run has finished.
Test Markers¶
Static markers are registered in the root pytest.ini (--strict-markers rejects any unregistered marker) plus a per-package override in tests/ingestion/pytest.ini. Exact inference markers are registered by tests/fixtures/inference.py. The commonly used ones:
unit,integration,e2e,system— test tier.unitandintegrationare also applied automatically from the test's location (tests/<pkg>/unit/ortests/<pkg>/integration/) bytests/fixtures/markers.py::apply_location_markers, wired intopytest_collection_modifyitemsintests/conftest.pyand the nested-root conftests (tests/ingestion/,tests/routing/) — so a file that forgets the marker cannot silently fall out of its directory's CI-mselection. A file markedlocal_onlyopts out of the location marker (a declared, visible CI exclusion).tests/runtime/unit/test_marker_coverage.pyverifies every file against the actual workflow selections.ci_fast— small, essential subset run on every CI pushci_safe— safe to run in CI (no local-only dependencies)local_only— should only run locally, not in CI/CDslow,benchmark— long-running or performance testse2e_heavy— full CronWorkflow-driven e2e (DSPy optimization, synthetic data generation, distillation); opt-in only, 5-10+ minutesbrowser— Playwright-driven browser e2e teststelemetry— requires the telemetry/Phoenix providerno_shared_vespa— runtime integration module owns another real boundary; its module fixture leaves the dead backend sentinel in place and does not start the shared Vespa containerrequires_vespa,requires_docker,requires_gpu,requires_ollama,requires_gliner,requires_models,requires_cv2,requires_ffmpeg— infrastructure/model dependenciesrequires_inference("vllm_colpali")— exact ColPali/ColQwen HTTP embedding service;requires_inference("video_embed")— exact X-CLIP service. Collection also requests every service a shipped profile using the named embedding service resolves at pipeline init, derived fromconfigs/config.json(vllm_asrforvllm_colpali,video_embedandcolbert_pylate). Cluster discovery reads--revisionfrom rendered workload args: a pinned workload is taggedidentity_evidence=DEPLOYMENT, and an unpinned one staysENDPOINTand must report the exact revision from/v1/models.requires_modal_inference("vllm_llm_student")— explicitly opt an exact service into paid Modal provisioning; ordinaryrequires_inferencetests resolve the cluster's endpoint (see "Model endpoints" below).requires_whisper— Whisper model dependency outside the exact inference fixturerequires_teacher_model— scales up the vllm-llm-teacher pod (off by default)requires_optimizer_data— exercises non-router optimizers end-to-end against the live cluster (off by default)phoenix,inspect,ragas— evaluation-framework dependenciestimeout— tests with custom timeout values
local_only is an explicit run-on-demand policy, not a CI pass. In particular, the real image-search encoder test and real ColPali image-encoding test (both need the cluster's ColPali service) are intentionally excluded from normal CI and also carry slow plus their precise resource markers. Run them locally when changing their boundaries:
uv run pytest -m local_only \
tests/agents/integration/test_image_search_real_encoder_e2e.py \
tests/core/integration/test_colpali_encode_image_real.py
# Run only the fast CI subset
JAX_PLATFORM_NAME=cpu uv run pytest tests/agents/integration/ -m ci_fast -v
# Run integration tests but skip anything needing Ollama
uv run pytest tests/agents/integration/ -m "integration and not requires_ollama" -v
With Timeout¶
# 30-minute timeout for long tests (using system timeout command)
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/ -v
# For per-test timeout, install pytest-timeout first:
# uv pip install pytest-timeout
# Then use:
# uv run pytest tests/ -v --timeout=300
Parallel Execution¶
Note: pytest-xdist is not currently installed. For parallel execution, install it first:
# Install pytest-xdist
uv pip install pytest-xdist
# Run with 4 workers
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v -n 4
# Auto-detect workers
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v -n auto
With Output¶
# Show stdout/stderr
uv run pytest tests/ -v -s
# Log to file
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v 2>&1 | tee /tmp/test_output.log
Finding Affected Tests¶
Before committing, find all tests affected by your changes:
# Find tests for a module
grep -r "def test_" --include="*.py" tests/ | grep -i "search_agent"
# Find all tests in a directory
find tests/agents -name "test_*.py" -exec grep -l "def test_" {} \;
Writing Tests¶
Basic Test Structure¶
import pytest
from pathlib import Path
from cogniverse_agents.search_agent import SearchAgent, SearchInput, SearchOutput, SearchAgentDeps
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
class TestSearchAgent:
"""Tests for SearchAgent."""
@pytest.fixture
def config_manager(self):
"""Create config manager for agent."""
return create_default_config_manager()
@pytest.fixture
def schema_loader(self):
"""Create schema loader for agent."""
# FilesystemSchemaLoader requires path to schema directory
# Use relative path from project root (tests run from project root)
from pathlib import Path
schema_path = Path("configs/schemas")
return FilesystemSchemaLoader(base_path=schema_path)
@pytest.fixture
def agent(self, config_manager, schema_loader):
"""Create test agent instance."""
# Both config_manager and schema_loader are required for SearchAgent
deps = SearchAgentDeps(
backend_url="http://localhost",
backend_port=8080
)
return SearchAgent(
deps=deps,
schema_loader=schema_loader,
config_manager=config_manager,
port=8002 # A2A server port (optional, defaults to 8002)
)
@pytest.fixture
def valid_input(self):
"""Create valid test input."""
return SearchInput(
query="machine learning tutorial",
tenant_id="test-tenant",
modality="video",
top_k=10
)
def test_search_by_text_returns_results(self, agent, valid_input):
"""Test that search_by_text returns results."""
results = agent.search_by_text(
query=valid_input.query,
tenant_id="test-tenant",
modality=valid_input.modality,
top_k=valid_input.top_k
)
assert isinstance(results, list)
# Results is a list of dicts, each with id, score, video_id, frame_id, etc.
if len(results) > 0:
assert "video_id" in results[0] or "id" in results[0]
Async Tests¶
Use @pytest.mark.asyncio for async tests. Note that SearchAgent.search_by_text() is synchronous:
import pytest
from concurrent.futures import ThreadPoolExecutor
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
from cogniverse_foundation.config.manager import ConfigManager
from cogniverse_sdk.interfaces.config_store import ConfigStore
from unittest.mock import MagicMock
class TestSearchAgent:
@pytest.fixture
def mock_config_store(self):
"""Create mock ConfigStore."""
store = MagicMock(spec=ConfigStore)
store.get_config.return_value = None
return store
@pytest.fixture
def config_manager(self, mock_config_store):
"""Create ConfigManager with mock store."""
return ConfigManager(store=mock_config_store)
@pytest.fixture
def schema_loader(self):
"""Create schema loader for agent."""
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
return FilesystemSchemaLoader(Path("configs/schemas"))
@pytest.fixture
def agent(self, config_manager, schema_loader):
"""Create SearchAgent with config manager and schema loader."""
return SearchAgent(
deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
config_manager=config_manager,
schema_loader=schema_loader,
)
def test_search(self, agent):
"""Test search operation (synchronous)."""
result = agent.search_by_text(query="test query", tenant_id="test", top_k=5)
assert result is not None
def test_concurrent_searches(self, agent):
"""Test concurrent search operations using threads."""
with ThreadPoolExecutor(max_workers=3) as executor:
futures = [
executor.submit(agent.search_by_text, query="query1", tenant_id="test", top_k=5),
executor.submit(agent.search_by_text, query="query2", tenant_id="test", top_k=5),
executor.submit(agent.search_by_text, query="query3", tenant_id="test", top_k=5),
]
results = [f.result() for f in futures]
assert len(results) == 3
Parametrized Tests¶
import pytest
class TestValidation:
@pytest.mark.parametrize("query,expected_validation", [
("", "empty"),
(" ", "empty"),
("a" * 10001, "too long"),
])
def test_invalid_queries(self, agent, query, expected_validation):
"""Test validation of invalid queries."""
result = agent.search_by_text(query=query, tenant_id="test-tenant", modality="video", top_k=10)
# SearchAgent returns a list (may be empty for invalid queries)
assert isinstance(result, list)
Testing Exceptions¶
import pytest
from cogniverse_foundation.config.manager import ConfigManager
class TestConfigManager:
def test_missing_store_raises_error(self):
"""ConfigManager requires a store — passing None raises ValueError.
Note: get_system_config() itself does NOT raise when no system
config is stored — it logs a warning and returns SystemConfig()
defaults. Test the actual required-argument validation instead.
"""
with pytest.raises(ValueError):
ConfigManager(store=None)
Fixtures and Mocking¶
Shared Fixtures (conftest.py)¶
# tests/conftest.py
import pytest
@pytest.fixture
def config_manager(backend_config_env):
"""Create ConfigManager with backend store for testing.
Requires backend_config_env fixture to set environment variables.
"""
from cogniverse_foundation.config.utils import create_default_config_manager
return create_default_config_manager()
@pytest.fixture
def config_manager_memory():
"""Create ConfigManager with in-memory store for unit testing.
Does not require any backend infrastructure (Vespa, etc.).
"""
from cogniverse_foundation.config.manager import ConfigManager
from tests.utils.memory_store import InMemoryConfigStore
store = InMemoryConfigStore()
store.initialize()
return ConfigManager(store=store)
@pytest.fixture
def workflow_store(telemetry_manager_with_phoenix):
"""Resolve a workflow store via the registry — same path production uses."""
from cogniverse_core.registries import WorkflowStoreRegistry
provider = telemetry_manager_with_phoenix.get_provider(
tenant_id="workflow-store-test"
)
# Evict any instance cached under a stale provider so each test resolves clean.
WorkflowStoreRegistry.clear_cache()
store = WorkflowStoreRegistry.get(
name="telemetry",
config={"telemetry_provider": provider},
)
store.initialize()
return store
Mocking External Services¶
from unittest.mock import AsyncMock, MagicMock, patch
class TestSearchAgentWithMocks:
@pytest.fixture
def mock_backend(self):
"""Create mock backend."""
backend = MagicMock()
backend.search = AsyncMock(return_value=[
{"id": "1", "score": 0.9, "title": "Result 1"},
{"id": "2", "score": 0.8, "title": "Result 2"},
])
return backend
def test_search_calls_backend(self, mock_backend, config_manager, schema_loader):
"""Test that search calls backend with correct params."""
deps = SearchAgentDeps(
backend_url="http://localhost",
backend_port=8080
)
agent = SearchAgent(
deps=deps,
schema_loader=schema_loader,
config_manager=config_manager,
port=8002
)
result = agent.search_by_text(query="test", tenant_id="test-tenant", modality="video", top_k=5)
# Verify backend was called (if using mock backend in deps)
# mock_backend.search.assert_called()
def test_handles_backend_error(self, mock_backend, config_manager, schema_loader):
"""Test graceful handling of backend error."""
mock_backend.search.side_effect = ConnectionError("Backend down")
deps = SearchAgentDeps(
backend_url="http://localhost",
backend_port=8080
)
agent = SearchAgent(
deps=deps,
schema_loader=schema_loader,
config_manager=config_manager,
port=8002
)
result = agent.search_by_text(query="test", tenant_id="test-tenant", modality="video", top_k=10)
# Result is a list (may be empty or contain error info depending on implementation)
assert isinstance(result, list)
Patching¶
from unittest.mock import patch, MagicMock
class TestWithPatching:
@patch("cogniverse_agents.search_agent.get_backend_registry")
def test_with_patched_backend(self, mock_registry, config_manager, schema_loader):
"""Test with patched backend registry."""
mock_backend = MagicMock()
mock_registry.return_value.get_search_backend.return_value = mock_backend
deps = SearchAgentDeps(
backend_url="http://localhost",
backend_port=8080
)
agent = SearchAgent(
deps=deps,
schema_loader=schema_loader,
config_manager=config_manager
)
# Verify backend registry was accessed
assert agent is not None
Integration Tests¶
With Real Services¶
import pytest
@pytest.mark.integration
class TestVespaIntegration:
"""Integration tests requiring Vespa."""
@pytest.fixture
def search_backend(self, config_manager, schema_loader):
"""Create real VespaSearchBackend."""
import os
from cogniverse_vespa.search_backend import VespaSearchBackend
backend_url = os.getenv("BACKEND_URL", "http://localhost")
backend_port = int(os.getenv("BACKEND_PORT", "8080"))
return VespaSearchBackend(
config={
"url": backend_url,
"port": backend_port,
"profiles": {"video_colpali_smol500_mv_frame": {}},
"default_profiles": {"video": "video_colpali_smol500_mv_frame"},
},
config_manager=config_manager,
schema_loader=schema_loader,
)
def test_real_search(self, search_backend):
"""Test search against real Vespa."""
# VespaSearchBackend.search() takes a query_dict and returns
# List[SearchResult]. strategy is a rank-profile-name string.
results = search_backend.search({
"query": "test",
"type": "video",
"profile": "video_colpali_smol500_mv_frame",
"strategy": "bm25_only",
"top_k": 5,
"tenant_id": "test:unit",
})
assert isinstance(results, list)
Skipping Without Services¶
import pytest
import os
requires_vespa = pytest.mark.skipif(
os.getenv("BACKEND_URL") is None,
reason="BACKEND_URL not set"
)
@requires_vespa
class TestVespaRequired:
"""Tests that require Vespa."""
def test_vespa_operation(self):
...
LLM-driven integration tests (llm_endpoint fixture)¶
Tests that need a real LLM (visual judges, Inspect-AI eval tasks) consume the provider-agnostic llm_endpoint fixture defined in tests/evaluation/integration/conftest.py. The fixture resolves the endpoint from, in order:
- Env vars
COGNIVERSE_TEST_LLM_PROVIDER_URI(required) andCOGNIVERSE_TEST_LLM_BASE_URL(optional). When a base URL is set, it is propagated to the provider's own env var — onlyopenai/vllm(→OPENAI_BASE_URL) andanthropic(→ANTHROPIC_BASE_URL) prefixes are mapped today; a matching*_API_KEYof"not-required"is also injected so Inspect AI's provider client doesn't refuse to initialize against a keyless local endpoint.TEST_LLM_MODEL/TEST_LLM_API_BASE(as set by theensure_host_ollamafixture intests/conftest.py) are also accepted as a same-priority alternative. - JSON file at
tests/evaluation/integration/resources/test_llm.json(gitignored; atest_llm.example.jsonis checked in). llm_config.primaryof the session configensure_host_ollamawrites, which points at the Modal Gemma (see "Model endpoints" below).- Otherwise, the fixture fails the test with a message explaining how to configure an endpoint.
Hermetic session configs publish only roles whose exact model endpoint was resolved and verified. A primary-only session points the teacher entry at a dead port instead of exposing an endpoint the fixture did not check. Tests marked requires_teacher_model resolve and verify the distinct teacher before adding that role. Teacher-dependent code must call resolve_teacher(), which fails explicitly when the role was not configured; it never falls back to the primary. Real optimizer integration tests that reach resolve_teacher() carry this marker so collection records the teacher role before the session config is materialized.
Model endpoints¶
No test starts a model on this host: no model container, no vLLM, Ollama or model-server process. Parallel test runs share the host's memory, and a local model copy holds gigabytes of it. Every model a test uses is served remotely:
- Chat LLMs (the Gemma student, the Qwen teacher) are served on Modal.
ensure_llm(tests/utils/hermetic_llm.py), whichensure_host_ollama,gemma_inference_endpointandrequires_inference("vllm_llm_student")use, reads the role's deployment throughModalInferenceLifecycle.statusand accepts it only when its authenticated/v1/modelsnames the exact model and revision. It does not depend on any cluster. AnINFERENCE_SERVICE_URLSentry forvllm_llm_studentorvllm_llm_teacherreplaces the Modal lookup. A missingCOGNIVERSE_INFERENCE_API_KEY, an undeployed app or a wrong model raisesModalLlmNotDeployedErrororLlmEndpointMismatchErrornaming the deploy command; the test fails, never skips or falls back. - Every other model (ColPali/Tomoro, DenseOn, PyLate LateOn and the code encoder, GLiNER, ASR, CLAP, video-embed, face-embed) is served by the
cogniverse-e2ecluster on GPU.remote_inference(tests/fixtures/inference.py) and therequires_inferencemarker resolve a service from an explicitINFERENCE_SERVICE_URLSentry, else the workload thecogniverse-e2ecluster (then thecogniversedev cluster) publishes through its load balancer, and validate its model identity. With no endpoint it raisesRemoteServiceUnavailablenaming the service.
Containers that serve no model (Vespa, Redis, MinIO, Phoenix, the semantic router stack) are still started by the tests' own fixtures.
Before a run that needs models:
# Modal credentials and the inference key live in .env/ (loaded by tests/conftest.py)
ls .env/MODAL_TOKEN_ID.env .env/MODAL_TOKEN_SECRET.env .env/COGNIVERSE_INFERENCE_API_KEY.env
# The chat models must be deployed on Modal
uv run cogniverse inference modal status vllm_llm_student vllm_llm_teacher
uv run cogniverse inference modal deploy vllm_llm_student # if status fails
# The other models come from the cogniverse-e2e cluster
kubectl --context k3d-cogniverse-e2e -n cogniverse get deploy,svc
When the deploy serves the chat models from Modal, the e2e session runs cogniverse inference modal warm vllm_llm_student before its first test and release at teardown: the query rewrite gives the student 2.8 s, which a runner starting from zero cannot meet. A failed warm fails the session. The deploy itself refuses to build when the host disk holding docker's data is at or above 75%, Vespa's feed-block limit; free space (docker builder prune -f) and rerun.
The cluster is deployed through the e2e path (tests/e2e/deployment), not a plain cogniverse up. tests/common/unit/test_no_local_model_servers.py fails any test code that runs a model image's container, launches a vLLM/Ollama server, or loads a model in-process (SentenceTransformer, PyLate ColBERT, from_pretrained on a model class, Whisper, faster-whisper, FaceAnalysis).
No test loads model weights into the pytest process either: the session plugin tests/fixtures/no_local_models.py replaces the weight-loading entry points of transformers, sentence-transformers (and so PyLate), GLiNER, the Hugging Face hub mixin, Whisper, faster-whisper, InsightFace and open_clip with a refusal (LocalModelLoadForbidden), however the call is reached. Product code that degrades when a model is missing keeps degrading; a test whose result needs the model fails on its assertions. Every refused load is listed, by test, in the session's refused in-process model loads summary. A test that needs a model's output uses the cluster service that serves it (remote_inference, gliner_url, served_semantic_embedder).
Parity tests compare the served model against reference outputs recorded once, on CPU, from the model's own library at the pinned revision: tests/fixtures/model_references/{lateon,denseon}.json carry the model, revision, library versions, device and inputs beside the outputs. Tests never run the recorder; re-record after a pinned revision changes:
Every pytest session prints a test sidecars section in its terminal summary, captured output or not: each model resolution (resolved-remote with its endpoint and provider, or refused with the reason) and the dead-owner containers the session reaped when it started.
Examples:
# OpenAI for inspect_eval-driven retrieval tests; OPENAI_API_KEY picked up
# automatically by the provider SDK.
COGNIVERSE_TEST_LLM_PROVIDER_URI="openai/gpt-4o-mini" \
uv run pytest tests/evaluation/integration/test_end_to_end.py
# vLLM behind an OpenAI-compatible proxy:
COGNIVERSE_TEST_LLM_PROVIDER_URI="vllm/Qwen/Qwen2.5-7B-Instruct" \
COGNIVERSE_TEST_LLM_BASE_URL="http://vllm.internal:8000/v1" \
uv run pytest tests/evaluation/integration/
# No env vars set: the fixture uses the Modal Gemma ensure_host_ollama resolves.
uv run pytest tests/evaluation/integration/test_visual_judge_e2e.py
Test classes never reference any specific provider, model, or container manager — they only consume llm_endpoint["provider_uri"] and llm_endpoint["base_url"]. API keys are resolved by the provider SDK from its own conventional env var (OPENAI_API_KEY, ANTHROPIC_API_KEY, ...) — the test infrastructure does not re-implement that lookup.
MediaLocator-driven integration tests (media_root_uri)¶
Tests that exercise the unified-MediaLocator rollout (ingestion, audio transcribe, visual judge) read media via MediaLocator instead of bare filesystem paths. The pipeline accepts a media_root_uri config (e.g. s3://corpus/ or file:///abs/path/) that the locator joins with each video's relative path to produce the canonical source_url written into Vespa. See tests/ingestion/integration/test_pipeline_minio_round_trip.py for the end-to-end pattern (MinIO docker fixture managed by MinIOTestManager); the schema-level field is documented in docs/modules/common.md under "MediaLocator".
Test Coverage¶
Running with Coverage¶
# Basic coverage
uv run pytest tests/ --cov=cogniverse_core --cov-report=term
# HTML report
uv run pytest tests/ --cov=cogniverse_core --cov-report=html
# Multiple packages
uv run pytest tests/ \
--cov=cogniverse_core \
--cov=cogniverse_agents \
--cov=cogniverse_foundation \
--cov-report=html
Coverage Configuration¶
Coverage is not configured in pyproject.toml or the root pytest.ini — it's specified per invocation via --cov=<path> flags, either on the command line or in a workflow's test step. The per-package tests/ingestion/pytest.ini explicitly notes this in its addopts comment ("no coverage here — handled by CI/Make targets").
Coverage Targets¶
Most CI workflows collect and report coverage (--cov-report=term-missing, --cov-report=xml) without enforcing a minimum. Only two currently gate the build on a coverage floor via --cov-fail-under:
| Workflow | Package | Enforced Minimum |
|---|---|---|
evaluation-tests.yml | cogniverse-evaluation | 50% |
routing-tests.yml | cogniverse-agents (routing) | 15% |
All other workflows (agents, core, finetuning, ingestion, synthetic, telemetry, vespa, runtime) report coverage but do not fail the build below any specific percentage.
CI/CD Testing¶
GitHub Actions Workflows¶
The project has 19 GitHub workflow files: 14 per-module test workflows plus chart-validation.yml, test-integrity.yml, docs.yml, publish-packages.yml, and two manual/release workflows not tied to a single module:
| Workflow | Module | Tests | Docker Services |
|---|---|---|---|
agents-tests.yml | cogniverse-agents | unit + integration | Vespa (ci_fast subset) |
chart-validation.yml | Helm chart (charts/cogniverse) | lint + template + kubeconform | None |
cli-tests.yml | cogniverse-cli | unit + integration | None |
core-tests.yml | cogniverse-core (incl. tests/core/*, tests/memory/* and the ci_fast files under tests/utils/; the rest of memory integration is local-tier — it needs the cluster's DenseOn service) | unit + integration | Vespa |
evaluation-tests.yml | cogniverse-evaluation | unit + integration | Phoenix |
finetuning-tests.yml | cogniverse-finetuning | unit + integration | Vespa |
ingestion-tests.yml | cogniverse-runtime (ingestion) | unit + integration | Vespa |
messaging-tests.yml | cogniverse-messaging | unit + integration | None |
routing-tests.yml | cogniverse-agents (routing) | unit + integration | Vespa |
runtime-tests.yml | cogniverse-runtime, cogniverse-foundation, cogniverse-cli, events, cogniverse-messaging (unit for all; integration for runtime + the small events/foundation/messaging suites; the web client's browser suites, test_web_*.py and test_ag_ui_threads.py, run in five parallel web-ops-integration-tests-* jobs and the Optimization framework suite in web-ops-optimization-framework-tests) | unit + integration | Vespa |
synthetic-tests.yml | cogniverse-synthetic | unit + integration | Phoenix |
telemetry-tests.yml | cogniverse-telemetry-phoenix | unit + integration | Phoenix |
test-integrity.yml | Whole-repo test guards (no paths filter) | assertion strength + CI coverage | None |
vespa-tests.yml | cogniverse-vespa | unit + integration | Vespa |
docs.yml | Documentation | mkdocs build + deploy to GitHub Pages | None |
publish-packages.yml | Package publishing (PyPI/TestPyPI) | N/A | None |
mirror-third-party.yml | Air-gapped support (workflow_dispatch only) | Mirrors third-party images (Vespa, Phoenix, Ollama, …) to a target registry | None |
release-images.yml | Release (v* tag push or workflow_dispatch) | Builds/pushes every first-party image + Helm chart OCI artifact | None |
Workflow Structure¶
Each test workflow typically has these jobs:
- unit-tests - Unit tests; unit-dir jobs select
-m unit(satisfied by every test via the location-derived marker) so no file can silently fall out of the selection - integration-tests - Integration tests (with ci_fast subset for quick feedback)
- lint - Code linting with ruff/black
- test-cli or test-imports - Verify package imports work
- coverage-report - Combined coverage (if applicable)
Note: Workflows don't have separate "fast-integration-tests" jobs. Instead, integration-tests jobs use -m ci_fast to run essential tests quickly.
Guards on CI coverage itself¶
A test CI never runs reports as absence, which reads like success. Three guards close that:
tests/runtime/unit/test_marker_coverage.py— every file under a selected path carries markers its selection's-mexpression keeps.tests/common/unit/test_ci_coverage_guard.py— everyci_fasttest that is notlocal_onlyis named by some commit-gating selection whose expression keeps it; every selection's own paths are inside its workflow'spathsfilter; and every test that walks a tree outside its own package runs in a workflow that fires on changes to that tree.tests/common/unit/test_ci_fast_excludes_model_spawners.py— noci_fastselection reaches a fixture that needs a model server.
All three read the selections from tests/fixtures/ci_workflows.py, which parses .github/workflows/*.yml; a tag-only workflow gates no commit and its selections do not count. test-integrity.yml runs the whole-tree guards with no paths filter, since any filter would skip them on the commits they exist to catch.
Assertion-strength guard and its waivers¶
tests/common/unit/test_assertion_strength_guard.py fails a change that leaves any tests/ file with fewer assertions than before (net of assertions moved verbatim into a file the change creates), or that adds a skip, an xfail or an unbounded assertion form. CI compares against the pull request's base or the push's previous commit (ASSERTION_GUARD_BASE); locally it defaults to HEAD~1:
ASSERTION_GUARD_BASE=$(git merge-base HEAD main) \
uv run pytest tests/common/unit/test_assertion_strength_guard.py
The one accepted loss is of assertions that tested code the change deletes. Each is declared in tests/common/assertion_waivers.toml:
[[waiver]]
file = "tests/utils/test_vllm_sidecar.py" # the test file that lost assertions
max_net_loss = 173 # its largest accepted net loss
removed_symbols = ["tests/utils/vllm_sidecar.py:VllmSidecarFactory"] # path:Name
reason = "Tests of the local vLLM sidecar launch, removed with it."
The guard checks every waiver against the compared range: each named top-level symbol must be gone at HEAD and must have existed in HEAD's history, the file's net loss must not exceed max_net_loss, and a waiver for a file that lost nothing fails as stale. A waiver applies only when at least one of its symbols existed at the base; one whose symbols were all gone before the range covered an earlier change and waives nothing.
CI Fast Integration Tests¶
Workflows run integration tests with the ci_fast marker on every push to provide quick feedback:
integration-tests:
runs-on: ubuntu-latest
timeout-minutes: 60
needs: unit-tests
steps:
- uses: actions/checkout@v4
- name: Set up Python 3.12
uses: actions/setup-python@v4
with:
python-version: '3.12'
- name: Install uv
uses: astral-sh/setup-uv@v4
with:
version: "latest"
- name: Install dependencies
run: |
uv sync --all-packages --all-extras
uv pip install pytest-cov
- name: Free up disk space for Vespa (if needed)
run: |
# Vespa requires <75% disk usage
sudo rm -rf /usr/share/dotnet
sudo rm -rf /usr/local/lib/android
sudo rm -rf /opt/ghc
sudo rm -rf /opt/hostedtoolcache/CodeQL
sudo docker image prune -af
- name: Run integration tests (CI Fast subset)
run: |
JAX_PLATFORM_NAME=cpu uv run python -m pytest \
tests/module/integration/ \
-m ci_fast \
-v --tb=long \
--cov=libs/module/cogniverse_module
Disk Cleanup for Vespa¶
Vespa requires disk usage below 75%. GitHub runners have ~14GB disk, so cleanup is required:
# Remove ~30GB of unused packages
sudo rm -rf /usr/share/dotnet # .NET SDK (~6GB)
sudo rm -rf /usr/local/lib/android # Android SDK (~10GB)
sudo rm -rf /opt/ghc # Haskell compiler (~2GB)
sudo rm -rf /opt/hostedtoolcache/CodeQL # CodeQL (~5GB)
sudo docker image prune -af # Unused Docker images
Docker Service Management¶
Tests share a single session-scoped Vespa container and manage their own Phoenix containers via fixtures.
Vespa (shared_vespa):
All Vespa integration tests use the shared_vespa session-scoped fixture defined in tests/conftest.py. Each per-package conftest.py re-exports it and provides a compatibility shim under whatever fixture name its tests expect:
# e.g. tests/backends/integration/conftest.py
from tests.conftest import shared_vespa # noqa: F401
@pytest.fixture(scope="module")
def vespa_instance(shared_vespa):
"""Shim: yields the dict shape tests expect, backed by shared_vespa."""
yield {
"http_port": shared_vespa["http_port"],
"config_port": shared_vespa["config_port"],
"base_url": shared_vespa["base_url"],
"container_name": shared_vespa["container_name"],
}
# No teardown — shared_vespa owns the container lifecycle.
For tests that need to deploy their own schemas, use the helpers in tests/utils/vespa_test_helpers.py:
from tests.utils.vespa_test_helpers import deploy_tenant_schema, make_config_manager
from tests.utils.tenant_helpers import tenant_id_for_test
@pytest.fixture
def my_schema(shared_vespa, request):
tenant_id = tenant_id_for_test(request)
config_manager = make_config_manager(shared_vespa)
deploy_tenant_schema(
shared_vespa,
tenant_id=tenant_id,
base_schema_name="video_colpali_smol500_mv_frame",
config_manager=config_manager,
)
yield tenant_id
Phoenix:
The phoenix_container fixture is defined in tests/conftest.py (module-scoped). It allocates non-default, per-process ports to avoid conflicts both with local Phoenix instances and with other concurrent pytest sweeps:
- HTTP:
16006 + port_offset(instead of 6006) - gRPC:
14317 + port_offset(instead of 4317)
where port_offset = (os.getpid() % 1000) * 10, giving each process a distinct 10-port-spaced slot in a ~1000-process range.
Image: arizephoenix/phoenix:20.16.0 pinned by digest (never :latest).
Containers are named phoenix_test_pid<pid>_<timestamp> and tagged with the owning pid; on startup the fixture only kills leftover containers matching phoenix_test_pid<its own pid>_* from a prior crashed run of that same process — never another concurrent session's container.
Pre-commit Checklist¶
Before every commit:
# 1. Find affected tests
grep -r "def test_" --include="*.py" tests/ | grep -i "<module>"
# 2. Run tests
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/ -v 2>&1 | tee /tmp/test_output.log
# 3. Verify 100% pass
grep -E "passed|failed" /tmp/test_output.log
# 4. Lint
uv run make lint-all
Troubleshooting¶
JAX Platform Errors¶
# Always set JAX platform for CPU
JAX_PLATFORM_NAME=cpu uv run pytest tests/ -v
# Or in conftest.py
import os
os.environ["JAX_PLATFORM_NAME"] = "cpu"
Async Test Not Running¶
# Ensure pytest-asyncio is installed and configured
# In pytest.ini (already configured):
# asyncio_mode = auto
# asyncio_default_fixture_loop_scope = function
# asyncio_default_test_loop_scope = function # pinned so a leaked coroutine
# in one async test can't fail every subsequent async test in the same
# module/session scope (pytest-asyncio 1.3+ default)
# Use decorator for async tests
@pytest.mark.asyncio
async def test_async_function():
...
Import Errors¶
# Ensure packages are installed
uv sync
# Check import path
uv run python -c "import cogniverse_core; print(cogniverse_core.__file__)"
# Run from project root
cd /path/to/cogniverse
uv run pytest tests/ -v
Flaky Tests¶
For retries or timeout decorators, install the required packages first:
# Install pytest-timeout for timeout decorator
uv pip install pytest-timeout
# Install pytest-rerunfailures for flaky test retries
uv pip install pytest-rerunfailures
Then use:
# Use retries for network-dependent tests (requires pytest-rerunfailures)
@pytest.mark.flaky(reruns=3)
def test_network_operation():
...
# Or add explicit timeout (requires pytest-timeout)
@pytest.mark.timeout(30)
def test_slow_operation():
...
Test Isolation¶
# Each test should be independent
# Use fixtures with appropriate scope
@pytest.fixture # Default: function scope - fresh for each test
def config():
return Config()
@pytest.fixture(scope="class") # Shared within class
def expensive_resource():
return create_expensive_resource()
@pytest.fixture(scope="session") # Shared for entire session
def database():
return setup_database()
Related Documentation¶
Several tests/ subpackages carry their own more detailed testing guide; this file covers cross-cutting practices, these cover package-specific scenarios and fixtures:
tests/README.md/tests/CONSOLIDATED_TEST_README.md— full test suite overviewtests/agents/README.md— multi-agent routing, A2A protocol, DSPy/GEPA testingtests/agents/e2e/README.md— real-LLM end-to-end tests (Ollama, Vespa, Phoenix)tests/ingestion/README.md— ingestion pipeline test suitetests/routing/README.md— query routing and classification test suite
Summary¶
- Organize tests by package (agents, system, ingestion, evaluation, etc.)
- Use fixtures for common setup
- Run with JAX_PLATFORM_NAME=cpu to avoid GPU issues
- 100% pass rate required before commit
- Find affected tests with grep before committing
- Use mocks for external services in unit tests
- Mark integration tests that require real services