Skip to content

Pytest Best Practices


Async Testing Configuration

Threading Issues with Async Tests

Async tests in pytest can encounter threading conflicts when loading ML models. The codebase uses several libraries that spawn background threads:

  • tqdm (from transformers): Progress bars during model downloads
  • posthog (from mem0ai): Telemetry/analytics background threads
  • torch: Multi-threaded tensor operations

These background threads can cause segmentation faults during pytest cleanup when combined with async event loops.

Solution: Single-Threaded Mode

The test suite is configured for single-threaded execution to avoid threading conflicts:

File: tests/conftest.py

import os

# Configure torch and tokenizers to avoid threading issues
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["OMP_NUM_THREADS"] = "1"
os.environ["MKL_NUM_THREADS"] = "1"

# Import torch and configure threading before any tests run
try:
    import torch

    torch.set_num_threads(1)
except ImportError:
    pass

File: pytest.ini

[pytest]
asyncio_mode = auto
asyncio_default_fixture_loop_scope = function
asyncio_default_test_loop_scope = function

asyncio_default_test_loop_scope pins the test loop to function scope too — without it, pytest-asyncio 1.3+ defaults the test loop to module/session scope, so an async test that leaves the loop "running" (e.g. a backgrounded coroutine that never fully awaited) makes every subsequent async test in the same scope fail with Runner.run() cannot be called from a running event loop. Function-scoped loops get torn down per-test, isolating leaks.

Background Thread Cleanup

The conftest also includes automatic cleanup for background threads:

from tests.utils.async_polling import simulate_processing_delay

def cleanup_background_threads():
    """
    Clean up background threads from tqdm (transformers) and posthog (mem0ai).

    These libraries create daemon threads that can cause segfaults during pytest
    cleanup in async tests. We need to give them time to finish and exit cleanly.
    """
    max_wait = 2.0  # seconds
    start_time = time.time()

    while time.time() - start_time < max_wait:
        background_threads = [
            t
            for t in threading.enumerate()
            if t != threading.current_thread()
            and t.daemon
            and any(name in t.name.lower() for name in ["tqdm", "posthog", "monitor"])
        ]

        if not background_threads:
            break

        # Give threads time to finish their work
        simulate_processing_delay(delay=0.1, description="test processing")

    # Force garbage collection to clean up any remaining references
    gc.collect()

Running Tests

Basic Test Execution

# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest

# Run specific test file with 30-minute timeout
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/agents/test_routing.py

# Run with verbose output
JAX_PLATFORM_NAME=cpu uv run pytest -xvs

Test Markers

# Run only unit tests
uv run pytest -m unit

# Run only integration tests
uv run pytest -m integration

# Run fast CI tests (subset for quick feedback)
uv run pytest -m ci_fast

# Skip slow tests
uv run pytest -m "not slow"

# Run tests declaring an exact inference service
uv run pytest -m requires_inference
uv run pytest -m requires_ollama

# Skip tests requiring Ollama (for CI)
uv run pytest -m "not requires_ollama"

Available Markers

Defined in pytest.ini:

[pytest]
markers =
    unit: Unit tests for individual components
    integration: Integration tests with multiple components
    slow: Slow tests that take significant time
    requires_ollama: Tests that require Ollama to be running
    requires_lm: Tests that require the configured test LM endpoint
    requires_gliner: Tests that require GLiNER models
    phoenix: Tests requiring Phoenix
    inspect: Tests requiring Inspect AI
    ragas: Tests requiring RAGAS
    asyncio: Async tests
    local_only: Tests that should only run locally (not in CI/CD)
    requires_models: Tests that require actual ML models to be available
    benchmark: Performance benchmarking tests
    ingestion: Tests for ingestion pipeline
    requires_vespa: Tests that require Vespa backend to be running
    no_shared_vespa: Integration modules that own a non-Vespa boundary
    requires_docker: Tests that require Docker
    requires_gpu: Tests that require GPU availability
    requires_whisper: Tests that require Whisper models
    requires_teacher_model: Tests that scale up the vllm-llm-teacher pod (long-running, off by default)
    requires_optimizer_data: Tests that exercise non-router optimizers (workflow / modality / xgboost) end-to-end against the live cluster (slow, off by default)
    requires_cv2: Tests that require OpenCV
    requires_ffmpeg: Tests that require FFmpeg
    ci_safe: Tests that are safe to run in CI environment
    ci_fast: Fast, essential tests for CI (subset of most important functionality)
    timeout: Tests with custom timeout values
    e2e: End-to-end integration tests with real services
    e2e_heavy: Heavy end-to-end tests that submit a full CronWorkflow execution and may take 5-10+ minutes (DSPy optimization, synthetic data generation, distillation); opt-in via -m e2e_heavy
    browser: Browser-based E2E tests requiring Playwright
    system: System-level end-to-end tests with full infrastructure
    telemetry: Tests for telemetry and observability system

Exact inference markers are registered by tests/fixtures/inference.py, not pytest.ini. Name the canonical service in the decorator:

@pytest.mark.requires_inference("vllm_colpali")
def test_colpali_or_colqwen_boundary():
    ...


@pytest.mark.requires_inference("video_embed")
def test_xclip_boundary():
    ...

ColPali and ColQwen use the same vllm_colpali HTTP service. The collector also requests every service a shipped profile using the named embedding service resolves at pipeline init, derived from configs/config.json (vllm_asr for vllm_colpali, video_embed and colbert_pylate). Use @pytest.mark.requires_modal_inference("vllm_llm_student") only when the test explicitly requires that exact service through paid Modal provisioning.

The ci_fast Marker

The ci_fast marker identifies tests run in the CI Fast subset steps across workflows:

# tests/backends/integration/test_tenant_schema_lifecycle.py
@pytest.mark.integration
@pytest.mark.ci_fast
class TestSchemaRegistryDeployment:
    """Test schema deployment via SchemaRegistry"""

    def test_deploy_single_schema(self, get_backend):
        """Test deploying a single schema for a tenant"""
        backend = get_backend("acme")
        backend.schema_registry.deploy_schema("acme", "video_colpali_smol500_mv_frame")

        schemas = backend.schema_registry.get_tenant_schemas("acme")
        assert len(schemas) == 1
        assert schemas[0].base_schema_name == "video_colpali_smol500_mv_frame"

Guidelines for ci_fast tests:

  • Essential functionality that must work

  • Complete in under 2 minutes

  • No external API calls (Ollama, OpenAI, etc.)

  • Can use Docker containers (Vespa, Phoenix)

Async Test Timeout

Some async tests carry a @pytest.mark.timeout(N) marker (registered in pytest.ini as timeout: Tests with custom timeout values) with a literal per-test second count:

# tests/routing/integration/test_deep_research_integration.py
@pytest.mark.asyncio
@pytest.mark.timeout(120)
async def test_full_research_cycle(self, real_search_fn, seeded_outdoor_corpus):
    """Decompose -> search Vespa -> evaluate -> synthesize against the configured LM."""
    ...

Note the pytest-timeout plugin that enforces this marker is not a project dependency — install it (uv pip install pytest-timeout) for the marker to actually abort a hung test; without it, the marker is inert and only the shell-level timeout wrapper below bounds runtime.

Conventional shell-level timeouts used across the project's own test commands and CI workflows:

  • Individual test files: 30 minutes (timeout 1800)

  • Full test suite: 120 minutes (timeout 7200)


Common Issues

Segmentation Faults in Async Tests

Symptoms:

Fatal Python error: Segmentation fault
Thread 0x000000033614f000 (most recent call first):
  File "/path/to/threading.py", line 359 in wait

Cause: Threading conflict between pytest async event loop and background threads from model loading.

Solution: Already handled by tests/conftest.py configuration. If you still see segfaults:

  1. Check that tests use @pytest.mark.asyncio correctly
  2. Verify pytest.ini has asyncio_mode = auto
  3. Ensure no manual thread creation in test code

Model Loading Errors

Symptoms:

Fetching 5 files: 100%
[Segfault or hang]

Cause: Locally downloading and loading a large embedding/VLM model inside the pytest process can trigger the same background-thread conflicts described above.

Solution: The primary embedding model (TomoroAI/tomoro-colqwen3-embed-4b, configured as colpali_model / per-profile embedding_model in configs/config.json) has no in-process loader — ColPaliModelLoader and ColQwenModelLoader (libs/core/cogniverse_core/common/models/model_loaders.py) raise immediately for any ColQwen3/Tomoro model name, directing callers to RemoteColPaliLoader instead:

ColQwen3/Tomoro models are remote-only — serve via vLLM and set
inference_service_url (profile inference_services.embedding). Local
in-process loading is unsupported (requires transformers>=4.57, blocked
by the pylate cap).

Tests marked requires_inference("vllm_colpali") therefore exercise ColPali and ColQwen through that exact service over HTTP rather than downloading weights into the test process — avoid adding a local from_pretrained(...) call for these models in new tests.

Import Timing Issues

Symptoms:

ModuleNotFoundError: No module named 'cogniverse_core'

Cause: UV workspace not synced or tests run outside virtual environment.

Solution: Always use uv run pytest which ensures workspace packages are available.

Example Fix:

# ❌ Bad: Direct pytest without uv
pytest tests/agents/

# ✅ Good: Use uv run to activate workspace
uv run pytest tests/agents/

Package Import Patterns:

# ✅ Good: Absolute imports from workspace packages
from cogniverse_foundation.telemetry.manager import TelemetryManager
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent
from cogniverse_vespa.backend import VespaBackend

# ❌ Bad: Old src-style imports (not valid in uv workspace)
from src.agents.orchestrator_agent import OrchestratorAgent  # ❌ Will fail


Package Testing

Testing Package Imports

Example Test Patterns - Verify Package Structure:

# Example: tests/test_imports.py (create this file to verify imports)
import pytest

def test_sdk_package_imports():
    """Verify cogniverse_sdk package imports work"""
    from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
    from pathlib import Path
    from cogniverse_sdk import SearchResult, Document

    assert FilesystemSchemaLoader is not None
    assert SearchResult is not None
    assert Document is not None

def test_foundation_package_imports():
    """Verify cogniverse_foundation package imports work"""
    from cogniverse_foundation.telemetry.manager import TelemetryManager

    assert TelemetryManager is not None

def test_core_package_imports():
    """Verify cogniverse_core package imports work"""
    from cogniverse_foundation.config.unified_config import SystemConfig
    from cogniverse_core.registries.backend_registry import BackendRegistry

    assert SystemConfig is not None
    assert BackendRegistry is not None

def test_agents_package_imports():
    """Verify cogniverse_agents package imports work"""
    from cogniverse_agents.orchestrator_agent import OrchestratorAgent
    from cogniverse_agents.search_agent import SearchAgent

    assert OrchestratorAgent is not None
    assert SearchAgent is not None

def test_retrieval_package_imports():
    """Verify cogniverse_vespa package imports work"""
    from cogniverse_vespa.backend import VespaBackend
    from cogniverse_vespa.vespa_schema_manager import VespaSchemaManager

    assert VespaBackend is not None
    assert VespaSchemaManager is not None

def test_processing_package_imports():
    """Verify cogniverse_runtime.ingestion package imports work"""
    from cogniverse_runtime.ingestion.pipeline import VideoIngestionPipeline
    from cogniverse_runtime.ingestion.pipeline_builder import VideoIngestionPipelineBuilder

    assert VideoIngestionPipeline is not None
    assert VideoIngestionPipelineBuilder is not None

def test_evaluation_package_imports():
    """Verify cogniverse_evaluation package imports work"""
    from cogniverse_evaluation.core.experiment_tracker import ExperimentTracker
    from cogniverse_evaluation.metrics import calculate_mrr, calculate_ndcg

    assert ExperimentTracker is not None
    assert calculate_mrr is not None
    assert calculate_ndcg is not None

Package Dependency Testing

Example Test Patterns - Cross-Package Dependencies:

# Example: tests/test_package_dependencies.py (create this file to verify dependencies)
import pytest
from cogniverse_foundation.telemetry.manager import TelemetryManager
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent

def test_agents_depends_on_foundation_and_core():
    """Verify agents package can use foundation and core packages"""
    from cogniverse_agents.orchestrator_agent import OrchestratorDeps
    from cogniverse_core.registries.agent_registry import AgentRegistry
    from cogniverse_foundation.config.manager import ConfigManager
    from cogniverse_foundation.config.unified_config import LLMEndpointConfig
    from cogniverse_foundation.telemetry.config import TelemetryConfig
    from tests.utils.memory_store import InMemoryConfigStore

    telemetry_config = TelemetryConfig()
    deps = OrchestratorDeps(
        telemetry_config=telemetry_config,
        llm_config=LLMEndpointConfig(
            model="openai/google/gemma-4-e4b-it",
            api_base="http://localhost:11434/v1",
        ),
    )

    # OrchestratorAgent requires deps, registry, and config_manager
    store = InMemoryConfigStore()
    store.initialize()
    config_manager = ConfigManager(store=store)
    registry = AgentRegistry(tenant_id="test-tenant", config_manager=config_manager)
    agent = OrchestratorAgent(deps=deps, registry=registry, config_manager=config_manager)

    assert agent.deps is not None
    assert agent.deps == deps

def test_runtime_depends_on_all():
    """Verify runtime package can use all dependencies"""
    from cogniverse_foundation.config.unified_config import SystemConfig
    from cogniverse_agents.orchestrator_agent import OrchestratorAgent
    from cogniverse_vespa.backend import VespaBackend

    # Runtime should be able to import all packages
    assert OrchestratorAgent is not None
    assert VespaBackend is not None
    assert SystemConfig is not None

def test_layered_architecture_dependencies():
    """Verify proper layering - lower layers don't import higher layers"""
    # SDK should not import from other packages
    # Foundation can import SDK
    # Core can import SDK and Foundation
    # Agents can import SDK, Foundation, and Core
    # Implementation can import SDK, Foundation, and Core
    # Services can import all
    pass


Multi-Tenant Testing

Tenant Isolation Tests

Test Tenant-Specific Configuration:

# tests/test_tenant_isolation.py
import pytest
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps

def test_tenant_config_isolation():
    """Verify tenant IDs are distinct — isolation is enforced per-operation."""
    # Tenant IDs are plain strings passed to operations (search, schema deploy,
    # memory add). SystemConfig does not carry a tenant_id field; the backend
    # URL/port are shared across tenants in the same deployment.
    tenant_a = "acme_corp"
    tenant_b = "globex_inc"

    assert tenant_a != tenant_b

def test_tenant_schema_naming():
    """Verify tenant schemas use correct naming convention"""
    from cogniverse_foundation.config.utils import create_default_config_manager
    from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
    from pathlib import Path

    config_manager = create_default_config_manager()
    schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
    agent = SearchAgent(
        deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
        config_manager=config_manager,
        schema_loader=schema_loader,
    )

    # Agent's search service should target tenant-specific schema
    # tenant_id is per-request on search_by_text()
    results = agent.search_by_text(
        query="test",
        tenant_id="acme_corp",
        top_k=5,
    )

def test_tenant_phoenix_project_isolation():
    """Verify Phoenix projects can be registered per-tenant"""
    from cogniverse_foundation.telemetry.manager import TelemetryManager

    # TelemetryManager is a singleton, configure per-tenant projects
    telemetry = TelemetryManager()

    # Register tenant-specific projects
    telemetry.register_project(tenant_id="acme_corp", project_name="acme-project")
    telemetry.register_project(tenant_id="globex_inc", project_name="globex-project")

    # Verify registration succeeded (projects are stored internally)
    # Actual project isolation is enforced when creating spans
    assert telemetry is not None

Multi-Tenant Data Isolation

Test Cross-Tenant Data Boundaries:

# tests/integration/test_multi_tenant_isolation.py
import pytest
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_runtime.ingestion.pipeline import VideoIngestionPipeline
from cogniverse_core.registries.backend_registry import BackendRegistry
from cogniverse_vespa.backend import VespaBackend

@pytest.mark.integration
async def test_tenant_data_isolation(sample_video, config_manager, schema_loader):
    """Verify tenants cannot access each other's data"""
    from cogniverse_runtime.ingestion.pipeline import PipelineConfig

    # Create pipeline config for tenant A
    pipeline_config_a = PipelineConfig(
        extract_keyframes=True,
        generate_embeddings=True,
        search_backend="vespa"
    )

    # Ingest video for tenant A
    pipeline_a = VideoIngestionPipeline(
        tenant_id="acme_corp",
        config=pipeline_config_a,
        config_manager=config_manager,
        schema_loader=schema_loader
    )
    result_a = await pipeline_a.process_video_async(sample_video)
    assert result_a["status"] == "completed"

    # Search as tenant B (should get no results from tenant A)
    backend = BackendRegistry.get_search_backend(
        name="vespa",
        config_manager=config_manager,
        schema_loader=schema_loader
    )

    results = backend.search(
        query_dict={
            "query": "test",
            "type": "video",
            "profile": "video_colpali_smol500_mv_frame",
            "tenant_id": "globex_inc",
        }
    )

    # Tenant B should not see tenant A's documents
    assert len(results) == 0

@pytest.mark.integration
def test_tenant_memory_isolation(config_manager, schema_loader):
    """Verify tenant memories are isolated"""
    from cogniverse_core.memory.manager import Mem0MemoryManager

    memory_a = Mem0MemoryManager(tenant_id="acme_corp")
    memory_a.initialize(
        backend_host="localhost",
        backend_port=8080,
        llm_model="openai/google/gemma-4-e4b-it",
        embedding_model="lightonai/DenseOn",
        llm_base_url="http://localhost:11434",
        embedder_base_url="http://localhost:11434",
        config_manager=config_manager,
        schema_loader=schema_loader,
    )

    memory_b = Mem0MemoryManager(tenant_id="globex_inc")
    memory_b.initialize(
        backend_host="localhost",
        backend_port=8080,
        llm_model="openai/google/gemma-4-e4b-it",
        embedding_model="lightonai/DenseOn",
        llm_base_url="http://localhost:11434",
        embedder_base_url="http://localhost:11434",
        config_manager=config_manager,
        schema_loader=schema_loader,
    )

    # Add memory for tenant A
    memory_a.add_memory(
        content="Secret message",
        tenant_id="acme_corp",
        agent_name="test_agent"
    )

    # Search as tenant B (should not find tenant A's memory)
    results_b = memory_b.search_memory(
        query="Secret message",
        tenant_id="globex_inc",
        agent_name="test_agent",
        top_k=5
    )

    assert len(results_b) == 0

Tenant-Aware Fixtures

Example Patterns (not in conftest.py; create as needed for multi-tenant tests):

Tenant isolation in Cogniverse is enforced at the operation level (schema deployment, search, memory add) via a tenant_id string argument. SystemConfig carries backend connection details (URL, port) shared across all tenants — it does not hold a tenant_id field.

import pytest
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps
from cogniverse_core.registries.agent_registry import AgentRegistry

@pytest.fixture
def backend_config():
    """Shared backend config (same Vespa endpoint for all tenants)."""
    return SystemConfig(
        backend_url="http://localhost",
        backend_port=8080,
    )

@pytest.fixture
def multi_tenant_ids():
    """Tenant ID strings for cross-tenant tests."""
    return {
        "tenant_a": "acme_corp",
        "tenant_b": "globex_inc",
    }

@pytest.fixture
def tenant_agent():
    """Create orchestrator agent (tenant_id is passed per-operation)."""
    from cogniverse_foundation.config.manager import ConfigManager
    from cogniverse_foundation.config.unified_config import LLMEndpointConfig
    from cogniverse_foundation.telemetry.config import TelemetryConfig
    from tests.utils.memory_store import InMemoryConfigStore

    store = InMemoryConfigStore()
    store.initialize()
    config_manager = ConfigManager(store=store)
    deps = OrchestratorDeps(
        telemetry_config=TelemetryConfig(),
        llm_config=LLMEndpointConfig(
            model="openai/google/gemma-4-e4b-it",
            api_base="http://localhost:11434/v1",
        ),
    )
    registry = AgentRegistry(tenant_id="test-tenant", config_manager=config_manager)
    return OrchestratorAgent(deps=deps, registry=registry, config_manager=config_manager)

# Use in tests:
def test_with_tenant_fixtures(multi_tenant_ids, tenant_agent):
    assert tenant_agent.deps is not None
    assert multi_tenant_ids["tenant_a"] != multi_tenant_ids["tenant_b"]

Test Isolation

State Cleanup Between Tests

Tests use auto-fixtures to clean up state:

@pytest.fixture(autouse=True, scope="function")
def cleanup_dspy_state():
    """Clean up DSPy state between tests to prevent isolation issues"""
    yield

    # Clean up any DSPy state after each test
    try:
        import dspy

        # Reset ALL DSPy settings attributes to prevent any state pollution
        if hasattr(dspy, "settings"):
            if hasattr(dspy.settings, "lm"):
                dspy.settings.lm = None
            if hasattr(dspy.settings, "adapter"):
                dspy.settings.adapter = None
            if hasattr(dspy.settings, "rm"):
                dspy.settings.rm = None
            if hasattr(dspy.settings, "experimental"):
                dspy.settings.experimental = False

        # Clear any context stack from async tests
        if hasattr(dspy, "_context_stack"):
            if hasattr(dspy._context_stack, "clear"):
                dspy._context_stack.clear()
            elif isinstance(dspy._context_stack, list):
                dspy._context_stack.clear()

    except (ImportError, AttributeError, RuntimeError):
        pass

    cleanup_background_threads()

Best Practices

  1. Always use fixtures for shared state
  2. Clean up resources in fixture teardown
  3. Don't rely on test execution order - tests should be independent
  4. Use unique IDs for multi-tenant tests (tenant_id, user_id)

Test Config Isolation (cogniverse_test_config)

tests/conftest.py ships a session-scoped autouse=True fixture named cogniverse_test_config. It isolates configuration mutations from the tracked configs/config.json; it does not change model identity or silently substitute a smaller model.

What it does:

  1. Clones configs/config.json unchanged into a tmpdir and links its sibling schemas/ directory so relative schema discovery still works.
  2. Sets COGNIVERSE_CONFIG=<tmpdir>/config.json for the session.
  3. Tests that request ensure_host_ollama then require the production Gemma primary model and, when marked, the distinct teacher model, both served on Modal. The resolver accepts an endpoint only when its authenticated /v1/models lists the exact model and revision, and publishes the verified endpoints in a second temporary config. Nothing is started on this host; see "Model endpoints" in docs/testing/TESTING_GUIDE.md.

Override the Modal lookup via env vars before pytest starts:

INFERENCE_SERVICE_URLS='{"vllm_llm_student":"https://my-gemma.example"}' \
uv run pytest

Skip it by setting COGNIVERSE_CONFIG yourself (e.g. CI matrix):

COGNIVERSE_CONFIG=/path/to/ci-config.json uv run pytest

If your unit test exercises ConfigUtils' cwd-based auto-discovery (chdir to a tmp dir that has its own configs/config.json), clear COGNIVERSE_CONFIG first — otherwise ConfigUtils honours the env var and bypasses the chdir:

def test_something(self, monkeypatch):
    monkeypatch.delenv("COGNIVERSE_CONFIG", raising=False)
    monkeypatch.chdir(tmp_path)
    # ...

tests/utils/llm_config.py honours the same COGNIVERSE_CONFIG priority: env var first, then configs/config.json. Use get_llm_model() / get_llm_base_url() from there in test code so the autouse fixture's overrides flow through.


Performance

Test Execution Time

Optimize test performance:

# Parallel execution (be careful with async tests) — requires pytest-xdist,
# not a project dependency by default: `uv pip install pytest-xdist` first
uv run pytest -n auto

# Show the 10 slowest test/setup durations after the run (built into pytest core)
uv run pytest --durations=10

Model Caching

Local model loaders cache loaded weights across calls to speed up tests. get_or_load_model in libs/core/cogniverse_core/common/models/model_loaders.py keys the module-level _model_cache dict by model_name (or model_name@remote_inference_url when remote inference is configured), holds a per-key threading.Lock so concurrent from_pretrained calls can't corrupt PyTorch/accelerate's meta-tensor dispatch state, and evicts a cached entry whose parameters have degraded to a meta device before re-loading:

def get_or_load_model(
    model_name: str,
    config: Dict[str, Any],
    logger: Optional[logging.Logger] = None,
    force_reload: bool = False,
) -> Tuple[Any, Any]:
    """Get model from cache or load it. Thread-safe via a per-key lock."""
    ...

Debugging Tests

Debug Output

# Show print statements
uv run pytest -s

# Show full error traceback
uv run pytest --tb=long

# Drop into debugger on failure
uv run pytest --pdb

Logging

Enable detailed logging:

import logging
logging.basicConfig(level=logging.DEBUG)

# Or in tests
@pytest.fixture
def debug_logging():
    logging.getLogger().setLevel(logging.DEBUG)

CI/CD Considerations

GitHub Actions Configuration

- name: Run Tests
  env:
    JAX_PLATFORM_NAME: cpu
    TOKENIZERS_PARALLELISM: false
    OMP_NUM_THREADS: 1
  run: |
    timeout 7200 uv run pytest -v --tb=long

Docker Testing

# Ensure single-threaded mode
ENV TOKENIZERS_PARALLELISM=false
ENV OMP_NUM_THREADS=1
ENV MKL_NUM_THREADS=1

RUN pytest

Docker Fixtures for Integration Tests

Integration tests use self-managed Docker containers via pytest fixtures.

shared_vespa — Single Session-Scoped Vespa Container

The project uses a single session-scoped shared_vespa fixture defined in tests/conftest.py. All integration tests share this one Vespa container; isolation between tests comes from unique tenant IDs rather than separate containers. This eliminates the RAM pressure that caused OOM-kills when every package spawned its own container.

shared_vespa yields a dict:

{
    "http_port": <int>,       # Vespa data port
    "config_port": <int>,     # Vespa config-server port
    "container_name": <str>,
    "base_url": "http://localhost:<http_port>",
}

Per-package compatibility shims

Each per-package conftest.py re-exports shared_vespa and provides a thin shim under the fixture name that tests in that package already reference:

# e.g. tests/backends/integration/conftest.py
from tests.conftest import shared_vespa  # noqa: F401

@pytest.fixture(scope="module")
def vespa_instance(shared_vespa):
    """Shim: exposes shared_vespa under the name backends tests expect."""
    yield {
        "http_port": shared_vespa["http_port"],
        "config_port": shared_vespa["config_port"],
        "base_url": shared_vespa["base_url"],
        "container_name": shared_vespa["container_name"],
    }
    # No teardown — shared_vespa owns the container lifecycle.

Tenant isolation helpers

tests/utils/tenant_helpers.py provides two functions for deriving Vespa-safe unique tenant IDs from the running test's request object:

  • tenant_id_for_module(request) — all tests in the same module share a tenant (use when a fixture is module-scoped).
  • tenant_id_for_test(request) — unique per test function, including parametrize variants (use when each test needs a blank-slate schema).

tests/utils/vespa_test_helpers.py provides helpers for building config managers and deploying schemas against shared_vespa:

  • make_config_manager(shared_vespa) — returns a ConfigManager bound to the shared container's ports.
  • deploy_tenant_schema(shared_vespa, *, tenant_id, base_schema_name) — deploys a base schema scoped to tenant_id; returns the full schema name.
  • schema_full_name(base_schema_name, tenant_id) — computes the deployed schema name without going through deploy.
  • load_raw_schema_json(base_schema_name) — reads a base schema definition from configs/schemas/.

Example fixture using the helpers:

# Usage in tests
@pytest.mark.integration
async def test_vespa_operations(shared_vespa, request):
    """Test with the shared Vespa container."""
    from tests.utils.tenant_helpers import tenant_id_for_test
    from tests.utils.vespa_test_helpers import deploy_tenant_schema, make_config_manager

    tenant_id = tenant_id_for_test(request)
    config_manager = make_config_manager(shared_vespa)
    deploy_tenant_schema(
        shared_vespa,
        tenant_id=tenant_id,
        base_schema_name="video_colpali_smol500_mv_frame",
        config_manager=config_manager,
    )
    http_port = shared_vespa["http_port"]
    base_url = shared_vespa["base_url"]
    # ... test operations

Phoenix Docker Fixtures

For telemetry and evaluation tests:

The phoenix_container fixture is defined in tests/conftest.py (module-scoped). Key details:

  • Image: arizephoenix/phoenix:20.16.0 pinned by digest (never :latest)
  • Ports: HTTP 16006 + per-process offset (→ 6006), gRPC 14317 + per-process offset (→ 4317) — port_offset = (os.getpid() % 1000) * 10 so concurrent pytest sweeps don't collide
  • Env var: Sets TELEMETRY_OTLP_ENDPOINT (not OTLP_ENDPOINT) and TELEMETRY_SYNC_EXPORT
  • Leftover cleanup: Kills only its own process's leftover phoenix_test_pid<pid>_* containers from a prior crashed run — never touches another concurrent sweep's containers
  • Scope: Module-scoped — one Phoenix instance per test module

Test containers carry cogniverse-test-owner-pid=<pytest pid>. When every pytest session starts, the shared reaper removes containers whose owner process is gone and already-exited containers; the session's test sidecars summary lists what it removed, or why docker could not be asked. A live owner's created, restarting, paused, or running container is preserved: docker create and docker start are separate operations, so treating the short created interval as abandoned can kill another concurrent test's service before it starts.

Companion fixtures also defined in tests/conftest.py:

  • phoenix_client — phoenix.client.Client pointed at the container's http_endpoint (http://localhost:16006 plus the per-process port offset)
  • telemetry_config_with_phoenix — TelemetryConfig pre-configured for the container
  • telemetry_manager_with_phoenix — session-installed TelemetryManager singleton
  • telemetry_manager_without_phoenix — function-scoped manager with mock endpoints (use for unit/integration tests that don't export real spans)

Port Management

Vespa test containers get their host ports when they start, in tests/utils/docker_utils.py:

  • start_docker_container_with_port_retry(module_name, name_prefix=..., image=..., container_ports=(8080, 19071), ...) picks a free (http_port, config_port) pair with generate_unique_ports (config_port is always http_port + 10991, the offset the runtime re-derives), runs the container under a name unique to the process, thread and port, and removes the container a failed start created. Docker's bind-conflict errors (address already in use, port is already allocated) are retried on a fresh pair; other failures raise with Docker's stderr, and exhausted retries raise Docker container allocation failed after N attempts.
  • VespaDockerManager.start_container(module_name) in tests/utils/vespa_docker.py (and VespaTestManager, which uses it) starts the Vespa image that way and returns the ports it got.

generate_unique_ports only probes: another process can take a probed port before Docker binds it. Never compute ports at import time or pass a pre-chosen pair to a container start; read the ports from the started container.

CI Disk Space Requirements

Vespa requires disk usage below 75%. In GitHub Actions:

- name: Free up disk space for Vespa
  run: |
    # Remove ~30GB of unused packages
    sudo rm -rf /usr/share/dotnet           # .NET SDK
    sudo rm -rf /usr/local/lib/android      # Android SDK
    sudo rm -rf /opt/ghc                    # Haskell
    sudo rm -rf /opt/hostedtoolcache/CodeQL # CodeQL
    sudo docker image prune -af
    df -h  # Verify disk usage

- name: Pre-pull Vespa Docker image
  run: docker pull vespaengine/vespa:8.668.5

Summary

This guide covers comprehensive testing for Cogniverse multi-agent system:

  1. Async Testing: Thread-safe configuration for async tests with ML models
  2. SDK Package Testing: Verify cross-package imports and dependencies
  3. Multi-Tenant Testing: Ensure tenant isolation at all layers
  4. Test Execution: Use uv run pytest for workspace packages
  5. Performance: Model caching and parallel execution strategies
  6. CI/CD: Configuration for automated testing

Key Testing Principles:

  • Always use uv run pytest to activate workspace

  • Test tenant isolation at schema, project, and memory levels

  • Use fixtures for reusable tenant configurations

  • Verify SDK package imports work correctly

  • Maintain test independence (no shared state)

  • Clean up background threads to avoid segfaults

Test Organization by Package:

graph TD
    root["<span style='color:#000'><b>tests/</b></span>"]

    admin["<span style='color:#000'><b>admin/</b><br/>profile & tenant management tests</span>"]
    admin_unit["<span style='color:#000'>unit/</span>"]

    agents["<span style='color:#000'><b>agents/</b><br/>cogniverse_agents tests</span>"]
    agents_unit["<span style='color:#000'>unit/</span>"]
    agents_int["<span style='color:#000'>integration/</span>"]
    agents_e2e["<span style='color:#000'>e2e/</span>"]

    backends["<span style='color:#000'><b>backends/</b><br/>cogniverse_vespa tests</span>"]
    backends_unit["<span style='color:#000'>unit/</span>"]
    backends_int["<span style='color:#000'>integration/</span>"]

    charts["<span style='color:#000'><b>charts/</b><br/>Helm chart tests</span>"]

    cli["<span style='color:#000'><b>cli/</b><br/>cogniverse_cli tests</span>"]
    cli_unit["<span style='color:#000'>unit/</span>"]

    common["<span style='color:#000'><b>common/</b><br/>cogniverse_core.common tests</span>"]
    common_unit["<span style='color:#000'>unit/</span>"]
    common_int["<span style='color:#000'>integration/</span>"]

    core["<span style='color:#000'><b>core/</b><br/>cogniverse_core tests</span>"]
    core_unit["<span style='color:#000'>unit/</span>"]
    core_int["<span style='color:#000'>integration/</span>"]

    e2e["<span style='color:#000'><b>e2e/</b><br/>cross-package end-to-end tests</span>"]
    e2e_deployment["<span style='color:#000'>deployment/</span>"]

    evaluation["<span style='color:#000'><b>evaluation/</b><br/>cogniverse_evaluation tests</span>"]
    evaluation_unit["<span style='color:#000'>unit/</span>"]
    evaluation_int["<span style='color:#000'>integration/</span>"]

    events["<span style='color:#000'><b>events/</b><br/>event system tests</span>"]
    events_unit["<span style='color:#000'>unit/</span>"]
    events_int["<span style='color:#000'>integration/</span>"]

    finetuning["<span style='color:#000'><b>finetuning/</b><br/>fine-tuning pipeline tests</span>"]
    finetuning_int["<span style='color:#000'>integration/</span>"]

    fixtures["<span style='color:#000'><b>fixtures/</b><br/>shared LLM/sidecar fixtures</span>"]

    foundation["<span style='color:#000'><b>foundation/</b><br/>cogniverse_foundation tests</span>"]
    foundation_unit["<span style='color:#000'>unit/</span>"]
    foundation_int["<span style='color:#000'>integration/</span>"]

    ingestion["<span style='color:#000'><b>ingestion/</b><br/>cogniverse_runtime.ingestion tests</span>"]
    ingestion_unit["<span style='color:#000'>unit/</span>"]
    ingestion_int["<span style='color:#000'>integration/</span>"]

    memory["<span style='color:#000'><b>memory/</b><br/>memory system tests</span>"]
    memory_unit["<span style='color:#000'>unit/</span>"]
    memory_int["<span style='color:#000'>integration/</span>"]

    messaging["<span style='color:#000'><b>messaging/</b><br/>cogniverse_messaging tests</span>"]
    messaging_unit["<span style='color:#000'>unit/</span>"]
    messaging_int["<span style='color:#000'>integration/</span>"]

    routing["<span style='color:#000'><b>routing/</b><br/>routing-specific tests</span>"]
    routing_unit["<span style='color:#000'>unit/</span>"]
    routing_int["<span style='color:#000'>integration/</span>"]

    runtime["<span style='color:#000'><b>runtime/</b><br/>cogniverse_runtime tests</span>"]
    runtime_unit["<span style='color:#000'>unit/</span>"]
    runtime_int["<span style='color:#000'>integration/</span>"]

    synthetic["<span style='color:#000'><b>synthetic/</b><br/>synthetic data tests</span>"]
    synthetic_int["<span style='color:#000'>integration/</span>"]

    system["<span style='color:#000'><b>system/</b><br/>end-to-end system tests</span>"]

    telemetry["<span style='color:#000'><b>telemetry/</b><br/>cogniverse_foundation.telemetry tests</span>"]
    telemetry_unit["<span style='color:#000'>unit/</span>"]
    telemetry_int["<span style='color:#000'>integration/</span>"]

    utils["<span style='color:#000'><b>utils/</b><br/>shared test utilities</span>"]

    root --> admin
    root --> agents
    root --> backends
    root --> charts
    root --> cli
    root --> common
    root --> core
    root --> e2e
    root --> evaluation
    root --> events
    root --> finetuning
    root --> fixtures
    root --> foundation
    root --> ingestion
    root --> memory
    root --> messaging
    root --> routing
    root --> runtime
    root --> synthetic
    root --> system
    root --> telemetry
    root --> utils

    admin --> admin_unit

    e2e --> e2e_deployment

    agents --> agents_unit
    agents --> agents_int
    agents --> agents_e2e

    backends --> backends_unit
    backends --> backends_int

    cli --> cli_unit

    common --> common_unit
    common --> common_int

    core --> core_unit
    core --> core_int

    evaluation --> evaluation_unit
    evaluation --> evaluation_int

    events --> events_unit
    events --> events_int

    finetuning --> finetuning_int

    foundation --> foundation_unit
    foundation --> foundation_int

    ingestion --> ingestion_unit
    ingestion --> ingestion_int

    messaging --> messaging_unit
    messaging --> messaging_int

    runtime --> runtime_unit
    runtime --> runtime_int

    memory --> memory_unit
    memory --> memory_int

    routing --> routing_unit
    routing --> routing_int

    synthetic --> synthetic_int

    telemetry --> telemetry_unit
    telemetry --> telemetry_int

    style root fill:#b0bec5,stroke:#546e7a,color:#000
    style admin fill:#b0bec5,stroke:#546e7a,color:#000
    style agents fill:#ce93d8,stroke:#7b1fa2,color:#000
    style backends fill:#90caf9,stroke:#1565c0,color:#000
    style charts fill:#b0bec5,stroke:#546e7a,color:#000
    style cli fill:#b0bec5,stroke:#546e7a,color:#000
    style common fill:#a5d6a7,stroke:#388e3c,color:#000
    style core fill:#ce93d8,stroke:#7b1fa2,color:#000
    style e2e fill:#b0bec5,stroke:#546e7a,color:#000
    style evaluation fill:#ffcc80,stroke:#ef6c00,color:#000
    style events fill:#ffcc80,stroke:#ef6c00,color:#000
    style finetuning fill:#ffcc80,stroke:#ef6c00,color:#000
    style fixtures fill:#b0bec5,stroke:#546e7a,color:#000
    style foundation fill:#a5d6a7,stroke:#388e3c,color:#000
    style ingestion fill:#ffcc80,stroke:#ef6c00,color:#000
    style memory fill:#90caf9,stroke:#1565c0,color:#000
    style messaging fill:#90caf9,stroke:#1565c0,color:#000
    style routing fill:#ce93d8,stroke:#7b1fa2,color:#000
    style runtime fill:#90caf9,stroke:#1565c0,color:#000
    style synthetic fill:#ffcc80,stroke:#ef6c00,color:#000
    style system fill:#b0bec5,stroke:#546e7a,color:#000
    style telemetry fill:#a5d6a7,stroke:#388e3c,color:#000
    style utils fill:#b0bec5,stroke:#546e7a,color:#000
    style admin_unit fill:#b0bec5,stroke:#546e7a,color:#000
    style e2e_deployment fill:#b0bec5,stroke:#546e7a,color:#000
    style agents_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
    style agents_int fill:#ba68c8,stroke:#7b1fa2,color:#000
    style agents_e2e fill:#ba68c8,stroke:#7b1fa2,color:#000
    style backends_unit fill:#64b5f6,stroke:#1565c0,color:#000
    style backends_int fill:#64b5f6,stroke:#1565c0,color:#000
    style cli_unit fill:#b0bec5,stroke:#546e7a,color:#000
    style common_unit fill:#81c784,stroke:#388e3c,color:#000
    style common_int fill:#81c784,stroke:#388e3c,color:#000
    style core_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
    style core_int fill:#ba68c8,stroke:#7b1fa2,color:#000
    style evaluation_unit fill:#ffb74d,stroke:#ef6c00,color:#000
    style evaluation_int fill:#ffb74d,stroke:#ef6c00,color:#000
    style events_unit fill:#ffb74d,stroke:#ef6c00,color:#000
    style events_int fill:#ffb74d,stroke:#ef6c00,color:#000
    style finetuning_int fill:#ffb74d,stroke:#ef6c00,color:#000
    style foundation_unit fill:#81c784,stroke:#388e3c,color:#000
    style foundation_int fill:#81c784,stroke:#388e3c,color:#000
    style ingestion_unit fill:#ffb74d,stroke:#ef6c00,color:#000
    style ingestion_int fill:#ffb74d,stroke:#ef6c00,color:#000
    style memory_unit fill:#64b5f6,stroke:#1565c0,color:#000
    style memory_int fill:#64b5f6,stroke:#1565c0,color:#000
    style messaging_unit fill:#64b5f6,stroke:#1565c0,color:#000
    style messaging_int fill:#64b5f6,stroke:#1565c0,color:#000
    style routing_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
    style routing_int fill:#ba68c8,stroke:#7b1fa2,color:#000
    style runtime_unit fill:#64b5f6,stroke:#1565c0,color:#000
    style runtime_int fill:#64b5f6,stroke:#1565c0,color:#000
    style synthetic_int fill:#ffb74d,stroke:#ef6c00,color:#000
    style telemetry_unit fill:#81c784,stroke:#388e3c,color:#000
    style telemetry_int fill:#81c784,stroke:#388e3c,color:#000

Related Documentation:


References