Pytest Best Practices¶
Async Testing Configuration¶
Threading Issues with Async Tests¶
Async tests in pytest can encounter threading conflicts when loading ML models. The codebase uses several libraries that spawn background threads:
- tqdm (from transformers): Progress bars during model downloads
- posthog (from mem0ai): Telemetry/analytics background threads
- torch: Multi-threaded tensor operations
These background threads can cause segmentation faults during pytest cleanup when combined with async event loops.
Solution: Single-Threaded Mode¶
The test suite is configured for single-threaded execution to avoid threading conflicts:
File: tests/conftest.py
import os
# Configure torch and tokenizers to avoid threading issues
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["OMP_NUM_THREADS"] = "1"
os.environ["MKL_NUM_THREADS"] = "1"
# Import torch and configure threading before any tests run
try:
import torch
torch.set_num_threads(1)
except ImportError:
pass
File: pytest.ini
[pytest]
asyncio_mode = auto
asyncio_default_fixture_loop_scope = function
asyncio_default_test_loop_scope = function
asyncio_default_test_loop_scope pins the test loop to function scope too — without it, pytest-asyncio 1.3+ defaults the test loop to module/session scope, so an async test that leaves the loop "running" (e.g. a backgrounded coroutine that never fully awaited) makes every subsequent async test in the same scope fail with Runner.run() cannot be called from a running event loop. Function-scoped loops get torn down per-test, isolating leaks.
Background Thread Cleanup¶
The conftest also includes automatic cleanup for background threads:
from tests.utils.async_polling import simulate_processing_delay
def cleanup_background_threads():
"""
Clean up background threads from tqdm (transformers) and posthog (mem0ai).
These libraries create daemon threads that can cause segfaults during pytest
cleanup in async tests. We need to give them time to finish and exit cleanly.
"""
max_wait = 2.0 # seconds
start_time = time.time()
while time.time() - start_time < max_wait:
background_threads = [
t
for t in threading.enumerate()
if t != threading.current_thread()
and t.daemon
and any(name in t.name.lower() for name in ["tqdm", "posthog", "monitor"])
]
if not background_threads:
break
# Give threads time to finish their work
simulate_processing_delay(delay=0.1, description="test processing")
# Force garbage collection to clean up any remaining references
gc.collect()
Running Tests¶
Basic Test Execution¶
# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest
# Run specific test file with 30-minute timeout
JAX_PLATFORM_NAME=cpu timeout 1800 uv run pytest tests/agents/test_routing.py
# Run with verbose output
JAX_PLATFORM_NAME=cpu uv run pytest -xvs
Test Markers¶
# Run only unit tests
uv run pytest -m unit
# Run only integration tests
uv run pytest -m integration
# Run fast CI tests (subset for quick feedback)
uv run pytest -m ci_fast
# Skip slow tests
uv run pytest -m "not slow"
# Run tests declaring an exact inference service
uv run pytest -m requires_inference
uv run pytest -m requires_ollama
# Skip tests requiring Ollama (for CI)
uv run pytest -m "not requires_ollama"
Available Markers¶
Defined in pytest.ini:
[pytest]
markers =
unit: Unit tests for individual components
integration: Integration tests with multiple components
slow: Slow tests that take significant time
requires_ollama: Tests that require Ollama to be running
requires_lm: Tests that require the configured test LM endpoint
requires_gliner: Tests that require GLiNER models
phoenix: Tests requiring Phoenix
inspect: Tests requiring Inspect AI
ragas: Tests requiring RAGAS
asyncio: Async tests
local_only: Tests that should only run locally (not in CI/CD)
requires_models: Tests that require actual ML models to be available
benchmark: Performance benchmarking tests
ingestion: Tests for ingestion pipeline
requires_vespa: Tests that require Vespa backend to be running
no_shared_vespa: Integration modules that own a non-Vespa boundary
requires_docker: Tests that require Docker
requires_gpu: Tests that require GPU availability
requires_whisper: Tests that require Whisper models
requires_teacher_model: Tests that scale up the vllm-llm-teacher pod (long-running, off by default)
requires_optimizer_data: Tests that exercise non-router optimizers (workflow / modality / xgboost) end-to-end against the live cluster (slow, off by default)
requires_cv2: Tests that require OpenCV
requires_ffmpeg: Tests that require FFmpeg
ci_safe: Tests that are safe to run in CI environment
ci_fast: Fast, essential tests for CI (subset of most important functionality)
timeout: Tests with custom timeout values
e2e: End-to-end integration tests with real services
e2e_heavy: Heavy end-to-end tests that submit a full CronWorkflow execution and may take 5-10+ minutes (DSPy optimization, synthetic data generation, distillation); opt-in via -m e2e_heavy
browser: Browser-based E2E tests requiring Playwright
system: System-level end-to-end tests with full infrastructure
telemetry: Tests for telemetry and observability system
Exact inference markers are registered by tests/fixtures/inference.py, not pytest.ini. Name the canonical service in the decorator:
@pytest.mark.requires_inference("vllm_colpali")
def test_colpali_or_colqwen_boundary():
...
@pytest.mark.requires_inference("video_embed")
def test_xclip_boundary():
...
ColPali and ColQwen use the same vllm_colpali HTTP service. The collector also requests every service a shipped profile using the named embedding service resolves at pipeline init, derived from configs/config.json (vllm_asr for vllm_colpali, video_embed and colbert_pylate). Use @pytest.mark.requires_modal_inference("vllm_llm_student") only when the test explicitly requires that exact service through paid Modal provisioning.
The ci_fast Marker¶
The ci_fast marker identifies tests run in the CI Fast subset steps across workflows:
# tests/backends/integration/test_tenant_schema_lifecycle.py
@pytest.mark.integration
@pytest.mark.ci_fast
class TestSchemaRegistryDeployment:
"""Test schema deployment via SchemaRegistry"""
def test_deploy_single_schema(self, get_backend):
"""Test deploying a single schema for a tenant"""
backend = get_backend("acme")
backend.schema_registry.deploy_schema("acme", "video_colpali_smol500_mv_frame")
schemas = backend.schema_registry.get_tenant_schemas("acme")
assert len(schemas) == 1
assert schemas[0].base_schema_name == "video_colpali_smol500_mv_frame"
Guidelines for ci_fast tests:
-
Essential functionality that must work
-
Complete in under 2 minutes
-
No external API calls (Ollama, OpenAI, etc.)
-
Can use Docker containers (Vespa, Phoenix)
Async Test Timeout¶
Some async tests carry a @pytest.mark.timeout(N) marker (registered in pytest.ini as timeout: Tests with custom timeout values) with a literal per-test second count:
# tests/routing/integration/test_deep_research_integration.py
@pytest.mark.asyncio
@pytest.mark.timeout(120)
async def test_full_research_cycle(self, real_search_fn, seeded_outdoor_corpus):
"""Decompose -> search Vespa -> evaluate -> synthesize against the configured LM."""
...
Note the pytest-timeout plugin that enforces this marker is not a project dependency — install it (uv pip install pytest-timeout) for the marker to actually abort a hung test; without it, the marker is inert and only the shell-level timeout wrapper below bounds runtime.
Conventional shell-level timeouts used across the project's own test commands and CI workflows:
-
Individual test files: 30 minutes (
timeout 1800) -
Full test suite: 120 minutes (
timeout 7200)
Common Issues¶
Segmentation Faults in Async Tests¶
Symptoms:
Fatal Python error: Segmentation fault
Thread 0x000000033614f000 (most recent call first):
File "/path/to/threading.py", line 359 in wait
Cause: Threading conflict between pytest async event loop and background threads from model loading.
Solution: Already handled by tests/conftest.py configuration. If you still see segfaults:
- Check that tests use
@pytest.mark.asynciocorrectly - Verify
pytest.inihasasyncio_mode = auto - Ensure no manual thread creation in test code
Model Loading Errors¶
Symptoms:
Cause: Locally downloading and loading a large embedding/VLM model inside the pytest process can trigger the same background-thread conflicts described above.
Solution: The primary embedding model (TomoroAI/tomoro-colqwen3-embed-4b, configured as colpali_model / per-profile embedding_model in configs/config.json) has no in-process loader — ColPaliModelLoader and ColQwenModelLoader (libs/core/cogniverse_core/common/models/model_loaders.py) raise immediately for any ColQwen3/Tomoro model name, directing callers to RemoteColPaliLoader instead:
ColQwen3/Tomoro models are remote-only — serve via vLLM and set
inference_service_url (profile inference_services.embedding). Local
in-process loading is unsupported (requires transformers>=4.57, blocked
by the pylate cap).
Tests marked requires_inference("vllm_colpali") therefore exercise ColPali and ColQwen through that exact service over HTTP rather than downloading weights into the test process — avoid adding a local from_pretrained(...) call for these models in new tests.
Import Timing Issues¶
Symptoms:
Cause: UV workspace not synced or tests run outside virtual environment.
Solution: Always use uv run pytest which ensures workspace packages are available.
Example Fix:
# ❌ Bad: Direct pytest without uv
pytest tests/agents/
# ✅ Good: Use uv run to activate workspace
uv run pytest tests/agents/
Package Import Patterns:
# ✅ Good: Absolute imports from workspace packages
from cogniverse_foundation.telemetry.manager import TelemetryManager
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent
from cogniverse_vespa.backend import VespaBackend
# ❌ Bad: Old src-style imports (not valid in uv workspace)
from src.agents.orchestrator_agent import OrchestratorAgent # ❌ Will fail
Package Testing¶
Testing Package Imports¶
Example Test Patterns - Verify Package Structure:
# Example: tests/test_imports.py (create this file to verify imports)
import pytest
def test_sdk_package_imports():
"""Verify cogniverse_sdk package imports work"""
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
from cogniverse_sdk import SearchResult, Document
assert FilesystemSchemaLoader is not None
assert SearchResult is not None
assert Document is not None
def test_foundation_package_imports():
"""Verify cogniverse_foundation package imports work"""
from cogniverse_foundation.telemetry.manager import TelemetryManager
assert TelemetryManager is not None
def test_core_package_imports():
"""Verify cogniverse_core package imports work"""
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_core.registries.backend_registry import BackendRegistry
assert SystemConfig is not None
assert BackendRegistry is not None
def test_agents_package_imports():
"""Verify cogniverse_agents package imports work"""
from cogniverse_agents.orchestrator_agent import OrchestratorAgent
from cogniverse_agents.search_agent import SearchAgent
assert OrchestratorAgent is not None
assert SearchAgent is not None
def test_retrieval_package_imports():
"""Verify cogniverse_vespa package imports work"""
from cogniverse_vespa.backend import VespaBackend
from cogniverse_vespa.vespa_schema_manager import VespaSchemaManager
assert VespaBackend is not None
assert VespaSchemaManager is not None
def test_processing_package_imports():
"""Verify cogniverse_runtime.ingestion package imports work"""
from cogniverse_runtime.ingestion.pipeline import VideoIngestionPipeline
from cogniverse_runtime.ingestion.pipeline_builder import VideoIngestionPipelineBuilder
assert VideoIngestionPipeline is not None
assert VideoIngestionPipelineBuilder is not None
def test_evaluation_package_imports():
"""Verify cogniverse_evaluation package imports work"""
from cogniverse_evaluation.core.experiment_tracker import ExperimentTracker
from cogniverse_evaluation.metrics import calculate_mrr, calculate_ndcg
assert ExperimentTracker is not None
assert calculate_mrr is not None
assert calculate_ndcg is not None
Package Dependency Testing¶
Example Test Patterns - Cross-Package Dependencies:
# Example: tests/test_package_dependencies.py (create this file to verify dependencies)
import pytest
from cogniverse_foundation.telemetry.manager import TelemetryManager
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent
def test_agents_depends_on_foundation_and_core():
"""Verify agents package can use foundation and core packages"""
from cogniverse_agents.orchestrator_agent import OrchestratorDeps
from cogniverse_core.registries.agent_registry import AgentRegistry
from cogniverse_foundation.config.manager import ConfigManager
from cogniverse_foundation.config.unified_config import LLMEndpointConfig
from cogniverse_foundation.telemetry.config import TelemetryConfig
from tests.utils.memory_store import InMemoryConfigStore
telemetry_config = TelemetryConfig()
deps = OrchestratorDeps(
telemetry_config=telemetry_config,
llm_config=LLMEndpointConfig(
model="openai/google/gemma-4-e4b-it",
api_base="http://localhost:11434/v1",
),
)
# OrchestratorAgent requires deps, registry, and config_manager
store = InMemoryConfigStore()
store.initialize()
config_manager = ConfigManager(store=store)
registry = AgentRegistry(tenant_id="test-tenant", config_manager=config_manager)
agent = OrchestratorAgent(deps=deps, registry=registry, config_manager=config_manager)
assert agent.deps is not None
assert agent.deps == deps
def test_runtime_depends_on_all():
"""Verify runtime package can use all dependencies"""
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent
from cogniverse_vespa.backend import VespaBackend
# Runtime should be able to import all packages
assert OrchestratorAgent is not None
assert VespaBackend is not None
assert SystemConfig is not None
def test_layered_architecture_dependencies():
"""Verify proper layering - lower layers don't import higher layers"""
# SDK should not import from other packages
# Foundation can import SDK
# Core can import SDK and Foundation
# Agents can import SDK, Foundation, and Core
# Implementation can import SDK, Foundation, and Core
# Services can import all
pass
Multi-Tenant Testing¶
Tenant Isolation Tests¶
Test Tenant-Specific Configuration:
# tests/test_tenant_isolation.py
import pytest
from cogniverse_agents.search_agent import SearchAgent, SearchAgentDeps
def test_tenant_config_isolation():
"""Verify tenant IDs are distinct — isolation is enforced per-operation."""
# Tenant IDs are plain strings passed to operations (search, schema deploy,
# memory add). SystemConfig does not carry a tenant_id field; the backend
# URL/port are shared across tenants in the same deployment.
tenant_a = "acme_corp"
tenant_b = "globex_inc"
assert tenant_a != tenant_b
def test_tenant_schema_naming():
"""Verify tenant schemas use correct naming convention"""
from cogniverse_foundation.config.utils import create_default_config_manager
from cogniverse_core.schemas.filesystem_loader import FilesystemSchemaLoader
from pathlib import Path
config_manager = create_default_config_manager()
schema_loader = FilesystemSchemaLoader(Path("configs/schemas"))
agent = SearchAgent(
deps=SearchAgentDeps(profile="video_colpali_smol500_mv_frame"),
config_manager=config_manager,
schema_loader=schema_loader,
)
# Agent's search service should target tenant-specific schema
# tenant_id is per-request on search_by_text()
results = agent.search_by_text(
query="test",
tenant_id="acme_corp",
top_k=5,
)
def test_tenant_phoenix_project_isolation():
"""Verify Phoenix projects can be registered per-tenant"""
from cogniverse_foundation.telemetry.manager import TelemetryManager
# TelemetryManager is a singleton, configure per-tenant projects
telemetry = TelemetryManager()
# Register tenant-specific projects
telemetry.register_project(tenant_id="acme_corp", project_name="acme-project")
telemetry.register_project(tenant_id="globex_inc", project_name="globex-project")
# Verify registration succeeded (projects are stored internally)
# Actual project isolation is enforced when creating spans
assert telemetry is not None
Multi-Tenant Data Isolation¶
Test Cross-Tenant Data Boundaries:
# tests/integration/test_multi_tenant_isolation.py
import pytest
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_runtime.ingestion.pipeline import VideoIngestionPipeline
from cogniverse_core.registries.backend_registry import BackendRegistry
from cogniverse_vespa.backend import VespaBackend
@pytest.mark.integration
async def test_tenant_data_isolation(sample_video, config_manager, schema_loader):
"""Verify tenants cannot access each other's data"""
from cogniverse_runtime.ingestion.pipeline import PipelineConfig
# Create pipeline config for tenant A
pipeline_config_a = PipelineConfig(
extract_keyframes=True,
generate_embeddings=True,
search_backend="vespa"
)
# Ingest video for tenant A
pipeline_a = VideoIngestionPipeline(
tenant_id="acme_corp",
config=pipeline_config_a,
config_manager=config_manager,
schema_loader=schema_loader
)
result_a = await pipeline_a.process_video_async(sample_video)
assert result_a["status"] == "completed"
# Search as tenant B (should get no results from tenant A)
backend = BackendRegistry.get_search_backend(
name="vespa",
config_manager=config_manager,
schema_loader=schema_loader
)
results = backend.search(
query_dict={
"query": "test",
"type": "video",
"profile": "video_colpali_smol500_mv_frame",
"tenant_id": "globex_inc",
}
)
# Tenant B should not see tenant A's documents
assert len(results) == 0
@pytest.mark.integration
def test_tenant_memory_isolation(config_manager, schema_loader):
"""Verify tenant memories are isolated"""
from cogniverse_core.memory.manager import Mem0MemoryManager
memory_a = Mem0MemoryManager(tenant_id="acme_corp")
memory_a.initialize(
backend_host="localhost",
backend_port=8080,
llm_model="openai/google/gemma-4-e4b-it",
embedding_model="lightonai/DenseOn",
llm_base_url="http://localhost:11434",
embedder_base_url="http://localhost:11434",
config_manager=config_manager,
schema_loader=schema_loader,
)
memory_b = Mem0MemoryManager(tenant_id="globex_inc")
memory_b.initialize(
backend_host="localhost",
backend_port=8080,
llm_model="openai/google/gemma-4-e4b-it",
embedding_model="lightonai/DenseOn",
llm_base_url="http://localhost:11434",
embedder_base_url="http://localhost:11434",
config_manager=config_manager,
schema_loader=schema_loader,
)
# Add memory for tenant A
memory_a.add_memory(
content="Secret message",
tenant_id="acme_corp",
agent_name="test_agent"
)
# Search as tenant B (should not find tenant A's memory)
results_b = memory_b.search_memory(
query="Secret message",
tenant_id="globex_inc",
agent_name="test_agent",
top_k=5
)
assert len(results_b) == 0
Tenant-Aware Fixtures¶
Example Patterns (not in conftest.py; create as needed for multi-tenant tests):
Tenant isolation in Cogniverse is enforced at the operation level (schema deployment, search, memory add) via a tenant_id string argument. SystemConfig carries backend connection details (URL, port) shared across all tenants — it does not hold a tenant_id field.
import pytest
from cogniverse_foundation.config.unified_config import SystemConfig
from cogniverse_agents.orchestrator_agent import OrchestratorAgent, OrchestratorDeps
from cogniverse_core.registries.agent_registry import AgentRegistry
@pytest.fixture
def backend_config():
"""Shared backend config (same Vespa endpoint for all tenants)."""
return SystemConfig(
backend_url="http://localhost",
backend_port=8080,
)
@pytest.fixture
def multi_tenant_ids():
"""Tenant ID strings for cross-tenant tests."""
return {
"tenant_a": "acme_corp",
"tenant_b": "globex_inc",
}
@pytest.fixture
def tenant_agent():
"""Create orchestrator agent (tenant_id is passed per-operation)."""
from cogniverse_foundation.config.manager import ConfigManager
from cogniverse_foundation.config.unified_config import LLMEndpointConfig
from cogniverse_foundation.telemetry.config import TelemetryConfig
from tests.utils.memory_store import InMemoryConfigStore
store = InMemoryConfigStore()
store.initialize()
config_manager = ConfigManager(store=store)
deps = OrchestratorDeps(
telemetry_config=TelemetryConfig(),
llm_config=LLMEndpointConfig(
model="openai/google/gemma-4-e4b-it",
api_base="http://localhost:11434/v1",
),
)
registry = AgentRegistry(tenant_id="test-tenant", config_manager=config_manager)
return OrchestratorAgent(deps=deps, registry=registry, config_manager=config_manager)
# Use in tests:
def test_with_tenant_fixtures(multi_tenant_ids, tenant_agent):
assert tenant_agent.deps is not None
assert multi_tenant_ids["tenant_a"] != multi_tenant_ids["tenant_b"]
Test Isolation¶
State Cleanup Between Tests¶
Tests use auto-fixtures to clean up state:
@pytest.fixture(autouse=True, scope="function")
def cleanup_dspy_state():
"""Clean up DSPy state between tests to prevent isolation issues"""
yield
# Clean up any DSPy state after each test
try:
import dspy
# Reset ALL DSPy settings attributes to prevent any state pollution
if hasattr(dspy, "settings"):
if hasattr(dspy.settings, "lm"):
dspy.settings.lm = None
if hasattr(dspy.settings, "adapter"):
dspy.settings.adapter = None
if hasattr(dspy.settings, "rm"):
dspy.settings.rm = None
if hasattr(dspy.settings, "experimental"):
dspy.settings.experimental = False
# Clear any context stack from async tests
if hasattr(dspy, "_context_stack"):
if hasattr(dspy._context_stack, "clear"):
dspy._context_stack.clear()
elif isinstance(dspy._context_stack, list):
dspy._context_stack.clear()
except (ImportError, AttributeError, RuntimeError):
pass
cleanup_background_threads()
Best Practices¶
- Always use fixtures for shared state
- Clean up resources in fixture teardown
- Don't rely on test execution order - tests should be independent
- Use unique IDs for multi-tenant tests (tenant_id, user_id)
Test Config Isolation (cogniverse_test_config)¶
tests/conftest.py ships a session-scoped autouse=True fixture named cogniverse_test_config. It isolates configuration mutations from the tracked configs/config.json; it does not change model identity or silently substitute a smaller model.
What it does:
- Clones
configs/config.jsonunchanged into a tmpdir and links its siblingschemas/directory so relative schema discovery still works. - Sets
COGNIVERSE_CONFIG=<tmpdir>/config.jsonfor the session. - Tests that request
ensure_host_ollamathen require the production Gemma primary model and, when marked, the distinct teacher model, both served on Modal. The resolver accepts an endpoint only when its authenticated/v1/modelslists the exact model and revision, and publishes the verified endpoints in a second temporary config. Nothing is started on this host; see "Model endpoints" indocs/testing/TESTING_GUIDE.md.
Override the Modal lookup via env vars before pytest starts:
Skip it by setting COGNIVERSE_CONFIG yourself (e.g. CI matrix):
If your unit test exercises ConfigUtils' cwd-based auto-discovery (chdir to a tmp dir that has its own
configs/config.json), clearCOGNIVERSE_CONFIGfirst — otherwiseConfigUtilshonours the env var and bypasses the chdir:
tests/utils/llm_config.py honours the same COGNIVERSE_CONFIG priority: env var first, then configs/config.json. Use get_llm_model() / get_llm_base_url() from there in test code so the autouse fixture's overrides flow through.
Performance¶
Test Execution Time¶
Optimize test performance:
# Parallel execution (be careful with async tests) — requires pytest-xdist,
# not a project dependency by default: `uv pip install pytest-xdist` first
uv run pytest -n auto
# Show the 10 slowest test/setup durations after the run (built into pytest core)
uv run pytest --durations=10
Model Caching¶
Local model loaders cache loaded weights across calls to speed up tests. get_or_load_model in libs/core/cogniverse_core/common/models/model_loaders.py keys the module-level _model_cache dict by model_name (or model_name@remote_inference_url when remote inference is configured), holds a per-key threading.Lock so concurrent from_pretrained calls can't corrupt PyTorch/accelerate's meta-tensor dispatch state, and evicts a cached entry whose parameters have degraded to a meta device before re-loading:
def get_or_load_model(
model_name: str,
config: Dict[str, Any],
logger: Optional[logging.Logger] = None,
force_reload: bool = False,
) -> Tuple[Any, Any]:
"""Get model from cache or load it. Thread-safe via a per-key lock."""
...
Debugging Tests¶
Debug Output¶
# Show print statements
uv run pytest -s
# Show full error traceback
uv run pytest --tb=long
# Drop into debugger on failure
uv run pytest --pdb
Logging¶
Enable detailed logging:
import logging
logging.basicConfig(level=logging.DEBUG)
# Or in tests
@pytest.fixture
def debug_logging():
logging.getLogger().setLevel(logging.DEBUG)
CI/CD Considerations¶
GitHub Actions Configuration¶
- name: Run Tests
env:
JAX_PLATFORM_NAME: cpu
TOKENIZERS_PARALLELISM: false
OMP_NUM_THREADS: 1
run: |
timeout 7200 uv run pytest -v --tb=long
Docker Testing¶
# Ensure single-threaded mode
ENV TOKENIZERS_PARALLELISM=false
ENV OMP_NUM_THREADS=1
ENV MKL_NUM_THREADS=1
RUN pytest
Docker Fixtures for Integration Tests¶
Integration tests use self-managed Docker containers via pytest fixtures.
shared_vespa — Single Session-Scoped Vespa Container¶
The project uses a single session-scoped shared_vespa fixture defined in tests/conftest.py. All integration tests share this one Vespa container; isolation between tests comes from unique tenant IDs rather than separate containers. This eliminates the RAM pressure that caused OOM-kills when every package spawned its own container.
shared_vespa yields a dict:
{
"http_port": <int>, # Vespa data port
"config_port": <int>, # Vespa config-server port
"container_name": <str>,
"base_url": "http://localhost:<http_port>",
}
Per-package compatibility shims¶
Each per-package conftest.py re-exports shared_vespa and provides a thin shim under the fixture name that tests in that package already reference:
# e.g. tests/backends/integration/conftest.py
from tests.conftest import shared_vespa # noqa: F401
@pytest.fixture(scope="module")
def vespa_instance(shared_vespa):
"""Shim: exposes shared_vespa under the name backends tests expect."""
yield {
"http_port": shared_vespa["http_port"],
"config_port": shared_vespa["config_port"],
"base_url": shared_vespa["base_url"],
"container_name": shared_vespa["container_name"],
}
# No teardown — shared_vespa owns the container lifecycle.
Tenant isolation helpers¶
tests/utils/tenant_helpers.py provides two functions for deriving Vespa-safe unique tenant IDs from the running test's request object:
tenant_id_for_module(request)— all tests in the same module share a tenant (use when a fixture is module-scoped).tenant_id_for_test(request)— unique per test function, including parametrize variants (use when each test needs a blank-slate schema).
tests/utils/vespa_test_helpers.py provides helpers for building config managers and deploying schemas against shared_vespa:
make_config_manager(shared_vespa)— returns aConfigManagerbound to the shared container's ports.deploy_tenant_schema(shared_vespa, *, tenant_id, base_schema_name)— deploys a base schema scoped totenant_id; returns the full schema name.schema_full_name(base_schema_name, tenant_id)— computes the deployed schema name without going through deploy.load_raw_schema_json(base_schema_name)— reads a base schema definition fromconfigs/schemas/.
Example fixture using the helpers:
# Usage in tests
@pytest.mark.integration
async def test_vespa_operations(shared_vespa, request):
"""Test with the shared Vespa container."""
from tests.utils.tenant_helpers import tenant_id_for_test
from tests.utils.vespa_test_helpers import deploy_tenant_schema, make_config_manager
tenant_id = tenant_id_for_test(request)
config_manager = make_config_manager(shared_vespa)
deploy_tenant_schema(
shared_vespa,
tenant_id=tenant_id,
base_schema_name="video_colpali_smol500_mv_frame",
config_manager=config_manager,
)
http_port = shared_vespa["http_port"]
base_url = shared_vespa["base_url"]
# ... test operations
Phoenix Docker Fixtures¶
For telemetry and evaluation tests:
The phoenix_container fixture is defined in tests/conftest.py (module-scoped). Key details:
- Image:
arizephoenix/phoenix:20.16.0pinned by digest (never:latest) - Ports: HTTP 16006 + per-process offset (→ 6006), gRPC 14317 + per-process offset (→ 4317) —
port_offset = (os.getpid() % 1000) * 10so concurrent pytest sweeps don't collide - Env var: Sets
TELEMETRY_OTLP_ENDPOINT(notOTLP_ENDPOINT) andTELEMETRY_SYNC_EXPORT - Leftover cleanup: Kills only its own process's leftover
phoenix_test_pid<pid>_*containers from a prior crashed run — never touches another concurrent sweep's containers - Scope: Module-scoped — one Phoenix instance per test module
Test containers carry cogniverse-test-owner-pid=<pytest pid>. When every pytest session starts, the shared reaper removes containers whose owner process is gone and already-exited containers; the session's test sidecars summary lists what it removed, or why docker could not be asked. A live owner's created, restarting, paused, or running container is preserved: docker create and docker start are separate operations, so treating the short created interval as abandoned can kill another concurrent test's service before it starts.
Companion fixtures also defined in tests/conftest.py:
phoenix_client—phoenix.client.Clientpointed at the container'shttp_endpoint(http://localhost:16006plus the per-process port offset)telemetry_config_with_phoenix—TelemetryConfigpre-configured for the containertelemetry_manager_with_phoenix— session-installedTelemetryManagersingletontelemetry_manager_without_phoenix— function-scoped manager with mock endpoints (use for unit/integration tests that don't export real spans)
Port Management¶
Vespa test containers get their host ports when they start, in tests/utils/docker_utils.py:
start_docker_container_with_port_retry(module_name, name_prefix=..., image=..., container_ports=(8080, 19071), ...)picks a free(http_port, config_port)pair withgenerate_unique_ports(config_portis alwayshttp_port + 10991, the offset the runtime re-derives), runs the container under a name unique to the process, thread and port, and removes the container a failed start created. Docker's bind-conflict errors (address already in use,port is already allocated) are retried on a fresh pair; other failures raise with Docker's stderr, and exhausted retries raiseDocker container allocation failed after N attempts.VespaDockerManager.start_container(module_name)intests/utils/vespa_docker.py(andVespaTestManager, which uses it) starts the Vespa image that way and returns the ports it got.
generate_unique_ports only probes: another process can take a probed port before Docker binds it. Never compute ports at import time or pass a pre-chosen pair to a container start; read the ports from the started container.
CI Disk Space Requirements¶
Vespa requires disk usage below 75%. In GitHub Actions:
- name: Free up disk space for Vespa
run: |
# Remove ~30GB of unused packages
sudo rm -rf /usr/share/dotnet # .NET SDK
sudo rm -rf /usr/local/lib/android # Android SDK
sudo rm -rf /opt/ghc # Haskell
sudo rm -rf /opt/hostedtoolcache/CodeQL # CodeQL
sudo docker image prune -af
df -h # Verify disk usage
- name: Pre-pull Vespa Docker image
run: docker pull vespaengine/vespa:8.668.5
Summary¶
This guide covers comprehensive testing for Cogniverse multi-agent system:
- Async Testing: Thread-safe configuration for async tests with ML models
- SDK Package Testing: Verify cross-package imports and dependencies
- Multi-Tenant Testing: Ensure tenant isolation at all layers
- Test Execution: Use
uv run pytestfor workspace packages - Performance: Model caching and parallel execution strategies
- CI/CD: Configuration for automated testing
Key Testing Principles:
-
Always use
uv run pytestto activate workspace -
Test tenant isolation at schema, project, and memory levels
-
Use fixtures for reusable tenant configurations
-
Verify SDK package imports work correctly
-
Maintain test independence (no shared state)
-
Clean up background threads to avoid segfaults
Test Organization by Package:
graph TD
root["<span style='color:#000'><b>tests/</b></span>"]
admin["<span style='color:#000'><b>admin/</b><br/>profile & tenant management tests</span>"]
admin_unit["<span style='color:#000'>unit/</span>"]
agents["<span style='color:#000'><b>agents/</b><br/>cogniverse_agents tests</span>"]
agents_unit["<span style='color:#000'>unit/</span>"]
agents_int["<span style='color:#000'>integration/</span>"]
agents_e2e["<span style='color:#000'>e2e/</span>"]
backends["<span style='color:#000'><b>backends/</b><br/>cogniverse_vespa tests</span>"]
backends_unit["<span style='color:#000'>unit/</span>"]
backends_int["<span style='color:#000'>integration/</span>"]
charts["<span style='color:#000'><b>charts/</b><br/>Helm chart tests</span>"]
cli["<span style='color:#000'><b>cli/</b><br/>cogniverse_cli tests</span>"]
cli_unit["<span style='color:#000'>unit/</span>"]
common["<span style='color:#000'><b>common/</b><br/>cogniverse_core.common tests</span>"]
common_unit["<span style='color:#000'>unit/</span>"]
common_int["<span style='color:#000'>integration/</span>"]
core["<span style='color:#000'><b>core/</b><br/>cogniverse_core tests</span>"]
core_unit["<span style='color:#000'>unit/</span>"]
core_int["<span style='color:#000'>integration/</span>"]
e2e["<span style='color:#000'><b>e2e/</b><br/>cross-package end-to-end tests</span>"]
e2e_deployment["<span style='color:#000'>deployment/</span>"]
evaluation["<span style='color:#000'><b>evaluation/</b><br/>cogniverse_evaluation tests</span>"]
evaluation_unit["<span style='color:#000'>unit/</span>"]
evaluation_int["<span style='color:#000'>integration/</span>"]
events["<span style='color:#000'><b>events/</b><br/>event system tests</span>"]
events_unit["<span style='color:#000'>unit/</span>"]
events_int["<span style='color:#000'>integration/</span>"]
finetuning["<span style='color:#000'><b>finetuning/</b><br/>fine-tuning pipeline tests</span>"]
finetuning_int["<span style='color:#000'>integration/</span>"]
fixtures["<span style='color:#000'><b>fixtures/</b><br/>shared LLM/sidecar fixtures</span>"]
foundation["<span style='color:#000'><b>foundation/</b><br/>cogniverse_foundation tests</span>"]
foundation_unit["<span style='color:#000'>unit/</span>"]
foundation_int["<span style='color:#000'>integration/</span>"]
ingestion["<span style='color:#000'><b>ingestion/</b><br/>cogniverse_runtime.ingestion tests</span>"]
ingestion_unit["<span style='color:#000'>unit/</span>"]
ingestion_int["<span style='color:#000'>integration/</span>"]
memory["<span style='color:#000'><b>memory/</b><br/>memory system tests</span>"]
memory_unit["<span style='color:#000'>unit/</span>"]
memory_int["<span style='color:#000'>integration/</span>"]
messaging["<span style='color:#000'><b>messaging/</b><br/>cogniverse_messaging tests</span>"]
messaging_unit["<span style='color:#000'>unit/</span>"]
messaging_int["<span style='color:#000'>integration/</span>"]
routing["<span style='color:#000'><b>routing/</b><br/>routing-specific tests</span>"]
routing_unit["<span style='color:#000'>unit/</span>"]
routing_int["<span style='color:#000'>integration/</span>"]
runtime["<span style='color:#000'><b>runtime/</b><br/>cogniverse_runtime tests</span>"]
runtime_unit["<span style='color:#000'>unit/</span>"]
runtime_int["<span style='color:#000'>integration/</span>"]
synthetic["<span style='color:#000'><b>synthetic/</b><br/>synthetic data tests</span>"]
synthetic_int["<span style='color:#000'>integration/</span>"]
system["<span style='color:#000'><b>system/</b><br/>end-to-end system tests</span>"]
telemetry["<span style='color:#000'><b>telemetry/</b><br/>cogniverse_foundation.telemetry tests</span>"]
telemetry_unit["<span style='color:#000'>unit/</span>"]
telemetry_int["<span style='color:#000'>integration/</span>"]
utils["<span style='color:#000'><b>utils/</b><br/>shared test utilities</span>"]
root --> admin
root --> agents
root --> backends
root --> charts
root --> cli
root --> common
root --> core
root --> e2e
root --> evaluation
root --> events
root --> finetuning
root --> fixtures
root --> foundation
root --> ingestion
root --> memory
root --> messaging
root --> routing
root --> runtime
root --> synthetic
root --> system
root --> telemetry
root --> utils
admin --> admin_unit
e2e --> e2e_deployment
agents --> agents_unit
agents --> agents_int
agents --> agents_e2e
backends --> backends_unit
backends --> backends_int
cli --> cli_unit
common --> common_unit
common --> common_int
core --> core_unit
core --> core_int
evaluation --> evaluation_unit
evaluation --> evaluation_int
events --> events_unit
events --> events_int
finetuning --> finetuning_int
foundation --> foundation_unit
foundation --> foundation_int
ingestion --> ingestion_unit
ingestion --> ingestion_int
messaging --> messaging_unit
messaging --> messaging_int
runtime --> runtime_unit
runtime --> runtime_int
memory --> memory_unit
memory --> memory_int
routing --> routing_unit
routing --> routing_int
synthetic --> synthetic_int
telemetry --> telemetry_unit
telemetry --> telemetry_int
style root fill:#b0bec5,stroke:#546e7a,color:#000
style admin fill:#b0bec5,stroke:#546e7a,color:#000
style agents fill:#ce93d8,stroke:#7b1fa2,color:#000
style backends fill:#90caf9,stroke:#1565c0,color:#000
style charts fill:#b0bec5,stroke:#546e7a,color:#000
style cli fill:#b0bec5,stroke:#546e7a,color:#000
style common fill:#a5d6a7,stroke:#388e3c,color:#000
style core fill:#ce93d8,stroke:#7b1fa2,color:#000
style e2e fill:#b0bec5,stroke:#546e7a,color:#000
style evaluation fill:#ffcc80,stroke:#ef6c00,color:#000
style events fill:#ffcc80,stroke:#ef6c00,color:#000
style finetuning fill:#ffcc80,stroke:#ef6c00,color:#000
style fixtures fill:#b0bec5,stroke:#546e7a,color:#000
style foundation fill:#a5d6a7,stroke:#388e3c,color:#000
style ingestion fill:#ffcc80,stroke:#ef6c00,color:#000
style memory fill:#90caf9,stroke:#1565c0,color:#000
style messaging fill:#90caf9,stroke:#1565c0,color:#000
style routing fill:#ce93d8,stroke:#7b1fa2,color:#000
style runtime fill:#90caf9,stroke:#1565c0,color:#000
style synthetic fill:#ffcc80,stroke:#ef6c00,color:#000
style system fill:#b0bec5,stroke:#546e7a,color:#000
style telemetry fill:#a5d6a7,stroke:#388e3c,color:#000
style utils fill:#b0bec5,stroke:#546e7a,color:#000
style admin_unit fill:#b0bec5,stroke:#546e7a,color:#000
style e2e_deployment fill:#b0bec5,stroke:#546e7a,color:#000
style agents_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
style agents_int fill:#ba68c8,stroke:#7b1fa2,color:#000
style agents_e2e fill:#ba68c8,stroke:#7b1fa2,color:#000
style backends_unit fill:#64b5f6,stroke:#1565c0,color:#000
style backends_int fill:#64b5f6,stroke:#1565c0,color:#000
style cli_unit fill:#b0bec5,stroke:#546e7a,color:#000
style common_unit fill:#81c784,stroke:#388e3c,color:#000
style common_int fill:#81c784,stroke:#388e3c,color:#000
style core_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
style core_int fill:#ba68c8,stroke:#7b1fa2,color:#000
style evaluation_unit fill:#ffb74d,stroke:#ef6c00,color:#000
style evaluation_int fill:#ffb74d,stroke:#ef6c00,color:#000
style events_unit fill:#ffb74d,stroke:#ef6c00,color:#000
style events_int fill:#ffb74d,stroke:#ef6c00,color:#000
style finetuning_int fill:#ffb74d,stroke:#ef6c00,color:#000
style foundation_unit fill:#81c784,stroke:#388e3c,color:#000
style foundation_int fill:#81c784,stroke:#388e3c,color:#000
style ingestion_unit fill:#ffb74d,stroke:#ef6c00,color:#000
style ingestion_int fill:#ffb74d,stroke:#ef6c00,color:#000
style memory_unit fill:#64b5f6,stroke:#1565c0,color:#000
style memory_int fill:#64b5f6,stroke:#1565c0,color:#000
style messaging_unit fill:#64b5f6,stroke:#1565c0,color:#000
style messaging_int fill:#64b5f6,stroke:#1565c0,color:#000
style routing_unit fill:#ba68c8,stroke:#7b1fa2,color:#000
style routing_int fill:#ba68c8,stroke:#7b1fa2,color:#000
style runtime_unit fill:#64b5f6,stroke:#1565c0,color:#000
style runtime_int fill:#64b5f6,stroke:#1565c0,color:#000
style synthetic_int fill:#ffb74d,stroke:#ef6c00,color:#000
style telemetry_unit fill:#81c784,stroke:#388e3c,color:#000
style telemetry_int fill:#81c784,stroke:#388e3c,color:#000 Related Documentation: