Skip to content

Troubleshooting Guide


Table of Contents

  1. Test Failures
  2. Import Errors
  3. Model Loading Issues
  4. Memory and Performance
  5. Vespa Issues
  6. Deployment and Rollout
  7. Agent Communication
  8. Quick Reference

Test Failures

Segmentation Faults in Async Tests

Symptoms:

Fatal Python error: Segmentation fault

Thread 0x000000033614f000 (most recent call first):
  File "/path/to/threading.py", line 359 in wait
  File "/path/to/tqdm/_monitor.py", line 60 in run

Cause: Threading conflicts between pytest async event loops and background threads from:

  • tqdm (transformers progress bars)

  • posthog (mem0ai telemetry)

  • torch (multi-threaded operations)

Solution: The test suite is configured for single-threaded mode in tests/conftest.py:

# Already configured - no action needed
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["OMP_NUM_THREADS"] = "1"
os.environ["MKL_NUM_THREADS"] = "1"
torch.set_num_threads(1)

If you still see segfaults:

  1. Check test markers: Ensure async tests use @pytest.mark.asyncio
  2. Verify pytest.ini: Must have asyncio_mode = auto
  3. Check manual threading: Don't create threads manually in tests
  4. Update conftest.py: Ensure using latest version with background thread cleanup

Prevention:

  • Always run tests with: JAX_PLATFORM_NAME=cpu uv run pytest

  • Don't override threading environment variables

  • Use smaller models in tests (e.g. ColPali TomoroAI/tomoro-colqwen3-embed-4b rather than larger ColQwen variants)


DSPy Training Data Errors

Symptoms:

AttributeError: 'Example' object has no attribute 'primary_intent'

Cause: Training examples missing required output fields that metrics try to access.

Solution: Add all required output fields to your DSPy Examples:

# ❌ Bad: Missing output fields
example = dspy.Example(
    query="test query",
).with_inputs("query")

# ✅ Good: All output fields present (with both input fields)
example = dspy.Example(
    query="test query",
    context="relevant context for the query",
    primary_intent="search",
    complexity_level="simple",
    needs_video_search="true",
    needs_text_search="false",
    multimodal_query="false",
    temporal_pattern="none",
).with_inputs("query", "context")

Required Fields by Module:

  • Query Analysis: primary_intent, complexity_level, needs_video_search, needs_text_search, multimodal_query, temporal_pattern

  • Agent Routing: recommended_workflow, primary_agent, routing_confidence

Prevention:

  • Validate training data before optimization (see docs/modules/optimization.md)

  • Use example templates from libs/agents/cogniverse_agents/optimizer/dspy_agent_optimizer.py:328-386

  • Run unit tests for training data loading


Model Loading Hangs or Crashes

Symptoms:

Fetching 5 files: 100%|██████████| 5/5 [00:00<00:00, 80000.00it/s]
[Test hangs or segfaults]

Cause: Large models (1B+ parameters) can cause threading issues or memory exhaustion in test environment.

Solution: Use the provisioned production encoder endpoint; do not load a different model in the test process:

# The integration fixture verifies this exact identity through /v1/models.
model_name = "TomoroAI/tomoro-colqwen3-embed-4b"

Default Models:

  • ColPali: TomoroAI/tomoro-colqwen3-embed-4b (recommended)

  • X-CLIP: microsoft/xclip-large-patch14

  • ColQwen: TomoroAI/tomoro-colqwen3-embed-4b

Prevention:

  • Check ingestion pipeline configuration for default models

  • Update documentation when changing models

  • Run ingestion tests before committing model changes


Import Errors

Module Import Timing Issues

Symptoms:

KeyError: Module not found in sys.modules
ImportError: Cannot import module at function call time

Cause: Function tries to access sys.modules or import modules after the module has already started loading, or imports modules in the wrong order.

Solution: Import system modules (sys, os, logging) at module level, not inside functions:

# ✅ Correct: Import at module level
import sys
import os

def function_using_modules():
    # Now can safely use sys, os, etc.
    sys.modules["some.module"] = ...

Affected Files:

  • Configuration and memory management modules in core package

  • Note: With layered architecture, ensure imports from correct layers (foundation, core, implementation, application)

Prevention:

  • Import system modules (sys, os, logging) at module level

  • Only use function-level imports for optional dependencies

  • Run test collection before committing: pytest --collect-only


Missing Dependencies

Symptoms:

ImportError: No module named 'colpali_engine'
ModuleNotFoundError: No module named 'mem0'

Solution:

# Sync all dependencies
uv sync

# Dependencies are managed in pyproject.toml
# All required packages will be installed via uv sync

Common Missing Dependencies:

  • colpali-engine: ColPali/ColQwen models

  • mem0ai: Memory management (includes posthog)

  • gliner: Relationship extraction

  • arize-phoenix-client, arize-phoenix-otel: Telemetry

Prevention:

  • Always run uv sync after pulling changes

  • Check pyproject.toml for required dependencies

  • Use uv run instead of direct python execution


Model Loading Issues

First-Time Model Download

Symptoms:

Fetching 5 files:   0%|          | 0/5 [00:00<?, ?it/s]
[Very slow or timeout]

Cause: First-time model downloads from HuggingFace can be large (several GB).

Solution:

  1. Be patient: Initial download takes time

  2. Check disk space: Models cache in ~/.cache/huggingface/

  3. Reuse the verified inference service: the test fixture discovers the exact deployed endpoint and starts the same model only when none is available

Model Sizes:

  • TomoroAI/tomoro-colqwen3-embed-4b (shared visual encoder): ~8GB of bfloat16 weights plus runtime/KV-cache overhead

  • microsoft/xclip-large-patch14: ~3GB

Prevention:

  • Pre-download models before running tests

  • Use Docker with pre-cached models

  • Set HF_HOME for custom cache location


Model Cache Corruption

Symptoms:

RuntimeError: Error loading model
OSError: Unable to load weights

Solution: Clear the HuggingFace cache:

# Remove corrupted cache for the affected model
rm -rf ~/.cache/huggingface/hub/models--TomoroAI--tomoro-colqwen3-embed-4b

# Re-run ingestion or tests to re-download
JAX_PLATFORM_NAME=cpu uv run pytest

Prevention:

  • Don't interrupt model downloads

  • Ensure sufficient disk space

  • Use stable internet connection


Memory and Performance

Out of Memory (OOM)

Symptoms:

RuntimeError: CUDA out of memory
MemoryError: Unable to allocate array

Solution:

  1. Reduce batch size:

    # In config
    batch_size = 8  # Instead of 32
    

  2. Use CPU instead of GPU:

    JAX_PLATFORM_NAME=cpu uv run python script.py
    

  3. Enable gradient checkpointing (for PyTorch model training only):

    model.gradient_checkpointing_enable()  # PyTorch models only; not applicable to JAX models
    

Prevention:

  • Monitor memory usage: nvidia-smi (GPU) or htop (CPU)

  • Use smaller models for development

  • Batch processing for large datasets


Slow Test Execution

Symptoms:

  • Test suite takes >30 minutes

  • Individual tests timeout

Solution:

  1. Run tests in parallel (requires pytest-xdist: uv add --dev pytest-xdist):

    uv run pytest -n auto
    

  2. Skip slow tests:

    uv run pytest -m "not slow"
    

  3. Use test markers:

    # Only fast unit tests
    uv run pytest -m unit -m "not slow"
    

Prevention:

  • Mark slow tests with @pytest.mark.slow

  • Use mocks for external dependencies

  • Cache model loading in test fixtures


Vespa Issues

Connection Refused

Symptoms:

ConnectionError: [Errno 111] Connection refused
requests.exceptions.ConnectionError: http://localhost:8080

Cause: Vespa container not running.

Solution:

# Check if Vespa is running
docker ps | grep vespa

# Start Vespa
docker run -d -p 8080:8080 -p 19071:19071 vespaengine/vespa

# Verify
curl http://localhost:8080/state/v1/health

Prevention:

  • Ensure Vespa is running via cogniverse up

  • Use health checks in tests

  • Document Vespa requirement in README


Schema Deployment Failures

Symptoms:

RuntimeError: Schema deployment failed
400 Bad Request: Unknown field 'embeddings'

Cause: Mismatch between code expectations and deployed schema.

Solution:

GET /admin/schemas/drift lists tenant schemas still on a definition other than the shipped one, and the reason Vespa refused any the startup migration could not redeploy (see Schema changes in a release).

# Re-deploy schema for the affected tenant
RUNTIME_URL=http://localhost:8000
curl -sfX POST "$RUNTIME_URL/admin/profiles/<profile>/deploy" \
  -H 'Content-Type: application/json' \
  -d '{"tenant_id": "<tenant_id>", "force": true}'

# Verify deployment
curl http://localhost:19071/application/v2/tenant/default/application/default

Prevention:

  • Version your schemas

  • Test schema changes before deploying

  • Use schema validation in CI/CD


Vespa JVM Heaps

The vespa pod runs seven JVMs. The chart sets two of them through vespa.env; the launchers in vespaengine/vespa read exactly these names, and nothing in the image reads any other.

JVM Heap Set by
config server -Xms128m -Xmx2g VESPA_CONFIGSERVER_JVMARGS (just-start-configserver)
config proxy -Xms32m -Xmx512m VESPA_CONFIGPROXY_JVMARGS (just-run-configproxy)
container -Xms1536m -Xmx1536m Vespa default; <container><nodes><jvm options> in the rendered services.xml
logserver, metrics proxy -Xms32m -Xmx256m Vespa default
logserver container, cluster controller -Xms32m -Xmx128m Vespa default

The config server holds the active application model: 93 MiB live after a full GC at 180 schemas (generation 3340), 0 full GCs across a 21-activation sweep on the 2 GiB ceiling. Measure before changing it:

kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
  'PID=$(pgrep -f "^java.*jdisc.logger.tag=configserver"); jcmd $PID GC.run >/dev/null; jcmd $PID GC.heap_info'

tests/charts/test_memory_qos_budget.py pins the vespa container's env to exactly the variables above.


Config Proxy Heap Exhaustion

Symptoms:

configproxy  stdout  Terminating due to java.lang.OutOfMemoryError: Java heap space
runserver    event   stopped/1 name="libexec/vespa/vespa-wrapper just-run-configproxy" exitcode=3
in /opt/vespa/logs/vespa/vespa.log on the vespa pod, one second after a Session activated line. The deploy that triggered it converges in ~80s instead of ~6s while the sentinel restarts the proxy, and every Vespa service is without a config source for that window.

Cause: The config proxy caches every config it has served, so its live heap grows with the schema population (76 MiB after a full GC at 180 schemas). Vespa's default proxy heap is 128 MiB, and the JVM runs with -XX:+ExitOnOutOfMemoryError: one activation's allocation burst on top of the live set exceeds the ceiling and the process exits.

Solution: The chart sets the proxy heap through vespa.env.VESPA_CONFIGPROXY_JVMARGS (-Xmx512m; just-run-configproxy reads this variable). Sized for up to roughly twice the current schema population; raise it in the values overlay beyond that:

vespa:
  env:
    VESPA_CONFIGPROXY_JVMARGS: "-Xmx1g"
Measure the live set before choosing a value:
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
  'PID=$(pgrep -f config.proxy.ProxyServer); jcmd $PID GC.run >/dev/null; jcmd $PID GC.heap_info'

Prevention:

  • tests/charts/test_memory_qos_budget.py pins the rendered value

Vespa Restart Cost

Symptoms: The vespa pod reads Ready within a minute of a restart, but queries and feeds fail for far longer while /opt/vespa/logs/vespa/vespa.log fills with transactionlog.replay.start events and DocumentDB(<schema>): Replayed config … serialNum=<n> lines, one per retained config operation, at roughly 20 ms each.

Cause: Every application activation appends a config operation and a config snapshot (documents/<schema>/config/config-<serial>) to every DocumentDB. Proton prunes a DocumentDB's transaction log only after all of its flush targets have flushed; a DocumentDB that holds no documents never flushes on memory pressure, so the retained operations are bounded only by flush.memory.maxage.time. On restart proton replays the retained operations sequentially across all DocumentDBs.

The readiness probe targets the config server (:19071/state/v1/health), which is up before proton has finished replaying.

What the deploy funnel sets: build_services_config (libs/vespa/cogniverse_vespa/vespa_schema_manager.py) renders services.xml for every package that passes through VespaSchemaManager._deploy_package or VespaBackend._deploy_package, with <flushstrategy><native><component><maxage>1800</maxage>. Proton applies it live on activation (no restart); once unflushed data is older than 1800 s the flush engine flushes it, prunes the transaction log and removes the superseded config snapshots, so a restart replays at most the operations of the last 30 minutes.

Reading replay progress on a live pod:

kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
  'vespa-logfmt -l all -s time,message | grep -cE "transactionlog.replay.(start|complete)"'
replay.start events carry "serialnum":{"first":F,"last":L} per domain; L - F is the number of operations still to replay for that DocumentDB. The last Replayed config … serialNum=<n> line per DocumentDB against that domain's last shows how far replay has come.

Confirming the log was pruned:

kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
  'vespa-logfmt -l all -s time,message | grep "transactionlog.prune.complete" | tail'
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
  'ls /opt/vespa/var/db/vespa/search/cluster.cogniverse_content/n0/documents/*/config | grep -c config-'
After the flush bound has taken effect the second command prints the number of DocumentDBs (one config-<serial> per DocumentDB).

Pod stop budget: vespa-stop-services runs each service's pre-shutdown command from the sentinel config: the container drains with prepareStop d:360 under a 370 s RPC timeout, and the searchnode runs vespa-proton-cmd prepareRestart under a 600 s budget, which flushes every DocumentDB and prunes the log so the next start replays nothing. vespa.terminationGracePeriodSeconds is 1200 (970 s for both commands run in sequence plus process exit) so the kubelet does not kill proton mid-flush.

Prevention:

  • tests/backends/unit/test_services_config_flush_tuning.py pins the rendered services.xml
  • tests/backends/integration/test_proton_flush_maxage_effective.py reads the effective proton config and observes the prune on a real Vespa
  • tests/charts/test_termination_grace_period.py pins the grace period

Backend Profile Not Found

Symptoms:

404 Not Found: Profile 'my_profile' not found for tenant 'acme_corp'

Cause: Backend profiles are stored per-tenant under the config service's backend service namespace (ConfigManager.get_backend_profile / list_backend_profiles / create_backend_profile all default to service="backend"). A 404 usually means the profile was created for a different tenant_id, or it was never created for this tenant at all.

Solution:

# List profiles that actually exist for the tenant
curl "http://localhost:8000/admin/profiles?tenant_id=acme_corp"

# Inspect one profile
curl "http://localhost:8000/admin/profiles/my_profile?tenant_id=acme_corp"

Prevention:

  • Always pass the same tenant_id used at profile-creation time to every subsequent get/deploy/delete call (runtime admin API and web client both go through the same service="backend" config namespace)

  • List profiles for a tenant before assuming one is missing vs. misnamed


Deployment and Rollout

Runtime Pod Stuck 0/1 NotReady After Deploy

Symptoms:

NAME                       READY   STATUS    RESTARTS   AGE
cogniverse-runtime-xxxxx   0/1     Running   0          12m

Cause: The runtime's FastAPI lifespan waits for Vespa, bootstraps the metadata schemas, and blocks on deploy convergence before binding port 8000 — worst case ~810s on a loaded cluster. The startupProbe budget (30s initial delay + 80 × 15s = 1230s, ~21 min) is sized for exactly this, so Running + 0/1 Ready inside that window is a pod doing its job, not a hung one.

Solution: Wait. Deleting the pod restarts schema convergence from zero and makes the rollout strictly slower. Only once the startup budget is exhausted — the probe itself restarts the pod — is it genuinely stuck:

# Watch startup progress (Vespa wait, schema bootstrap, convergence)
kubectl logs -n cogniverse deploy/cogniverse-runtime -f

# Check whether the startupProbe has started killing the pod
kubectl describe pod -n cogniverse \
  -l app.kubernetes.io/component=runtime | grep -A5 Events

Ingestion Jobs in the Dead-Letter Stream

Symptoms:

# Submit reports a terminal failure without pipeline errors in the logs
{"state": "failed", ...}

# The dead stream is non-empty
kubectl exec -n cogniverse deploy/cogniverse-redis -- \
  redis-cli XLEN ingest:queue:dead

Cause: A job redelivered more than INGEST_REAPER_MAX_DELIVERIES (default 5) times without completing is moved to ingest:queue:dead. The delivery counter never resets and cannot distinguish two very different causes:

  1. Genuine poison — the job crashes its worker every run (e.g. an OOM-sized video). Worker logs show a crash/OOM per delivery.
  2. Operational restarts — each worker restart (rollout, node drain, manual delete) while the job was mid-processing costs one redelivery after the reaper's idle threshold (INGEST_REAPER_MIN_IDLE_MS, default 5 min). Enough restarts in a row dead-letter a healthy job.

Solution: Inspect the dead entries — each carries ingest_id, source_url, profile, tenant_id, sha, and times_delivered:

kubectl exec -n cogniverse deploy/cogniverse-redis -- \
  redis-cli XRANGE ingest:queue:dead - +

For poison jobs, fix the cause first (raise worker memory, cap video size). For restart-victims, no fix is needed. Either way, dead-lettering clears the job's in-flight idempotency marker and writes no done marker, so re-submitting the same source through POST /ingestion/upload re-enqueues it as a fresh job with a fresh delivery count.

Prevention: Avoid repeated worker restarts while long jobs are in flight — one rollout is one redelivery, but restart loops burn through the cap.


Agent Communication

Agent Input Validation Errors

Symptoms:

ValidationError: Invalid input format
KeyError: 'query' in agent request
pydantic.ValidationError: Field required

Cause: Request doesn't match expected AgentInput schema or missing required fields.

Solution: Ensure requests follow the agent's input model:

# ✅ Correct format - inherit from AgentInput
from cogniverse_core.agents.base import AgentInput

class SearchInput(AgentInput):
    query: str
    top_k: int = 10

# Create valid input
input = SearchInput(query="user query here", top_k=5)

# ❌ Missing required fields
input = SearchInput(top_k=5)  # ValidationError: query is required

Prevention:

  • Use AgentInput subclasses for type-safe inputs

  • Check agent interface documentation for required fields

  • Add Pydantic validation in tests


Agent Not Found (404)

Symptoms:

GET /agents/kg_traversal_agent -> 404 {"detail": "Agent 'kg_traversal_agent' not found"}

Cause: Agents self-register with the runtime by calling POST /agents/register on startup; GET /agents/{agent_name} 404s for any name that hasn't registered. Of the 23 agents defined in configs/config.json under agents.*, 8 (the knowledge-graph and federation agents — citation_tracing_agent, contradiction_reconciliation_agent, multi_document_synthesis_agent, kg_traversal_agent, cross_tenant_comparison_agent, federated_query_agent, temporal_reasoning_agent, knowledge_summarization_agent) ship with enabled: false by default, so their processes are not started and they never register.

Solution:

  1. List registered agents:

    curl http://localhost:8000/agents/
    

  2. Check the agent's enabled flag in configs/config.json under agents.<agent_name>.enabled — flip it to true and start that agent's process if you need it.

  3. Discover by capability instead of by name if you're unsure which agent provides it:

    curl http://localhost:8000/agents/by-capability/<capability>
    

Prevention:

  • Don't assume every agent in configs/config.json is running — check enabled before wiring a caller to it

  • Use /agents/ or /agents/by-capability/{capability} for discovery instead of hardcoding agent names


Health Check Failures

Symptoms:

GET /health/ready returns {"status": "not_ready", "reason": "No backends registered"}
Health check shows 0 agents or backends

Cause: Runtime initialization incomplete or dependencies unavailable.

Solution:

  1. Check service logs:

    docker logs <container-name>
    # Or for local development
    tail -f outputs/logs/*.log
    

  2. Verify dependencies:

  3. Backend (Vespa/Elasticsearch) running and accessible

  4. Environment variables set (BACKEND_URL, BACKEND_PORT)
  5. Configuration manager initialized

  6. Check health endpoints:

    # Basic health check
    curl http://localhost:8000/health
    
    # Kubernetes readiness probe
    curl http://localhost:8000/health/ready
    
    # Kubernetes liveness probe
    curl http://localhost:8000/health/live
    

  7. Restart service:

    docker restart <container-name>
    

Prevention:

  • Use /health/ready for readiness probes in Kubernetes

  • Use /health/live for liveness probes

  • Monitor backend and agent registry status


Quick Reference

Essential Commands

# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest

# Run with debugging
JAX_PLATFORM_NAME=cpu uv run pytest -xvs --pdb

# Check test collection
pytest --collect-only

# Clear caches
rm -rf ~/.cache/huggingface/
rm -rf .pytest_cache/

# Restart services
cogniverse down && cogniverse up

Log Locations

Test logs:       outputs/logs/*.log
Agent logs:      docker logs <container>
Vespa logs:      docker logs vespa
Phoenix traces:  http://localhost:6006

Support

  • Documentation: docs/
  • Issues: GitHub Issues
  • Tests: tests/README.md
  • API Reference: docs/modules/