Troubleshooting Guide¶
Table of Contents¶
- Test Failures
- Import Errors
- Model Loading Issues
- Memory and Performance
- Vespa Issues
- Deployment and Rollout
- Agent Communication
- Quick Reference
Test Failures¶
Segmentation Faults in Async Tests¶
Symptoms:
Fatal Python error: Segmentation fault
Thread 0x000000033614f000 (most recent call first):
File "/path/to/threading.py", line 359 in wait
File "/path/to/tqdm/_monitor.py", line 60 in run
Cause: Threading conflicts between pytest async event loops and background threads from:
-
tqdm (transformers progress bars)
-
posthog (mem0ai telemetry)
-
torch (multi-threaded operations)
Solution: The test suite is configured for single-threaded mode in tests/conftest.py:
# Already configured - no action needed
os.environ["TOKENIZERS_PARALLELISM"] = "false"
os.environ["OMP_NUM_THREADS"] = "1"
os.environ["MKL_NUM_THREADS"] = "1"
torch.set_num_threads(1)
If you still see segfaults:
- Check test markers: Ensure async tests use
@pytest.mark.asyncio - Verify pytest.ini: Must have
asyncio_mode = auto - Check manual threading: Don't create threads manually in tests
- Update conftest.py: Ensure using latest version with background thread cleanup
Prevention:
-
Always run tests with:
JAX_PLATFORM_NAME=cpu uv run pytest -
Don't override threading environment variables
-
Use smaller models in tests (e.g. ColPali
TomoroAI/tomoro-colqwen3-embed-4brather than larger ColQwen variants)
DSPy Training Data Errors¶
Symptoms:
Cause: Training examples missing required output fields that metrics try to access.
Solution: Add all required output fields to your DSPy Examples:
# ❌ Bad: Missing output fields
example = dspy.Example(
query="test query",
).with_inputs("query")
# ✅ Good: All output fields present (with both input fields)
example = dspy.Example(
query="test query",
context="relevant context for the query",
primary_intent="search",
complexity_level="simple",
needs_video_search="true",
needs_text_search="false",
multimodal_query="false",
temporal_pattern="none",
).with_inputs("query", "context")
Required Fields by Module:
-
Query Analysis: primary_intent, complexity_level, needs_video_search, needs_text_search, multimodal_query, temporal_pattern
-
Agent Routing: recommended_workflow, primary_agent, routing_confidence
Prevention:
-
Validate training data before optimization (see docs/modules/optimization.md)
-
Use example templates from
libs/agents/cogniverse_agents/optimizer/dspy_agent_optimizer.py:328-386 -
Run unit tests for training data loading
Model Loading Hangs or Crashes¶
Symptoms:
Cause: Large models (1B+ parameters) can cause threading issues or memory exhaustion in test environment.
Solution: Use the provisioned production encoder endpoint; do not load a different model in the test process:
# The integration fixture verifies this exact identity through /v1/models.
model_name = "TomoroAI/tomoro-colqwen3-embed-4b"
Default Models:
-
ColPali:
TomoroAI/tomoro-colqwen3-embed-4b(recommended) -
X-CLIP:
microsoft/xclip-large-patch14 -
ColQwen:
TomoroAI/tomoro-colqwen3-embed-4b
Prevention:
-
Check ingestion pipeline configuration for default models
-
Update documentation when changing models
-
Run ingestion tests before committing model changes
Import Errors¶
Module Import Timing Issues¶
Symptoms:
Cause: Function tries to access sys.modules or import modules after the module has already started loading, or imports modules in the wrong order.
Solution: Import system modules (sys, os, logging) at module level, not inside functions:
# ✅ Correct: Import at module level
import sys
import os
def function_using_modules():
# Now can safely use sys, os, etc.
sys.modules["some.module"] = ...
Affected Files:
-
Configuration and memory management modules in core package
-
Note: With layered architecture, ensure imports from correct layers (foundation, core, implementation, application)
Prevention:
-
Import system modules (
sys,os,logging) at module level -
Only use function-level imports for optional dependencies
-
Run test collection before committing:
pytest --collect-only
Missing Dependencies¶
Symptoms:
Solution:
# Sync all dependencies
uv sync
# Dependencies are managed in pyproject.toml
# All required packages will be installed via uv sync
Common Missing Dependencies:
-
colpali-engine: ColPali/ColQwen models -
mem0ai: Memory management (includes posthog) -
gliner: Relationship extraction -
arize-phoenix-client,arize-phoenix-otel: Telemetry
Prevention:
-
Always run
uv syncafter pulling changes -
Check
pyproject.tomlfor required dependencies -
Use
uv runinstead of direct python execution
Model Loading Issues¶
First-Time Model Download¶
Symptoms:
Cause: First-time model downloads from HuggingFace can be large (several GB).
Solution:
-
Be patient: Initial download takes time
-
Check disk space: Models cache in
~/.cache/huggingface/ -
Reuse the verified inference service: the test fixture discovers the exact deployed endpoint and starts the same model only when none is available
Model Sizes:
-
TomoroAI/tomoro-colqwen3-embed-4b(shared visual encoder): ~8GB of bfloat16 weights plus runtime/KV-cache overhead -
microsoft/xclip-large-patch14: ~3GB
Prevention:
-
Pre-download models before running tests
-
Use Docker with pre-cached models
-
Set HF_HOME for custom cache location
Model Cache Corruption¶
Symptoms:
Solution: Clear the HuggingFace cache:
# Remove corrupted cache for the affected model
rm -rf ~/.cache/huggingface/hub/models--TomoroAI--tomoro-colqwen3-embed-4b
# Re-run ingestion or tests to re-download
JAX_PLATFORM_NAME=cpu uv run pytest
Prevention:
-
Don't interrupt model downloads
-
Ensure sufficient disk space
-
Use stable internet connection
Memory and Performance¶
Out of Memory (OOM)¶
Symptoms:
Solution:
-
Reduce batch size:
-
Use CPU instead of GPU:
-
Enable gradient checkpointing (for PyTorch model training only):
Prevention:
-
Monitor memory usage:
nvidia-smi(GPU) orhtop(CPU) -
Use smaller models for development
-
Batch processing for large datasets
Slow Test Execution¶
Symptoms:
-
Test suite takes >30 minutes
-
Individual tests timeout
Solution:
-
Run tests in parallel (requires
pytest-xdist:uv add --dev pytest-xdist): -
Skip slow tests:
-
Use test markers:
Prevention:
-
Mark slow tests with
@pytest.mark.slow -
Use mocks for external dependencies
-
Cache model loading in test fixtures
Vespa Issues¶
Connection Refused¶
Symptoms:
ConnectionError: [Errno 111] Connection refused
requests.exceptions.ConnectionError: http://localhost:8080
Cause: Vespa container not running.
Solution:
# Check if Vespa is running
docker ps | grep vespa
# Start Vespa
docker run -d -p 8080:8080 -p 19071:19071 vespaengine/vespa
# Verify
curl http://localhost:8080/state/v1/health
Prevention:
-
Ensure Vespa is running via
cogniverse up -
Use health checks in tests
-
Document Vespa requirement in README
Schema Deployment Failures¶
Symptoms:
Cause: Mismatch between code expectations and deployed schema.
Solution:
GET /admin/schemas/drift lists tenant schemas still on a definition other than the shipped one, and the reason Vespa refused any the startup migration could not redeploy (see Schema changes in a release).
# Re-deploy schema for the affected tenant
RUNTIME_URL=http://localhost:8000
curl -sfX POST "$RUNTIME_URL/admin/profiles/<profile>/deploy" \
-H 'Content-Type: application/json' \
-d '{"tenant_id": "<tenant_id>", "force": true}'
# Verify deployment
curl http://localhost:19071/application/v2/tenant/default/application/default
Prevention:
-
Version your schemas
-
Test schema changes before deploying
-
Use schema validation in CI/CD
Vespa JVM Heaps¶
The vespa pod runs seven JVMs. The chart sets two of them through vespa.env; the launchers in vespaengine/vespa read exactly these names, and nothing in the image reads any other.
| JVM | Heap | Set by |
|---|---|---|
| config server | -Xms128m -Xmx2g | VESPA_CONFIGSERVER_JVMARGS (just-start-configserver) |
| config proxy | -Xms32m -Xmx512m | VESPA_CONFIGPROXY_JVMARGS (just-run-configproxy) |
| container | -Xms1536m -Xmx1536m | Vespa default; <container><nodes><jvm options> in the rendered services.xml |
| logserver, metrics proxy | -Xms32m -Xmx256m | Vespa default |
| logserver container, cluster controller | -Xms32m -Xmx128m | Vespa default |
The config server holds the active application model: 93 MiB live after a full GC at 180 schemas (generation 3340), 0 full GCs across a 21-activation sweep on the 2 GiB ceiling. Measure before changing it:
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
'PID=$(pgrep -f "^java.*jdisc.logger.tag=configserver"); jcmd $PID GC.run >/dev/null; jcmd $PID GC.heap_info'
tests/charts/test_memory_qos_budget.py pins the vespa container's env to exactly the variables above.
Config Proxy Heap Exhaustion¶
Symptoms:
configproxy stdout Terminating due to java.lang.OutOfMemoryError: Java heap space
runserver event stopped/1 name="libexec/vespa/vespa-wrapper just-run-configproxy" exitcode=3
/opt/vespa/logs/vespa/vespa.log on the vespa pod, one second after a Session activated line. The deploy that triggered it converges in ~80s instead of ~6s while the sentinel restarts the proxy, and every Vespa service is without a config source for that window. Cause: The config proxy caches every config it has served, so its live heap grows with the schema population (76 MiB after a full GC at 180 schemas). Vespa's default proxy heap is 128 MiB, and the JVM runs with -XX:+ExitOnOutOfMemoryError: one activation's allocation burst on top of the live set exceeds the ceiling and the process exits.
Solution: The chart sets the proxy heap through vespa.env.VESPA_CONFIGPROXY_JVMARGS (-Xmx512m; just-run-configproxy reads this variable). Sized for up to roughly twice the current schema population; raise it in the values overlay beyond that:
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
'PID=$(pgrep -f config.proxy.ProxyServer); jcmd $PID GC.run >/dev/null; jcmd $PID GC.heap_info'
Prevention:
tests/charts/test_memory_qos_budget.pypins the rendered value
Vespa Restart Cost¶
Symptoms: The vespa pod reads Ready within a minute of a restart, but queries and feeds fail for far longer while /opt/vespa/logs/vespa/vespa.log fills with transactionlog.replay.start events and DocumentDB(<schema>): Replayed config … serialNum=<n> lines, one per retained config operation, at roughly 20 ms each.
Cause: Every application activation appends a config operation and a config snapshot (documents/<schema>/config/config-<serial>) to every DocumentDB. Proton prunes a DocumentDB's transaction log only after all of its flush targets have flushed; a DocumentDB that holds no documents never flushes on memory pressure, so the retained operations are bounded only by flush.memory.maxage.time. On restart proton replays the retained operations sequentially across all DocumentDBs.
The readiness probe targets the config server (:19071/state/v1/health), which is up before proton has finished replaying.
What the deploy funnel sets: build_services_config (libs/vespa/cogniverse_vespa/vespa_schema_manager.py) renders services.xml for every package that passes through VespaSchemaManager._deploy_package or VespaBackend._deploy_package, with <flushstrategy><native><component><maxage>1800</maxage>. Proton applies it live on activation (no restart); once unflushed data is older than 1800 s the flush engine flushes it, prunes the transaction log and removes the superseded config snapshots, so a restart replays at most the operations of the last 30 minutes.
Reading replay progress on a live pod:
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
'vespa-logfmt -l all -s time,message | grep -cE "transactionlog.replay.(start|complete)"'
replay.start events carry "serialnum":{"first":F,"last":L} per domain; L - F is the number of operations still to replay for that DocumentDB. The last Replayed config … serialNum=<n> line per DocumentDB against that domain's last shows how far replay has come. Confirming the log was pruned:
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
'vespa-logfmt -l all -s time,message | grep "transactionlog.prune.complete" | tail'
kubectl exec -n cogniverse statefulset/cogniverse-vespa -- sh -c \
'ls /opt/vespa/var/db/vespa/search/cluster.cogniverse_content/n0/documents/*/config | grep -c config-'
config-<serial> per DocumentDB). Pod stop budget: vespa-stop-services runs each service's pre-shutdown command from the sentinel config: the container drains with prepareStop d:360 under a 370 s RPC timeout, and the searchnode runs vespa-proton-cmd prepareRestart under a 600 s budget, which flushes every DocumentDB and prunes the log so the next start replays nothing. vespa.terminationGracePeriodSeconds is 1200 (970 s for both commands run in sequence plus process exit) so the kubelet does not kill proton mid-flush.
Prevention:
tests/backends/unit/test_services_config_flush_tuning.pypins the renderedservices.xmltests/backends/integration/test_proton_flush_maxage_effective.pyreads the effective proton config and observes the prune on a real Vespatests/charts/test_termination_grace_period.pypins the grace period
Backend Profile Not Found¶
Symptoms:
Cause: Backend profiles are stored per-tenant under the config service's backend service namespace (ConfigManager.get_backend_profile / list_backend_profiles / create_backend_profile all default to service="backend"). A 404 usually means the profile was created for a different tenant_id, or it was never created for this tenant at all.
Solution:
# List profiles that actually exist for the tenant
curl "http://localhost:8000/admin/profiles?tenant_id=acme_corp"
# Inspect one profile
curl "http://localhost:8000/admin/profiles/my_profile?tenant_id=acme_corp"
Prevention:
-
Always pass the same
tenant_idused at profile-creation time to every subsequentget/deploy/deletecall (runtime admin API and web client both go through the sameservice="backend"config namespace) -
List profiles for a tenant before assuming one is missing vs. misnamed
Deployment and Rollout¶
Runtime Pod Stuck 0/1 NotReady After Deploy¶
Symptoms:
Cause: The runtime's FastAPI lifespan waits for Vespa, bootstraps the metadata schemas, and blocks on deploy convergence before binding port 8000 — worst case ~810s on a loaded cluster. The startupProbe budget (30s initial delay + 80 × 15s = 1230s, ~21 min) is sized for exactly this, so Running + 0/1 Ready inside that window is a pod doing its job, not a hung one.
Solution: Wait. Deleting the pod restarts schema convergence from zero and makes the rollout strictly slower. Only once the startup budget is exhausted — the probe itself restarts the pod — is it genuinely stuck:
# Watch startup progress (Vespa wait, schema bootstrap, convergence)
kubectl logs -n cogniverse deploy/cogniverse-runtime -f
# Check whether the startupProbe has started killing the pod
kubectl describe pod -n cogniverse \
-l app.kubernetes.io/component=runtime | grep -A5 Events
Ingestion Jobs in the Dead-Letter Stream¶
Symptoms:
# Submit reports a terminal failure without pipeline errors in the logs
{"state": "failed", ...}
# The dead stream is non-empty
kubectl exec -n cogniverse deploy/cogniverse-redis -- \
redis-cli XLEN ingest:queue:dead
Cause: A job redelivered more than INGEST_REAPER_MAX_DELIVERIES (default 5) times without completing is moved to ingest:queue:dead. The delivery counter never resets and cannot distinguish two very different causes:
- Genuine poison — the job crashes its worker every run (e.g. an OOM-sized video). Worker logs show a crash/OOM per delivery.
- Operational restarts — each worker restart (rollout, node drain, manual delete) while the job was mid-processing costs one redelivery after the reaper's idle threshold (
INGEST_REAPER_MIN_IDLE_MS, default 5 min). Enough restarts in a row dead-letter a healthy job.
Solution: Inspect the dead entries — each carries ingest_id, source_url, profile, tenant_id, sha, and times_delivered:
For poison jobs, fix the cause first (raise worker memory, cap video size). For restart-victims, no fix is needed. Either way, dead-lettering clears the job's in-flight idempotency marker and writes no done marker, so re-submitting the same source through POST /ingestion/upload re-enqueues it as a fresh job with a fresh delivery count.
Prevention: Avoid repeated worker restarts while long jobs are in flight — one rollout is one redelivery, but restart loops burn through the cap.
Agent Communication¶
Agent Input Validation Errors¶
Symptoms:
ValidationError: Invalid input format
KeyError: 'query' in agent request
pydantic.ValidationError: Field required
Cause: Request doesn't match expected AgentInput schema or missing required fields.
Solution: Ensure requests follow the agent's input model:
# ✅ Correct format - inherit from AgentInput
from cogniverse_core.agents.base import AgentInput
class SearchInput(AgentInput):
query: str
top_k: int = 10
# Create valid input
input = SearchInput(query="user query here", top_k=5)
# ❌ Missing required fields
input = SearchInput(top_k=5) # ValidationError: query is required
Prevention:
-
Use
AgentInputsubclasses for type-safe inputs -
Check agent interface documentation for required fields
-
Add Pydantic validation in tests
Agent Not Found (404)¶
Symptoms:
Cause: Agents self-register with the runtime by calling POST /agents/register on startup; GET /agents/{agent_name} 404s for any name that hasn't registered. Of the 23 agents defined in configs/config.json under agents.*, 8 (the knowledge-graph and federation agents — citation_tracing_agent, contradiction_reconciliation_agent, multi_document_synthesis_agent, kg_traversal_agent, cross_tenant_comparison_agent, federated_query_agent, temporal_reasoning_agent, knowledge_summarization_agent) ship with enabled: false by default, so their processes are not started and they never register.
Solution:
-
List registered agents:
-
Check the agent's
enabledflag inconfigs/config.jsonunderagents.<agent_name>.enabled— flip it totrueand start that agent's process if you need it. -
Discover by capability instead of by name if you're unsure which agent provides it:
Prevention:
-
Don't assume every agent in
configs/config.jsonis running — checkenabledbefore wiring a caller to it -
Use
/agents/or/agents/by-capability/{capability}for discovery instead of hardcoding agent names
Health Check Failures¶
Symptoms:
GET /health/ready returns {"status": "not_ready", "reason": "No backends registered"}
Health check shows 0 agents or backends
Cause: Runtime initialization incomplete or dependencies unavailable.
Solution:
-
Check service logs:
-
Verify dependencies:
-
Backend (Vespa/Elasticsearch) running and accessible
- Environment variables set (BACKEND_URL, BACKEND_PORT)
-
Configuration manager initialized
-
Check health endpoints:
-
Restart service:
Prevention:
-
Use /health/ready for readiness probes in Kubernetes
-
Use /health/live for liveness probes
-
Monitor backend and agent registry status
Quick Reference¶
Essential Commands¶
# Run all tests
JAX_PLATFORM_NAME=cpu uv run pytest
# Run with debugging
JAX_PLATFORM_NAME=cpu uv run pytest -xvs --pdb
# Check test collection
pytest --collect-only
# Clear caches
rm -rf ~/.cache/huggingface/
rm -rf .pytest_cache/
# Restart services
cogniverse down && cogniverse up
Log Locations¶
Test logs: outputs/logs/*.log
Agent logs: docker logs <container>
Vespa logs: docker logs vespa
Phoenix traces: http://localhost:6006
Support¶
- Documentation: docs/
- Issues: GitHub Issues
- Tests: tests/README.md
- API Reference: docs/modules/