Coding Agent CLI¶
Interactive coding agent accessible from the terminal. Plan, generate, and execute code changes against the cogniverse runtime with real-time streaming and multi-turn conversation.
Pi client¶
The clients/pi-cogniverse extension connects Pi's terminal and local workspace tools to the Cogniverse runtime. It registers the cogniverse provider using models returned by GET /v1/models when the extension loads. Model discovery failures stop extension loading with an error identifying the endpoint.
Set COGNIVERSE_API_KEY to the runtime's harness key and COGNIVERSE_BASE_URL to its OpenAI-compatible base URL (default: http://localhost:8000/v1). See the client setup.
The extension prompts before bash, write, and edit: allow once, always allow for the current session, or deny. A permission dialog resolved after switching sessions is discarded with a notification. The command /cogniverse-search <query> inserts retrieval results as visible custom context without starting a model turn.
Commands¶
cogniverse code¶
Interactive REPL that streams the coding agent's plan → generate → execute → evaluate loop.
| Option | Default | Description |
|---|---|---|
--tenant | $COGNIVERSE_TENANT_ID (required) | Tenant identifier — must be passed or set in the env var |
-l, --language | python | Primary programming language |
-n, --iterations | 5 | Max plan-code-execute iterations per task |
-c, --codebase | (none) | Indexed codebase path for context search |
REPL commands¶
| Command | Action |
|---|---|
| free text | Send a coding task to the agent |
/apply | Write the last generated code changes to local files |
/diff | Show a diff between proposed changes and local files |
/plan | Re-display the last plan |
/language <lang> | Change language mid-session |
/codebase <path> | Set codebase path for context search |
/iterations <n> | Set max iterations |
/clear | Clear conversation history |
/help | Show all commands |
/exit or Ctrl+D | Exit the REPL |
Example session¶
$ export COGNIVERSE_TENANT_ID=acme
$ cogniverse code
Cogniverse Coding Agent (tenant: acme, lang: python)
Type a coding task, or /help for commands. Ctrl+D to exit.
>>> write a retry decorator with exponential backoff
>> Searching code context...
>> Planning implementation...
## Plan
1. Create utils/retry.py with a retry decorator
2. Support max_retries, base_delay, max_delay parameters
3. Use random jitter to avoid thundering herd
>> Generating code (iteration 1/5)...
>> Executing in sandbox...
Iteration 1: passed
## Summary
Created retry decorator with exponential backoff and random jitter.
Files: utils/retry.py
>>> /apply
utils/retry.py (new)
Applied 1 file(s)
>>> make max_retries configurable via an env var
>> Searching code context...
>> Generating code (iteration 1/5)...
...
>>> /exit
cogniverse index¶
Index local files into Vespa so the coding agent can find relevant context when generating code.
| Option | Default | Description |
|---|---|---|
<path> | required | Directory to index |
--type | code | Content type: code (only code is currently implemented) |
--tenant | $COGNIVERSE_TENANT_ID (required) | Tenant identifier |
--profile | (auto from type) | Override Vespa profile |
Code files go to code_lateon_mv (tree-sitter AST chunking, LateOn-Code multi-vector embeddings). The --type docs and --type video choices are accepted by the CLI but not yet implemented — the command prints a warning and returns without indexing.
Knowledge graph extraction — in addition to content indexing, cogniverse index extracts a knowledge graph of entities and relationships from code and text files and writes it to a separate schema. Query it with cogniverse graph. See Knowledge Graph for full details.
The indexer walks the directory, respects .gitignore, skips node_modules / .venv / __pycache__, and uploads via the runtime's /ingestion/upload endpoint. Re-indexing the same files is idempotent.
cogniverse sandbox status¶
Show the state of the OpenShell gateway that runs sandboxed code execution.
Prints the active gateway name, its config directory, whether it's running, and whether its certs are synced into the cluster.
cogniverse sandbox sync¶
Re-sync the OpenShell gateway's mTLS certs into the cluster after cert rotation (host mode only).
Reads the current certs from ~/.config/openshell/gateways/<name>/ and updates the openshell-mtls Secret and openshell-metadata / openshell-active ConfigMaps. Restart the runtime pod afterwards so the Python client picks up the new certs.
Architecture¶
The CLI is a thin HTTP client. All agent logic runs inside the cogniverse runtime.
flowchart LR
CLI["<span style='color:#000'><b>cogniverse code</b><br/>REPL</span>"]
subgraph RT["<span style='color:#000'><b>Cogniverse Runtime</b><br/>k3d / prod k8s</span>"]
DISP["<span style='color:#000'>AgentDispatcher<br/>create_streaming_agent("coding_agent", ...)</span>"]
AGENT["<span style='color:#000'>CodingAgent<br/>DSPy planning, code generation, evaluation</span>"]
SBM["<span style='color:#000'>SandboxManager</span>"]
DISP --> AGENT --> SBM
end
subgraph GW["<span style='color:#000'><b>OpenShell Gateway</b></span>"]
PODS["<span style='color:#000'>Sandbox pods</span>"]
end
CLI -->|"POST /a2a/ (message/stream)"| RT
RT -.->|"SSE: status, partial, final"| CLI
SBM -->|"gRPC (mTLS)"| GW
GW --> PODS
style CLI fill:#90caf9,stroke:#1565c0,color:#000
style RT fill:#ce93d8,stroke:#7b1fa2,color:#000
style DISP fill:#ba68c8,stroke:#7b1fa2,color:#000
style AGENT fill:#ba68c8,stroke:#7b1fa2,color:#000
style SBM fill:#ba68c8,stroke:#7b1fa2,color:#000
style GW fill:#b0bec5,stroke:#546e7a,color:#000
style PODS fill:#b0bec5,stroke:#546e7a,color:#000 - REPL loop sends each turn as a JSON-RPC
message/streamrequest to/a2a/withconversation_historyso the agent sees prior plans and code. - Streaming is done via Server-Sent Events. Each
statusevent updates the current phase label (search,plan,generate,execute,evaluate). The final event carries theCodingOutputwith plan, code changes, execution results, and summary. - Sandbox execution happens inside an OpenShell sandbox pod. Code is written to
/tmp/coding_workspace/solution.<ext>, run with the generated test command, and the stdout/stderr/exit code come back to the agent for evaluation. If the exit code is non-zero, the agent iterates with the error as context.
Sandbox Deployment Modes¶
The coding agent requires an OpenShell sandbox gateway for code execution. The gateway comes in two flavors depending on where cogniverse runs.
Host Mode (local dev with k3d)¶
cogniverse up auto-installs the OpenShell CLI if missing, starts the gateway on the host, and syncs its mTLS certs into the cluster as k8s Secrets.
What happens under the hood:
- The
openshellCLI is downloaded to~/.local/binif not already installed. openshell gateway start --port $OPENSHELL_GATEWAY_HOST_PORTlaunches theghcr.io/nvidia/openshell/cluster:0.0.13Docker container on the host. This image bundles a mini-k3s cluster that in turn runs the gateway as a pod inside itself.OPENSHELL_GATEWAY_HOST_PORTdefaults to28080(openshell's own default of8080collides with k3d's serverlb container). Step 4 below reads the port back from the gateway's ownmetadata.json(gateway_port) when it rewrites the pod-facing endpoint, so the two can never drift.- The gateway generates mTLS certs at
~/.config/openshell/gateways/<name>/mtls/. cogniverse_cli.sandbox.sync_gateway_certs_to_cluster()reads those certs and creates k8s resources in thecogniversenamespace:- Secret
openshell-mtls— containsca.crt,tls.crt,tls.key - ConfigMap
openshell-metadata— the gateway'smetadata.jsonwith the endpoint rewritten tohttps://host.docker.internal:<gateway_port>(taken from the samemetadata.json) so the pod can reach the host gateway on the port it actually listens on; the runtime pod also needsruntime.sandbox.hostGatewayIPsohost.docker.internalresolves inside k3d - ConfigMap
openshell-active— points tocogniverseas the active gateway name - The runtime pod mounts all three at
/home/cogniverse/.config/openshell/gateways/cogniverse/. - When the coding agent needs to execute code, it calls
SandboxClient.from_active_cluster()which reads the mounted metadata and certs, opens a gRPC connection tohost.docker.internal:<gateway_port>with mTLS, and creates a sandbox.
Host mode is a single-machine setup — one host, one gateway, one developer. The sandboxes run inside the inner k3s cluster that the openshell image bundles.
If the openshell CLI can't be installed or the gateway fails to start, cogniverse up doesn't abort — it logs a warning and continues without the coding agent's execution sandbox.
Cert rotation: OpenShell regenerates certs if the gateway is destroyed and restarted. Run cogniverse sandbox sync to copy the new certs into the cluster, then restart the runtime pod.
Sandbox task sessions¶
A coding task leases one OpenShell session for its whole run through SandboxManager.task_session(agent_type, tenant_id). The session is created on entry, owned exclusively for the duration of the task, and destroyed on release, so each task pays a cold sandbox start and no sandbox state crosses tasks. SandboxSessionPool tracks the live leases, counts them against COGNIVERSE_SANDBOX_POOL_SIZE, and raises SandboxCapacityError when every slot is taken — container creation is bounded by the gateway's capacity, not by request concurrency.
close_all destroys every live task session immediately. It runs on runtime shutdown and on the mTLS reconnect path (SandboxManager._drop_stale_pool), both of which close the gateway client right after; a task session outlives any single call, so deferring it to release would leave the container alive at the gateway. A lease whose session is still being created stays with its owner, which destroys it on release.
| Env var | Default | Effect |
|---|---|---|
COGNIVERSE_SANDBOX_POOL_SIZE | 8 | Ceiling on concurrent task sessions. |
Sandbox lifecycle telemetry¶
Taking a lease emits sandbox.create_session and sandbox.wait_ready; releasing it emits sandbox.delete. Each SandboxTaskSession.exec emits a parent sandbox.task_exec span with a child sandbox.exec span. The sandbox.exec span carries:
| Attribute | Meaning |
|---|---|
openshell.agent_type | Agent name (e.g. coding_agent) |
openshell.command_first | First token of the command (audit aid) |
openshell.timeout_seconds | The exec timeout |
openshell.exit_code | Subprocess exit code |
openshell.wall_ms | Wall-clock duration of the exec |
openshell.oom | True when exit_code ∈ {137, 139} or stderr matches OOM markers |
openshell.policy_denied | True when stderr matches permission denied / syscall denied / blocked by policy |
openshell.error | Exception class name (parent span only, on hard failure) |
The parent span also carries openshell.tenant_id and openshell.session_name. These spans become children of whichever agent span is active when the task runs, so Phoenix shows the sandbox call inline with the rest of the agent's processing trace.
Application-layer egress enforcement¶
In addition to kernel-layer NetworkPolicy enforcement (in-cluster mode), the runtime enforces each agent's network_policies.egress allow-list at the httpx transport layer. Agents whose dispatcher path stamps a policy obtain their httpx client via SandboxManager.make_http_client(agent_type) — the returned client wraps every outbound request in a PolicyEnforcingTransport that raises EgressDeniedError for non-allow-listed (host, port). A rule naming a service's SystemConfig default address (localhost:8000 for the runtime, localhost:8080 for Vespa) also admits the address the deployment configures for that service, passed as endpoint_bindings.
This is defence-in-depth: kernel policy stops out-of-process bypass; the transport surfaces the violation in application logs with the offending endpoint and the operator-actionable allow-list.
| Env var | Default | Effect |
|---|---|---|
COGNIVERSE_OPENSHELL_HTTP_ENFORCEMENT | unset | Set to disabled to bypass the transport check (useful while iterating on policies in dev). |
Today wired:
| Agent | Status |
|---|---|
coding_agent | Code execution sandboxed via the existing OpenShell SDK exec path. |
orchestrator_agent | A2A sub-agent calls flow through make_http_client("orchestrator_agent"). |
search_agent | Policy file in place; outbound httpx client to be migrated through the dispatcher's make_http_client per the same pattern as orchestrator. |
summarizer_agent | Policy file in place; LLM endpoint allow-listed via Ollama (port 11434). |
routing_agent | Policy file in place; same pattern as summarizer. |
The remaining agents (search/summarizer/routing) keep the existing httpx clients today; make_http_client(<agent>) is the migration path. Adding the wrapper to a new agent is a one-line change at the agent's construction site in agent_dispatcher.py (mirror the OrchestratorAgent example).
Gateway health probe¶
When sandboxing is not disabled, the runtime starts a background probe that calls SandboxClient.health() every 30 s (configurable via COGNIVERSE_SANDBOX_PROBE_INTERVAL). Each probe emits an OpenTelemetry span named openshell.gateway_health with attributes:
| Attribute | Meaning |
|---|---|
openshell.gateway_available | 1 when health() returns SERVICE_STATUS_HEALTHY/SERVICE_STATUS_UNSPECIFIED (or a response with no status at all); 0 when the gateway reports SERVICE_STATUS_UNHEALTHY/SERVICE_STATUS_DEGRADED, or health() raises (including a timeout) |
openshell.gateway_latency_ms | Round-trip probe latency |
openshell.gateway_error | Sick status name, exception class name, or no_client (only set when unavailable) |
Phoenix holds these spans as the gateway's status history. The probe runs as part of the FastAPI lifespan; stop() is awaited at shutdown so the runtime can exit cleanly.
mTLS cert rotation¶
Production clusters that rotate the OpenShell client certs (cert-manager, Vault PKI, manual openshell auth refresh) need cogniverse to pick up the new TLS material without a process restart. The runtime ships an opt-in :class:CertRotator that watches the active gateway's cert directory:
| Watched file | Purpose |
|---|---|
~/.config/openshell/gateways/<name>/metadata.json | Endpoint + name |
~/.config/openshell/gateways/<name>/mtls/ca.crt | Gateway CA bundle |
~/.config/openshell/gateways/<name>/mtls/tls.crt | Client cert |
~/.config/openshell/gateways/<name>/mtls/tls.key | Client key |
from cogniverse_runtime.openshell_cert_rotator import CertRotator
rotator = CertRotator(sandbox_manager=mgr, interval_seconds=300)
mgr.attach_cert_rotator(rotator)
rotator.start()
# … later, on shutdown:
await rotator.stop()
The rotator polls mtimes on interval_seconds (default 300 s — slow enough to be free, fast enough to catch rotations inside typical cert grace windows). When any watched file changes, it calls SandboxManager.reconnect() so the next exec uses the new client.
The rotator is also wired into the task-session error path: an auth/TLS-shaped error from taking a lease or from SandboxTaskSession.exec (matched on auth, x509, tls, ssl, certificate, permission, unauthenticated, unauthorized) eagerly calls rotator.trigger_on_auth_failure() so rotation visibility doesn't have to wait for the next polling tick. The trigger is rate-limited (one reconnect per 5 s) so a burst of failing requests can't thrash the gateway with handshake attempts.
The rotator emits openshell.cert_rotation spans with attributes openshell.cert_rotation_detected (0/1), openshell.cert_rotation_reason (one of baseline_capture, unchanged, rotation_detected, auth_failure_reconnect, auth_trigger_rate_limited, no_gateway_dir), and openshell.cert_rotation_changed_paths (comma-separated when detected=1).
Sandbox boot policy¶
The runtime resolves a single sandbox.policy knob with three values:
| Value | Behaviour at boot when gateway is unreachable |
|---|---|
required | Refuse to start with SandboxGatewayUnavailableError. Use for production tenants where egress isolation is a compliance requirement. |
optional | Log a warning and continue without sandbox enforcement. Default; suitable for dev and staging. |
disabled | Do not even attempt to connect; SandboxManager.available is permanently False. Use when sandboxing is intentionally off. |
Resolution order (first non-empty wins):
COGNIVERSE_SANDBOX_POLICYenv var —required/optional/disabled.config["sandbox"]["policy"]fromconfigs/config.json(or per-tenant config).COGNIVERSE_SANDBOX_ENABLED+ presence ofOPENSHELL_GATEWAY_ENDPOINT→ maps tooptional(true) ordisabled(false).
Default when none are set: optional.
In-Cluster Mode (production)¶
Production clusters (EKS, GKE, AKS, bare-metal k8s) don't have a "host" to run things on. For these, the gateway is deployed as a k8s StatefulSet inside the same cluster as cogniverse.
helm install cogniverse charts/cogniverse \
--set runtime.sandbox.enabled=true \
--set runtime.sandbox.inCluster.enabled=true \
--set openshell.server.sshHandshakeSecret=$(openssl rand -hex 32)
What the Helm chart deploys:
| Resource | Purpose |
|---|---|
Job: cogniverse-openshell-cert-gen | Pre-install hook that generates CA, server, and client mTLS certs using openssl. Stores them as four k8s Secrets: openshell-server-tls, openshell-server-client-ca, openshell-client-tls, openshell-client-ca. Runs once per Helm install/upgrade. |
StatefulSet: openshell | Runs ghcr.io/nvidia/openshell/gateway:0.0.13 as a non-root pod. Uses the in-cluster k8s API to create sandbox pods in the cogniverse namespace. |
Service: openshell | ClusterIP on port 8080. In-cluster DNS: openshell.cogniverse.svc.cluster.local:8080. |
ServiceAccount + Role + RoleBinding | Permissions for the gateway to create/delete sandbox pods. |
NetworkPolicy | Restricts sandbox SSH ingress to the gateway only. |
The runtime pod's env var OPENSHELL_GATEWAY_ENDPOINT is auto-set to openshell.<namespace>.svc.cluster.local:8080 and it mounts openshell-client-tls + openshell-client-ca as volumes. There is no host dependency — nothing runs outside the cluster.
Differences Between the Two Modes¶
| Host mode | In-cluster mode | |
|---|---|---|
| Docker image | openshell/cluster (k3s-in-Docker wrapper) | openshell/gateway (just the gateway) |
| Bootstrap | openshell gateway start CLI | Helm subchart + pre-install Job |
| Where it runs | Host as a plain Docker container | K8s StatefulSet in cogniverse namespace |
| Sandbox isolation | Inner k3s cluster inside the container | Sandbox pods alongside cogniverse |
| Certs generated by | openshell CLI on first start | openssl pre-install Job in-cluster |
| Runtime endpoint | host.docker.internal:<gateway_port> (from the active gateway's metadata.json) | openshell.cogniverse.svc.cluster.local:8080 |
| Portable? | Local dev only | Any k8s cluster |
Why two modes? The gateway needs a k8s API to schedule sandboxes into. On a dev machine there's no k8s available to the host, so NVIDIA ships openshell/cluster which bundles k3s and runs it inside a Docker container. In production that's redundant — you already have a real k8s cluster, so you run the gateway image directly and it uses the cluster's existing control plane.
Requirements¶
cogniverse upmust have already provisioned the stack (host mode) or the production Helm release must haveruntime.sandbox.enabled=true(in-cluster mode).- The
openshell==0.0.13Python package is pinned incogniverse-runtimeand installed automatically — no manual setup. - Host mode requires Docker (for the gateway container) and downloads the
openshellCLI binary to~/.local/binon firstcogniverse up. - Host mode needs
fs.inotify.max_user_instancesof at least 512 (sysctl -w fs.inotify.max_user_instances=512, persisted under/etc/sysctl.d/). Every K3s on the host — each k3d node and eachopenshell/clustergateway — holds about 35 inotify instances from root's per-user budget, so the Linux default of 128 is exhausted by one k3d node plus one gateway.
Troubleshooting¶
openshell gateway start fails with K8s namespace not ready and the gateway container exits 0 — the gateway's K3s logged failed to create image import watcher ... too many open files: root's inotify instance budget is exhausted. Raise fs.inotify.max_user_instances (see Requirements).
Cannot connect to runtime. Run 'cogniverse up' first. — The REPL can't reach http://localhost:28000. Verify the runtime is healthy: cogniverse status.
Coding agent returns 500 with a SandboxManager error — The gateway isn't reachable from the runtime pod. Check cogniverse sandbox status. In host mode, ensure openshell gateway info reports the gateway is running. In in-cluster mode, check the openshell StatefulSet is ready: kubectl get statefulset openshell -n cogniverse.
Indexed code isn't showing up in search — Code search uses the code_lateon_mv profile. This requires the LateOn-Code query encoder to be registered, which isn't enabled by default. The coding agent falls back to empty context and still generates code without it — search is optional.
REPL commands not recognized — The REPL only recognizes commands starting with /. Free text is always sent to the agent. Use /help to list commands.