Skip to main content

Realtime runs

POST /v1/realtime/runs (Agent Runs) and POST /v1/super-agent/realtime/runs (Super Agent) return application/x-ndjson. Use them when your product needs live assistant output, thinking, search and todo cards, opaque tool progress, artifacts, and a terminal event. The event vocabulary is the same. This is a breaking change from the previous text/event-stream framing. Existing EventSource / SSE parsers will not work. There is no dual-frame compatibility shim.
Super Agent uses the same stream without agent_id:

Event format

Each line is one JSON object with a type field:
Parse incrementally — do not buffer the body as a single JSON document. Ignore unknown type values. run.created includes render_protocol: 1 so clients can detect this catalog.

Events

One card wins. When a dedicated search or todo card is emitted for a tool_call_id, the stream does not also emit generic tool.started / tool.completed for that id. Tool input, output, bash commands, file-read paths, and raw object-store URLs never appear on this wire.

Terminal events

Realtime runs end with either run.completed or run.failed. Both include usage.credits when the platform computed a cost greater than zero. usage.models is present only for Super Agent–allowlisted accounts: one aggregated row per (model, kind) — chat plus any vision, image, video, or speech models. Token fields appear only when that capability is token-metered; image-gen and Veo rows use generated_images or duration_seconds instead. Webhooks stay usage-free.
GET /v1/runs/{run_id} or GET /v1/super-agent/runs/{run_id} matches the terminal usage for that kind. If a failure happens after stream headers have been sent, the API reports it as a terminal stream event instead of a normal JSON error response.

GET snapshot

Completed runs persist a compact result.blocks array alongside result.text. The stream is full fidelity; the snapshot stored on the run row is not. GET /v1/runs/{run_id} and GET /v1/super-agent/runs/{run_id} may add thinking content onto { "type": "thinking" } markers — one block per thinking span, in the same order as the live stream. List and webhook payloads stay marker-only. Artifact url and url_expires_at are added only on GET hydrate. Persist thinking.delta yourself if you need thinking mid-run or after a disconnect before the run completes. Block types: text, thinking (optional content on GET), tool, search, todo, artifact. The same snapshot is written for async POST /v1/runs.

Client notes

  • Reconnect behavior is your responsibility. If the stream disconnects, fetch the run with GET /v1/runs/{run_id} or GET /v1/super-agent/runs/{run_id} when you have a run ID.
  • Realtime runs do not support webhook_url; the terminal result is delivered through the stream.
  • Keep rendering tolerant of new event fields and unknown type values.