Skip to main content

Reconnect and resume

If a client’s connection to /chat drops mid-stream, it can pick up exactly where it left off:
This works because a chat run is never tied to the HTTP response that started it. POST /chat constructs an internal RunDriver that starts pumping chat.stream() the moment the run begins and keeps running independently of whether anyone is still listening — buffering every formatted SSE frame with a sequence number and broadcasting to live subscribers. GET /chat/:runId/resume just attaches a new subscriber to that same driver: it replays whatever’s already buffered after since, then streams new frames live until the run finishes. Lifecycle hooks (afterRun, onError, etc.) fire exactly once per run regardless of how many times a client reconnects — they belong to the driver, not to any one HTTP response. POST /architect and GET /architect/:runId/resume (behind the architect capability) use the identical mechanism.

Run retention

Finished runs stay resumable for 5 minutes by default before an internal eviction sweep (running every 60s) removes them; the registry also caps out at 1000 tracked runs by default, evicting the oldest-finished ones first if a host’s traffic pattern leaves many runs unclaimed. Both are configurable:
A resume request for an evicted, unknown, or someone-else’s run returns 404 RUN_NOT_FOUND (never 403 — a 404 doesn’t confirm whether the id ever existed).

Honest limitation: single-process, in-memory only

A RunDriver holds a live upstream connection and a JS closure over its subscribers — it cannot be represented in Redis or shared across separate runtime instances. This closes the reconnect gap for the common single-instance deployment (the client dropped and came back, same server process still running). If you run more than one instance behind a load balancer — multiple Kubernetes pods, a PM2 cluster, etc. — a reconnect that lands on a different instance than the one running the original pump won’t find the run (404 RUN_NOT_FOUND) even though it’s still live elsewhere.
Today’s mitigation (deployment-level, no code change): configure your load balancer for session affinity / sticky sessions so reconnects land back on the instance actually running the pump. This doesn’t survive that instance crashing or being redeployed, and doesn’t apply to serverless (no persistent instance to stick to), but covers the common case for free.
Planned, not yet built: a pluggable RunBroker interface — publish/subscribe/claim-ownership for run frames — so a host can back it with Redis (or anything else) and get resume working across instances, including a resume request landing on an instance that never touched the original pump. The in-memory behavior above would remain the zero-config default; a Redis (or similar) implementation would be opt-in, not a dependency this package forces on everyone. Deliberately deferred rather than built speculatively — track issue #229 or open a new one if you need this now. createRuntime()’s close() stops the eviction timer; the timer is also unref’d so it won’t itself keep a Node process alive, but call close() if you construct runtimes repeatedly in a long-lived process (e.g. per-test-suite setup) to avoid accumulating timers.

Heartbeats

POST /chat, GET /chat/:runId/resume, POST /architect, and GET /architect/:runId/resume all send an SSE comment-line heartbeat (: heartbeat\n\n) during any gap between real AG-UI events — e.g. a long-running tool call with no token output — so intermediary proxies and load balancers with an idle-connection timeout don’t kill the stream. Comment lines are invisible to any data:-only SSE parser (including @personaai/sdk’s own parseAguiEventStream), so a consumer never sees them as part of the event sequence.
Heartbeats only cover gaps after the first event of a run — headers can’t be sent until the runtime has already peeked that first event to decide whether the run started successfully (a 401/400/500 has to be a normal buffered response, not a stream), so there’s no way to keep a connection alive with heartbeats before that point. In practice this matters little: the gap heartbeats exist for is a stalled middle of a run (a slow tool call), not the initial time-to-first-token.

Backpressure — a deliberate tradeoff

Backpressure has a real, deliberate tradeoff as of reconnect support. Before reconnect existed, a slow or disconnected consumer propagated backpressure all the way back to Persona’s server — the runtime never pulled a frame it hadn’t been asked for. That’s no longer true: a RunDriver’s pump starts draining chat.stream() the moment the run begins and keeps going regardless of subscriber speed, because resumability requires buffering whatever a reconnecting client might ask to replay. You cannot have both “backpressure all the way to the source” and “a disconnected client can come back and get what it missed” — they’re in direct tension, and this runtime chose resumability. What’s still true and tested (test/runDriver.test.ts):
  • The pump drains the upstream generator exactly once, strictly in order, with no duplicate or skipped next() calls, no matter how many subscribers attach or how slowly they read.
  • Per-run buffers are bounded by that one run’s event count (not indefinite) and released after the grace period described above.
  • The Node bridge in examples/ still layers transport-level backpressure via res.write()’s return value and the drain event — that protects against one slow subscriber blocking the Node process’s memory, but it no longer protects against the runtime itself buffering an in-progress run that nobody is currently reading.