Retries and rate limits
Every resource call goes through the shared transport, which automatically retries429 Too Many Requests responses: it reads the response’s Retry-After header (seconds), waits that long, and
retries — up to max_retries times (default 2) before giving up and raising PersonaApiError
with status_code: 429. If Retry-After is absent it falls back to a 1-second wait. This is a
fixed per-attempt wait, not exponential backoff, matching the API’s own Retry-After contract
rather than guessing at a backoff curve. The same retry logic covers chat.stream()’s streaming
requests, not just regular JSON calls.
Only 429 is retried automatically. Any other error (400, 401, 403, 404, 500, etc.)
raises immediately on the first attempt, since retrying those wouldn’t change the outcome.
Pagination
Everylist()-style method (agents.list(), skills.list(), knowledge.list(), mcps.list(),
threads.list(), files.list(), audit_logs.list()) returns a PaginatedResult[T]:
- Defaults:
page=1,limit=20(per endpoint; passed explicitly to override). pagination["total"]is the total matching count across all pages — use it, notlen(items).pagination["pages"]isceil(total / limit)— check it to know whether there’s more, rather than inferring fromlen(items)vs.limit.providers.list()is the one exception: a plain bare list, no params at all — Providers have no discovery concept, so there’s nothing to page or filter by.search(where offered) is free-text against name/description (plustaglineon Agents), andscope: "mine"restricts to the asserted external user’s own items (runtime context only — ignored or empty on a control-plane client).
Idempotency keys
Every resource’screate() (and files.upload()) accepts an optional idempotency_key keyword
argument, sent as the Idempotency-Key request header:
create() call might get retried after an ambiguous failure — a network timeout, a
dropped connection — where you can’t tell whether the original request actually reached the server.
Retrying with the same key returns the exact original response instead of creating a second
resource; retrying with a different (or omitted) key creates a new one as normal. Keys are
scoped per (domain, credential, external user if asserted) and expire after 24 hours. This is
fully opt-in — omitting the argument is identical to every version before this feature, no behavior
change.
_id vs id
The backend has a real, documented inconsistency the SDK reflects rather than papers over:
Use the right key per resource:
provider["id"]/file["id"] but agent["_id"]/skill["_id"]/
thread["_id"]/kb["_id"]/mcp["_id"]. Passing a wrong-shape id to a method simply 404s.
Populated vs. bare fields
Several resources return different field shapes depending on which method produced the value:
These are typed loosely (
list[object] / str | ThreadAgentRef) to reflect the real difference
rather than picking one shape and being wrong for the other calls.
Empty results and existence-hiding
providers.list()on a runtime-plane client returns[]— a Provider’s ownership can only ever be"PersonaUser"or"Project", never"ExternalUser", so a client withexternal_user_idset can never own any. Same foraudit_logs.list(). Silent empty result, not an error.get()/update()/delete()never distinguish “not found” from “not authorized” — a caller that can’t own a resource gets the same404 NOT_FOUNDas a caller asking for a nonexistent id. Don’t build error-driven existence probes.ProviderTestConnectionResult["success"] == Falseis a result, not a thrown error — the call reached the endpoint but the credentials/URL were rejected.KnowledgeSearchResult["score"]can beNoneif the vector store didn’t return a similarity score.ResourceUsage["agents"]is a preview capped at 20 —agentCountis the real total.AuditLogEntry["actorIdentity"]/["targetResourceId"]areNonewhen not applicable.
Bulk delete
bulk_delete(ids) is best-effort: it never raises on per-id failures, and one blocked/not-found
id doesn’t abort the rest of the batch. Check the returned BulkDeleteResult:
reason is generic (existence-hiding: never distinguishes “not found” from “not authorized” from
“blocked by a dependency”). Up to 100 ids per call — a request over that limit is rejected with
a 400 before anything is deleted. Deletes that can be blocked by dependencies (Provider/Skill/MCP
still referenced by an Agent) show up in failed rather than raising.
Ordering
threads.list()returns Threads most recently active first — the only ordering guarantee in the SDK. Everything else has no documented ordering contract; don’t rely on the order oflist()items beyond the search relevance ofsearch.
Timeouts and cancellation
There is no request-cancellation parameter (noAbortSignal equivalent) in v1. For the async
client, wrap a call in asyncio.wait_for(...) or cancel the enclosing asyncio.Task yourself if
you need a timeout/cancellation — this is a documented gap, not a hidden one. There is no default
per-request timeout configured by the SDK beyond whatever the underlying httpx client defaults
to; pass your own http_client= (with timeout=... set on it) to control timeouts globally.
HTTP client lifecycle
- The SDK creates its own
httpx.Client/httpx.AsyncClientif you don’t passhttp_client=, and closes it when youclose()/use the context manager. - If you pass
http_client=, the SDK never closes it (not even viaclose()/context-manager exit) — ownership stays with you. This makes sharing onehttpxclient across many per-requestPersonaClientinstances safe. - Non-JSON success responses (file downloads) return the raw
httpx.Response— read.content/.text/.iter_bytes()/.aiter_bytes()yourself. Non-JSON failure responses raisePersonaApiErrorwithcode == "NON_JSON_ERROR_RESPONSE".
Validation and envelope behavior
- Envelope decoding: a response body with
success: Falseraises even on a 2xx status; adatakey is unwrapped (json_body["data"]) so methods return the payload directly, not the envelope. /whoamiand similar endpoints that omit the fuller envelope are returned as-is.- Request bodies are serialized as JSON with camelCase field names matching the API
(
providerId,systemPrompt, …) even though method parameters are snake_case. - Multipart uploads (
files.upload(),knowledge.upload_documents()) set their ownContent-Typewith a boundary — never forceapplication/jsonon them. - Query params with
Nonevalues are omitted (build_urlfilters them out).