Limits
Use discovery and response headers for deployment-specific limits. The tables below give the implementation defaults and the failure each boundary produces.
Rate limits
Section titled “Rate limits”Requests are metered with a token bucket per durable authority, not per IP, per key, or per connection. Rotating a credential does not get you a fresh bucket.
| Bucket | Partition | Burst | Sustained per minute |
|---|---|---|---|
Authenticated /v1 |
Your installation (or membership) within one organization | 120 | 600 |
| Pre-authentication edge | Client IP | 600 | 3000 |
| Token issuance | Client IP | 60 | 60 |
Those are the shipped defaults, and an operator can tune every one of them. Do
not hardcode them. The buckets report under different policy names —
wamp-cloud for the authenticated bucket, wamp-cloud-auth for the
pre-authentication edge. Every authenticated response tells you where you
stand:
RateLimit-Policy: "wamp-cloud";q=600;w=60;wamp-burst=120RateLimit: "wamp-cloud";r=117;t=1q is the sustained quota, w the window in seconds, wamp-burst the bucket
capacity, r your remaining tokens. t is seconds until the bucket is
fully replenished at the current rate — not the wait until the next token,
which is usually much shorter. One request costs one token regardless of what
it does.
When the bucket is empty you get 429 with Retry-After, Cache-Control: no-store, and retryAfterSeconds in the body. If the limiter itself faults, the
API fails closed with 503 cloud_rate_limit_unavailable and Retry-After: 1
rather than becoming unmetered.
The two discovery endpoints — GET /v1/openapi.json and GET /v1/docs — are
unauthenticated and are not metered by either bucket.
This is operational flow control, not a commercial entitlement or an SLA.
Quotas
Section titled “Quotas”| Quota | Limit | Enforcement |
|---|---|---|
| Durable turns per session | 1,000 | Checked after exact replay under the session lock; 409 cloud_turn_limit_reached for new work |
| Live organization environments | 200 | Creating or unarchiving above the cap returns 409 quota_exceeded; archived environments do not count |
| Environment archive bytes per organization | 1 GiB | An upload or a listing that copies an archive above the cap returns 409 quota_exceeded |
Retained Session history is not creation admission-limited. Catalogs are
keyset-paginated instead. Capabilities expose request and page limits, not a
Session count.
Archiving (DELETE /v1/sessions/{sessionId} → 204) stops commands on the
session but leaves every read working, as described in
Sessions, turns, runs.
There is no quota on runs over time, artifacts per session, or events per session. The turn boundary keeps every durable conversation and client snapshot finite; it never invalidates an exact retry of a turn already accepted.
Request payload maximums
Section titled “Request payload maximums”| Input | Limit | On violation |
|---|---|---|
| Any JSON request body | 1 MB | 413 request_entity_too_large, with the standard error body on /v1 |
Transcription file |
20 MiB (20 971 520 bytes), FLAC / MP3 / MP4 / M4A / OGG / WAV / WebM | 413 audio_too_large |
Transcription durationMs |
100–600 000 ms (10 minutes) | 400 invalid_audio_duration |
initialTurn.message on session creation |
100 000 characters | 400 invalid_request |
message on a turn |
100 000 characters | 400 invalid_request |
title on a session |
160 characters | 400 invalid_request |
model |
128 characters | 400 invalid_request |
runtime |
64 characters, ^[a-z0-9][a-z0-9-]*$ |
400 invalid_request |
Publication title |
256 characters, no CR or LF | 400 invalid_request |
Publication body |
65 536 characters and 65 536 UTF-8 bytes | 400 invalid_request on characters, 400 invalid_input on bytes |
Merge commitTitle |
256 characters, no CR or LF | 400 invalid_request |
Merge commitMessage |
65 536 characters and 65 536 UTF-8 bytes | 400 invalid_request on characters, 400 invalid_input on bytes |
| Any HTTPS URL in a request | 2048 characters | 400 invalid_request |
origin.tenantKey |
128 characters | 400 invalid_request |
origin.objectType |
64 characters | 400 invalid_request |
origin.objectId |
256 characters | 400 invalid_request |
origin.endUserId |
128 characters | 400 invalid_request |
origin.label |
256 characters | 400 invalid_request |
source.label (public source) |
512 characters | 400 invalid_request |
source.baseBranch (GitHub source) |
255 characters, no NUL, CR or LF | 400 invalid_request |
| Environment document | 256 KiB serialized JSON; name 1–64 characters, description 500 characters, setup 64 KiB UTF-8, at most 200 variables with values up to 32 KiB each | 400 environment_invalid |
| Environment cache paths | At most 16 canonical home or workspace-relative paths | 400 environment_invalid |
| Environment repositories | At most 20; each path one directory name, distinct ignoring case; requires github: true |
400 environment_invalid |
| Environment repository clones | Five minutes for all clones of one sandbox | The apply fails with environment_setup_failed |
| Environment extensions | At most 20, each listed once | 400 environment_invalid |
| Environment archive | A .zip of at most 32 MiB, 10,000 entries and 200 MiB expanded, with an extension.json of at most 1 MiB |
413 request_entity_too_large over 32 MiB; 400 environment_invalid otherwise |
| Environment secret value | 32 KiB UTF-8 | 400 environment_invalid |
replyTo.interactionId |
255 characters | 400 invalid_request |
Opaque paging cursor |
1024 characters (512 on GET /v1/repositories) |
400 invalid_request (invalid_input on repositories) |
Two limits interact and both apply: an initial message of 100 000 characters is legal by
character count, but if it is heavily multibyte or heavily escaped the encoded
JSON can still exceed the 1 MB body cap and fail with 413 before validation
ever runs. capabilities.limits.maxMessageCharacters reports the character cap.
The two long free-text fields — publication body and merge commitMessage —
are checked twice, and the second check is the binding one. The route validates
65 536 characters; the service behind it then re-validates 65 536 bytes of
UTF-8. Text that is legal by character count is still rejected if its encoded
form is larger, so size those two fields in bytes.
Interaction projection maximums
Section titled “Interaction projection maximums”An ask_user Interaction retains at most 4 questions and 20 options per
question. A projected question or prompt is limited to 2000 characters, a
header to 80, an option label or description to 500, and a structured default
or free-text answer to 4000. These are output-sanitization bounds: excess or
malformed runtime fields are omitted or truncated instead of failing the Turn.
Page sizes
Section titled “Page sizes”| Endpoint | limit range |
Default |
|---|---|---|
GET /v1/sessions |
1–100 | 50 |
GET /v1/repositories |
1–100 | 50 |
GET …/publications |
1–100 | 50 |
GET …/publications/{publicationId}/merges |
1–100 | 50 |
GET …/events |
1–100 | 100 |
GET …/turns |
1–100 | 100 |
GET …/artifacts |
1–100 | 100 |
limit=0, limit=101 and a non-numeric limit are each 400, not clamped —
invalid_request on the session, publication, merge, event, turn and artifact
lists, invalid_input on GET /v1/repositories. GET …/interactions is not
paged and returns the whole open set.
The integer cursors are bounded too: GET …/turns requires after ≥ −1
(default −1), and GET …/events requires after ≥ 0 (default 0). Discovery
reports the ceiling as capabilities.limits.maxPageSize.
GET …/events/stream and GET /v1/sessions/stream each hold one live SSE
response, and they share one budget: each authority may hold 64 streams of
either kind, and the Cloud process 256. A refused open returns 429 or 503
with Retry-After. An open stream ends after a random 5–10 minutes or at
credential expiry, whichever comes first, with a retry: of at most 1 second; a
Cloud shutdown ends every stream with a retry: spread over 0–30 seconds.
Keep-alive comments arrive at least
every 15 seconds. More than 1 MiB of unflushed output closes a slow reader;
resume from the last Event or watermark your handler committed. Neither stream
issues idle page reads. Appends to one Session within 250 ms reach the change
stream as one notice.
Artifact sizes
Section titled “Artifact sizes”These are enforced when the server retains an artifact, so exceeding them does
not fail your API call — it means the artifact is not kept, and you receive a
wamp.artifact.omitted event instead.
| Artifact | Limit |
|---|---|
| Presented file, each | 32 MiB (33 554 432 bytes) |
| Presented files, total per run | 64 MiB (67 108 864 bytes) |
| Presented files, count per run | 24 |
Run summary (result.summary) |
100 000 bytes |
| Conversation checkpoint | 256 MiB, with a 16 KiB manifest |
| Workspace checkpoint | 256 MiB (268 435 456 bytes), with a 64 KiB manifest |
| ACP runtime continuation | 256 MiB compressed, at most 1 GiB source data, a 32 KiB manifest and 32 files |
The conversation and workspace checkpoint figures are reported by discovery as
capabilities.artifacts.conversationCheckpointMaxBytes and
capabilities.artifacts.workspaceCheckpointMaxBytes; the workspace figure also
appears as capabilities.environments[0].workspace.maxCheckpointBytes. The
conversation log and ACP continuation are private artifacts streamed between the engine and
object storage; their bytes never share the 64 MiB JSON-RPC frame. The
ACP and presented-file caps are not advertised in the public API, which is why they
are worth reading here: a run that writes a 40 MiB build output produces an
omission event rather than an artifact, and nothing about the API call tells you
that in advance.
A workspace checkpoint over its cap is the difference between a session that can
continue after its sandbox ends and one that cannot — see continuation on the
session resource.
Two more bounds apply to presented files and fail the same quiet way. Their
metadata is limited — a path of 1 024 characters, a label of 256, a
description of 1 024 — and a file whose metadata exceeds a cap is silently
omitted, not rejected: it is counted in wamp.artifact.omitted like an
oversized one. And each file is read out of the sandbox with a 30-second
timeout; a read that does not finish in time is omitted as well.
Sandbox timeouts
Section titled “Sandbox timeouts”A lease has two deadlines, and the earlier one wins. They answer different questions and neither is configurable by a caller.
| Limit | Value | Measured from |
|---|---|---|
| Runaway ceiling on one sandbox lease | 86 400 seconds (24 hours) | when the lease started |
| Idle hold past the last sign of use | 1500 seconds (25 minutes) | the last touch |
The ceiling is unmovable: nothing extends a lease past 24 hours from its own start, and a session that outlives it rolls onto a new lease rather than extending the old one.
The idle hold is what usually ends a lease. Three things count as a touch, and each pushes the underlying lease out again — attaching to the workspace, opening a connection to it, and the platform observing a run the engine still reports as live. All three are re-derived facts, not heartbeats. An abandoned session releases its capacity 25 minutes later instead of holding it to the ceiling.
Neither number is advertised in the API. What you read is
workspace.expiresAt, and it is a live view of the hold the sandbox is
actually under: every touch that pushes the lease out is reflected in it, so
it does move forward while the session is in use, up to the 24-hour
ceiling. Read it when it matters rather than caching it, and see
Sandboxes and environments for what that means for a
long-running turn.
One more lease bound is visible as an error: work will not start on a lease
with less than 10 minutes of life remaining. A turn submitted against one
answers 409 cloud_workspace_expiring instead of starting.
You observe the lease through the session’s workspace object, which has
exactly three states:
workspace.state |
Meaning |
|---|---|
unavailable |
No sandbox is attached |
attached |
A sandbox is attached; expiresAt is its deadline |
expired |
The lease deadline has passed |
An expired workspace does not destroy the session. Read continuation to
find out whether another turn can resume the work, and at what fidelity.
There is no per-run wall-clock limit and no per-turn timeout you can set, and
neither sandbox deadline is tunable per request. An agent that stops on its own
budget reports it as stopReason on the terminal run event (time_budget,
cost_budget, turn_budget, max_turns). A run killed because its lease
deadline passed is different: it arrives as wamp.run.crashed with
stopReason: "execution_deadline". See Events.
Concurrency ceilings
Section titled “Concurrency ceilings”These are enforced by database constraints, which means they hold under concurrent requests from any number of your processes — not just per process.
| Ceiling | Value | Result of exceeding |
|---|---|---|
Executing runs per session (one more may wait queued behind a finalizing or awaiting run) |
1 | 409 cloud_run_in_progress |
| Non-terminal publications per session | 1 | 409 cloud_publication_in_progress |
| Live merges per publication | 1 | 409 cloud_publication_merge_conflict |
| Successful merges per publication, ever | 1 | 409 |
| Turns resolving one interaction | 1 | 409 cloud_interaction_conflict |
What the two merge rows mean for a publishing flow — why “one at a time” and “once, ever” are different guarantees, and how to recover a merge id you lost — is Publications.
The first row is the one that shapes your architecture: one run at a time per session means one session per concurrent task. Do not multiplex unrelated work into one session merely to reduce catalog rows.
There is no documented cap on concurrent sessions executing at once and no separate sandbox count you can query.
Retry behavior the server applies for you
Section titled “Retry behavior the server applies for you”| Command | Attempts | Backoff |
|---|---|---|
| Run dispatch to a sandbox | 10, then the run fails with dispatch_attempts_exhausted |
2 s doubling, capped at 5 minutes |
| Run cancellation | 6 cancel RPC attempts, counted separately from dispatch. After the sixth backoff, the next claim force-stops a Run that has not ended as cancelled and releases its sandbox lease |
Same curve, using the cancellation counter |
| Publication | 6 ordinary claims. Before a provider write, exhaustion fails the publication. After a write was attempted, an uncertain outcome stays pending until GitHub confirms what happened | At least 30 seconds initially; unresolved attempts after the sixth are observed at most every 15 minutes |
Merge (mode: when_ready) |
Not capped | Re-attempted about every 30 seconds while the pull request stays open and unchanged, or after the wait GitHub asks for when one applies, clamped to 15 minutes |
A run exposes dispatch attempts; cancellation has its own internal counter.
A merge also exposes attempts; a publication does not. The two retrying objects behave
differently at the end:
- A publication ends on its own. After six attempts it stops and reports
terminal
failedwithlastError.code. Poll for a terminal status and handle it; there is nothing to keep waiting for. - A
when_readymerge does not. It is uncapped, so your side owns the timeout: decide how long you are willing to wait for the pull request to become mergeable and stop polling when that elapses. The server will keep re-attempting indefinitely.
Credential lifetimes
Section titled “Credential lifetimes”| Credential | Lifetime |
|---|---|
| App assertion (the JWT you sign) | Up to 600 seconds. A longer one is rejected |
| Installation access token | 600 seconds |
| Clock skew tolerated on an assertion | 30 seconds |
Cache an installation token for its lifetime minus a margin — a minute is
comfortable — and re-mint on 401. Do not mint one per request.
Repository review budgets
Section titled “Repository review budgets”GET /v1/sessions/{sessionId}/repository/review is bounded, and it tells you
when it hit a bound rather than silently shortening the diff.
| Budget | Value |
|---|---|
| Files in one review | 500 |
| Total path bytes | 128 KiB |
| Total patch bytes | 1 MiB |
When the patch budget is exhausted, patchTruncated is true and individual
files lose their patch field; the file list stays complete. Exceeding the file
budget is 413 review_too_large.
Transcript budgets
Section titled “Transcript budgets”| Budget | Value |
|---|---|
| Transcript facts captured per run | 10 000 |
wamp.message.created text |
16 384 characters |
| Activity or runtime name | 128 characters |
Beyond the fact budget you receive one wamp.timeline.truncated event carrying
the count that was dropped. Text over its cap arrives with truncated: true.
Retention
Section titled “Retention”Checkpoint retention is explicit; deletion windows for user data are not:
| Data | What is verifiable |
|---|---|
| Sessions | Archive stops commands; reads keep working. No physical-deletion window is published |
| Events | No active-Session retention window is published. Persist events your product must keep |
| User-visible artifact bytes | Nothing is reclaimed while its manifest exists. A manifest you can list is a manifest whose bytes you can fetch — see Artifacts |
| Conversation, workspace and ACP checkpoints | The latest three complete restore points are retained per session. A failed or result-only capture does not evict a good restore point. ACP state is optional; the portable conversation remains the fallback |
| Orphaned checkpoint objects | A stored checkpoint object that no live artifact references is reclaimed 24 hours after it is orphaned |
Checkpoint retention is count-based, not time-based. The robust transcript pattern is unchanged: if you need transcripts in your own system, page events into your own store; checkpoint objects are private restore machinery, not an export or archive API.
If you have a data-retention obligation for events, sessions, or user-visible artifacts, archive the session and confirm the deployment’s deletion policy with your operator. Checkpoint pruning is not user-data deletion.
Related
Section titled “Related”- Errors — the codes each of these limits produces.
- API conventions — paging parameters and the rate-limit headers.
- Capabilities operations — reading the advertised limits at runtime instead of hardcoding them.