Runs
A run is one invocation of an agent. Last night's cron fire is a run. One @agent call is a run. Submitting a pile of tickets is a run. A run processes N items — one item per thing in the pile — tracks them together under a shared concurrency cap, and reports aggregate progress as the items drain.
If an item is "execute this agent against one input," a run is "execute this agent across N inputs, track them as a unit, and tell me what worked."
Where runs come from
Every path that starts an agent mints exactly one run:
- Calling an
@agentfunction —python agent.py. Each call mints one run over one item; loop it yourself for N. papayya run <agent> <input>— invoke the deployed agent in the cloud.--item-idand--partition-keydeclare identity and tenant.- Run now, in the dashboard — the same invocation, with Input and Item ID fields.
papayya runs submit --agent <id> --file items.jsonl— submit a pile to the hosted worker pool; each JSONL line is one item.- A schedule or trigger — each cron fire and each inbound trigger mints its own run.
A run is the right grain whenever you have N inputs that belong to the same logical job: a nightly review of all new signups, a 5,000-item classification pass, a queue of support tickets. For a single ad-hoc execution you still get a run — it just holds one item.
Lifecycle
materializing → queued → running → completed
→ partial
→ failed
→ cancelled
→ budget_exceeded- materializing — Papayya is inserting one item row per input. Transient; you'll see this briefly after submission.
- queued — every item is inserted, waiting for worker capacity.
- running — at least one item is executing.
- completed — every item finished and none needs triage.
- partial — every item resolved, but at least one didn't work — a raised failure or a
degradedoutcome. The run keeps those items visible for operator action. A partial run promotes tocompletedonce every not-ok item is replayed or dismissed. - failed — every item failed, or the run was stopped by a run-level constraint (budget exceeded, too many consecutive failures).
- cancelled — cancelled via the API or dashboard.
- budget_exceeded — the run's aggregate spend crossed its budget cap; remaining queued items stop dispatching.
Why partial exists
A run with 9,998 items that worked and 2 that didn't has finished, but it hasn't succeeded. Calling that state completed would bury the two failures under a green badge. partial keeps them visible and actionable — the triage surface lists each not-ok item with its captured input and reason so you can replay, dismiss, or acknowledge it. Once triage is done, the run transitions to completed on its own.
Outcome roll-up
A run doesn't just count what crashed — it rolls up whether each item actually worked. Every item carries an outcome (ok, degraded, or failed); the run summarizes them:
| Field | Description |
|---|---|
worst_outcome_status | The worst outcome across all items — ok, degraded, or failed. This is the badge you see first. |
degraded_count | How many items returned a 200 but didn't actually work (a refusal, empty content, or a mark_degraded call). |
This is the run-list line from the quickstart:
triage-ticket — completed · 2 of 6 degraded · 2 tenants · worst: degraded
Six calls returned. Two didn't work. The run's execution status is completed and its worst_outcome_status is degraded — the ran-vs-worked verdict, rolled up.
Progress fields
Every run tracks the pile in real time:
| Field | Description |
|---|---|
total_items | How many items were in the pile at submission |
completed | Count of items that finished successfully |
failed | Count of items that raised or hit budget_exceeded |
degraded_count | Count of items that returned but didn't work |
cancelled | Count of items that were cancelled |
paused | Count of items currently paused (e.g. credit exhaustion) |
aggregate_cost_cents | Sum of spend across all items |
Run-level controls
Budget cap
A run inherits its agent's budget from @agent(budget_usd=...). If aggregate spend across the run's items crosses that cap, remaining queued items stop dispatching and the run transitions to budget_exceeded. Budgets pause the pile — they don't kill in-flight items.
Concurrency cap
A concurrency cap limits how many items execute at once. It defaults to your project's limit; lower it when the downstream API has its own rate limits.
Callback URL
When a run reaches a terminal state, Papayya can POST a summary to a callback URL — 3 retries with exponential backoff, fire-and-forget from the run's perspective. Ideal for queue-worker architectures that want a push instead of polling.
Submitting a run
Locally, from your own code
Each call is a run. Loop it yourself, and every element becomes one item with its own steps, outcome and cost:
from papayya import agent
@agent(name="triage-ticket", model="claude-sonnet-5")
def triage(run, ticket: dict) -> str:
...
for ticket in tickets:
triage(
ticket,
item_id=ticket["id"], # your identity for this item
partition_key=ticket["tenant"], # whose item it is (the tenant)
)item_id= and partition_key= are consumed by Papayya at the call site and are not forwarded to your function — so identity and tenant are declared where the work starts, not threaded through a signature. Declaring either in your own signature still works if you want to read the value.
To the hosted worker pool
Submit a JSONL of items — one line per item — to a deployed agent:
papayya runs submit --agent research-bot --file items.jsonlEach line becomes one item; the whole submission is one run you can watch drain in the dashboard.
Monitoring
- Dashboard — the runs list shows every run with live progress counters and its
worst_outcome_status. Click a run to see its items, each with a worked/degraded badge, its tenant, and its step trace. Runs, items, and steps all land in the hosted dashboard at https://app.getpapayya.com (opens in a new tab), whether you runpython agent.pylocally or deploy to the worker pool. - CLI —
papayya runs listfor the runs,papayya items listfor the hosted items,papayya items stream <item_id>to live-tail one. - API —
GET /v1/durable/runs/{runId}for one item's execution record,GET /v1/durable/runsto list item records. (The wire path is transitional; the product noun is the run/item.) - Callbacks — register a callback URL to get a push notification when the run finishes.
Downloading results
The results of a run are its items — there is no separate results export. List the item execution records as newline-delimited JSON (NDJSON) from the durable-runs surface:
GET /v1/durable/runs
Accept: application/x-ndjson(The wire path is transitional; each durable/runs item is one item.) Each item carries item_id, status, outcome_status, degraded_reason, input_snapshot, output, total_cost_cents, total_input_tokens, total_output_tokens, error_code, error_message, created_at, completed_at, and the like. Fetch a single item's item with GET /v1/durable/runs/{itemId}.
CLI
papayya items listPrints every hosted item as NDJSON, one per line — pipe it into jq or redirect to a file. papayya items list takes no flags; scope to a run or filter by outcome in the dashboard, or via papayya runs list + drill-down. To live-tail one item the moment its steps land — e.g. wiring downstream processing as it finishes — use papayya items stream <item_id>.
For very large runs, export in shards by partitioning on partition_key.
Failure clustering
When a run has failures, the run detail page groups them into clusters so you don't have to scroll through hundreds of item pages to spot a pattern.
Each cluster represents items that share both:
- The same
error_code(e.g.budget_exceeded,tool_call_error,step_timeout), and - A hash of the first 200 characters of the canonicalized input prompt.
So fifteen items that all failed with tool_call_error on the same prompt template collapse into one cluster of size 15. A one-off failure — a single item with a prompt no other item shares — rolls into an "other" bucket below the clustered rows, so the reported counts always reconcile (sum(clusters.count) == total_failures).
Why it matters
Production runs fail in patterns. "47 of your 500 items failed" is useless on its own; "22 failed with tool_call_error on refund-processing prompts, 18 failed with budget_exceeded on research queries, 7 were one-offs" is actionable. You fix the two patterns, replay the failed items, and the run drains.
API
GET /v1/durable/runs/clusters(Failure clusters are a top-level query over the durable-runs surface, not a per-run path.)
Response:
{
"clusters": [
{
"cluster_key": "tool_call_error:a3f1b2c4d5e6f789",
"error_code": "tool_call_error",
"sample_prompt": "process refund for customer",
"count": 22,
"item_ids": ["item_1", "item_2", "item_3", "item_4", "item_5"]
},
{
"cluster_key": "other",
"error_code": "",
"sample_prompt": "",
"count": 3,
"item_ids": []
}
],
"total_failures": 25
}item_idsis capped at 5 per cluster as a representative sample. To drill into every item in a cluster, filter the items list byerror_code.sample_promptis the first 200 characters of the canonicalized input — good enough to recognize the pattern, not enough to leak a full user prompt into summaries.- Results are cached server-side for 30 seconds with automatic invalidation when a new failure lands, so polling the endpoint while a run is in flight is cheap.
Algorithm limitations (v1)
The current grouping is a hash bucket, not semantic similarity. That means:
- Prompts that differ by even one character in the first 200 chars won't cluster together, even if they mean the same thing semantically.
- Agents that construct prompts by prepending per-item data (IDs, timestamps, customer names) will see lots of singletons in the "other" bucket. Template-based prompts (a constant preamble > 200 chars, varying detail at the end) cluster cleanly.
Embedding-based clustering — which groups by meaning rather than exact prefix — is on the post-launch roadmap. For v1, hash-bucket gets you the pattern for most runs where agents reuse prompt templates.