Core Concepts
Runs

Runs

A run is one invocation of an agent. Last night's cron fire is a run. One @agent call is a run. Submitting a pile of tickets is a run. A run processes N items — one item per thing in the pile — tracks them together under a shared concurrency cap, and reports aggregate progress as the items drain.

If an item is "execute this agent against one input," a run is "execute this agent across N inputs, track them as a unit, and tell me what worked."

Where runs come from

Every path that starts an agent mints exactly one run:

  • Calling an @agent function — python agent.py. Each call mints one run over one item; loop it yourself for N.
  • papayya run <agent> <input> — invoke the deployed agent in the cloud. --item-id and --partition-key declare identity and tenant.
  • Run now, in the dashboard — the same invocation, with Input and Item ID fields.
  • papayya runs submit --agent <id> --file items.jsonl — submit a pile to the hosted worker pool; each JSONL line is one item.
  • A schedule or trigger — each cron fire and each inbound trigger mints its own run.

A run is the right grain whenever you have N inputs that belong to the same logical job: a nightly review of all new signups, a 5,000-item classification pass, a queue of support tickets. For a single ad-hoc execution you still get a run — it just holds one item.

Lifecycle

materializing → queued → running → completed
                                 → partial
                                 → failed
                                 → cancelled
                                 → budget_exceeded
  • materializing — Papayya is inserting one item row per input. Transient; you'll see this briefly after submission.
  • queued — every item is inserted, waiting for worker capacity.
  • running — at least one item is executing.
  • completed — every item finished and none needs triage.
  • partial — every item resolved, but at least one didn't work — a raised failure or a degraded outcome. The run keeps those items visible for operator action. A partial run promotes to completed once every not-ok item is replayed or dismissed.
  • failed — every item failed, or the run was stopped by a run-level constraint (budget exceeded, too many consecutive failures).
  • cancelled — cancelled via the API or dashboard.
  • budget_exceeded — the run's aggregate spend crossed its budget cap; remaining queued items stop dispatching.

Why partial exists

A run with 9,998 items that worked and 2 that didn't has finished, but it hasn't succeeded. Calling that state completed would bury the two failures under a green badge. partial keeps them visible and actionable — the triage surface lists each not-ok item with its captured input and reason so you can replay, dismiss, or acknowledge it. Once triage is done, the run transitions to completed on its own.

Outcome roll-up

A run doesn't just count what crashed — it rolls up whether each item actually worked. Every item carries an outcome (ok, degraded, or failed); the run summarizes them:

FieldDescription
worst_outcome_statusThe worst outcome across all items — ok, degraded, or failed. This is the badge you see first.
degraded_countHow many items returned a 200 but didn't actually work (a refusal, empty content, or a mark_degraded call).

This is the run-list line from the quickstart:

triage-ticket — completed · 2 of 6 degraded · 2 tenants · worst: degraded

Six calls returned. Two didn't work. The run's execution status is completed and its worst_outcome_status is degraded — the ran-vs-worked verdict, rolled up.

Progress fields

Every run tracks the pile in real time:

FieldDescription
total_itemsHow many items were in the pile at submission
completedCount of items that finished successfully
failedCount of items that raised or hit budget_exceeded
degraded_countCount of items that returned but didn't work
cancelledCount of items that were cancelled
pausedCount of items currently paused (e.g. credit exhaustion)
aggregate_cost_centsSum of spend across all items

Run-level controls

Budget cap

A run inherits its agent's budget from @agent(budget_usd=...). If aggregate spend across the run's items crosses that cap, remaining queued items stop dispatching and the run transitions to budget_exceeded. Budgets pause the pile — they don't kill in-flight items.

Concurrency cap

A concurrency cap limits how many items execute at once. It defaults to your project's limit; lower it when the downstream API has its own rate limits.

Callback URL

When a run reaches a terminal state, Papayya can POST a summary to a callback URL — 3 retries with exponential backoff, fire-and-forget from the run's perspective. Ideal for queue-worker architectures that want a push instead of polling.

Submitting a run

Locally, from your own code

Each call is a run. Loop it yourself, and every element becomes one item with its own steps, outcome and cost:

from papayya import agent
 
@agent(name="triage-ticket", model="claude-sonnet-5")
def triage(run, ticket: dict) -> str:
    ...
 
for ticket in tickets:
    triage(
        ticket,
        item_id=ticket["id"],            # your identity for this item
        partition_key=ticket["tenant"],  # whose item it is (the tenant)
    )

item_id= and partition_key= are consumed by Papayya at the call site and are not forwarded to your function — so identity and tenant are declared where the work starts, not threaded through a signature. Declaring either in your own signature still works if you want to read the value.

To the hosted worker pool

Submit a JSONL of items — one line per item — to a deployed agent:

papayya runs submit --agent research-bot --file items.jsonl

Each line becomes one item; the whole submission is one run you can watch drain in the dashboard.

Monitoring

  • Dashboard — the runs list shows every run with live progress counters and its worst_outcome_status. Click a run to see its items, each with a worked/degraded badge, its tenant, and its step trace. Runs, items, and steps all land in the hosted dashboard at https://app.getpapayya.com (opens in a new tab), whether you run python agent.py locally or deploy to the worker pool.
  • CLI — papayya runs list for the runs, papayya items list for the hosted items, papayya items stream <item_id> to live-tail one.
  • API — GET /v1/durable/runs/{runId} for one item's execution record, GET /v1/durable/runs to list item records. (The wire path is transitional; the product noun is the run/item.)
  • Callbacks — register a callback URL to get a push notification when the run finishes.

Downloading results

The results of a run are its items — there is no separate results export. List the item execution records as newline-delimited JSON (NDJSON) from the durable-runs surface:

GET /v1/durable/runs
Accept: application/x-ndjson

(The wire path is transitional; each durable/runs item is one item.) Each item carries item_id, status, outcome_status, degraded_reason, input_snapshot, output, total_cost_cents, total_input_tokens, total_output_tokens, error_code, error_message, created_at, completed_at, and the like. Fetch a single item's item with GET /v1/durable/runs/{itemId}.

CLI

papayya items list

Prints every hosted item as NDJSON, one per line — pipe it into jq or redirect to a file. papayya items list takes no flags; scope to a run or filter by outcome in the dashboard, or via papayya runs list + drill-down. To live-tail one item the moment its steps land — e.g. wiring downstream processing as it finishes — use papayya items stream <item_id>.

For very large runs, export in shards by partitioning on partition_key.

Failure clustering

When a run has failures, the run detail page groups them into clusters so you don't have to scroll through hundreds of item pages to spot a pattern.

Each cluster represents items that share both:

  1. The same error_code (e.g. budget_exceeded, tool_call_error, step_timeout), and
  2. A hash of the first 200 characters of the canonicalized input prompt.

So fifteen items that all failed with tool_call_error on the same prompt template collapse into one cluster of size 15. A one-off failure — a single item with a prompt no other item shares — rolls into an "other" bucket below the clustered rows, so the reported counts always reconcile (sum(clusters.count) == total_failures).

Why it matters

Production runs fail in patterns. "47 of your 500 items failed" is useless on its own; "22 failed with tool_call_error on refund-processing prompts, 18 failed with budget_exceeded on research queries, 7 were one-offs" is actionable. You fix the two patterns, replay the failed items, and the run drains.

API

GET /v1/durable/runs/clusters

(Failure clusters are a top-level query over the durable-runs surface, not a per-run path.)

Response:

{
  "clusters": [
    {
      "cluster_key": "tool_call_error:a3f1b2c4d5e6f789",
      "error_code": "tool_call_error",
      "sample_prompt": "process refund for customer",
      "count": 22,
      "item_ids": ["item_1", "item_2", "item_3", "item_4", "item_5"]
    },
    {
      "cluster_key": "other",
      "error_code": "",
      "sample_prompt": "",
      "count": 3,
      "item_ids": []
    }
  ],
  "total_failures": 25
}
  • item_ids is capped at 5 per cluster as a representative sample. To drill into every item in a cluster, filter the items list by error_code.
  • sample_prompt is the first 200 characters of the canonicalized input — good enough to recognize the pattern, not enough to leak a full user prompt into summaries.
  • Results are cached server-side for 30 seconds with automatic invalidation when a new failure lands, so polling the endpoint while a run is in flight is cheap.

Algorithm limitations (v1)

The current grouping is a hash bucket, not semantic similarity. That means:

  • Prompts that differ by even one character in the first 200 chars won't cluster together, even if they mean the same thing semantically.
  • Agents that construct prompts by prepending per-item data (IDs, timestamps, customer names) will see lots of singletons in the "other" bucket. Template-based prompts (a constant preamble > 200 chars, varying detail at the end) cluster cleanly.

Embedding-based clustering — which groups by meaning rather than exact prefix — is on the post-launch roadmap. For v1, hash-bucket gets you the pattern for most runs where agents reuse prompt templates.