Triage (the needs-attention surface)
Triage is where the items a run couldn't finish clean wait for a decision. This is Papayya's dead-letter queue: retry an item, dismiss it, or fix your code and replay the whole not-ok slice. Once every flagged item in a run is resolved, the run promotes from partial to completed.
Why
A run over 10,000 leads finishes. 47 items didn't work — some transient rate limits, some bad inputs, one real bug in the agent code, and a handful that returned a 200 but came back empty or refused (an item's outcome is degraded, not failed, but it still didn't do the job). Without triage you get a green "completed" badge and the failures quietly age out. With triage the run stays partial, those 47 items show up in a list with their captured input and what went wrong, and you decide what to do with each (or each cluster).
Dead-letter queues are a standard primitive in queue systems (SQS, RabbitMQ, Kafka). Papayya carries the same idea one level down, at the item: failures and silent degrades don't disappear — they pile up somewhere you can see them and act on them.
What lands in triage
An item is awaiting triage when both are true:
- Its outcome isn't
ok— it either failed (raised) or came back degraded (ran but didn't work:empty_none,empty_string,refusal, or a reason you set withpapayya.mark_degraded(...)). - No operator decision has been recorded on it yet.
The captured input snapshot on the item is the replay source — Papayya grabbed it when the run minted the item, so it can be re-driven with the exact same payload.
Actions
Each flagged item has three operator moves.
Retry
Re-drives the single item through your live agent code. papayya triage retry <item_id> mints a new run for that one item, linked back to the source via replayed_from; the original item is marked resolved.
- If the retry works, the run progresses toward
completed. - If it comes back not-ok again, the new item lands back in triage. Resolve it independently.
Retry runs the agent code that's live right now — so fixing a real bug before you retry is how you recover from it.
Dismiss
papayya triage dismiss <item_id> resolves the item with no re-execution. Use it for items that will never work (corrupt data, out-of-scope input) or that you've already handled outside Papayya. Dismissing an item removes it from triage without claiming it succeeded — the run's rolled-up outcome still remembers it didn't.
Replay the slice
Retry is per item; replay re-drives every not-ok item in a run at once. This is the recovery verb from the quickstart, and it's the fastest way to recover from a fix that affects the whole run:
papayya replay --run <run_id> # re-drive all not-ok items into a NEW run
papayya replay --run <run_id> --tenant acme # just one tenant's slice
papayya replay --item <item_id> # a single itemReplay selects the items whose outcome wasn't ok, mints a new run linked to the old one via replayed_from, and re-drives each through your current code. Items that already worked are never touched.
Run status and triage
all items ok → completed (nothing in triage)
mixed → partial (items awaiting triage)
all items not-ok → failed (items in triage; status stays failed)A run also rolls its items up into worst_outcome_status (ok | degraded | failed) and a degraded_count, so the run row tells you the shape of the damage before you open it.
When a partial run's triage drains — every not-ok item retried or dismissed — the run promotes to completed automatically. A failed run stays failed even after triage: losing every item isn't a success, no matter how you close the books.
Dashboard
On the run detail page, a Needs attention section lists every unresolved not-ok item with:
- The item ID (links to the item's trace timeline)
- Its tenant (the
partition_keyit was minted with) - What went wrong — the raised error, or the degrade reason
- The input snapshot the item received
- Retry, Dismiss, and Replay controls
Locally, retry and replay invoke the papayya CLI in a subprocess so your agent code runs in an isolated process — the dashboard stays loosely coupled to your code.
In the hosted dashboard, the same actions dispatch through the platform's runtime worker pool, executing your deployed image — no host-side process required.
CLI
The whole triage workflow is reachable from the terminal, both locally with the free SDK and against a deployed project:
papayya triage list # items awaiting a decision
papayya triage retry <item_id> # re-drive one item into a new run
papayya triage dismiss <item_id> # resolve without re-runningAnd the slice-level recovery verb:
papayya replay --run <run_id> # re-drive every not-ok item in the runLocally, replay reads the run from the local durable store, discovers your agent file (agent.py in cwd by default; pass --file otherwise), looks up the matching @agent registration by name, and re-invokes it with each item's captured input snapshot. The original items are marked resolved and their new run carries replayed_from back to the source.
Retry one vs. replay the run
papayya triage retry <item_id>is per item, operator-triggered, and tracked — for when you want a decision on each failure.papayya replay --run <run_id>re-drives the whole not-ok slice in one shot — for when a code fix affects every failure and you want them all recovered at once. Add--tenant <key>to scope it to one tenant.
Both link each new run to its source via replayed_from, so the lineage is always traceable.
Hosted API
The hosted control-pane exposes the same workflow over HTTP for automation — list the not-ok items of a run, and retry, dismiss, or replay them. Retry and replay mint a new run cloned from the item's input snapshot and set replayed_from on it; dismiss records the decision without re-running. The mark is single-write: a second concurrent decision on the same item loses the race and returns 409 Conflict rather than overwriting the first, and on the hosted side the source mark and the new-run insert commit in the same transaction, so an observer never sees a half-replayed source. See the API reference for the exact endpoints and payloads.
Requirements for replay
Replay and retry only work when:
- The item has a captured input snapshot. Items minted before snapshot support don't have one and can't be re-driven — dismiss them instead.
- Your agent is discoverable (in
agent.pyin cwd, or pointed at via--file) and hasn't been removed.
The input snapshot is captured against your function's signature when the run mints the item: if the snapshot is a dict whose keys bind to the agent's parameters, it's unpacked as kwargs (fn(**snapshot)); otherwise it's passed as a single positional argument.
Data model
Triage state is a small set of fields on each item's durable item:
| Field | Meaning | Populated when |
|---|---|---|
input_snapshot | The payload the item received | When the run mints the item |
outcome | ok / degraded (with reason) / failed | When the item finishes |
| disposition | The operator decision: retried or dismissed | When someone triages the item |
resolved_at | Timestamp of that decision | Same transition |
replayed_from | On a run created by retry/replay — points back at the source run | On the new run |
An item leaves triage when a disposition is recorded. Because the mark is single-write, the durable item is the one place lineage and decisions can't drift out of sync.