Core Concepts
Triage & DLQ

Triage (the needs-attention surface)

Triage is where the items a run couldn't finish clean wait for a decision. This is Papayya's dead-letter queue: retry an item, dismiss it, or fix your code and replay the whole not-ok slice. Once every flagged item in a run is resolved, the run promotes from partial to completed.

Why

A run over 10,000 leads finishes. 47 items didn't work — some transient rate limits, some bad inputs, one real bug in the agent code, and a handful that returned a 200 but came back empty or refused (an item's outcome is degraded, not failed, but it still didn't do the job). Without triage you get a green "completed" badge and the failures quietly age out. With triage the run stays partial, those 47 items show up in a list with their captured input and what went wrong, and you decide what to do with each (or each cluster).

Dead-letter queues are a standard primitive in queue systems (SQS, RabbitMQ, Kafka). Papayya carries the same idea one level down, at the item: failures and silent degrades don't disappear — they pile up somewhere you can see them and act on them.

What lands in triage

An item is awaiting triage when both are true:

  • Its outcome isn't ok — it either failed (raised) or came back degraded (ran but didn't work: empty_none, empty_string, refusal, or a reason you set with papayya.mark_degraded(...)).
  • No operator decision has been recorded on it yet.

The captured input snapshot on the item is the replay source — Papayya grabbed it when the run minted the item, so it can be re-driven with the exact same payload.

Actions

Each flagged item has three operator moves.

Retry

Re-drives the single item through your live agent code. papayya triage retry <item_id> mints a new run for that one item, linked back to the source via replayed_from; the original item is marked resolved.

  • If the retry works, the run progresses toward completed.
  • If it comes back not-ok again, the new item lands back in triage. Resolve it independently.

Retry runs the agent code that's live right now — so fixing a real bug before you retry is how you recover from it.

Dismiss

papayya triage dismiss <item_id> resolves the item with no re-execution. Use it for items that will never work (corrupt data, out-of-scope input) or that you've already handled outside Papayya. Dismissing an item removes it from triage without claiming it succeeded — the run's rolled-up outcome still remembers it didn't.

Replay the slice

Retry is per item; replay re-drives every not-ok item in a run at once. This is the recovery verb from the quickstart, and it's the fastest way to recover from a fix that affects the whole run:

papayya replay --run <run_id>              # re-drive all not-ok items into a NEW run
papayya replay --run <run_id> --tenant acme  # just one tenant's slice
papayya replay --item <item_id>            # a single item

Replay selects the items whose outcome wasn't ok, mints a new run linked to the old one via replayed_from, and re-drives each through your current code. Items that already worked are never touched.

Run status and triage

all items ok        → completed         (nothing in triage)
mixed               → partial           (items awaiting triage)
all items not-ok    → failed            (items in triage; status stays failed)

A run also rolls its items up into worst_outcome_status (ok | degraded | failed) and a degraded_count, so the run row tells you the shape of the damage before you open it.

When a partial run's triage drains — every not-ok item retried or dismissed — the run promotes to completed automatically. A failed run stays failed even after triage: losing every item isn't a success, no matter how you close the books.

Dashboard

On the run detail page, a Needs attention section lists every unresolved not-ok item with:

  • The item ID (links to the item's trace timeline)
  • Its tenant (the partition_key it was minted with)
  • What went wrong — the raised error, or the degrade reason
  • The input snapshot the item received
  • Retry, Dismiss, and Replay controls

Locally, retry and replay invoke the papayya CLI in a subprocess so your agent code runs in an isolated process — the dashboard stays loosely coupled to your code.

In the hosted dashboard, the same actions dispatch through the platform's runtime worker pool, executing your deployed image — no host-side process required.

CLI

The whole triage workflow is reachable from the terminal, both locally with the free SDK and against a deployed project:

papayya triage list                     # items awaiting a decision
papayya triage retry <item_id>          # re-drive one item into a new run
papayya triage dismiss <item_id>        # resolve without re-running

And the slice-level recovery verb:

papayya replay --run <run_id>           # re-drive every not-ok item in the run

Locally, replay reads the run from the local durable store, discovers your agent file (agent.py in cwd by default; pass --file otherwise), looks up the matching @agent registration by name, and re-invokes it with each item's captured input snapshot. The original items are marked resolved and their new run carries replayed_from back to the source.

Retry one vs. replay the run

  • papayya triage retry <item_id> is per item, operator-triggered, and tracked — for when you want a decision on each failure.
  • papayya replay --run <run_id> re-drives the whole not-ok slice in one shot — for when a code fix affects every failure and you want them all recovered at once. Add --tenant <key> to scope it to one tenant.

Both link each new run to its source via replayed_from, so the lineage is always traceable.

Hosted API

The hosted control-pane exposes the same workflow over HTTP for automation — list the not-ok items of a run, and retry, dismiss, or replay them. Retry and replay mint a new run cloned from the item's input snapshot and set replayed_from on it; dismiss records the decision without re-running. The mark is single-write: a second concurrent decision on the same item loses the race and returns 409 Conflict rather than overwriting the first, and on the hosted side the source mark and the new-run insert commit in the same transaction, so an observer never sees a half-replayed source. See the API reference for the exact endpoints and payloads.

Requirements for replay

Replay and retry only work when:

  • The item has a captured input snapshot. Items minted before snapshot support don't have one and can't be re-driven — dismiss them instead.
  • Your agent is discoverable (in agent.py in cwd, or pointed at via --file) and hasn't been removed.

The input snapshot is captured against your function's signature when the run mints the item: if the snapshot is a dict whose keys bind to the agent's parameters, it's unpacked as kwargs (fn(**snapshot)); otherwise it's passed as a single positional argument.

Data model

Triage state is a small set of fields on each item's durable item:

FieldMeaningPopulated when
input_snapshotThe payload the item receivedWhen the run mints the item
outcomeok / degraded (with reason) / failedWhen the item finishes
dispositionThe operator decision: retried or dismissedWhen someone triages the item
resolved_atTimestamp of that decisionSame transition
replayed_fromOn a run created by retry/replay — points back at the source runOn the new run

An item leaves triage when a disposition is recorded. Because the mark is single-write, the durable item is the one place lineage and decisions can't drift out of sync.