Items & Steps
An item is one thing a run processed — one ticket, one document, one item. Each item has its own execution status, its own outcome (did it actually work?), its own step trace, and its own cost. An item is the unit you replay: when a run comes back partial, you re-drive the items that didn't work and leave the ones that did untouched.
A step is one node inside an item's trace — a single model call or tool call, checkpointed so a crash resumes from it instead of re-paying for it.
Two things an item carries: its input, and your name for it
They are separate fields and neither is derived from the other.
inputis what your@agentfunction is called with — the ticket, the document, the item. Papayya stores it verbatim, and it is what a replay re-executes.item_idis your id for this item: an order id, a ticket id, whatever you key your world on. Optional, at most 256 bytes, and never inferred from the input.
client.runs.create(agent_id="ag_...", items=[
{"input": {"text": "..."}, "item_id": "co_42", "partition_key": "acme"},
])Declaring item_id costs one field and buys two things: your own id appears everywhere Papayya shows you an item — the dashboard, papayya pull, the cohort table — instead of a uuid you have to cross-reference, and it is the key a signal will resolve on, so a thumbs-down handler on your side can say "this one was wrong" using the id it already has.
Leave it out and it is simply absent; nothing infers one for you, and every surface falls back to Papayya's own item id.
Item lifecycle
Every item goes through an execution state machine:
queued → running → completed
→ failed
→ cancelled
→ budget_exceeded- queued — the item is in the work queue, waiting for a worker to pick it up.
- running — a worker is actively executing the item's steps.
- completed — the item finished and produced its result.
- failed — an unrecoverable error was raised.
- cancelled — the item was cancelled via the API or dashboard.
- budget_exceeded — the item's run hit its cost cap before this item finished.
Ran vs. worked: the outcome
Execution status only tells you whether the item ran. A separate outcome tells you whether it worked — because a model can return a clean 200 that's still useless.
| Outcome | Meaning |
|---|---|
ok | The item ran and the result passed every inspector. |
degraded | The item ran (no exception, status completed) but the result didn't actually work. Carries a reason token. |
failed | The item raised — same as execution status failed. |
A degraded item is the wedge: no exception, a green execution status, and yet the work didn't land. The reason token says why:
empty_none— the model returnedNone.empty_string— the model returned empty or whitespace-only content.refusal— the model refused ("I can't help with that").- …plus any reason you declare yourself with
papayya.mark_degraded("reason")from anywhere under the loop, for the things an inspector can't see.
Outcomes are set two ways: @papayya.llm runs inspectors on every model return automatically, and papayya.mark_degraded(...) lets your own logic flag an item explicitly. The run rolls all of this up into worst_outcome_status and degraded_count — see Runs.
What is a step?
A step is one node inside an item's trace — the smallest unit Papayya checkpoints. Each step items its input, its output, its duration, and its status, so a crash mid-item resumes from the last completed step instead of replaying the whole item.
LLM step
The function that calls your model, marked with @papayya.llm (or created with .llm_step(...) on an explicit item handle). It items:
- The model, tokens, and timing of the call.
- The response, run through the outcome inspectors — a refusal or empty result flips the item to
degraded, with no check written anywhere.
Tool / work step
Any other checkpointed unit of work — a fetch, an enrichment, a transform. Created with .step(...) on an explicit item handle. Items its input, output, duration, and status.
.step()is the current spelling;.task()still works as an alias for the legacy name. New code should use.step().
Checkpointing and replay-on-crash
Every step's result is cached in the ledger keyed by its id. Re-running an item with the same step ids replays completed steps from cache instead of re-executing them — so a process that dies after an expensive model call doesn't pay for that call twice on resume. This is why steps exist: they're the seam where durability lives.
item = papayya().item("enrich", item_id="co_42")
fetch = item.step("fetch", fetch_fn) # checkpointed
extract = item.llm_step("extract", extract_fn) # checkpointed + token/outcome capture
snippet = fetch(domain)
fields = extract(name, snippet)
item.complete(fields)Reach for the explicit handle when you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls. For the common case, @papayya.llm on your model function plus run.step(...) inside an @agent body gives you the same step capture without the handle.
Tracking
Every item records:
| Field | Description |
|---|---|
outcome_status | ok, degraded, or failed |
degraded_reason | The reason token, when degraded |
total_input_tokens | Cumulative input tokens across all steps |
total_output_tokens | Cumulative output tokens across all steps |
total_cost_cents | Cumulative cost in cents (integer math) |
started_at | When the first step began |
completed_at | When the item reached a terminal state |
Replaying an item
Because each item is durable and individually addressable, you replay at the item grain:
papayya replay --item <item_id> # re-drive one item
papayya replay --run <run_id> # re-drive every not-ok item in a run
papayya replay --run <run_id> --tenant acme # just one tenant's sliceReplay selects the items whose outcome wasn't ok, mints a new run linked to the old one via replayed_from, and re-drives each through your current code. Items that already worked are never touched.
Viewing items
- Dashboard — a run's detail page lists its items, each with a worked/degraded badge, its tenant, and its full step trace, in the hosted dashboard at https://app.getpapayya.com (opens in a new tab) — whether you ran
python agent.pylocally or deployed to the worker pool. - CLI —
papayya items list,papayya items get <item_id>,papayya items stream <item_id>to live-tail one. - API —
GET /v1/durable/runs/{itemId}for the item,GET /v1/durable/runs/{itemId}/checkpointsfor its step trace. (The wire path is transitional; eachdurable/runsitem is one item.)