Core Concepts
Items & Steps

Items & Steps

An item is one thing a run processed — one ticket, one document, one item. Each item has its own execution status, its own outcome (did it actually work?), its own step trace, and its own cost. An item is the unit you replay: when a run comes back partial, you re-drive the items that didn't work and leave the ones that did untouched.

A step is one node inside an item's trace — a single model call or tool call, checkpointed so a crash resumes from it instead of re-paying for it.

Two things an item carries: its input, and your name for it

They are separate fields and neither is derived from the other.

  • input is what your @agent function is called with — the ticket, the document, the item. Papayya stores it verbatim, and it is what a replay re-executes.
  • item_id is your id for this item: an order id, a ticket id, whatever you key your world on. Optional, at most 256 bytes, and never inferred from the input.
client.runs.create(agent_id="ag_...", items=[
    {"input": {"text": "..."}, "item_id": "co_42", "partition_key": "acme"},
])

Declaring item_id costs one field and buys two things: your own id appears everywhere Papayya shows you an item — the dashboard, papayya pull, the cohort table — instead of a uuid you have to cross-reference, and it is the key a signal will resolve on, so a thumbs-down handler on your side can say "this one was wrong" using the id it already has.

Leave it out and it is simply absent; nothing infers one for you, and every surface falls back to Papayya's own item id.

Item lifecycle

Every item goes through an execution state machine:

queued → running → completed
                 → failed
                 → cancelled
                 → budget_exceeded
  • queued — the item is in the work queue, waiting for a worker to pick it up.
  • running — a worker is actively executing the item's steps.
  • completed — the item finished and produced its result.
  • failed — an unrecoverable error was raised.
  • cancelled — the item was cancelled via the API or dashboard.
  • budget_exceeded — the item's run hit its cost cap before this item finished.

Ran vs. worked: the outcome

Execution status only tells you whether the item ran. A separate outcome tells you whether it worked — because a model can return a clean 200 that's still useless.

OutcomeMeaning
okThe item ran and the result passed every inspector.
degradedThe item ran (no exception, status completed) but the result didn't actually work. Carries a reason token.
failedThe item raised — same as execution status failed.

A degraded item is the wedge: no exception, a green execution status, and yet the work didn't land. The reason token says why:

  • empty_none — the model returned None.
  • empty_string — the model returned empty or whitespace-only content.
  • refusal — the model refused ("I can't help with that").
  • …plus any reason you declare yourself with papayya.mark_degraded("reason") from anywhere under the loop, for the things an inspector can't see.

Outcomes are set two ways: @papayya.llm runs inspectors on every model return automatically, and papayya.mark_degraded(...) lets your own logic flag an item explicitly. The run rolls all of this up into worst_outcome_status and degraded_count — see Runs.

What is a step?

A step is one node inside an item's trace — the smallest unit Papayya checkpoints. Each step items its input, its output, its duration, and its status, so a crash mid-item resumes from the last completed step instead of replaying the whole item.

LLM step

The function that calls your model, marked with @papayya.llm (or created with .llm_step(...) on an explicit item handle). It items:

  • The model, tokens, and timing of the call.
  • The response, run through the outcome inspectors — a refusal or empty result flips the item to degraded, with no check written anywhere.

Tool / work step

Any other checkpointed unit of work — a fetch, an enrichment, a transform. Created with .step(...) on an explicit item handle. Items its input, output, duration, and status.

.step() is the current spelling; .task() still works as an alias for the legacy name. New code should use .step().

Checkpointing and replay-on-crash

Every step's result is cached in the ledger keyed by its id. Re-running an item with the same step ids replays completed steps from cache instead of re-executing them — so a process that dies after an expensive model call doesn't pay for that call twice on resume. This is why steps exist: they're the seam where durability lives.

item = papayya().item("enrich", item_id="co_42")
fetch   = item.step("fetch", fetch_fn)          # checkpointed
extract = item.llm_step("extract", extract_fn)  # checkpointed + token/outcome capture
snippet = fetch(domain)
fields  = extract(name, snippet)
item.complete(fields)

Reach for the explicit handle when you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls. For the common case, @papayya.llm on your model function plus run.step(...) inside an @agent body gives you the same step capture without the handle.

Tracking

Every item records:

FieldDescription
outcome_statusok, degraded, or failed
degraded_reasonThe reason token, when degraded
total_input_tokensCumulative input tokens across all steps
total_output_tokensCumulative output tokens across all steps
total_cost_centsCumulative cost in cents (integer math)
started_atWhen the first step began
completed_atWhen the item reached a terminal state

Replaying an item

Because each item is durable and individually addressable, you replay at the item grain:

papayya replay --item <item_id>        # re-drive one item
papayya replay --run <run_id>          # re-drive every not-ok item in a run
papayya replay --run <run_id> --tenant acme  # just one tenant's slice

Replay selects the items whose outcome wasn't ok, mints a new run linked to the old one via replayed_from, and re-drives each through your current code. Items that already worked are never touched.

Viewing items

  • Dashboard — a run's detail page lists its items, each with a worked/degraded badge, its tenant, and its full step trace, in the hosted dashboard at https://app.getpapayya.com (opens in a new tab) — whether you ran python agent.py locally or deployed to the worker pool.
  • CLI — papayya items list, papayya items get <item_id>, papayya items stream <item_id> to live-tail one.
  • API — GET /v1/durable/runs/{itemId} for the item, GET /v1/durable/runs/{itemId}/checkpoints for its step trace. (The wire path is transitional; each durable/runs item is one item.)