Core Concepts
Durability

Durability

Papayya's wedge is making silent, partial, and degraded failure visible and recoverable — per tenant, per item. Durability is the mechanism that serves that wedge: because Papayya owns each item's execution and checkpoints every step, a run's items survive crashes, deploys, and restarts — and resume without re-paying for the work already done. That same owned execution is what lets Papayya run your work out of band — cron batches and long-running agent loops, where no one is watching the request — and still see whether each item actually worked.

Crash-resume happens at the step level, inside one item. A run processes many items; each item has its own step trace; each completed step is cached. Re-running an item with the same item_id replays completed steps from cache instead of re-executing them.

Execution guarantees

Papayya provides at-least-once execution. This means:

  • Every step will execute at least once
  • If a crash occurs between executing a step and saving its checkpoint, the step may execute again on resume
  • Steps that have side effects (sending emails, charging credit cards, writing to databases) should be idempotent — safe to run more than once

Exactly-once execution is not possible in distributed systems without two-phase commit.

Rule of thumb: Design your steps so that running them twice produces the same result as running them once.

A retry and a re-execution are not the same thing

Two different events put a step through your code a second time, and Papayya tells them apart. The difference is the money.

  • A retry is a step whose execution left no item — the worker died after the LLM call returned but before the checkpoint was saved. Papayya has no way to know the call happened, so on resume the step runs again.
  • A re-execution is a step Papayya does have an item for, run again deliberately — you replayed it because the result was wrong.

item.idempotency_key() hands your provider a token, and the token tracks that difference:

key = item.idempotency_key("extract")
extract = item.llm_step("extract", lambda: client.messages.create(
    ..., extra_headers={"Idempotency-Key": key}))
fields = extract()

A retry sends the same token, so the provider replays its stored response and you are not billed twice for a call that already succeeded — that is what the token is for. A deliberate re-execution sends a new one, so the provider genuinely redoes the work, which is what you asked for by replaying it.

Call it immediately before the step it protects. This is a seam to your provider's own idempotency, not exactly-once: Papayya cannot dedupe a side effect it does not own.

How checkpointing works

Cloud

Worker leases an item from the queue
  → Launches the agent code in the worker pool
    → Runs the item, reporting every step to the control plane
      → Step persisted to the durable store
        → If the worker crashes, it stops heartbeating and its lease expires
          → The reaper re-queues that item under a fresh lease
            → A fresh worker picks it up and replays from the last persisted step

Recovery is per item, not per run: the lease is the item's, so a crash re-queues only the items that worker held in flight. Everything already checkpointed replays from cache, so a crash costs you the step that was mid-flight and nothing else.

Local

item.step("label", fn) called
  → Check cache: if this label already ran for this item_id, return the cached result
  → Execute fn()
  → Save the checkpoint to the local durable store
  → On process restart with the same item_id: cached steps skip, uncached steps re-execute

Important: The checkpoint is saved after the function executes. If your process crashes between execution and checkpoint save, the function runs again on the next attempt.

The explicit handle

When you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls — reach for the explicit handle:

import papayya
 
item = papayya().item("enrich", item_id="co_42")
fetch   = item.step("fetch", fetch_fn)          # checkpointed
extract = item.llm_step("extract", extract_fn)  # checkpointed + token/outcome capture
snippet = fetch(domain)
fields  = extract(name, snippet)
item.complete(fields)

Each step's result is cached in the ledger keyed by item_id + label. Re-run with the same item_id and the completed steps replay from cache; only the uncached ones re-execute. item.step() is the current spelling of the legacy .task(); item.llm_step() adds token and outcome capture on top of the checkpoint.

Worker model (Cloud)

The cloud path uses stateless workers:

  • Workers pull runs from the queue
  • The worker pool loads the agent once and drives its items
  • Each item reports its steps via HTTP to the control plane
  • If a worker dies, heartbeat-based detection marks the affected in-flight items for recovery within 60 seconds
  • The recovery sweep re-enqueues orphaned work

Workers can be scaled horizontally, deployed, or restarted without losing progress on any run.

Locking (Cloud)

To prevent two workers from executing the same run concurrently:

  • A worker acquires an exclusive, time-bounded lease on a run before picking it up
  • The lease is acquired atomically, so no two workers can ever hold the same run at once
  • If a worker crashes, its lease expires and a recovery sweep re-enqueues the run for a fresh worker

What gets persisted

DataStoragePurpose
Run stateDurable storeSource of truth for the run lifecycle over N items
Item stateDurable storePer-item execution status, outcome, and cost
Step results / checkpointsDurable storeExecution trace and replay
Tool call I/ODurable storeDebugging
Queue notificationsThe queueWork distribution (rebuildable from the durable store)

The queue is a notification layer only. If it's lost, it can be rebuilt from the durable store. No run or item data is lost.

Per-item snapshots

Every step belongs to an item, and Papayya captures the input and output state of that item alongside the checkpoint:

Column (on steps)SourcePurpose
item_idThe item's identity (your_agent(input, item_id=...), run.step(..., item_id=), or .item(name, item_id=))Group steps across runs by the item they operated on
input_snapshotAuto-captured from the wrapped fn's call args (or snapshot=... to override / snapshot=False to opt out)The item's state before the step ran
output_snapshotThe step function's return valueThe item's state after the step ran

Snapshots are JSON-encoded and capped at 64 KB per column — pass a reference (e.g. an S3 key) rather than a large payload if you need more. The dashboard surfaces these as an /item/:id timeline with a field-level diff between input and output.

Budget enforcement

Budget is enforced at step boundaries:

  • Before each step, the system checks whether the run's budget is exhausted
  • A single step (one LLM call) may exceed the remaining budget — the check happens before the call, but the call's cost isn't known until it completes
  • After the step, if the budget is exceeded, the run pauses and notifies (status paused, pause_reason="budget") and the SDK raises WorkloadPaused at the next step boundary — the completed step's checkpoint is preserved, never discarded

This means the actual spend may exceed the budget by the cost of one step. For most use cases this is a few cents. Set your budget with this margin in mind. See Budget Enforcement for the full lifecycle.

Pause and resume

A pause is not a failure and not a stop. It is Papayya declining to keep spending on your behalf until you look — a budget breach, a streak of degraded steps, or a provider reporting exhausted credits. It always lands on a step boundary, after the step that triggered it was safely checkpointed.

What happens, in order:

  1. The step completes and its checkpoint is written. Nothing in flight is discarded.
  2. The run's status becomes paused with a pause_reason, and the SDK raises WorkloadPaused (or CreditExhausted) at the next step boundary.
  3. The worker hands the item back to the queue, parked. It does not report the item finished. A parked item is not leasable — no worker will pick it up, and it will not time out and quietly restart.
  4. Resume unparks it. A worker leases it and continues from exactly where the pause landed — re-running the steps the fence objected to, and replaying everything else from its checkpoint.

Two properties follow from step 3, and they are the reason the pause is worth having:

  • Nothing runs while you decide. The item is out of the pool, not in it with a timer.
  • Nothing is re-paid for. Resume is a continuation, not a restart. Steps that already cost you money and were fine are cache hits.

Resume the run when you've addressed the cause — raised the budget, fixed the prompt, topped up the provider account. There is no timer that resumes it for you: Papayya cannot know when the thing that stopped it is fixed, so the resume is the signal.

What re-runs, and what doesn't

Fix the cause first, then resume. The re-run is what carries your fix into the work — a resume taken before the fix just reproduces the same outputs.

A degradation pause re-runs exactly the steps that formed the streak it tripped on. Not the whole run: steps the fence was happy with stay cache hits, so a resume never re-pays for work that was already good. The set is decided when the fence trips, not when you resume, so it is the work that was actually objected to rather than whatever the run looks like later.

Two kinds of step are counted by the fence but deliberately not re-run:

  • Steps you marked yourself with papayya.mark_degraded — there is no function to call again.
  • Steps that returned None. A step with no return value is usually a side effect (send an email, write a row), and re-running it would do that thing twice. Papayya has no way to know whether your side effect is safe to repeat, so it doesn't guess. If you want a re-run, return something.

A budget pause re-runs nothing at all. Its outputs were fine; the cap was the problem. Raise the cap and the run continues from where it stopped.

The resume response reports how many steps will re-run, so you can tell a resume that will change something from one that won't before you take it. If it's zero and you need different output, replay instead.

Resuming also clears the degraded-step streak. Without that, a run resumed after a degradation fence would re-pause on its very next imperfect step, which is a loop rather than a recovery.

Setting the threshold

The degradation fence pauses after 3 consecutive degraded steps by default. It is a per-agent setting, not a per-run one — it is an operator control, and the dashboard shows it on the agent:

papayya agents update <agent-id> --config '{"pause_after_degraded": 5}'

Set it to 0 to turn the run-level fence off for that agent. Passing pause_after_degraded to an individual run has no effect on the hosted path; the SDK logs a warning if you try.

Resume vs replay. Resume continues the same run from its last checkpoint — reach for it when the work is still valid and something external needs fixing. Replay mints a new run that re-drives the original's item from the start — reach for it when you've changed the code or the prompt and want the work done differently. See Recovery.

Durability by path

FeatureCloud (@agent + deploy)Local (item.step())
CheckpointingEvery step to the durable storeEvery item.step() to the local durable store
Crash recoveryHeartbeat timeout + recovery sweepResume the item with the same item_id
Budget enforcementAutomatic — pauses + notifies on exceedPer-step check, pauses and raises WorkloadPaused
Execution traceFull (steps with tokens, cost, tool calls)Checkpoints (labels + results)
Resume after restartAutomaticSame item_id