Durability
Papayya's wedge is making silent, partial, and degraded failure visible and recoverable — per tenant, per item. Durability is the mechanism that serves that wedge: because Papayya owns each item's execution and checkpoints every step, a run's items survive crashes, deploys, and restarts — and resume without re-paying for the work already done. That same owned execution is what lets Papayya run your work out of band — cron batches and long-running agent loops, where no one is watching the request — and still see whether each item actually worked.
Crash-resume happens at the step level, inside one item. A run processes many items; each item has its own step trace; each completed step is cached. Re-running an item with the same item_id replays completed steps from cache instead of re-executing them.
Execution guarantees
Papayya provides at-least-once execution. This means:
- Every step will execute at least once
- If a crash occurs between executing a step and saving its checkpoint, the step may execute again on resume
- Steps that have side effects (sending emails, charging credit cards, writing to databases) should be idempotent — safe to run more than once
Exactly-once execution is not possible in distributed systems without two-phase commit.
Rule of thumb: Design your steps so that running them twice produces the same result as running them once.
A retry and a re-execution are not the same thing
Two different events put a step through your code a second time, and Papayya tells them apart. The difference is the money.
- A retry is a step whose execution left no item — the worker died after the LLM call returned but before the checkpoint was saved. Papayya has no way to know the call happened, so on resume the step runs again.
- A re-execution is a step Papayya does have an item for, run again deliberately — you replayed it because the result was wrong.
item.idempotency_key() hands your provider a token, and the token tracks that difference:
key = item.idempotency_key("extract")
extract = item.llm_step("extract", lambda: client.messages.create(
..., extra_headers={"Idempotency-Key": key}))
fields = extract()A retry sends the same token, so the provider replays its stored response and you are not billed twice for a call that already succeeded — that is what the token is for. A deliberate re-execution sends a new one, so the provider genuinely redoes the work, which is what you asked for by replaying it.
Call it immediately before the step it protects. This is a seam to your provider's own idempotency, not exactly-once: Papayya cannot dedupe a side effect it does not own.
How checkpointing works
Cloud
Worker leases an item from the queue
→ Launches the agent code in the worker pool
→ Runs the item, reporting every step to the control plane
→ Step persisted to the durable store
→ If the worker crashes, it stops heartbeating and its lease expires
→ The reaper re-queues that item under a fresh lease
→ A fresh worker picks it up and replays from the last persisted stepRecovery is per item, not per run: the lease is the item's, so a crash re-queues only the items that worker held in flight. Everything already checkpointed replays from cache, so a crash costs you the step that was mid-flight and nothing else.
Local
item.step("label", fn) called
→ Check cache: if this label already ran for this item_id, return the cached result
→ Execute fn()
→ Save the checkpoint to the local durable store
→ On process restart with the same item_id: cached steps skip, uncached steps re-executeImportant: The checkpoint is saved after the function executes. If your process crashes between execution and checkpoint save, the function runs again on the next attempt.
The explicit handle
When you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls — reach for the explicit handle:
import papayya
item = papayya().item("enrich", item_id="co_42")
fetch = item.step("fetch", fetch_fn) # checkpointed
extract = item.llm_step("extract", extract_fn) # checkpointed + token/outcome capture
snippet = fetch(domain)
fields = extract(name, snippet)
item.complete(fields)Each step's result is cached in the ledger keyed by item_id + label. Re-run with the same item_id and the completed steps replay from cache; only the uncached ones re-execute. item.step() is the current spelling of the legacy .task(); item.llm_step() adds token and outcome capture on top of the checkpoint.
Worker model (Cloud)
The cloud path uses stateless workers:
- Workers pull runs from the queue
- The worker pool loads the agent once and drives its items
- Each item reports its steps via HTTP to the control plane
- If a worker dies, heartbeat-based detection marks the affected in-flight items for recovery within 60 seconds
- The recovery sweep re-enqueues orphaned work
Workers can be scaled horizontally, deployed, or restarted without losing progress on any run.
Locking (Cloud)
To prevent two workers from executing the same run concurrently:
- A worker acquires an exclusive, time-bounded lease on a run before picking it up
- The lease is acquired atomically, so no two workers can ever hold the same run at once
- If a worker crashes, its lease expires and a recovery sweep re-enqueues the run for a fresh worker
What gets persisted
| Data | Storage | Purpose |
|---|---|---|
| Run state | Durable store | Source of truth for the run lifecycle over N items |
| Item state | Durable store | Per-item execution status, outcome, and cost |
| Step results / checkpoints | Durable store | Execution trace and replay |
| Tool call I/O | Durable store | Debugging |
| Queue notifications | The queue | Work distribution (rebuildable from the durable store) |
The queue is a notification layer only. If it's lost, it can be rebuilt from the durable store. No run or item data is lost.
Per-item snapshots
Every step belongs to an item, and Papayya captures the input and output state of that item alongside the checkpoint:
Column (on steps) | Source | Purpose |
|---|---|---|
item_id | The item's identity (your_agent(input, item_id=...), run.step(..., item_id=), or .item(name, item_id=)) | Group steps across runs by the item they operated on |
input_snapshot | Auto-captured from the wrapped fn's call args (or snapshot=... to override / snapshot=False to opt out) | The item's state before the step ran |
output_snapshot | The step function's return value | The item's state after the step ran |
Snapshots are JSON-encoded and capped at 64 KB per column — pass a reference (e.g. an S3 key) rather than a large payload if you need more. The dashboard surfaces these as an /item/:id timeline with a field-level diff between input and output.
Budget enforcement
Budget is enforced at step boundaries:
- Before each step, the system checks whether the run's budget is exhausted
- A single step (one LLM call) may exceed the remaining budget — the check happens before the call, but the call's cost isn't known until it completes
- After the step, if the budget is exceeded, the run pauses and notifies (status
paused,pause_reason="budget") and the SDK raisesWorkloadPausedat the next step boundary — the completed step's checkpoint is preserved, never discarded
This means the actual spend may exceed the budget by the cost of one step. For most use cases this is a few cents. Set your budget with this margin in mind. See Budget Enforcement for the full lifecycle.
Pause and resume
A pause is not a failure and not a stop. It is Papayya declining to keep spending on your behalf until you look — a budget breach, a streak of degraded steps, or a provider reporting exhausted credits. It always lands on a step boundary, after the step that triggered it was safely checkpointed.
What happens, in order:
- The step completes and its checkpoint is written. Nothing in flight is discarded.
- The run's status becomes
pausedwith apause_reason, and the SDK raisesWorkloadPaused(orCreditExhausted) at the next step boundary. - The worker hands the item back to the queue, parked. It does not report the item finished. A parked item is not leasable — no worker will pick it up, and it will not time out and quietly restart.
- Resume unparks it. A worker leases it and continues from exactly where the pause landed — re-running the steps the fence objected to, and replaying everything else from its checkpoint.
Two properties follow from step 3, and they are the reason the pause is worth having:
- Nothing runs while you decide. The item is out of the pool, not in it with a timer.
- Nothing is re-paid for. Resume is a continuation, not a restart. Steps that already cost you money and were fine are cache hits.
Resume the run when you've addressed the cause — raised the budget, fixed the prompt, topped up the provider account. There is no timer that resumes it for you: Papayya cannot know when the thing that stopped it is fixed, so the resume is the signal.
What re-runs, and what doesn't
Fix the cause first, then resume. The re-run is what carries your fix into the work — a resume taken before the fix just reproduces the same outputs.
A degradation pause re-runs exactly the steps that formed the streak it tripped on. Not the whole run: steps the fence was happy with stay cache hits, so a resume never re-pays for work that was already good. The set is decided when the fence trips, not when you resume, so it is the work that was actually objected to rather than whatever the run looks like later.
Two kinds of step are counted by the fence but deliberately not re-run:
- Steps you marked yourself with
papayya.mark_degraded— there is no function to call again. - Steps that returned
None. A step with no return value is usually a side effect (send an email, write a row), and re-running it would do that thing twice. Papayya has no way to know whether your side effect is safe to repeat, so it doesn't guess. If you want a re-run, return something.
A budget pause re-runs nothing at all. Its outputs were fine; the cap was the problem. Raise the cap and the run continues from where it stopped.
The resume response reports how many steps will re-run, so you can tell a resume that will change something from one that won't before you take it. If it's zero and you need different output, replay instead.
Resuming also clears the degraded-step streak. Without that, a run resumed after a degradation fence would re-pause on its very next imperfect step, which is a loop rather than a recovery.
Setting the threshold
The degradation fence pauses after 3 consecutive degraded steps by default. It is a per-agent setting, not a per-run one — it is an operator control, and the dashboard shows it on the agent:
papayya agents update <agent-id> --config '{"pause_after_degraded": 5}'Set it to 0 to turn the run-level fence off for that agent. Passing pause_after_degraded to an individual run has no effect on the hosted path; the SDK logs a warning if you try.
Resume vs replay. Resume continues the same run from its last checkpoint — reach for it when the work is still valid and something external needs fixing. Replay mints a new run that re-drives the original's item from the start — reach for it when you've changed the code or the prompt and want the work done differently. See Recovery.
Durability by path
| Feature | Cloud (@agent + deploy) | Local (item.step()) |
|---|---|---|
| Checkpointing | Every step to the durable store | Every item.step() to the local durable store |
| Crash recovery | Heartbeat timeout + recovery sweep | Resume the item with the same item_id |
| Budget enforcement | Automatic — pauses + notifies on exceed | Per-step check, pauses and raises WorkloadPaused |
| Execution trace | Full (steps with tokens, cost, tool calls) | Checkpoints (labels + results) |
| Resume after restart | Automatic | Same item_id |