Papayya
Run your AI agents and batch pipelines over hundreds of items without losing money on failures you can't see.
Papayya is ready-made infrastructure for AI pipelines — KB ingestion, nightly evals, lead enrichment, document extraction, conversation post-processing, codemod runs. It makes silent, partial, non-deterministic failure visible and recoverable, per item. Checkpoint-and-resume of the execution it owns is the mechanism — it's what lets this work run out-of-band (cron batches and long-running agent loops, where no one is watching) and still actually finish when 3% of items quietly fail.
A handful of words carry all of Papayya, and these docs use them precisely: an agent is loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. (A step is one node inside an item's trace; a tenant is whoever an item belongs to, declared via partition_key.)
Without Papayya: container exited 1.
With Papayya: this item spent $4.20 thrashing on step 7 because the tool returned malformed JSON — and 3% of items in last night's run hit the same pattern.
Why Papayya?
Periodic LLM jobs fail in ways your monitoring misses. Each one turns a 5-minute root cause into a multi-week support thread:
- Silent partial failure — 30 of last night's 1,000 items failed. The cron exited 0. You can't tell which 30.
- Rate-limit poisoning — one tenant always runs last, so that tenant always has bad data. Nobody connects the dots.
- Cost runaway at the long tail — one bad input retried until $4 became $400. You found out at month-end.
- "Why didn't X happen?" — you can't replay last Tuesday's run against today's prompt.
- Starting over at turn 1 — an agent fails at turn 18 of 20, and the retry re-runs, and re-pays for, the 17 that worked. Long-running agents measures the difference.
- Backfill blast radius — one CLI re-runs N days of LLM cost with no dry-run, no estimate, no confirmation.
Papayya makes the consequences visible and recoverable in seconds instead of hours. That's the wedge.
How it works
Wrap the loop you already have. One pip install, no SDK lock-in — Papayya never touches your LLM client.
import papayya
from papayya import agent
@papayya.llm # mark the function that calls your model
def classify(text: str) -> dict:
... # your provider, your key
@agent(name="triage-ticket", model="claude-sonnet-5")
def triage(run, ticket: dict) -> str:
result = run.step("classify", classify, item_id=ticket["id"])(ticket["text"])
label = result["content"]
if not label:
papayya.mark_degraded("classifier returned nothing")
run.complete({"ticket": ticket["id"], "label": label})
return label
for ticket in TICKETS: # your loop, unchanged
print(triage(
ticket,
item_id=ticket["id"], # your identity for this item
partition_key=ticket["tenant"], # whose item it is
) or "(refused)")Each triage(...) call is one run over one item, with its own steps, outcome and cost. item_id= and partition_key= are consumed by Papayya and are not passed to your function, so you declare identity and tenant at the call site without threading them through a signature.
@papayya.llm marks the leaf that calls your model — it records each call as an LLM step (model, tokens, timing) and inspects the response, so a refusal or empty result flips the item to degraded even though the call returned a 200 and raised nothing. papayya.mark_degraded(reason) is the same verdict, stated by you.
Run it with python agent.py. papayya deploy then hands the same file to the managed worker pool, where papayya run triage-ticket '{"id": "t-9", ...}' --item-id t-9 --partition-key acme invokes it.
When you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls — reach for the explicit handle:
from papayya.durable import papayya
item = papayya().item("enrich", item_id="co_42")
fetch = item.step("fetch", fetch_fn) # checkpointed
extract = item.llm_step("extract", extract_fn) # checkpointed + token/outcome capture
snippet = fetch(item_domain)
fields = extract(item_name, snippet)
item.complete(fields)Each step's result is cached in the ledger. If your process crashes and restarts, completed steps replay from cache instead of re-executing — you don't re-pay for LLM calls that already landed. (item.step() is the current spelling; the legacy .task() is kept as an alias.)
When PAPAYYA_API_KEY is set, the ledger round-trips through the Papayya Cloud control plane. For local iteration, python agent.py runs against a durable store with the free SDK and consumes no hosted compute.
Run it locally
Run python agent.py locally with the free SDK — it consumes no hosted compute:
pip install papayya
papayya example
python agent.pyFor a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint. View runs, items, and steps in the hosted dashboard at app.getpapayya.com (opens in a new tab).
Deploy to the cloud
Register the function as a deployable agent, then push it to Papayya's hosted worker pool:
# agent.py
from papayya import agent
from openai import OpenAI
@agent(name="research-bot")
def research_bot(prompt: str) -> str:
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
return response.choices[0].message.contentpapayya signup
papayya secrets set OPENAI_API_KEY sk-...
papayya deploy # auto-discovers @agent functions, no flags needed
papayya run research-bot "Tell me about Anthropic"@agent(name=...) registers the function as a deployable agent. The runtime reserves cost before each LLM call and pauses the run when a budget cap is hit, preserving in-flight items for you to resume. The hosted dashboard at app.getpapayya.com (opens in a new tab) shows the runs → items → steps hierarchy.
@papayya.durable(name=..., budget_usd=...) is an alternate spelling that also sets a per-agent budget.
What you get
| Feature | Description |
|---|---|
| Per-item visibility | Every item carries its own item_id, outcome, trace, and cost, so you can answer "why didn't X happen?" without grep |
| Ran-vs-worked outcomes | Each item is ok or degraded (with a reason like empty_none or refusal) — a 200 that didn't actually work is still caught |
| Failure clustering | When 47 of 1,000 items don't work, see them grouped into a few patterns — not 47 separate stack traces |
| Per-item cost outliers | Find the one bad input that ate $4 in retries before your bill does |
| Crash recovery | Checkpoint after every step; resume mid-item after crashes without re-paying completed calls |
| Budget enforcement | Pause-and-notify on cap, not just alert-on-cap (cloud only) |
| Slice replay | Re-drive just the not-ok items of a run — or one tenant's slice, or a single item — into a new run linked via replayed_from |
| Schedules + triggers | Run on cron or fire from external events, with per-agent budgets |
| Bring your own framework | OpenAI, Anthropic, Bedrock, raw HTTP — Papayya never wraps your LLM client |
Get started
Jump to the Quickstart to have an agent running in under a minute. Or browse runnable starter agents at github.com/papayya-zero/examples (opens in a new tab) — lead enrichment, document extraction, eval harness.