Introduction

Papayya

Run your AI agents and batch pipelines over hundreds of items without losing money on failures you can't see.

Papayya is ready-made infrastructure for AI pipelines — KB ingestion, nightly evals, lead enrichment, document extraction, conversation post-processing, codemod runs. It makes silent, partial, non-deterministic failure visible and recoverable, per item. Checkpoint-and-resume of the execution it owns is the mechanism — it's what lets this work run out-of-band (cron batches and long-running agent loops, where no one is watching) and still actually finish when 3% of items quietly fail.

A handful of words carry all of Papayya, and these docs use them precisely: an agent is loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. (A step is one node inside an item's trace; a tenant is whoever an item belongs to, declared via partition_key.)

Without Papayya: container exited 1.

With Papayya: this item spent $4.20 thrashing on step 7 because the tool returned malformed JSON — and 3% of items in last night's run hit the same pattern.

Why Papayya?

Periodic LLM jobs fail in ways your monitoring misses. Each one turns a 5-minute root cause into a multi-week support thread:

  • Silent partial failure — 30 of last night's 1,000 items failed. The cron exited 0. You can't tell which 30.
  • Rate-limit poisoning — one tenant always runs last, so that tenant always has bad data. Nobody connects the dots.
  • Cost runaway at the long tail — one bad input retried until $4 became $400. You found out at month-end.
  • "Why didn't X happen?" — you can't replay last Tuesday's run against today's prompt.
  • Starting over at turn 1 — an agent fails at turn 18 of 20, and the retry re-runs, and re-pays for, the 17 that worked. Long-running agents measures the difference.
  • Backfill blast radius — one CLI re-runs N days of LLM cost with no dry-run, no estimate, no confirmation.

Papayya makes the consequences visible and recoverable in seconds instead of hours. That's the wedge.

How it works

Wrap the loop you already have. One pip install, no SDK lock-in — Papayya never touches your LLM client.

import papayya
from papayya import agent
 
@papayya.llm                       # mark the function that calls your model
def classify(text: str) -> dict:
    ...                            # your provider, your key
 
@agent(name="triage-ticket", model="claude-sonnet-5")
def triage(run, ticket: dict) -> str:
    result = run.step("classify", classify, item_id=ticket["id"])(ticket["text"])
    label = result["content"]
    if not label:
        papayya.mark_degraded("classifier returned nothing")
    run.complete({"ticket": ticket["id"], "label": label})
    return label
 
for ticket in TICKETS:             # your loop, unchanged
    print(triage(
        ticket,
        item_id=ticket["id"],            # your identity for this item
        partition_key=ticket["tenant"],  # whose item it is
    ) or "(refused)")

Each triage(...) call is one run over one item, with its own steps, outcome and cost. item_id= and partition_key= are consumed by Papayya and are not passed to your function, so you declare identity and tenant at the call site without threading them through a signature.

@papayya.llm marks the leaf that calls your model — it records each call as an LLM step (model, tokens, timing) and inspects the response, so a refusal or empty result flips the item to degraded even though the call returned a 200 and raised nothing. papayya.mark_degraded(reason) is the same verdict, stated by you.

Run it with python agent.py. papayya deploy then hands the same file to the managed worker pool, where papayya run triage-ticket '{"id": "t-9", ...}' --item-id t-9 --partition-key acme invokes it.

When you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls — reach for the explicit handle:

from papayya.durable import papayya
 
item = papayya().item("enrich", item_id="co_42")
fetch   = item.step("fetch", fetch_fn)          # checkpointed
extract = item.llm_step("extract", extract_fn)  # checkpointed + token/outcome capture
snippet = fetch(item_domain)
fields  = extract(item_name, snippet)
item.complete(fields)

Each step's result is cached in the ledger. If your process crashes and restarts, completed steps replay from cache instead of re-executing — you don't re-pay for LLM calls that already landed. (item.step() is the current spelling; the legacy .task() is kept as an alias.)

When PAPAYYA_API_KEY is set, the ledger round-trips through the Papayya Cloud control plane. For local iteration, python agent.py runs against a durable store with the free SDK and consumes no hosted compute.

Run it locally

Run python agent.py locally with the free SDK — it consumes no hosted compute:

pip install papayya
papayya example
python agent.py

For a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint. View runs, items, and steps in the hosted dashboard at app.getpapayya.com (opens in a new tab).

Deploy to the cloud

Register the function as a deployable agent, then push it to Papayya's hosted worker pool:

# agent.py
from papayya import agent
from openai import OpenAI
 
@agent(name="research-bot")
def research_bot(prompt: str) -> str:
    client = OpenAI()
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    )
    return response.choices[0].message.content
papayya signup
papayya secrets set OPENAI_API_KEY sk-...
papayya deploy        # auto-discovers @agent functions, no flags needed
papayya run research-bot "Tell me about Anthropic"

@agent(name=...) registers the function as a deployable agent. The runtime reserves cost before each LLM call and pauses the run when a budget cap is hit, preserving in-flight items for you to resume. The hosted dashboard at app.getpapayya.com (opens in a new tab) shows the runs → items → steps hierarchy.

@papayya.durable(name=..., budget_usd=...) is an alternate spelling that also sets a per-agent budget.

What you get

FeatureDescription
Per-item visibilityEvery item carries its own item_id, outcome, trace, and cost, so you can answer "why didn't X happen?" without grep
Ran-vs-worked outcomesEach item is ok or degraded (with a reason like empty_none or refusal) — a 200 that didn't actually work is still caught
Failure clusteringWhen 47 of 1,000 items don't work, see them grouped into a few patterns — not 47 separate stack traces
Per-item cost outliersFind the one bad input that ate $4 in retries before your bill does
Crash recoveryCheckpoint after every step; resume mid-item after crashes without re-paying completed calls
Budget enforcementPause-and-notify on cap, not just alert-on-cap (cloud only)
Slice replayRe-drive just the not-ok items of a run — or one tenant's slice, or a single item — into a new run linked via replayed_from
Schedules + triggersRun on cron or fire from external events, with per-agent budgets
Bring your own frameworkOpenAI, Anthropic, Bedrock, raw HTTP — Papayya never wraps your LLM client

Get started

Jump to the Quickstart to have an agent running in under a minute. Or browse runnable starter agents at github.com/papayya-zero/examples (opens in a new tab) — lead enrichment, document extraction, eval harness.