SDK Reference
Overview

SDK Reference

The Papayya SDK is a Python library installed alongside the CLI.

pip install papayya

Four words carry the whole SDK: an agent is loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. (A step is one node inside an item's trace; a tenant is whoever an item belongs to, declared via partition_key.)

Papayya does not ship LLM provider adapters. You bring your own LLM SDK (anthropic, openai, bedrock, ollama, ...) and call it directly — Papayya never wraps your provider client. You mark the leaf that calls your model with @papayya.llm, and Papayya records the call as a step and inspects what came back. This keeps provider SDK churn out of your critical path: we never break a run because a provider shipped a breaking SDK change.

The adoption ladder

There are three ways to reach for Papayya, in increasing order of commitment. They share one vocabulary and one ledger — start at the top and climb only when you need what the next rung buys you. Nothing injects a run argument into your functions; your business code stays Papayya-unaware.

1. Ambient — wrap the loop you already have

The lightest rung. Decorate the function that calls your model with @papayya.llm, and decorate the function that processes one item with @agent. Each call is one run over one item, with its own outcome, trace, and cost. Your loop stays yours.

import papayya
from papayya import agent
 
@papayya.llm                       # mark the function that calls your model
def classify(text: str) -> dict:
    ...                            # your provider, your key
 
@agent(name="triage-ticket", model="claude-sonnet-5")
def triage(run, ticket: dict) -> str:
    result = run.step("classify", classify, item_id=ticket["id"])(ticket["text"])
    return result["content"]
 
for ticket in TICKETS:             # your loop, unchanged
    print(triage(
        ticket,
        item_id=ticket["id"],            # your identity for this item
        partition_key=ticket["tenant"],  # whose item it is (the tenant)
    ) or "(refused)")
  • @papayya.llm records each call as an LLM step (model, tokens, timing) and runs outcome inspectors on the return value — a refusal or empty result flips the item to degraded with no check written anywhere. Called outside an active run it runs bare, unrecorded: adoption is rewarded, never required.
  • item_id= and partition_key= are declared at the CALL SITE and consumed by Papayya — they are not forwarded to your function, so triage needs no parameter for either. Declaring partition_key in your signature still works if you want to read it.
  • run as the first parameter is what records the run. Remove it and your code still works — nothing reaches the dashboard.
  • papayya.mark_degraded("reason") — when something the inspectors can't see makes an item not actually work, say so explicitly from anywhere inside the body. (papayya.mark_outcome("ok" | "degraded" | "failed", reason=...) is the general form.) The mark is recorded against the step that raised it.

2. Explicit item handle — checkpoint steps inside one item

When you want checkpointed steps inside one item — crash mid-item and resume without re-paying for completed calls — reach for the explicit handle. papayya() returns a client; .item(agent, item_id=...) opens one durable item you wrap your steps with.

import papayya
 
item = papayya().item("enrich", item_id="co_42")   # one durable item
fetch   = item.step("fetch", fetch_fn)             # checkpointed step
extract = item.llm_step("extract", extract_fn)     # checkpointed + token/outcome capture
snippet = fetch(domain)
fields  = extract(name, snippet)
item.complete(fields)
  • item.step(label, fn) returns a wrapped function. On first call it executes and checkpoints; on replay (same item id) it returns the cached result instead of re-executing. item.llm_step(label, fn) is the same machinery plus token / model / stop-reason extraction and credit-error classification. item.task(...) is a retained alias of item.step(...).
  • item.complete(output) marks the item finished successfully; item.fail(error) marks it failed.
  • Every step's result is cached in the ledger; re-running with the same id replays completed steps from cache rather than re-executing them. An LLM call is the most expensive side effect your code makes — durable memoization is the difference between a cheap resume and a costly one.

3. Deployable agent — hand it to the hosted worker pool

When you want Papayya to own execution — cron schedules, inbound triggers, budget caps, the hosted dashboard — register the function as an agent with @agent(name=...). Durability is ambient inside the body — an @papayya.llm call or mark_degraded resolves against the item the decorator opened.

# agent.py
import papayya
from papayya import agent
from openai import OpenAI
 
# TODO(verify): budget_usd kwarg + @papayya.durable alias not seen in examples; confirm against papayya-python SDK
@agent(name="research-bot", budget_usd=1.00)
def research_bot(prompt: str) -> str:
    client = OpenAI()
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    )
    return resp.choices[0].message.content
papayya deploy
papayya run research-bot "Tell me about Anthropic"

Attach cron fires with @papayya.schedule(...) and inbound triggers with @papayya.trigger(...), stacked above the decorator. See Agent for the full decorator surface.


Where the data lands

Everything lands in a checkpoint ledger. Run python agent.py locally with the free SDK — it consumes no hosted compute; pass an explicit store=SQLiteStore(...) to keep checkpoints on disk for an offline run. For a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint. When PAPAYYA_API_KEY is set, checkpoints persist to the cloud. View runs, items, and steps in the hosted dashboard at app.getpapayya.com (opens in a new tab). See Engine for the store internals.

Budget and cost caps are a hosted feature. The local path gives you durability and observability — outcomes, traces, and step timings — with no cost accounting. Deploy the agent to get per-run budget caps enforced by the worker pool. See Budget.

Replay the ones that didn't work

papayya replay --run <run_id>                # re-drive the run's not-ok items into a NEW run
papayya replay --run <run_id> --tenant acme  # just one tenant's slice
papayya replay --item <item_id>              # a single item

Slice replay selects the items whose outcome wasn't ok, mints a new run linked to the old one via replayed_from, and re-drives each through your current code. Items that already worked are never touched.


Papayya client (hosted operations)

The Papayya class exposes resource namespaces for managing runs, items, agents, secrets, schedules, and triggers against the hosted API. Construct once with an api_key (or let it resolve from env / CLI config) and use the durable surface and the resource surface from the same instance.

from papayya import Papayya
 
client = Papayya(api_key="cpk_...")
 
run = client.runs.create(agent_id="ag_...", items=[         # submit an invocation
    {"input": {"text": "..."}, "item_id": "co_42", "partition_key": "acme"},
])
client.schedules.create(agent_id="ag_...", cron="0 9 * * *")
usage = client.usage.summary(from_date="2026-03-01", to_date="2026-03-31")
 
# Live-tail an item's steps as they execute.
for event in client.items.stream(item_id):
    if event["event"] == "step":
        print(f"step {event['id']}")
    elif event["event"] == "terminal":
        print(f"item ended: {event['data']['status']}")
        break

Resources

papayya.runs addresses invocations (one map()/iter() call or one cron fire, over N items); papayya.items addresses the per-item items inside them.

ResourceMethods
client.runscreate, create_stream (submit invocations)
client.itemscreate, get, list, steps, stream, quarantine, release, discard
client.triagelist, iter (the DLQ / needs-attention surface)
client.agentscreate, list, get, update
client.schedulescreate, list, get, update, delete, enable, disable
client.webhookscreate, list, delete (the inbound-trigger transport)
client.deploymentscreate, get, list
client.secretsset, list, delete
client.projectscreate, list, get, update, delete
client.api_keyscreate, list, revoke
client.usagesummary, breakdown

Noun shift in 0.3.0: client.runs used to be the per-item resource. It now addresses invocations (the old batches surface); per-item access moved to client.items. client.batches forwards to client.runs as a deprecated alias — don't reach for it in new code.


Detailed references

ModuleDescription
Agent@agent (and the @papayya.durable alias), @papayya.llm, @schedule, @trigger, @tool
BudgetHosted budget enforcement and cost tracking
EngineThe local engine and the offline checkpoint store
BYOF ObservabilityCapturing steps from unpatched providers (Gemini, Bedrock, ...)
TypesPapayya, Client, RunResult, Item, PapayyaRun, and the durable dataclasses