Quickstart
Write an ordinary Python function, deploy it, and every call it handles becomes an item you can inspect, price and re-drive.
A handful of words carry all of Papayya, and this page uses them precisely: an agent is your function, loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. (A step is one node inside an item's trace; a tenant is whoever the item belongs to, declared via partition_key.)
Connect
Accounts live in the dashboard. papayya signup points you at it; papayya login connects this terminal.
pip install papayya
papayya signup # opens the dashboard — create an account and a project
papayya login # approve a code in your browserlogin prints a short code, opens your browser, and waits:
Visit https://app.getpapayya.com/device and enter 4Y9F-KYGQ
Waiting for you to approve it…Approve it, pick a project, and the CLI is connected. There is no key to copy — it goes straight into ~/.papayya/config.json. On CI, or a machine with no browser, PAPAYYA_API_KEY=cpk_... still works and skips this entirely.
Write an agent
One decorator. The function underneath is ordinary Python and stays that way.
# agent.py
from papayya import agent
@agent(name="hello", model="claude-sonnet-5")
def hello(run, name, partition_key=None):
result = {"greeting": f"Hello, {name}!"}
run.complete(result)
return result
if __name__ == "__main__":
print(hello("world", partition_key="acme"))python agent.py
# {'greeting': 'Hello, world!'}run is handed to you by the decorator — it is the item's durable item. partition_key declares whose item this is, and it is the axis every later question is asked along: per-tenant cost, per-tenant failure rate, per-tenant replay.
Deploy and run it
papayya deploy
papayya run hello "world"deploy bundles the directory, uploads it, and prints the version it registered. run submits one item and waits:
Run triggered: cdeba672-104a-44b6-8b9b-6e9a7a082447
Status: queued
Waiting for completion...
0 step(s) — completed
Final status: completedZero steps is correct here: a step is a run.step(...) call, and this agent makes none. Add them when you want a crash mid-item to resume without re-paying for the calls that already succeeded.
Keep that run id, or don't — papayya runs list answers "what have I run?" without one.
See what it did
papayya runs list
papayya status <run-id>
papayya logs <run-id>RUN AGENT STATUS ITEMS COST STARTED
cdeba672-104a-44b6-8b9b-6e9a7a082447 hello completed 1 $0.00 2m agostatus is the verdict: the run's state, why it failed if it did, its worst step outcome, and its cost. On a run that has not finished it also says how long it has been waiting — queued for four seconds is a healthy submission, queued for three days is a dead worker pool, and the word is the same in both.
logs is the trace, one line per step with tokens and cost, and the run's verdict underneath — so a failure is never invisible in the command named logs. On an agent with no steps it says so plainly rather than printing an empty list you cannot tell from a typo.
A cost of — means unpriced, not free: no rate card entry for that model, so we will not pretend to a number.
Run a pile
One JSONL line per item. item_id is your own name for the item — an order id, a ticket number — and is what you will look it up by later.
cat > items.jsonl <<'JSONL'
{"input": "Ada", "item_id": "greet-1"}
{"input": "Grace","item_id": "greet-2"}
JSONL
papayya runs submit --agent <agent-id> --file items.jsonlThroughput is a property of the worker pool, not of the submission: the hosted pool autoscales, and locally you add workers with docker compose up -d --scale worker=4 worker.
Capture the model call
Put @papayya.llm on the leaf — the function that actually calls your provider. Papayya never wraps your client; you keep your own SDK, your own key.
import papayya
from papayya import agent
from openai import OpenAI
client = OpenAI()
@papayya.llm
def classify(text: str) -> dict:
return client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": text}],
)
@agent(name="triage", model="gpt-4o-mini")
def triage(run, ticket, partition_key=None):
resp = classify(ticket)
label = resp.choices[0].message.content
run.complete({"label": label})
return {"label": label}Return the provider's response object as-is — that is the shape stored, as its own canonical form. The step items model, tokens, stop reason and cost, and runs outcome inspectors on the response: a refusal or an empty result flips the item to degraded with no check written anywhere.
Your provider key reaches the pool through papayya secrets:
papayya secrets set OPENAI_API_KEY sk-...The ones that returned 200 and didn't work
This is the part Papayya is for. A run can be completed and still not have worked, and both commands above say so rather than making you notice:
$ papayya status <run-id>
Status: completed
Outcome: degraded (2 degraded step(s))Point at one bad item and ask what selects everything like it:
papayya pull --like <record-id>Records like 9c21e8b4-… — agent triage, all tenants
3 every item of "triage" that failed with "KeyError: 'body'"
papayya pull --probe da196683-…
41 every item of "triage" whose "classify" step was graded "empty_content"
papayya pull --probe bf88c012-…Each line is a predicate and its blast radius. Pick one, pull it to disk as fixtures, fix your code, and prove the fix offline before spending a cent of inference:
papayya pull --probe bf88c012-…
papayya verify --fixtures ./fixtures --agent-module agentThen re-drive only that cohort. A run that ENDED badly is replayed; a run something DECIDED to stop — a budget or degraded-streak fence — is resumed:
papayya replay <run-id> # mints a new run linked via replayed_from
papayya replay <run-id> --tenant acme # just one tenant's slice
papayya resume <run-id> # clear a fence and re-drive what it stoppedA re-drive does not start the item over. The steps that already worked are read back from the ledger, not run again, and replay says how many:
$ papayya replay <run-id>
Reusing: 17 completed step(s) — not re-executed
43 step(s) — completedThat is a 40-page document that failed at page 17: 24 pages run, not 40. Long-running agents has the wall-clock numbers.
Add a schedule
Stack a @schedule decorator above the agent. Schedules and triggers reconcile on the next papayya deploy.
from papayya import agent, schedule, trigger
@trigger(name="hello-hook", secret_env="HELLO_HOOK_SECRET") # an HTTP door
@schedule(cron="0 9 * * 1-5") # weekdays 9am UTC — each fire is one run
@agent(name="hello", model="claude-sonnet-5")
def hello(run, name, partition_key=None):
...deploy previews the change before applying it, and prints a webhook's signing secret exactly once:
managed_by='code' diff (PUT-replace preview):
agent: hello
managed_by='code' schedules: 1 to create, 0 to update, 0 to delete (0 unmanaged rows untouched)
+ schedule 0 9 * * 1-5 create
managed_by='code' webhooks: 1 to create, 0 to update, 0 to delete (0 unmanaged rows untouched)
+ webhook hello-hook create
Webhook 'hello-hook' created:
URL: https://api.getpapayya.com/v1/webhooks/6539e5fc-…/trigger
secret: whk_f25d0513_… (store in $HELLO_HOOK_SECRET — only shown once)Schedules and webhooks you create in the dashboard are managed_by='api' and a deploy never touches them.
Environments
Every command takes --env, and an env is a project plus its key:
papayya login --env prod
papayya --env prod deploy
papayya --env prod logs <run-id>An --env that names nothing is refused rather than silently falling back — a typo aimed at prod must not quietly hit dev.
What's next
- Core Concepts — runs, items, steps, durability, budgets
- SDK Reference — full API surface and patterns
- Environments — multi-env setup, dev → prod promotion