Quickstart

Quickstart

Write an ordinary Python function, deploy it, and every call it handles becomes an item you can inspect, price and re-drive.

A handful of words carry all of Papayya, and this page uses them precisely: an agent is your function, loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. (A step is one node inside an item's trace; a tenant is whoever the item belongs to, declared via partition_key.)


Connect

Accounts live in the dashboard. papayya signup points you at it; papayya login connects this terminal.

pip install papayya
papayya signup        # opens the dashboard — create an account and a project
papayya login         # approve a code in your browser

login prints a short code, opens your browser, and waits:

  Visit https://app.getpapayya.com/device and enter 4Y9F-KYGQ

  Waiting for you to approve it…

Approve it, pick a project, and the CLI is connected. There is no key to copy — it goes straight into ~/.papayya/config.json. On CI, or a machine with no browser, PAPAYYA_API_KEY=cpk_... still works and skips this entirely.


Write an agent

One decorator. The function underneath is ordinary Python and stays that way.

# agent.py
from papayya import agent
 
@agent(name="hello", model="claude-sonnet-5")
def hello(run, name, partition_key=None):
    result = {"greeting": f"Hello, {name}!"}
    run.complete(result)
    return result
 
if __name__ == "__main__":
    print(hello("world", partition_key="acme"))
python agent.py
# {'greeting': 'Hello, world!'}

run is handed to you by the decorator — it is the item's durable item. partition_key declares whose item this is, and it is the axis every later question is asked along: per-tenant cost, per-tenant failure rate, per-tenant replay.


Deploy and run it

papayya deploy
papayya run hello "world"

deploy bundles the directory, uploads it, and prints the version it registered. run submits one item and waits:

Run triggered: cdeba672-104a-44b6-8b9b-6e9a7a082447
  Status: queued
Waiting for completion...
  0 step(s) — completed

Final status: completed

Zero steps is correct here: a step is a run.step(...) call, and this agent makes none. Add them when you want a crash mid-item to resume without re-paying for the calls that already succeeded.

Keep that run id, or don't — papayya runs list answers "what have I run?" without one.


See what it did

papayya runs list
papayya status <run-id>
papayya logs <run-id>
RUN                                    AGENT     STATUS       ITEMS      COST  STARTED
cdeba672-104a-44b6-8b9b-6e9a7a082447   hello     completed        1     $0.00  2m ago

status is the verdict: the run's state, why it failed if it did, its worst step outcome, and its cost. On a run that has not finished it also says how long it has been waiting — queued for four seconds is a healthy submission, queued for three days is a dead worker pool, and the word is the same in both.

logs is the trace, one line per step with tokens and cost, and the run's verdict underneath — so a failure is never invisible in the command named logs. On an agent with no steps it says so plainly rather than printing an empty list you cannot tell from a typo.

A cost of — means unpriced, not free: no rate card entry for that model, so we will not pretend to a number.


Run a pile

One JSONL line per item. item_id is your own name for the item — an order id, a ticket number — and is what you will look it up by later.

cat > items.jsonl <<'JSONL'
{"input": "Ada",  "item_id": "greet-1"}
{"input": "Grace","item_id": "greet-2"}
JSONL
 
papayya runs submit --agent <agent-id> --file items.jsonl

Throughput is a property of the worker pool, not of the submission: the hosted pool autoscales, and locally you add workers with docker compose up -d --scale worker=4 worker.


Capture the model call

Put @papayya.llm on the leaf — the function that actually calls your provider. Papayya never wraps your client; you keep your own SDK, your own key.

import papayya
from papayya import agent
from openai import OpenAI
 
client = OpenAI()
 
@papayya.llm
def classify(text: str) -> dict:
    return client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": text}],
    )
 
@agent(name="triage", model="gpt-4o-mini")
def triage(run, ticket, partition_key=None):
    resp = classify(ticket)
    label = resp.choices[0].message.content
    run.complete({"label": label})
    return {"label": label}

Return the provider's response object as-is — that is the shape stored, as its own canonical form. The step items model, tokens, stop reason and cost, and runs outcome inspectors on the response: a refusal or an empty result flips the item to degraded with no check written anywhere.

Your provider key reaches the pool through papayya secrets:

papayya secrets set OPENAI_API_KEY sk-...

The ones that returned 200 and didn't work

This is the part Papayya is for. A run can be completed and still not have worked, and both commands above say so rather than making you notice:

$ papayya status <run-id>
Status:  completed
Outcome: degraded (2 degraded step(s))

Point at one bad item and ask what selects everything like it:

papayya pull --like <record-id>
Records like 9c21e8b4-… — agent triage, all tenants

      3  every item of "triage" that failed with "KeyError: 'body'"
         papayya pull --probe da196683-…
     41  every item of "triage" whose "classify" step was graded "empty_content"
         papayya pull --probe bf88c012-…

Each line is a predicate and its blast radius. Pick one, pull it to disk as fixtures, fix your code, and prove the fix offline before spending a cent of inference:

papayya pull --probe bf88c012-…
papayya verify --fixtures ./fixtures --agent-module agent

Then re-drive only that cohort. A run that ENDED badly is replayed; a run something DECIDED to stop — a budget or degraded-streak fence — is resumed:

papayya replay <run-id>              # mints a new run linked via replayed_from
papayya replay <run-id> --tenant acme  # just one tenant's slice
papayya resume <run-id>             # clear a fence and re-drive what it stopped

A re-drive does not start the item over. The steps that already worked are read back from the ledger, not run again, and replay says how many:

$ papayya replay <run-id>
  Reusing: 17 completed step(s) — not re-executed
  43 step(s) — completed

That is a 40-page document that failed at page 17: 24 pages run, not 40. Long-running agents has the wall-clock numbers.


Add a schedule

Stack a @schedule decorator above the agent. Schedules and triggers reconcile on the next papayya deploy.

from papayya import agent, schedule, trigger
 
@trigger(name="hello-hook", secret_env="HELLO_HOOK_SECRET")  # an HTTP door
@schedule(cron="0 9 * * 1-5")   # weekdays 9am UTC — each fire is one run
@agent(name="hello", model="claude-sonnet-5")
def hello(run, name, partition_key=None):
    ...

deploy previews the change before applying it, and prints a webhook's signing secret exactly once:

managed_by='code' diff (PUT-replace preview):

agent: hello
  managed_by='code' schedules: 1 to create, 0 to update, 0 to delete (0 unmanaged rows untouched)
    + schedule 0 9 * * 1-5            create
  managed_by='code' webhooks: 1 to create, 0 to update, 0 to delete (0 unmanaged rows untouched)
    + webhook hello-hook             create

Webhook 'hello-hook' created:
  URL:    https://api.getpapayya.com/v1/webhooks/6539e5fc-…/trigger
  secret: whk_f25d0513_…  (store in $HELLO_HOOK_SECRET — only shown once)

Schedules and webhooks you create in the dashboard are managed_by='api' and a deploy never touches them.


Environments

Every command takes --env, and an env is a project plus its key:

papayya login --env prod
papayya --env prod deploy
papayya --env prod logs <run-id>

An --env that names nothing is refused rather than silently falling back — a typo aimed at prod must not quietly hit dev.


What's next