CLI Reference
The Papayya CLI is a Python command-line tool installed alongside the SDK.
pip install papayyaThe nouns are the same everywhere: an agent is the deployable unit; each run is one invocation of that agent and processes items; every item has an outcome, a cost, and a step trace; you replay the items that didn't work. A tenant is whoever an item belongs to (declared via partition_key). This page uses those words precisely.
Global flags
Every server-hitting command accepts a global --env <name> flag (or PAPAYYA_ENV env var) that selects which environment from ~/.papayya/config.json to use. Omitting --env falls back to envs.current_env (set via papayya envs use <name>). See envs for the full model.
papayya --env staging run my-agent "input"
papayya --env prod status <run-id>
PAPAYYA_ENV=staging papayya secrets listCommands
signup
Point you at the dashboard, where accounts are created.
papayya signupAccount creation lives in the dashboard, not the CLI. Once you have an account, papayya login connects this terminal — you never mint or paste a key by hand.
login
Connect this terminal to your account.
papayya loginPrints a short code and opens your browser:
Visit https://app.getpapayya.com/device and enter 4Y9F-KYGQ
Waiting for you to approve it…Approve the code in the browser, pick the project it should connect to, and the CLI writes an API key + project ID to ~/.papayya/config.json. The key is never displayed and never touches your clipboard — it goes straight from the control plane into the config file.
The key it creates is an ordinary project key, named after the machine that asked for it. It appears in the dashboard under Project → API keys, and revoking it there disconnects that terminal.
| Flag | Description |
|---|---|
--key <key> | Skip the browser and use a key you already hold. For CI and machines with no browser. |
--no-browser | Print the code instead of opening a browser. |
--project-id <pid> | Pin the project. Must match the one approved in the browser. |
--env <name> | Connect a specific env (global flag, goes before login). |
PAPAYYA_API_KEY skips the CLI's config entirely and still takes precedence.
Codes expire after 15 minutes. A code that expires, is rejected in the browser, or has already been collected by another terminal ends the login with a message naming what to do next.
logout
Remove the saved CLI config.
papayya logoutDeletes ~/.papayya/config.json. Useful before signup when switching accounts, or before testing a fresh install.
envs
Manage environments. An env is a named project + API key pair stored locally — papayya signup creates one called dev automatically; papayya envs create staging provisions another. See Environments for the full model.
papayya envs list # list all configured envs (current marked with *)
papayya envs use <name> # set the default env for subsequent commands
papayya envs create <name> # provision a new project + API key
papayya envs link <name> --project-id <pid> --api-key <key> # link an existing dashboard projectenvs list
Prints each configured env, its project ID, and marks the current one with *.
* dev (project: proj_abc...)
staging (project: proj_def...)If no envs are configured, prints a hint to run papayya signup or papayya envs link.
envs use <name>
Switches current_env in ~/.papayya/config.json. Future commands without an explicit --env flag use this env's credentials.
envs create <name>
Provisions a new project and API key on the control plane and persists them under envs.<name>. Requires an account-level session (JWT) — run papayya login if the command rejects your stored credentials. Sets the new env as current on success.
envs link <name>
Records an existing project + API key (typically created in the dashboard) as a local env. No server call.
| Flag | Required | Description |
|---|---|---|
--project-id <pid> | Yes | Existing project ID |
--api-key <key> | Yes | Project-scoped API key (cpk_...) |
--base-url <url> | No | Override the control plane URL for this env |
example
Scaffold a runnable demo agent in the current directory.
papayya exampleWrites agent.py — the triage-ticket demo that classifies six support tickets against canned data. It runs offline (no provider key needed) and deliberately includes two model refusals so you can see the ran-vs-worked verdict: every call returns a 200, but two items still land degraded because Papayya inspected what came back, not just whether the call returned. Run it to feel the iteration loop, then view its runs in the hosted dashboard at app.getpapayya.com (opens in a new tab).
Pass --print to write the source to stdout instead of disk — handy for piping into a different filename.
Local iteration
There is no separate local dashboard command. Run python agent.py locally with the free SDK — it consumes no hosted compute — and view its runs, items, and steps in the hosted dashboard at https://app.getpapayya.com (opens in a new tab). For a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint.
deploy
Deploy an agent to the Papayya cloud. Auto-discovers @agent functions in the file, looks up or creates each agent by slug, bundles the code, builds a container image in the control plane, uploads it, and reconciles the schedules and triggers declared via @schedule / @trigger decorators against the selected env's project.
# Zero-arg — discovers agent.py in cwd and all agent functions in it
papayya deploy
# Explicit file
papayya deploy agents.py
# Pick which env to deploy to (resolved from ~/.papayya/config.json)
papayya deploy --env prod
# Preview the trigger reconciliation plan without applying
papayya deploy --dry-run
# CI/CD — use env vars instead of interactive config
PAPAYYA_API_KEY=cpk_... PAPAYYA_PROJECT_ID=... papayya deploy| Flag | Required | Description |
|---|---|---|
[file] | No | Python file to deploy (default: agent.py in cwd) |
--dry-run | No | Print the trigger reconciliation plan and exit without applying triggers (the bundle is still built and uploaded) |
--api-key <key> | No | API key. Env: PAPAYYA_API_KEY |
--project-id <id> | No | Project ID. Env: PAPAYYA_PROJECT_ID |
--runtime <type> | No | "python" (default) or "node" |
--entrypoint <file> | No | Entrypoint file name (default: auto-detected) |
--agent-id <uuid> | No | UUID escape hatch (single-agent files only) |
Deploy harvests the @schedule / @trigger decorators off the discovered agents and reconciles them against the selected env's project after the bundle upload. The env is resolved from --env (or PAPAYYA_ENV), falling back to current_env in ~/.papayya/config.json. An agent with no automation decorators just bundles + uploads and skips the reconcile phase.
Heads-up on
--dry-run: the flag only short-circuits the trigger apply phase. The bundle is still built and the container image is still pushed. To preview reconcile changes without touching the registry, run against a throwaway env.
The CLI discovers agent names from the decorators and looks each up by slug, creating any that don't exist yet — so --agent-id is only needed for single-agent escape hatches.
run
Trigger a cloud run of a deployed agent by slug (or UUID). One run invocation is one run of the agent.
# Primary form — slug + input, agent.py auto-detected in cwd
papayya run my-agent "Research AI trends"
# Explicit file
papayya run my-agent "Research AI trends" --file agents.py
# UUID still works
papayya run 11111111-1111-4111-8111-111111111111 "Research AI trends"
# Legacy flag form (kept for scripts)
papayya run --file agent.py --input "Research AI trends" --agent-id <uuid>The slug is resolved in the selected env's project via list_agents — so papayya --env staging run ops-bot "..." hits the staging project's ops-bot, not dev's. Local execution is not handled by this command — run your agent file directly with python agent.py, since your code owns the LLM call. To run a whole pile of inputs as one run, use papayya runs submit.
| Argument / Flag | Required | Description |
|---|---|---|
AGENT (positional) | Yes* | Agent slug or UUID. * may be omitted if --agent-id is passed |
INPUT (positional) | Yes* | Input for the agent. * may be passed via --input instead |
--file <path> | No | Agent definition file (default: agent.py in cwd) |
--input <text> | No | Input (alt to positional) |
--agent-id <uuid> | No | UUID escape hatch; wins over the positional slug |
--name <name> | No | Pick one agent when the file declares multiple |
--api-key <key> | No | Local-run LLM key; ignored for cloud runs |
A slug miss prints the available slugs in the current env so you can recover without a dashboard trip.
papayya run --localis rejected with an error explaining the bring-your-own-function model — execute your agent file withpython agent.pyinstead.
status
Check the status of a cloud run.
papayya status <run-id>Prints the run ID, its rollup status (materializing, queued, running, completed, partial, failed, cancelled, or budget_exceeded), its item counts, the degraded_count and worst_outcome_status (ok | degraded | failed), and total cost in cents. For a single-input run this is one item; for a submitted pile it is the whole run rolled up.
logs
Show the step trace for a run.
papayya logs <run-id>For each step prints the step number, type, status, input/output token counts, duration, output content, and any tool calls with their arguments. For a multi-item run, drill into one item's steps with papayya items get <item-id> or tail them live with papayya items stream <item-id>.
replay
Re-drive one item that didn't work. Mints a new run linked to the original via replayed_from and re-drives the item through your current code; the original is never mutated. For a whole incident rather than one item, see release.
papayya replay <run_id> # re-drive the item into a NEW run
papayya replay <run_id> --latest # on the agent's CURRENT version
papayya replay <run_id> --tenant acme # only if the item's partition matches
papayya replay <run_id> --force # re-drive even a item that worked| Flag | Description |
|---|---|
--run <run_id> | The item to replay (or pass it positionally) |
--tenant <key> | Only replay if the item's partition_key matches |
--latest | Re-run on the agent's current version instead of the captured one |
--force | Re-drive an item that already worked (by default only a not-ok item is eligible) |
--wait / --no-wait | Poll to completion (default) or return immediately |
Only a terminal item (completed / failed / quarantined) is replayable. For the interactive needs-attention surface — inspecting failures and dismissing the ones you accept — see papayya triage.
pull
Materialize a cohort of failed items as local fixtures — the first verb of the recovery loop. One JSON file per item: the input, the real bad output, the verdict, the tenant, the timestamp, and the whole step trace.
papayya pull --agent enrich --tenant acme --since 2026-08-01T00:00:00Z
papayya pull --agent enrich --dry-run # what would this predicate select?
papayya pull --like 9c21e8b4-… # what selects everything like this item?
papayya pull --probe bf88-… # pull the proposal you picked| Flag | Description |
|---|---|
--like <record-id> | Derive the predicates that fit one known-bad item and print them with the size of each cohort. Writes nothing — pick one and pass --probe. See search by example |
--probe <probe-id> | Pull the cohort a derived predicate selects (from --like) |
--drift-episode <id> | Pull the items behind a detected change |
--agent <slug> | Narrow to one agent |
--tenant <key> | Narrow to one partition_key |
--run <run_id> | Narrow to one run's items — one predicate term, not the addressing scheme |
--outcome <not_ok|any|degraded|failed> | Verdict axis (default not_ok) |
--flagged | Only items a person said were wrong — a thumbs-down, a ticket, a refund. Implies --outcome any unless you pass one: what people flag usually completed fine. See backflow |
--since / --until | Window bounds (RFC3339) |
--include-triaged | Re-admit items someone already dispositioned |
--limit <n> | Cap the number of fixtures written |
--out <dir> | Output directory (default ./fixtures) |
--dry-run | Show the selection without writing anything |
When the cohort is larger than --limit, pull says so on stderr rather than letting "wrote 100 fixtures" read as the whole incident.
--like, --probe and --drift-episode each select on their own: they carry their whole predicate, so passing --agent, --tenant, --run or --outcome alongside one of them is refused rather than merged. --since / --until are the exception on --probe — they intersect with the frozen window and can only narrow it.
--flagged is the other exception, and it works beside both selectors: papayya pull --probe bf88-… --flagged is everything like this item that a person also complained about. It can only remove items from a selection, never add to one.
verify
Run your patched function over pulled fixtures, offline, and report whether the fix worked. No control-plane call, no store, no API key. Exits non-zero if any fixture is still not ok.
papayya verify --fixtures ./fixtures # against ./agent.py
papayya verify --fixtures ./fixtures --strict # in CI| Flag | Description |
|---|---|
--fixtures <path> | Fixture file or directory (default ./fixtures) |
--agent-module <path> | Module your @agent lives in (default ./agent.py) |
--strict | Also fail when a fixture could not be verified at all |
--json | Emit the full result as JSON (and nothing else on stdout) |
On API spend: verify stops Papayya spending and stops the re-drive touching production. It does not stop your function from calling your LLM provider — your code, your keys, your process.
release
Re-drive a whole cohort, then show the diff: how many recovered, how many are still broken, how many were newly broken, and what it cost.
papayya release --agent enrich --tenant acme --since 2026-08-01T00:00:00Z --latest| Flag | Description |
|---|---|
| (predicate flags) | Same as pull — --agent, --tenant, --run, --outcome, --since, --until, --include-triaged |
--flagged | Only items a person flagged. On a hand-written predicate this requires --outcome: unlike pull, a re-drive will not widen its own selection for you. Refused before anything is previewed or released |
--probe <probe-id> | Release the cohort a derived predicate selects. Excludes items that are themselves re-drives, so it may act on fewer items than the same probe's pull — the prompt says so before you confirm |
--drift-episode <id> | Release the items behind a detected change |
--latest | Re-drive the whole cohort on the agent's current version — the "I shipped the fix" flag |
--wait / --no-wait | Poll the re-driven items and print the diff (default), or return the manifest |
--timeout <s> | Seconds to wait for the cohort (default 600) |
-y, --yes | Skip the confirmation prompt |
--json | Emit the diff as JSON |
Quota is reserved for the whole cohort up front: if it doesn't fit, nothing is released and the error says how far the quota went. Items still running are skipped, counted, and left in the cohort so a later release picks them up.
runs
Operate on hosted runs as a resource. The top-level run command triggers a single run and tails it; runs is the resource group that lists runs and submits a whole pile of items as one run.
papayya runs list [--limit N] [--db <path>] # NDJSON list of runs
papayya runs submit --agent <id> --file items.jsonl # submit a pile as ONE runruns submit
Submits one run over many items from a JSONL file — the primary way to drive an agent across a large pile of inputs under a shared budget and concurrency cap. Each line is a JSON object of the form {"input": ..., "metadata"?: ...} — input is whatever your agent consumes, metadata is an optional tag-along object echoed back on the item. The CLI streams the file as NDJSON, so there is no item-count ceiling (only a 1 GiB byte guard). One submission mints one run; each line becomes one item with its own durable item, outcome, and cost.
| Flag | Required | Description |
|---|---|---|
--agent <id> | Yes | Agent every item runs against |
--file <path> | Yes | JSONL file, one item per line |
--budget <dollars> | No | Total run budget in whole dollars (converted to cents) |
--concurrency <n> | No | Max concurrent items the dispatcher will launch |
--name <str> | No | Human-readable label for the run |
--callback-url <url> | No | Callback fired when the run reaches a terminal status |
--idempotency-key <str> | No | Client key to dedupe duplicate submissions |
When the budget cap is hit, remaining queued items flip to paused (not failed) so you can raise the cap and resume — Papayya pauses degrading work rather than hard-cutting it.
runs list
Lists every run (NDJSON, one per line). --limit caps the page size; --db points at a local ledger. Inspect a run's items with papayya items.
items
Operate on the items a run processed. An item is one thing a run handled — it has an execution status, an outcome (ok or degraded, with a reason token like empty_none, empty_string, or refusal; or failed if it raised), a cost, and a step trace.
papayya items list # NDJSON list of hosted items (no flags)
papayya items get <item_id> # full item: outcome, cost, step trace
papayya items stream <item_id> [--from-step N] # tail the item's step events (SSE → NDJSON)items list
Lists hosted items as NDJSON, one per line — pipe it into jq or redirect to a file. It takes no flags. To scope to one run's items or filter by outcome, use the dashboard, or papayya runs list to find the run and drill into it. This replaces the old per-run results export: the results of a run are its items.
items get <id>
Prints the full item: its outcome and reason token, its cost, and its step trace (each step's type, tokens, timing, and output).
items stream <id>
Yields one JSON object per Server-Sent Event as the item's steps land: {"event": "step" | "terminal" | "error", "data": {...}, "id": <step_number>}. The stream exits when the item reaches a terminal status. Pass --from-step with the highest step number already observed to resume after a transient disconnect.
triage
The needs-attention surface — the CLI face of the dashboard's dead-letter / needs-attention view. triage lists the items that didn't work (failed or degraded) and lets you resolve each one: re-drive it, or accept it as terminal. Once every not-ok item in a run has a disposition, the run promotes from partial to completed.
papayya triage list [--tenant <key>] [--kind failed|degraded] [--limit N] # items awaiting a disposition
papayya triage retry <item_id> # re-drive this item through your current code
papayya triage dismiss <item_id> # accept the outcome as terminal; stop surfacing it-
triage listprints one JSON object per line (NDJSON), so it pipes cleanly:papayya triage list --tenant <key> | jq -r .item_id | xargs -I {} papayya triage dismiss {} -
triage retry <item_id>re-drives a single not-ok item. For bulk recovery — every item matching a predicate over a window — reach forpapayya release, which re-drives the whole cohort and diffs it. -
triage dismiss <item_id>accepts the failure or degrade as reviewed and terminal, so it stops showing up as needing attention.
Predicting the cost of a bulk replay before you run it lives in the dashboard's confirmation gate, not as a CLI command — see Preview replay cost in the API reference.
agents
CRUD over hosted agents. The papayya deploy command creates the agent item implicitly the first time it sees a new slug; this group is for inspecting and updating agents directly.
papayya agents create --name X --slug x --project-id <id> [--description ...] [--config '{...}']
papayya agents list [--project-id <id>]
papayya agents get <agent-id>
papayya agents update <agent-id> [--name ...] [--description ...] [--config '{...}']--config accepts a JSON object string. update sends only the fields you pass — omitted flags stay untouched, and --config merges into the stored object rather than replacing it, so setting one key can't silently drop the others. Pass null as a value to clear a key. At least one update field is required.
config is where an agent's fences live:
| Key | Default | Effect |
|---|---|---|
pause_after_degraded | 3 | Pause a run after N consecutive degraded steps. 0 disables it. See Durability. |
budget_usd | — | Per-run budget cap. Pauses the run when consumed spend crosses it. |
workload_pause_pct | 50 | Pause the whole agent when this % of its last workload_pause_window runs were degraded. |
workload_pause_window | 20 | How many recent runs the workload fence looks at. |
workload_pause_min_degraded | 5 | Floor before the percentage means anything. |
papayya agents update <agent-id> --config '{"pause_after_degraded": 5}'schedules
Manage cron schedules over hosted agents. Each fire of a schedule is one run.
papayya schedules create --agent <id> --cron "0 */6 * * *" \
[--timezone America/Toronto] [--input "..."] [--budget $D]
papayya schedules list [--agent <id>]
papayya schedules get <schedule-id>
papayya schedules update <schedule-id> [--cron ...] [--timezone ...] [--input ...] \
[--budget $D]
papayya schedules delete <schedule-id>
papayya schedules enable <schedule-id>
papayya schedules disable <schedule-id>--budget is whole dollars (converted to cents on the wire to match the rest of the CLI). --timezone is an IANA zone name and is honoured as wall-clock time in that zone. update sends only the fields you pass; at least one is required. enable/disable toggle the schedule's enabled flag without deleting it — re-enabling recomputes the next occurrence, and fails loudly if the stored cron will not parse.
There is no --max-steps. The step ceiling is read from the deployed bundle, so it lives on @agent(max_steps=N); the flag used to exist and store a value the runtime never applied.
To see whether a schedule's occurrences actually produced runs, use GET /v1/schedules/{id}/events or the schedules page in the dashboard — see Schedules.
triggers
Manage inbound triggers per agent. A trigger is an invocation hook: an inbound HTTP call fires one run of the agent, with the request body becoming the run's input.
papayya triggers create --agent <id> --name "GitHub Push" [--description "..."]
papayya triggers list <agent-id>
papayya triggers delete <trigger-id>triggers create returns the trigger including its signing secret — capture it on creation; it's not re-shown on list. The signing secret authenticates the inbound webhook that fires the trigger (the HTTP transport detail); the product noun is the trigger.
projects (hosted)
Hosted project CRUD. Plural is deliberate — the singular project group handles local run history (export/import).
papayya projects list # NDJSON
papayya projects get <project-id>
papayya projects update <project-id> [--name ...] [--slug ...]
papayya projects delete <project-id> # prompts for confirmationTo create a hosted project from the CLI, use papayya envs create <name> — it provisions the project and an API key in one step and writes them into the local env config.
deployments
Inspect hosted deployments. Creation stays on papayya deploy (which bundles + uploads in one step); this group is for read-only inspection of deployment history.
papayya deployments list <agent-id> # NDJSON, newest first
papayya deployments get <deployment-id>api-keys
Inspect and revoke project API keys. Creation stays on papayya envs create — it provisions a project and a key in one step and writes the key into the local env config.
papayya api-keys list --project-id <id> # NDJSON, prefixes only
papayya api-keys revoke <key-id> --project-id <id> # prompts for confirmationlist returns key metadata (prefix, name, created_at) — never the secret. Secrets are returned exactly once on creation; item them then or rotate. revoke is immediate and irreversible.
secrets
Manage project secrets. Secrets are injected as environment variables into the cloud agent's container at runtime — this is how you provide your LLM API keys to deployed agents.
# Set a secret in the current env's project
papayya secrets set ANTHROPIC_API_KEY sk-ant-...
# Target a different env
papayya --env staging secrets set ANTHROPIC_API_KEY sk-ant-...
# Override the project explicitly (still wins over env config)
papayya secrets set OPENAI_API_KEY sk-proj-... --project-id proj_abc123
# List secrets (names only — values never returned)
papayya secrets list
# Delete a secret
papayya secrets delete ANTHROPIC_API_KEYThe argument order is NAME VALUE, with --project-id as an optional trailing flag. When omitted, the project is resolved from the selected env (envs.<current_env>.project_id in ~/.papayya/config.json) or from PAPAYYA_PROJECT_ID.
| Subcommand | Args |
|---|---|
set | NAME VALUE [--project-id <id>] |
list | [--project-id <id>] |
delete | NAME [--project-id <id>] |
usage
Usage rollups over an optional date range. Useful for billing reconciliation, alerts, or piping into scripts.
papayya usage summary [--from <date>] [--to <date>] # aggregate (pretty JSON)
papayya usage breakdown [--from <date>] [--to <date>] # per-dimension (NDJSON)--from / --to accept either ISO dates (YYYY-MM-DD) or RFC3339 timestamps. Both are optional; omitting them returns the platform's default window.
Configuration
After signup or login, the CLI persists config to ~/.papayya/config.json. Credentials are scoped per env; signup provisions a dev env automatically.
{
"version": 2,
"current_env": "dev",
"auth": {
"jwt": "eyJhbGc...",
"email": "you@example.com"
},
"envs": {
"dev": {
"api_key": "cpk_...",
"base_url": "https://api.getpapayya.com",
"project_id": "proj_...",
"email": "you@example.com"
},
"staging": {
"api_key": "cpk_...",
"base_url": "https://api.getpapayya.com",
"project_id": "proj_..."
}
}
}Each command resolves its credentials by:
- Explicit flag (
--api-key,--project-id,--base-url) wins outright. - Otherwise, fall back to env-var (
PAPAYYA_API_KEY,PAPAYYA_PROJECT_ID,PAPAYYA_BASE_URL). - Otherwise, read from
envs.<selected env>— where the selected env is--env <name>, thenPAPAYYA_ENV, thencurrent_env.
auth.jwt is the account-level session used by papayya envs create to provision new projects.
Environment variables
| Variable | Description |
|---|---|
PAPAYYA_API_KEY | Papayya platform API key for cloud operations |
PAPAYYA_BASE_URL | Control plane URL (default: https://api.getpapayya.com) |
PAPAYYA_PROJECT_ID | Default project ID for secrets, deploy, etc. |
PAPAYYA_ENV | Default env name (overrides current_env; equivalent to passing --env <name> on every command) |
Papayya does not read provider keys (
ANTHROPIC_API_KEY,OPENAI_API_KEY, …) directly. Those are read by your LLM SDK, which you import and call inside your agent code. Store them as project secrets viapapayya secrets setso they're injected into your deployed container.