Core Concepts
Overview

Core Concepts

A handful of words carry all of Papayya: an agent is loaded once by the worker pool; each run processes items; every item shows what it did and what it cost, and you can replay the ones that didn't work. A step is one node inside an item's trace; a tenant is whoever an item belongs to, declared via partition_key.

The nouns

  • An agent is the deployable unit — your function plus its name, version, schedule, triggers, and budget. Register it with @agent(name=...). Agents and batch pipelines both live here: if you don't think of your code as an "agent," it's still one to Papayya.
  • A run is one invocation of an agent. One @agent call, one cron fire, one submitted pile — each is a run that processes N items. Runs have a lifecycle: materializing → queued → running → completed | partial | failed | cancelled | budget_exceeded, and roll up worst_outcome_status + degraded_count.
  • An item is one thing a run processed. It has an execution status, an outcome (ok / degraded / failed), a step trace, and a cost — and it's the unit you replay.
  • A step is one node inside an item's trace — a model call or tool call, checkpointed so a crash resumes from it instead of re-paying for it.

Two execution paths

Both paths share the same runs → items → steps hierarchy, the same outcome inspection, and the same cost control.

Local (you run, Papayya tracks)

Run python agent.py locally with the free SDK — it consumes no hosted compute. Papayya provides step checkpointing, outcome inspection, and cost tracking. For a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint. View runs, items, and steps in the hosted dashboard at app.getpapayya.com (opens in a new tab).

Your process
  → your_agent(item, item_id=…, partition_key=…)   ← one run over one item
    → run.step("label", fn, item_id=…) checkpoints each unit of work
      → @papayya.llm leaf records an LLM step and inspects the result  ← checkpointed to the durable store
  → the dashboard shows runs → items → steps, cost, and worked/degraded verdicts

Best for: developers with existing pipelines or agents, local resources, or anyone who wants the loop before they want the cloud.

Cloud (Papayya runs your code)

You deploy agent code with papayya deploy. Papayya runs it in the hosted worker pool with full lifecycle management.

Your code → papayya deploy → worker pool loads the agent once
                                             ↓
Schedule / trigger / runs submit → one run minted per fire
                                             ↓
                             Worker executes each item, records steps + usage
                             Inspects every model return for outcome
                                             ↓
                             Runs → items → steps → durable store → dashboard

Best for: scheduled runs, trigger-invoked agents, and long-running work that shouldn't depend on your process staying alive.

Key principles

Both paths share these guarantees:

PrincipleWhat it means
Step-based executionEvery item is a sequence of discrete steps, not a black box
Checkpoint after every stepEach step's result is persisted, so a crash resumes from the last completed step
Resume, don't restartRe-running an item replays completed steps from cache instead of re-executing them
Ran vs. workedEvery item carries an outcome — a clean 200 that didn't actually work is flagged degraded, no check written
Budgets pause, they don't killA run over budget stops dispatching new items; it doesn't kill in-flight work
Full traceEvery step, tool call, token count, and cost is recorded and queryable

Read on

  • Runs — one invocation of an agent over N items: lifecycle, outcome roll-up, and results
  • Items & Steps — the per-item outcome, the step trace, and replay
  • Search — querying across step traces
  • Triggers — three ways to start a run (the API, schedules, and triggers), with the declarative YAML and reconciliation rules in Schedules & Triggers
  • Durability — step checkpointing and crash recovery
  • Budget Enforcement — how cost caps pause a run
  • Usage & Billing — how metering and period limits work
  • Recovery — selecting the items that didn't work as a cohort, then fixing and re-driving them
  • Replay — re-driving the items that didn't work
  • Change Detection — when outputs move without raising, and whether your deploy explains it
  • Environments — how envs map to projects, API keys, and credentials