Budget
Budget enforcement is a hosted feature. When you deploy an agent and run it in the Papayya worker pool, the runtime reserves cost against your declared cap before every LLM call and holds back calls that would exceed it. The local Item path is durability-only and does not track cost.
Why no local budget? Papayya does not ship provider rate cards — LLM pricing changes per customer contract and per provider release. Asking you to compute cost from your own pricing table locally means asking you to own a problem we explicitly don't ship. The hosted runtime solves this: it observes real provider responses and applies a conservative default rate card, reconciled against actual usage after each call.
Pause, don't cut off
A budget cap does not hard-kill a run mid-flight. When the cap is reached the run pauses and notifies — in-flight state is preserved, and an operator resumes it (with a bumped budget) or accepts the partial result. The point of owning execution is that hitting a limit is a decision surface, not a dropped result.
Setting a budget
Declare the cap on the agent decorator. Every run of that agent inherits it.
from papayya import agent
from openai import OpenAI
# TODO(verify): budget_usd kwarg not seen in examples; confirm against papayya-python SDK
@agent(name="research-bot", budget_usd=1.00)
def research_bot(run, prompt: str) -> str:
client = OpenAI()
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
answer = resp.choices[0].message.content
run.complete(answer)
return answerDeploy the agent and every run is capped at the declared budget. You can also set budget_cents per submission at trigger time via the API or SDK, which overrides the decorator default.
How the runtime enforces the cap
- Before each LLM call the runtime reserves the worst-case cost against the run's
budget_cents. If the reservation would exceed the cap, the call is held back before it hits the provider and the run pauses withbudget_exceeded. - After each LLM call the runtime reads the actual
usagefield from the response, releases the reservation, and commits the real cost via the control plane. - When the cap is reached the run pauses; no further LLM calls execute. Bump the budget and resume, or accept the partial result.
Budget auto-pause and WorkloadPaused
When a run crosses its cap, Papayya pauses it — it does not reject the work. The step that just completed is checkpointed first (never discarded), the run transitions to paused with pause_reason="budget", and the SDK raises WorkloadPaused at the next step boundary so your process unwinds with a named, catchable exception:
from papayya import WorkloadPaused
try:
... # your run body
except WorkloadPaused as e:
# e.reason == "budget"; the run is already 'paused' server-side and every
# completed step is preserved. Resume from the dashboard (or the resume
# endpoint) and replay picks up exactly where it stopped.
...You rarely catch it — letting it bubble up is the intended flow: bump the budget and resume the run, and every step already saved is a cache hit. WorkloadPaused is also what the degradation fences raise (reason carries the trigger); budget is one of the fences that pause a run.
It's the budget sibling of CreditExhausted, which pauses a run when your provider (not your Papayya cap) runs out of credits. Both pause-and-preserve rather than fail.
Why integer cents
Hosted runs store budgets and aggregate cost in integer cents, not floats. Floating-point addition drifts over hundreds of steps; integer math stays exact. When you write budget_usd=1.00 in the SDK it becomes budget_cents: 100 on the wire.
Setting budgets via the API
POST /v1/agents/{agent_id}/runs
{
"input": "...",
"budget_cents": 200
}Setting budgets on a schedule
Per-run budgets live on the agent decorator, not on the schedule. The schedule only controls when runs fire; every fire inherits the agent's configured budget.
from papayya import agent, schedule
@schedule(cron="0 9 * * *") # 9am UTC daily — each run gets the $5 cap
@agent(name="daily-report", budget_usd=5.00)
def daily_report(run, payload):
...Choosing a budget
A few rules of thumb, where "steps" is the number of LLM calls one run makes:
| Shape | Steps | Suggested cap |
|---|---|---|
| Simple classification | 1–5 | $0.05 – $0.20 |
| Tool-using research | 10–30 | $0.50 – $2.00 |
| Long-context summarization | 5–15 | $0.50 – $5.00 |
| Multi-document extraction | 30–100 | $5.00 – $20.00 |
Start conservative and increase based on actual usage you observe in the dashboard. A run that hits budget_exceeded pauses long before it affects your bill — check the per-run cost histogram after a few real runs to dial it in.