Core Concepts
Long-running Agents

Long-running agents

An agent with a lot of tool calls is slow because it waits. It waits on the model, then on the tool, then on the model again, and every turn needs the one before it. The CPU is idle for almost all of it.

Papayya does not make a turn faster. It does three things around the turns:

  • A failure costs the step that failed, not every step before it.
  • A slow item is not mistaken for a dead one, and a stuck one does not run forever.
  • Many items run at once, so a thousand three-minute items do not take fifty hours.

A failure costs the step, not the item

Every run.step(...) is checkpointed when it returns. When an item fails at turn 18 and you re-drive it, the 17 turns that worked are read back from the ledger instead of being run again, and the item resumes at turn 18.

Wrap each model call and each tool call as its own step:

import papayya
from papayya import agent
 
@agent(name="research")
def research(run, task: dict) -> dict:
    messages = [{"role": "user", "content": task["question"]}]
    for _ in range(30):
        reply = run.step("model", call_model, item_id=task["id"])(messages)
        messages.append(reply["message"])
        if not reply["tool_calls"]:
            return {"answer": reply["message"]["content"]}
        for call in reply["tool_calls"]:
            result = run.step(f"tool:{call['name']}", run_tool, item_id=task["id"])(call)
            messages.append({"role": "tool", "content": result})
    papayya.mark_degraded("hit the turn limit without an answer")

Repeating a label in a loop is fine: the second "model" step is a different checkpoint from the first. On a re-drive, the reused turns hand back exactly what they returned the first time, so the conversation that reaches the failed turn is the same conversation that failed there, not a new one sampled from a non-deterministic model.

Measured

A 40-page document with a model-shaped 2.5 seconds per page, failing at page 17, from document-processing (opens in a new tab) on a local stack (three runs, within 0.1s of each other):

Wall-clockSteps runSteps reused
First pass, fails at page 1768.0s160
Re-drive60.3s2416
The same document from scratch100.6s400
  • The re-drive finished 40.2s sooner than starting over. That is the 16 pages that already worked.
  • Skipping those 16 pages took 0.02s.
  • Across the whole re-drive, 0.32s was spent on anything other than the pages themselves.

The saving is proportional to where the failure lands: a failure at page 17 of 40 saves 40%, and at page 39 it saves 95%. Money follows the same split, because a reused step makes no model call at all.

When nothing is reused

Reuse is for re-driving the same code. A replay on a new version (papayya replay --latest, papayya release --latest) reuses nothing and runs every step again. That is deliberate: reusing steps from the old code under the new would give you an answer built from both, with nothing on the record saying so. See Replay.

Retries before failure

A step that raises is retried before the item fails: 5 attempts, waiting 1, 2, 4 and 8 seconds between them. That covers a provider blip, a rate limit or a dropped socket without an operator ever seeing it. Every retry is written to the worker log, and a step that succeeds after retrying records how many it took.

It also means a failure that does not go away costs about 15 seconds plus five attempts before the item gives up. For work that cannot succeed on a second try, say so:

from papayya import NonRetriable
 
run.step("charge-card", charge, retries=0)(order)       # never retried
 
def classify(doc):
    if doc["kind"] not in KNOWN:
        raise NonRetriable(f"unknown document kind: {doc['kind']}")  # this raise only

A retry is not a re-drive. Retries happen inside one attempt at the item; once they are exhausted the item fails, and re-driving it is what reuses the steps that worked.

Slow is not dead

A worker holds a lease on the item it is running. The lease is 30 seconds long, and the SDK renews it with a heartbeat from a separate process, so it keeps renewing even while a step is busy. An item can run far longer than the lease: the example above has a 150-page document that holds one lease for 47 seconds with no duplicate execution.

If the worker dies, the heartbeat stops, the lease expires, and the item goes back on the queue for another worker to resume from its last checkpoint. See Durability.

The ceiling

Every item gets 30 minutes of wall-clock by default. Past that, the worker stops it and the item fails with timeout: agent ran for >1800s. The watchdog is a signal, so it fires even when a step is holding the GIL, which is exactly when a heartbeat thread would not.

Declare a different ceiling on the agent, and every item gets it, however it was submitted:

@agent(name="research", max_duration_seconds=3 * 3600)

Up to 24 hours. A single run started through the API (POST /v1/agents/{agentId}/runs) can set its own with timeout_seconds, which outranks the agent's. See API Reference.

An agent that genuinely needs more than 30 minutes per input is usually better split: one item per document section, per customer, or per task, rather than one item for the whole job.

Many items at once

Each worker runs one item at a time, and throughput comes from the number of workers. Items are leased with FOR UPDATE SKIP LOCKED, so workers never contend for the same one, and adding workers scales close to linearly. Six four-second items on three local workers finish in 8.1 seconds, in two waves of three.

The hosted pool autoscales. Locally, add workers with:

docker compose up -d --scale worker=4 worker

This is why a slow agent is mostly a throughput question rather than a latency one. A 20-turn agent that takes three minutes per input is fine when a thousand of them are running in parallel and none of them has to start over.

Per-tenant caps

Running everything at once is not always what you want. One tenant's thousand items can starve everyone else's, or push one customer's key past its provider limit. Cap it per partition_key:

@agent(name="research", concurrency_per_key=2, rate_limit="30/min")
  • concurrency_per_key=2: at most two of one tenant's items run at a time. The rest wait in the queue, and other tenants' items run around them.
  • rate_limit="30/min": at most 30 of one tenant's items start in any 60-second window. "N/sec" is accepted and converted to a per-minute count, so "1/sec" allows a burst of 60.

Both are checked when a worker picks up an item, and apply per tenant: two tenants each get their own two. Driven on three workers: with concurrency_per_key=1, six items of one tenant ran strictly one after another while two other tenants' items ran alongside; with rate_limit="2/min", four items started at 0s, 0s, 60s and 60s.

These live on the agent, so papayya deploy is what changes them, and it prints what it changed. Removing one from the decorator removes it from the agent. A change applies to items queued after the deploy succeeds; items already waiting keep the caps they were queued with.

A slot is held by a heartbeat, not by a counter. If the worker running a capped item dies, its slot frees within a minute of its last heartbeat, and the tenant's next item starts. Measured: 59.9 seconds from killing the worker to the next item of a tenant capped at 1 starting.

Not built yet

Said plainly, so you do not build on it:

  • Scheduling around your provider's rate limit. The caps above are numbers you choose. Papayya does not yet adjust how many items run at once from your model provider's actual token budget. Today, a 429 is retried like any other raise, and a run that saturates your key will see more of them.
  • Parallel tool calls within one turn. Running a turn's independent tool calls at the same time is up to your code or your agent framework. Each still goes through its own run.step.