Core Concepts
Recovery

Recovery: pull → verify → release

When a batch goes wrong, the work is not "re-run everything." It is four questions, in order:

  1. Which items went wrong? Not which run — which items.
  2. Can I reproduce one locally?
  3. Does my fix actually fix them — before I spend a cent re-driving production?
  4. Did the re-drive work, on which items, at what cost — and did it break anything that was fine?

Papayya answers them with three verbs over one noun.

A cohort is a predicate, not a run

A cohort is every item where a predicate held over a window:

--agent enrich --tenant acme --outcome not_ok --since 2026-08-01T00:00:00Z

It is deliberately not "the failed items of run X". A run id is one optional term in the predicate (--run), not the addressing scheme. An incident is rarely bounded by one invocation — it is bounded by a window and a symptom — and a real incident's items are usually spread across many.

Every term is optional. Narrow by agent, by tenant, by verdict, by time, or by run. Records an operator already triaged are excluded by default, so a sweep never resurrects work someone already dealt with; --include-triaged re-admits them.

A cohort can also come from a detected change

You do not always have to write the predicate. When change detection reports that a step's outputs moved, it hands you the items that moved as a ready-made cohort — the ones that crossed the learned floor, in the window the change happened, excluding the ones that were fine. In the dashboard it is one button on the change; over the API it is ?drift_episode=<id> on any of the three verbs below.

That path is worth preferring when it exists. The selection is resolved from the change on our side rather than retyped, so the set you preview is provably the set you re-drive.

Or from one bad item

At 2am you usually do not have a predicate. You have one item — from a customer email, an alert, or a row you were staring at — and the question is whether it is one-off or the tip of something.

Point at it and Papayya derives the predicates that fit it, each with the size of the cohort it selects:

papayya pull --like 9c21e8b4-3f21-4b7e-9d0a-1c4e7f2b8a55
Records like 9c21e8b4-… — agent summarizer, tenant acme
  window 2026-06-14T18:50:59Z → 2026-08-13T18:50:59Z

      126  every item whose "summarize" step was graded "conservation:field:summary"
           papayya pull --probe 9645e945-3de9-4ce9-a612-4f8b7b8a865e
      126  every item whose "summarize" step returned no "summary"
           papayya pull --probe 3d74b531-3a1d-497d-84db-2e45d1389953

Not available for this item:
  below_size_floor: no size_floor baseline exists for this step yet, so there is
    nothing to compare records against

You pick one. Papayya never picks for you — the proposals mean genuinely different things, and which one describes your incident is a judgement about your domain. See Search by example below for what each kind of proposal can and cannot say.

In the dashboard the same thing is a button on any item — Find everything like this — and on the Recovery page you can paste an item id straight in.

Or from what the world told you

The three sources above all read something Papayya recorded. The fourth reads something only a person knows: that a well-formed, correct-looking output was wrong.

Every item page has a second button beside Find everything like this — This one was wrong. It writes a signal: one piece of evidence, from the world, about one item.

papayya pull --agent summarizer --flagged
papayya pull --probe 9645e945-… --flagged   # like this item, AND flagged

A signal is evidence, not a grade. It overwrites nothing: the verdicts Papayya computes stay computed, and stay recomputable. It is stored precisely because it is the one judgement here that we did not compute — a thumbs-down exists nowhere else, and unrecorded it is simply gone. Everything else on this page can be re-derived from the item forever.

It is also the only source that does not decay. A learned baseline moves, a threshold ages, a predicate stops fitting; "a customer asked for their money back on this one" is true in a year.

--flagged selects at any outcome

Items people flag are overwhelmingly the ones that completed ok — that is the entire premise of a product about silent failure. So --flagged selects at --outcome any unless you pass one; the default didn't work filter would return almost none of them and read as good news.

release is the exception. It will not widen its own selection for you:

papayya release --agent summarizer --flagged                # refused
papayya release --agent summarizer --flagged --outcome any  # re-drives them

A read may widen what it shows you. A verb that re-executes production work may not widen what it acts on without you saying so.

What it is for, and what is not built yet

Today a signal does one thing: it makes an item selectable, so the flags people leave become a cohort you pull, verify against, and re-drive — the same loop as the rest of this page.

One honest limit: good signals are stored but nothing reads them yet. They are there because an accumulated suite needs the passes too.

From your own app, with your own id

The button flags an item you are already looking at. Your thumbs-down handler is not looking at anything of ours — it has an order id — so it gets its own door:

import papayya
 
def on_thumbs_down(order):
    papayya.signal(order.id, agent="triage")

One line, and it never blocks: the send is off-thread, one attempt, under a second, and a control plane that is down cannot break your handler. Two things do raise, both immediately, because both are bugs in the call rather than in the network — no API key, and a verdict or source outside the allowed set.

order.id is the item_id you declared when you submitted the work. Papayya resolves it to whichever item had already produced output when the world judged — so after a re-drive, the complaint lands on the item that actually answered, not on the first attempt. That resolution happens every time someone asks, not once at write time: a signal that arrives before we finished recording the item is still stored, and still counted the moment it can be.

If you submitted under a partition_key, pass it. A signal without one names only items submitted without one — Papayya will not pick between your tenants, and the response tells you when that is what happened.

papayya.signal(order.id, agent="triage", partition_key=order.tenant,
               verdict="corrected", source="edit", reason="agent named the wrong policy")

Then the flags become the same cohort as everything else on this page:

papayya pull --agent triage --flagged

1. pull — the incident, as files

papayya pull --agent enrich --tenant acme --since 2026-08-01T00:00:00Z --out ./fixtures

One JSON fixture per item, holding what the item actually was: the input, the real bad output, the verdict and reason, the tenant, the timestamp, and the whole step trace — not just the item's boundary. The trace is what lets the fixture catch a regression that changes which steps run, which is the point of keeping it in your repo forever.

Where the input comes from, and why the fixture says so. Every fixture carries an input_source:

input_sourceMeaning
submitted_inputWhat you actually submitted, verbatim — the same value the worker handed your function. The one source that is the original rather than a reconstruction of it.
item_snapshotThe item's own recorded input.
first_checkpoint_snapshotA reconstruction — the bound arguments of the item's first durable step. Items submitted before Papayya persisted the request separately resolve this way, and a function that checkpointed no steps has nothing to reconstruct from.
noneNothing was captured. The fixture is kept for its trace and cannot be re-run.

A fixture that quietly claimed a reconstruction was the original request would be a lie in a file you keep for years, so it never does.

If the cohort is larger than --limit, pull says so on stderr — "cohort is 812 items, pulling 100" — rather than letting you read "wrote 100 fixtures" as the whole incident.

2. verify — prove the fix, offline

papayya verify --fixtures ./fixtures            # against ./agent.py
papayya verify --fixtures ./fixtures --strict   # in CI

verify runs your function over each fixture's recorded input, in your process, and re-derives the verdict with the same inspectors production used. "Fixed" here means what "ok" means in the dashboard — not a second opinion computed by different code.

It reports one of these per fixture:

VerdictMeaning
FIXEDWas not ok, is ok now.
STILL NOT OKWas not ok, still isn't. If the output is also unchanged, verify says so — the fix never reached this item's path.
NEWLY BROKENWas fine before and is not now.
okWas fine, still is.
NO VERDICTYour function ran but produced no inspected step, so there is nothing to compare.

Exit code is 0 only if every fixture that could be answered passed and at least one was answered. --strict also fails on fixtures that could not be answered at all. A "pass" over a set where nothing ran is the silent partial success this product exists to catch, so it is never reported as one.

On API spend. verify makes no control-plane call, touches no store, writes nothing, and needs no API key — so Papayya spends nothing and no production item is re-driven. It does not stop your own function from calling your LLM provider: that is your code, your keys, your process. Verifying 500 fixtures against a live provider key will produce a provider bill. Fixtures are free to verify only when the failure reproduces without the provider.

verify deliberately does not gate on agent version the way replay does. Its whole purpose is to run changed code against an old failure; gating on the mismatch would refuse every real use. The version shift is reported per fixture instead.

3. release — re-drive the cohort, and diff it

papayya release --agent enrich --tenant acme --since 2026-08-01T00:00:00Z --latest

release re-drives the whole cohort and then shows what changed:

  RECOVERED     co_0041  [degraded -> ok]        $0.0135 -> $0.0141
  STILL NOT OK  co_0042  [degraded -> degraded]  $0.0135 -> $0.0139
  NEWLY BROKEN  co_0043  [ok -> degraded]        $0.0121 -> $0.0128

3 item(s): 1 recovered, 1 still not ok, 1 newly broken, 0 still ok
Cost of the re-drive: $0.0408 (the originals cost $0.0391)

--latest is the flag you will usually want — it re-drives the whole cohort on the agent's current version, which is the "I shipped the fix" case. Without it, each item replays on the version it originally ran, per item: a cohort selected by a predicate over time legitimately spans deploys, and resolving one version for all of it would silently re-drive half the cohort on code it never ran.

Each source item is marked replayed and drains out of the triage feed, and each new run links back to its source via replayed_from.

It fails whole, or not at all

A re-drive is N runs of work and N runs of plan quota. release reserves for every item before submitting any of them; if the cohort doesn't fit, it releases nothing and tells you how far the quota went:

Error: cohort is 500 releasable item(s); only 120 trigger reservation(s)
remain. Nothing was released.

A half-released cohort in the middle of an incident is the worst available state, so it is not a state release can leave you in.

What it skips is named, never silent

Items still running are not re-driven — that would double-drive work about to finish — and they are not marked, so they stay in the cohort and a later release picks them up. Items whose agent no longer exists have nowhere to route. Both are counted in the output.

Search by example: paste one known-bad item

Writing a predicate against history is the power-user path. Pointing at a bad one is what actually happens at 2am — so this is the door built for that hour.

papayya pull --like <record-id>          # what selects everything like it?
papayya pull --probe <probe-id>          # pull the one you picked
papayya verify --fixtures ./fixtures     # prove the fix, offline
papayya release --probe <probe-id> --latest

The whole loop, from one id to a diff, without naming a step, a window or a threshold.

What a proposal can say

Papayya only proposes predicates in the vocabulary it will commit to. There is no similarity score and no "items that look like this one" — an item is selected because something decidable is true of it:

A proposal can sayDerived from
this step was graded with the same verdictthe reason your own check or contract recorded
this step returned no <field> — absent or emptythe item's own output
this step omitted <field> entirelythe fields your other items carry
this step came back shorter than its peersthe population's learned floor
this step came back with fewer items than its peersthe population's learned floor
this step was cut off at the token limitthe provider's stop reason

The verdict proposal is the best one, and it only exists if you declared something. It is a plain equality on the reason your own code recorded — a run.step(..., expect_fields=…) contract, a check you wrote, a conservation violation. If your agent declares nothing, this proposal never appears and you fall back to the shape-based ones. That is the strongest practical argument for declaring contracts: they make your incidents addressable later, not just visible now.

An item that fits nothing is a real answer. If Papayya says "no predicate fits this item", it means nothing about the recorded output is expressible in the vocabulary above. That is worth knowing plainly rather than dressed up as a failure to try — and it is usually the prompt to add a check.

Proposals overlap, and one pair overlaps by definition

"returned no summary" and "omitted summary entirely" are nested, not alternatives: an item missing the key is also an item with no value there. The first strictly contains the second. When both fit, Papayya proposes the narrower one only if it says something the broader one does not.

This matters beyond the picker, because change detection measures the narrow one. field_missing:<key> is key absence, so a step that starts returning {"summary": ""} — a well-formed object with a hollow value — opens no episode at all. Nothing is missing; the key is right there. That gap is exactly why this door exists: it reaches the shapes the detector is structurally blind to.

When a proposal is refused

Two of the proposals need a population baseline — how long your outputs normally are, how many items they normally carry — and a baseline is only borrowed once it is armed. Until then Papayya refuses rather than comparing your item against a floor learned from a handful of others, which would select a cohort of about one and call it a pattern.

Refusals are printed, not swallowed, because they are not the same answer:

RefusalWhat to do
"the size_floor baseline for this step is still forming (4 reference / 2 recent items)"Come back later. It is learning.
"no size_floor baseline exists for this step yet"Not later — not until this step has been watched at all.
"this step's baseline carries no floor to test items against"The baseline is live but cannot answer this question.

The selection is frozen

Picking a proposal saves it. Every later verb takes --probe <id>, and the predicate behind it — the step, the window, the threshold — is resolved on our side from that id.

That is deliberate: the population floors it may have borrowed keep moving as your workload runs, and pull → fix → verify → release spans hours or days. A predicate re-derived at release time would select a different cohort than the one you previewed. Freezing it means the set you approved is the set that re-drives.

You can still narrow a frozen probe with --since / --until, which intersect with its window and can only ever shrink it. You cannot widen it — that would be a different question, and the answer to it is a new --like.

A preview shows more than a release will re-drive

On the by-example path only, these two numbers differ on purpose:

  • The preview includes items that are themselves already re-drives. You asked what else looks like this; the honest answer is your whole history, and a recurring incident is precisely the case where the earlier re-drives matter.
  • The release excludes them. Nobody should be handed a re-drive of a re-drive.

So the confirmation says "re-drive up to N", and the dashboard says the same above its count. The gap can only ever remove items from what you previewed, never add — you cannot release something you did not see.

How far back it reaches, honestly

The window defaults to 60 days. That is a default, not a retention guarantee. Today nothing purges recorded steps, so in practice the reach is all recorded history; the plan object separately declares a retention figure that nothing currently enforces. Until those are reconciled, treat 60 days as the question Papayya asks by default and not as a promise about what exists.

A derivation over that window is a heavy read: it scans every recorded step for the agent in range, and the cost grows with the window rather than with the number of proposals — all of them are counted in one pass. If you already know the incident started Tuesday, passing --since makes it markedly cheaper and the cohort more honest.

In the dashboard

Recovery in the sidebar runs the same loop: name the predicate — or paste a item you know is bad — preview what it selects, release it, and read the roll-up: how many recovered, how many are still broken, how many were newly broken, and what the re-drive cost. Each row links into the per-item step diff, source beside re-drive.

Every item page also carries Find everything like this, on every item — including ones that completed ok. That is not an oversight: an item that returned successfully and was wrong anyway is the whole reason this product exists, and hiding the button behind a failure badge would hide it from exactly the item it was built for.

Previewing a cohort shows you the items, not just a count. "Is this really the same thing I'm looking at" is the question that precedes "should I re-drive it", and a number cannot answer it.

A note on the cost columns

Every cost here is an estimate: token counts × your project rate card. It is not a bill — billing meters separately — and 0 means "we could not price it" (no model, or a model absent from your rate card) as often as it means free.