Recovery: pull → verify → release
When a batch goes wrong, the work is not "re-run everything." It is four questions, in order:
- Which items went wrong? Not which run — which items.
- Can I reproduce one locally?
- Does my fix actually fix them — before I spend a cent re-driving production?
- Did the re-drive work, on which items, at what cost — and did it break anything that was fine?
Papayya answers them with three verbs over one noun.
A cohort is a predicate, not a run
A cohort is every item where a predicate held over a window:
--agent enrich --tenant acme --outcome not_ok --since 2026-08-01T00:00:00ZIt is deliberately not "the failed items of run X". A run id is one
optional term in the predicate (--run), not the addressing scheme. An
incident is rarely bounded by one invocation — it is bounded by a window and
a symptom — and a real incident's items are usually spread across many.
Every term is optional. Narrow by agent, by tenant,
by verdict, by time, or by run. Records an operator already
triaged are excluded by default, so a sweep never resurrects
work someone already dealt with; --include-triaged re-admits them.
A cohort can also come from a detected change
You do not always have to write the predicate. When
change detection reports that a step's outputs
moved, it hands you the items that moved as a ready-made cohort — the ones
that crossed the learned floor, in the window the change happened, excluding
the ones that were fine. In the dashboard it is one button on the change; over
the API it is ?drift_episode=<id> on any of the three verbs below.
That path is worth preferring when it exists. The selection is resolved from the change on our side rather than retyped, so the set you preview is provably the set you re-drive.
Or from one bad item
At 2am you usually do not have a predicate. You have one item — from a customer email, an alert, or a row you were staring at — and the question is whether it is one-off or the tip of something.
Point at it and Papayya derives the predicates that fit it, each with the size of the cohort it selects:
papayya pull --like 9c21e8b4-3f21-4b7e-9d0a-1c4e7f2b8a55Records like 9c21e8b4-… — agent summarizer, tenant acme
window 2026-06-14T18:50:59Z → 2026-08-13T18:50:59Z
126 every item whose "summarize" step was graded "conservation:field:summary"
papayya pull --probe 9645e945-3de9-4ce9-a612-4f8b7b8a865e
126 every item whose "summarize" step returned no "summary"
papayya pull --probe 3d74b531-3a1d-497d-84db-2e45d1389953
Not available for this item:
below_size_floor: no size_floor baseline exists for this step yet, so there is
nothing to compare records againstYou pick one. Papayya never picks for you — the proposals mean genuinely different things, and which one describes your incident is a judgement about your domain. See Search by example below for what each kind of proposal can and cannot say.
In the dashboard the same thing is a button on any item — Find everything like this — and on the Recovery page you can paste an item id straight in.
Or from what the world told you
The three sources above all read something Papayya recorded. The fourth reads something only a person knows: that a well-formed, correct-looking output was wrong.
Every item page has a second button beside Find everything like this — This one was wrong. It writes a signal: one piece of evidence, from the world, about one item.
papayya pull --agent summarizer --flagged
papayya pull --probe 9645e945-… --flagged # like this item, AND flaggedA signal is evidence, not a grade. It overwrites nothing: the verdicts Papayya computes stay computed, and stay recomputable. It is stored precisely because it is the one judgement here that we did not compute — a thumbs-down exists nowhere else, and unrecorded it is simply gone. Everything else on this page can be re-derived from the item forever.
It is also the only source that does not decay. A learned baseline moves, a threshold ages, a predicate stops fitting; "a customer asked for their money back on this one" is true in a year.
--flagged selects at any outcome
Items people flag are overwhelmingly the ones that completed ok — that is
the entire premise of a product about silent failure. So --flagged selects at
--outcome any unless you pass one; the default didn't work filter would
return almost none of them and read as good news.
release is the exception. It will not widen its own selection for you:
papayya release --agent summarizer --flagged # refused
papayya release --agent summarizer --flagged --outcome any # re-drives themA read may widen what it shows you. A verb that re-executes production work may not widen what it acts on without you saying so.
What it is for, and what is not built yet
Today a signal does one thing: it makes an item selectable, so the flags people leave become a cohort you pull, verify against, and re-drive — the same loop as the rest of this page.
One honest limit: good signals are stored but nothing reads them yet. They are
there because an accumulated suite needs the passes too.
From your own app, with your own id
The button flags an item you are already looking at. Your thumbs-down handler is not looking at anything of ours — it has an order id — so it gets its own door:
import papayya
def on_thumbs_down(order):
papayya.signal(order.id, agent="triage")One line, and it never blocks: the send is off-thread, one attempt, under a
second, and a control plane that is down cannot break your handler. Two things
do raise, both immediately, because both are bugs in the call rather than in
the network — no API key, and a verdict or source outside the allowed set.
order.id is the item_id you declared when you submitted the work. Papayya
resolves it to whichever item had already produced output when the world
judged — so after a re-drive, the complaint lands on the item that actually
answered, not on the first attempt. That resolution happens every time someone
asks, not once at write time: a signal that arrives before we finished recording
the item is still stored, and still counted the moment it can be.
If you submitted under a partition_key, pass it. A signal without one names
only items submitted without one — Papayya will not pick between your tenants,
and the response tells you when that is what happened.
papayya.signal(order.id, agent="triage", partition_key=order.tenant,
verdict="corrected", source="edit", reason="agent named the wrong policy")Then the flags become the same cohort as everything else on this page:
papayya pull --agent triage --flagged1. pull — the incident, as files
papayya pull --agent enrich --tenant acme --since 2026-08-01T00:00:00Z --out ./fixturesOne JSON fixture per item, holding what the item actually was: the input, the real bad output, the verdict and reason, the tenant, the timestamp, and the whole step trace — not just the item's boundary. The trace is what lets the fixture catch a regression that changes which steps run, which is the point of keeping it in your repo forever.
Where the input comes from, and why the fixture says so. Every fixture
carries an input_source:
input_source | Meaning |
|---|---|
submitted_input | What you actually submitted, verbatim — the same value the worker handed your function. The one source that is the original rather than a reconstruction of it. |
item_snapshot | The item's own recorded input. |
first_checkpoint_snapshot | A reconstruction — the bound arguments of the item's first durable step. Items submitted before Papayya persisted the request separately resolve this way, and a function that checkpointed no steps has nothing to reconstruct from. |
none | Nothing was captured. The fixture is kept for its trace and cannot be re-run. |
A fixture that quietly claimed a reconstruction was the original request would be a lie in a file you keep for years, so it never does.
If the cohort is larger than --limit, pull says so on stderr —
"cohort is 812 items, pulling 100" — rather than letting you read "wrote
100 fixtures" as the whole incident.
2. verify — prove the fix, offline
papayya verify --fixtures ./fixtures # against ./agent.py
papayya verify --fixtures ./fixtures --strict # in CIverify runs your function over each fixture's recorded input, in your
process, and re-derives the verdict with the same inspectors production
used. "Fixed" here means what "ok" means in the dashboard — not a second
opinion computed by different code.
It reports one of these per fixture:
| Verdict | Meaning |
|---|---|
FIXED | Was not ok, is ok now. |
STILL NOT OK | Was not ok, still isn't. If the output is also unchanged, verify says so — the fix never reached this item's path. |
NEWLY BROKEN | Was fine before and is not now. |
ok | Was fine, still is. |
NO VERDICT | Your function ran but produced no inspected step, so there is nothing to compare. |
Exit code is 0 only if every fixture that could be answered passed and at
least one was answered. --strict also fails on fixtures that could not be
answered at all. A "pass" over a set where nothing ran is the silent partial
success this product exists to catch, so it is never reported as one.
On API spend.
verifymakes no control-plane call, touches no store, writes nothing, and needs no API key — so Papayya spends nothing and no production item is re-driven. It does not stop your own function from calling your LLM provider: that is your code, your keys, your process. Verifying 500 fixtures against a live provider key will produce a provider bill. Fixtures are free to verify only when the failure reproduces without the provider.
verify deliberately does not gate on agent version the way
replay does. Its whole purpose is to run changed code
against an old failure; gating on the mismatch would refuse every real use.
The version shift is reported per fixture instead.
3. release — re-drive the cohort, and diff it
papayya release --agent enrich --tenant acme --since 2026-08-01T00:00:00Z --latestrelease re-drives the whole cohort and then shows what changed:
RECOVERED co_0041 [degraded -> ok] $0.0135 -> $0.0141
STILL NOT OK co_0042 [degraded -> degraded] $0.0135 -> $0.0139
NEWLY BROKEN co_0043 [ok -> degraded] $0.0121 -> $0.0128
3 item(s): 1 recovered, 1 still not ok, 1 newly broken, 0 still ok
Cost of the re-drive: $0.0408 (the originals cost $0.0391)--latest is the flag you will usually want — it re-drives the whole
cohort on the agent's current version, which is the "I shipped the fix"
case. Without it, each item replays on the version it originally ran, per
item: a cohort selected by a predicate over time legitimately spans
deploys, and resolving one version for all of it would silently re-drive half
the cohort on code it never ran.
Each source item is marked replayed and drains out of the triage feed, and
each new run links back to its source via replayed_from.
It fails whole, or not at all
A re-drive is N runs of work and N runs of plan quota. release reserves for
every item before submitting any of them; if the cohort doesn't fit, it
releases nothing and tells you how far the quota went:
Error: cohort is 500 releasable item(s); only 120 trigger reservation(s)
remain. Nothing was released.A half-released cohort in the middle of an incident is the worst available
state, so it is not a state release can leave you in.
What it skips is named, never silent
Items still running are not re-driven — that would double-drive work about
to finish — and they are not marked, so they stay in the cohort and a
later release picks them up. Items whose agent no longer exists have
nowhere to route. Both are counted in the output.
Search by example: paste one known-bad item
Writing a predicate against history is the power-user path. Pointing at a bad one is what actually happens at 2am — so this is the door built for that hour.
papayya pull --like <record-id> # what selects everything like it?
papayya pull --probe <probe-id> # pull the one you picked
papayya verify --fixtures ./fixtures # prove the fix, offline
papayya release --probe <probe-id> --latestThe whole loop, from one id to a diff, without naming a step, a window or a threshold.
What a proposal can say
Papayya only proposes predicates in the vocabulary it will commit to. There is no similarity score and no "items that look like this one" — an item is selected because something decidable is true of it:
| A proposal can say | Derived from |
|---|---|
| this step was graded with the same verdict | the reason your own check or contract recorded |
this step returned no <field> — absent or empty | the item's own output |
this step omitted <field> entirely | the fields your other items carry |
| this step came back shorter than its peers | the population's learned floor |
| this step came back with fewer items than its peers | the population's learned floor |
| this step was cut off at the token limit | the provider's stop reason |
The verdict proposal is the best one, and it only exists if you declared
something. It is a plain equality on the reason your own code recorded — a
run.step(..., expect_fields=…) contract, a
check you wrote, a conservation
violation. If your agent declares nothing, this proposal never appears and you
fall back to the shape-based ones. That is the strongest practical argument for
declaring contracts: they make your incidents addressable later, not just
visible now.
An item that fits nothing is a real answer. If Papayya says "no predicate fits this item", it means nothing about the recorded output is expressible in the vocabulary above. That is worth knowing plainly rather than dressed up as a failure to try — and it is usually the prompt to add a check.
Proposals overlap, and one pair overlaps by definition
"returned no summary" and "omitted summary entirely" are nested, not
alternatives: an item missing the key is also an item with no value there.
The first strictly contains the second. When both fit, Papayya proposes the
narrower one only if it says something the broader one does not.
This matters beyond the picker, because
change detection measures the narrow one.
field_missing:<key> is key absence, so a step that starts returning
{"summary": ""} — a well-formed object with a hollow value — opens no
episode at all. Nothing is missing; the key is right there. That gap is
exactly why this door exists: it reaches the shapes the detector is
structurally blind to.
When a proposal is refused
Two of the proposals need a population baseline — how long your outputs normally are, how many items they normally carry — and a baseline is only borrowed once it is armed. Until then Papayya refuses rather than comparing your item against a floor learned from a handful of others, which would select a cohort of about one and call it a pattern.
Refusals are printed, not swallowed, because they are not the same answer:
| Refusal | What to do |
|---|---|
| "the size_floor baseline for this step is still forming (4 reference / 2 recent items)" | Come back later. It is learning. |
| "no size_floor baseline exists for this step yet" | Not later — not until this step has been watched at all. |
| "this step's baseline carries no floor to test items against" | The baseline is live but cannot answer this question. |
The selection is frozen
Picking a proposal saves it. Every later verb takes --probe <id>, and the
predicate behind it — the step, the window, the threshold — is resolved on our
side from that id.
That is deliberate: the population floors it may have borrowed keep moving as your workload runs, and pull → fix → verify → release spans hours or days. A predicate re-derived at release time would select a different cohort than the one you previewed. Freezing it means the set you approved is the set that re-drives.
You can still narrow a frozen probe with --since / --until, which
intersect with its window and can only ever shrink it. You cannot widen it —
that would be a different question, and the answer to it is a new --like.
A preview shows more than a release will re-drive
On the by-example path only, these two numbers differ on purpose:
- The preview includes items that are themselves already re-drives. You asked what else looks like this; the honest answer is your whole history, and a recurring incident is precisely the case where the earlier re-drives matter.
- The release excludes them. Nobody should be handed a re-drive of a re-drive.
So the confirmation says "re-drive up to N", and the dashboard says the same above its count. The gap can only ever remove items from what you previewed, never add — you cannot release something you did not see.
How far back it reaches, honestly
The window defaults to 60 days. That is a default, not a retention guarantee. Today nothing purges recorded steps, so in practice the reach is all recorded history; the plan object separately declares a retention figure that nothing currently enforces. Until those are reconciled, treat 60 days as the question Papayya asks by default and not as a promise about what exists.
A derivation over that window is a heavy read: it scans every recorded step for
the agent in range, and the cost grows with the window rather than with the
number of proposals — all of them are counted in one pass. If you already know
the incident started Tuesday, passing --since makes it markedly cheaper and
the cohort more honest.
In the dashboard
Recovery in the sidebar runs the same loop: name the predicate — or paste a item you know is bad — preview what it selects, release it, and read the roll-up: how many recovered, how many are still broken, how many were newly broken, and what the re-drive cost. Each row links into the per-item step diff, source beside re-drive.
Every item page also carries Find everything like this, on every item —
including ones that completed ok. That is not an oversight: an item that
returned successfully and was wrong anyway is the whole reason this product
exists, and hiding the button behind a failure badge would hide it from exactly
the item it was built for.
Previewing a cohort shows you the items, not just a count. "Is this really the same thing I'm looking at" is the question that precedes "should I re-drive it", and a number cannot answer it.
A note on the cost columns
Every cost here is an estimate: token counts × your
project rate card. It is not a bill — billing meters
separately — and 0 means "we could not price it" (no model, or a model
absent from your rate card) as often as it means free.