Change detection
Nothing raised. Nothing retried. Every item came back ok. And the outputs
are half the length they were on Tuesday.
That is the failure this exists for. A step that throws gets caught by durability; a step that returns something different gets caught by nobody, because there is no error to catch. Papayya watches the shape of what your steps return and tells you when it moves — and whether one of your own deploys explains it.
What it watches
Per agent, per step label, over a rolling window. Each signal is phrased so that "more" is always the direction worth telling you about:
| Signal | What moved |
|---|---|
size_floor | Outputs are shorter than they were |
cardinality_floor | Outputs carry fewer items than they did |
field_missing:<key> | A field the outputs used to carry is now absent |
truncation | More responses are being cut off mid-answer |
degraded_rate | More items are coming back degraded |
volume | Fewer items arrived than usual |
Notice what is not in that list: there is no score, no grade, and no judgement about whether an output is right. Papayya's vocabulary is changed, missing, short, absent — every one of which is decidable from the item. Whether the output is correct is your word for your domain, and we do not borrow it.
What this does not see
field_missing:<key> measures key absence. A step that starts returning
{"summary": ""} — or null, or [] — has not lost the key, so no episode
opens for it. The object is well-formed and hollow, and that is one of the most
common ways an LLM step degrades.
Two things cover the gap, and both are yours to reach for:
- Declare a contract. A
run.step(..., expect_fields=…)contract or a check of your own grades the itemdegradedat the moment it happens, whichdegraded_ratedoes watch. - Point at one. Search by example
selects on emptiness directly — "returned no
summary" means absent or hollow — so an item you find by hand still yields the whole population like it, even where no episode ever opened.
Rates, not averages
Each signal is the fraction of items that fell below a learned floor, not a moving average. That is not a stylistic choice: if 8% of your outputs collapse, an average barely twitches — the good 92% absorb it — and by the time a median moves, roughly 40% of your workload has already shifted with it.
A rate below a floor sees the same 8% as a large, obvious shift. So the floor is learned once from a healthy week and frozen, and what is tested afterwards is how many items fall under it.
The three answers
When a signal moves, the first question is whether you moved it. Papayya knows, because the agent version is stamped on every item:
- Your deploy is in the window → the notification is a question: "this deploy moved the output distribution — expected?" Often it is exactly what you shipped, so this routes quietly, in-app by default.
- No deploy is in the window → this is the alarm. The code was held constant and the output moved anyway, so the cause is upstream: the provider, the inputs, or one of your customers. Nobody watching from outside your execution can draw that line.
- The version axis carries no information → we say so, and alarm without naming a cause.
That third answer is not a degraded case. If your agent runs from a container
with no version tag and no git metadata, every item is stamped unknown —
and a system with only two answers reads that as "nothing was deployed,
therefore it's the provider's fault," which would blame your provider for every
regression you ever ship yourself. Tagging your versions is worth the minute
precisely because it buys you the first two answers. On the managed worker pool
you get them automatically.
The three route to two alert signals — deploy-changed and output-changed —
so you can route or mute them separately in the dashboard's
alert rules, per project or per agent.
Baselines form in the open
A new agent has no history, so there is nothing to compare against. Papayya says that in as many words rather than showing you an empty panel:
baseline forming, 40/200 items
Every baseline starts in observe-only, items what it sees without notifying anyone, and is promoted only after it has been quiet for a stretch of real items. A signal that pages you on its first day gets muted on its first false alarm, and a muted signal is worth less than one you never shipped.
Two things follow that are worth knowing:
- The count is in items, never in elapsed time. A nightly job that lands 4,000 items in an hour and idles for 23 does not earn a baseline by sitting still overnight.
- A quiet week does not cost you your progress. If traffic dips below the window's floor the baseline pauses rather than resetting, and the surface tells you which side it is waiting on — history, or recent volume.
One incident is one notification
A rolling window's halves overlap almost completely from one evaluation to the next, so a single real change is present in every evaluation until it ages out. Reported naively that is dozens of identical alerts for one event, which trains you to mute the channel — the exact outcome the design exists to avoid.
So a change opens an episode: one notification when it starts, one when it clears, and nothing in between. The all-clear matters as much as the alert. A tool that tells you something broke and never tells you it stopped teaches you to distrust it just as fast as duplicates do.
Two more things stay quiet on purpose. A workload that is already paused does not keep alerting — its window is frozen with the problem still in it, and re-reporting a stopped workload forever is noise. And items produced by a re-drive are excluded from the comparison, so fixing an incident never fires the detector on your fix.
From "something changed" to "these items changed"
An alarm you cannot act on is a worse version of no alarm. Every episode that identifies items carries a link to exactly those items — the ones that crossed the floor, in the window the change happened, excluding the ones that were fine:
GET /v1/durable/cohorts?drift_episode=<episode-id>That is an ordinary cohort, so every recovery verb works on it unchanged: preview it, pull it as fixtures, fix and verify offline, then release the re-drive and diff it. In the dashboard it is one button on the change — "show the items that changed" — landing on the recovery page with the selection already made.
The selection is resolved from the episode on our side, not rebuilt from the link. That is what makes the set you preview and the set you re-drive provably the same one.
volume is the exception, and it says so. Its signal is items that did
not arrive — there is no row to hand you. Rather than returning "every item
in the window" and letting that look like an answer, the change reports plainly
that it has no cohort, and why.
Which of your customers
A change reported across a whole agent is a change you have to go looking for. The common shape is worse than that: one of your tenants regresses and the rest are fine. If that tenant is 2.5% of your volume, the agent-wide rate moves about 1.3% — real, and far too small to fire. The signal is not weak, it is diluted.
So every signal is also evaluated per tenant, against the same floor the agent-level baseline already learned. That borrowing is what makes it work on small tenants: a tenant with ten items cannot learn its own baseline, but ten items are plenty to say it sits well below one that already exists.
When it finds one, the change names the tenant, and its cohort is that tenant's items — not the agent's:
enrich / extract / tenant acme — outputs are shorter than they were (60% of recent records), with no deploy in the window.
Two things keep this from becoming noise. A tenant needs at least 10 measurable items in the window before it is tested at all, because "100% of three items" is not a finding. And because testing forty tenants at once means forty chances to be wrong, the results are corrected for multiple comparisons before any of them is reported — so "these two customers" means what it says, rather than being the two that got unlucky this tick.
What it does not do
- It never stops your work. Change detection notifies; it does not pause, fence or cancel. Budget enforcement is the surface that pauses, and it is deliberately a different one.
- It never says "wrong." See the vocabulary above.
Where to look
Agent → Change detection in the dashboard shows both halves: what has changed recently, and every baseline with how far along it is. An empty change list above a populated baseline list means something specific — we are watching, and nothing has moved — which is a different answer from a blank page, and the reason both are shown.
The same data is on GET /v1/durable/drift.