Dashboard

Dashboard

The Papayya dashboard gives you full observability into your agents, runs, items, and usage — organized around one hierarchy: runs → items → steps. A run is one invocation of an agent; it processes many items; each item has a trace of steps.

The dashboard is hosted at app.getpapayya.com (opens in a new tab). Runs from python agent.py (free SDK, no hosted compute) and from deployed agents render the same runs → items → steps hierarchy — what you learn on one you read on the other. For a prod-like local environment, use docker-compose, which mirrors production and differs only by endpoint.

Runs

URL: /runs

The runs page shows all agent invocations across your account.

Run list

A table of all runs. Each row leads with the interesting thing — not just that the run finished, but how much of it actually worked:

triage-ticket — completed · 2 of 6 degraded · 2 tenants · worst: degraded

  • Agent — which agent was invoked
  • Status — the run-level lifecycle: materializing, queued, running, completed, partial, failed, cancelled, budget_exceeded
  • Outcome rollup — N of M degraded plus the number of tenants touched, and the worst outcome across the run's items (ok / degraded / failed)
  • Items — how many items the run processed
  • Cost — total cost in USD, rolled up from the run's items
  • Created — when the run was triggered

Click any run to drill into its items.

Run detail

URL: /runs/{runId}

The run detail page shows:

  • Header — status, agent name, timing, total tokens, total cost, and the outcome rollup (degraded_count, worst_outcome_status)
  • Items — the list of items this run processed, each with an outcome badge (ok or degraded, with the reason token — e.g. empty_none, empty_string, refusal — on degraded ones; a failed item is one that raised), its tenant, and its per-item cost
  • Budget progress — visual progress bar showing spend vs. budget
  • Actions — Cancel (for a running run) or Replay (re-drive the not-ok items, or one tenant's slice)

Item detail

URL: /runs/{runId}/items/{itemId}

Drill into a single item to see its step trace — a step-by-step timeline of that item's execution. Each step shows:

  • Step type (tool call, LLM step, response, error)
  • For tool calls: tool name, input parameters, output
  • For LLM steps: model, token counts, and — if a rate card is configured — a per-step ≈ $X cost estimate
  • Duration

Per-item and per-step cost both come from the same rate-card token math applied at that level. The trace view is virtualized, so it handles items with 100+ steps smoothly.

Find everything like this

Every item carries one button: Find everything like this. It answers the question you have while looking at a bad item — is this one-off, or is there more of it? — with the predicates that fit this item and the size of the cohort each selects, scoped to the tenant it belongs to.

It is on every item, including ones that completed ok. An item that returned successfully and was wrong anyway is the whole reason this product exists; hiding the button behind a failure badge would hide it from exactly the item it was built for.

Picking a proposal takes you to Recovery with that selection applied, where you can see the items, narrow the window, and re-drive. The full loop — and what each kind of proposal can and cannot say — is in search by example.

This one was wrong

Beside it is the other half of the same question: This one was wrong. Find everything like this asks what else looks like this item; this one records the fact only a person has — that a well-formed, correct-looking output was wrong anyway.

Add a reason if you have one ("refunded — the summary named the wrong policy"), and it is stored as a signal against this item. It changes no verdict and re-runs nothing. What it does is make the item selectable: on Recovery, Only items a person flagged turns every flag anyone has left into a cohort you can pull, verify against and re-drive — and it composes with the two selections above, so everything like this item that a person also complained about is one checkbox.

An item accumulates signals rather than carrying one verdict, and the ones already on it are listed under the button. See backflow.

Needs attention — slice replay

Items whose outcome wasn't ok are the recovery surface. From a run, Replay re-drives just those not-ok items into a new run linked to the old one via replayed_from; items that already worked are never touched. You can replay the whole run's not-ok items, one tenant's slice, or a single item.

The same surface is available from the CLI:

papayya replay --run <run_id>              # re-drive the run's not-ok items
papayya replay --run <run_id> --tenant acme  # just one tenant's slice
papayya replay --item <item_id>            # a single item
papayya triage list                        # the needs-attention / DLQ view
papayya triage retry <item_id>             # retry one flagged item
papayya triage dismiss <item_id>           # acknowledge and clear it

Before any bulk replay fires, the dashboard shows the projected cost of re-running the selected items — total, p50, and p95 — so you can tell "$3 to retry 47 transient timeouts" apart from "$4,000 to retry 500 prompts that failed because the model is wrong," and confirm or cancel before spend.


Recovery

URL: /recovery

The cohort loop in the browser — see Recovery for what a cohort is and why it is a predicate rather than a run.

Three ways to select one, and you never type a predicate for the last two:

  • Name it — agent, tenant, verdict, window.
  • Paste an item you know is bad, and pick from the proposals.
  • Arrive from a detected change or from an item's Find everything like this, with the selection already applied.
  • Take what people flagged — Only items a person flagged, on its own or beside any of the above.

Ticking flagged moves the verdict selector to any, and says so. The items people flag are overwhelmingly the ones that completed ok, so leaving it on didn't work would return almost none of them — a confident, empty answer, which is the failure this product exists to catch.

A preview shows the count and the items — id, outcome, tenant, cost, and a marker on any item that is already a re-drive. Release re-drives the cohort, and the roll-up underneath answers the question you actually had: how many recovered, how many are still broken, how many were newly broken, and what it cost. Each row links into the per-item step diff, source beside re-drive.

Two things the page tells you rather than smoothing over: a by-example cohort's count is an upper bound on what a release will act on (items that are themselves re-drives are shown but not re-driven), and a cohort larger than the release cap says so, with the window controls to narrow it sitting directly above.


Agents

URL: /agents

Lists all registered agents with their name, description, model, and version.

Agent detail (/agents/{agentId}) shows the agent's configuration and recent runs.


Schedules

URL: /schedules

Manage cron schedules for your agents. Create, enable/disable, update, or delete schedules. Each fire is one run.

Each row shows the cron, the timezone it is evaluated in, Next Run, Last Run, and an Outcome chip for the most recent occurrence. Last Run moves only when a run was actually submitted, so the two together answer "was there a run, and did the last occurrence produce one?".

  • Expand a row for the occurrence history: one line per occurrence with the scheduled time, when we acted, the outcome, and why. This is where a skipped or missed occurrence explains itself.
  • Run now fires the schedule by hand using its own input, budget and deployed version. It works on a disabled schedule too, so a broken cron expression does not block running the thing while you fix it.
  • Status distinguishes a schedule a human turned off from one the scheduler disabled itself, and names the reason in the latter case.

Alerting on a schedule that stops firing is the schedule-missed signal — on by default, muted per agent in the alert rules. See Schedules.


Projects & API Keys

URL: /projects

Manage projects — logical groupings for agents and API keys.

API Keys (/projects/{projectId}/api-keys) — create and revoke API keys scoped to a project.


Usage

URL: /usage

View usage stats across your account:

  • Total runs, items, and cost
  • Breakdown by project
  • Time-based usage trends

See Usage & Billing for how metering works.


Navigation

The sidebar provides quick access to all sections. The dashboard uses a dark developer aesthetic with no UI framework — fast, minimal, and focused on the data.

Authentication uses JWT with automatic token refresh. If your session expires, you'll be redirected to the login page.