SDK Reference
BYOF Observability

BYOF observability

Papayya items every model call as a step — model, tokens, duration, stop reason — and inspects what came back. There are two ways to mark a model call, and one auto-path:

  • @papayya.llm on the leaf function that calls your provider. Works inside @agent bodies and under an explicit papayya().item(...) handle — the ambient path. Your orchestration never changes.
  • item.llm_step(label, fn) on the explicit item handle, when you're already wrapping steps by hand.
  • Auto-patch — in the hosted runtime a shim wraps OpenAI and Anthropic clients at import time, so openai.chat.completions.create(...) and anthropic.messages.create(...) are captured even without a decorator.

Every other provider (Gemini, Bedrock, Groq, Mistral, Cohere, a raw httpx call, a proprietary internal API) is invisible to the auto-patch path. For these, mark the leaf with @papayya.llm (or wrap it with item.llm_step) and the same telemetry lands — one step row per call, same dashboard, same rollups.

The one-line change

import papayya
from papayya import agent
 
@papayya.llm
def generate(prompt: str):
    import google.generativeai as genai
    model = genai.GenerativeModel("gemini-2.0-flash")
    return model.generate_content(prompt)
 
@agent(name="gemini-agent")
def gemini_agent(run, prompt: str) -> str:
    return generate(prompt).text

Explicit-handle equivalent. On a papayya().item(...) handle, wrap the call and then invoke it: gen = item.llm_step("generate", generate); resp = gen(prompt). Same telemetry, same durability. (item.step(..., kind="llm") is a deprecated form — use llm_step.)

What @papayya.llm / llm_step does on each call:

  1. Runs the wrapped function normally.
  2. Inspects the response shape for usage metadata. Recognized shapes: OpenAI (response.usage.prompt_tokens), Anthropic (response.usage.input_tokens), Gemini (response.usage_metadata.prompt_token_count), and OpenAI-compat / Anthropic-compat dict responses.
  3. Items a step row with prompt_tokens, completion_tokens, model, and stop_reason when any of those shapes match.
  4. Falls back to provider_shape="unknown" when nothing matches — the step still items that the call happened and how long it took; you just lose token granularity.
  5. Runs outcome inspectors on the return value — a refusal, an empty result, or a degenerate stop reason flips the item to degraded, with no check written anywhere.
  6. Classifies any exception raised by the call. Credit-exhaustion errors (HTTP 402, insufficient_quota, credit_balance_too_low, and similar) become CreditExhausted, which pauses the run instead of failing it. Transient errors (rate limits, 5xx) and permanent errors (401, 404) propagate unchanged.

Three examples

Gemini

@papayya.llm
def gemini_call(prompt: str):
    import google.generativeai as genai
    return genai.GenerativeModel("gemini-2.0-flash").generate_content(prompt)

Extracts: prompt_tokens (from usage_metadata.prompt_token_count), completion_tokens (from candidates_token_count), model (from model_version), stop_reason (from candidates[0].finish_reason).

Bedrock (boto3)

import boto3, json
 
@papayya.llm
def bedrock_claude(prompt: str) -> dict:
    bedrock = boto3.client("bedrock-runtime")
    raw = bedrock.invoke_model(
        modelId="anthropic.claude-3-5-sonnet-20241022-v2:0",
        body=json.dumps({"anthropic_version": "bedrock-2023-05-31",
                         "max_tokens": 1024,
                         "messages": [{"role": "user", "content": prompt}]}),
    )
    return json.loads(raw["body"].read())

Bedrock returns a dict shaped like the underlying Anthropic response, so the dict branch of the extractor matches — you get Anthropic-style token counts without any per-provider code.

Raw httpx against an internal gateway

import httpx
 
@papayya.llm
def internal_gateway(messages: list[dict]) -> dict:
    return httpx.post(
        "https://internal-llm.corp.local/v1/chat/completions",
        json={"model": "internal-llama", "messages": messages},
        headers={"Authorization": f"Bearer {token}"},
        timeout=60.0,
    ).raise_for_status().json()

As long as the response is a dict with a usage field (either prompt_tokens or input_tokens), token counts land. Otherwise the step items duration-only with provider_shape="unknown".

Provider-support matrix

Provider familyAuto-patchedCaptured via @papayya.llm / llm_stepHow
OpenAI✓✓ (dedupe keeps one row)SDK patch
Anthropic✓✓ (dedupe keeps one row)SDK patch
Gemini (google-generativeai)—✓usage_metadata shape
OpenAI-compat (Groq, Fireworks, Together, vLLM, Ollama /v1/*, Mistral)—✓OpenAI shape or dict
Anthropic-compat (Bedrock Claude)—✓Anthropic shape or dict
Any httpx / requests / custom—✓ when response is a dict with a usage fielddict fallback
Exotic / binary responses—✓ (items duration only)provider_shape="unknown"

When both paths apply (e.g. an OpenAI call wrapped with @papayya.llm), the shim emits one step row — the interceptor's, not a duplicate.

What the dashboard shows

Per step:

  • prompt_tokens / completion_tokens / total_tokens (when extractable)
  • model (the provider-reported identifier, e.g. gpt-4o-mini-2024-07-18, claude-sonnet-4-20250514, gemini-2.0-flash-001)
  • stop_reason (e.g. stop, length, end_turn, max_tokens, STOP) — useful for spotting silent truncation
  • duration_ms
  • the step's outcome (ok or degraded, with a reason token)

Cost for unpatched providers is recorded as 0 cents by default. Papayya does not ship a pricing table for arbitrary providers; a ballpark-wrong number is worse than none. To get real cost on the dashboard for these calls, set PAPAYYA_COST_FN=my_module:my_cost_fn on the deployed agent.

Custom outcome checks

The built-in inspectors are structural — empty result, degenerate embedding, degraded stop reason. They catch failures that look wrong regardless of your domain. For failures only you can define — "an answer under 20 characters is degraded", "the JSON is missing a required field", "the tone is wrong" — register a check.

A check is a plain callable that receives the step's result and returns a verdict, or None to pass:

import papayya
from papayya import agent, CheckVerdict
 
def not_too_short(result):
    if isinstance(result, str) and len(result) < 20:
        return CheckVerdict("degraded", "too_short")
    return None   # pass
 
@agent(name="summarizer", checks=[not_too_short])
def summarize(run, doc): ...

Checks run in the same pipeline as the built-in inspectors, on every step. The worst verdict across built-in and custom wins the run's outcome — one pipeline, one worst_outcome_status, one dashboard. Your reason token is namespaced under user: so the Outcome quality view can group custom checks. A check is an observer: if it raises, it's logged and treated as a pass — a bug in a check never fails your run.

This is the same BYO posture as BYO LLM keys and BYO retrieve/embed/search: Papayya owns execution and the pipeline; you bring the check.

LLM-as-judge

For quality only a model can assess, papayya.llm_judge is a check scaffold you parameterize with a rubric and your own model-invoking callable — run on your key. Papayya formats the rubric and result into a judge prompt, calls your callable, parses a PASS/FAIL, and maps FAIL to a degraded verdict (user:judge:<name>):

from papayya import agent, llm_judge
 
def call_my_model(prompt: str) -> str:
    # YOUR model, YOUR key — Papayya never spends its tokens judging your traffic.
    return openai.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
    ).choices[0].message.content
 
tone = llm_judge(
    name="tone",
    model=call_my_model,
    rubric="Is the reply polite and on-topic?",
    sample_rate=0.2,   # judge 20% of runs — bounds cost
    timeout=10.0,
)
 
@agent(name="support-bot", checks=[tone])
def reply(run, ticket): ...

Two honesty boundaries are baked in:

  • BYO key only. The callable is yours; Papayya never spends its own tokens judging your traffic.
  • Sampled + bounded. The judge runs inline on a configurable fraction of runs (sample_rate, sampled per run) with a hard timeout. A slow judge, an error, or an unparseable response is a contained pass — never a run failure. Accept the inline cost; sampling is what bounds it.

(A truly asynchronous, off-hot-path judge is a later addition — it needs a step-update + rollup-rewrite path this scaffold deliberately doesn't build.)

Outcome quality

A single run's badge answers "did this one work". The Outcome quality section on each agent's detail page answers the question the wedge promises — is this agent getting worse? — by rendering outcome as a rate, not a flag:

  • a headline ran-vs-worked rate ("94% of runs worked · last 30 days");
  • a trend of the non-ok rate by day;
  • a reason breakdown — which inspector or custom-check token dominates (custom checks grouped);
  • a worst-offenders list of recent non-ok runs, into their step detail.

Every custom-check verdict surfaces here alongside the built-in ones — the user:-namespaced reasons roll up into the "custom checks" bucket.

OpenTelemetry

Papayya stamps the active OpenTelemetry span (and baggage) with the agent name, item_id, and partition_key while your body runs. If you already export OTel traces, your spans carry the same agent → item → tenant correlation Papayya uses internally — no extra wiring. Absent an OTel pipeline this is a no-op; steps still land in the Papayya ledger.

Error semantics

from papayya import CreditExhausted
 
try:
    result = generate(prompt)          # an @papayya.llm leaf
except CreditExhausted:
    # The runtime auto-pauses the run when this bubbles up. Catch only if
    # you have a deliberate fallback path (e.g. a secondary provider).
    ...

You rarely catch CreditExhausted yourself — letting it bubble up pauses the run, and you top up the provider and resume. It's importable from papayya if you need alternate-provider fallback logic, or to raise it manually from an exotic provider's exception path:

from papayya import CreditExhausted
 
try:
    response = exotic_sdk.generate(prompt)
except exotic_sdk.OutOfCreditsError as e:
    raise CreditExhausted(f"exotic provider out of credits: {e}") from e

CreditExhausted is the provider-side sibling of Papayya's budget auto-pause: when a run crosses your Papayya budget cap, the fence pauses it (run status paused) rather than failing it, and the SDK raises WorkloadPaused at the next step boundary. Both pause-and-preserve rather than fail.

Durability interaction

An LLM step is still a checkpointed step — the call result is durably stored. On resume after a crash, the cached response is returned; the model is not called again. An LLM call is the most expensive side effect your code makes, so durable memoization is often the difference between a cheap resume and a costly one.

What gets stored

Papayya serializes the step's return value to JSON for durable storage. Provider response objects — OpenAI ChatCompletion, Anthropic Message, Gemini SDK objects, SimpleNamespace, dataclasses, plain classes — are not JSON-native, so Papayya runs them through a shape ladder:

  1. json.dumps — JSON-native values pass through unchanged.
  2. .model_dump() — Pydantic v2 models.
  3. .dict() — Pydantic v1 models.
  4. dataclasses.asdict — dataclasses.
  5. vars(obj) — any object with __dict__, including SimpleNamespace.
  6. repr(obj) — opaque objects as a last resort.

The token/model/stop_reason columns are extracted before this ladder runs, so they always reflect the real response regardless of how the raw object degrades. On replay the cached result comes back as whatever the ladder produced — typically a dict, not the original SDK class. Write code that doesn't depend on the return value being a specific class type across a crash boundary.

When not to use it

  • For OpenAI / Anthropic calls in the hosted runtime, marking the leaf is optional — the auto-patch already items the same telemetry. Marking them is fine (dedupe keeps it clean) but adds no new information. Do it anyway if you want the same code to item identically when run locally.
  • For non-LLM side effects (HTTP calls, DB writes, tool execution), wrap the call with item.step("label", fn) (the verified per-item form). The token extractor and credit classifier do nothing useful on non-LLM responses, but you still get a durable, outcome-inspected step.