- An external scorer POSTs each sampled trace to your own HTTP endpoint and uses its verdict.
- A code scorer runs a script you write (Python or JavaScript) inside the engine itself - no endpoint to host.
custom-evaluators.)
Code scorers
Scorers → New scorer → Code scorer. The script defines one function:async function handler(...),
await trace.getSpans({ spanType: ["llm"] })). A returned score below the scorer’s alert
threshold (default 0.5) raises a signal; every scored check is recorded in the scorer’s event
history either way, and None/null skips the trace entirely.
trace.get_spans() returns the trace’s own span subtree (root first, start-time ordered). Each
span carries span_id, parent_span_id, name, input, output, error, model,
latency_ms, input_tokens, output_tokens, tool_calls, metadata, started_at, and a
derived type - the engine-wide span kind (agent, llm, tool,
retrieval, chain, embedding, reranker, guardrail, evaluator, prompt, memory).
Every span, the root included, classifies through the same ladder: a producer-stated
span_kind always wins, a span with a model infers llm, one with tool calls infers tool,
and an unclassifiable span is chain; the older "span" value is retired and never emitted,
the Dry run synthetic sample included (its root is "agent").
Match the root by its parent_span_id being null, not by type.
Dry run in the dialog executes the script against a synthetic two-span sample (one root, one
llm span) and shows the score it returns.
External scorers
Request (engine → your endpoint,POST, JSON, schemaVersion: 2):
trace (the v1 keys -
input/output/error/toolCalls - are exactly where they always were, so v1 endpoints keep
working untouched), and the trace’s span subtree under spans with the same per-span fields and
derived type values code scorers see (the full span-kind vocabulary - llm, tool, retrieval, chain, memory, …).
Response your endpoint must return:
matches is the only field that decides anything. reason and score are both optional and purely informational - recorded and shown alongside the resulting signal, but neither affects whether one is raised. A minimal example endpoint (Flask):
Setup and behavior
Build it from the dashboard: Scorers → New scorer → External scorer. The dialog documents the request/response contract in place, and Dry run sends a synthetic v2 payload to your URL so you can confirm it responds correctly before saving.
External scorer URLs pass the engine’s egress guard:
http(s) only, cloud metadata endpoints
are always refused, and private/loopback targets are blocked on multi-tenant deployments (on
regular self-host they stay allowed - pointing a scorer at your own local endpoint is the
normal deployment shape). A blocked URL is a 400 at save time, not a silent failure at check
time.
Distinct from
monitor_profiles.channels’ "webhook:<url>" notification targets - those fire-and-forget a message after a signal is raised and never consume a response. A custom evaluator’s response is the detection result, awaited synchronously (up to the 8-second timeout).client.monitor.scorers:
sample_ratedefaults to 0.1 for both kinds - every check is a real outbound HTTP call or a script execution, so sampling starts conservative.- The REST equivalents live on
/agent-monitoring/custom-evaluators:GET/POSTon the collection,PUT/DELETEon/:id, andGET /:id/eventsfor the check history (code scorers pass{"kind": "code", "language": "python"|"javascript", "script": "...", "alertBelow": 0.5}). POST /agent-monitoring/custom-evaluators/dry-runis transient -{"url": ...}for external,{"kind": "code", "language": ..., "script": ...}for code - the live response back out, nothing persisted.

