Skip to main content
A pattern is a detection rule checked against your traffic: when a trace matches, a signal is raised (or a healthy tally recorded) in the same triage queue everything else feeds. Patterns are the deterministic kind on the Scorers page - cheap and exact - alongside the six shipped built-in templates, while LLM judge scorers are the judgment kind, scoring quality continuously with a real LLM. Use a pattern when you can say precisely what bad looks like (“promises a refund”, “mentions a competitor”, “apologizes more than once”); use an LLM judge when you can only describe it.

Building a pattern

Scorers → New scorer → Pattern scorer. A pattern is one or more condition rows, evaluated top to bottom:
  • Detector - each row is one of three kinds:
    • Phrase: plain-text contains match.
    • Regex: a regular expression body. Don’t want to write one? The row’s Generate with AI popover drafts the regex from a plain-English description with an LLM call, right in the dialog - review and test it before saving.
    • Semantic: a rubric an LLM judges the text against (“The response promises a refund.”). The only detector kind that needs a provider key configured.
  • Case sensitivity - a per-condition-row toggle, applied to phrase and regex rows alike.
  • Match target - where the row looks: response (the traced output), userMessage (the traced input), or trace (output plus error plus every recorded tool call’s name/input/output, flattened - the right target for “did any tool mention X”).
  • Negate - flips a single row’s verdict before it joins the others.
  • Connector - rows combine top to bottom with AND, OR, or NOR (joins as “and not”), so “contains ‘refund’ AND NOT matches the approved-refund-template regex” is two rows.
Rows short-circuit: evaluation stops calling detectors once the verdict is already decided (false AND anything is false, true OR anything is true). Order matters for cost - put cheap phrase and regex rows before semantic rows, and the LLM call only happens when the cheap rows haven’t already settled the answer.

Behavior settings

Agent scoping (scopeMode / agentIds - all agents, or only the ones you name) exists on the wire and is enforced at detection time, but the pattern dialog doesn’t expose a picker for it yet - set it via the API if you need it. Phrase and regex detection work with no provider keys at all - only semantic rows need one.

Regex safety

Regex rows compile through RE2, which guarantees linear-time matching against attacker-influenced text. Two consequences:
  • Lookaround and backreferences never compile - RE2 does not support them. A row using them is refused at save time with a 400 that names the problem, so a pattern can never be stored with a regex that would silently never match.
  • Nested unbounded quantifiers (e.g. (a+)+) are rejected at save time, so that footgun can’t be stored at all.

Where matches go

Detection stops at the first matching pattern per trace - one signal per trace - and matches are deduped by pattern and agent: a recurring issue accumulates an occurrence count on one signal instead of a new row per trace. From a signal’s row on the Review tab you can record a verdict (Confirm, Fixed, Ignore, Wrong judgement), open the matched trace with View trace, or Add to dataset - turning the failure into a regression case, with an LLM assist that drafts the human feedback and expected results for you to edit. Finished with a signal? Archive it (per row, or “Archive selected” in bulk) - archived signals leave every filter except Review → Review signals’ explicit Archived filter, so the queue never grows unbounded. Archiving is a shelf, not a grave: if the same issue fires again, the signal reopens into the active list automatically.

From the SDK

Patterns are first-class SDK resources: client.monitor.patterns.builder(...).publish() returns an id you can pass at trace time via pattern_ids - a trace sent with pattern_ids is checked against only those patterns (built-in checks and sampling are bypassed for it). The Monitor page documents the builder’s full parameter table, the built-in checks that run alongside your custom patterns, and reading signals back with client.monitor.signals.

LLM Judge Scorers

The judgment half: continuous LLM scoring of live traffic

Custom Evaluators

Delegate the verdict to your own HTTP endpoint