fail_under- the run’s average judge rating must be at or above the floor. Works from the very first run.no_regression- the average must not drop more thantolerance(default0.5) below the dataset’s previous completed run. The tolerance exists because judge scores are noisy: an exact comparison would flake builds on variance, not regressions. Requires run history, so point CI at a persistent self-host instance rather than an ephemeral one.
evaluationSettingsId, scorerGroupId, and split qualify, single-trace evaluations are excluded, and only the 5 newest candidates are considered. With no matching run the check passes explicitly (“no previous completed run - nothing to regress against”). That means changing the grading config, the scorer group, or the split silently resets the baseline and no_regression passes on the next run - by design, but know it when you rework grading.
On a multi-judge run, scorer gates a named additional judge scorer’s average instead of the primary rating - run.gate(fail_under=8, scorer="Safety") fails the build when Safety is low even if the blended average looks fine. The name (or id) must match an additional judge scorer on the run; an unknown name is a 400, and deterministic scorer-group members (pattern/code kinds) have no per-run average to gate.
gate() prints a per-check verdict into the CI log and returns a GateResult (passed, exit_code, average_rating, baseline_average, checks). The same check is a plain HTTP call if you’d rather script it raw - add record=true (what the SDK sends by default) so the verdict lands in the dashboard’s gate history, with an optional caller label. Note that with record=True the SDK deliberately disables transport retry: despite the GET verb, recording is a server-side write, and a timeout after the row was stored would be retried into a duplicate verdict:
GitHub Actions
Pointing at a persistent self-host instance (recommended - run history enablesno_regression):
What to gate on
Every recorded gate lands in the dashboard’s CI Gates history - the same list is readable from the SDK for scripting (e.g. a weekly “how often did gates fire” report):fail_under (an absolute quality bar) with no_regression (this PR made nothing worse) - they catch different failure modes.
Gate history in the dashboard
Every gate the SDK runs is recorded by default (record=True, with a caller label like "github-actions"), and the dashboard’s CI Gates tab (in the sidebar under Automations) lists that history newest-first: verdict, dataset, average vs baseline, which checks ran, who called. The same page has a preview (“would the latest run pass these thresholds?”) that is never recorded - so exploring thresholds can’t pollute the history real CI jobs write - plus the copy-paste setup snippets. Pass record=False to gate_run() for an unrecorded check from code.
Datasets from Production
Grow the golden dataset the gate runs against
Hosted CI/CD API
The hosted platform’s equivalent pipeline
TypeScript CI SDK
The same gate from a TypeScript pipeline, via @agentx/eval

