Skip to main content
Run your golden dataset against the PR’s version of your agent, then gate the run: an absolute rating floor, a no-regression check against the dataset’s previous run, or both. The gate is a plain exit-code contract, so it drops into any CI system.
The two checks:
  • fail_under - the run’s average judge rating must be at or above the floor. Works from the very first run.
  • no_regression - the average must not drop more than tolerance (default 0.5) below the dataset’s previous completed run. The tolerance exists because judge scores are noisy: an exact comparison would flake builds on variance, not regressions. Requires run history, so point CI at a persistent self-host instance rather than an ephemeral one.
The baseline must have measured the same thing as the gated run: only completed runs with the same evaluationSettingsId, scorerGroupId, and split qualify, single-trace evaluations are excluded, and only the 5 newest candidates are considered. With no matching run the check passes explicitly (“no previous completed run - nothing to regress against”). That means changing the grading config, the scorer group, or the split silently resets the baseline and no_regression passes on the next run - by design, but know it when you rework grading. On a multi-judge run, scorer gates a named additional judge scorer’s average instead of the primary rating - run.gate(fail_under=8, scorer="Safety") fails the build when Safety is low even if the blended average looks fine. The name (or id) must match an additional judge scorer on the run; an unknown name is a 400, and deterministic scorer-group members (pattern/code kinds) have no per-run average to gate. gate() prints a per-check verdict into the CI log and returns a GateResult (passed, exit_code, average_rating, baseline_average, checks). The same check is a plain HTTP call if you’d rather script it raw - add record=true (what the SDK sends by default) so the verdict lands in the dashboard’s gate history, with an optional caller label. Note that with record=True the SDK deliberately disables transport retry: despite the GET verb, recording is a server-side write, and a timeout after the row was stored would be retried into a duplicate verdict:

GitHub Actions

Pointing at a persistent self-host instance (recommended - run history enables no_regression):
No persistent instance reachable from CI? Launch an ephemeral engine inside the job and gate on the floor only (a fresh database has no baseline to regress against):

What to gate on

Every recorded gate lands in the dashboard’s CI Gates history - the same list is readable from the SDK for scripting (e.g. a weekly “how often did gates fire” report):
The strongest gate is a dataset curated from production: real failures as regression cases mean “the PR doesn’t re-break what production already taught you.” Pair fail_under (an absolute quality bar) with no_regression (this PR made nothing worse) - they catch different failure modes.

Gate history in the dashboard

Every gate the SDK runs is recorded by default (record=True, with a caller label like "github-actions"), and the dashboard’s CI Gates tab (in the sidebar under Automations) lists that history newest-first: verdict, dataset, average vs baseline, which checks ran, who called. The same page has a preview (“would the latest run pass these thresholds?”) that is never recorded - so exploring thresholds can’t pollute the history real CI jobs write - plus the copy-paste setup snippets. Pass record=False to gate_run() for an unrecorded check from code.

Datasets from Production

Grow the golden dataset the gate runs against

Hosted CI/CD API

The hosted platform’s equivalent pipeline

TypeScript CI SDK

The same gate from a TypeScript pipeline, via @agentx/eval