Two backends, two flows.
run_eval() and the create_ci_run family below use the hosted
platform’s CI Ingest API (/ingest/ci-runs) and 404 against a self-host engine. On
self-host, a CI run is an ordinary evaluation run plus the
run gate - report.gate(...) /
GET /runs/:id/gate - documented in the
API reference.Prerequisites
- Create an evaluation dataset in AgentX with at least one question.
- Enable CI/CD in the dataset settings and set a pass rate threshold.
- Export
AGENTX_API_KEYin your environment.
High-level: run_eval()
The easiest path: one call handles the entire lifecycle:
Parameters
Return: CIRunResult
Low-level: step-by-step
For custom orchestration: parallel execution, streaming, or external agents:Exception handling
Parallel question execution
Run multiple questions concurrently to speed up large datasets:Inspecting scores
Polling a run
If you submit results asynchronously, poll until finalized:GitHub Actions
See the GitHub Actions integration guide for a complete workflow template.Gating a Custom Agent Evaluation run
The flow above is the dedicated CI/CD run type (pass-rate gate over per-question pass/fail). A regular Custom Agent Evaluation run can be gated too, on its average judge rating - useful when the same golden dataset serves both deep evaluation and the merge gate (self-host):Set
AGENTX_EVAL_QUIET=1 in CI to silence this runner’s interactive progress UI (spinners,
per-case lines) - gate verdicts, results, and errors still print, keeping a gated run’s log
to a few lines. It applies to the self-host custom-agent runner only, not tracer.run_eval.
For
no_regression, the baseline must have measured the same thing: only completed runs with the
same evaluationSettingsId, scorerGroupId, and split qualify, trace evaluations are excluded,
and only the 5 newest candidates are considered. With no matching baseline the check passes
explicitly (“nothing to regress against”).
At least one check is required. gate() prints a CI-log-friendly per-check verdict and
returns a GateResult with .passed, .exit_code (0/1), .average_rating,
.baseline_average, .baseline_run_id, and the per-check .checks list; the caller decides
whether to exit.
Every recorded verdict lands in the dashboard’s CI Gates history. Two standalone forms cover
runs created elsewhere:
pytest-native assertions
agentx.testing.assert_evaluation is the same gate as above in pytest ergonomics - a plain
AssertionError subclass, so it works in any runner with no plugin registration, and the
verdict still lands in the dashboard’s CI gate history (caller="pytest"):

