Skip to main content
The CI/CD evaluation feature lets you run your agent against a test dataset in your CI pipeline and receive a binary PASS / FAIL gate result. If the gate fails, the pipeline exits with a non-zero code and blocks the merge or deploy.
Two backends, two flows. run_eval() and the create_ci_run family below use the hosted platform’s CI Ingest API (/ingest/ci-runs) and 404 against a self-host engine. On self-host, a CI run is an ordinary evaluation run plus the run gate - report.gate(...) / GET /runs/:id/gate - documented in the API reference.

Prerequisites

  1. Create an evaluation dataset in AgentX with at least one question.
  2. Enable CI/CD in the dataset settings and set a pass rate threshold.
  3. Export AGENTX_API_KEY in your environment.

High-level: run_eval()

The easiest path: one call handles the entire lifecycle:

Parameters

Return: CIRunResult

Low-level: step-by-step

For custom orchestration: parallel execution, streaming, or external agents:

Exception handling

Parallel question execution

Run multiple questions concurrently to speed up large datasets:

Inspecting scores

Polling a run

If you submit results asynchronously, poll until finalized:

GitHub Actions

See the GitHub Actions integration guide for a complete workflow template.

Gating a Custom Agent Evaluation run

The flow above is the dedicated CI/CD run type (pass-rate gate over per-question pass/fail). A regular Custom Agent Evaluation run can be gated too, on its average judge rating - useful when the same golden dataset serves both deep evaluation and the merge gate (self-host):
Set AGENTX_EVAL_QUIET=1 in CI to silence this runner’s interactive progress UI (spinners, per-case lines) - gate verdicts, results, and errors still print, keeping a gated run’s log to a few lines. It applies to the self-host custom-agent runner only, not tracer.run_eval.
For no_regression, the baseline must have measured the same thing: only completed runs with the same evaluationSettingsId, scorerGroupId, and split qualify, trace evaluations are excluded, and only the 5 newest candidates are considered. With no matching baseline the check passes explicitly (“nothing to regress against”). At least one check is required. gate() prints a CI-log-friendly per-check verdict and returns a GateResult with .passed, .exit_code (0/1), .average_rating, .baseline_average, .baseline_run_id, and the per-check .checks list; the caller decides whether to exit. Every recorded verdict lands in the dashboard’s CI Gates history. Two standalone forms cover runs created elsewhere:

pytest-native assertions

agentx.testing.assert_evaluation is the same gate as above in pytest ergonomics - a plain AssertionError subclass, so it works in any runner with no plugin registration, and the verdict still lands in the dashboard’s CI gate history (caller="pytest"):
On failure the test output carries the per-check verdict plus the run id, so the red CI job links straight to the run in the dashboard: