version just needs to be a key inside metadata, or a top-level version key on the subject itself (checked first when both are present).
Reading the verdict
From the dashboard: the sidebar’s Manage → Datasets → a dataset’s row menu → Compare versions. It shows every version that’s been run against that dataset (run count, rated count, average rating) and a headline verdict comparing the two most recent versions, the samecandidateAvg >= baselineAvg check autotune’s own validation step uses. The per-version average is a true average over every rated result in that version’s runs - not an average of per-run averages, which would misweight versions whose runs have different result counts - and single-trace evaluations are excluded from the version buckets entirely. Nothing gets merged or shipped automatically - you read the verdict and go edit your own code.
The dialog also includes a Case by case drill-down: each version’s latest run diffed per question, with rating deltas, regressions highlighted, and both outputs (plus the judge’s justification) one click away - the averages say whether a version regressed, this says where.
To gate a deploy on this same comparison automatically, see the self-host CI gate’s no_regression check.
