version just needs to be a key inside metadata.
Reading the verdict
From the dashboard: Governance → Evaluate → Datasets → a dataset’s row menu → Compare versions. It shows every version that’s been run against that dataset (run count, rated count, average rating) and a headline verdict comparing the two most recent versions, the samecandidateAvg >= baselineAvg check autotune’s own validation step uses. Nothing gets merged or shipped automatically - you read the verdict and go edit your own code.
The dialog also includes a Case by case drill-down: each version’s latest run diffed per question, with rating deltas, regressions highlighted, and both outputs (plus the judge’s justification) one click away - the averages say whether a version regressed, this says where.
