Dataset versions
A dataset version snapshots the full editable recipe, not just the display fields:name, description, status, questions, acceptanceCriteria, rejectionCriteria, evaluationCriteria, numberOfRequests, similarityConfig, and codeScorers. Restoring an old version therefore brings back the whole grading recipe as it was, criteria and scorers included.
Judge scorer versions
A judge scorer version snapshots the rubric and offline config: name, description, number of requests, the acceptance/rejection/evaluation criteria, judge prompt and judge model, the similarity metric toggles (vector, Jaccard, BLEU, ROUGE), code scorers, sovereignty settings, status, and tool context.Using the history
- Versions are shown newest-first, each with a
v{N}badge and a summary of what changed ("Updated questions","Updated acceptance criteria, judge model"). The summary is a synchronous field diff computed at save time, not an LLM-written description. - On any older version, click Apply this version to pull its fields back into the form - you still have to hit Save to actually restore it, nothing is overwritten on its own - or Delete to remove that version from the log. The current version offers neither.
- The very first save is always recorded as
"Created", so a freshly-made dataset never shows an empty history. - A no-op save (open the editor, change nothing, hit Save) records nothing - only genuine field changes create a version.
- When judge tuning publishes a rubric change, the resulting version’s summary carries a provenance prefix, so a measured, validated rubric change is distinguishable from a manual tweak in the log.

