Skip to main content
Every edit to a dataset or LLM judge scorer is saved as a new version, in a log of its own per entity, browsable from a Version history panel (for a judge scorer, the panel lives in the editor’s header).

Dataset versions

A dataset version snapshots the full editable recipe, not just the display fields: name, description, status, questions, acceptanceCriteria, rejectionCriteria, evaluationCriteria, numberOfRequests, similarityConfig, and codeScorers. Restoring an old version therefore brings back the whole grading recipe as it was, criteria and scorers included.

Judge scorer versions

A judge scorer version snapshots the rubric and offline config: name, description, number of requests, the acceptance/rejection/evaluation criteria, judge prompt and judge model, the similarity metric toggles (vector, Jaccard, BLEU, ROUGE), code scorers, sovereignty settings, status, and tool context.

Using the history

  • Versions are shown newest-first, each with a v{N} badge and a summary of what changed ("Updated questions", "Updated acceptance criteria, judge model"). The summary is a synchronous field diff computed at save time, not an LLM-written description.
  • On any older version, click Apply this version to pull its fields back into the form - you still have to hit Save to actually restore it, nothing is overwritten on its own - or Delete to remove that version from the log. The current version offers neither.
  • The very first save is always recorded as "Created", so a freshly-made dataset never shows an empty history.
  • A no-op save (open the editor, change nothing, hit Save) records nothing - only genuine field changes create a version.
  • When judge tuning publishes a rubric change, the resulting version’s summary carries a provenance prefix, so a measured, validated rubric change is distinguishable from a manual tweak in the log.