.analyze() is the last step in the evaluation chain. It runs the same analysis as the dashboard’s “Analyze” button. On self-host the whole analysis runs synchronously in one request (the response reports mode: "sync"): the engine re-judges a sample of the rated results - the 12 worst plus the 5 best - with up to 3 judges, then produces the qualitative report in one final judge call, returned as a Report object.
Numeric scores (
average_rating, similarity metrics) are ready right after .finalize(), so you don’t need to wait for .analyze() just to see how the run went. Because .analyze() polls a job until it completes, it can take noticeably longer than a single LLM call for larger runs; progress is shown in the terminal while it waits.Controlling the analysis
Check on a long-running analysis without calling
.analyze() again, even from a separate script execution:
Fields
Each
rating sub-field is one of "high", "medium", "low". recommendations[].category is one of "instructions", "tools", "knowledge", "reasoning", "consistency", "other"; priority is "high", "medium", or "low".
Reading it in your terminal
run_context.run_id and open that run under Evaluate > Runs.
Per-result rows without an analysis
results() returns the scored rows - no analysis job, no extra judge calls - for asserting
on individual results in scripts and CI. Rows are typed RunResultRow objects with snake_case
attributes (dict-style access still works but is deprecated and warns; row.raw keeps the
full wire dict):
rating/justification, a status ("scored", "skipped" when the judge
could not score it - rating stays None, not 0 - or "failed" when the result carried an
error). A failed row is stored with rating 0 and IS counted into averageRating/ratedCount
and any CI gate, while a skipped row stays None and is not. Rows also carry the linked
trace_id, latency_ms and token counts, and the similarity metrics
(cosine_similarity, jaccard_similarity, bleu_score, rouge_score).
client.evaluations.get_run(run_id) fetches the same payload later, from a different process.
