session_id. Sessions catch
the failure mode single-trace monitoring can’t see: every individual reply looks fine, but the
conversation as a whole goes in circles, contradicts itself, or never resolves.
Grouping is opt-in per call - pass any stable per-conversation id:
The Sessions view
Observe → Sessions lists each conversation with:
Failed spans anywhere in the conversation surface as an error-count badge inside the Session
cell. Opening a session shows every turn in order; each turn’s full span tree is one click away.

A session's detail view: turns in order, judge verdicts, and Add to dataset.
Conversation-level judging
Three kinds of scorer judge a session as a whole, all triggered automatically when the conversation goes quiet:- Session Baseline Judge - built-in: goal progression, consistency, non-repetition. Also runs on demand from the session’s detail view. Its rubric lives in an LLM judge scorer (tunable via Tune judge); pause it from the Scorers page’s row toggle.
- Session-scoped judge scorers - your own criteria with
scope="session", judged on idle and re-judged if the conversation resumes. See Multi-Turn Session Evaluation. - Session-scoped scorer groups - several scorers (judges, patterns, code) blended into one 0-10 verdict per conversation, on the same idle trigger, with must-pass gates.

