Skip to main content
Distinct from Model portability (a what-if replay of one trace): Overview’s Model Comparison card aggregates what your traffic already shows - every model observed in real traces over the window, with average judge rating, latency, token usage, and estimated cost side by side. No LLM calls involved, it’s a pure read of data you already have; use it to notice that the cheaper model your team switched half the traffic to is also scoring a point lower.