Live Traffic Scoring
Model Comparison
Per-model health, failure rate, latency, and cost aggregated from real traffic
Distinct from Model portability (a what-if replay of one trace): Overview’s Model comparison card aggregates what your traffic already shows - every model observed in real traces over the window, with trace count, health rate, failure rate, p95 latency, and estimated cost side by side. Health and failure come from Monitor’s own event ledger, cost from measured token usage priced against the model catalog. No LLM calls involved, it’s a pure read of data you already have; a metric with no data shows ”-” rather than a fake 100% or $0.00. Use it to notice that the cheaper model your team switched half the traffic to is also failing twice as often.

