What it measures
Per model observed in production traffic:- Trace count - how much of the window’s traffic ran on this model.
- Health rate / failure rate - from Monitor’s own event ledger. Both divide by the traces that actually produced a classification event, not by the raw trace count - so monitoring coverage doesn’t dilute the rates.
- p95 latency - from the traces’ recorded latencies.
- Estimated cost - measured token usage priced against the model catalog.
What it excludes
Evaluation-run traffic is excluded - the card reads production traces only, so a big offline run never skews the comparison. Models missing from the pricing catalog reportnull cost.
The window
24h, 7d (the default), or 30d. Rows sort by trace count descending - the top row is the
model most of the traffic actually runs on.
Reading it back
GET /api/v1/agent-monitoring/model-comparison returns the card’s data (window as a query
parameter). There is no dedicated SDK method today.
