Skip to main content
Distinct from Model portability (a what-if replay of one trace): Overview’s Model comparison card aggregates what your traffic already shows - every model observed in real traces over the window, side by side. No LLM calls are involved; it is a pure read of data you already have. Use it to notice that the cheaper model your team switched half the traffic to is also failing twice as often. There is deliberately no routing or canary infrastructure behind it: AgentX automates the comparison, not the split.

What it measures

Per model observed in production traffic:
  • Trace count - how much of the window’s traffic ran on this model.
  • Health rate / failure rate - from Monitor’s own event ledger. Both divide by the traces that actually produced a classification event, not by the raw trace count - so monitoring coverage doesn’t dilute the rates.
  • p95 latency - from the traces’ recorded latencies.
  • Estimated cost - measured token usage priced against the model catalog.
A metric with no data shows ”-” rather than a fabricated number: a model none of whose traces have a Monitor verdict shows no health rate (never a fake 100%), and an unpriced model shows no cost (never a fake $0.00 that reads as “this model is free”).

What it excludes

Evaluation-run traffic is excluded - the card reads production traces only, so a big offline run never skews the comparison. Models missing from the pricing catalog report null cost.

The window

24h, 7d (the default), or 30d. Rows sort by trace count descending - the top row is the model most of the traffic actually runs on.

Reading it back

GET /api/v1/agent-monitoring/model-comparison returns the card’s data (window as a query parameter). There is no dedicated SDK method today.