
How the core objects relate: traces group into sessions; scorers raise signals; datasets drive eval runs; findings become validated registry proposals.
Tracing
| Term | Meaning |
|---|---|
| Trace | One agent interaction: a root record with input, output, latency, model, tokens, and tool calls. See Tracing. |
| Span | One step inside a trace - a graph node, LLM call, tool execution, or retrieval - linked by spanId/parentSpanId into a tree. The trace dialog’s Timeline and Graph views render this tree. See Span kinds. |
| Session | The conversation a trace belongs to: every trace sharing a session_id. Sessions get their own Observe view, turn counts, a conversation-level coherence score, and session-scoped judging once the conversation goes idle. See Sessions. |
| Tool call | A recorded tool invocation (name, arguments, result, success flag). Feeds tool-quality metrics, the Tool-failure check, and trajectory matching. |
| Agent | The named producer of traces. Auto-registered on first ingest; carries a monitoring profile and a healthy rate in the Agents tab. |
Monitoring
| Term | Meaning |
|---|---|
| Scorer | Anything that can pass judgment on traffic, managed from the one Scorers tab: patterns, LLM judge scorers, and custom scorers. All opt-in - nothing is judged until a scorer is enabled. |
| Pattern | A deterministic scorer: one of the six shipped templates (secrets, PII, prompt-injection echo, profanity, refusal / non-answer, malformed JSON) or a rule you author from phrase, regex, and semantic (LLM-judged) conditions. Polarity is failure (raises an issue) or proper (an informational match, counted as neither a failure nor a healthy run). |
| Operational outcome | NOT a scorer: a fact the trace itself recorded - an error, a failed tool call, an empty response. Always classified into the KPI failure metrics, never raises a triage signal, nothing to enable. |
| Signal | One deduped finding: repeated matches of the same scorer for the same agent accumulate on one row (occurrence count, last-seen) instead of flooding the list. Triage with statuses, archive or resolve what’s done - a re-firing archived or resolved signal reopens automatically. |
| LLM judge scorer | ONE judge entity usable everywhere: a rubric (criteria + judge prompt + judge model + tool-context level), an offline profile (dataset-run grading - the scorer’s id IS the scorer_id runs take), and an optional online profile whose single Score live traffic switch is the only gate (per trace, or per session once a conversation goes idle). Low live scores raise signals. (Its halves are “evaluation settings” and “online evaluator” in the legacy SDK/API surfaces - same entity, older names.) |
| Custom scorer | Your own logic per sampled trace: an external HTTP endpoint returning a verdict, or a code scorer - a Python or JavaScript handler() you write, run in-engine, scoring 0-1. See Custom scorers. (“Custom evaluator” in the SDK/API.) |
| User feedback | An end user’s thumbs up/down forwarded from your app (client.feedback.report) - human ground truth attached to the trace: raises a triage signal on a downvote, drives the Downvote rate KPI, and feeds judge calibration. Not a scorer; nothing to configure. See User feedback. |
| Topic | Automatic clustering of traffic by subject: the Topics tab and the Overview topic map, and the traffic side of dataset coverage. |
| Outcome | A real-world result you report back (client.outcomes.report) - refund issued, ticket reopened - used to compare judge verdicts against reality. See Outcomes. |
| Judge calibration | Overview’s scoreboard for the judges themselves: how often automated verdicts agreed with recorded reality (reported outcomes, user votes, human re-scores). See Outcomes. |
| Judge tuning | Rewriting an online judge’s grading criteria from its recorded disagreements with reality - validated by exact re-judging (fixes the cases it got wrong, preserves a control set it got right) before a human publishes. |
| Alert rule | A threshold on an aggregate - failure rate, p95 latency, estimated cost, judge failures, trace count - over a sliding window, paged to Slack, Teams, PagerDuty, email, or a webhook with a firing/resolved lifecycle (Alert rules). Distinct from a scorer’s per-verdict alert threshold and from an automation rule’s per-trace routing. |
| Model comparison | An Overview card aggregating quality, cost, and latency per model from real traffic - which model is actually earning its bill. See Model comparison. |
Evaluation
| Term | Meaning |
|---|---|
| Dataset | A versioned, reusable set of test cases plus scoring configuration (criteria, metrics, code scorers, CI settings). Runs snapshot the dataset version they used. See Build a dataset; cases can also be curated from production traces. |
| Question | One test case: query, optional expectedResults (the reference answer), optional per-case judgeGuideline, optional expected trajectory, optional smoke-test variants (LLM-paraphrased rewordings to catch phrasing brittleness). |
| Evaluator config | Legacy name for an LLM judge scorer’s rubric + offline profile (the evaluation settings record). Not a separate entity since the unification - managed from the Scorers tab, and its id is the scorer’s id. |
| Evaluation run | One execution of a dataset against an agent: every output, its scores, the locked dataset version, and the subject description. Statuses: in_progress → completed or failed. |
| Judge rating | The LLM judge’s 0-10 score with a written justification. When a result links its trace, the judge also sees the execution trajectory - tools called, order, failures - not just the final text. |
| Similarity metrics | Optional per-dataset scorers against expectedResults: vector (cosine) similarity, Jaccard token overlap, BLEU, ROUGE-L. |
| Code scorer | A JavaScript function run in-engine against each result - deterministic checks (format, keywords, tool behavior) that need no LLM. See Code scorers. |
| Trajectory match | A deterministic pass/fail comparing the tool calls a linked trace actually made against the case’s expectedTrajectory (a tools list plus a mode): strict (same calls, same order), unordered (same calls, any order), superset (all expected present, extras allowed), subset (nothing unexpected, missing allowed). |
| Model portability | Replay a captured trace’s input against alternative models for a quick cost/latency/quality comparison, without touching your agent. See Model portability. |
| Evaluation subject | Metadata describing what was evaluated (kind, displayName, framework, optional version for version comparison). |
CI/CD
| Term | Meaning |
|---|---|
| Gate | The pass/fail decision for a pipeline: rating floor (fail_under) and/or no-regression versus a baseline run. Recorded gates land in the dashboard’s CI Gates view. See CI/CD. |
| Pass rate | Fraction of cases passing every threshold rule (0.0-1.0). |
| Threshold | A per-metric rule on the dataset (“judge rating ≥ 7”, “cosine ≥ 0.8”) counted into the pass rate. |
| Fail fast | Finalize the run as failed on the first threshold violation instead of waiting for remaining cases. |
Insights and Improve
The old Improve tab merged into Insights, which holds three views: Suggestions, Dataset coverage, and Auto-improve (confirmed failures become an improvement report that becomes a code fix - see Auto-improve). The Prompts and Tools & MCPs registries live under the sidebar’s Manage section.| Term | Meaning |
|---|---|
| Suggestion | One queued improvement proposal in Insights → Suggestions (the improvement inbox): a background sweep notices fresh failure evidence on a prompt or tool schema, generates the proposal, runs its baseline-vs-candidate validation, and queues it with the measured verdict attached. Humans review, then publish or dismiss. |
| Dataset coverage | Insights → Dataset coverage: does your dataset look anything like your traffic? Joins production topics to dataset cases and reports traffic-weighted coverage, topic breadth, and risk-weighted coverage, plus a probe for whether one specific query (or a pasted list) is covered. |
| Prompt registry | Versioned prompts your agent pulls at runtime (client.evaluations.prompts), with judge-proposed, human-approved rewrites. See Prompt Management. |
| Tool schema registry | Versioned tool definitions with failure evidence gathered from traces and playground runs, and “Suggest improvement” proposals. See Tool schemas. |
| Proposal validation | Before adopting a proposed change, AgentX runs candidate-vs-current against a dataset and shows the measured difference. See Validating proposals. |

