Skip to main content
AgentX is a governance platform for AI agents. It records what your agents actually do in production, judges that behavior continuously, scores changes against test datasets before they ship, and turns everything it learns into concrete improvements. Tracing and scoring never touch your agent’s own code or prompts; the Improve loop proposes new prompt versions, and those ship only when you explicitly publish them.
Everything documented on this site runs against a self-hosted instance. Start at Self-Host Overview.
It is built for teams running AI agents in production (or about to): engineers who need to see what an agent did, reviewers who need to judge whether it was right, and release owners who need proof that a change made things better before it ships.
  • Framework-agnostic: LangChain/LangGraph, OpenAI Agents SDK, CrewAI, AutoGen, Google ADK, Google GenAI (Gemini), LlamaIndex, LiteLLM, direct OpenAI/Anthropic clients, Databricks/MLflow, Moveworks, plain Python, or any OpenTelemetry-instrumented app.
  • Online and offline evaluation: judges score live traffic as it arrives (traces, whole conversations, tool behavior), and dataset runs score candidate changes before release.
  • Trajectory-aware: evaluation sees the path the agent took - which tools, in what order, with what failures - not just the final answer.
  • Self-hosted: one local install or container, bring your own LLM keys. SQLite by default, Postgres for teams, Postgres + ClickHouse for enterprise telemetry volume. Traces stay in your own database; only the judge features send trace content to the LLM provider whose key you supply, and rule-based checks need no key at all.
circular diagram with four nodes - Trace (agent icon, arrows from LangChain/OpenAI/OTel logos), Monitor (judge/magnifier icon producing Signals), Evaluate (dataset + checkmark icon), Improve (wrench icon feeding back into the agent). Arrows flow clockwise: Trace to Monitor to Evaluate to Improve and back to the agent.

The governance loop: production traffic in, judged findings out, improvements validated before they ship.

Choose your path

Quickstart

Self-host the engine and send your first trace in five minutes

Trace your agent

Decorators, context managers, sessions, and per-framework integrations

Evaluate offline

Dataset runs on demand: LLM-as-judge scoring, similarity metrics, trajectory matching

Evaluate online

Continuous scoring of live traffic: patterns, LLM judge scorers, deduped signals

Gate releases in CI

Fail the pipeline when quality drops below your bar

Improve the agent

Prompt and tool registries with evidence-backed, validated proposals

How it fits together

Four capabilities share one SDK, one trace store, and one dataset model. Each works alone; together they form a loop. AgentX sits outside your agent in every case: tracing wraps your code with a decorator, context manager, or framework callback; evaluation calls your agent function and scores what comes back. Tracing and scoring inject no prompts and change none of your agent’s logic; the Improve loop proposes new prompt versions, but a proposal ships only when you explicitly publish it.

Deployment model

The self-hosted stack (AgentX-trace-eval) is a single engine binary plus the real AgentX dashboard, installed via curl | bash, the Python launcher (pip install agentx-python, then agentx-trace-eval --dev), or Docker. One binary, three storage tiers: local SQLite by default, your Postgres with one env var, or Postgres + ClickHouse for enterprise telemetry volume. You bring provider keys (OpenAI/Anthropic/Gemini) only for the LLM-judge features - trace ingest and rule-based detection run with no keys at all. Licensing: Elastic License 2.0 for the engine and CLI, Apache-2.0 for the Python SDK - the details live in Self-Host Overview.