Everything documented on this site runs against a self-hosted instance. Start at
Self-Host Overview.
- Framework-agnostic: LangChain/LangGraph, OpenAI Agents SDK, CrewAI, AutoGen, Google ADK, Google GenAI (Gemini), LlamaIndex, LiteLLM, direct OpenAI/Anthropic clients, Databricks/MLflow, Moveworks, plain Python, or any OpenTelemetry-instrumented app.
- Online and offline evaluation: judges score live traffic as it arrives (traces, whole conversations, tool behavior), and dataset runs score candidate changes before release.
- Trajectory-aware: evaluation sees the path the agent took - which tools, in what order, with what failures - not just the final answer.
- Self-hosted: one local install or container, bring your own LLM keys. SQLite by default, Postgres for teams, Postgres + ClickHouse for enterprise telemetry volume. Traces stay in your own database; only the judge features send trace content to the LLM provider whose key you supply, and rule-based checks need no key at all.

The governance loop: production traffic in, judged findings out, improvements validated before they ship.
Choose your path
Quickstart
Self-host the engine and send your first trace in five minutes
Trace your agent
Decorators, context managers, sessions, and per-framework integrations
Evaluate offline
Dataset runs on demand: LLM-as-judge scoring, similarity metrics, trajectory matching
Evaluate online
Continuous scoring of live traffic: patterns, LLM judge scorers, deduped signals
Gate releases in CI
Fail the pipeline when quality drops below your bar
Improve the agent
Prompt and tool registries with evidence-backed, validated proposals
How it fits together
Four capabilities share one SDK, one trace store, and one dataset model. Each works alone; together they form a loop.
AgentX sits outside your agent in every case: tracing wraps your code with a decorator,
context manager, or framework callback; evaluation calls your agent function and scores what
comes back. Tracing and scoring inject no prompts and change none of your agent’s logic; the
Improve loop proposes new prompt versions, but a proposal ships only when you explicitly
publish it.
Deployment model
The self-hosted stack (AgentX-trace-eval) is a single engine binary plus the real AgentX dashboard, installed viacurl | bash, the Python
launcher (pip install agentx-python, then agentx-trace-eval --dev), or Docker. One binary,
three storage tiers: local SQLite by default, your Postgres with one env var, or Postgres +
ClickHouse for enterprise telemetry volume. You bring provider keys (OpenAI/Anthropic/Gemini)
only for the LLM-judge features - trace ingest and rule-based detection run with no keys at all.
Licensing: Elastic License 2.0 for the engine and CLI, Apache-2.0 for the Python SDK - the
details live in Self-Host Overview.
