Skip to main content

AgentX

The top-level client. Create one per process.

Constructor

Class methods

Attributes

ping()

Verifies the client can actually reach AgentX and that the API key is accepted. The constructor is deliberately lazy (no network call) and trace delivery is fire-and-forget, so a wrong base_url or api_key otherwise surfaces only as a one-time log warning while traces silently go nowhere - call this once at startup of a long-running service to fail fast. Raises AgentXConnectionError when the URL is unreachable, AgentXAuthError when the key is rejected, and AgentXAPIError on any other non-OK response. Returns {"ok": True, "base_url": ...} on success.

close()

Flushes queued traces and stops the tracer’s background ingest worker; returns True when everything drained before the deadline. Optional - an atexit hook already flushes queued traces on interpreter shutdown - but a long-running service that tears clients down mid-process should call it, or use the client as a context manager (with AgentX(...) as client:), so worker threads don’t accumulate.

Tracer

Accessed via client.tracer. Handles both tracing and CI/CD evaluation.

trace()

Returns a _TraceSpan that works as a decorator or context manager.

flush()

Blocks until all queued traces have been delivered, or until timeout seconds elapse. Returns True when everything drained, False on deadline (a warning is logged and undelivered traces keep sending in the background). Call before process exit in scripts or one-shot jobs.

use_span()

Makes span (created on another thread) the active span for the duration of the block, on the calling thread. The active-span stack is context-local (a ContextVar): bare threads start with an empty stack, while asyncio tasks inherit a copy of their creator’s. A span opened with tracer.trace(...) on the main thread therefore isn’t automatically visible inside a ThreadPoolExecutor worker (or any other thread). Wrap the worker’s body in use_span() to attach its LLM and tool calls to the parent span instead of starting an independent trace.
Safe to call concurrently from multiple threads for the same span: each thread pushes and pops on its own stack.

current_span

Property: the innermost with tracer.trace(...) span active in the current context, if any. The manual recorders below attach their child spans to it.

record_tool_call()

Manually records a tool call that an auto-instrumented framework integration can’t see - e.g. a hand-rolled tool-use loop where the tool executes in plain Python between two model calls. Sent as a tool child span of the active span (see current_span), plus a summary on the root trace’s flat tool_calls list - the surface the engine’s built-in “Tool failure” Monitor check and the dashboard’s Tool quality column read. With no active span, the record is queued onto the next trace this tracer sends (tracer-wide, so wrap in tracer.trace() when concurrent no-span use matters). success=False marks a failed call; leaving success unset means “unknown” and the dashboard falls back to its output-text heuristic.

trace_tool_call()

Context manager that times the block and records it via record_tool_call() on exit. An exception escaping the block records the call as failed (success=False plus the error text) and then propagates unchanged.

record_retrieval()

Manually records a knowledge-base / vector-store retrieval as a retrieval child span of the active span - the spans the engine reads retrieved chunks from for trajectory-aware judging and the RAG metrics.

trace_retrieval()

Context manager that times a retrieval and records it via record_retrieval() on exit. An escaping exception records the retrieval as failed (error set) instead of as a clean empty retrieval, then propagates unchanged.

record_memory()

Manually records a long-term-memory operation (a Mem0/Zep/Letta-style recall or store) as a memory child span of the active span. operation is free text - conventionally "read" or "write" - carried in the span’s metadata. record_memory/trace_memory drop when no span is active (warned once per process, then logged at debug), unlike record_tool_call/record_retrieval, which queue onto the next trace.

trace_memory()

Context manager that times a memory operation and records it via record_memory() on exit. An escaping exception records the operation as failed, then propagates unchanged.

run_eval()

High-level CI/CD evaluation: creates a run, calls agent_fn for every test case, submits results, finalizes, and returns the gate decision. Raises CIGateFailure (containing the CIRunResult) when fail_on_gate=True and the gate is "fail".

create_ci_run()

submit_result()

finalize_ci_run()

get_ci_run()

evaluate_trace()

Score a previously-recorded trace against a dataset. The agent is not re-run. Returns { run_id, trace_id, rating, justification, status }. trace_id comes from a prior tracer.trace(..., sync=True) call’s span.trace_id.

_TraceSpan

Returned by tracer.trace(). Can be used as a decorator or context manager.

Attributes

Methods

add_tool_call()

Records a tool call made during this span. Safe to call multiple times.

set_error()

Marks this span as failed. Takes precedence over any exception automatically caught by __exit__.

Usage patterns


CI/CD Dataclasses

CIRun

Returned by create_ci_run().

CITestCase

CIQuestionScore

Returned by submit_result().

CIRunResult

Returned by finalize_ci_run() and run_eval().

CIRunStatus

Returned by get_ci_run().

ThresholdViolation


Exceptions

All exceptions inherit from AgentXError.

CIGateFailure


Environment variables