Response shapes
Almost every error from the self-host engine is a single-field object:- Rate limiter responses carry the shape the dashboard’s interceptor reads:
{ "statusCode": 429, "message": "Too many requests" }. - Trace ingest validation (
422) adds the field-level breakdown:{ "error": "Invalid trace payload", "details": { "fieldErrors": ..., "formErrors": ... } }. - The global error handler - anything a route threw or rejected with, rather than answered
itself - uses the same statusCode-carrying shape:
{ "statusCode": ..., "message": ... }. That covers unexpected500s, body-parser rejections (400unparseable JSON,413body too large,415bad charset), and the read-side503below.
{ "status": "error", "statusCode": ..., "message": ..., "errors": [] }). The Python SDK
normalizes all of these - AgentXEvaluationsError and friends carry the message either way,
so application code never needs to branch on the shape.
Status codes
The three flavors of 429
- Rate limiting (
{ "statusCode": 429, "message": "Too many requests" }): per-IP, per-minute ceilings - 120/min on credential/auth routes, 6000/min on data-plane routes (tunable viaAGENTX_RATE_LIMIT_CREDENTIAL/AGENTX_RATE_LIMIT_DATA_PLANE;AGENTX_RATE_LIMIT=offdisables). StandardRateLimitheaders (draft-7) are sent. Retry with backoff. - Ingest queue full (
{ "error": "Ingest queue is full - retry with backoff (Retry-After: 1s)." }): explicit backpressure onPOST /ingest/traces; nothing was stored. Honor theRetry-Afterheader (1s) and redeliver - span ids keep the redelivery idempotent. - Daily trace quota (
{ "error": "Daily trace quota reached (...)" }):AGENTX_QUOTA_TRACES_PER_DAYcaps root-trace ingest per project; resets at midnight UTC, so retrying before then won’t help. (The judge-call quota,AGENTX_QUOTA_JUDGE_CALLS_PER_DAY, never returns a 429 - a capped judge call degrades like a judge outage: the affected result is stored asskippedwith the quota message in itsjustification.)
Common cases
401 on every request (self-host) - the key is wrong or absent. Copy the
Default project API key: agtx_local_... line from the engine’s startup log; the dashboard
prompts for it on the connect screen, SDK/CI callers set AGENTX_API_KEY. The body is always
{ "error": "Invalid or missing API key" }. In AGENTX_AUTH=enabled mode, dashboard routes
want a signed-in session instead - see Authentication.
409 from POST /api/v1/custom-agent-evaluations/runs/:id/results - the run is already
in a terminal state (completed or failed); the body is
{ "error": "Run is already in a terminal state" }.
Check GET /api/v1/custom-agent-evaluations/runs/:id before submitting more batches. Note
that POST /api/v1/custom-agent-evaluations/runs/:id/finalize itself never 409s: it is idempotent on completed runs and returns status: "failed" (not an
error) for failed ones.
400 "Batch size must not exceed 10" - result submission is capped at 10 per batch. The
SDK’s execute() batches for you; hand-rolled callers should chunk.
503 from POST /ingest/traces - the telemetry store is down or the disk is full; the
span was not stored. The response carries Retry-After: 2. Redeliver with the same spanId.
null ratings instead of errors - a judge that cannot score (missing judge provider key,
provider outage) does not fail the /results call: the result is stored with
status: "skipped", rating: null, and the reason in justification, and it shows up in
liveStatistics.skippedCount. Check skippedCount before trusting an average.
Retrying
- Safe on
500,503, and rate-limit/queue-full429s, with exponential backoff (honorRetry-Afterwhen present). Quota429s reset at midnight UTC - don’t spin on them. - Trace ingest is idempotent when a
spanIdis supplied - replaying the same span stores nothing new and triggers no duplicate judging. Importers (agentx-moveworks,agentx-databricks) rely on exactly this. GET /api/v1/export/:entitycan400on a badsinceand404on an unknown entity, and a failure mid-stream destroys the socket instead of ending the response cleanly - treat a truncated.ndjsonas a failed export and re-run it.- Result submission is idempotent per
idempotencyKey(the SDK sends one automatically), and the engine treats a racing duplicate as a duplicate, never an error - so a retried batch never scores the same case twice. Still, avoid blind transport-level retries of/results: scoring is synchronous and slow, and a retry fired while the first attempt is still scoring just wastes a request.
The OTLP endpoint differs
POST /api/v1/otel/v1/traces speaks OTLP/HTTP conventions, not this page’s envelope: its
shed/unavailable answers use { message } rather than { error } (429 queue-full with
Retry-After: 1, 503 storage-down with Retry-After: 2 - both mean “redeliver the whole
export”), and its daily-trace-quota 429 keeps the { error } envelope but carries Retry-After: 60. Schema-invalid spans are
reported via OTLP partialSuccess and must not be retried.
