Skip to main content

Response shapes

Almost every error from the self-host engine is a single-field object:
Three exceptions:
  • Rate limiter responses carry the shape the dashboard’s interceptor reads: { "statusCode": 429, "message": "Too many requests" }.
  • Trace ingest validation (422) adds the field-level breakdown: { "error": "Invalid trace payload", "details": { "fieldErrors": ..., "formErrors": ... } }.
  • The global error handler - anything a route threw or rejected with, rather than answered itself - uses the same statusCode-carrying shape: { "statusCode": ..., "message": ... }. That covers unexpected 500s, body-parser rejections (400 unparseable JSON, 413 body too large, 415 bad charset), and the read-side 503 below.
The hosted API returns a structured envelope instead ({ "status": "error", "statusCode": ..., "message": ..., "errors": [] }). The Python SDK normalizes all of these - AgentXEvaluationsError and friends carry the message either way, so application code never needs to branch on the shape.

Status codes

The three flavors of 429

  1. Rate limiting ({ "statusCode": 429, "message": "Too many requests" }): per-IP, per-minute ceilings - 120/min on credential/auth routes, 6000/min on data-plane routes (tunable via AGENTX_RATE_LIMIT_CREDENTIAL / AGENTX_RATE_LIMIT_DATA_PLANE; AGENTX_RATE_LIMIT=off disables). Standard RateLimit headers (draft-7) are sent. Retry with backoff.
  2. Ingest queue full ({ "error": "Ingest queue is full - retry with backoff (Retry-After: 1s)." }): explicit backpressure on POST /ingest/traces; nothing was stored. Honor the Retry-After header (1s) and redeliver - span ids keep the redelivery idempotent.
  3. Daily trace quota ({ "error": "Daily trace quota reached (...)" }): AGENTX_QUOTA_TRACES_PER_DAY caps root-trace ingest per project; resets at midnight UTC, so retrying before then won’t help. (The judge-call quota, AGENTX_QUOTA_JUDGE_CALLS_PER_DAY, never returns a 429 - a capped judge call degrades like a judge outage: the affected result is stored as skipped with the quota message in its justification.)

Common cases

401 on every request (self-host) - the key is wrong or absent. Copy the Default project API key: agtx_local_... line from the engine’s startup log; the dashboard prompts for it on the connect screen, SDK/CI callers set AGENTX_API_KEY. The body is always { "error": "Invalid or missing API key" }. In AGENTX_AUTH=enabled mode, dashboard routes want a signed-in session instead - see Authentication. 409 from POST /api/v1/custom-agent-evaluations/runs/:id/results - the run is already in a terminal state (completed or failed); the body is { "error": "Run is already in a terminal state" }. Check GET /api/v1/custom-agent-evaluations/runs/:id before submitting more batches. Note that POST /api/v1/custom-agent-evaluations/runs/:id/finalize itself never 409s: it is idempotent on completed runs and returns status: "failed" (not an error) for failed ones. 400 "Batch size must not exceed 10" - result submission is capped at 10 per batch. The SDK’s execute() batches for you; hand-rolled callers should chunk. 503 from POST /ingest/traces - the telemetry store is down or the disk is full; the span was not stored. The response carries Retry-After: 2. Redeliver with the same spanId. null ratings instead of errors - a judge that cannot score (missing judge provider key, provider outage) does not fail the /results call: the result is stored with status: "skipped", rating: null, and the reason in justification, and it shows up in liveStatistics.skippedCount. Check skippedCount before trusting an average.

Retrying

  • Safe on 500, 503, and rate-limit/queue-full 429s, with exponential backoff (honor Retry-After when present). Quota 429s reset at midnight UTC - don’t spin on them.
  • Trace ingest is idempotent when a spanId is supplied - replaying the same span stores nothing new and triggers no duplicate judging. Importers (agentx-moveworks, agentx-databricks) rely on exactly this.
  • GET /api/v1/export/:entity can 400 on a bad since and 404 on an unknown entity, and a failure mid-stream destroys the socket instead of ending the response cleanly - treat a truncated .ndjson as a failed export and re-run it.
  • Result submission is idempotent per idempotencyKey (the SDK sends one automatically), and the engine treats a racing duplicate as a duplicate, never an error - so a retried batch never scores the same case twice. Still, avoid blind transport-level retries of /results: scoring is synchronous and slow, and a retry fired while the first attempt is still scoring just wastes a request.

The OTLP endpoint differs

POST /api/v1/otel/v1/traces speaks OTLP/HTTP conventions, not this page’s envelope: its shed/unavailable answers use { message } rather than { error } (429 queue-full with Retry-After: 1, 503 storage-down with Retry-After: 2 - both mean “redeliver the whole export”), and its daily-trace-quota 429 keeps the { error } envelope but carries Retry-After: 60. Schema-invalid spans are reported via OTLP partialSuccess and must not be retried.