curl -X POST http://localhost:4700/api/v1/custom-agent-evaluations/runs/rK7dP2qWx9TzB4mV6nJcE/finalize \
-H "x-api-key: agtx_local_0f3c9a17d2b84e6a5c01b9f4e7d8a2c6431b5f97a0e2d4c8"
result = client.evaluations.finalize_run("rK7dP2qWx9TzB4mV6nJcE")
print(result["liveStatistics"]["averageRating"])
{
"runId": "rK7dP2qWx9TzB4mV6nJcE",
"status": "completed",
"liveStatistics": {
"averageRating": 7.8,
"minRating": 5,
"maxRating": 10,
"ratedCount": 6,
"skippedCount": 0,
"failedCount": 0
}
}
{
"error": "Run not found"
}
Custom Evaluations
Finalize Run
Mark a run as complete and get the final rating statistics
POST
/
api
/
v1
/
custom-agent-evaluations
/
runs
/
{runId}
/
finalize
curl -X POST http://localhost:4700/api/v1/custom-agent-evaluations/runs/rK7dP2qWx9TzB4mV6nJcE/finalize \
-H "x-api-key: agtx_local_0f3c9a17d2b84e6a5c01b9f4e7d8a2c6431b5f97a0e2d4c8"
result = client.evaluations.finalize_run("rK7dP2qWx9TzB4mV6nJcE")
print(result["liveStatistics"]["averageRating"])
{
"runId": "rK7dP2qWx9TzB4mV6nJcE",
"status": "completed",
"liveStatistics": {
"averageRating": 7.8,
"minRating": 5,
"maxRating": 10,
"ratedCount": 6,
"skippedCount": 0,
"failedCount": 0
}
}
{
"error": "Run not found"
}
Marks the run
"completed" and returns the authoritative rating aggregate, recomputed from
every stored result. After finalizing, no more results are accepted (further /results calls
return 409), and the run is ready for
Analyze Run or the
CI gate.
Finalizing is idempotent: calling it again on a completed run returns the same
"completed" response. Finalizing a "failed" run returns status: "failed" (with its
statistics) rather than flipping it to completed - and rather than an error, so retried
finalize calls never crash a pipeline. You can finalize with partial results; use
Get Missing Results first to check coverage.Authentication
string
required
Project API key.
Path Parameters
string
required
Run ID.
Body
No body required.Response
string
Run ID.
string
"completed", or "failed" if the run had previously failed.object
Final rating aggregate, recomputed from every stored result:
{ averageRating, minRating, maxRating, ratedCount, skippedCount, failedCount }. Same shape
as Submit Results’ liveStatistics. Available
without calling Analyze Run; analysis only adds the
LLM-written qualitative report on top.Errors
| Status | Body | Meaning |
|---|---|---|
404 | { "error": "Run not found" } | No run with this id in the project |
curl -X POST http://localhost:4700/api/v1/custom-agent-evaluations/runs/rK7dP2qWx9TzB4mV6nJcE/finalize \
-H "x-api-key: agtx_local_0f3c9a17d2b84e6a5c01b9f4e7d8a2c6431b5f97a0e2d4c8"
result = client.evaluations.finalize_run("rK7dP2qWx9TzB4mV6nJcE")
print(result["liveStatistics"]["averageRating"])
{
"runId": "rK7dP2qWx9TzB4mV6nJcE",
"status": "completed",
"liveStatistics": {
"averageRating": 7.8,
"minRating": 5,
"maxRating": 10,
"ratedCount": 6,
"skippedCount": 0,
"failedCount": 0
}
}
{
"error": "Run not found"
}

