Skip to main content
POST
Run an evaluation
The server must expose development evaluation endpoints and have evaluation dependencies installed. Prepare a set and cases, then select evalMetrics; query List evaluation metrics for available metrics Use evalCaseIds to select cases. When omitted or empty, it falls back to the deprecated evalIds; if neither selects cases, the entire set is evaluated
Evaluation runs the agent and selected metrics again. It can call models and tools, incur charges, and perform external actions. Check samples, tool permissions, and metric settings before running it
HTTP 200 means evaluation results were returned, not that all cases passed. Inspect finalEvalStatus and metric results for each item in runEvalResults

Authorizations

Authorization
string
header
required

Optional locally without a gateway; cloud deployments use the Runtime API key or user-pool JWT required by that deployment, never the model API key

Path Parameters

app_name
string
required

Application name; harness_agent for the CLI deployment

eval_set_id
string
required

Evaluation set ID

Body

application/json
evalMetrics
EvalMetric · object[]
required

Metrics to apply in this evaluation

evalIds
string[]
deprecated

Deprecated; use evalCaseIds

evalCaseIds
string[]

Case IDs to evaluate; omission or an empty array falls back to evalIds. Evaluate all cases when neither selects cases

Response

Successful response

runEvalResults
RunEvalResult · object[]
required

Per-case evaluation results

Last modified on September 19, 2026