Skip to main content
eval run is the core command that ties the evaluation loop together: it takes an evaluation dataset and one or more evaluators, then submits an experiment against a deployed Runtime or TEA source evaluation target. The dataset, evaluators, and target can all be specified by ID or name, and field mapping is automatic by default. Prepare a dataset, evaluators, and an available target on the same backend; submit evaluator versions first on TEA. The first example previews the request. --dry-run still resolves resources over the network, but does not submit an experiment or auto-publish a dataset draft
Removing --dry-run invokes the target and judge models, may incur usage charges, and sends cases and outputs to the evaluation platform. Without a committed dataset version, submission can also publish the current draft automatically
The CLI first resolves the evaluation backend for the current account. On TEA, every --evaluator must have a corresponding --evaluator-version; on Coze, the command continues to use the evaluator’s current version. To inspect the current backend, run eval backend.

eval run

Automatic Field Mapping

An experiment has three layers of data to align: dataset fields, target input and output, and evaluator inputs. eval run connects them automatically by convention:
  • The target Runtime typically uses user_input as its input field and actual_output as its output field; the dataset’s primary input field, such as input, is passed to the target’s user_input.
  • The evaluator field that represents the model answer, such as output, is supplied by the target output actual_output; the remaining fields, such as input and reference_output, are taken from the dataset by matching name.

--map Syntax

The TEA backend uses prefixed arrow syntax:
The Coze backend uses equals syntax:

Check With Dry Run First

It is recommended to add --dry-run before your first run, to confirm the resolved dataset version, evaluator versions, target version, and field mapping:
An experiment requires a committed dataset version. When --dataset-version is omitted and the dataset only has an uncommitted draft, eval run auto-publishes a version before submitting. When --dataset-version is provided, that version is used.
Once submitted successfully it returns an experiment ID; on TEA it also returns a run ID. Track it with eval experiment show:
Last modified on September 19, 2026