eval run is the core command that ties the evaluation loop together: it takes an evaluation dataset and one or more evaluators, then submits an experiment against a deployed Runtime or TEA source evaluation target. The dataset, evaluators, and target can all be specified by ID or name, and field mapping is automatic by default.
Prepare a dataset, evaluators, and an available target on the same backend; submit evaluator versions first on TEA. The first example previews the request. --dry-run still resolves resources over the network, but does not submit an experiment or auto-publish a dataset draft
The CLI first resolves the evaluation backend for the current account. On TEA, every
--evaluator must have a corresponding --evaluator-version; on Coze, the command continues to use the evaluator’s current version. To inspect the current backend, run eval backend.eval run
Automatic Field Mapping
An experiment has three layers of data to align: dataset fields, target input and output, and evaluator inputs.eval run connects them automatically by convention:
- The target Runtime typically uses
user_inputas its input field andactual_outputas its output field; the dataset’s primary input field, such asinput, is passed to the target’suser_input. - The evaluator field that represents the model answer, such as
output, is supplied by the target outputactual_output; the remaining fields, such asinputandreference_output, are taken from the dataset by matching name.
--map Syntax
The TEA backend uses prefixed arrow syntax:
Check With Dry Run First
It is recommended to add--dry-run before your first run, to confirm the resolved dataset version, evaluator versions, target version, and field mapping:
An experiment requires a committed dataset version. When
--dataset-version is omitted and the dataset only has an uncommitted draft, eval run auto-publishes a version before submitting. When --dataset-version is provided, that version is used.eval experiment show: