Skip to main content
eval run is the core command that ties the evaluation loop together: it takes an evaluation dataset and one or more evaluators, and submits an experiment against a deployed runtime (the eval target). The dataset, evaluators, and target can all be specified by id or name, and field mapping is done automatically by default.

Flags and arguments

Automatic field mapping

An experiment has three layers of data to align: dataset fields → target input / output → evaluator inputs. eval run connects them automatically by convention:
  • The target runtime conventionally uses user_input as its input field and actual_output as its output field; the dataset’s primary input field (such as input) is passed to the target’s user_input.
  • The evaluator field that represents the “model answer” (such as output) is supplied by the target output actual_output; the remaining fields (such as input, reference_output) are taken from the dataset by matching name.
Most scenarios need no manual mapping. When field names differ, use --map to override.

--map syntax

Check with dry-run first

It is strongly recommended to add --dry-run before your first run, to confirm the resolved dataset version, target, and field mapping are correct:
An experiment requires the dataset to have a committed version. If the dataset only has an uncommitted draft, eval run will automatically publish a version before submitting — version management is transparent to you and requires no manual action.
Once submitted successfully it returns an experiment id, which you can track with eval experiment get:
Last modified on September 19, 2026