> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# Run an experiment

`eval run` is the core command that ties the evaluation loop together: it takes an evaluation dataset and one or more evaluators, then submits an experiment against a deployed Runtime or TEA source evaluation target. The dataset, evaluators, and target can all be specified by ID or name, and field mapping is automatic by default.

```bash lines theme={null}
agentkit eval run \
  --dataset qa-set \
  --evaluator relevance \
  --evaluator-version 0.0.1 \
  --target my-agent
```

<Note>
  The CLI first resolves the evaluation backend for the current account. On TEA, every `--evaluator` must have a corresponding `--evaluator-version`; on Coze, the command continues to use the evaluator's current version. To inspect the current backend, run [`eval backend`](/productions/agentkit-cli/preview/en/commands/eval/backend).
</Note>

## eval run

| Flag / Argument | Description | Default |
| - | - | - |
| `--dataset <id\|name>` | Evaluation set ID or exact name. | Required |
| `--dataset-version <id\|name>` | TEA dataset version ID or version string; when omitted, the CLI uses an existing committed version and auto-publishes a draft version when needed. | — |
| `--evaluator <id\|name>` | Evaluator ID or exact name; repeatable. | Required |
| `--evaluator-version <id\|name>` | TEA evaluator version ID or version string; must be provided once for each `--evaluator`. | — |
| `--target <runtime name\|id>` | Deployed Runtime or TEA source evaluation target to evaluate. | Required |
| `--target-type <n>` | TEA source evaluation target type. | `101` |
| `--target-version <version>` | TEA source evaluation target version; when omitted, the CLI uses the latest version from the version list. | Latest version |
| `--name <name>` | Experiment name. | `<dataset>-<timestamp>` |
| `--description <text>` | Experiment description. | — |
| `--concurrency <n>` | Number of cases run concurrently. | `5` |
| `--map <spec>` | Override field mapping, repeatable; syntax below. | Automatic |
| `--dry-run` | Print the request that would be submitted without starting an experiment. | `false` |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit eval run \
  --dataset qa-set \
  --dataset-version 0.0.1 \
  --evaluator relevance \
  --evaluator-version 0.0.1 \
  --target my-agent \
  --target-version <target-version> \
  --concurrency 8
```

## Automatic Field Mapping

An experiment has three layers of data to align: dataset fields, target input and output, and evaluator inputs. `eval run` connects them automatically by convention:

* The target Runtime typically uses `user_input` as its input field and `actual_output` as its output field; the dataset's primary input field, such as `input`, is passed to the target's `user_input`.
* The evaluator field that represents the model answer, such as `output`, is supplied by the target output `actual_output`; the remaining fields, such as `input` and `reference_output`, are taken from the dataset by matching name.

Most scenarios need no manual mapping. When field names differ, use `--map` to override.

### `--map` Syntax

The TEA backend uses prefixed arrow syntax:

| Form | Meaning |
| - | - |
| `evaluator.<field> <- dataset.<field>` | Evaluator input comes from a dataset field. |
| `evaluator.<field> <- target.<field>` | Evaluator input comes from a target output field. |
| `target.<field> <- dataset.<field>` | Target input comes from a dataset field. |

```bash lines theme={null}
agentkit eval run --dataset qa-set --evaluator relevance --evaluator-version 0.0.1 --target my-agent \
  --map "evaluator.output <- target.actual_output" \
  --map "target.user_input <- dataset.question"
```

The Coze backend uses equals syntax:

| Form | Meaning |
| - | - |
| `<evaluatorField>=<datasetField>` | Evaluator input comes from a dataset field. |
| `<evaluatorField>=target:<targetOutput>` | Evaluator input comes from a target output field. |
| `target:<targetInput>=<datasetField>` | Target input comes from a dataset field. |

## Check With Dry Run First

It is recommended to add `--dry-run` before your first run, to confirm the resolved dataset version, evaluator versions, target version, and field mapping:

```bash lines theme={null}
agentkit eval run \
  --dataset qa-set \
  --evaluator relevance \
  --evaluator-version 0.0.1 \
  --target my-agent \
  --dry-run
```

<Note>
  An experiment requires a committed dataset version. When `--dataset-version` is omitted and the dataset only has an uncommitted draft, `eval run` auto-publishes a version before submitting. When `--dataset-version` is provided, that version is used.
</Note>

Once submitted successfully it returns an experiment ID; on TEA it also returns a run ID. Track it with [`eval experiment show`](/productions/agentkit-cli/preview/en/commands/eval/experiment#eval-experiment-show):

```bash lines theme={null}
agentkit eval run --dataset qa-set --evaluator relevance --evaluator-version 0.0.1 --target my-agent --json
# → { "experimentId": "75901...", "runId": "75902...", "name": "qa-set-<timestamp>" }
```
