> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# Manage datasets

The `eval dataset` command group manages evaluation sets: list, show, create, update, and delete datasets, add or remove cases, and manage TEA dataset versions. A dataset is the data source for an experiment; its fields, such as `input` and `reference_output`, are mapped to the evaluation target and evaluator during [`eval run`](/productions/agentkit-cli/preview/en/commands/eval/run).

<Note>
  These subcommands are also available as top-level aliases, so the `eval` prefix may be omitted: `agentkit eval dataset list` is equivalent to `agentkit dataset list`. The examples below use the shorter `agentkit dataset ...` form. `--project` applies only to the Coze evaluation backend; the TEA backend ignores it.
</Note>

Run `agentkit eval backend` first to confirm the backend and workspace. Adding cases writes to a remote dataset; exclude user data that is not approved for evaluation

### Group options

Place these options between `dataset` and its subcommand. They also apply to `eval dataset`:

| Flag / argument | Description | Default |
| - | - | - |
| `--cache-refresh` | Bypass local cache and refresh evaluation backend data | `false` |
| `--verbose` | Print full TEA requests to stderr with the session cookie masked; business content may remain in logs | `false` |

```bash lines theme={null}
agentkit dataset --cache-refresh list
```

## dataset list

List evaluation datasets.

| Flag / Argument | Description | Default |
| - | - | - |
| `-r, --region <region>` | Volcengine region. | Auto-detect |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset list
```

## dataset show

Show details of an evaluation dataset, including schema, version information, and case items.

| Flag / Argument | Description | Default |
| - | - | - |
| `<id\|name>` | Evaluation set ID or exact name. | Required |
| `--dataset-version <id\|name>` | TEA dataset version ID or version string, such as `0.0.1`. | — |
| `--items <n>` | List up to N case items; pass `0` to skip the item list. | `20` |
| `-r, --region <region>` | Volcengine region. | Auto-detect |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset show qa-set --items 50

agentkit dataset show qa-set --dataset-version 0.0.1 --json
```

## dataset create

Create an evaluation dataset. The schema is fixed once created, and a case's keys must match it.

| Flag / Argument | Description | Default |
| - | - | - |
| `--name <name>` | Dataset name. | Required |
| `--schema <fields>` | Comma-separated field names. | `input,reference_output,output` |
| `--description <text>` | Dataset description. | — |
| `-r, --region <region>` | Volcengine region. | Environment and current cloud context |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset create --name qa-set --schema "input,reference_output"
```

<Tip>
  `output` is usually produced by the evaluated target Runtime during the experiment, so it does not need to be part of the dataset schema. A dataset typically only needs `input` and `reference_output`.
</Tip>

## dataset update

Update the name or description of a TEA evaluation dataset.

<Note>
  `dataset update` is supported only on the TEA evaluation backend.
</Note>

| Flag / Argument | Description | Default |
| - | - | - |
| `<id>` | Evaluation set ID. | Required |
| `--name <name>` | New dataset name. | — |
| `--description <text>` | New dataset description. | — |
| `-r, --region <region>` | Volcengine region. | Environment |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset update <dataset-id> --name qa-set-v2 --description "Customer-support QA set"
```

## dataset add

Add one or more cases to an evaluation dataset. The Coze backend accepts flat `{field: value}` objects; the TEA backend accepts full turn-shaped items, and can also build a single-turn text item from `--field`.

| Flag / Argument | Description | Default |
| - | - | - |
| `<id\|name>` | Evaluation set ID or exact name. | Required |
| `--field <key=value>` | Case field, repeatable; on TEA this becomes a single-turn text item. | — |
| `--file <path>` | JSON file; Coze uses a flat object or object array, TEA uses full item objects or arrays. | — |
| `--items-json <json>` | Raw TEA item object array; takes precedence over `--file`. | — |
| `-r, --region <region>` | Volcengine region. | Auto-detect |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset add qa-set --field "input=What is the capital of France?" --field "reference_output=Paris"
```

The following `items.json` uses the TEA item format. Coze files use flat objects, such as `[{"input":"Question","reference_output":"Answer"}]`

```json title="items.json" lines theme={null}
[
  {
    "turns": [
      {
        "field_data_list": [
          {
            "key": "input",
            "name": "input",
            "content": {
              "content_type": "Text",
              "format": 1,
              "text": "What is the capital of France?"
            }
          }
        ]
      }
    ]
  }
]
```

```bash lines theme={null}
agentkit dataset add qa-set --file ./items.json
```

## dataset remove

Remove one or more cases from an evaluation dataset.

<Warning>
  This command removes the selected cases from the dataset. Check the dataset and case IDs first; the CLI prompts for confirmation unless you pass `--yes`.
</Warning>

| Flag / Argument | Description | Default |
| - | - | - |
| `<id\|name>` | Evaluation set ID or exact name. | Required |
| `<itemIds...>` | Case item IDs to remove; accepts one or more. | Required |
| `-r, --region <region>` | Volcengine region. | Auto-detect |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `-y, --yes` | Skip the confirmation prompt. | `false` |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset remove qa-set item-1 item-2 -y
```

## dataset delete

Delete an evaluation dataset.

<Warning>
  This command deletes the entire dataset. Check the target first; the CLI prompts for confirmation unless you pass `--yes`.
</Warning>

| Flag / Argument | Description | Default |
| - | - | - |
| `<id\|name>` | Evaluation set ID or exact name. | Required |
| `-r, --region <region>` | Volcengine region. | Auto-detect |
| `-p, --project <name>` | Coze project name; ignored on TEA. | `default` |
| `-y, --yes` | Skip the confirmation prompt. | `false` |

```bash lines theme={null}
agentkit dataset delete qa-set -y
```

## dataset version list

List TEA dataset versions.

<Note>
  `dataset version` subcommands are supported only on the TEA evaluation backend.
</Note>

| Flag / Argument | Description | Default |
| - | - | - |
| `--dataset <id\|name>` | Evaluation set ID or exact name. | Required |
| `-r, --region <region>` | Volcengine region. | Environment |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset version list --dataset qa-set
```

## dataset version create

Create a committed version for a TEA evaluation dataset.

| Flag / Argument | Description | Default |
| - | - | - |
| `<version>` | Version string, such as `0.0.2`. | Required |
| `--dataset <id\|name>` | Evaluation set ID or exact name. | Required |
| `--description <text>` | Version description. | — |
| `--desc <text>` | Alias for `--description`. | — |
| `-r, --region <region>` | Volcengine region. | Environment |
| `--json` | Output raw JSON. | `false` |

```bash lines theme={null}
agentkit dataset version create 0.0.2 --dataset qa-set --description "Add edge cases"
```

After adding cases, inspect fields and items with `agentkit dataset show qa-set --items 20`, then create a version on TEA. Editing cases does not change the version used by an existing experiment. Save and explicitly pass `--dataset-version` for reproducible evaluation
