eval dataset command group manages evaluation sets: list, show, create, update, and delete datasets, add or remove cases, and manage TEA dataset versions. A dataset is the data source for an experiment; its fields, such as input and reference_output, are mapped to the evaluation target and evaluator during eval run.
These subcommands are also available as top-level aliases, so the
eval prefix may be omitted: agentkit eval dataset list is equivalent to agentkit dataset list. The examples below use the shorter agentkit dataset ... form. --project applies only to the Coze evaluation backend; the TEA backend ignores it.agentkit eval backend first to confirm the backend and workspace. Adding cases writes to a remote dataset; exclude user data that is not approved for evaluation
Group options
Place these options betweendataset and its subcommand. They also apply to eval dataset:
dataset list
List evaluation datasets.dataset show
Show details of an evaluation dataset, including schema, version information, and case items.dataset create
Create an evaluation dataset. The schema is fixed once created, and a case’s keys must match it.dataset update
Update the name or description of a TEA evaluation dataset.dataset update is supported only on the TEA evaluation backend.dataset add
Add one or more cases to an evaluation dataset. The Coze backend accepts flat{field: value} objects; the TEA backend accepts full turn-shaped items, and can also build a single-turn text item from --field.
items.json uses the TEA item format. Coze files use flat objects, such as [{"input":"Question","reference_output":"Answer"}]
items.json
dataset remove
Remove one or more cases from an evaluation dataset.dataset delete
Delete an evaluation dataset.dataset version list
List TEA dataset versions.dataset version subcommands are supported only on the TEA evaluation backend.dataset version create
Create a committed version for a TEA evaluation dataset.agentkit dataset show qa-set --items 20, then create a version on TEA. Editing cases does not change the version used by an existing experiment. Save and explicitly pass --dataset-version for reproducible evaluation