Skip to main content
The eval dataset command group manages evaluation sets: list, show, create, and delete datasets, and add or remove cases within a dataset. A dataset is the data source for an experiment; its fields (such as input and reference_output) are mapped to the target and the evaluator during eval run.
These subcommands are also available as top-level aliases, so the eval prefix may be omitted: agentkit eval dataset list is equivalent to agentkit dataset list. The examples below use the shorter agentkit dataset ... form.

dataset list

List evaluation datasets.

dataset show

Show details of an evaluation dataset, including its case items.

dataset create

Create an evaluation dataset. The schema is fixed once created, and a case’s keys must match it.
output (the model’s actual output) is usually produced by the evaluated target runtime during the experiment, so it does not need to be part of the dataset schema. A dataset typically only needs input and reference_output.

dataset add

Add one or more cases to an evaluation dataset.

dataset remove

Remove one or more cases from an evaluation dataset.
This command removes the selected cases from the dataset. Check the dataset and case IDs first; the CLI prompts for confirmation unless you pass --yes.

dataset delete

Delete an evaluation dataset.
This command deletes the entire dataset. Check the target first; the CLI prompts for confirmation unless you pass --yes.
Last modified on September 19, 2026