The eval dataset command group manages evaluation sets: list, show, create, and delete datasets, and add or remove cases within a dataset. A dataset is the data source for an experiment; its fields (such as input and reference_output) are mapped to the target and the evaluator during eval run.
These subcommands are also available as top-level aliases, so the eval prefix may be omitted: agentkit eval dataset list is equivalent to agentkit dataset list. The examples below use the shorter agentkit dataset ... form.
dataset list
List evaluation datasets.
dataset show
Show details of an evaluation dataset, including its case items.
dataset create
Create an evaluation dataset. The schema is fixed once created, and a case’s keys must match it.
output (the model’s actual output) is usually produced by the evaluated target runtime during the experiment, so it does not need to be part of the dataset schema. A dataset typically only needs input and reference_output.
dataset add
Add one or more cases to an evaluation dataset.
dataset remove
Remove one or more cases from an evaluation dataset.
This command removes the selected cases from the dataset. Check the dataset and case IDs first; the CLI prompts for confirmation unless you pass --yes.
dataset delete
Delete an evaluation dataset.
This command deletes the entire dataset. Check the target first; the CLI prompts for confirmation unless you pass --yes.