Skip to main content
The eval experiment command group (alias exp) is used to view experiments submitted by eval run: list experiments, view an experiment’s status and aggregate scores, and fetch per-case results.
All subcommands share -p, --project <name> (project name, default default) and --json (output raw JSON).

experiment list

List the experiments in a project: id, name, status, start time, and aggregate score.

experiment get

Show details of an experiment: status, dataset, target, weighted overall score, and each evaluator’s average score and score distribution.

experiment results

Fetch per-case results: each dataset case, the target output, and each evaluator’s score and reasoning.

Experiment status

When a case fails but there is no experiment-level error message, it is usually because the target runtime did not respond correctly (not deployed, not ready, or authentication failed), rather than an evaluation configuration issue — first confirm that the runtime pointed to by --target is running.
Last modified on September 19, 2026