The eval experiment command group (alias exp) is used to view experiments submitted by eval run: list experiments, view an experiment’s status and aggregate scores, and fetch per-case results.
All subcommands share -p, --project <name> (project name, default default) and --json (output raw JSON).
experiment list
List the experiments in a project: id, name, status, start time, and aggregate score.
experiment get
Show details of an experiment: status, dataset, target, weighted overall score, and each evaluator’s average score and score distribution.
experiment results
Fetch per-case results: each dataset case, the target output, and each evaluator’s score and reasoning.
Experiment status
When a case fails but there is no experiment-level error message, it is usually because the target runtime did not respond correctly (not deployed, not ready, or authentication failed), rather than an evaluation configuration issue — first confirm that the runtime pointed to by --target is running.