The eval experiment command group (alias exp) is used to view experiments submitted by eval run: list experiments, inspect an experiment’s status and aggregate scores, and fetch per-case results. On TEA, details include targets, evaluator versions, aggregate scores, and field mappings; on Coze, the commands continue to query by project.
--project applies only to the Coze evaluation backend; the TEA backend ignores it.
eval experiment list
List experiments in a project or workspace: ID, name, status, start time, creator, and aggregate score.
eval experiment show
Show an experiment’s details: status, dataset, target, evaluators, aggregate scores, and field mappings. get is an alias of show.
eval experiment results
Fetch per-case results: each dataset case, the target output, and each evaluator’s score and reasoning.
Experiment Status
When a case fails but there is no experiment-level error message, it is usually because the target Runtime did not respond correctly (not deployed, not ready, or authentication failed), rather than an evaluation configuration issue. First confirm that the Runtime pointed to by --target is running.