Skip to main content
The eval experiment command group (alias exp) is used to view experiments submitted by eval run: list experiments, inspect an experiment’s status and aggregate scores, and fetch per-case results. On TEA, details include targets, evaluator versions, aggregate scores, and field mappings; on Coze, the commands continue to query by project.
--project applies only to the Coze evaluation backend; the TEA backend ignores it.

eval experiment list

List experiments in a project or workspace: ID, name, status, start time, creator, and aggregate score.

eval experiment show

Show an experiment’s details: status, dataset, target, evaluators, aggregate scores, and field mappings. get is an alias of show.

eval experiment results

Fetch per-case results: each dataset case, the target output, and each evaluator’s score and reasoning.

Experiment Status

For failed cases, inspect per-case errors, then check target readiness, authentication, field mappings, and the judge model. An aggregate result without an error message does not establish that the evaluation configuration is correct.
Obtain the experiment ID from eval run or experiment list. show does not wait continuously for completion; query again later while processing. results defaults to the first page, so retrieve subsequent pages for a complete analysis
Per-case results can include source cases, target outputs, and scoring reasons. Store and share them within the evaluation data’s access scope. Success indicates execution completed, not that scores satisfy business requirements
Last modified on September 19, 2026