review-llm-annotations-and-improve-prompt

Installation
SKILL.md

Improve a metric from reviewed disagreements

Use this for development, not to declare a judge calibrated. When the goal is a trust or release claim, use coval-calibrate-metric if installed. This workflow remains usable by itself with the boundaries below.

Confirm organization/workspace and one text judge metric. Read its definition, version, project and completed human annotations using current CLI context and --help. Binary, categorical and numerical text judges require different error analysis; don't silently coerce one type into another.

Collect actual completed human labels and reviewer notes, exact machine output IDs/versions and the relevant transcripts. Paginate the public reviews API when the CLI can't prove completeness. Zero is a valid label; null or pending is missing. AI-suggested labels are not human ground truth. Multiple reviewers on one call need adjudication, not duplicate counting. Annotation simulation_output_id can identify an uploaded conversation; retain the source collection and retrieve original metrics through its matching public API/CLI resource.

Installs
52
GitHub Stars
2
First Seen
Apr 9, 2026
review-llm-annotations-and-improve-prompt — coval-ai/coval-external-skills