review-llm-annotations-and-improve-prompt
Improve a metric from reviewed disagreements
Use this for development, not to declare a judge calibrated. When the goal is a
trust or release claim, use coval-calibrate-metric if installed. This workflow
remains usable by itself with the boundaries below.
Confirm organization/workspace and one text judge metric. Read its definition,
version, project and completed human annotations using current CLI context
and --help. Binary, categorical and numerical text judges require different
error analysis; don't silently coerce one type into another.
Collect actual completed human labels and reviewer notes, exact machine output
IDs/versions and the relevant transcripts. Paginate the public reviews API when
the CLI can't prove completeness. Zero is a valid label; null or pending is
missing. AI-suggested labels are not human ground truth. Multiple reviewers on
one call need adjudication, not duplicate counting. Annotation simulation_output_id can identify an uploaded
conversation; retain the source collection and retrieve original metrics through
its matching public API/CLI resource.