agent-evaluation

Installation
SKILL.md

Agent Evaluation

Evaluate observable agent behavior against task-specific cases. Modified by AAS maintainers on 2026-09-05 to remove unsupported benchmark claims, correct uncertainty/error reporting and separate optional architecture sketches from the operating procedure.

When to Use

Use when comparing a changed agent, prompt or tool configuration, reproducing an observed failure, or estimating reliability on a declared task distribution. Do not infer product readiness from a public benchmark percentage or a generic score threshold.

Prerequisites

  • A versioned case set with expected observable outcomes and permission boundaries.
  • A known baseline and candidate revision, including model, prompt, tools, configuration and runtime versions.
  • Authorized synthetic or redacted inputs, isolated targets and a bounded token, time and cost budget.
  • A verifier that distinguishes wrong outcomes, expected safe rejections, evaluator failures and infrastructure outages. Provider access is needed only if the declared evaluation calls that provider.

Evaluation procedure

Installs
1.0K
GitHub Stars
46.5K
First Seen
Jan 19, 2026
agent-evaluation — sickn33/agentic-awesome-skills