harbor
Harbor
Agent evaluation framework from the creators of Terminal-Bench.
Official Documentation
- Docs: https://harborframework.com/docs
- Getting Started: https://harborframework.com/docs/getting-started
- GitHub: https://github.com/laude-institute/harbor
Local Workspace & API Keys
.local-workspace/- Git-ignored directory for cloning PRs, temporary files, external repos, etc..local-workspace/.env- May containANTHROPIC_API_KEYand other API credentials. Check and use when running harbor with API access.
Quick Reference
# Install
uv tool install harbor
# Validate task
harbor tasks check tasks/<task-id>
# Run oracle (must pass 100%)
harbor run -p tasks/<task-id> -a oracle
# Run with agent (specify model with -m)
harbor run -p tasks/<task-id> -a claude-code -m 'anthropic/claude-opus-4-5'
# List datasets
harbor datasets list
# Cloud execution (parallel)
harbor run -d "<dataset@version>" -a "<agent>" -m "<model>" --env "daytona" -n 32
SkillsBench Task Structure
tasks/<task-id>/
task.toml # Metadata
instruction.md # Agent instructions
environment/
Dockerfile # Container + COPY skills to all agent locations
skills/ # Skills for agents
tests/
test.sh # Runs pytest, writes reward.txt
test_outputs.py # Test cases
solution/
solve.sh # Oracle solution (human-written)
Results Location
jobs/<timestamp>/<task-id>/:
trial.log- Execution logverifier/reward.txt- 0 (fail) or 1 (pass)verifier/ctrf.json- Test details
For task format details, see references/task-format.md
Agent Skill Support
Skills are copied to agent-specific locations in task Dockerfiles. Place skills in environment/skills/ and they'll be copied to:
Supported by Harbor (benchmarkable)
| Agent | Skills Directory | Docs |
|---|---|---|
| Claude Code | .claude/skills/ |
docs |
| Codex (OpenAI) | .codex/skills/ |
docs |
| OpenCode | .opencode/skill/ or .claude/skills/ |
docs |
| Goose | .goose/skills/ or .claude/skills/ |
docs |
| Factory | .factory/skills/ |
docs |
| Portable format | .agents/skills/ |
Used by Goose, Amp |
| GitHub Copilot | .github/skills/ |
docs |
Not yet supported by Harbor
| Agent | Skills Directory | Docs |
|---|---|---|
| Amp | .agents/skills/ or .claude/skills/ |
docs |
| Letta | .skills/ |
docs |
Adding Skills to Tasks
# Copy skills to ALL agent paths in Dockerfile
COPY skills /root/.claude/skills
COPY skills /root/.codex/skills
COPY skills /root/.opencode/skill
COPY skills /root/.goose/skills
COPY skills /root/.factory/skills
COPY skills /root/.agents/skills
COPY skills /root/.github/skills
More from neversight/skills_feed
ai-image-generation
|
7react-best-practices
Provides React patterns for hooks, effects, refs, and component design. Covers escape hatches, anti-patterns, and correct effect usage. Must use when reading or writing React components (.tsx, .jsx files with React imports).
7ui-designer
Use when user needs visual UI design, interface creation, component systems, design systems, interaction patterns, or accessibility-focused user interfaces.
7python-env
Fast Python environment management with uv (10-100x faster than pip). Triggers on: uv, venv, pip, pyproject, python environment, install package, dependencies.
7typescript-best-practices
Provides TypeScript patterns for type-first development, making illegal states unrepresentable, exhaustive handling, and runtime validation. Must use when reading or writing TypeScript/JavaScript files.
6ai-marketing-videos
|
6