langsmith-evaluator
Audited by ZeroLeaks on Apr 15, 2026
The skill's SKILL.md is mostly readable, but carries one transparency concern—remote script execution without a clear review boundary—which weakens pre-use reviewability and drives the AT_RISK verdict. Instruction/data separation stays reasonable; the scan did not surface strong prompt-injection vectors or patterns that encourage the agent to treat external content as policy. However, behavior analysis was not run, so it's not possible to confirm whether loading this skill materially changes downstream behavior compared to a no-skill baseline. Confidence is medium: the transparency gap around unreviewed remote execution is a real concern, and the absence of behavior testing leaves that risk dimension unvalidated.
The skill has 1 transparency concern that weaken pre-use reviewability, mainly around remote script execution without review boundary.
The scanned skill keeps data and instructions reasonably separate and does not strongly encourage the agent to treat external content as policy.
Behavior analysis was not run.
Remote script execution without review boundary