exploring-scouts
Exploring Signals scouts
A scout is a scheduled agent that wakes on its own interval, looks at one PostHog project, decides what's genuinely worth surfacing, and either writes it into the Signals inbox as a report or closes out empty (a real, valid outcome).
PostHog ships a fleet of canonical scouts — a cross-product generalist (signals-scout-general) plus per-surface specialists (error tracking, logs, AI observability, experiments, feature flags, session replay, web analytics, surveys, and more).
A project may also have custom scouts beyond the canonical fleet — any signals-scout-* skill a team authored (e.g. -brand-mentions, -mcp-feedback) shows up here too, so don't assume a fixed roster: scout-config-list is the authoritative roster for a project.
(One caveat: a just-authored scout has no config row until the coordinator's next tick auto-registers one — or until someone registers it via the write-side scout-config-create — so a brand-new scout may briefly be missing from the list.)
This skill helps you understand and explore what a project's scouts are doing and how they're performing — entirely through read-only MCP tools.
It is the observability counterpart to the authoring-scouts skill (which teaches writing and tuning) and to the inbox-exploration skill (which covers the inbox reports scouts feed into).
(The scout tools were recently renamed from signals-scout-* to scout-*; if a scout-* name comes back unknown, the server may still expose it under the legacy signals-scout-* name — search the tool catalog and call whichever name it returns.)
A scout's output is inbox reports, written 1:1. Scouts list emit_report / edit_report in their allowed_tools and author or edit inbox reports directly; a run's output shows up as emitted_report_ids (reports it authored) and edited_report_ids (reports it updated).
The run rows also carry emitted_count / emitted_finding_ids — legacy fields from the deprecated signal-emitting channel (weak emit_signal findings a pipeline consolidated). On a report-channel scout they stay 0 / empty even on a productive run; a non-zero tally means the run came from a scout still on the legacy channel (an old custom scout, or a canonical scout not yet ported) — real output for that run, not noise. When unsure of a scout's channel, check its allowed_tools via skill-get.
Never read emitted_count: 0 as "did nothing" — check the report columns and the run summary first.
A scout whose config carries a structured_output_schema has a third output channel next to reports: schema-validated measurement records, recorded as $scout_structured_output events in the project (only scalar top-level payload keys flatten to output_<key> properties — object and array fields live solely inside the full output property, so a missing output_<key> is not a missing value; subject names the judged entity) rather than as run-row columns.
The events are the ground truth — metadata.derived.has_structured_output says the run had at least one batch accepted, which is a fast per-run screen but not delivery confirmation (a rare capture failure after acceptance leaves it true with fewer or no events behind it), so count the events when the number of records matters. See references/scout-data-model.md for the event shape.
Each run also carries a metadata map. Top-level: the provenance set harness_prompt_version / report_channel (none, emit, edit, or both) / skill_origin / github_guidance, saying which instructions the run was given; plus routing keys (model / runtime_adapter / reasoning_effort) only when a gate or pin overrode the default. Nested under metadata.derived: booleans the harness computes at the end of the run (has_emit_report, has_edit_report, has_self_improvement, has_chart, has_self_validation, has_structured_output).
When comparing runs (before/after a prompt change, one model against another), segment on all four provenance values first: runs differing on any of harness_prompt_version, report_channel, skill_origin, or github_guidance were given different instructions and aren't a like-for-like population. Runs predating this field have none of them, so treat missing provenance as unknown and exclude those runs from a comparison rather than pooling them.
For "what kind of run was this?" questions — did it author a self-improvement report, did it validate its follow-up queue — read derived rather than parsing the prose summary. It's computed server-side from what the run actually did, so it can't disagree with the run's own output — with one exception: has_structured_output tracks batches the run had accepted, not events delivered, so it alone can be true with fewer or no records behind it (count the events, as above). No derived map at all means unknown, not "all false" — the run predates the field, failed before finishing, or its stamp failed. Most runs from before this shipped have no map, so don't read their absence as a finding.