Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 14 additions & 5 deletions context/skills/self-driving/references/6-scouts.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Scout runs are also budgeted server-side: a project gets up to **100 scout runs

The server also caps how many scouts a project can have **enabled at once** (`max_enabled_scouts` in `scout-metadata-get`). The code default is far above ten, but a project or fleet override can set it lower. The server rejects any enable past that cap, including the enables in this step and in step 6b. Every enabled config counts toward it: the operational scouts, custom scouts the team already has, and the troop. So the cap can bind before the ten-scout ceiling does — step 1c reads it.

**Operational scouts are not part of the troop.** A few rows come back from `scout-config-sync` with `scout_role: "operational"` — today `signals-scout-inbox-validation`. They watch the self-driving system itself rather than a product surface: inbox validation is the lane that runs every follow-up check a report schedules, and those checks start arriving with the first reports, so a fresh project needs it as much as an old one. The server seeds them enabled and outside the ten-scout ceiling, even when the project is at its enabled-scout cap. They still count toward that cap, so each enabled operational scout uses one of its slots. Keep them enabled: a disabled operational scout ends every check on its lane as errored. Earlier versions of this setup disabled inbox validation, and `scout-config-sync` does not reset an existing row, so on a rerun you may find one off — step 3 turns it back on. If a row has no `scout_role` field (older PostHog deploy), treat `signals-scout-inbox-validation` as operational by name.
**Operational scouts are not part of the troop.** A few rows come back from `scout-config-sync` with `scout_role: "operational"` — today `signals-scout-inbox-validation`. They watch the self-driving system itself rather than a product surface: inbox validation is the lane that runs every follow-up check a report schedules, and those checks start arriving with the first reports, so a fresh project needs it as much as an old one. The server seeds them enabled and outside the ten-scout ceiling, even when the project is at its enabled-scout cap. They still count toward that cap, so each enabled operational scout uses one of its slots. Keep them enabled: a disabled operational scout ends every check on its lane as errored. But never force one back on. A person can pause an operational scout on purpose, and the failure breaker pauses a scout whose runs keep failing. The server never resumes either pause, and setup does not overrule them. `scout-config-sync` already resumes the pauses the harness made itself (an inactivity-sweep pause, or a row the seed created disabled), so a row that is still off after the sync is one the server chose to leave off — step 3 reports it instead. If a row has no `scout_role` field (older PostHog deploy), treat `signals-scout-inbox-validation` as operational by name.

## Status

Expand All @@ -34,11 +34,11 @@ Reach the scout-config tools through the PostHog `exec` tool — `info` then `ca

1b. **Read the run budget**: call `scout-metadata-get`. It returns the enforced limits (`max_runs_per_day` — `null` means unbounded — plus `runs_today` and `runs_remaining_today`) and any announcement banner. Record the limits and the banner text for the report. At the default budget the ceiling that binds is the **ten-scout troop**, not the run budget, so size the pick against that and treat this read as a check that nothing unusual is in force. **Only if the budget is genuinely low, shrink the pick to fit**: when `max_runs_per_day` is under ~10 (remember step 6b adds custom scouts at the same cadence), enable fewer specialists so the whole enabled set fits inside it, leaving a run or two a day for the operational scouts (their check runs draw from the same budget) — `general` + one specialist is the usual floor, and only when even that exceeds the budget fall back to `general` alone and note it in the report. **Soft-degrade if the tool is missing or fails** (older PostHog deploy): size against the ten-scout ceiling as usual and continue. **Not an abort.** In the report, say the budget could not be read and quote 100/day as the published default rather than as this project's verified limit — you did not read it, so do not state it as fact.

1c. **Count the free enabled-scout slots.** Read `max_enabled_scouts` from the same `scout-metadata-get` response. From the `scout-config-sync` rows, count the enabled rows this setup keeps on whatever it picks: the operational scouts, and every enabled custom scout (`scout_origin: "custom"`). Then:
1c. **Count the free enabled-scout slots.** Read `max_enabled_scouts` from the same `scout-metadata-get` response. From the `scout-config-sync` rows, count the rows that need a slot whatever this setup picks: the enabled operational scouts, every operational scout the failure breaker paused (`status: "paused_by_system"`, `pause_reason: "repeated_failures"`), and every enabled custom scout (`scout_origin: "custom"`). A breaker-paused scout is off now, but the breaker resumes it after a successful probe, and the server refuses that resume when the project is at its cap. So it keeps its slot. Then:

`slots = max_enabled_scouts − (enabled operational scouts) − (enabled custom scouts)`
`slots = max_enabled_scouts − (enabled operational scouts) − (breaker-paused operational scouts) − (enabled custom scouts)`

`slots` is the number of non-operational, non-custom scouts this step and step 6b can have enabled **together**. Cap the pick in step 2 so `general` + the specialists fit inside `slots`, and remember step 6b needs room from the same number. At the default cap `slots` is far above ten, so the ten-scout ceiling still binds and nothing changes. When `slots` is small, drop specialists first (least-used first), then `general` only if `slots` is 1 or less — with `slots` at 1 keep `general` alone, and with `slots` at 0 enable no troop scout at all. Whenever `slots` removes a scout you would otherwise enable, record a follow-up: "This project can have at most N scouts enabled at once, and M slots were already in use. Ask PostHog to raise the project's `max_enabled_scouts`, or disable a scout you no longer need, then enable `<scout>`." Record `max_enabled_scouts` and `slots` for the report. **Soft-degrade** if the field is absent (older PostHog deploy): skip this bound and size against the ten-scout ceiling and the run budget as usual.
`slots` is the number of non-operational, non-custom scouts this step and step 6b can have enabled **together**. Cap the pick in step 2 so `general` + the specialists fit inside `slots`, and remember step 6b needs room from the same number. At the default cap `slots` is far above ten, so the ten-scout ceiling still binds and nothing changes. When `slots` is small, drop specialists first (least-used first), then `general` only if `slots` is 1 or less — with `slots` at 1 keep `general` alone, and with `slots` at 0 or below enable no troop scout at all. If `slots` is below 0 and an operational scout is breaker-paused, the scouts already enabled fill the cap and the breaker cannot resume it. Record this follow-up for that scout instead of its step 3 breaker follow-up: "The failure breaker paused `<scout>` after repeated failed runs, and the project has no free enabled-scout slot for it. The breaker cannot resume it until a slot is free. Ask PostHog to raise the project's `max_enabled_scouts`, or disable a scout you no longer need. Then look at the scout's recent runs in PostHog to find the failure." Whenever `slots` removes a scout you would otherwise enable, record a follow-up: "This project can have at most N scouts enabled at once, and M slots were already in use. Ask PostHog to raise the project's `max_enabled_scouts`, or disable a scout you no longer need, then enable `<scout>`." Record `max_enabled_scouts` and `slots` for the report. **Soft-degrade** if the field is absent (older PostHog deploy): skip this bound and size against the ten-scout ceiling and the run budget as usual.

2. **Decide the enabled set — the whole point of this step is to enable FEW scouts, not many.** Work from the rows `scout-config-sync` actually returned (the troop grows over time — ~19 scouts today — so never hardcode a list). The enabled set has exactly three parts:

Expand Down Expand Up @@ -72,7 +72,16 @@ Reach the scout-config tools through the PostHog `exec` tool — `info` then `ca
- **At least one.** Always end with a specialist enabled. If no product surface clearly stands out — e.g. the only products in use are error tracking / session replay (excluded in (b)), or the profile was unavailable and nothing is rankable — **fall back to one universal cross-product scout** (`signals-scout-anomaly-detection` or `signals-scout-health-checks`) as the stand-in.
- **A scout the table doesn't name** (posthog keeps adding them): treat it as a specialist candidate — read its description, judge whether its surface is among this project's most-used, and enable it only if it earns one of the ≤5 slots.

3. **Disable every scout you did NOT enable** in (a)–(c) — this is now most of the troop. Skip the operational scouts: they are not yours to turn off here, and they stay enabled whatever (a)–(c) picked. Also skip every custom scout (`scout_origin: "custom"`): the team wrote it, so it keeps its current state. An enabled custom scout uses a slot (step 1c), but this step never disables it to free that slot. If an operational scout comes back with `enabled: false`, update it to `{ enabled: true }` instead, and note in the report that setup turned it back on. Disable via `scout-config-update` with the config `id` and `{ enabled: false }` — that one field is the whole update, since `emit` (dry-run posture) and `run_interval_minutes` keep the server's defaults, which are the intended posture. A failed update is a follow-up, not an abort.
3. **Disable every scout you did NOT enable** in (a)–(c) — this is now most of the troop. Skip the operational scouts: they are not yours to turn off here, and they stay enabled whatever (a)–(c) picked. Also skip every custom scout (`scout_origin: "custom"`): the team wrote it, so it keeps its current state. An enabled custom scout uses a slot (step 1c), but this step never disables it to free that slot. Never send `scout-config-update` for an operational scout, not even `{ enabled: true }`. Disable via `scout-config-update` with the config `id` and `{ enabled: false }` — that one field is the whole update, since `emit` (dry-run posture) and `run_interval_minutes` keep the server's defaults, which are the intended posture. A failed update is a follow-up, not an abort.

**Leave a disabled operational scout as it is, and report it.** If an operational scout comes back from `scout-config-sync` with `enabled: false`, the sync already tried the recovery the server supports and left the row off. Read its `status` and `pause_reason`, and record one follow-up:

| Row state | Follow-up |
|---|---|
| `status: "paused_by_user"` | "`<scout>` is paused. A person paused it, or an earlier version of this setup did. Setup left it off. While it is off, every report follow-up check on its lane ends as errored. Turn it back on in PostHog if nobody paused it on purpose." |
| `status: "paused_by_system"`, `pause_reason: "repeated_failures"` | "The failure breaker paused `<scout>` after repeated failed runs. It probes the scout and resumes it after a successful run, and setup kept an enabled-scout slot free for that. Setup left it off. Look at the scout's recent runs in PostHog to find the failure." Use the step 1c text instead when step 1c found no slot for it. |
| `status: "paused_by_system"`, any other `pause_reason` | "`<scout>` stayed paused after the sync tried to resume it, usually because the project is at its enabled-scout cap. Free a slot, then turn it back on in PostHog." |
| no `status` field (older PostHog deploy) | "`<scout>` is disabled, and setup could not read why. Setup left it off. Turn it back on in PostHog if nobody paused it on purpose." |

**Do the disables first, then the enables.** The server checks the enabled-scout cap against the count at the time of each enable. `scout-config-sync` can leave scouts enabled that this step turns off, so an enable sent before those disables can fail even when the final set fits. Enable a picked scout that came back disabled via `scout-config-update` with `{ enabled: true }`. If an enable still fails with the cap error, record the scout as a follow-up with the same text as step 1c, and continue.

Expand Down
2 changes: 1 addition & 1 deletion context/skills/self-driving/references/7-report.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ Emit:
- **Products enabled** — from step 3: a short table of Session Replay / Error Tracking / Support, each as **enabled** / **already enabled** / **enabled but inert** (backend or mobile — the server flip is on, but it needs SDK code on this platform before it captures anything) / **not enabled** (a non-admin couldn't turn it on — the project-admin follow-up). The server flip happens regardless of platform, so a backend/mobile product is *enabled, just inert* — never "skipped". For a web app, note whether the `posthog.init` override check was clean or edited. This is the *product* toggle, distinct from the signal sources below. **Support row:** when Conversations is on, tickets only arrive once an inbound channel is connected — spell out that the user must connect a channel (email / inbox / Slack) in PostHog, and add a matching follow-up.
- **Signal sources** — a table of every source you touched or deliberately skipped: `source_product` / `source_type`, action taken (enabled / already enabled / skipped + why / failed).
- **Connected tools** — what the user picked, and per tool the step-5 class: "connected by this setup (source id …, first sync started)", "already connected" / "verified connected", "responder enabled but warehouse source not detected (dormant)", or "not used" (only for tools the user didn't pick). Never report a tool as connected unless this run created its source or saw it in `external-data-sources-list`. For sources this run created, note that only the responder-consumed table (issues / tickets) is syncing and more can be enabled in the UI. For a Linear source this run created, also say which Linear teams Self-driving reads: the team names picked in step 5b, or "all teams", and that this can be changed in the inbox (the Linear source's Filters button). Any tool the user picked but didn't connect — whether they said "done" or skipped — is "selected but no source detected (dormant)" with a follow-up, never "user confirmed connecting" and never "not used".
- **Scout troop** — kept-on scouts, the operational scouts the system keeps on (one line: they run report follow-up checks and sit outside the troop ceiling), disabled scouts with the one-line reason each, or the not-yet-materialized note from step 6. Include the run budget from step 6's `scout-metadata-get` read — max runs per day, runs used today, and the enabled-scout cap (`max_enabled_scouts`) with the free slots step 6 counted, plus any scout the cap kept off — plus the announcement banner text if one was returned (it says how to request more runs); if the metadata read soft-degraded or never ran, say the budget could not be read and give 100 runs a day as the published early-access default, not as this project's confirmed limit.
- **Scout troop** — kept-on scouts, the operational scouts the system keeps on (one line: they run report follow-up checks and sit outside the troop ceiling), plus any operational scout step 6 found paused and left off, with its follow-up, disabled scouts with the one-line reason each, or the not-yet-materialized note from step 6. Include the run budget from step 6's `scout-metadata-get` read — max runs per day, runs used today, and the enabled-scout cap (`max_enabled_scouts`) with the free slots step 6 counted, plus any scout the cap kept off — plus the announcement banner text if one was returned (it says how to request more runs); if the metadata read soft-degraded or never ran, say the budget could not be read and give 100 runs a day as the published early-access default, not as this project's confirmed limit.
- **Custom scouts** — from step 6b: each created scout (name, what it watches, its discriminator, and why no built-in scout covers it) or one line on why none was warranted; surfaces considered and ruled out, with the filter that killed each; declined proposals; and the noise escape hatch (set `emit: false` on a scout's config in PostHog to switch it to dry-run). Omit only if step 6b was skipped entirely.
- **Replay Vision scanners** — from step 6c: a row per brief (the breakage monitor, the frustration monitor) with its name, what it watches, the query scope you chose and — for the breakage monitor — why that's this product's completion flow, its `sampling_rate`, and its estimated monthly credit spend (with the observation count as context). Mark each **created** / **updated an earlier run's scanner** / **skipped** with the reason (not a web app, no identifiable completion flow, the team already covers it, the scanner API isn't available on this deploy). Say plainly what a scanner is — an LLM that watches individual session recordings on a schedule and pushes what it finds to the inbox — that it is the only thing in this setup which spends Replay Vision quota, and that findings arrive at half weight so they need corroboration before they're promoted into a report. If the project has no recordings yet, note that the scanners are armed and start working the day recordings begin. Omit only if step 6c was skipped entirely.
- **Follow-ups** — every follow-up recorded along the way, as a checklist. Omit the section if there are none.
Expand Down
Loading