feat(sid): derive policy_fp at OTel sketch ingest via content lookup - #205
Merged
Merged
Conversation
Wires up the OTel sketch-ingest path to populate
`SketchInstanceMetadata.policy_fp` instead of leaving it `UNSET`.
Sketch-backed sids now participate in the `policy_fp → {sids}`
reverse index (PR #203), so the analyzer's
`find_matching_policies` (PR #204) returns fingerprints whose sids
are O(1)-reachable.
## What
- New `control_plane::warm_tier_analysis::find_policy_by_content(
registry, metric, group_by_keys, agg_type, expected_params)` —
finds the policy whose contents match a freshly-ingested sketch's
shape. Returns `Some(fp)` on a unique match, `None` on zero or
multiple matches (ambiguous → stay UNSET).
- Two data-plane helpers in `data_plane/src/drivers/ingest/otel.rs`:
- `aggregation_type_for_sketch_handle(SketchKindHandle) → Option<AggregationType>`
- `sketch_config_to_params(&SketchConfig) → HashMap<String, Value>`
Both lock in the data-plane → control-plane wire-shape mapping so
drift surfaces as test failures, not silent lookup misses.
- New `derive_sketch_policy_fp(ingest_state, metric, kind, cfg, group_by_keys)`
threads the content match. Called from the OTel sketch ingest
registration site (`route_modified_otlp_sketches_to_precompute`).
Returns `PolicyFingerprint::UNSET` when no policy matches — sids
stay reachable via the legacy `instances_matching(metric, gbk)`
walk.
## Matching shape
Match requires all of:
1. `policy.metric == metric`
2. `policy.aggregation_type == agg_type` (mapped from `SketchKindHandle`)
3. `policy.grouping_labels` (as a set) == `group_by_keys`
4. For every key in `expected_params`, `policy.parameters` has the
same `serde_json::Value` (extra policy params tolerated)
5. `policy.spatial_filter_normalized.is_empty()` — OTLP sketches
don't carry a filter context
Ambiguous match (multiple policies → same shape) returns `None`
intentionally. Such policies would have collided on sid identity
anyway — surfacing as UNSET is the honest signal of a control-plane
bug.
## Param-key vocabulary
| `SketchConfig` | params key(s) |
|---|---|
| `DDSketch { relative_accuracy }` | `relative_accuracy` |
| `Kll { k }` | `k` |
| `Hll { precision }` | `precision` |
| `CountSketch { rows, cols }` | `rows`, `cols` |
| `CountMin { rows, cols }` | `rows`, `cols` |
Names must stay in sync with `AggregationConfig::from_yaml_data` in
`asap_types/src/aggregation_config.rs`. The new test
`sketch_config_to_params_uses_canonical_keys` locks them.
## What's still standing
- OTel sketch-ingest doesn't surface a spatial-filter context, so
filtered policies remain unreachable from this path. When the
agent's processor emits a filter shape in the OTLP DP, plumb it
through to the lookup.
- Keyed-group ExactAgg (`MultipleSum` / `MultipleIncrease`) policies
don't appear in `policy_capability` either; if/when those land an
L4 binder, both this lookup and `find_matching_policies` need
matching arms.
## Test plan
- [x] 2 new unit tests in `otel::policy_fp_lookup_tests` covering the
`SketchKindHandle → AggregationType` mapping and the
`SketchConfig → params` rendering
- [x] `cargo check --workspace` clean
- [x] `cargo test --workspace --lib --bins` green (pre-existing
`avg_finds_sum_and_count` HashMap-iteration flake unchanged)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2 tasks
zzylol
added a commit
that referenced
this pull request
May 14, 2026
…ryEngine.execute (#206) Switches the query engine's candidate → sid resolution to the content-addressed fast path. PRs #203, #204, #205 landed the primitives; this PR integrates them. ## Before vs. after Before, per candidate: 1. `idx.instances_matching(metric, gbk)` walks the `RwLock<HashMap<u64, SketchInstanceMetadata>>` and returns every sid whose metric + group-by-keys match. 2. Per-sid: classify, check capability satisfaction, push to `hit_sids`. After, per candidate: 1. Snapshot the streaming config; build the `PolicyRegistry`. 2. `find_matching_policies(registry, candidate)` — walks the small policy registry (≤ thousands of entries) with the full match predicate (metric + group_by + capability + window + filter). Returns `Vec<PolicyFingerprint>`. 3. For each fp: `idx.sids_for_policy(fp)` — O(1) hash lookup over the reverse index added in PR #203. 4. Same per-sid classify + capability check as before (defensive; fast-path-discovered sids already satisfy the candidate by construction, but UNSET sids reached via the fallback don't). ## Slow-path fallback retained When the fast path yields zero sids — because either: - no policy in the registry matches the candidate (control plane hasn't published one yet), OR - the candidate's sids were registered with `PolicyFingerprint::UNSET` (legacy paths that didn't carry an `AggregationConfig` at ingest, test fixtures, raw mode) — the engine falls back to the metadata walk `instances_matching(metric, gbk)`. The per-sid capability filter below catches mismatches the fast path would have rejected at policy-match time. The fallback can be deleted in a follow-up once every code path populates `policy_fp` and existing on-disk records have aged out. ## Snapshot semantics The streaming-config snapshot is pinned once per query (not per candidate). Hot-reload swaps the underlying `Arc<StreamingConfig>` mid-query are isolated by the snapshot: the query sees the policy set that was active at query start. Same isolation the legacy `streaming_config_snapshot()` call already gave the engine elsewhere. ## Test plan - [x] `cargo check --workspace` clean - [x] `cargo test --workspace --lib --bins` green (`asap_types::capability_matching::tests::avg_finds_sum_and_count` HashMap-iteration flake unchanged) - Existing query-engine tests cover the slow-path fallback (their fixtures register sids with UNSET fp); future tests on the fast path land alongside an end-to-end test that builds a streaming config with matching policies and verifies the fp-keyed lookup is used. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
zzylol
added a commit
that referenced
this pull request
May 14, 2026
…ty-matching flake (#208) Two related cleanups bundled because both touch deterministic sid lookup and they're each small. ## 1. Drop `instances_matching` fallback in ASAPQueryEngine.execute PR #206 wired the content-addressed fast path with a fallback to the legacy `idx.instances_matching(metric, gbk)` walk for sids registered with `PolicyFingerprint::UNSET`. With PRs #203 + #205 populating the fp on every production registration path, the fallback's only consumers are test fixtures and the raw-mode fast-path that the sink already drops. Removing it makes "capability miss" mean exactly one thing — no policy in the registry satisfies the candidate — instead of overloading the miss path between "no policy" and "no sid metadata". ## 2. Make `aggregation_priority` a total order `capability_matching::tests::avg_finds_sum_and_count` had been flaky because `aggregation_priority` returned `Equal` on equal `window_size`, leaving `Vec::sort_by` order dependent on the underlying `HashMap` iteration. The test inserts both a `Sum` and a `CountMinSketch` config for the same metric; `Statistic::Sum` matches both, and when the multi-pop `CountMinSketch` sorted first the downstream key-aggregation lookup (only needed for multi-pop value types) missed and the whole match returned `None`. Sort keys now: 1. Larger `window_size` (coarser windows answer finer-grained queries via re-aggregation). 2. Single-population types beat multi-population (avoids the key-aggregation hunt when both shapes serve the statistic). 3. `aggregation_id()` (the policy fingerprint u64) tie-break — deterministic across runs and hosts. The result for the failing test: `Sum` always wins the Sum-stat candidate, no key-aggregation lookup fires, the function returns `Some(...)` deterministically. ## Engine test update `execute_returns_capability_miss_when_classify_is_ghost` previously asserted the detail string contained "ghost" / "unknown". With the fallback removed, the engine short-circuits at the policy-resolution step (the test fixture's `dd_meta` helper registers with `PolicyFingerprint::UNSET`, so the policy lookup never finds the sid). The `CapabilityMiss` outcome is preserved; the detail-string assertion is dropped since it pinned implementation, not contract. ## Test plan - [x] `cargo check --workspace` clean - [x] `cargo test --workspace --lib --bins` green - [x] `avg_finds_sum_and_count` ran 5× in a row, all pass — flake gone 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Wires the OTel sketch-ingest path to populate
SketchInstanceMetadata.policy_fpinstead of leaving itUNSET. Sketch-backed sids now participate in thepolicy_fp → {sids}reverse index (PR #203); the analyzer'sfind_matching_policies(PR #204) returns fingerprints whose sids are O(1)-reachable.What
control_plane::warm_tier_analysis::find_policy_by_content(registry, metric, group_by_keys, agg_type, expected_params)— finds the policy whose contents match a freshly-ingested sketch. ReturnsSome(fp)on a unique match;Noneon zero or multiple (ambiguous → stay UNSET).Two data-plane helpers (
otel.rs):aggregation_type_for_sketch_handle(SketchKindHandle) → Option<AggregationType>sketch_config_to_params(&SketchConfig) → HashMap<String, Value>Both lock in the data-plane → control-plane wire-shape mapping.
derive_sketch_policy_fpwires them together at the sketch-ingest registration site.Matching predicate
policy.metric == metricpolicy.aggregation_type == agg_typepolicy.grouping_labels(set) ==group_by_keysexpected_params,policy.parametershas the same value (extra policy keys tolerated)policy.spatial_filter_normalized.is_empty()Ambiguous match →
None. Those policies would have collided on sid identity anyway; surfacing UNSET is the honest signal.Not in scope
MultipleSum/MultipleIncrease) policies — nopolicy_capabilityarm exists yet.Test plan
cargo check --workspacecleancargo test --workspace --lib --binsgreen🤖 Generated with Claude Code