feat(ai): priority classification for AI/LLM requests (#285) - #286
Merged
dan-petty merged 2 commits intoSep 19, 2026
Merged
Conversation
… Obs 18-20, Systems 09-10 New roadmap items derived from empirical vibes observations: P0 - Critical: - Lazy Domain-Gated MCP Tool Schema Hydration (Obs 18): 100+ tools consume 25% of context window; lazy hydration reclaims 85% - Pipeline Stage Context Budgeting & Invariant Pinning (Obs 19): sequential pipeline context bloat; invariant eviction under multi-turn drift - Capability-Gated Model Failover & AIMD Batch Recovery (Obs 20): 70B→14B failover cliff; one-way embedding batch ratchet P1 - High: - Lossless Structured Error Reflection for Schema Retries (Obs 19): 256-char truncation forces multi-turn retry loops - Background Shell Pipe Deadlock Fix & Output Contract (Systems Obs 09): 64KB pipe buffer deadlock on verbose commands - Structured Constraint Propagation Across Subagent Delegation (Systems Obs 09): 3-layer delegation retains only 61% fidelity - MCP Resource-First Data Access & Tool Output Sandboxing (Systems Obs 10): Resources 3x cheaper for read-heavy patterns - Lossless Structured Error Reflection (Obs 19): single-turn correction via structured field-path errors
dan-petty
added a commit
that referenced
this pull request
Sep 19, 2026
* feat(release): v0.2.21 (#279) * fix(reliability): harden exception handling, optimize telemetry & clean data tier (#280) (#281) * feat(reliability): harden exception handling, optimize telemetry & clean data tier (#280) - Harden exception handling across AI, Git, Security, and Core modules with explicit error types and context - Prevent silent error suppression and arbitrary default fallbacks across public APIs - Optimize Jaeger/OTel traces, Prometheus metrics, and FluentBit logging configurations - Enhance GitHub rate limiting client with header tracking and proactive quota safety - Clean up .data directory handling, configure dedicated data.dir path, and deduplicate config writes - Add forward-looking roadmap initiatives for process group management and merge readiness - Expand test suites and architectural invariants for data dir config and stray script prevention - Track deliverable completion in docs/agent/tasks/task-280-harden-exceptions-telemetry-data-tier.md * fix(security): sanitize CodeQL clear-text secret logging in settings and credentials (#280) * feat(ai): multi-scale semantic outline & inspectional scanner (#272) (#282) * feat(ai, telemetry): dynamic slot leasing, telemetry deduplication, and context-aware file review (#283) (#284) - Implement dynamic least-loaded slot leasing across candidate Ollama servers using condition variables to prevent Head-of-Line blocking. - Add multi-layered file classification (shebangs, MIME types, AST/JSON/YAML/TOML parsing, canonical filenames/extensions) to route documentation, configuration, and code files to specialized review task prompts and persona subsets. - Optimize review pre-analysis batching to run static AST extraction without redundant LLM chat calls. - Fix OTelTyper lazy proxy registration to attribute code.namespace and code.function to target commands and eliminate duplicate telemetry attributes. - Add comprehensive test suites for slot leasing, file classification, and telemetry deduplication. * feat(ai): priority classification for AI/LLM requests (#285) (#286) * docs(roadmap): expand v0.2.24 with 8 vibes-grounded improvements from Obs 18-20, Systems 09-10 New roadmap items derived from empirical vibes observations: P0 - Critical: - Lazy Domain-Gated MCP Tool Schema Hydration (Obs 18): 100+ tools consume 25% of context window; lazy hydration reclaims 85% - Pipeline Stage Context Budgeting & Invariant Pinning (Obs 19): sequential pipeline context bloat; invariant eviction under multi-turn drift - Capability-Gated Model Failover & AIMD Batch Recovery (Obs 20): 70B→14B failover cliff; one-way embedding batch ratchet P1 - High: - Lossless Structured Error Reflection for Schema Retries (Obs 19): 256-char truncation forces multi-turn retry loops - Background Shell Pipe Deadlock Fix & Output Contract (Systems Obs 09): 64KB pipe buffer deadlock on verbose commands - Structured Constraint Propagation Across Subagent Delegation (Systems Obs 09): 3-layer delegation retains only 61% fidelity - MCP Resource-First Data Access & Tool Output Sandboxing (Systems Obs 10): Resources 3x cheaper for read-heavy patterns - Lossless Structured Error Reflection (Obs 19): single-turn correction via structured field-path errors * feat(ai): priority classification for AI/LLM requests (#285) * feat(roadmap): add GitHub/VS Code agentic integrations and Grafana dashboards (#287, #289) (#288) * feat(roadmap): add GitHub and VS Code agentic integrations (#287) * feat(roadmap): add devops-cli Grafana dashboards suite (#289) * feat(ai): track approximate lifetime spend per backend service with prometheus and grafana observability (#290) (#291) * feat(ai): track approximate lifetime spend per backend service with prometheus and grafana observability (#290) * fix(security): resolve Bandit B608 by using static parameterized queries in SpendLedger * feat(ai, k8s): prioritize roadmap into 0.3.x-0.5.x lines and integrate LightLLM and Portkey AI routing (#292) * feat(ai, k8s): prioritize roadmap into 0.3.x-0.5.x lines and integrate LightLLM and Portkey AI routing * feat(roadmap, k8s): add turnkey Grafana dashboards for all Kubernetes stacks (#293) (#294) * feat(roadmap, k8s): add turnkey Grafana dashboards for all Kubernetes stacks (#293) * docs(sdlc): fix Phase 6 Mermaid release choreography diagram * feat(roadmap): synchronize milestones, issues #307-#326, release epics, and project fields (#327) * docs(release): compile v0.2.21 changelog notes (#328) * fix(release): query milestone deliverables and enforce branch protection in pre-commit (#331) * fix(release): query milestone deliverables and enforce branch protection in pre-commit (#330) * build(ci): move devops ci quality gate to pre-push hook
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #285.
Summary
Implements priority classification across AI / LLM requests and dynamic slot leasing:
RequestPriorityenum (HIGH,NORMAL,AS_AVAILABLE) with canonical weights and comparison operators.acquire_ollama_slotinnetwork.pywith priority queueing and capacity reservation:HIGHpriority requests (chat, interactive loops) take immediate precedence and preempt waiting lower-priority tasks for newly freed slots.NORMALpriority requests (reviews, analysis, standard CLI commands) lease capacity fairly.AS_AVAILABLEpriority requests (background tasks, prewarming, scheduled jobs) only lease when headroom is available (capping concurrency to leave a slot open whenmax_parallel > 1) and default to non-blocking timeout (0.0s) to avoid starving interactive workloads.current_request_priorityContextVar andrequest_priority_scope(...)context manager, integrated intodevops ai chat(scoped toHIGH),devops ai prewarm(scoped toAS_AVAILABLE), and LLM client dispatch methods.gen_ai.request.priorityattribute onto trace spans (ai.llm.dispatch,gen_ai.chat,gen_ai.stream).Verification
tests/test_ai_request_priority.pyverifying priority order, preemption, non-blocking background rejection, and contextvar scoping.uv run devops ci) passed with 100% passing checks (coverage >= 90%, 0 errors, 0 warnings).