Skip to content

feat(ai): track approximate lifetime spend per backend service with prometheus and grafana observability (#290) - #291

Merged
dan-petty merged 2 commits into
release/v0.2.21from
feat/290-ai-spend-tracking
Sep 19, 2026
Merged

dan-petty merged 2 commits into
release/v0.2.21from
feat/290-ai-spend-tracking

Conversation

@dan-petty

Copy link
Copy Markdown
Owner

Closes #290.

Overview & Objectives

Implements persistent approximate lifetime AI/LLM spend tracking per backend service/server and exposes spend, token, and request metrics for Prometheus scraping and Grafana visualization.

Key Changes

  • Persistent AI Spend Ledger (src/devops_cli/ai/spend/):
    • tracker.py: Thread-safe atomic ledger tracking cumulative input/output tokens, request counts, and estimated USD spend across servers and models.
    • pricing.py: Multi-tier pricing calculation engine using industrial average rate tables with local override capability.
    • pricing_data.py: Open source industrial benchmark pricing catalog covering OpenAI, Anthropic, DeepSeek, Google Gemini, and local inference ($0/token).
    • prometheus.py: Prometheus text exposition formatter generating standard counters and gauges (devops_cli_ai_spend_usd_total, devops_cli_ai_tokens_total, devops_cli_ai_requests_total, devops_cli_ai_active_servers, devops_cli_ai_active_models, devops_cli_ai_server_spend_usd).
  • REST & CLI Observability:
    • src/devops_cli/server/routes/telemetry.py: /metrics endpoint exports Prometheus spend metrics alongside system and telemetry counters.
    • src/devops_cli/commands/ai_cost.py: Added --format prometheus and subcommand devops ai cost prometheus.
  • Grafana Dashboards:
    • k8s/monitoring/dashboards/ai-spend.json: Dedicated 11-panel AI spend dashboard (uid: "ai-spend-telemetry").
    • k8s/monitoring/dashboards/devops-cli.json: Integrated "AI Spend & Economics" panels.
  • Verification:
    • Comprehensive unit test suite in tests/test_ai_spend_*.py.
    • 100% passing across all CI quality gates (uv run devops ci).

@dan-petty
dan-petty marked this pull request as ready for review September 19, 2026 16:46
@dan-petty
dan-petty merged commit 3e3cf25 into release/v0.2.21 Sep 19, 2026
4 checks passed
@dan-petty
dan-petty deleted the feat/290-ai-spend-tracking branch September 19, 2026 17:34
dan-petty added a commit that referenced this pull request Sep 19, 2026
* feat(release): v0.2.21 (#279)

* fix(reliability): harden exception handling, optimize telemetry & clean data tier (#280) (#281)

* feat(reliability): harden exception handling, optimize telemetry & clean data tier (#280)

- Harden exception handling across AI, Git, Security, and Core modules with explicit error types and context
- Prevent silent error suppression and arbitrary default fallbacks across public APIs
- Optimize Jaeger/OTel traces, Prometheus metrics, and FluentBit logging configurations
- Enhance GitHub rate limiting client with header tracking and proactive quota safety
- Clean up .data directory handling, configure dedicated data.dir path, and deduplicate config writes
- Add forward-looking roadmap initiatives for process group management and merge readiness
- Expand test suites and architectural invariants for data dir config and stray script prevention
- Track deliverable completion in docs/agent/tasks/task-280-harden-exceptions-telemetry-data-tier.md

* fix(security): sanitize CodeQL clear-text secret logging in settings and credentials (#280)

* feat(ai): multi-scale semantic outline & inspectional scanner (#272) (#282)

* feat(ai, telemetry): dynamic slot leasing, telemetry deduplication, and context-aware file review (#283) (#284)

- Implement dynamic least-loaded slot leasing across candidate Ollama servers using condition variables to prevent Head-of-Line blocking.
- Add multi-layered file classification (shebangs, MIME types, AST/JSON/YAML/TOML parsing, canonical filenames/extensions) to route documentation, configuration, and code files to specialized review task prompts and persona subsets.
- Optimize review pre-analysis batching to run static AST extraction without redundant LLM chat calls.
- Fix OTelTyper lazy proxy registration to attribute code.namespace and code.function to target commands and eliminate duplicate telemetry attributes.
- Add comprehensive test suites for slot leasing, file classification, and telemetry deduplication.

* feat(ai): priority classification for AI/LLM requests (#285) (#286)

* docs(roadmap): expand v0.2.24 with 8 vibes-grounded improvements from Obs 18-20, Systems 09-10

New roadmap items derived from empirical vibes observations:

P0 - Critical:
- Lazy Domain-Gated MCP Tool Schema Hydration (Obs 18): 100+ tools consume 25% of context window; lazy hydration reclaims 85%
- Pipeline Stage Context Budgeting & Invariant Pinning (Obs 19): sequential pipeline context bloat; invariant eviction under multi-turn drift
- Capability-Gated Model Failover & AIMD Batch Recovery (Obs 20): 70B→14B failover cliff; one-way embedding batch ratchet

P1 - High:
- Lossless Structured Error Reflection for Schema Retries (Obs 19): 256-char truncation forces multi-turn retry loops
- Background Shell Pipe Deadlock Fix & Output Contract (Systems Obs 09): 64KB pipe buffer deadlock on verbose commands
- Structured Constraint Propagation Across Subagent Delegation (Systems Obs 09): 3-layer delegation retains only 61% fidelity
- MCP Resource-First Data Access & Tool Output Sandboxing (Systems Obs 10): Resources 3x cheaper for read-heavy patterns
- Lossless Structured Error Reflection (Obs 19): single-turn correction via structured field-path errors

* feat(ai): priority classification for AI/LLM requests (#285)

* feat(roadmap): add GitHub/VS Code agentic integrations and Grafana dashboards (#287, #289) (#288)

* feat(roadmap): add GitHub and VS Code agentic integrations (#287)

* feat(roadmap): add devops-cli Grafana dashboards suite (#289)

* feat(ai): track approximate lifetime spend per backend service with prometheus and grafana observability (#290) (#291)

* feat(ai): track approximate lifetime spend per backend service with prometheus and grafana observability (#290)

* fix(security): resolve Bandit B608 by using static parameterized queries in SpendLedger

* feat(ai, k8s): prioritize roadmap into 0.3.x-0.5.x lines and integrate LightLLM and Portkey AI routing (#292)

* feat(ai, k8s): prioritize roadmap into 0.3.x-0.5.x lines and integrate LightLLM and Portkey AI routing

* feat(roadmap, k8s): add turnkey Grafana dashboards for all Kubernetes stacks (#293) (#294)

* feat(roadmap, k8s): add turnkey Grafana dashboards for all Kubernetes stacks (#293)

* docs(sdlc): fix Phase 6 Mermaid release choreography diagram

* feat(roadmap): synchronize milestones, issues #307-#326, release epics, and project fields (#327)

* docs(release): compile v0.2.21 changelog notes (#328)

* fix(release): query milestone deliverables and enforce branch protection in pre-commit (#331)

* fix(release): query milestone deliverables and enforce branch protection in pre-commit (#330)

* build(ci): move devops ci quality gate to pre-push hook
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant