Skip to content

fix(copilot): avoid float drift in live budget observations - #170

Merged
Yushangjinghong2 merged 3 commits into
devfrom
fix/copilot-integer-budget-observation-20260928
Oct 7, 2026
Merged

Yushangjinghong2 merged 3 commits into
devfrom
fix/copilot-integer-budget-observation-20260928

Conversation

@fanyangCS

Copy link
Copy Markdown
Collaborator

Summary

Fix Copilot live budget cost accumulation by summing integer nano-AIU before converting the total to USD once. This is a focused Python bug fix plus synthetic offline regression tests; no TypeScript, configuration, receipt validation, settlement, reconciliation or recovery code changes.

Problem and root cause

LiveBudgetMonitor.check() currently converts each provider row to a float and adds the floats. Canonical CopilotCallUsage.cost_usd instead sums integer billing units and converts once. Thirteen generated rows of 11,000,000,000 nano-AIU each reproduce an observation of 1.4300000000000004 versus the canonical 1.43 (two ULPs). A conservative observed-cost obligation should not be inflated by row-wise representation error. In runtimes with strict complete-receipt floors this mismatch can reject an otherwise truthful final receipt.

The new explicitly typed integer accumulator retains the existing current-day filter and optional-price handling. It does not use epsilon comparisons, round away genuine charges, weaken validators, erase obligations, change cached-token accounting or mutate existing records.

Tests and verification

  • Base: current dev at 34588003ce89db744fe1c6c11268d823b0c27726, per CONTRIBUTING.md; independent clean clone, not an installed-runtime branch.
  • Identical new test bytes over base production source (read-only overlay): 7 failed, all at exact observed-cost equality; existing monitor tests 5 passed. Patched focused suite: 216 passed.
  • Regression coverage: thirteen generated charges; growing cumulative query batches and repeated polls; forward/reverse/interleaved orders; mixed models; cached input without double counting; missing optional prices/tokens; old/undated rows; cold bound-session lookup; real cost/token budget enforcement and lower repeated observations cannot refund usage.
  • Focused command: python -m pytest tests/test_live_budget_rounding.py tests/test_live_budget_monitor.py tests/core/test_accounting_integrity.py tests/core/test_cost_settlement_floor.py tests/core/test_cost_control_live.py tests/core/test_cost_control.py tests/core/test_cost_recovery.py tests/core/test_daily_token_budget.py tests/core/test_usage.py tests/provider_integrations/test_copilot_usage.py -o addopts='' -p no:cacheprovider -q
  • ruff check argus argus_skill tests: passed.
  • python -m argus.release_tools.typecheck_gate --base 34588003ce89db744fe1c6c11268d823b0c27726: passed, 0 introduced diagnostics (1263 existing first-party diagnostics on both trees in the same isolated dependency environment). This is the repository's differential gate, not a claim of zero type debt.

All focused tests ran offline with network/PID isolation, read-only source and disposable state, using existing Python 3.12 dependencies. No provider/model calls. Initial system-Python attempt had four dependency/environment failures; rerunning the complete selected suite with the complete dependency environment produced the 216-pass result above.

Scope and limitations

Current dev does not contain the separate strict assert_complete_receipt_floor implementation from the locally diagnosed runtime, so this PR does not import that accounting stack or pretend to test that absent API on dev. The unchanged installed strict guard was separately checked offline with synthetic receipts: the truthful value was accepted and genuine cost/token shortfalls were rejected. That is separate installed-runtime evidence, not published-dev coverage.

The prior installed-runtime repair's 154 passed / 1 inherited fixture failure is also separate evidence, not this PR's test result. That fixture expected a resumed session to change; it failed identically on the original installed source. Current-dev monitor tests pass. No recovery scripts, private records, prompts, real call/project/session identifiers or provider logs are included here.

Documentation: no user-facing interface or policy change; rationale is documented inline and in this description, so no separate docs file is needed. Full repository/cross-platform results remain subject to CI. No merge, deployment, runtime installation or campaign restart is included.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Copilot review overview

Review effort: Lite
Findings: 2 Medium severity

Open (2)
What changed in this PR

Adds regression coverage and a rounding-safe implementation for Copilot usage cost observations so budget obligations match canonical integer billing.

Changes:

  • Add synthetic tests that exercise Copilot usage ordering, missing fields, date filtering, and budget “no-reopen” behavior.
  • Update LiveBudgetMonitor.check() to sum total_nano_aiu as integers and convert once to avoid float accumulation inflation.
File Description
tests/​test_live_budget_rounding.py New regression tests validating integer-sum cost observation and budget obligation behavior.
argus/​adapters/​agent_cli_backend/​_budget_monitor.py Fix cost accumulation by summing nano units as integers before converting to USD.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread argus/adapters/agent_cli_backend/_budget_monitor.py Outdated
Comment thread tests/test_live_budget_rounding.py Outdated
@Yushangjinghong2
Yushangjinghong2 merged commit 7657269 into dev Oct 7, 2026
5 checks passed
@Yushangjinghong2
Yushangjinghong2 deleted the fix/copilot-integer-budget-observation-20260928 branch October 7, 2026 08:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants