Bug Report: Agent continues indefinitely with regression tests after all features complete
🐛 Problem Description
The autonomous agent does not automatically stop when all features have been implemented and marked as passing. Instead, it continues running indefinitely, performing endless regression testing sessions that consume tokens without making any progress on the project.
Discovered During
- Project: CookAi (React Native recipe app)
- Runtime: Overnight run (~12+ hours)
- Features: 134 features total
- Completion Status: All 134 features marked as passing after Session 7
- Sessions 8-20: Agent ran 12+ additional sessions doing only regression tests
- Token Waste: Estimated 120,000+ tokens consumed on unnecessary regression testing
📊 Current Behavior
When all features are complete (passing = total and pending = 0), the agent:
- ✅ Calls
feature_get_next tool
- ✅ Receives response:
{"error": "All features are passing! No more work to do."}
- ❌ Ignores this signal and continues anyway
- ❌ Runs
feature_get_for_regression (gets 3 random passing features)
- ❌ Tests those features via Playwright
- ❌ Writes regression test report to
claude-progress.txt
- ❌ Returns
status = "continue"
- ❌ Loop repeats indefinitely (Steps 1-7)
Evidence from Logs
From claude-progress.txt after Session 7 completed all features:
## Session 8 - Verification Agent
### Feature Statistics
- **Total features: 134**
- **Passing: 134**
- **In Progress: 0**
- **Completion: 100.0%**
### Regression Tests Performed
1. Feature #70: Numeric fields reject letters ✅
2. Feature #46: Dropdown options from database ✅
3. Feature #18: Recipe edit button navigates correctly ✅
### No Code Changes Required
All features working as expected - no fixes needed this session.
## Session 9 - Verification Agent
[... identical regression testing ...]
## Session 10 - Verification Agent
[... identical regression testing ...]
[Sessions 11-20 continue with identical pattern]
🎯 Expected Behavior
When all features are complete, the agent should:
- ✅ Call
feature_get_next tool
- ✅ Receive
{"error": "All features are passing! No more work to do."}
- ✅ Recognize completion signal
- ✅ Write final summary to progress file
- ✅ Optionally perform one final stability check
- ✅ Commit final state
- ✅ Display completion message and EXIT gracefully
- ✅ Stop the agent loop to prevent token waste
🔍 Root Cause Analysis
After investigating the codebase, we identified three issues:
1. Missing Completion Check in agent.py
File: agent.py (Lines 167-212)
while True:
iteration += 1
# Check max iterations
if max_iterations and iteration > max_iterations:
break # ✅ This works
# ... run session ...
# Handle status
if status == "continue":
# ❌ NO CHECK FOR COMPLETION HERE!
print(f"\nAgent will auto-continue in {AUTO_CONTINUE_DELAY_SECONDS}s...")
await asyncio.sleep(AUTO_CONTINUE_DELAY_SECONDS)
Problem: The loop only breaks on max_iterations, never on completion.
2. No Completion Detection Function
File: progress.py
The file has count_passing_tests() but no function to determine if a project is complete.
3. Missing Instructions in Prompt Template
File: .claude/templates/coding_prompt.template.md
The prompt tells the agent to call feature_get_next, but doesn't explain what to do when it returns {"error": "All features are passing! No more work to do."}.
The agent interprets "no feature to work on" as "do regression tests instead" rather than "project is complete, stop working".
💡 Proposed Solution
We have implemented a fix (untested) with three components:
Fix 1: Add Completion Detection to agent.py
Location: agent.py Lines 199-208 (after status handling)
# Handle status
if status == "continue":
# Check if project is complete (all features passing)
if is_project_complete(project_dir):
print("\n" + "=" * 70)
print(" 🎉 ALL FEATURES COMPLETE!")
print("=" * 70)
print("\n✅ All features have been implemented and verified.")
print("✅ Project is production-ready!")
print("\nThe agent will now stop to avoid unnecessary token usage.")
break # Exit the loop - project is complete
print(f"\nAgent will auto-continue in {AUTO_CONTINUE_DELAY_SECONDS}s...")
print_progress_summary(project_dir)
await asyncio.sleep(AUTO_CONTINUE_DELAY_SECONDS)
Import added:
from progress import print_session_header, print_progress_summary, has_features, is_project_complete
Fix 2: Add Completion Helper Function to progress.py
Location: progress.py Lines 93-109
def is_project_complete(project_dir: Path) -> bool:
"""
Check if all features are passing (project is complete).
Args:
project_dir: Directory containing the project
Returns:
True if all features are passing, False otherwise
"""
passing, in_progress, total = count_passing_tests(project_dir)
# Project is complete if:
# 1. There are features (total > 0)
# 2. All features are passing (passing == total)
# 3. No features are in progress
return total > 0 and passing == total and in_progress == 0
Fix 3: Update Prompt Template
Location: .claude/templates/coding_prompt.template.md (after line 102)
Added explicit instructions:
**IMPORTANT:** If `feature_get_next` returns:
```json
{"error": "All features are passing! No more work to do."}
Then the project is complete! Take these steps:
-
✅ Write final summary to claude-progress.txt:
- Note that all features are implemented
- List any remaining recommendations
- Congratulate on completion
-
✅ Perform final verification (optional):
- Test 2-3 core features to ensure app is stable
- Check for console errors
- Verify app is production-ready
-
✅ Commit final state:
git add -A
git commit -m "chore: All features complete - project production-ready"
-
✅ End session gracefully - Do NOT continue with more regression tests
## ✅ Benefits of Proposed Solution
- 🎯 **Automatic stopping** when all features are complete
- 💰 **Prevents token waste** on endless regression loops
- 📢 **Clear completion message** for users
- 🔄 **Backward compatible** - existing behavior unchanged for incomplete projects
- 🛡️ **Works alongside** existing `--max-iterations` safety mechanism
- 🤖 **Agent understands** completion signal from prompt instructions
## ⚠️ Testing Status
**IMPORTANT:** This fix has been implemented locally but **NOT yet tested** in a real agent run.
We have verified:
- ✅ `is_project_complete()` function works correctly on completed project
- ✅ Returns `True` for project with 134/134 features passing
- ✅ Code compiles without errors
We have NOT verified:
- ❌ Agent actually stops when completion is detected
- ❌ Prompt instructions are followed by the agent
- ❌ No unintended side effects during normal operation
- ❌ Behavior with partially complete projects
**Testing needed:**
1. Run agent on a project with all features already complete
2. Verify it stops after 1 session instead of continuing indefinitely
3. Run agent on a project with pending features
4. Verify it continues normally until all features are done
## 📁 Implementation Details
Branch: `fix/endless-regression-loop`
Commit: `3c67471`
**Changed Files:**
- `agent.py` (+11 lines)
- `progress.py` (+17 lines)
- `.claude/templates/coding_prompt.template.md` (+27 lines)
Total changes: 3 files, 56 insertions(+), 2 deletions(-)
## 🔗 Related Information
**MCP Tool Response:** `mcp_server/feature_mcp.py` Line 165
```python
if feature is None:
return json.dumps({"error": "All features are passing! No more work to do."})
The MCP tool already provides the completion signal - the agent just needs to recognize and act on it.
📝 Additional Context
System Configuration:
- OS: Linux 6.14.0-37-generic
- Python: 3.x with venv
- Model: claude-opus-4-5-20251101 (as per DEFAULT_MODEL in autonomous_agent_demo.py)
- Agent Mode: Standard (not YOLO mode)
Project Structure:
testapp/
├── features.db # SQLite with 134 features (all passing)
├── claude-progress.txt # 1549 lines (Sessions 1-20+)
├── prompts/
│ ├── app_spec.txt
│ ├── initializer_prompt.md
│ └── coding_prompt.md
└── [generated app files]
Token Consumption Estimate per Regression Session:
- System prompt: ~5,000 tokens
- Playwright commands: ~2,000 tokens
- Screenshot analysis: ~3,000 tokens
- Total: ~10,000 tokens/session
12 unnecessary sessions = ~120,000 wasted tokens
💭 Questions for Maintainer
- Was this behavior intentional? (README says "until completion")
- Should regression testing continue indefinitely as a feature?
- Should there be a
--continuous-regression flag for users who want endless testing?
- Is our proposed solution aligned with the project's architecture?
- Should completion also trigger the N8N webhook notification?
🙋 Can We Help?
We're happy to:
- ✅ Submit a PR with our proposed fix
- ✅ Add comprehensive tests for completion detection
- ✅ Update documentation (README, CLAUDE.md)
- ✅ Test on multiple scenarios before merging
Please let us know if this solution approach makes sense or if you'd prefer a different implementation!
Environment:
Bug Report: Agent continues indefinitely with regression tests after all features complete
🐛 Problem Description
The autonomous agent does not automatically stop when all features have been implemented and marked as passing. Instead, it continues running indefinitely, performing endless regression testing sessions that consume tokens without making any progress on the project.
Discovered During
📊 Current Behavior
When all features are complete (
passing = totalandpending = 0), the agent:feature_get_nexttool{"error": "All features are passing! No more work to do."}feature_get_for_regression(gets 3 random passing features)claude-progress.txtstatus = "continue"Evidence from Logs
From
claude-progress.txtafter Session 7 completed all features:🎯 Expected Behavior
When all features are complete, the agent should:
feature_get_nexttool{"error": "All features are passing! No more work to do."}🔍 Root Cause Analysis
After investigating the codebase, we identified three issues:
1. Missing Completion Check in
agent.pyFile:
agent.py(Lines 167-212)Problem: The loop only breaks on
max_iterations, never on completion.2. No Completion Detection Function
File:
progress.pyThe file has
count_passing_tests()but no function to determine if a project is complete.3. Missing Instructions in Prompt Template
File:
.claude/templates/coding_prompt.template.mdThe prompt tells the agent to call
feature_get_next, but doesn't explain what to do when it returns{"error": "All features are passing! No more work to do."}.The agent interprets "no feature to work on" as "do regression tests instead" rather than "project is complete, stop working".
💡 Proposed Solution
We have implemented a fix (untested) with three components:
Fix 1: Add Completion Detection to
agent.pyLocation:
agent.pyLines 199-208 (after status handling)Import added:
Fix 2: Add Completion Helper Function to
progress.pyLocation:
progress.pyLines 93-109Fix 3: Update Prompt Template
Location:
.claude/templates/coding_prompt.template.md(after line 102)Added explicit instructions:
Then the project is complete! Take these steps:
✅ Write final summary to
claude-progress.txt:✅ Perform final verification (optional):
✅ Commit final state:
git add -A git commit -m "chore: All features complete - project production-ready"✅ End session gracefully - Do NOT continue with more regression tests
The MCP tool already provides the completion signal - the agent just needs to recognize and act on it.
📝 Additional Context
System Configuration:
Project Structure:
Token Consumption Estimate per Regression Session:
12 unnecessary sessions = ~120,000 wasted tokens
💭 Questions for Maintainer
--continuous-regressionflag for users who want endless testing?🙋 Can We Help?
We're happy to:
Please let us know if this solution approach makes sense or if you'd prefer a different implementation!
Environment:
1e20ba9from Jan 7, 2026)