Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
-
Updated
Aug 14, 2026 - Python
Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
OpenTelemetry-native SDK for AI agent observability, tracing, evaluation, debugging, governance, and runtime policy enforcement. Framework-agnostic and built for OpenAI Agents, LangGraph, CrewAI, and LLM applications in production.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
Local-first runtime for observable, debuggable AI-agent workflows with slidable autonomy and multi-agent/tool orchestration.
Local open-source dev tool to debug, secure, and evaluate LLM agents. Provides static analysis, dynamic security checks, and runtime monitoring - integrates with Cursor and Claude Code.
Root cause analysis for AI agents. Detects agent loops, retry storms, and optimization opportunities in LangSmith, Langfuse, Arize Phoenix, and OpenTelemetry traces.
Local replay debugger for Browser Use failures with screenshots, model I/O, failed-step timelines, and public-safe HTML exports.
Explain why your agent failed — root-cause debugging, memory attribution, and run divergence for LLM agents.
A real-time observability and debugging layer for AI agents.
AI agents fail like junior teammates, looping on bad ideas, ignoring feedback, and escalating commitment. vstack ports 34 of the most-cited organizational-behavior frameworks so you can diagnose your agents the same way you'd diagnose your team.
ChainWatch is a flight data recorder for multi-step AI systems. It's a CLI-based tool that records every step in an AI decision chain, links them together in order, prevents tampering, and allows you to verify the chain's integrity and replay the full decision flow.
A truth-first visual observatory for understanding, replaying, comparing, and monitoring AI agent runs.
Freeze, rewind, and attribute AI agent failures with replayable public traces.
Failure attribution for agent pipelines — find which span caused the failure and what kind of fix it needs.
Verify that AI agents actually executed API/tool calls they claim.
Preprint paper package — Agent Trajectory Replay for Debugging Tool-Using AI Workflow Regressions (Zenodo DOI 10.5281/zenodo.20073574)
Diff and regression-detect LLM agent execution traces
Android Agent Reliability Runtime A debugging and safety runtime for mobile GUI agents: detect readiness, block unsafe actions, verify progress, diagnose failures, and save reproducible traces.
Add a description, image, and links to the agent-debugging topic page so that developers can more easily learn about it.
To associate your repository with the agent-debugging topic, visit your repo's landing page and select "manage topics."