Repository navigation
Integration proposal: ast-guard as static code analysis policy for reward hacking detection #375
Description
Activity
Prototype: ast-guard policy for failproofai
Following up with a concrete implementation sketch. I studied the custom policy API and built a working prototype that shows exactly how ast-guard would integrate.
Policy Code
// ast-guard-policy.js // Drop into .failproofai/policies/ or install with: // failproofai policies --install --custom ./ast-guard-policy.js import { customPolicies, allow, deny, instruct } from "failproofai"; import { execSync } from "child_process"; import { readFileSync, existsSync } from "fs"; customPolicies.add({ name: "ast-guard-reward-hacking", description: "Deterministic reward hacking detection for LLM-generated Python code via AST analysis", match: { events: ["PreToolUse"], }, fn: async (ctx) => { // Only intercept Write and Edit on .py files if (!["Write", "Edit"].includes(ctx.toolName ?? "")) return allow(); const filePath = ctx.toolInput?.file_path ?? ""; if (!filePath.endsWith(".py")) return allow(); const newContent = ctx.toolInput?.content ?? ""; if (!newContent.trim()) return allow(); // Read the original file content (empty string if new file) let originalContent = ""; try { if (existsSync(filePath)) { originalContent = readFileSync(filePath, "utf-8"); } } catch { originalContent = ""; } // Skip if this is a brand new file (no original to compare against) if (!originalContent.trim()) return allow(); // Run ast-guard scan via Python try { const escapedOriginal = JSON.stringify(originalContent); const escapedNew = JSON.stringify(newContent); const script = ` import json from ast_guard import scan result = scan(${escapedOriginal}, ${escapedNew}, mode="strict") print(json.dumps({"verdict": result["verdict"], "findings": [ f["explanation"] for c in result["checks"].values() for f in c.get("findings", []) ]})) `; const output = execSync(`python3 -c '${script.replace(/'/g, "'\\''")}'`, { timeout: 10000, encoding: "utf-8", }).trim(); const result = JSON.parse(output); if (result.verdict === "CRITICAL") { const reasons = result.findings.slice(0, 3).join(" | "); return deny( `[ast-guard] Reward hacking detected: ${reasons}. ` + `The generated code shows structural cheating patterns. ` + `Rewrite the solution using legitimate algorithms.` ); } if (result.verdict === "WARNING") { const reasons = result.findings.slice(0, 3).join(" | "); return instruct( `[ast-guard] Suspicious patterns found: ${reasons}. ` + `Review your approach — avoid hardcoded lookup tables, ` + `unexplained complexity drops, or unnecessary system calls.` ); } return allow(); } catch (err) { // Fail-open: if ast-guard errors, don't block the agent return allow(); } }, });
How it works
- Intercepts PreToolUse events for Write and Edit on .py files only
- Reads the original file from disk before the agent overwrites it
- Runs ast_guard.scan() comparing original vs. new content
- Returns deny() on CRITICAL (blocks the write + tells the agent to rewrite), instruct() on WARNING (adds context without blocking), allow() on CLEAN
- Fails open — if ast-guard isn't installed or errors out, the agent continues normally
Prerequisites
The user needs ast-guard cloned or pip-installable on the machine where failproofai runs. Since ast-guard has zero dependencies (pure stdlib), this is just:
git clone https://github.com/Nick-is-building/ast-guard.git cd ast-guard && pip install -e .
Open questions for the team
-
Python bridge: This prototype shells out to Python via execSync. Is there a preferred pattern in failproofai for policies that need non-Node runtimes? I saw that errors are fail-open by default, which aligns with ast-guard's design.
-
Performance: ast-guard runs in <50ms on typical files (pure AST parsing, no I/O). The 10-second timeout is a safety net, not an expected duration. Is there a recommended timeout ceiling for custom policies?
-
Scope: Should this live as a community-contributed custom policy in examples/, or would you consider it for the built-in policy set given that reward hacking detection is a gap in the current 39 policies?
Happy to open a PR once we align on the approach.
Hey team, following up from the conversation on r/AI_Agents where @Big_Wonder7834 suggested raising a PR to integrate ast-guard into failproofai's coding harnesses.
What ast-guard does:
ast-guard is a deterministic reward hacking detector for LLM-generated Python code. It analyzes code structurally via Python's AST before execution. No LLM, zero dependencies, pure static analysis. It catches hardcoded lookup tables, if/else chains for test inputs, forbidden system calls (eval, exec, os, subprocess), obfuscation attempts, and unexplained complexity collapses.
Repo: https://github.com/Nick-is-building/ast-guard
How it could fit into failproofai:
I see ast-guard working as a custom policy that hooks into PreToolUse events for Write and Edit tool calls. When an agent writes or modifies a Python file, the policy would compare the original file content against the new content using ast-guard's scan() function and return deny() if structural cheating is detected, allow() with context if warnings are found, or allow() if the code is clean.
This would add a layer that failproofai currently doesn't cover: not just what the agent is doing (running sudo, deleting files, leaking secrets), but whether the code the agent produces is structurally honest.
What I'd need from you:
Some guidance on the best integration approach. I see two options: a standalone custom policy file that users drop into .failproofai/policies/, or a deeper integration as a built-in policy. Happy to start with option 1 and iterate from there.
Would also be great to know if there are any conventions for policies that shell out to Python (since ast-guard is Python-based and failproofai is Node/TypeScript).
Looking forward to collaborating on this.