Skip to content

Integration proposal: ast-guard as static code analysis policy for reward hacking detection #375

Description

@Nick-is-building

Hey team, following up from the conversation on r/AI_Agents where @Big_Wonder7834 suggested raising a PR to integrate ast-guard into failproofai's coding harnesses.

What ast-guard does:
ast-guard is a deterministic reward hacking detector for LLM-generated Python code. It analyzes code structurally via Python's AST before execution. No LLM, zero dependencies, pure static analysis. It catches hardcoded lookup tables, if/else chains for test inputs, forbidden system calls (eval, exec, os, subprocess), obfuscation attempts, and unexplained complexity collapses.

Repo: https://github.com/Nick-is-building/ast-guard

How it could fit into failproofai:
I see ast-guard working as a custom policy that hooks into PreToolUse events for Write and Edit tool calls. When an agent writes or modifies a Python file, the policy would compare the original file content against the new content using ast-guard's scan() function and return deny() if structural cheating is detected, allow() with context if warnings are found, or allow() if the code is clean.
This would add a layer that failproofai currently doesn't cover: not just what the agent is doing (running sudo, deleting files, leaking secrets), but whether the code the agent produces is structurally honest.

What I'd need from you:
Some guidance on the best integration approach. I see two options: a standalone custom policy file that users drop into .failproofai/policies/, or a deeper integration as a built-in policy. Happy to start with option 1 and iterate from there.
Would also be great to know if there are any conventions for policies that shell out to Python (since ast-guard is Python-based and failproofai is Node/TypeScript).
Looking forward to collaborating on this.

Activity

  1. Nick-is-building commented on May 23, 2026

    @Nick-is-building
    Author

    Prototype: ast-guard policy for failproofai

    Following up with a concrete implementation sketch. I studied the custom policy API and built a working prototype that shows exactly how ast-guard would integrate.

    Policy Code

    // ast-guard-policy.js
    // Drop into .failproofai/policies/ or install with:
    // failproofai policies --install --custom ./ast-guard-policy.js
    
    import { customPolicies, allow, deny, instruct } from "failproofai";
    import { execSync } from "child_process";
    import { readFileSync, existsSync } from "fs";
    
    customPolicies.add({
      name: "ast-guard-reward-hacking",
      description:
        "Deterministic reward hacking detection for LLM-generated Python code via AST analysis",
    
      match: {
        events: ["PreToolUse"],
      },
    
      fn: async (ctx) => {
        // Only intercept Write and Edit on .py files
        if (!["Write", "Edit"].includes(ctx.toolName ?? "")) return allow();
    
        const filePath = ctx.toolInput?.file_path ?? "";
        if (!filePath.endsWith(".py")) return allow();
    
        const newContent = ctx.toolInput?.content ?? "";
        if (!newContent.trim()) return allow();
    
        // Read the original file content (empty string if new file)
        let originalContent = "";
        try {
          if (existsSync(filePath)) {
            originalContent = readFileSync(filePath, "utf-8");
          }
        } catch {
          originalContent = "";
        }
    
        // Skip if this is a brand new file (no original to compare against)
        if (!originalContent.trim()) return allow();
    
        // Run ast-guard scan via Python
        try {
          const escapedOriginal = JSON.stringify(originalContent);
          const escapedNew = JSON.stringify(newContent);
    
          const script = `
    import json
    from ast_guard import scan
    result = scan(${escapedOriginal}, ${escapedNew}, mode="strict")
    print(json.dumps({"verdict": result["verdict"], "findings": [
        f["explanation"] for c in result["checks"].values()
        for f in c.get("findings", [])
    ]}))
    `;
          const output = execSync(`python3 -c '${script.replace(/'/g, "'\\''")}'`, {
            timeout: 10000,
            encoding: "utf-8",
          }).trim();
    
          const result = JSON.parse(output);
    
          if (result.verdict === "CRITICAL") {
            const reasons = result.findings.slice(0, 3).join(" | ");
            return deny(
              `[ast-guard] Reward hacking detected: ${reasons}. ` +
              `The generated code shows structural cheating patterns. ` +
              `Rewrite the solution using legitimate algorithms.`
            );
          }
    
          if (result.verdict === "WARNING") {
            const reasons = result.findings.slice(0, 3).join(" | ");
            return instruct(
              `[ast-guard] Suspicious patterns found: ${reasons}. ` +
              `Review your approach — avoid hardcoded lookup tables, ` +
              `unexplained complexity drops, or unnecessary system calls.`
            );
          }
    
          return allow();
        } catch (err) {
          // Fail-open: if ast-guard errors, don't block the agent
          return allow();
        }
      },
    });

    How it works

    1. Intercepts PreToolUse events for Write and Edit on .py files only
    2. Reads the original file from disk before the agent overwrites it
    3. Runs ast_guard.scan() comparing original vs. new content
    4. Returns deny() on CRITICAL (blocks the write + tells the agent to rewrite), instruct() on WARNING (adds context without blocking), allow() on CLEAN
    5. Fails open — if ast-guard isn't installed or errors out, the agent continues normally

    Prerequisites

    The user needs ast-guard cloned or pip-installable on the machine where failproofai runs. Since ast-guard has zero dependencies (pure stdlib), this is just:

    git clone https://github.com/Nick-is-building/ast-guard.git
    cd ast-guard && pip install -e .

    Open questions for the team

    1. Python bridge: This prototype shells out to Python via execSync. Is there a preferred pattern in failproofai for policies that need non-Node runtimes? I saw that errors are fail-open by default, which aligns with ast-guard's design.

    2. Performance: ast-guard runs in <50ms on typical files (pure AST parsing, no I/O). The 10-second timeout is a safety net, not an expected duration. Is there a recommended timeout ceiling for custom policies?

    3. Scope: Should this live as a community-contributed custom policy in examples/, or would you consider it for the built-in policy set given that reward hacking detection is a gap in the current 39 policies?

    Happy to open a PR once we align on the approach.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions