Repository navigation
fix(analyzer): keep Python and Perl syntax from exhausting the shell parser - #778
Merged
Merged
Conversation
…parser
Valid host-language source could make has_bounded_parse_exhaustion report
static_parse_limit. The file was then recorded as partially inspected and
every SKILL.md reference to it became a HIGH AE1 finding. These triggers
are new in 2.12.0:
- Perl `eval BLOCK` (`eval { require $f; 1 } or die $@;`) reached the
shell `eval` command-string reconstruction, and the dynamic `$f`
operand made the reparsed string unresolvable.
- In Python, a string, docstring or comment containing a Markdown fence,
an unpaired backtick or an apostrophe opened a shell quote or command
substitution that ran to the end of the view. Any file with more than
4096 characters after it was partial.
- Python code that uses a shell wrapper name (`signal.alarm(timeout)`,
`if (timeout > 0 ...`) was reparsed as a `timeout` command string, and
a `$0` in a comment was treated as a runtime-selected command whose
operands ran on into later code.
Recognize Perl `eval BLOCK` in .pl files: a clause-initial `eval` followed
by `{` is not a string reparse. `{` already starts a new clause, so the
statements in the block are still scanned. String eval (`eval $code`,
`eval "..."`, `eval qq{...}`) and shell `eval` are unchanged.
For a complete .py module that ast.parse accepts (no non-Python shebang,
no lone CR, within the AST size bound), compute string, f-string and
comment token spans with tokenize and verify them against the source.
Valid Python code outside those tokens contains no shell quote or
expansion opener, so:
- an unresolved shell word is charged only up to the first code-level
word break after it, which ends the string or comment token that
opened the quote, instead of the end of the view;
- a command-string clause (eval, sh -c, wrappers) that begins in Python
code is not reparsed as shell;
- a runtime-selected command inside a comment is rechecked against that
comment alone, because no shell ever receives a comment.
Spans are computed lazily, only when one of these checks would otherwise
fail closed, and they honor the runtime deadline. Fragments, malformed
source, shell-shebang polyglots, a string or comment that itself exceeds
the command-word budget, the destructive rm, brace-expansion and printf
checks, runtime commands built from strings, and all shell and Markdown
behavior keep their existing conservative results. Rule findings are
unchanged.
Refs #694
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995
commented
Oct 6, 2026
rng1995
left a comment
Collaborator
Author
There was a problem hiding this comment.
Review: the ownership model is sound. Requiring a full ast.parse, verifying spans against the source, and bounding at the first code-level break rather than at the token end are the right conservative choices. The Perl eval BLOCK gate on the exact bare word keeps quoted "eval {" in shell strings failing closed. I checked tab-indented, coding-cookie, Python 2 (falls back) and invalid-escape sources. Two P3 items, no blockers.
The module parse that proves Python token ownership repeated every compiler warning the runner's own parse had already reported, such as an invalid `"\d"` escape. Ordinary skill code printed duplicate warnings, and under `-W error::SyntaxWarning` the parse raised and silently dropped ownership, so the file fell back to static_parse_limit. Run that parse with warnings ignored. `warnings.catch_warnings` swaps the process-global filter list, and analyzer nodes run on concurrent graph worker threads, where two such blocks that exit out of order leave their `ignore` filter installed for the whole process. Serialize the block with a module lock. Pin the existing result for a module that starts with a UTF-8 BOM: the file cache decodes with utf-8, not utf-8-sig, so a leading U+FEFF fails the parse and ownership stays unproven. That gap belongs with a separate decoding fix. Refs #694 Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Valid Perl and Python source makes
has_bounded_parse_exhaustion(static_patterns_tool_misuse) reportstatic_parse_limit. The file is then recorded as partially inspected, and everySKILL.mdreference to it becomes a HIGHAE1, which blocks downstream signing gates even though the content is harmless. All triggers below are new in 2.12.0 (2.11.2 scans the same files complete).This PR recognizes Perl
eval BLOCK, and gives Python string, f-string and comment tokens ownership of their bytes when the whole module is proven to be valid Python. Refs #694 (Python part; the Rust, JavaScript, PowerShell and Markdown-table cases there are not addressed).Customer impact
Reported internally against SkillSpector 2.12.0 (via SkillEvaluator 1.6.0). Two internal skills with larger Python/Perl helper scripts fail their release gate with a critical "scan did not execute reliably" result.
static_parse_limitNeither skill needs to change.
Root cause
eval { require $f; 1 } or die "x: $@";_command_string_from_clausetreats a clause-initialevalas shelleval. The dynamic$foperand makes the reparsed string unresolvable (None), which returns exhaustion immediately.md = "\n```\n"followed by more than 4 KB of codeunresolved_end - start > _SHELL_COMMAND_WORD_CHARSwithunresolved_endat the end of the view.helper'sget_token(),# ... the caller's path)signal.alarm(timeout),if (timeout > 0 ...timeoutis reparsed as a command string.# Invoke with $0 = the launcher's pathFix
Perl. With
file_type == "perl", a clause-initial bareevalfollowed by{is not a string reparse.{already starts a new clause, so the statements in the block are still scanned.eval $code,eval "...",eval qq{...}, quoted"eval {"inside a shell string, and any shell-fileeval.eval { `$TOOL -rf /` },eval { qx(sh -c "$cmd") }and brace-expansion inside the block;eval { system("rm -rf /") }still produces TM1.Python. A lazy
_PythonSourceOwnershipis used only for a complete.pyview that:ast.parseaccepts (tokenize alone accepts$and backticks);MAX_PYTHON_AST_SOURCE_CHARS;Its string, f-string/t-string (outermost) and comment spans come from
tokenizeand are verified against the source. Valid Python code outside those tokens contains no shell quote or expansion opener, so:eval,sh -c, wrappers) that begins in Python code is not reparsed as shell;# $TOOL -rf /stays partial.Unchanged:
parsed.limited, the printf, destructive-rm, brace-expansion and root-glob checks, runtime commands built from strings, all shell and Markdown behavior, and every rule finding fromanalyze().The new section in
docs/ANALYSIS_RESOURCE_BOUNDS.mdrecords the ownership contract.Tests
tests/nodes/analyzers/test_host_language_shell_ownership.pyhas 56 tests; 32 of them fail onmain, and all fail-closed controls pass on both.end="", implicit concatenation,timeoutwrapper names, comment$0`. Also CRLF, no trailing newline, non-ASCII, and exact span offsets.#!/bin/sh, lone CR, NUL,complete_context=False, a single literal/adjacent literal/comment over 4096 chars, a comment with$TOOL -rf /, and a.shfile with an unmatched backtick plus a long tail (ledger PARTIAL).graph.invoke): aSKILL.mdreferencing a Python and a Perl helper is complete with no AE1; the over-budget negative control stays partial with AE1.make lintandmake format-checkpass. The unit suite (pytest -m "not integration and not provider" tests/) gives 10343 passed, 15 skipped, 4 xfailed, 0 failed.main, the only changes are the removedstatic_parse_limitrows and their AE1s; no rule finding changes.Residual
eval { `sh -c "$cmd"` }was partial onmainonly incidentally, through the eval-operand path. It is now complete, the same asmy $x = `sh -c "$cmd"`already is. The underlying gap is that command strings inside backtick substitutions are not recognized as clauses; it predates this PR.utf-8, notutf-8-sig, so the leading U+FEFF failsast.parse, andmainalready reports such a file assyntax_errorin the AST analyzers. A test pins today's conservative result; stripping the BOM belongs with a separate decoding fix.🤖 Generated with Claude Code