Skip to content

fix(parsers): rst_parser בודק את הקלט בכניסה, וברירת המחדל של max_sections מיושרת ל-MAX_SECTIONS (#3421, #3420) - #3436

Closed
amirbiron wants to merge 1 commit into
mainfrom
claude/gracious-einstein-sevk8p-3421
Closed

amirbiron wants to merge 1 commit into
mainfrom
claude/gracious-einstein-sevk8p-3421

Conversation

@amirbiron

@amirbiron amirbiron commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

✨ תיאור קצר

📦 שינויים עיקריים

  • קוד (Backend)
  • בוט טלגרם
  • מסד נתונים/מיגרציות
  • תיעוד (docs/)
  • DevOps/CI/CD

פירוט נקודות (רשימת תבליטים):

  • יישור בדיקת הכניסה של rst_parser.parse_document לקלט שאינו מחרוזת #3421 — בדיקת הכניסה. פונקציה אחת לשני הפארסרים, doc_sections.require_str, שמרימה TypeError עם אותן מילים בדיוק (parse_document expects str, got NoneType). לא שתי שורות זהות בשני מודולים — זה מה שנסחף בתיקון הבא (R6). ב-rst_parser הוסר ה-(text or "") שהפך None למסמך ריק בשקט. מי נשען על הצורה הישנה — נבדק, אף אחד: docs_handlers מעביר content = res.get("content") or "" (תמיד מחרוזת), הסורק מעביר טקסט מפוענח, והסקריפטים קוראים קבצים. ותשובה לשאלה השנייה באישו: docs_get_section אינו ממיר את ה-TypeError לתשובת error — בדיוק כפי שלא המיר אותה מ-md_parser עד היום: זה "חוזה נשבר" ולא "הקלט נדחה", ועטיפה הייתה widened-exception-scope. ההערה בקוד מעודכנת.
  • refactor: לשקול יישור ברירת המחדל של max_sections ב-rst_parser.parse_document #3420 — ההכרעה: יישור, לא תיעוד כהחלטה סגורה. הסיבה: הכלל בריפו הוא fail-closed (_declares_write, _ceiling.py), שני הצרכנים בייצור כבר העבירו תקרה ולכן ברירת המחדל משפיעה רק על מי ששכח, והמדידה הראתה שהמחיר על קלט אמיתי הוא אפס. MAX_SECTIONS עבר ל-services.doc_sections (המודול המשותף שכבר מחזיק את TooManySections), מיוצא מחדש משני הפארסרים (אותו אובייקט), והוא ברירת המחדל בשניהם; None מכבה במפורש.
  • המדידה (על כל 208 קובצי ה-RST ב-docs/, לפני השינוי, נספר כמו שהתקרה סופרת): 1,384 סקשנים בסך הכול; הקובץ העשיר ביותר, docs/mcp-server.rst, נושא 50 מול תקרה של 50,000 — אחד לאלף; ואף קובץ אינו נחסם תחת MAX_SECTIONS. פרסור כולם: 0.05 שניות.
  • אפס-דיף על docs_get_section: scripts/docs_section_zero_diff.py על אותו קורפוס (--corpus מצביע לאותו עץ לשתי ההרצות) — הקוד הישן ב-worktree מנותק על origin/main והקוד החדש: 5,931 רשומות, sha256 זהה (9f12d0ed…5cba9), cmp בייט-בייט זהה. (הרצה ראשונה הראתה 196 רשומות שונות — כי ערכתי את docs/mcp-server.rst ו-whats-new.rst בין שתי ההרצות, והקורפוס הוא הקבצים החיים; ההרצה החוזרת של הישן על הקורפוס העדכני זהה. מציין את זה כי זה בדיוק מה שהדוקסטרינג של הסקריפט מזהיר מפניו.)
  • תוצאה של היישור ב-mcp_server/docs_handlers.py: הכלי אינו מעביר תקרה לאף פארסר — parse_kwargs וה-parser is rst_parser נעלמו (ההערה שם נימקה את ההעברה המפורשת ב"ברירת המחדל של rst_parser היא None", וזה כבר לא נכון; והנימוק שלה עצמה נגד העברה ל-Markdown — "עותק שני של החלטה שהפארסר כבר הכריע" — חל עכשיו על שניהם). "max" בסירוב הוא doc_sections.MAX_SECTIONS, המספר שהפרסר באמת השתמש בו, ולא _ceiling.MAX_SYMBOLS ("נכון ל-Markdown רק מפני שהטסט קושר"). הייבוא של _ceiling מהמטפל הוסר.
  • הסורק אינו מושפע, וזה נאמר בטסט: mcp_server/outline_scanners/rst.py ממשיך להעביר max_sections=_ceiling.MAX_SYMBOLS במפורש (המספר של המפה — כותרות ותוויות יחד, Capped), והדוקסטרינג אומר למה גם כשהמספרים שווים. הטסט החדש ב-test_mcp_outline.py מרגל על הקריאה ומקבע את הארגומנט.
  • scripts/measure_md_parse_cost.py: שלוש הצורות רצות עכשיו על ברירת המחדל בשני הפרסרים — כלומר בדיוק כמו הכלי; הצורה העוינת של RST נעצרת על התקרה (outcome: too_many_sections, 19.9MiB, מול 44.0MiB בלי תקרה). מתועד בדוקסטרינג, והסקריפט הורץ מקצה לקצה (10 שניות).
  • תיעוד וקוד שהצהירו על ההבדל, כולם עודכנו: הדוקסטרינגים של שלושת המודולים (כולל ה-.. important:: בשניהם שהצהיר "ברירת המחדל כאן הפוכה"), ההערה ב-_ceiling.py על העותק השני, docs/mcp-server.rst (הפסקה של too_many_sections), ו-docs/whats-new.rst. ההיסטוריה נשארת כתובה ליד הקבוע — למה זה היה None, ולמה זה השתנה.

🧪 בדיקות

  • חדש ב-tests/test_rst_parser.py: None/17/b"x" → TypeError שנוקב בטיפוס (match על שם הטיפוס, כי bytes נפל גם קודם ב-TypeError אחר מתוך .replace); שני הפארסרים מרימים את אותן מילים; ברירת המחדל היא אותו אובייקט doc_sections.MAX_SECTIONS בשניהם (טענת O(1), כמו הטסט המקביל ב-test_md_parser — לא 50,001 סקשנים בכל CI), "MAX_SECTIONS" ב-__all__, ו-None מכבה.
  • נערכו: test_the_two_parsers_export_the_same_names_but_two ← ..._but_one (ההפרש עכשיו {"InconsistentLineEndings"} בלבד); test_the_default_ceiling_is_the_documented_constant מקבע שזה ייצוא-מחדש ולא עותק; ב-test_mcp_docs_handlers.py הטסט שקיבע את ההעברה המפורשת ל-RST הפך ל-test_the_rst_reader_is_capped_by_the_parser_default_and_refuses_above_it (הגלגול השלישי שלו — fix: תקרת הסימבולים נוסעת לתוך פרסור ה-RST במקום לסנן את הפלט #3378 ← fix(mcp): מאגר הקריאות נגזר ממכסת הזיכרון של הקונטיינר, לא ממעבדי המארח (#3391) #3429 ← refactor: לשקול יישור ברירת המחדל של max_sections ב-rst_parser.parse_document #3420, מתועד בדוקסטרינג): passed == {}, ברירת המחדל היא הקבוע, סירוב מעל התקרה דרך functools.partial על הפרסר, "max" הוא הקבוע; test_the_two_refusals_are_mapped_whichever_parser_raised_them חזר ל-partial (עכשיו זה עובד, כי הכלי אינו מעביר ארגומנט שדורס); test_the_two_section_ceilings_are_the_same_number משווה _ceiling.MAX_SYMBOLS ל-doc_sections.MAX_SECTIONS ומקבע ששלושת השמות הם אותו אובייקט.
  • חדש ב-tests/test_mcp_outline.py: test_the_rst_scanner_passes_its_own_ceiling_and_does_not_lean_on_the_parser_default — פין, לפי דרישת refactor: לשקול יישור ברירת המחדל של max_sections ב-rst_parser.parse_document #3420 ("בטסט ולא בהיגיון").
  • על הקוד הישן (worktree מנותק על origin/main = 1a45c28c, עם קובצי הטסטים החדשים): 12 נופלים (שלושת סוגי הקלט × שני הטסטים, ברירת המחדל, הייצוא, הקבוע המשותף, שלושת טסטי המטפל), ואחד עובר — הפין של הסורק, בכוונה ומתועד.
  • מוטציות ב-worktree על הקוד החדש: הסורק בלי הארגומנט ← הפין נופל; המטפל שמעביר תקרה שוב ← test_the_rst_reader_is_capped_by_the_parser_default… נופל; require_str שמחזיר text or "" ← שלושת טסטי הכניסה נופלים.
  • סוויטות: 894 טסטים ב-test_mcp_to_thread, test_mcp_server_build, test_mcp_docs_handlers, test_mcp_outline, test_mcp_repo_backend, test_mcp_primer, test_rst_parser, test_md_parser, test_md_parser_oracle, test_doc_sections, test_docs_headings_carry_no_identifier, test_mcp_logging_visible, test_mcp_analytics_privacy — ירוקים. flake8 מלא (max-line-length=127) על כל הקבצים שנגעתי בהם — נקי. docutils על שני עמודי התיעוד — 0 אזהרות (בלי בנייה מלאה, לפי CLAUDE.md).
  • לא אימתתי: בניית Sphinx מלאה (RTD יתפוס ב-PR); ה-:data:/:func: roles החדשים בדוקסטרינגים מצביעים ל-services.doc_sections, ובדקתי שאין nitpicky ב-conf.py ואין automodule לשלושת המודולים, כלומר הפניה שלא תיפתר אינה מפילה את הבילד.
  • Unit
  • Integration
  • Manual

🧪 בדיקות נדרשות ב‑PR

  • 🔍 Code Quality & Security
  • Unit Tests (3.11)
  • Unit Tests (3.12)

📝 סוג שינוי

  • feat: פיצ'ר חדש
  • fix: תיקון באג
  • docs: שינוי תיעוד בלבד
  • refactor: שינוי קוד ללא שינוי התנהגות
  • perf: שיפור ביצועים
  • chore/ci: תשתית/CI
  • breaking change: שינוי שובר תאימות

✅ צ'קליסט

  • הקוד עוקב אחרי הסגנון (Black/isort/flake8/mypy)
  • בדיקות רצות ועוברות
  • תיעוד עודכן (README/Docs)
  • אם נוספו ג'ובים חדשים (Background Jobs) – לא נוספו
  • אם נוספו/שונו משתני סביבה – לא נוספו
  • אם נוספו/השתנו טוקנים – לא רלוונטי
  • אין סודות/מפתחות בקוד
  • אין מחיקות מסוכנות/פעולות על root (ראו .cursorrules)
  • הודעת הקומיט תואמת Conventional Commits (ע"פ הטבלה)
  • CHANGELOG עודכן אם נדרש — docs/whats-new.rst עודכן
  • כל ה‑Required Checks לעיל ירוקים — ייבדק ב-CI
  • צילום/וידאו UI מצורף אם רלוונטי — לא רלוונטי
  • עיינתי במסמכי אתר התיעוד — נתיב: AI-MAP.md, docs/mcp-server.rst ("שני סירובים שמגיעים מהפרסור עצמו"), docs/development/scripts.rst (הסקריפט שמודד את עלות הפרסור), docs/doc-authoring.rst, docs/versioning-stable-anchors.rst | המשפט: "ו-RST דרך התקרה שהכלי מעביר במפורש (_ceiling.MAX_SYMBOLS, מאז fix(mcp): מאגר הקריאות נגזר ממכסת הזיכרון של הקונטיינר, לא ממעבדי המארח (#3391) #3429), כי ברירת המחדל של rst_parser היא ללא תקרה" — הפסקה שהשתנתה
  • לא נדרש עיון — התנאי התקיים (התנהגות מתועדת)

דפוסי באגים שנקראו ומה שקבעו בקוד: CORE-PATTERNS.md U3 + bugbot-rules/external-input-isinstance.md — isinstance(text, str) לפני .replace על ערך שמגיע מחוץ לתהליך, וזה כל #3421; bugbot-rules/silent-fallback-to-worse-path.md — ה-(text or "") היה "המחבוא הנפוץ": תוצאה ריקה במקום חריגה, ש"לא נמצא" ו"נשבר" נראים בה זהים — הוסר; RECURRING-PATTERNS.md R6 — grep על "expects str" לפני הכתיבה: עותק אחד ב-md_parser ← אוחד ל-require_str, ו-MAX_SECTIONS עבר למודול המשותף במקום להיות מוקלד פעם שלישית; bugbot-rules/state-record-without-state-change.md — כל תיאור של ברירת המחדל (שלושה דוקסטרינגים, הערה ב-_ceiling, שני עמודי תיעוד, הערת המטפל) עודכן יחד עם השינוי, כדי שלא תישאר רשומה שמתארת מצב שאינו קיים; bugbot-rules/line-number-coupling.md — בלי מספרי שורות בתיעוד; claude-md-snippets/testing.md §5 + TESTING-PATTERNS.md T2 — הטסטים הורצו על הקוד הישן ונפלו, והמוטציות מתועדות; bugbot-rules/widened-exception-scope.md — לא הורחב אף except, וה-TypeError נשאר לא-נתפס במטפל בכוונה.

🧩 השפעות/סיכונים

  • שינוי התנהגות על מסלול חי, במפורש: קורא של rst_parser.parse_document שמעביר ערך שאינו מחרוזת מקבל TypeError במקום מסמך ריק; קורא שאינו מעביר תקרה מקבל תקרה של 50,000 סקשנים. בקוד הריפו אין אף קורא משני הסוגים (נבדק: המטפל, הסורק, הסקריפטים והטסטים), ואפס-הדיף מוכיח שהכלי הציבורי מחזיר בדיוק אותן תשובות.
  • מה לא השתנה: המספר (50,000), _ceiling.MAX_SYMBOLS ומה שהסורק מעביר, md_parser (אותה בדיקה, אותה ברירת מחדל, רק דרך המודול המשותף), ותשובות docs_get_section על כל 208 העמודים.

🔗 קישורים

🧯 סיכון / החזרה לאחור (Rollback)

  • revert של הקומיט היחיד מחזיר את שני הפארסרים למצבם הקודם, כולל ההעברה המפורשת מהמטפל. אין מיגרציה ואין משתנה סביבה.

🤖 Generated with Claude Code

https://claude.ai/code/session_01SfJTSpDAhDr2yhtmpFkwTx


Generated by Claude Code

Review in cubic

…tions מיושרת ל-MAX_SECTIONS (#3421, #3420)

שני הפארסרים מתועדים כבני-החלפה, ועל קלט שאינו מחרוזת הם התנהגו אחרת:
rst_parser.parse_document(None) החזיר מסמך ריק בשקט (מפה ריקה שמתחזה למפה
של קובץ בלי כותרות), ו-17 נפל ב-AttributeError גולמי מתוך .replace, בעוד
md_parser זרק TypeError שאומר מה התקבל. וברירת המחדל של max_sections הייתה
None ב-rst_parser ו-MAX_SECTIONS ב-md_parser, כלומר קורא ששכח להעביר תקרה
קיבל הגנה מהפארסר האחד ולא מהשני.

- #3421: בדיקת כניסה אחת לשניהם, doc_sections.require_str, שמרימה TypeError
  עם אותן מילים ("parse_document expects str, got NoneType"). אף קורא לא
  נשען על הצורה הישנה: המטפל, הסורק והסקריפטים מעבירים תמיד מחרוזת.
  docs_get_section אינו תופס את החריגה, כמו שלא תפס אותה מ-md_parser:
  "חוזה נשבר" ולא "הקלט נדחה", ו-content שם תמיד מחרוזת.
- #3420, ההכרעה: יישור. MAX_SECTIONS עובר ל-services.doc_sections (המודול
  המשותף), מיוצא מחדש משני הפארסרים, והוא ברירת המחדל בשניהם. נמדד לפני
  השינוי על כל 208 קובצי ה-RST ב-docs/: 1,384 סקשנים בסך הכול, הקובץ העשיר
  ביותר (docs/mcp-server.rst) נושא 50 מול תקרה של 50,000, אחד לאלף, ואף
  קובץ אינו נחסם. אפס-דיף על docs_get_section: 5,931 רשומות זהות בית-בית
  לפני ואחרי על אותו קורפוס.
- בעקבות היישור docs_get_section אינו מעביר תקרה לאף פארסר (עד עכשיו העביר
  ל-rst_parser את _ceiling.MAX_SYMBOLS במפורש, כי ברירת המחדל שם הייתה
  None), ו-"max" בסירוב הוא doc_sections.MAX_SECTIONS, המספר שהפרסר באמת
  השתמש בו. הייבוא של _ceiling מהמטפל הוסר.
- outline_scanners/rst.py ממשיך להעביר את _ceiling.MAX_SYMBOLS במפורש
  (המספר של המפה, כותרות ותוויות יחד), וטסט חדש מקבע זאת: מוטציה שמוחקת
  את הארגומנט מפילה אותו.
- scripts/measure_md_parse_cost.py: שלוש הצורות רצות עכשיו על ברירת המחדל
  בשני הפרסרים, כמו הכלי; מתועד בדוקסטרינג, והסקריפט רץ מקצה לקצה.

טסטים: 12 טסטים חדשים או שנערכו נופלים על origin/main; הפין של הסורק עובר
שם בכוונה. שלוש מוטציות ב-worktree (הסורק בלי הארגומנט, המטפל שמעביר תקרה
שוב, require_str שממיר None ל-"") מפילות כל אחת את הטסט שלה.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SfJTSpDAhDr2yhtmpFkwTx
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@qodo-code-review

Copy link
Copy Markdown

ⓘ Qodo reviews are paused because your trial has ended. Ask your workspace admin to add credits to resume reviews. Manage billing

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry @amirbiron, you've used your own review budget of 250,000 diff characters for the last 7 days.

You can request another review in 6 hours and 5 minutes by commenting @sourcery-ai review. Upgrade to get a review now.

@coderabbitai

coderabbitai Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

Next included review available in 3 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 4250d5f5-3c30-4260-992d-b332d73b84e9

📥 Commits

Reviewing files that changed from the base of the PR and between 1a45c28 and 30f42c6.

📒 Files selected for processing (13)
  • docs/mcp-server.rst
  • docs/whats-new.rst
  • mcp_server/docs_handlers.py
  • mcp_server/outline_scanners/_ceiling.py
  • mcp_server/outline_scanners/rst.py
  • scripts/measure_md_parse_cost.py
  • services/doc_sections.py
  • services/md_parser.py
  • services/rst_parser.py
  • tests/test_mcp_docs_handlers.py
  • tests/test_mcp_outline.py
  • tests/test_md_parser.py
  • tests/test_rst_parser.py

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

🧯 Dangerous deletes guard report

Policy: see .cursorrules — dangerous deletions are blocked unless wrapped safely.

Summary:

  • Flagged findings (blocking): 0
    0
  • Excluded matches (not blocking): 15
  • Total matches (all files): 128

Flagged findings (file:line:snippet):
(none)

Excluded matches (by path pattern)
./webapp/static/js/md_preview.bundle.js.map:4:  "sourcesContent": ["// Markdown-it plugin to render GitHub-style task lists; see\n//\n// https://github.com/blog/1375-task-lists-in-gfm-issues-pulls-comments\n// https://github.com/blog/1825-t … [truncated]
./docs/DOCUMENTATION_GUIDE.md:453:rm -rf _build
./docs/Makefile:24:	rm -rf $(BUILDDIR)
./Dockerfile:42:    rm -rf /var/lib/apt/lists/*
./Dockerfile:121:    rm -rf /var/lib/apt/lists/*
./node_modules/katex/package.json:153:    "build": "rimraf dist/ && mkdirp dist && cp README.md dist && rollup -c --failAfterWarnings && webpack && node update-sri.js package dist/README.md",
./node_modules/katex/src/fonts/Makefile:139:	rm -rf pfa ff otf ttf woff woff2
./node_modules/mermaid/dist/mermaid.js.map:4:  "sourcesContent": ["/**\n* Default values for dimensions\n*/\nconst defaultIconDimensions = Object.freeze({\n\tleft: 0,\n\ttop: 0,\n\twidth: 16,\n\theight: 16\n});\n/**\n* Default values for tr … [truncated]
./node_modules/mermaid/dist/chunks/mermaid.esm/chunk-2M32CCKP.mjs.map:4:  "sourcesContent": ["{\n  \"name\": \"mermaid\",\n  \"version\": \"11.12.0\",\n  \"description\": \"Markdown-ish syntax for generating flowcharts, mindmaps, sequence d … [truncated]
./node_modules/mermaid/dist/chunks/mermaid.esm.min/chunk-4HFYJGYH.mjs.map:4:  "sourcesContent": ["{\n  \"name\": \"mermaid\",\n  \"version\": \"11.12.0\",\n  \"description\": \"Markdown-ish syntax for generating flowcharts, mindmaps, sequen … [truncated]
./node_modules/mermaid/dist/chunks/mermaid.esm.min/chunk-4HFYJGYH.mjs:1:var r={name:"mermaid",version:"11.12.0",description:"Markdown-ish syntax for generating flowcharts, mindmaps, sequence diagrams, class diagrams, gantt charts, git graph … [truncated]
./node_modules/mermaid/dist/chunks/mermaid.core/chunk-KS23V3DP.mjs.map:4:  "sourcesContent": ["{\n  \"name\": \"mermaid\",\n  \"version\": \"11.12.0\",\n  \"description\": \"Markdown-ish syntax for generating flowcharts, mindmaps, sequence  … [truncated]
./node_modules/mermaid/dist/mermaid.min.js:1524:`,"getStyles"),c1e=RQe});var h1e={};dr(h1e,{diagram:()=>NQe});var NQe,f1e=N(()=>{"use strict";$ge();a1e();l1e();u1e();NQe={parser:Fge,db:n1e,renderer:o1e,styles:c1e}});var m1e,g1e=N(()=>{"use  … [truncated]
./node_modules/mermaid/dist/mermaid.min.js.map:4:  "sourcesContent": ["/**\n* Default values for dimensions\n*/\nconst defaultIconDimensions = Object.freeze({\n\tleft: 0,\n\ttop: 0,\n\twidth: 16,\n\theight: 16\n});\n/**\n* Default values fo … [truncated]
./README.md:842:find . -name "__pycache__" -exec rm -rf {} +

@sourcery-ai

sourcery-ai Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Reviewer's Guide

The PR makes RST and Markdown parser contracts consistent by sharing strict string validation and a fail-closed MAX_SECTIONS default, while removing redundant ceiling forwarding from the document handler, preserving the outline scanner’s explicit limit, and updating tests and documentation to verify compatibility.

Sequence diagram for shared parser validation and section limits

sequenceDiagram
    participant Caller
    participant DocsHandler as docs_get_section
    participant Parser
    participant DocSections as doc_sections

    Caller->>DocsHandler: docs_get_section(...)
    DocsHandler->>Parser: parse_document(content)
    Parser->>DocSections: require_str(text)
    alt text is not str
        DocSections-->>Parser: TypeError
        Parser-->>DocsHandler: TypeError
    else text is str
        Parser->>Parser: apply MAX_SECTIONS default
        alt section limit exceeded
            Parser-->>DocsHandler: TooManySections
            DocsHandler-->>Caller: error: too_many_sections, max=MAX_SECTIONS
        else within limit
            Parser-->>DocsHandler: Document
            DocsHandler-->>Caller: section response
        end
    end
Loading

File-Level Changes

Change Details Files
Centralized parser input validation and aligned both parsers on a shared section ceiling default.
  • Added shared require_str validation with consistent TypeError messages for non-string input.
  • Changed RST parsing to reject invalid input instead of silently treating falsy values as empty text.
  • Moved MAX_SECTIONS to doc_sections and re-exported the same object from both parsers.
  • Changed the RST parser default from None to MAX_SECTIONS while preserving explicit None as unlimited parsing.
services/doc_sections.py
services/md_parser.py
services/rst_parser.py
tests/test_rst_parser.py
tests/test_md_parser.py
Simplified document-section handling so parsers own their default protection while preserving explicit scanner limits.
  • Removed parser-specific max_sections forwarding and the _ceiling dependency from docs_get_section.
  • Reported doc_sections.MAX_SECTIONS in too_many_sections responses.
  • Kept the RST outline scanner’s explicit _ceiling.MAX_SYMBOLS argument and added a regression test to enforce it.
  • Updated handler tests for parser defaults, refusal mapping, and reported ceilings.
mcp_server/docs_handlers.py
mcp_server/outline_scanners/rst.py
mcp_server/outline_scanners/_ceiling.py
tests/test_mcp_docs_handlers.py
tests/test_mcp_outline.py
Updated parser documentation, measurement tooling, and release notes to reflect the unified behavior and its compatibility rationale.
  • Revised module and API docstrings to remove the former parser-default discrepancy and document shared validation and ceiling ownership.
  • Updated the parser-cost measurement script to exercise default ceiling behavior.
  • Recorded the change in MCP documentation and whats-new.
scripts/measure_md_parse_cost.py
docs/mcp-server.rst
docs/whats-new.rst

Possibly linked issues


Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@github-actions

Copy link
Copy Markdown
Contributor

⏱️ Performance report

(No performance test durations collected. Mark tests with @pytest.mark.performance.)

@github-actions

Copy link
Copy Markdown
Contributor

📖 Documentation Preview

The documentation has been built successfully!

To view locally:

  1. Download the artifacts
  2. Extract the zip file
  3. Open index.html in your browser

@codecov

codecov Bot commented Sep 20, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@amirbiron

Copy link
Copy Markdown
Owner Author

נכנס דרך #3443

@amirbiron amirbiron closed this Sep 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants