Repository navigation
feat(mcp): פארסר Markdown ל-docs_get_section, מעל markdown-it-py #3418
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
8 commits
Select commit
Hold shift + click to select a range
a6269b7
feat(mcp): פארסר Markdown ל-docs_get_section, מעל markdown-it-py
claude 0990e80
test(mcp): הטסט על התקרה מבחין בין עצירה בפרסור לסינון אחריו
claude 6dd764f
test(mcp): הסרת קורפוס ה-fixtures — הצורות המחוללות הן הרשת, לא העותק
claude 57cb845
fix(mcp): טבלת GFM אינה כותרת setext, והנעיצה מיושרת ל-myst-parser
claude 2425632
md_parser: תיקוני סקירת הקוד — מרוץ האתחול, התקרה, והמספרים בפרוזה
claude faeed7f
CLAUDE.md: שורת טריגר לדפוס שנתפס בסקירה של #3418
claude 7784b95
CLAUDE.md: ניסוח חדש לשורת הטריגר, ורמז הדדי לשורת import-time
claude 15ab1ae
md_parser: BOM מוסר בכניסה, פער סוג 7 מקובע ולא מושתק, וטענת הייצוא מ…
claude File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,153 @@ | ||
| #!/usr/bin/env python3 | ||
| """משווה את ``services/md_parser`` מול cmark-gfm על **כל** קובצי ה-``.md`` | ||
| בריפו שמעבירים לו. | ||
|
|
||
| **מה מושווה, כדי שהתוצאה לא תיקרא רחבה ממה שהיא:** לכל כותרת ברמת | ||
| המסמך — **הרמה ומספר השורה**. טקסט הכותרת אינו מושווה כאן, כי אצלנו | ||
| הוא המקור הגולמי ומ-cmark חוזר HTML מרונדר. הצד הזה נבדק בנפרד ועל | ||
| תת-קבוצה, ב-``tests/test_md_parser_oracle.py``. | ||
|
|
||
| **למה הוא קיים, ולמה זה לא טסט.** ``tests/test_md_parser_oracle.py`` רץ | ||
| ב-CI על **צורות מחוללות** ולא על קבצים אמיתיים, וזו החלטה: קורפוס | ||
| אמיתי הוא בדיקת שפיות חד-פעמית ולא רשת שתופסת רגרסיה עתידית, והוא היה | ||
| דורש להחזיק בריפו הזה עותק של תוכן שאינו שלו — עותק שמתיישן ברגע | ||
| שהמקור זז. הסקריפט הוא מה שמחליף אותו: אותה השוואה בדיוק, על הקורפוס | ||
| החי, **על פי דרישה**. | ||
|
|
||
| **שלוש הנקודות בזמן שבהן כדאי להריץ אותו:** לפני שלב 2 (האאוטליין | ||
| ל-``.md``), אחרי שדרוג של ``markdown-it-py``, וכשנוגעים ב- | ||
| ``services/md_parser.py``. | ||
|
|
||
| **ואותה השוואה פירושה אותו קוד.** האורקל אינו נכתב כאן מחדש אלא נטען | ||
| מקובץ הטסטים, כי שני עותקים של "מה נחשב הסכמה" היו נסחפים זה מזה — | ||
| והראשון שהיה נשבר הוא זה שרץ לעיתים רחוקות, כלומר הסקריפט. | ||
|
|
||
| שימוש:: | ||
|
|
||
| python scripts/compare_md_parser_to_cmark.py /path/to/amir-bug-patterns | ||
| python scripts/compare_md_parser_to_cmark.py /path/to/repo --out /tmp/report.txt | ||
|
|
||
| דורש ``cmarkgfm``, שנעוץ ב-``requirements/development.txt``. | ||
| """ | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| import argparse | ||
| import importlib.util | ||
| import sys | ||
| from pathlib import Path | ||
|
|
||
| _REPO = Path(__file__).resolve().parent.parent | ||
| # לפני כל ייבוא מ-``services``: הסקריפט רץ כנקודת כניסה עצמאית, | ||
| # ו-``scripts`` אינה חבילה. | ||
| sys.path.insert(0, str(_REPO)) | ||
|
|
||
| from services.doc_sections import InconsistentLineEndings, TooManySections # noqa: E402 | ||
|
|
||
|
|
||
| def _load_oracle(): | ||
| """טוען את האורקל מקובץ הטסטים — הגדרה אחת לשני המריצים.""" | ||
| spec = importlib.util.spec_from_file_location( | ||
| "md_parser_oracle", _REPO / "tests" / "test_md_parser_oracle.py" | ||
| ) | ||
| module = importlib.util.module_from_spec(spec) | ||
| spec.loader.exec_module(module) | ||
| return module | ||
|
|
||
|
|
||
| def main(argv: list[str] | None = None) -> int: | ||
| parser = argparse.ArgumentParser(description=__doc__.split("\n")[0]) | ||
| # הנתיב הוא ארגומנט **חובה** ובלי ברירת מחדל: ברירת מחדל ל-``.`` | ||
| # הייתה גורמת לסקריפט לסרוק את הריפו שהוא עצמו יושב בו, ולדווח | ||
| # "אפס פערים" על קורפוס שאינו הקורפוס. | ||
| parser.add_argument("repo", type=Path, help="שורש הריפו שקובצי ה-.md שלו ייסרקו") | ||
| parser.add_argument("--out", type=Path, default=None, help="קובץ דוח (ברירת מחדל: stdout)") | ||
| args = parser.parse_args(argv) | ||
|
|
||
| root: Path = args.repo.expanduser().resolve() | ||
| if not root.is_dir(): | ||
| parser.error(f"אינו תיקייה: {root}") | ||
|
|
||
| oracle = _load_oracle() | ||
| files = sorted(p for p in root.rglob("*.md") if ".git" not in p.parts) | ||
| if not files: | ||
| parser.error(f"אפס קובצי .md תחת {root} — זה כישלון, לא ריצה ריקה") | ||
|
|
||
| lines: list[str] = [] | ||
| skipped: list[str] = [] | ||
| mismatched = 0 | ||
| compared = 0 | ||
| total_headings = 0 | ||
| for path in files: | ||
| name = path.relative_to(root) | ||
| # **``read_bytes().decode`` ולא ``read_text``.** האחרון פותח את | ||
| # הקובץ במצב טקסט עם universal newlines וממיר כל ``\r`` ל- | ||
| # ``\n`` לפני שהמחרוזת מגיעה לפארסר — כלומר | ||
| # ``InconsistentLineEndings`` לא הייתה יכולה להידלק כאן על שום | ||
| # קובץ, והסקריפט היה מדווח "אפס אי-הסכמות" גם על מחלקת קלט | ||
| # שבורה לגמרי. ``newline=""`` אינו פתרון: הפרמטר נוסף ל- | ||
| # ``Path.read_text`` רק בפייתון 3.13, וה-CI רץ על 3.11 ו-3.12. | ||
| # | ||
| # **וכל קובץ עומד בפני עצמו.** הסקריפט מכוון על ריפו זר, ולכן | ||
| # קובץ מוזר הוא המקרה הצפוי ולא החריג — נפילה עליו הייתה מוחקת | ||
| # גם את התוצאות של כל מה שכבר נסרק, כי הדוח נבנה אחרי הלולאה. | ||
| # **``utf-8`` ולא ``utf-8-sig`` — בכוונה.** הסקריפט הוא בדיקה של | ||
| # הפארסר, ו-``parse_document`` הוא זה שמסיר BOM. אילו הפענוח כאן | ||
| # היה מסיר אותו קודם, רגרסיה בדיוק בהתנהגות הזאת הייתה בלתי | ||
| # נראית מכאן. הייצור מפענח אחרת — ``git_mirror_service. | ||
| # _try_decode_content`` משתמש ב-``utf-8-sig`` — וזה בסדר: שם | ||
| # המטרה היא תוכן נקי, כאן המטרה היא לראות מה הפארסר עושה. | ||
| try: | ||
| text = path.read_bytes().decode("utf-8") | ||
| ours = oracle._ours(text) | ||
| theirs = oracle._oracle_sections(text) | ||
| except UnicodeDecodeError: | ||
| skipped.append(f"⊘ {name} — אינו UTF-8") | ||
| continue | ||
| except OSError as exc: | ||
| skipped.append(f"⊘ {name} — לא ניתן לקריאה: {exc.strerror}") | ||
| continue | ||
| except InconsistentLineEndings: | ||
| skipped.append(f"⊘ {name} — \\r בודד, הפארסר סירב") | ||
| continue | ||
| except TooManySections as exc: | ||
| skipped.append(f"⊘ {name} — מעל התקרה, הפארסר סירב בשורה {exc.args[0]}") | ||
| continue | ||
|
|
||
| compared += 1 | ||
| total_headings += len(theirs) | ||
| if ours != theirs: | ||
| mismatched += 1 | ||
| lines.append(f"✘ {name}") | ||
| lines.append(f" שלנו : {ours}") | ||
| lines.append(f" cmark: {theirs}") | ||
|
|
||
| header = [ | ||
| f"ריפו: {root}", | ||
| f"קבצים: {len(files)} (הושוו {compared}, סורבו {len(skipped)})", | ||
| f"כותרות: {total_headings} (לפי cmark-gfm, ברמת המסמך)", | ||
| f"אי-הסכמות: {mismatched} — ההשוואה היא על **רמה ומספר שורה** בלבד,", | ||
| " ולא על טקסט הכותרת. ההנמקה בראש קובץ האורקל.", | ||
| "", | ||
| ] | ||
| if skipped: | ||
| header.extend(skipped) | ||
| header.append("") | ||
| report = "\n".join(header + lines) + "\n" | ||
| if args.out: | ||
| args.out.write_text(report, encoding="utf-8") | ||
| print(f"הדוח נכתב ל-{args.out}") | ||
| print(report) | ||
|
|
||
| # **"אפס אי-הסכמות" על אפס קבצים שהושוו אינו הצלחה.** מאותו נימוק | ||
| # בדיוק שכתוב למעלה על ריפו בלי קובצי ``.md``: מספר שנראה טוב כי | ||
| # לא נבדק דבר הוא אישור שקרי, ומי שיראה ``exit 0`` יסיק שהפארסר | ||
| # מסכים עם cmark על הקורפוס הזה. | ||
| if not compared: | ||
| print("כל הקבצים סורבו — לא הושווה דבר, וזה כישלון ולא ריצה נקייה.") | ||
| return 1 | ||
| return 1 if mismatched else 0 | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| raise SystemExit(main()) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win
קובץ בודד שנכשל מפיל את כל הריצה.
parse_documentמרימהInconsistentLineEndingsעל\rבודד ו-TooManySectionsעל קובץ שחוצה את תקרת ברירת המחדל.read_text(encoding="utf-8")מרימהUnicodeDecodeErrorעל קובץ שאינו UTF-8. הסקריפט סורק ריפו חיצוני שרירותי, ולכן כל אחד משלושת המקרים ריאלי. היום החריגה מבעבעת החוצה, והמשתמש מקבל traceback במקום דוח — גם על מאות הקבצים שכן נסרקו.אם קובץ נכשל, רשום אותו בדוח והמשך. שמור על קוד יציאה שאינו אפס.
🛡️ תיקון מוצע: כישלון פר-קובץ נרשם ואינו עוצר
📝 Committable suggestion
🤖 Prompt for AI Agents