Skip to content
Discussion options

You must be logged in to vote

The PYQ parser pipeline in \scripts/pyq_parser.py\ processes questions as follows:

  1. PDF Text Extraction: Uses \pdfplumber\ / OCR to extract raw text and diagrams.
  2. Regex & Layout Analysis: Identifies question boundaries, shift timestamps, and subject headers.
  3. Structured JSON Output: Formats each question into standard schema with options, correct answer keys, subject tag, and difficulty rating.

Replies: 1 comment

Comment options

itzzdev09
Aug 27, 2026
Maintainer Author

You must be logged in to vote
0 replies
Answer selected by itzzdev09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
1 participant