Repository navigation
fix(pdf): errors name the next step — password, no pages, over 10 000 pages - #199
Merged
Merged
Conversation
… pages
A PDF that needs a password to open got the answer for a broken file on
every PDF path ("Could not read the PDF. Verify the file is valid.", the
compression failure on /pdf/compress, a 500 on PDF → PDF/A) — a dead end,
since checking the file can't help. pypdf's FileNotDecryptedError and
pikepdf's PasswordError now become EncryptedPdfError, an InvalidInputError
with a fixed message that names the fix. /pdf/extract, /pdf/split,
/pdf/compress and /convert answer 400 with the new X-FileMorph-Error-Code
pdf_encrypted; pdf-tools.js and app.js map it to the localized
FM_I18N.pdfEncrypted (DE catalog + .mo); /convert/batch reports the
message per file. PDFs that open without a password are untouched.
/pdf/extract sent invalid_page_selection for a PDF without pages, so its
page blamed the user's selection; parse_page_ranges now raises the
UnreadablePdfError subclass for it and the route sends invalid_pdf, as
/pdf/split did. PDF → PDF without a selection keeps every page and is
held to the 10 000-page selection cap; over it, the message now names the
document's size and points to /api/v1/pdf/extract instead of "Too many
pages selected."
The pypdf paths were the original scope; compress and PDF/A (pikepdf) are
included because they had the same dead end and share the code and text.
Rejected: keeping invalid_pdf / invalid_input and showing `detail` in the
UI — server details are English-only, so German users would lose the
localized text. Rejected: refusing every PDF with reader.is_encrypted —
PDFs with only an owner password open without one and must keep working
(pinned by tests).
Full suite 1728 green (95 skipped; the pikepdf cases run on Linux CI);
ruff + i18n-drift + Tailwind + pip-audit clean. UI checked in a browser
against a local server (DE/EN: extract, split, converter).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…it renders it without the password CI (Linux, Ghostscript 10.02.1) showed PDF → PDF/A answering 200 for a PDF that needs a password. Ghostscript doesn't fail on such a file: it renders it without the password and exits 0, so the PasswordError catch after it never fired and the markup pass shipped a PDF/A that cannot hold the protected content — a false success. PdfToPdfaConverter now opens the input with pikepdf first and raises EncryptedPdfError on PasswordError before Ghostscript sees the file. Any other open error still goes on to Ghostscript, which repairs some files pikepdf can't open. The catch after Ghostscript is gone (unreachable now). Two Linux-only tests fake Ghostscript, so they hold whether or not it is installed: a locked PDF must not reach it, and a file pikepdf can't open still does. The changelog fragment now describes the old PDF/A behaviour correctly (200 without the content; a 500 only where Ghostscript is missing). Full suite 1728 green locally (123 skipped); test_pdfa, test_pdf_compress and test_pdf_error_messages 106 green via a local pikepdf-first run; ruff + i18n-drift + Tailwind clean. pip-audit is red on main for a new WeasyPrint advisory (CVE-2026-106443), handled separately. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MrChengLen
enabled auto-merge
October 8, 2026 08:44
MrChengLen
added a commit
that referenced
this pull request
Oct 8, 2026
…s code /pdf/split answered a PDF over the 10,000-page cap with invalid_pdf, so the split page told the user that a valid 12,000-page file could not be read, and offered no way forward. /pdf/extract answered a selection over the cap with invalid_page_selection, and its page showed the page-number syntax hint instead of the limit. The cap now raises TooManyPagesError, a PageSelectionError subclass, and both routes send X-FileMorph-Error-Code: pdf_too_many_pages with a message that names the cap (built from _MAX_SELECTION_PAGES) and the next step. pdf-tools.js picks the localized text by tool: split says to extract up to 10,000 pages at a time and split each part, and un-hides a link to the extract page next to it (product-ux rule 2: no dead end); extract says to select fewer pages or to work in parts. DE/EN catalogs updated; a test pins the number in both texts to the constant. /convert is unchanged on purpose: PDF -> PDF over the cap keeps invalid_input, whose message already names the cap (#199). A new code there would change a public contract with no UI behind it. Full suite 1755 green (128 skipped on Windows); ruff + i18n-drift + pip-audit clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #186: three PDF errors that didn't name the user's next step.
What
/pdf/extract,/pdf/split,/convertPDF → TXT / PDF), "PDF compression failed. …" (/pdf/compress), and PDF → PDF/A answered 200 with a PDF/A that lacks the content: Ghostscript, which runs first where installed, renders such a file without the password and exits 0 (a 500 where Ghostscript is missing). Every path now answers400+X-FileMorph-Error-Code: pdf_encrypted+ "This PDF is password-protected. Remove the password (e.g. open the file and print it to a new PDF) and try again." Both UIs (pdf-tools.js,app.js) show it localized (DE/EN);/convert/batchreports it per file./pdf/extract:invalid_page_selection→invalid_pdf, as on/pdf/split(the page used to blame a selection that was fine)./convert, no selection): "Too many pages selected." → names the document's size and points to/api/v1/pdf/extract.How
EncryptedPdfError(InvalidInputError)inapp/converters/base.py(fixed message, copy/pickle-safe), raised by the pypdf guards (_reading_pdf,PdfToTxtConverter) onFileNotDecryptedErrorand by the pikepdf paths (compress_pdf_to_target,PdfToPdfaConverter) onPasswordError. The routes map it topdf_encrypted.parse_page_rangesraises theUnreadablePdfErrorsubclass for a PDF without pages.pdfEncryptedinFM_I18N, DE translated,.mocompiled, no fuzzy entries.docs/api-reference.md), batch messages (docs/api-usage-guide.md), honest limits (docs/formats.md). Changelog:changelog.d/2026-10-06-pdf-error-next-step.md.Scope: the pypdf paths were the original scope; compress and PDF/A (pikepdf) are included because they had the same dead end — PDF/A even a 500 — and share the code and the text.
Verification
tests/test_pdf_error_messages.py(73 cases): RC4-128 / AES-128 / AES-256 on all six routes and batch, owner-password-only PDFs still processed, PDF without pages, page-cap boundary, DE/ENFM_I18N, JS wiring. Checked to fail on the pre-fix code, apart from the guard cases.bea9437:PdfToPdfaConverterprobes for a user password with pikepdf before Ghostscript runs; two fake-Ghostscript tests pin the order (a locked PDF never reaches it, a file pikepdf can't open still does). The PR's merge ref tree equals the locally tested tree.Not in this PR
/pdf/splitover 10 000 pages and a too-large/pdf/extractselection still show generic texts on the tool pages (invalid_pdf/invalid_page_selectionmapping inpdf-tools.js)./Adobe.PubSec) still answer 500: pypdf raisesNotImplementedErroron open.🤖 Generated with Claude Code