Skip to content

fix(pdf): errors name the next step — password, no pages, over 10 000 pages - #199

Merged
MrChengLen merged 3 commits into
mainfrom
pr-pdf-error-next-step
Oct 8, 2026
Merged

MrChengLen merged 3 commits into
mainfrom
pr-pdf-error-next-step

Conversation

@MrChengLen

@MrChengLen MrChengLen commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Follow-up to #186: three PDF errors that didn't name the user's next step.

What

  1. Password-protected PDF (needs a password to open; RC4 or AES) got the answer for a broken file on every PDF path: "Could not read the PDF. Verify the file is valid." (/pdf/extract, /pdf/split, /convert PDF → TXT / PDF), "PDF compression failed. …" (/pdf/compress), and PDF → PDF/A answered 200 with a PDF/A that lacks the content: Ghostscript, which runs first where installed, renders such a file without the password and exits 0 (a 500 where Ghostscript is missing). Every path now answers 400 + X-FileMorph-Error-Code: pdf_encrypted + "This PDF is password-protected. Remove the password (e.g. open the file and print it to a new PDF) and try again." Both UIs (pdf-tools.js, app.js) show it localized (DE/EN); /convert/batch reports it per file.
  2. PDF without pages on /pdf/extract: invalid_page_selection → invalid_pdf, as on /pdf/split (the page used to blame a selection that was fine).
  3. PDF → PDF pass-through over 10 000 pages (/convert, no selection): "Too many pages selected." → names the document's size and points to /api/v1/pdf/extract.

How

  • EncryptedPdfError(InvalidInputError) in app/converters/base.py (fixed message, copy/pickle-safe), raised by the pypdf guards (_reading_pdf, PdfToTxtConverter) on FileNotDecryptedError and by the pikepdf paths (compress_pdf_to_target, PdfToPdfaConverter) on PasswordError. The routes map it to pdf_encrypted.
  • parse_page_ranges raises the UnreadablePdfError subclass for a PDF without pages.
  • i18n: pdfEncrypted in FM_I18N, DE translated, .mo compiled, no fuzzy entries.
  • Docs: error-code table + 400 row (docs/api-reference.md), batch messages (docs/api-usage-guide.md), honest limits (docs/formats.md). Changelog: changelog.d/2026-10-06-pdf-error-next-step.md.

Scope: the pypdf paths were the original scope; compress and PDF/A (pikepdf) are included because they had the same dead end — PDF/A even a 500 — and share the code and the text.

Verification

  • Full suite 1728 passed / 95 skipped locally (Windows). The pikepdf + ghostscript cases (compress, PDF/A, batch PDF/A) only run here on Linux CI.
  • New tests/test_pdf_error_messages.py (73 cases): RC4-128 / AES-128 / AES-256 on all six routes and batch, owner-password-only PDFs still processed, PDF without pages, page-cap boundary, DE/EN FM_I18N, JS wiring. Checked to fail on the pre-fix code, apart from the guard cases.
  • UI checked in a browser against a local server (DE/EN: extract, split, converter).
  • The first CI run caught that Ghostscript behaviour (6 PDF/A cases answered 200). Fixed in bea9437: PdfToPdfaConverter probes for a user password with pikepdf before Ghostscript runs; two fake-Ghostscript tests pin the order (a locked PDF never reaches it, a file pikepdf can't open still does). The PR's merge ref tree equals the locally tested tree.
  • CI is currently blocked by main's red pip-audit (WeasyPrint CVE-2026-106443), handled separately; the PR needs an update-branch once main is green.
  • ruff, i18n drift-check, Tailwind gate, pip-audit clean; gitleaks + pre-commit scope guard clean; security review PASS, code review approved.

Not in this PR

  • /pdf/split over 10 000 pages and a too-large /pdf/extract selection still show generic texts on the tool pages (invalid_pdf / invalid_page_selection mapping in pdf-tools.js).
  • Certificate-encrypted PDFs (/Adobe.PubSec) still answer 500: pypdf raises NotImplementedError on open.

🤖 Generated with Claude Code

MrChengLen and others added 3 commits October 6, 2026 16:14
… pages

A PDF that needs a password to open got the answer for a broken file on
every PDF path ("Could not read the PDF. Verify the file is valid.", the
compression failure on /pdf/compress, a 500 on PDF → PDF/A) — a dead end,
since checking the file can't help. pypdf's FileNotDecryptedError and
pikepdf's PasswordError now become EncryptedPdfError, an InvalidInputError
with a fixed message that names the fix. /pdf/extract, /pdf/split,
/pdf/compress and /convert answer 400 with the new X-FileMorph-Error-Code
pdf_encrypted; pdf-tools.js and app.js map it to the localized
FM_I18N.pdfEncrypted (DE catalog + .mo); /convert/batch reports the
message per file. PDFs that open without a password are untouched.

/pdf/extract sent invalid_page_selection for a PDF without pages, so its
page blamed the user's selection; parse_page_ranges now raises the
UnreadablePdfError subclass for it and the route sends invalid_pdf, as
/pdf/split did. PDF → PDF without a selection keeps every page and is
held to the 10 000-page selection cap; over it, the message now names the
document's size and points to /api/v1/pdf/extract instead of "Too many
pages selected."

The pypdf paths were the original scope; compress and PDF/A (pikepdf) are
included because they had the same dead end and share the code and text.

Rejected: keeping invalid_pdf / invalid_input and showing `detail` in the
UI — server details are English-only, so German users would lose the
localized text. Rejected: refusing every PDF with reader.is_encrypted —
PDFs with only an owner password open without one and must keep working
(pinned by tests).

Full suite 1728 green (95 skipped; the pikepdf cases run on Linux CI);
ruff + i18n-drift + Tailwind + pip-audit clean. UI checked in a browser
against a local server (DE/EN: extract, split, converter).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…it renders it without the password

CI (Linux, Ghostscript 10.02.1) showed PDF → PDF/A answering 200 for a PDF
that needs a password. Ghostscript doesn't fail on such a file: it renders
it without the password and exits 0, so the PasswordError catch after it
never fired and the markup pass shipped a PDF/A that cannot hold the
protected content — a false success.

PdfToPdfaConverter now opens the input with pikepdf first and raises
EncryptedPdfError on PasswordError before Ghostscript sees the file. Any
other open error still goes on to Ghostscript, which repairs some files
pikepdf can't open. The catch after Ghostscript is gone (unreachable now).

Two Linux-only tests fake Ghostscript, so they hold whether or not it is
installed: a locked PDF must not reach it, and a file pikepdf can't open
still does. The changelog fragment now describes the old PDF/A behaviour
correctly (200 without the content; a 500 only where Ghostscript is
missing).

Full suite 1728 green locally (123 skipped); test_pdfa, test_pdf_compress
and test_pdf_error_messages 106 green via a local pikepdf-first run; ruff +
i18n-drift + Tailwind clean. pip-audit is red on main for a new WeasyPrint
advisory (CVE-2026-106443), handled separately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MrChengLen
MrChengLen enabled auto-merge October 8, 2026 08:44
@MrChengLen
MrChengLen merged commit 7fb5631 into main Oct 8, 2026
7 checks passed
@MrChengLen
MrChengLen deleted the pr-pdf-error-next-step branch October 8, 2026 08:50
MrChengLen added a commit that referenced this pull request Oct 8, 2026
…s code

/pdf/split answered a PDF over the 10,000-page cap with invalid_pdf, so the
split page told the user that a valid 12,000-page file could not be read, and
offered no way forward. /pdf/extract answered a selection over the cap with
invalid_page_selection, and its page showed the page-number syntax hint
instead of the limit.

The cap now raises TooManyPagesError, a PageSelectionError subclass, and both
routes send X-FileMorph-Error-Code: pdf_too_many_pages with a message that
names the cap (built from _MAX_SELECTION_PAGES) and the next step.
pdf-tools.js picks the localized text by tool: split says to extract up to
10,000 pages at a time and split each part, and un-hides a link to the
extract page next to it (product-ux rule 2: no dead end); extract says to
select fewer pages or to work in parts. DE/EN catalogs updated; a test pins
the number in both texts to the constant.

/convert is unchanged on purpose: PDF -> PDF over the cap keeps
invalid_input, whose message already names the cap (#199). A new code there
would change a public contract with no UI behind it.

Full suite 1755 green (128 skipped on Windows); ruff + i18n-drift + pip-audit clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant