Skip to content

Move to the paged document support in Verify 33.3: text by page, and a sheet is a page - #160

Merged
SimonCropp merged 4 commits into
mainfrom
paged-documents
Oct 5, 2026
Merged

SimonCropp merged 4 commits into
mainfrom
paged-documents

Conversation

@SimonCropp

@SimonCropp SimonCropp commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Moves Verify.OpenXml onto the source and derived targets that Verify 33.3 added for paged documents (Verify#1951), and takes the package from 1.28.0 to 2.0.0.

The converter now says which of its targets is the document and which were computed from it. Verify compares the document first, so a file derived from it is only given the benefit of its comparer while the document itself is unchanged, and DiffEngine's viewer shows the files derived from a document beneath it and accepts them with it.

It is a major version because what an existing test suite has to change is listed below. Verify and its adapter are pinned at 33.3.0.

Beyond the move

  • Text by page for a docx. With a rendering backend the text of a Word document is under the page it is on, read with DocumentConverter.GetPageTexts of Morph 1.20.0 (Morph#186), so PagesToInclude limits it as it does the images. A paragraph or a table row that runs over the end of a page is divided where the page ends. Without a backend, or below net10.0, there is no layout, and the text is that of the whole document as before.
  • A workbook is drawn a sheet to a page, with SheetPagination.OnePagePerSheet of Morph 1.19.0: a sheet is the one image, however long it is and whatever paper its page setup names. It was a page for each page of the print layout.
  • PagesToInclude leaves out the csv of a sheet with its image. A page is a sheet in tab order, which does not depend on a renderer, so the csv is left out the same way where there is none.
  • Hidden sheets are pages as any other, counted by PagesToInclude, with a png and a csv. Morph draws a workbook as it prints, so one with a hidden sheet is drawn from a copy in which every sheet is shown. They are named under HiddenSheets in the info file.

Migrating from 1.x

Version 2 moves to the paged document support in Verify 33.3. Pages and single sheets are renamed, the text moves into the info file, and OpenXmlOutputs gives way to Verify's own settings.

Settings

Initialize no longer takes an OpenXmlOutputs. Each output that could be left out of it is now left out by a setting of Verify, for every test on VerifierSettings or for a single verification:

1.x: not in OpenXmlOutputs 2.x
Png ExcludeDerivedTargets("png")
Text PageText(PageTextPlacement.None)
Csv ExcludeDerivedTargets("csv")

So Initialize(OpenXmlOutputs.Csv) becomes:

VerifyOpenXml.Initialize();
VerifierSettings.PageText(PageTextPlacement.None);
VerifierSettings.ExcludeDerivedTargets("png");

Files

For a test Tests.Report:

1.x 2.x
Tests.Report.verified.png, the page of a document with one Tests.Report#page_0001.verified.png
Tests.Report#00.verified.png, #01, the pages of a document with several Tests.Report#page_0001.verified.png, #page_0002
Tests.Report.verified.csv, the sheet of a workbook with one Tests.Report#Sheet1.verified.csv, by the name of the sheet
Tests.Report#Sheet1.verified.csv, a sheet of a workbook with several The same
Tests.Report#00.verified.txt, the info of a docx or pptx Tests.Report.verified.txt
Tests.Report#01.verified.txt, the text of a docx or pptx In the info file, or with PageText(PageTextPlacement.PerPage) in #page_0001.verified.txt for each slide of a pptx and each page of a docx, or #text.verified.txt for a docx read without a rendering backend
Tests.Report.verified.txt, the info of an xlsx The same
Tests.Report.verified.docx, .xlsx, .pptx The same

A renamed snapshot shows as a new file and a pending delete. Accepting both, or running once with AutoVerify, moves a test over. The content of a sheet is unchanged, so source control shows it as a rename. The page images are drawn by Morph 1.20.0 where they were by 1.18.0, and those of a workbook are of a sheet where they were of a printed page, so they differ.

The info file

What was at the top of the info file is now under Document, with the page count and the text beside it. For a docx:

{                                    {
  Properties: {                        Document: {
    Title: Sample Document               Properties: {
  },                                       Title: Sample Document
  Fonts: [                               },
    Aptos                                Fonts: [
  ]                                        Aptos
}                                        ]
                                       },
                                       PageCount: 1,
                                       Text: The text of the document
                                     }

With a rendering backend the text of a docx is the Text of each page instead, as it is for a pptx.

For an xlsx the properties of the workbook, with its Sheets, move under Document the same way.

For a pptx SlideCount is now PageCount, and the text is the Text of each page, where it was one text with --- between the slides. Slides are read in the order they are shown in, that of p:sldIdLst, so that the text of a slide is on the page it is drawn on. They were read in the order of the slide parts, which is not the order of a deck whose slides have been moved.

Testing

Run locally in Debug against Verify 33.3.0 and Morph 1.20.0 from nuget.org: Verify.OpenXml.Tests, Tests.Skia, Tests.ImageSharp and StaticSettingsTests all pass.

Not done: skipping the drawing of the pages PagesToInclude leaves out. Every page is still drawn and the unwanted ones dropped. Morph's ImageExportOptions.Pages could limit that to a range.

With a rendering backend the text of a docx is under the page it is on, so PagesToInclude limits it as it does the images. Each paragraph and table row is put on the page a bookmark at its start lands on.

Morph 1.19.0, and workbooks are rendered with SheetPagination.OnePagePerSheet: a page is a visible sheet, drawn whole. The Word and PowerPoint page images change with the Morph version.
Morph 1.20.0 reports the text of each page, which replaces putting a bookmark on every paragraph and asking where each landed. A paragraph that runs over the end of a page is now divided where the page ends, and the page count is in the info file when no page is drawn.
PagesToInclude leaves out the csv of a sheet with its image, with or without a renderer. A hidden sheet is drawn and exported as any other, from a copy of the workbook in which every sheet is shown, and is named under HiddenSheets in the info file.
@SimonCropp SimonCropp added this to the 2.0.0 milestone Oct 5, 2026
@SimonCropp SimonCropp changed the title Move to the paged document support in Verify 33.3 Move to the paged document support in Verify 33.3: text by page, and a sheet is a page Oct 5, 2026
@SimonCropp
SimonCropp merged commit 15b529c into main Oct 5, 2026
6 checks passed
@SimonCropp
SimonCropp deleted the paged-documents branch October 5, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant