Skip to content

Issue/2183/4 notice retention - #2189

Open
CarloMendola wants to merge 3 commits into
MobilityData:masterfrom
CarloMendola:issue/2183/4-notice-retention
Open

CarloMendola wants to merge 3 commits into
MobilityData:masterfrom
CarloMendola:issue/2183/4-notice-retention

Conversation

@CarloMendola

Copy link
Copy Markdown

Summary:

Fourth of the six PRs #2183 is split into, in the grouping requested here. This is the group with the deliberate behaviour change, and the one you said you would want to discuss most.

Stacked on PR 1, which it needs to compile: the first commit of this diff is PR 1's and disappears from it once PR 1 merges. The two commits to review here are the last two.

1. NoticeContainer.addAll bypassed the retention limits. addAll copied every notice of the source container into the destination with validationNotices.addAll(...), applying no limits. That is not an edge case: CsvFileLoader creates one NoticeContainer per parsed row and merges it, so each source container holds a handful of notices and the per-container limits never trigger. A feed where every row of stop_times.txt produces a warning — trailing whitespace or a non-ASCII character in an id, both common — retains one notice per row and grows the heap without bound. The addAll javadoc documented this explicitly ("the final NoticeContainer may contain more than the maximum amount…"), so the behaviour was known; this PR makes the limits actually hold, and updates that javadoc.

The limits now live in one place, canRetain, used by both addValidationNoticeWithSeverity and addAll. Two properties are preserved on purpose:

  • Counts are merged in full, regardless of retention, so totalNotices in report.json — and the counts the HTML report shows after PR 1 — stay exact. A second map tracks how many notices of each type and severity were actually retained.
  • The first notice of a type and severity is retained even once the total limit is reached. A notice type is described in the exported report only if one of its notices was retained, so dropping them all would hide the type, and its exact count, from the report entirely — and since the total limit is reached in merge order, which types disappeared would depend on file order.

The defaults are lowered accordingly:

before after
MAX_VALIDATION_NOTICES_TYPE_AND_SEVERITY 100 000 2 000
MAX_TOTAL_VALIDATION_NOTICES 10 000 000 200 000
MAX_EXPORTS_PER_NOTICE_TYPE_AND_SEVERITY 1 000 1 000 (unchanged)

At most 1 000 notices per type and severity are ever exported and the HTML report lists 50, so retaining 100 000 of them was pure overhead. 2 000 keeps a margin over both.

2. Build only the notice views the report lists. ReportSummary created a NoticeView for every retained notice, each holding the complete JSON tree of its notice, while the report lists at most 50 records per code. On a feed with many notices that is hundreds of megabytes built and thrown away, allocated at the very end of validation when the heap is already at its peak. Views are now built only up to the limit, defined once as ReportSummary.MAX_NOTICES_PER_CODE and read by the template through summary.maxNoticesPerCode instead of being duplicated as the literal 50 in two places in report.html.

The consequence for HtmlReportGenerator.getUniqueFieldsForCodes, as you asked. That method derives the column set of a notice table from the views of the code. With the views capped, the columns come from exactly the rows that get rendered: a field only ever set on a record past the 50th no longer gets a column — a column that, before this change, existed and was N/A on every rendered row. I think that is the better behaviour and the javadoc now states it, but it is a rendering difference on feeds with heterogeneous notices of one code, so it should be a conscious decision rather than a side effect. If you would rather keep the column set derived from all retained notices, that is a small change to that method and it does not affect the memory win — say the word.

Implementation report for the whole series: https://github.com/CarloMendola/gtfs-validator/blob/9f409204bcefff7387f05a3f70118fb03134443b/prompts/memory-optimization/MEMORY_OPTIMIZATION_REPORT.md

Expected behavior:

The observable change: on a pathological feed, a caller that iterates getValidationNotices() gets a smaller sample than before. totalNotices is unaffected, the exported sampleNotices are unaffected (1 000 ≤ 2 000), and the HTML report is unaffected (50 listed). Callers needing different bounds can use the three-argument constructor, which is unchanged.

On a normal feed nothing changes. Verified on the CTA Chicago feed (95 MB zip, 413 MB of CSV, 168 443 notices): the notices object of report.json is identical to master's, because every notice type there is well below the new limits — the truncation only appears on feeds producing hundreds of thousands of notices of a single type.

The rendered page is unchanged: same rows, same cells, same text (344 <tr>, 1 689 <td> on both). The file does get much smaller — 6 587 188 B on master against 272 071 B on that feed — but that is whitespace, not content: for every notice skipped by the old th:if="${iterStat.index < 50}", Thymeleaf still emitted the indentation of the iteration, 168 255 blank lines of it. Capping the list means there are no skipped iterations left.

Tests:

  • addAll_shouldRespectMaxPerNoticeTypeAndSeverity, addAll_shouldRespectMaxTotalValidationNotices — the row-by-row merge pattern of CsvFileLoader, reproduced.
  • exportNoticesAfterAddAll_shouldReportExactTotalCountBeyondRetentionLimit — totalNotices is still exact when notices were dropped, asserted on the exported JSON.
  • addAll_shouldKeepValidationErrorAndWarningFlagsWhenNoticesAreDropped — the error/warning flags are not lost with the notices.
  • addAll_totalLimitReached_stillRetainsTheFirstNoticeOfEachType — every type still appears in the exported report, with its exact count.
  • ReportSummaryTest.noticesMapTest_isTruncatedButCountsAreNot — the map is capped, the counts are not.
  • HtmlReportGeneratorTest.generateReport_listsAtMostMaxNoticesPerCodeButReportsExactTotal and generateReport_fewerNoticesThanTheLimit_listsThemAll — what the page shows, on both sides of the limit.

Nothing under docs/ describes the retention limits, so no documentation change is needed; the addAll javadoc that did describe them is updated.

Please make sure these boxes are checked before submitting your pull request - thanks!

  • Run the unit tests with gradle test to make sure you didn't break anything
  • Add or update any needed documentation to the repo
  • Format the title like "feat: [new feature short description]". Title must follow the Conventional Commit Specification(https://www.conventionalcommits.org/en/v1.0.0/).
  • Linked all relevant issues
  • Include screenshot(s) showing how this pull request works and fixes the issue(s) — the rendered page is unchanged on any feed below the limits; the two tests above assert what the page shows above them

CarloMendola and others added 3 commits September 3, 2026 11:52
`ReportSummary` derived every number it shows from the list of notices the
container retained, but `NoticeContainer` stops retaining notices of a type once
`MAX_VALIDATION_NOTICES_TYPE_AND_SEVERITY` is reached while going on counting
them. The HTML report then contradicted the JSON report generated from the same
run: on a feed producing 200 000 `stop_too_far_from_shape` warnings, the page
reported 100 000 of them where `report.json` reported 200 000, both in the total
of the heading and in the per-code total of the table.

The counts now come from the container, which counts every notice it is given
whether or not it retains it. `NoticeContainer` exposes that number per notice
code and severity, and the key those counts are stored under is built in one
place instead of being spelled out twice. The template asks the summary for the
count of a code instead of taking the size of the list of notices it renders.

Adds the first test for `HtmlReportGenerator`, rendering a report from a
container that dropped notices and asserting on the numbers the page shows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`NoticeContainer.addAll` copied every notice of the source container into the
destination without applying the retention limits. `CsvFileLoader` creates one
container per parsed row and merges it, so each of them held at most a handful
of notices and the per-container limits never triggered: a feed where every row
of stop_times.txt produces a warning (trailing whitespace or a non-ASCII
character in an id, both common) retained one notice per row, growing the heap
without bound.

The limits are now enforced in a single place, `canRetain`, used both by
`addValidationNoticeWithSeverity` and by `addAll`. Notice counts are still
merged in full, so `totalNotices` in the validation report stays exact; a
separate map tracks how many notices of each type and severity are actually
retained. The first notice of a type and severity is retained even once the
total limit is reached, because a notice type is described in the exported
report only if one of its notices was retained: dropping them all would hide
the type, and its exact count, from the report entirely.

The defaults are lowered accordingly: at most 1000 notices per type and
severity are ever exported and the HTML report renders 50, so retaining 100 000
of them was pure overhead. This is a behaviour change for callers that iterate
`getValidationNotices()` expecting every notice of a pathological feed:
`totalNotices` is unaffected, the retained sample is smaller. Callers needing
different bounds can use the three-argument constructor.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`ReportSummary` created a `NoticeView` for every retained notice, each holding
the complete JSON tree of its notice, while the report lists at most 50 records
per notice code. On a feed with many notices that is hundreds of megabytes built
and thrown away, allocated at the end of validation when the heap is already at
its peak.

The views are now built only up to the limit the report lists, which is defined
in one place (`ReportSummary.MAX_NOTICES_PER_CODE`) instead of being duplicated
as a literal in the template. Notice counts are unaffected: the per-code total
shown in the table comes from `getNoticeCountForCode`, which counts all notices,
so the rendered report is unchanged.

`HtmlReportGenerator.getUniqueFieldsForCodes` derives the columns of a notice
table from the views of that code, so the column set is now derived from exactly
the rows the report renders: a field only ever set on a record past the limit no
longer gets a column of its own, which would have been empty on every rendered
row anyway.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@CLAassistant

CLAassistant commented Sep 3, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants