Skip to content

Normalize text-file captions before training cache lookups - #3262

Merged
bghira merged 1 commit into
mainfrom
fix/3251-textfile-caption-cache
Oct 9, 2026
Merged

bghira merged 1 commit into
mainfrom
fix/3251-textfile-caption-cache

Conversation

@bghira

@bghira bghira commented Oct 9, 2026 •

Copy link
Copy Markdown
Owner

Text-file captions with surrounding whitespace are stripped during precaching but used verbatim during training. The resulting hash mismatch aborts training with a missing text embedding cache file; the filename-only fix in #3244 does not cover this path.

Reuse the existing caption normalization and shape restoration helpers when reading text-file captions. Single captions remain strings and multiline captions remain lists.

Validation:

  • .venv/bin/python -m unittest -v -f tests.test_prompthandler: 13 passed. The regression verifies precache/training agreement for Unicode, bytes/CRLF, blank multiline entries, disabled newline splitting, and instance-prompt prefixes; all five cases fail on main.
  • Full discovery reached 6,328 tests before an unrelated webhook fixture failed to bind shared test port 8998. The 478 tests from webhook authentication onward passed when run separately, with 27 skips.
  • Formatting and commit hooks passed. The existing environment uses Python 3.14.7; full discovery used TORCHINDUCTOR_CPP_CACHE_PRECOMPILE_HEADERS=0 to avoid a stale Inductor header cache.

Closes #3251.

Additional validation on a Linux H100 host with CUDA PyTorch 2.11.0+cu128: all 13 focused caption tests passed.

@bghira
bghira merged commit 11a352c into main Oct 9, 2026
2 checks passed
@bghira
bghira deleted the fix/3251-textfile-caption-cache branch October 9, 2026 18:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Text cache file not found

1 participant