Skip to content

fix(optimizer): preserve Unicode heading keys - #546

Open
rudycelekli wants to merge 1 commit into
VectifyAI:mainfrom
rudycelekli:fix/unicode-optimization-headings-20261001
Open

rudycelekli wants to merge 1 commit into
VectifyAI:mainfrom
rudycelekli:fix/unicode-optimization-headings-20261001

Conversation

@rudycelekli

Copy link
Copy Markdown

Flash expansion currently strips every non-ASCII letter from its comparison keys. Printed Japanese subsection titles then become empty strings and are rejected as identical to the parent; an absent non-ASCII title can also match any page beneath an ASCII parent. Preserve Unicode letters and numbers while retaining the existing lowercase, whitespace, punctuation and underscore handling.

Add four public page_index_flash regressions using generated Japanese and ASCII PDFs. They exercise native PDF parsing, bookmarks and optimization, with only external completion responses stubbed. The cases check printed heading extraction, duplicate/parent/absent-title rejection, and page ranges when a subsection starts at the top or midway through a page.

Validation:

  • Four focused native PDF cases pass; restoring the old normalization expression makes all four fail.
  • Full local suite: 589 passed, 218 skipped on Python 3.12/PDFium 5.13 without optional agent frameworks. This used the same production source; the focused cases and mutation check were repeated after moving the printed-text assertion earlier in the tests.
  • git diff --check passed. No real model calls were made.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant