Skip to content

Keep flash summaries within the model's context and on their own section - #550

Open
justi wants to merge 15 commits into
VectifyAI:mainfrom
justi:pr/summary-context-and-scope
Open

justi wants to merge 15 commits into
VectifyAI:mainfrom
justi:pr/summary-context-and-scope

Conversation

@justi

@justi justi commented Oct 2, 2026 •

Copy link
Copy Markdown

Problem

I index long PDFs in flash mode with a local model (Ollama, 8k context). As the document grows:

  1. Prompts overrun the context. The document description gets the whole tree, a leaf summary gets all of its pages, and an expand prompt gets all of its pages. Ollama does not fail on an oversized prompt - it silently cuts it (keeps the first few tokens and the tail), so the model summarizes a fragment.
  2. Model JSON with LaTeX breaks the summary. Models write LaTeX inside JSON strings without doubling the backslashes. An invalid escape (\alpha) makes the reply unparseable and the raw reply, fence and all, becomes the summary. A valid one (\times, \text) is decoded by json into a TAB or a carriage return.
  3. A leaf summary covers its neighbours. A leaf is summarized from whole pages, so a section that starts mid-page also gets the end of the previous section and the start of the next one.

Changes

New behaviour is opt-in; without the two new options every prompt is byte-for-byte the same as on main (checked over the whole flash pipeline on two PDFs). The JSON fixes always apply.

client = PageIndexLocalClient(summary_max_input_tokens=8192,   # the indexing model's context
                              summary_scope="section")         # flash mode only
# or: PageIndexLocalClient(index={"summary_max_input_tokens": 8192, "summary_scope": "section"})
  • JSON replies (parse_summary, parse_title): a field whose JSON does not parse is read from the raw reply; escapes are decoded with LaTeX in mind (inside $...$ a backslash before a letter is kept, outside it \t, \r, \b, \f before a lowercase letter are kept); a parsed field that came back with control characters is read raw instead.
  • summary_max_input_tokens - the context size of the indexing model:
    • the document description is cut from the deepest tree level up until it fits,
    • a leaf too long for one call is summarized in parts (split at line boundaries) and the part summaries are combined,
    • an expand prompt that would overrun it is skipped with a warning (the node stays collapsed),
    • the size check counts digits one by one: per-digit tokenizers (Gemma, Llama) need more tokens for an index full of page numbers than tiktoken reports.
  • summary_scope="section" (flash only): a leaf is summarized from the layout blocks between its heading and the next located heading.
    • headings of bookmark trees are located on the start page, or on a heading block one page off,
    • numbered headings are kept in the blocks,
    • an intro node starts at its parent's heading,
    • a leaf falls back to its pages when a section without a located heading may start inside it,
    • mode="standard" refuses it, like the other flash-only knobs.

Results

Flash mode, gemma4 E4B via Ollama on an RTX 4090 (8192 context). main vs this branch (with #549) and summary_max_input_tokens=8192, summary_scope="section". The book numbers are means over 3 runs per side.

Llama 3 paper (92 p.) DL theory book (471 p.)
main branch main branch
prompts cut by Ollama 4 0 7 0
summaries that are raw JSON 7 0 157 0
control characters from LaTeX 0 0 18 0
leaf summary words from its own section 43% 67% 45% 61%
leaf summary words found only in neighbouring sections 26.4% 3.6% 16.1% 2.7%
indexing time 74 s 70 s 266 s 242 s

A leaf's own section is the text from its heading block to the next heading block. The two content rows are the share of a summary's distinctive words found only in that text, or only in the rest of the leaf's pages.

Known limits:

  • a JSON reply that does not parse and has a paragraph break between two dollar amounts ($5 ... $10) keeps a literal \n there (as on main);
  • a real TAB before a lowercase letter in such a reply is read as a LaTeX command (none in the 294 replies I recorded);
  • section text needs a located heading; an unlocated one falls back to pages, as on main.

Tests

New tests in tests/test_summary_replies.py, tests/test_summary_budget.py and tests/test_summary_scope.py; the full suite passes on every commit.

🤖 Generated with Claude Code

justi and others added 15 commits October 2, 2026 12:27
parse_summary returned the whole raw reply, fence included, whenever the
JSON did not parse. Small local models trigger this regularly: unescaped
quotes inside the text, or a stray backslash from LaTeX ($\alpha$, Q\&A).
Indexing a 92-page PDF with gemma 4 E4B via Ollama left 3-4 of 70 node
summaries as the raw ```json reply.

When the JSON does not parse, take the string after "summary" (or "title")
instead. One helper, shared by parse_summary and parse_title. Replies that
parse, prose replies and replies without the field behave as before.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Models write LaTeX inside JSON strings without doubling the backslashes.
Invalid escapes (\alpha) made the reply unparseable and the raw fallback
kept every escape literally; valid ones (\times, \boldsymbol) were decoded
by json into TAB and backspace characters. Escapes are now decoded with
LaTeX in mind: inside $...$ a backslash before a letter is kept, and a
parsed field that came back with control characters is read raw instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
generate_doc_description sends the whole tree with every summary in one
prompt. On a 92-page PDF that is 12.9k tokens; a local server with an 8192
token context truncates it silently, keeping only the first 5 tokens and the
tail, so the description covered one subsection or came back as raw JSON.

New option summary_max_input_tokens (default None, no change). When set, the
structure is cut from its deepest level until it fits the budget
(max_input_tokens * 0.85 - 200). Structures that already fit are passed
through untouched.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A leaf prompt carries all the text of the node's pages. A 18-page reference
list is 25k tokens; a local server with a smaller context truncates it
silently, keeping the tail only.

With summary_max_input_tokens set, a leaf over the budget is split at line
boundaries (a line of the text extracted by flash is a layout block), each
part is summarized with the existing leaf prompt, and the part summaries are
combined by one more call, in levels if they do not fit either. A line over
the budget is cut at spaces. Leaves within the budget, nodes merged from
same-page siblings and runs without the option are unchanged: the prompts
are byte-identical to upstream main.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
propose_children sends the text of every page of a collapsed node to the
model. For an 18-page reference list that is 25k tokens, and a local server
with a smaller context truncates it silently, so the model proposes
subsections from the tail only.

With summary_max_input_tokens set, the call is skipped when the pages exceed
the budget, with a warning; the node stays collapsed, as when the model finds
no subsections. Without the option nothing changes.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A leaf is summarized from the text of the pages of its node, so the text also
holds the end of the previous section and the start of the next. On a 92-page
paper, 74 of 81 leaves carried at least 30% text from outside their section,
and the summaries described the neighbours (a "Standard Benchmarks" summary
about "Steerability").

New option summary_scope ("pages" by default, no change). With "section",
extract_toc records the position of each heading's layout block, found in the
detection or, for nodes that came from the PDF's bookmarks, by locating the
title among the blocks of its start page. A leaf is then summarized from the
blocks between its heading and the next heading in the document. A leaf without
a position falls back to its pages. The positions are internal and are dropped
with the other internal keys; the tree on disk is unchanged.

On the 92-page paper 51 of 58 leaves get section text, the share of a summary
supported by its own section rises from 43% to 66%, and the share that only
the neighbouring text supports falls from 26% to 3%.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ts inside it

The section scope ends a leaf at the next located heading. When the next
section's heading was not located, the leaf took that whole section too. On a
92-page paper 10 of 51 leaves did so, and the share of their summaries taken
from the swallowed section rose from 3.5% with page text to 14.6%.

A leaf now falls back to its pages when a node without a position starts
within its page range, as the page scope does. The share returns to 4.1%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…st match

Section scope needs the block of each node's heading. On a 92-page paper
13 of 159 nodes from the PDF's bookmarks stayed without one: in 11 the
heading block is on the page after the one the bookmark names, and one was
refused by a tree-order check, although a bookmark tree with grafted detected
nodes is not in document order.

The heading is now searched on the start page, then among heading blocks on
the next and the previous page; the tree-order check is gone. Without it the
first matching block on a page could win over a better one ("Contributors and
Acknowledgements" for "Contributors", "6.1. Pre-training Evaluations" for
"6.2. Post-training Evaluations"), so the closest match wins: exact, then
with a numbering prefix, then as a prefix, then by spelling.

Nodes without a heading block: 19 -> 1 over six PDFs; positions agreeing with
the detection where it is known: 150 of 150, as before. On the paper all 58
leaves get their section text, none swallows a neighbour.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The budget is checked with tiktoken, which packs up to three digits into one
token, while Gemma's tokenizer (and other per-digit tokenizers) gives every
digit its own token. On a 471-page book the index is about 20% digits (page
numbers), so two of its parts measured 6.9k tokens with tiktoken and 9.5k in
the model: past the 8192 context, and truncated.

Budget checks now add the digits tiktoken folds together. On those two
prompts the estimate is within 1% of the real count; on prose it is 1-3%
higher, so a long leaf may be split slightly earlier. The 471-page book now
indexes with no truncated prompt; Llama 3 is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…h-only knobs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mark_outline_block_types gives a numbered heading type 8 before the blocks
are collected, so it fell out of the section text and its node lost its
position, falling back to whole pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An intro node holds the parent's text before its first child but has no
heading of its own, so it had no position and every leaf whose pages it
starts on fell back to whole pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
\text and \ref outside $...$ were decoded into TAB and carriage return.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant