fix: guard oversized JSON responses in code-generation prompt - #76
Open
tkatta-stack wants to merge 1 commit into
Open
tkatta-stack wants to merge 1 commit into
tkatta-stack wants to merge 1 commit into
Conversation
generate_code() (util/print.py) embeds the full HTTP response body in the prompt it sends to the LLM when writing the parsing code for each node. The text/html and application/javascript branch already guards against oversized responses (>100000 chars) by substituting short context snippets instead of the full body. The application/json branch had no such guard: it always embedded the complete raw response text verbatim, in addition to the already-resolved key_paths that pinpoint where each extracted variable lives in the parsed JSON. A single large JSON API response (a bulk data dump, a paginated listing, etc.) captured in the HAR file is therefore embedded in full, and can by itself push a single prompt well past the model's context window -- this matches the reported symptom of "messages exceeded available capacity by roughly 3x" after a normal integuru run. This mirrors the existing 100000-character threshold from the HTML/JS branch onto the JSON branch: past that threshold, the prompt drops the raw response text and relies solely on the key_paths already computed, which are sufficient for the model to navigate the parsed JSON and write correct extraction code without needing to see the raw payload. Added tests/test_generate_code.py covering both the oversized-response path (raw text excluded, key path included) and the normal small- response path (behavior unchanged). Fixes Integuru-AI#30 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WV3dWyboqjHodswTzHTBKr
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
generate_code()(ininteguru/util/print.py) builds a single LLM prompt per DAG node to write that node's request-parsing function. For atext/html/application/javascriptresponse it already guards against oversized bodies (>100000 chars) by substituting short context snippets instead of the full response. Theapplication/jsonbranch had no such guard — it always embedded the complete raw response text verbatim, on top of thekey_pathsalready computed that pinpoint exactly where each extracted variable lives in the parsed JSON.Why this causes #30
A single large JSON API response captured in the HAR file (a bulk data dump, a paginated listing, etc.) gets embedded in full. By itself this can push one prompt well past the model's context window — matching the reported "messages exceeded available capacity by roughly 3x" after a normal run, and the second reporter's identical experience.
Fix
Mirrors the existing 100000-character threshold from the HTML/JS branch onto the JSON branch. Past that threshold, the prompt drops the raw response text and relies solely on the already-computed
key_paths, which are sufficient for the model to navigate the parsed JSON and write correct extraction code without seeing the raw payload. Small/typical JSON responses are unaffected — same prompt as before.Testing
Added
tests/test_generate_code.py:Both pass with the fix; I also confirmed the large-response test fails against the pre-fix code (i.e. it genuinely catches the regression).
Note:
tests/test_integration_agent.pycurrently fails locally/in CI independent of this change — it needs atest.harfixture file that isn't committed to the repo. Not touched here since it's unrelated to this fix; flagging for visibility.Fixes #30