Skip to content

Commit 511652e

Browse files
committed
Increase bounded PDF source limit to 100 MiB
1 parent 2693685 commit 511652e

6 files changed

Lines changed: 21 additions & 4 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,10 @@
22

33
All notable changes to ManagedCode.FileContext are documented here.
44

5+
## 1.0.10
6+
7+
- Raise the bounded PDF source read limit from 25 MiB to 100 MiB for scanned documents while retaining per-page pixel and PNG output limits.
8+
59
## 1.0.9
610

711
- Add a native bounded DOCX text tool with paragraph and character cursors for long documents.

‎Directory.Build.props‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,7 @@
1212
<AnalysisMode>Recommended</AnalysisMode>
1313
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
1414
<NoWarn>$(NoWarn);CS1591;MAAI001</NoWarn>
15-
<Version>1.0.9</Version>
15+
<Version>1.0.10</Version>
1616
<PackageVersion>$(Version)</PackageVersion>
1717
</PropertyGroup>
1818

‎README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -187,7 +187,7 @@ Results include `StartLine`, `EndLine`, `HasMore`, and `TotalLines` when the end
187187

188188
`IFileContextPdf.ReadPdfTextAsync(path)` returns bounded text, `PageCount`, and one-based `PagesWithoutText`. It does not perform OCR. A scanned page can instead be rendered with `RenderPdfPageAsync(path, pageNumber)`, which returns PNG `DataContent`. Use `CountPdfPageImagesAsync` and `ExtractPdfImageAsync` when the original embedded pictures are needed rather than the complete page. The four read-only `file_context_pdf_*` tools expose the same operations from scoped storage.
189189

190-
For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. Storage reads enforce `MaximumPdfReadBytes` (25 MiB by default); page rasterization caps pixels and PNG size. A host must pass image `DataContent` to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
190+
For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. PDF source reads are capped at 100 MiB by default; page rasterization caps pixels and PNG size. A host must pass image `DataContent` to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
191191

192192
`file_context_docx_text(path, startParagraph?, startCharacter?, paragraphCount?)` reads ordinary paragraph and table text from a scoped DOCX package. The result contains numbered paragraph segments and `nextParagraph`/`nextCharacter`; use that cursor to continue a long document. Reads are limited to 50 paragraphs and 20,000 characters per call, with a configurable 25 MiB source limit (`MaximumDocxReadBytes`). It does not OCR embedded images. DOCX and XLSX packages are excluded from generic text reads and grep.
193193

‎src/ManagedCode.FileContext/FileContextDefaults.cs‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,11 +1,13 @@
1+
using ManagedCode.FileContext.Pdf;
2+
13
namespace ManagedCode.FileContext;
24

35
/// <summary>Default limits and selectors used by <see cref="FileContextOptions" />.</summary>
46
public static class FileContextDefaults
57
{
68
public const long MaximumGeneratedFileBytes = 64L * 1024L * 1024L;
79
public const int FirstLineNumber = 1;
8-
public const int MaximumPdfReadBytes = 25 * 1024 * 1024;
10+
public const int MaximumPdfReadBytes = FileContextPdfImages.MaximumPdfBytes;
911
public const int MaximumPdfTextCharacters = 50_000;
1012
public const int MaximumDocxReadBytes = 25 * 1024 * 1024;
1113
public const int MaximumDocxTextCharacters = 20_000;

‎src/ManagedCode.FileContext/Pdf/FileContextPdfImages.cs‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,7 @@ namespace ManagedCode.FileContext.Pdf;
66
/// <summary>Renders complete PDF pages and extracts embedded page images as PNG.</summary>
77
public static class FileContextPdfImages
88
{
9-
public const int MaximumPdfBytes = 25 * 1024 * 1024;
9+
public const int MaximumPdfBytes = 100 * 1024 * 1024;
1010
public const int MaximumImageBytes = 8 * 1024 * 1024;
1111
public const int MaximumPixels = 4_000_000;
1212
public const int MaximumImagesPerPage = 20;

‎tests/ManagedCode.FileContext.Tests/FileContextPdfTests.cs‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -93,6 +93,17 @@ public void Page_image_enforces_page_scale_and_input_limits()
9393
Should.Throw<ArgumentOutOfRangeException>(() => FileContextPdfImages.ExtractPageImagesPng(pdf, 2));
9494
}
9595

96+
[Fact]
97+
public void Pdf_over_the_former_25_mib_limit_reaches_parsing()
98+
{
99+
var source = new byte[25 * 1024 * 1024 + 1];
100+
101+
var exception = Record.Exception(() => FileContextPdfImages.RenderPagePng(source, 1));
102+
103+
exception.ShouldNotBeNull();
104+
exception.Message.ShouldNotContain("The PDF exceeds the read limit.");
105+
}
106+
96107
[Fact]
97108
public async Task Pdf_service_enforces_byte_and_image_index_limits()
98109
{

0 commit comments

Comments
 (0)