diff --git a/docs/about.rst b/docs/about.rst index 159922ea3..1bfdecc94 100644 --- a/docs/about.rst +++ b/docs/about.rst @@ -60,29 +60,11 @@ The following table illustrates how |PyMuPDF| compares with other typical soluti Therefore input files are mostly in a form that's useful for text extraction. - If faithful reproduction of layout is important, then consider using :ref:`PyMuPDF Pro `. + If faithful reproduction of layout is important, then consider using :ref:`PyMuPDF Office `. ---- -.. _About_PyMuPDF_Product_Suite: - -PyMuPDF Product Suite ------------------------------------------------ - -|PyMuPDF| is the standard version of the library, however there are a family of additional products each with different features and functionality. - -**Additional products** in the |PyMuPDF| product suite are: - -- |PyMuPDF Pro| adds support for Office document formats. -- |PyMuPDF4LLM| is optimized for large language model (LLM) applications, providing enhanced text extraction and processing capabilities. - It focuses on layout analysis and semantic understanding, ideal for document conversion and formatting tasks with enhanced results. - -.. note:: - All of the products above depend on the same core product - |PyMuPDF| and therefore have full access to all of its features. - These additional products can be seen as optional extras to the enhance the core |PyMuPDF| library. - - .. _About_PyMuPDF_Products_Comparison: PyMuPDF Products Comparison @@ -91,51 +73,40 @@ PyMuPDF Products Comparison The following table illustrates what features the products offer: .. list-table:: PyMuPDF Products Comparison - :widths: 10 30 30 30 + :widths: 10 45 45 :header-rows: 1 * - - PyMuPDF - - PyMuPDF Pro - - PyMuPDF4LLM + - PyMuPDF Office * - **Input Documents** - `PDF`, `XPS`, `EPUB`, `CBZ`, `MOBI`, `FB2`, `SVG`, `TXT`, `MD`, Images (*standard document types*) - *as PyMuPDF* and: `DOC`/`DOCX`, `XLS`/`XLSX`, `PPT`/`PPTX`, `HWP`/`HWPX` - - *as PyMuPDF* * - **Output Documents** - - Can convert any input document to `PDF`, `SVG` or Image + - Can convert any input document to `PDF`, Markdown (`MD`), `JSON`,`TXT`,`SVG`, Image - *as PyMuPDF* - - *as PyMuPDF* and: - Markdown (`MD`), `JSON` or `TXT` * - **Page Analysis** - - Basic page analysis to return document structure - - *as PyMuPDF* - Advanced Page Analysis with trained data for enhanced results - * - **Data extraction** - - Basic data extraction with structured layout information and bounding box data - *as PyMuPDF* + * - **Data extraction** - Advanced data extraction including layout analysis with semantic understanding and enhanced bounding box data - * - **Table extraction** - - Basic table extraction as part of text extraction - *as PyMuPDF* + * - **Table extraction** - Advanced table extraction with cell structure, including support for merged cells and complex layouts - * - **Image extraction** - - Basic image extraction - *as PyMuPDF* + * - **Image extraction** - Advanced detection and rendering of image areas on page saving them to disk or embedding in MD output - * - **Vector extraction** - - Vector extraction and clustering - *as PyMuPDF* + * - **Vector extraction** - Superior detection of "picture" areas + - *as PyMuPDF* * - **Popular RAG Integrations** - Langchain, LlamaIndex - *as PyMuPDF* - - *as PyMuPDF* and with some additional help methods for RAG workflows * - **OCR** - - On-demand invocation of built-in Tesseract for text detection on pages or images + - Hybrid OCR based on page content analysis. OCR adapators for popular OCR engines available - *as PyMuPDF* - - Automatic OCR based on page content analysis. OCR adapators for popular OCR engines available ---- diff --git a/docs/converting-files.rst b/docs/converting-files.rst index efb3707fd..aeb906a4c 100644 --- a/docs/converting-files.rst +++ b/docs/converting-files.rst @@ -158,19 +158,24 @@ For example, assuming you have access to the source files for the "Comic Sans" f PDF to Markdown ~~~~~~~~~~~~~~~~~ -By utlilizing the :doc:`PyMuPDF4LLM API ` we are able to convert PDF to a Markdown representation. +By utlilizing the :meth:`Document.to_markdown` method we are able to convert any of the following document types to a Markdown representation: + +- PDFs +- Image documents (because they are internally converted to 1-page PDFs) +- Office documents opened using PyMuPDF Office (because they are converted to PDFs via :ref:`to_pdf() ` before processing them) **Example** .. code-block:: python - import pymupdf4llm + import pymupdf import pathlib - md_text = pymupdf4llm.to_markdown("test.pdf") + doc = pymupdf.open("test.pdf") + md_text = doc.to_markdown() print(md_text) - pathlib.Path("4llm-output.md").write_bytes(md_text.encode()) + pathlib.Path("output.md").write_bytes(md_text.encode()) PDF to SVG diff --git a/docs/document.rst b/docs/document.rst index 6c3d9a2ce..c75e1e116 100644 --- a/docs/document.rst +++ b/docs/document.rst @@ -1,5 +1,23 @@ .. include:: header.rst + +.. |PyMuPDFLayoutMode_Ignored| raw:: html + + use_layout() must be False + +.. |PyMuPDFLayoutMode_Valid| raw:: html + + + +.. |PyMuPDFLayoutMode_EmptyList| raw:: html + + Only if use_layout() is False + +.. |PyMuPDFLayoutMode_Unavailable| raw:: html + + Only if use_layout() is False + + .. _Document: ================ @@ -120,6 +138,9 @@ For details on **embedded files** refer to Appendix 3. :meth:`Document.set_xml_metadata` PDF only: create or update document XML metadata :meth:`Document.subset_fonts` PDF only: create font subsets :meth:`Document.switch_layer` PDF only: activate OC configuration +:meth:`Document.to_json` PDF, Image & Office documents: convert the document to JSON +:meth:`Document.to_markdown` PDF, Image & Office documents: convert the document to Markdown +:meth:`Document.to_text` PDF, Image & Office documents: convert the document to plain text :meth:`Document.tobytes` PDF only: writes document to memory :meth:`Document.xref_copy` PDF only: copy a PDF dictionary to another :data:`xref` :meth:`Document.xref_get_key` PDF only: get the value of a dictionary key @@ -148,6 +169,7 @@ For details on **embedded files** refer to Appendix 3. :attr:`Document.permissions` permissions to access the document :attr:`Document.pagemode` PDF PageMode value :attr:`Document.pagelayout` PDF PageLayout value +:attr:`Document.use_layout` whether the PyMuPDF Layout module is used for page analysis :attr:`Document.version_count` PDF count of versions ======================================= ========================================================== @@ -295,6 +317,239 @@ For details on **embedded files** refer to Appendix 3. Activates the ON / OFF states of OCGs as defined in the identified layer. If ``as_default=True``, then additionally all layers, including the standard one, are merged and the result is written back to the standard layer, and **all optional layers are deleted**. + .. method:: to_markdown(detect_bg_color: bool = True, \ + dpi: int = 150, \ + embed_images: bool = False, \ + extract_words: bool = False, \ + filename: str | None = None, \ + fontsize_limit: float = 3, \ + footer: bool = True, \ + force_ocr: bool = False, \ + force_text: bool = True, \ + graphics_limit: int = None, \ + hdr_info: Any = None, \ + header: bool = True, \ + ignore_alpha: bool = False, \ + ignore_code: bool = False, \ + ignore_graphics: bool = False, \ + ignore_images: bool = False, \ + image_format: str = "png", \ + image_path: str = "", \ + image_size_limit: float = 0.05, \ + margins: float | list = 0, \ + ocr_dpi: int = 300, \ + ocr_function: callable = None, \ + ocr_language: str = "eng", \ + page_chunks: bool = False, \ + page_height: float = None, \ + page_separators: bool = False, \ + page_width: float = 612, \ + pages: list | range | None = None, \ + show_progress: bool = False, \ + table_strategy: str = "lines_strict", \ + use_glyphs: bool = False, \ + use_ocr: bool = True, \ + write_images: bool = False) -> str | list[dict] + + Reads the pages of the file and outputs the text of its pages in |Markdown| format. How this should happen in detail can be influenced by a number of parameters. Please note that **support for building page chunks** from the |Markdown| text is supported. + + :arg bool detect_bg_color: |PyMuPDFLayoutMode_Ignored| does a simple check for the general background color of the pages (default is ``True``). If any text or vector has this color it will be ignored. May increase detection accuracy. + + :arg int dpi: specify the desired image resolution in dots per inch. Relevant only if `write_images=True` or `embed_images=True`. Default value is 150. + + :arg bool embed_images: like `write_images`, but images will be included in the markdown text as base64-encoded strings. Mutually exclusive with `write_images` and ignores `image_path`. This may drastically increase the size of your markdown text. + + :arg bool extract_words: |PyMuPDFLayoutMode_Ignored| a value of `True` enforces `page_chunks=True` and adds key "words" to each page dictionary. Its value is a list of words as delivered by PyMuPDF's `Page` method `get_text("words")`. The sequence of the words in this list is the same as the extracted text. + + :arg str filename: Overwrites or sets the desired image file name of written images. Useful when the document is provided as a memory object (which has no inherent file name). + + :arg float fontsize_limit: |PyMuPDFLayoutMode_Ignored| limit the font size to consider for text extraction. If the font size is lower than what is set then the text won't be considered for extraction. Default is `3`, meaning only text with a font size `>= 3` will be considered for extraction. + + :arg bool footer: |PyMuPDFLayoutMode_Valid| boolean to switch on/off page footer content. This parameter controls whether to include or omit footer text from all the document pages. Useful if the document has repetitive footer content which doesn't add any value to the overall extraction data. Default is `True` meaning that footer content will be considered. + + :arg bool force_ocr: |PyMuPDFLayoutMode_Valid| if `True`, OCR will be applied to all pages regardless of their content. + + This may be useful for documents which are known to be image-based and thus profit from OCR, but which do not meet the default criteria for applying OCR. Default is `False` meaning that OCR will only be applied to pages which meet the default criteria. + + .. warning:: + Requires that either one of the default supported OCR engines is installed or `ocr_function` specifies a callable OCR function. Otherwise, an exception will be raised. + + :arg bool force_text: generate text output even when overlapping images / graphics. This text then appears after the respective image. + + :arg int graphics_limit: |PyMuPDFLayoutMode_Ignored| use this to limit dealing with excess amounts of vector graphics elements. Scientific documents, or pages simulating text via graphics commands may contain tens of thousands of these objects. As vector graphics are analyzed for multiple purposes, runtime may quickly become intolerable. With this parameter, all vector graphics will be ignored if their count exceeds the threshold. + + :arg hdr_info: |PyMuPDFLayoutMode_Ignored| use this if you want to provide your own header detection logic. This may be a callable or an object having a method named `get_header_id`. It must accept a text span (a span dictionary as contained in :meth:`~.extractDICT`) and a keyword parameter "page" (which is the owning :ref:`Page ` object). It must return a string "" or up to 6 "#" characters followed by 1 space. If omitted (`None`), a full document scan will be performed to find the most popular font sizes and derive header levels based on them. To completely avoid this behavior specify `hdr_info=lambda s, page=None: ""` or `hdr_info=False`. + + :arg bool header: |PyMuPDFLayoutMode_Valid| boolean to switch on/off page header content. This parameter controls whether we want to include or omit the header content from all the document pages. Useful if the document has repetitive header content which doesn't add any value to the overall extraction data. Default is `True` meaning that header content will be considered. + + :arg bool ignore_alpha: |PyMuPDFLayoutMode_Ignored| if ``True`` includes text even when completely transparent. Default is ``False``: transparent text will be ignored which usually increases detection accuracy. + + :arg bool ignore_code: if `True` then mono-spaced text lines do not receive special formatting. Code blocks will no longer be generated. This value is set to `True` if `extract_words=True` is used. + + :arg bool ignore_graphics: |PyMuPDFLayoutMode_Ignored| (New in v.0.0.20) Disregard vector graphics on the page. This may help detecting text correctly when pages are very crowded (often the case for documents representing presentation slides). Also speeds up processing time. This automatically prevents table detection. + + :arg bool ignore_images: |PyMuPDFLayoutMode_Ignored| (New in v.0.0.20) Disregard images on the page. This may help detecting text correctly when pages are very crowded (often the case for documents representing presentation slides). Also speeds up processing time. + + :arg str image_format: specify the desired image format via its extension. Default is "png" (portable network graphics). Another popular format may be "jpg". Possible values are all :ref:`supported output formats `. + + :arg str image_path: store images in this folder. Relevant if `write_images=True`. Default is the path of the script directory. + + :arg float image_size_limit: |PyMuPDFLayoutMode_Ignored| this must be a ``0 <= value < 1``. Images are ignored if `width / page.rect.width <= image_size_limit` or `height / page.rect.height <= image_size_limit`. For instance, the default value 0.05 means that to be considered for inclusion, an image's width and height must be larger than 5% of the page's width and height, respectively. + + :arg float,list margins: |PyMuPDFLayoutMode_Ignored| a float or a sequence of 2 or 4 floats specifying page borders. Only objects inside the margins will be considered for output. + + * `margin=f` yields `(f, f, f, f)` for `(left, top, right, bottom)`. + * `(top, bottom)` yields `(0, top, 0, bottom)`. + * To always read full pages **(default)**, use `margins=0`. + + :arg int ocr_dpi: |PyMuPDFLayoutMode_Valid| specify the desired image resolution in dots per inch for applying OCR to the intermediate image of the page. Default value is 300. Only relevant if the page has been determined to profit from OCR (no or few text, most of the page covered by images or character-like vectors, etc.). Larger values do not usually increase the OCR precision. There also is a risk of over-sharpening the image which may decrease OCR precision. So the default value should probably be sufficiently high - in many cases you should see satisfactory results already with values of 150 or 200. Be aware that processing time and memory requirements grow quadratically with this value (an O(ocr_dpi²) impact). + + :arg callable ocr_function: |PyMuPDFLayoutMode_Valid| if you want to provide your own :ref:`OCR function `, specify it here. If omitted (`None`), one of the available built-in OCR engines will be used. + + :arg str ocr_language: |PyMuPDFLayoutMode_Valid| specify the language to be used by the Tesseract OCR engine. Default is "eng" (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign "+", for example "eng+deu" for English and German. + + :arg bool page_chunks: if `True` the output will be a list of `Document.page_count` dictionaries (one per page). Each dictionary has the following structure: + + - **"metadata"** - a dictionary consisting of the document's metadata :attr:`Document.metadata`, enriched with additional keys **"file_path"** (the file name), **"page_count"** (number of pages in document), and **"page_number"** (1-based page number). + + - **"toc_items"** - a list of Table of Contents items pointing to this page. Each item of this list has the format `[lvl, title, pagenumber]`, where `lvl` is the hierarchy level, `title` a string and `pagenumber` as a 1-based page number. + + - **"tables"** - |PyMuPDFLayoutMode_EmptyList| a list of tables on this page. Each item is a dictionary with keys "bbox", "row_count" and "col_count". Key "bbox" is a `pymupdf.Rect` in tuple format of the table's position on the page. + + - **"images"** - |PyMuPDFLayoutMode_EmptyList| a list of images on the page. This a copy of page method :meth:`Page.get_image_info`. + + - **"graphics"** - |PyMuPDFLayoutMode_EmptyList| a list of vector graphics rectangles on the page. This is a list of boundary boxes of clustered vector graphics as delivered by method :meth:`Page.cluster_drawings`. + + - **"text"** - page content as |Markdown| text. + + - **"words"** - |PyMuPDFLayoutMode_EmptyList| if `extract_words=True` was used. This is a list of tuples `(x0, y0, x1, y1, "wordstring", bno, lno, wno)` as delivered by `page.get_text("words")`. The **sequence** of these tuples however is the same as produced in the markdown text string and thus honors multi-column text. This is also true for text in tables: words are extracted in the sequence of table row cells. + + - **"text"** - page content as |Markdown| text. + + - **"page_boxes"** - |PyMuPDFLayoutMode_Valid| a list of dictionaries representing the layout boundary boxes. Each dictionary has the following structure:: + + { + "index": int, # 0-based integer index of the box in reading sequence + "class": str, # one of "text", "picture", "table", etc. + "bbox": [x0, y0, x1, y1], # boundary box coordinates + "pos": (start, stop), # 0-based integers: bbox_text = chunk["text"][start:stop] + } + + See: :ref:`box classes ` + + :arg float page_height: specify a desired page height. For relevance see the `page_width` parameter. If using the default `None`, the document will appear as one large page with a width of `page_width`. Consequently in this case, no markdown page separators will occur (except the final one), respectively only one page chunk will be returned. + + :arg bool page_separators: if ``True`` inserts a string ``--- end of page=n ---`` at the end of each page output. Intended for debugging purposes. The page number is 0-based. The separator string is wrapped with line breaks. Default is ``False``. + + :arg float page_width: specify a desired page width. This is ignored for documents with a fixed page width like PDF, XPS etc. **Reflowable** documents however, like e-books, office [#f2]_ or text files have no fixed page dimensions. They by default are assumed to have Letter format width (612) and an **unlimited** page height. This means that the **full document is treated as one large page.** + + :arg list pages: optional, the pages to consider for output (caution: specify 0-based page numbers). If omitted (`None`) all pages are processed. Any Python sequence with integer items is accepted. The sequence is sorted and processed to only contain unique items. + + :arg bool show_progress: Default is `False`. A value of `True` displays a progress bar as pages are being converted. Package `tqdm `_ is used if installed, otherwise the built-in text based progress bar is used. + + :arg str table_strategy: |PyMuPDFLayoutMode_Ignored| see: :meth:`table detection strategy `. Default is `"lines_strict"` which ignores background colors. In some occasions, other strategies may be more successful, for example `"lines"` which uses all vector graphics objects for detection. + + :arg bool use_glyphs: |PyMuPDFLayoutMode_Ignored| (New in v.0.0.19) Default is `False`. A value of `True` will use the glyph number of the characters instead of the character itself if the font does not store the Unicode value. + + :arg bool use_ocr: |PyMuPDFLayoutMode_Valid| use :ref:`OCR capability ` to help analyse the page. This will OCR pages as determined by the default criteria. + + :arg bool write_images: when encountering images or vector graphics, images will be created from the respective page area and stored in the specified folder. |Markdown| references will be generated pointing to these images. Any text contained in these areas will not be included in the text output (but appear as part of the images). Therefore, if for instance your document has text written on full page images, make sure to set this parameter to `False`. + + If using :ref:`PyMuPDF Layout `, boundary boxes that are classified as "picture" by the layout module will be treated as images - independent from the mixture of text, images or vector graphics they may be covering. If `force_text=True` is used, text will still be extracted from these areas and included in the output after the respective image reference. + + :returns: Either a string of the combined text of all selected document pages, or a list of dictionaries if `page_chunks=True`. + + + .. method:: to_json(**kwargs) -> str + + Parses the document and the specified pages and converts the result into a `JSON formatted string `_. + + :arg bool use_ocr: |PyMuPDFLayoutMode_Valid| use :ref:`OCR capability ` to help analyse the page. + + :arg str ocr_language: |PyMuPDFLayoutMode_Valid| specify the language to be used by the Tesseract OCR engine. Default is "eng" (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign "+", for example "eng+deu" for English and German. + + :arg int ocr_dpi: |PyMuPDFLayoutMode_Valid| specify the desired image resolution in dots per inch for applying OCR to the intermediate image of the page. Default value is 400. Only relevant if the page has been determined to profit from OCR (no or few text, most of the page covered by images or character-like vectors, etc.). Large values may increase the OCR precision but increase memory requirements and processing time. There also is a risk of over-sharpening the image which may decrease OCR precision. So the default value should probably be sufficiently high. + + :arg int image_dpi: specify the desired image resolution in dots per inch. Default value is 150. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + + :arg str image_format: specify the desired image format via its extension. Default is "png" (portable network graphics). Another popular format may be "jpg". Possible values are all :ref:`supported output formats `. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + + :arg str image_path: store images in this folder. Relevant if `write_images=True`. Default is the path of the script directory. Page areas classified as "picture" will be written as image files to the specified location. The image file names will be of the format `{image_path}/{filename}-pagenumber-image_number.{image_format}`. + + :arg bool force_text: generate text output for text that is written upon areas that are classified as "picture" by the layout module. This may be especially be useful when picture content is not stored. + + :arg bool show_progress: display a progress bar during processing. + + :arg bool embed_images: store image binaries for "picture" boundary boxes. Base64-encoded images are included in the JSON output. Ignores `image_path` if used. This may drastically increase the size of your JSON text. + + :arg bool write_images: store image files "picture" boundary boxes. When encountering images, image files will be created from the respective page area and stored in the specified folder. Any text contained in these areas will still be included in the text output. + + :arg list pages: optional, the pages to consider for output (caution: specify 0-based page numbers). If omitted (`None`) all pages are processed. Specify any valid Python sequence containing integers between `0` and `page_count - 1`. + + :rtype: str + + See `JSON Schema `_ for the structure of the output JSON string. + + + + + + .. method:: to_text(**kwargs) -> str + + Reads the pages of the file and outputs the text of its pages in plain text (|TXT|) format. + + :arg Document,str doc: the file, to be specified either as a file path string, or as a |PyMuPDF| :class:`Document` (created via `pymupdf.open`). In order to use `pathlib.Path` specifications, Python file-like objects, documents in memory etc. you **must** use a |PyMuPDF| :class:`Document`. + + :arg bool use_ocr: |PyMuPDFLayoutMode_Valid| use :ref:`OCR capability ` to help analyse the page. + + :arg str ocr_language: |PyMuPDFLayoutMode_Valid| specify the language to be used by the Tesseract OCR engine. Default is "eng" (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign "+", for example "eng+deu" for English and German. + + :arg int ocr_dpi: |PyMuPDFLayoutMode_Valid| specify the desired image resolution in dots per inch for applying OCR to the intermediate image of the page. Default value is 400. Only relevant if the page has been determined to profit from OCR (no or few text, most of the page covered by images or character-like vectors, etc.). Large values may increase the OCR precision but increase memory requirements and processing time. There also is a risk of over-sharpening the image which may decrease OCR precision. So the default value should probably be sufficiently high. + + :arg bool header: boolean to switch on/off page header content. This parameter controls whether to include or omit the header content from all the document pages. Useful if the document has repetitive header content which doesn't add any value to the overall extraction data. Default is `True` meaning that header content will be written. + + :arg bool footer: boolean to switch on/off page footer content. This parameter controls whether to include or omit the footer content from all the document pages. Useful if the document has repetitive footer content which doesn't add any value to the overall extraction data. Default is `True` meaning that footer content will be written. + + :arg bool ignore_code: if `True` then mono-spaced text lines do not receive special formatting. No blocks will be written and text lines will be written continuously. + + :arg list pages: optional, the pages to consider for output (caution: specify 0-based page numbers). If omitted (`None`) all pages are processed. Any Python sequence with integer items is accepted. The sequence is sorted and processed to only contain unique items. + + :arg bool force_text: generate text output also when overlapping images / graphics. This text then appears after the respective image reference. Images (i.e. "picture" areas) however will not be written to the text output but appear as a text line in the output like `==> picture [width x height] <==`. + + :arg bool show_progress: Default is `False`. A value of `True` displays a progress bar as pages are being converted. Package `tqdm `_ is used if installed, otherwise the built-in text based progress bar is used. + + :arg bool page_chunks: if `True` the output will be a list of `Document.page_count` dictionaries (one per page). Each dictionary has the following structure: + + - **"metadata"** - a dictionary consisting of the document's metadata :attr:`Document.metadata`, enriched with additional keys **"file_path"** (the file name), **"page_count"** (number of pages in document), and **"page_number"** (1-based page number). + + - **"toc_items"** - a list of Table of Contents items pointing to this page. Each item of this list has the format `[lvl, title, pagenumber]`, where `lvl` is the hierarchy level, `title` a string and `pagenumber` as a 1-based page number. + + - **"tables"** - empty list. + - **"images"** - empty list. + - **"graphics"** - empty list. + - **"words"** - empty list. + + - **"text"** - page content as plain text. + + - **"page_boxes"** - a list of dictionaries representing the layout boundary boxes. Each dictionary has the following structure:: + + { + "index": int, # 0-based integer index of the box in reading sequence + "class": str, # one of "text", "picture", "table", etc. + "bbox": [x0, y0, x1, y1], # boundary box coordinates + "pos": (start, stop), # 0-based integers: bbox_text = chunk["text"][start:stop] + } + + See: :ref:`box classes ` + + + .. method:: use_layout(yes: bool = True) + + Switch on/off the use of the :ref:`PyMuPDF Layout module `. + + If `yes=True` (default), the layout module will be used for page analysis for optimal results. If `yes=False`, the layout module will not be used. + + .. method:: add_ocg(name, config=-1, on=True, intent="View", usage="Artwork") * New in v1.18.3 @@ -2291,6 +2546,34 @@ Obviously, similar ways can be found in more general situations. Just make sure >>> # put copied pages in front of doc1 >>> doc1.insert_pdf(doc2, from_page=21, to_page=25, start_at=0) + +.. _pymupdf4llm-api-boxclasses: + +.. note:: + + **About box classes** + + If `page_chunks = True` the return objects for :meth:`Document.to_markdown` & :meth:`Document.to_text` contains a list of dictionaries representing the layout boundary boxes `page_boxes`, within that a key ``class`` indicates the type of box content therein. + + The return object for :meth:`Document.to_json` contains a similar key called ``boxclass``. + + The possible string values are for this ``class`` / ``boxclass`` key are: + + .. code-block:: bash + + text + picture + table + caption + title + section-header + page-header + page-footer + list-item + footnote + formula + + Other Examples ---------------- **Extract all page-referenced images of a PDF into separate PNG files**:: diff --git a/docs/grounding.rst b/docs/grounding.rst new file mode 100644 index 000000000..e9f79fc67 --- /dev/null +++ b/docs/grounding.rst @@ -0,0 +1,1036 @@ +.. include:: header.rst + +.. _Grounding: + + +============================== +PyMuPDF & Grounding +============================== + +Grounding is about connecting output back to source evidence. This means ensuring that every extracted claim is traceable to a specific, verifiable location in the source document. In document data extraction, it is the practice of anchoring what a model says back to what a document actually contains — not what seems plausible given the context, but what is factually, verifiably present in the source. + +This section explores 3 grounding patterns and how |PyMuPDF| uses **deterministic reasoning** and grounding techniques to extract and connect information from documents. + +|PyMuPDF| helps you build **verifiable evidence objects**, that is: extracted content linked to its source document, page, and region, with enough context to inspect it and ground its origin. + +Grounding: 3 patterns +----------------------- + +These examples illustrate three common grounding patterns used in document data extraction. The scripts provided demonstrate each pattern and are self-contained for easy experimentation. + +---- + +1. Answer and citation grounding +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +This applies when an LLM answers questions about a document. The answer carries verbatim quotes, each one located and highlighted in the source PDF, so every claim can be traced to the passage behind it. Unverifiable quotes are flagged rather than trusted. + +Our example highlights how an answer is **grounded with citations directly linked to the source document** and we use a simulated LLM to demonstrate this process. + + +.. image:: images/grounding/citation.png + :alt: Citation grounding example + +.. note:: + + Citation example highlights source text by visually marking the grounding locations in the PDF. + +The code example creates a basic PDF with sample text content and asks PyMuPDF to highlight the areas that support the answer from the LLM response. + + + + +.. code-block:: python + + """ + Answer and citation grounding with PyMuPDF + ========================================== + + Goal: The goal of the answer and citation grounding sample is to make an AI-generated + answer checkable against the source PDF. Every claim in the answer should point to the + exact passage that supports it, and that passage should be highlighted in the original + document so a person can confirm it in seconds instead of rereading the whole file. + + Pipeline: + 1. Extract text chunks from a PDF, keeping page number + bounding box. + 2. Retrieve the chunks relevant to a question (simple keyword scoring here; + swap in your vector store). + 3. Ask an LLM to answer ONLY from those chunks and return verbatim quotes + with the chunk id they came from (LLM call is a pluggable stub). + 4. Locate each quote on the page (exact search, then a word-level fallback + that tolerates line breaks, hyphenation and whitespace differences). + 5. Highlight the evidence in a copy of the PDF, attach a note with the + citation number, and render a PNG snippet of each cited region. + + Requires: pip install pymupdf + """ + + import json + import re + import unicodedata + from dataclasses import dataclass, field + + import pymupdf + + # -------------------------------------------------------------------------- + # 1. Extraction: text chunks with page + position + # -------------------------------------------------------------------------- + @dataclass + class Chunk: + id: str + page: int # 0-based page index + bbox: tuple # (x0, y0, x1, y1) in PDF points + text: str + + + def extract_chunks(doc: pymupdf.Document) -> list[Chunk]: + """One chunk per text block. Use paragraphs/sections in production.""" + chunks = [] + for page in doc: + for x0, y0, x1, y1, text, block_no, block_type in page.get_text("blocks", sort=True): + if block_type != 0 or not text.strip(): # 0 = text, 1 = image + continue + chunks.append(Chunk( + id=f"p{page.number}-b{block_no}", + page=page.number, + bbox=(x0, y0, x1, y1), + text=" ".join(text.split()), + )) + return chunks + + + # -------------------------------------------------------------------------- + # 2. Retrieval (placeholder: keyword overlap) + # -------------------------------------------------------------------------- + def tokens(s: str) -> list[str]: + text = unicodedata.normalize("NFKC", s).casefold() + out = re.findall(r"[^\W_]+", text, flags=re.UNICODE) + hangul = "".join(re.findall(r"[가-힣]", text)) + out.extend(f"ko:{hangul[i:i + 2]}" for i in range(len(hangul) - 1)) + if len(hangul) == 1: + out.append(f"ko:{hangul}") + return out + + + def _normalise_match(text: str) -> str: + text = unicodedata.normalize("NFKC", text).casefold() + return "".join(char for char in text if char.isalnum()) + + + def retrieve(chunks: list[Chunk], question: str, k: int = 4) -> list[Chunk]: + q = set(tokens(question)) + scored = sorted(chunks, key=lambda c: len(q & set(tokens(c.text))), reverse=True) + return [c for c in scored[:k] if q & set(tokens(c.text))] + + + # -------------------------------------------------------------------------- + # 3. LLM call (stub): must return verbatim quotes tied to chunk ids + # -------------------------------------------------------------------------- + PROMPT = """Answer the question using ONLY the sources below. + Return JSON: {{"answer": "... [1] ... [2]", + "citations": [{{"n": 1, "chunk_id": "...", "quote": "exact words copied from the source"}}]}} + Quotes must be copied verbatim (no paraphrasing), ideally one sentence. + + Question: {question} + + Sources: + {sources}""" + + + def call_llm(prompt: str) -> str: + """Replace with your provider, e.g. the Anthropic SDK: + + import anthropic + msg = anthropic.Anthropic().messages.create( + model="claude-sonnet-4-6", max_tokens=1024, + messages=[{"role": "user", "content": prompt}]) + return msg.content[0].text + """ + raise NotImplementedError + + + def answer_with_citations(question: str, context: list[Chunk], llm=call_llm) -> dict: + sources = "\n".join(f"[{c.id}] {c.text}" for c in context) + raw = llm(PROMPT.format(question=question, sources=sources)) + raw = re.sub(r"^```(?:json)?|```$", "", raw.strip(), flags=re.M) + return json.loads(raw) + + + # -------------------------------------------------------------------------- + # 4. Locate the quote on the page + # -------------------------------------------------------------------------- + def locate_quote(page: pymupdf.Page, quote: str, clip=None) -> list: + """Return quads covering `quote`, or [] if it cannot be found.""" + quote = " ".join(quote.split()) + + # a) Exact search (case-insensitive, spans line breaks). + quads = page.search_for(quote, clip=clip, quads=True) + if quads: + return quads + + # b) Word-sequence fallback: normalise words on both sides and look for + # the quote as a contiguous run. Handles hyphenation, ligatures, + # punctuation and whitespace differences. + words = page.get_text("words", clip=clip, sort=True) # (x0,y0,x1,y1,word,...) + page_words = [(_normalise_match(w[4]), w) for w in words if _normalise_match(w[4])] + target = [t for t in (_normalise_match(w) for w in quote.split()) if t] + if not target: + return [] + + joined = [pw[0] for pw in page_words] + n = len(target) + for i in range(len(joined) - n + 1): + if joined[i:i + n] == target: + return [pymupdf.Rect(w[:4]).quad for _, w in page_words[i:i + n]] + + # Korean spacing and PDF word segmentation can differ. Compare the + # normalized character stream while retaining each character's word box. + target_text = _normalise_match(quote) + char_stream = "".join(word for word, _ in page_words) + start = char_stream.find(target_text) + if target_text and start >= 0: + char_word_indices = [ + index for index, (word, _) in enumerate(page_words) for _ in word + ] + matched_indices = dict.fromkeys(char_word_indices[start:start + len(target_text)]) + return [pymupdf.Rect(page_words[index][1][:4]).quad for index in matched_indices] + + # c) Partial match: longest leading run of the quote (min 4 words). + for size in range(n - 1, 3, -1): + head = target[:size] + for i in range(len(joined) - size + 1): + if joined[i:i + size] == head: + return [pymupdf.Rect(w[:4]).quad for _, w in page_words[i:i + size]] + return [] + + + # -------------------------------------------------------------------------- + # 5. Highlight, annotate and render evidence + # -------------------------------------------------------------------------- + @dataclass + class Grounding: + n: int + chunk_id: str + quote: str + page: int | None = None + rects: list = field(default_factory=list) + status: str = "not_found" # exact | fallback_to_chunk | not_found + snippet: str | None = None + + + def ground_citations(doc, result: dict, chunks: list[Chunk], + out_pdf: str, snippet_prefix: str | None = None) -> list[Grounding]: + by_id = {c.id: c for c in chunks} + grounded = [] + + for cit in result.get("citations", []): + g = Grounding(n=cit["n"], chunk_id=cit["chunk_id"], quote=cit["quote"]) + chunk = by_id.get(g.chunk_id) + if chunk is None: # hallucinated source id + grounded.append(g) + continue + + page = doc[chunk.page] + clip = pymupdf.Rect(chunk.bbox) + (-2, -2, 2, 2) # search inside the cited chunk + quads = locate_quote(page, g.quote, clip=clip) or locate_quote(page, g.quote) + + if quads: + annot = page.add_highlight_annot(quads) + g.status = "exact" + g.rects = [q.rect for q in quads] + else: + # Quote not verifiable: mark the whole chunk so a reviewer can check. + annot = page.add_rect_annot(clip) + annot.set_colors(stroke=(1, 0.5, 0)) + g.status = "fallback_to_chunk" + g.rects = [clip] + + annot.set_info(title=f"Citation [{g.n}]", content=g.quote) + annot.update() + g.page = page.number + + if snippet_prefix: + region = pymupdf.Rect(g.rects[0]) + for r in g.rects[1:]: + region |= r + region = (region + (-20, -20, 20, 20)) & page.rect + pix = page.get_pixmap(clip=region, dpi=150) # rendered after annotation + g.snippet = f"{snippet_prefix}_cite{g.n}.png" + pix.save(g.snippet) + + grounded.append(g) + + doc.save(out_pdf, garbage=3, deflate=True) + return grounded + + + # -------------------------------------------------------------------------- + # End-to-end + # -------------------------------------------------------------------------- + def run(pdf_path: str, question: str, llm=call_llm, + out_pdf: str = "grounded.pdf", snippet_prefix: str | None = "evidence") -> dict: + doc = pymupdf.open(pdf_path) + chunks = extract_chunks(doc) + context = retrieve(chunks, question) + result = answer_with_citations(question, context, llm=llm) + grounded = ground_citations(doc, result, chunks, out_pdf, snippet_prefix) + doc.close() + + return { + "answer": result["answer"], + "citations": [ + {"n": g.n, "page": None if g.page is None else g.page + 1, + "status": g.status, "quote": g.quote, "snippet": g.snippet, + "boxes": [tuple(round(v, 1) for v in r) for r in g.rects]} + for g in grounded + ], + "annotated_pdf": out_pdf, + } + + + # -------------------------------------------------------------------------- + # Self-contained demo (creates a sample PDF and uses a fake LLM) + # -------------------------------------------------------------------------- + if __name__ == "__main__": + sample = pymupdf.open() + page = sample.new_page() + page.insert_textbox(pymupdf.Rect(72, 72, 520, 300), ( + "Lease Agreement\n\n" + "The initial term of this lease is five years, starting on 1 March 2026. " + "The tenant may terminate the lease with six months' written notice.\n\n" + "Rent is payable monthly in advance and is indexed annually to the " + "consumer price index." + ), fontsize=12) + sample.save("sample_lease.pdf") + + def fake_llm(prompt: str) -> str: + ids = re.findall(r"^\[(p\d+-b\d+)\] (.*)$", prompt, flags=re.M) + term_id = next(i for i, t in ids if "five years" in t) + return json.dumps({ + "answer": "The lease runs for five years [1] and can be ended " + "with six months' written notice [2].", + "citations": [ + {"n": 1, "chunk_id": term_id, + "quote": "The initial term of this lease is five years"}, + {"n": 2, "chunk_id": term_id, + "quote": "may terminate the lease with six months' written notice"}, + ], + }) + + report = run("sample_lease.pdf", "How long is the lease and how can it be terminated?", + llm=fake_llm, out_pdf="sample_lease_grounded.pdf") + print(json.dumps(report, indent=2)) + + +---- + +2. Extracted-data grounding +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +This applies when values are pulled out of a document, whether by AI, OCR or rules. Each value is marked on the page it was read from and checked against the text actually there, so a reviewer sees the source rather than a bare field. + +Our example demonstrates how extracted data is grounded by marking each value on the page it was read from as well as highlighting mismatches or ambiguous data. + + +.. image:: images/grounding/extraction.png + :alt: Extraction grounding example + +.. note:: + + The colors used in the grounding PDF indicate the status of each extracted value: + + - **Green**: the value matches the source text. + - **Orange**: the value is ambiguous or partially matches. + - **Red**: the value does not match the source text. + +The validation produces a grounding PDF example with matches for validated areas (green), ambiguous areas (orange), and mismatches (red). + + +.. code-block:: python + + """ + Extracted-data grounding with PyMuPDF + ===================================== + + Goal: every extracted value (total, date, rate, ...) carries evidence of + WHERE it came from in the source PDF, and that evidence is checked. + + Two common situations are covered: + + A. Values come from an extractor that gives NO coordinates + (an LLM, a regex over plain text, a database record). + -> Locate each value on the page, using its label as an anchor to pick + the right occurrence when the value appears more than once. + + B. Values come from an OCR / document-AI service that DOES give + coordinates, but in its own space (inches, image pixels, a rotated view). + -> Convert those coordinates to PDF space. + + Both paths then: + - verify the value really is inside the box (re-read the text there), + - mark it on a copy of the PDF with a labelled, color-coded box, + - render an evidence snippet, + - return a JSON-ready record: field, value, page, bbox, status. + + Requires: pip install pymupdf + """ + + import json + import re + from dataclasses import dataclass, field, asdict + + import pymupdf + + # Coordinate notes (PyMuPDF): + # * Text extraction, search and annotation methods use UNROTATED page + # coordinates (points, origin top-left). + # * page.rect, rendered pixmaps and get_pixmap(clip=...) use the page AS + # DISPLAYED (rotated). + # * page.derotation_matrix maps displayed coordinates -> unrotated ones. + + GREEN, AMBER, RED = (0.1, 0.6, 0.2), (1.0, 0.6, 0.0), (0.85, 0.1, 0.1) + + + @dataclass + class Extracted: + """One value produced by any extractor.""" + field: str + value: str + label: str | None = None # text near the value, e.g. "Total due" + page: int | None = None # 0-based, if the extractor knows it + bbox: tuple | None = None # already in PDF points (unrotated), if known + + + @dataclass + class GroundedValue: + field: str + value: str + page: int | None = None + bbox: tuple | None = None + status: str = "not_found" # verified | ambiguous | mismatch | not_found + found_text: str | None = None + candidates: int = 0 + snippet: str | None = None + notes: list = field(default_factory=list) + + + # -------------------------------------------------------------------------- + # Value normalisation: "€ 1.250,00", "1,250.00" and "1250" should all match + # -------------------------------------------------------------------------- + def normalise(value: str) -> str: + v = value.strip().lower() + v = re.sub(r"[€$£%]|usd|eur|gbp", "", v) + num = re.sub(r"[\s'’]", "", v) + if re.fullmatch(r"[-+(]?[\d.,]+\)?", num): + num = num.strip("()+") + # decide which separator is the decimal one (last one with 1-2 digits after it) + m = re.search(r"[.,](\d{1,2})$", num) + if m: + whole, dec = num[:m.start()], m.group(1) + else: + whole, dec = num, "" + whole = re.sub(r"[.,]", "", whole) + return f"{int(whole or 0)}.{dec.ljust(2, '0')}" if dec else str(int(whole or 0)) + return re.sub(r"[^a-z0-9]", "", v) + + + def same_value(a: str, b: str) -> bool: + na, nb = normalise(a), normalise(b) + if na == nb: + return True + try: # 1250 vs 1250.00 + return abs(float(na) - float(nb)) < 1e-9 + except ValueError: + return False + + + # -------------------------------------------------------------------------- + # Locating a value on a page (path A) + # -------------------------------------------------------------------------- + def word_runs(page: pymupdf.Page, value: str) -> list[pymupdf.Rect]: + """Find `value` as a run of 1..n consecutive words, comparing normalised text. + Catches formatting differences an exact search would miss.""" + words = page.get_text("words", sort=True) + n_tokens = max(1, len(value.split())) + hits = [] + for size in {n_tokens, 1, 2, 3}: + for i in range(len(words) - size + 1): + run = words[i:i + size] + # keep runs on one line + if len({(w[5], w[6]) for w in run}) > 1: + continue + text = " ".join(w[4] for w in run) + if same_value(text.strip(":;"), value): + r = pymupdf.Rect(run[0][:4]) + for w in run[1:]: + r |= pymupdf.Rect(w[:4]) + hits.append(r) + # de-duplicate overlapping hits + unique = [] + for r in hits: + if not any(abs(r & u) > 0.8 * min(abs(r), abs(u)) for u in unique): + unique.append(r) + return unique + + + def distance_from_label(value_rect: pymupdf.Rect, label_rect: pymupdf.Rect) -> float: + """Values usually sit to the right of, or just below, their label.""" + dx = value_rect.x0 - label_rect.x1 + dy = value_rect.y0 - label_rect.y0 + same_line = abs((value_rect.y0 + value_rect.y1) - (label_rect.y0 + label_rect.y1)) / 2 < 4 + if same_line and dx >= -2: + return dx # best: same line, to the right + if 0 <= dy < 60: + return 200 + dy + abs(value_rect.x0 - label_rect.x0) # below the label + return 10_000 + abs(dx) + abs(dy) # anywhere else + + + def locate_value(doc, item: Extracted) -> tuple[int | None, pymupdf.Rect | None, int]: + """Return (page, rect, number_of_candidates).""" + pages = [doc[item.page]] if item.page is not None else list(doc) + candidates = [] # (score, page_no, rect) + for page in pages: + rects = page.search_for(item.value) or word_runs(page, item.value) + if not rects: + continue + labels = page.search_for(item.label) if item.label else [] + for r in rects: + score = min((distance_from_label(r, lr) for lr in labels), default=50_000) + candidates.append((score, page.number, r)) + + if not candidates: + return None, None, 0 + candidates.sort(key=lambda c: c[0]) + _, page_no, rect = candidates[0] + return page_no, rect, len(candidates) + + + # -------------------------------------------------------------------------- + # Converting external coordinates to PDF space (path B) + # -------------------------------------------------------------------------- + def from_inches(page: pymupdf.Page, polygon: list[float]) -> tuple: + """Document-AI style polygon [x1,y1,x2,y2,...] in inches on the displayed page + (e.g. Azure Document Intelligence for PDFs).""" + xs, ys = polygon[0::2], polygon[1::2] + shown = pymupdf.Rect(min(xs), min(ys), max(xs), max(ys)) * 72 + return tuple(shown * page.derotation_matrix) + + + def from_pixels(page: pymupdf.Page, box: tuple, dpi: int) -> tuple: + """(x0, y0, x1, y1) in pixels of an image rendered with page.get_pixmap(dpi=dpi), + e.g. boxes returned by an OCR or vision model that received that image.""" + shown = pymupdf.Rect(box) * (72 / dpi) + return tuple(shown * page.derotation_matrix) + + + # -------------------------------------------------------------------------- + # Verification, marking and evidence + # -------------------------------------------------------------------------- + def text_in_box(page: pymupdf.Page, rect: pymupdf.Rect) -> str: + """Words whose centre lies inside rect (more robust than clipping characters).""" + words = [] + for w in page.get_text("words", sort=True): + wr = pymupdf.Rect(w[:4]) + centre = pymupdf.Point((wr.x0 + wr.x1) / 2, (wr.y0 + wr.y1) / 2) + if centre in rect: + words.append(w[4]) + return " ".join(words) + + + def ground_values(doc, items: list[Extracted], out_pdf: str, + snippet_prefix: str | None = None) -> list[GroundedValue]: + results = [] + for item in items: + g = GroundedValue(field=item.field, value=item.value) + + if item.bbox is not None and item.page is not None: # path B + page_no, rect, g.candidates = item.page, pymupdf.Rect(item.bbox), 1 + g.notes.append("coordinates supplied by extractor") + else: # path A + page_no, rect, g.candidates = locate_value(doc, item) + + if rect is None: + results.append(g) + continue + + page = doc[page_no] + g.page, g.bbox = page_no, tuple(round(v, 1) for v in rect) + g.found_text = text_in_box(page, rect + (-1, -1, 1, 1)) + + if not same_value(g.found_text.strip(":;"), item.value): + g.status, color = "mismatch", RED + g.notes.append("text at this location does not match the extracted value") + elif g.candidates > 1 and not item.label: + g.status, color = "ambiguous", AMBER + g.notes.append("value occurs more than once; add a label to disambiguate") + else: + g.status, color = "verified", GREEN + + box = rect + (-2, -2, 2, 2) + annot = page.add_rect_annot(box) + annot.set_colors(stroke=color) + annot.set_border(width=1.5) + annot.set_info(title=item.field, + content=f"{item.field} = {item.value} [{g.status}]") + annot.update() + + # Small visible tag above the box. insert_text expects unrotated + # coordinates; `rotate` keeps the tag upright on rotated pages. + page.insert_text(box.tl + (0, -2), item.field, fontsize=6, + color=color, rotate=page.rotation) + + if snippet_prefix: + # get_pixmap's clip is in DISPLAYED coordinates, so rotate the box first. + region = (box * page.rotation_matrix + (-40, -25, 40, 15)) & page.rect + pix = page.get_pixmap(clip=region, dpi=150) + g.snippet = f"{snippet_prefix}_{item.field}.png" + pix.save(g.snippet) + + results.append(g) + + doc.save(out_pdf, garbage=3, deflate=True) + return results + + + # -------------------------------------------------------------------------- + # Self-contained demo + # -------------------------------------------------------------------------- + if __name__ == "__main__": + # Build a sample two-page PDF; page 2 is rotated like a landscape scan. + doc = pymupdf.open() + p1 = doc.new_page() + p1.insert_text((72, 80), "INVOICE No. INV-2026-0417", fontsize=16) + p1.insert_text((72, 110), "Invoice date: 14 September 2026", fontsize=11) + p1.insert_text((72, 160), "Consulting services 10 h 1,250.00", fontsize=11) + p1.insert_text((72, 180), "Licence renewal 1,250.00", fontsize=11) + p1.insert_text((72, 230), "VAT rate: 20%", fontsize=11) + p1.insert_text((72, 250), "Total due: EUR 3,000.00", fontsize=11) + + p2 = doc.new_page() + p2.insert_text((72, 80), "Payment terms", fontsize=14) + p2.insert_text((72, 110), "IBAN: DE89 3704 0044 0532 0130 00", fontsize=11) + p2.set_rotation(90) + doc.save("sample_invoice.pdf") + doc.close() + + doc = pymupdf.open("sample_invoice.pdf") + page2 = doc[1] + + # Pretend an OCR service rendered page 2 at 200 dpi and returned a pixel box + # for the IBAN (computed here from the real position, for the demo only). + real = page2.search_for("DE89 3704 0044 0532 0130 00")[0] + shown_px = tuple((real * page2.rotation_matrix) * (200 / 72)) + + # A box that points at the wrong place (simulated extractor error). + wrong_box = tuple(doc[0].search_for("Licence")[0]) + + items = [ + # Path A: no coordinates (e.g. from an LLM). Different number formats on purpose. + Extracted("invoice_number", "INV-2026-0417", label="Invoice"), + Extracted("invoice_date", "14 September 2026", label="Invoice date"), + Extracted("vat_rate", "20 %", label="VAT rate"), + Extracted("total_due", "3000", label="Total due"), + Extracted("line_amount", "1250.00"), # appears twice, no label + Extracted("po_number", "PO-99812", label="PO"), # not in the document + # Path B: coordinates from an OCR service, converted to PDF space. + Extracted("iban", "DE89 3704 0044 0532 0130 00", page=1, + bbox=from_pixels(page2, shown_px, dpi=200)), + # Path B with a wrong box (extractor error) -> should be flagged. + Extracted("bic", "COBADEFFXXX", page=0, bbox=wrong_box), + ] + + results = ground_values(doc, items, "sample_invoice_grounded.pdf", + snippet_prefix="evidence") + doc.close() + print(json.dumps([asdict(r) for r in results], indent=2)) + + +---- + + +3. Verification grounding +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +This applies when a document is checked against something trusted, either reference data or an earlier version. Each discrepancy is pinned to its location on the page, so reviewing means looking at the differences instead of reading everything. + +.. image:: images/grounding/verification.png + :alt: Verification grounding example + +.. note:: + + Verification grounding highlights discrepancies between the document and the trusted (CSV) input reference, making it easy to spot errors at a glance. + +Our example shows how a CSV input source is used as the trusted reference, with discrepancies in the document highlighted and linked back to their locations on the page. + +.. code-block:: python + + """ + Verification grounding with PyMuPDF + =================================== + + Goal: check a PDF against something trusted and pin every discrepancy to + its exact location on the page, so a reviewer sees WHAT is wrong and WHERE. + + Two common checks are covered: + + A. PDF vs reference data (Excel/CSV/database/ERP record) + For each expected field: find its label on the page, read the value + printed next to it, compare with the reference, mark the result. + green = matches red = differs (note shows expected vs found) + orange = label found but no value (label missing -> report only) + + B. PDF vs PDF (old version vs new version) + Word-level diff of the two documents. Changed/inserted words are + highlighted in the new version; deleted/replaced words are struck out + in the old version. Every change is reported with page and position. + + Both write annotated PDFs, evidence snippets and a JSON-ready report. + + Requires: pip install pymupdf (reading .xlsx references: pip install openpyxl) + """ + + import csv + import difflib + import json + import re + from dataclasses import dataclass, field, asdict + + import pymupdf + + # Coordinate notes: text search/extraction and annotations use UNROTATED page + # coordinates; page.rect and get_pixmap(clip=...) use the page AS DISPLAYED. + + GREEN, ORANGE, RED, BLUE = (0.1, 0.6, 0.2), (1.0, 0.55, 0.0), (0.85, 0.1, 0.1), (0.2, 0.4, 0.9) + + + # -------------------------------------------------------------------------- + # Shared helpers + # -------------------------------------------------------------------------- + def normalise(value: str) -> str: + """'EUR 1.250,00', '1,250.00' and '1250' -> '1250.00' / '1250'; text -> lowercase alnum.""" + v = re.sub(r"[€$£%]|\b(usd|eur|gbp)\b", "", str(value).strip().lower()) + num = re.sub(r"[\s'’]", "", v) + if re.fullmatch(r"[-+(]?[\d.,]+\)?", num) and re.search(r"\d", num): + neg = num.startswith(("-", "(")) + num = num.strip("()+-") + m = re.search(r"[.,](\d{1,2})$", num) + whole, dec = (num[:m.start()], m.group(1)) if m else (num, "") + whole = re.sub(r"[.,]", "", whole) or "0" + out = f"{int(whole)}.{dec.ljust(2, '0')}" if dec else str(int(whole)) + return f"-{out}" if neg else out + return re.sub(r"[^a-z0-9]", "", v) + + + def same_value(a, b, tolerance: float = 0.0) -> bool: + na, nb = normalise(a), normalise(b) + + if na == nb: + return True + try: + return abs(float(na) - float(nb)) <= tolerance + except ValueError: + return False + + + def mark(page: pymupdf.Page, rect: pymupdf.Rect, color, title: str, note: str): + annot = page.add_rect_annot(rect + (-2, -2, 2, 2)) + annot.set_colors(stroke=color) + annot.set_border(width=1.5) + annot.set_info(title=title, content=note) + annot.update() + + + def snippet(page: pymupdf.Page, rect: pymupdf.Rect, path: str, dpi: int = 150) -> str: + shown = (pymupdf.Rect(rect) * page.rotation_matrix + (-60, -25, 60, 25)) & page.rect + page.get_pixmap(clip=shown, dpi=dpi).save(path) + return path + + + def union(rects) -> pymupdf.Rect: + r = pymupdf.Rect(rects[0]) + for x in rects[1:]: + r |= pymupdf.Rect(x) + return r + + + # ========================================================================== + # A. PDF vs reference data + # ========================================================================== + @dataclass + class Check: + field: str # name in the reference data + label: str # text printed in the PDF before the value + expected: str # value from the reference data + pick: str = "first" # "first" value right of the label, or "last" (table rows) + page: int | None = None # restrict to a page (0-based) if known + tolerance: float = 0.0 # numeric tolerance, e.g. 0.01 for rounding + + + @dataclass + class CheckResult: + field: str + expected: str + found: str | None = None + status: str = "label_not_found" # match | mismatch | value_not_found | label_not_found + page: int | None = None + label_bbox: tuple | None = None + value_bbox: tuple | None = None + snippet: str | None = None + + + def read_value_right_of(page: pymupdf.Page, label: pymupdf.Rect, pick: str = "first", + max_gap: float = 40) -> tuple[str | None, pymupdf.Rect | None]: + """Read the words on the same line to the right of `label`. + pick='first': the first group of words (stops at a large horizontal gap). + pick='last' : the last group on the line (typical for table rows / amounts).""" + mid = (label.y0 + label.y1) / 2 + line = [w for w in page.get_text("words", sort=True) + if w[0] >= label.x1 - 1 and w[1] <= mid <= w[3]] + line.sort(key=lambda w: w[0]) + line = [w for w in line if w[4] not in (":", "-", "–")] + if not line: + return None, None + + groups, current = [], [line[0]] + for prev, w in zip(line, line[1:]): + if w[0] - prev[2] > max_gap: + groups.append(current) + current = [] + current.append(w) + groups.append(current) + + chosen = groups[-1] if pick == "last" else groups[0] + text = " ".join(w[4] for w in chosen).strip(":; ") + return text, union([w[:4] for w in chosen]) + + + def verify_against_reference(doc, checks: list[Check], out_pdf: str, + snippet_prefix: str | None = None) -> list[CheckResult]: + results = [] + for c in checks: + r = CheckResult(field=c.field, expected=str(c.expected)) + pages = [doc[c.page]] if c.page is not None else list(doc) + + # Every occurrence of the label is a candidate; prefer the one whose + # value matches, otherwise report the first one found. + candidates = [] + for page in pages: + for lr in page.search_for(c.label): + text, vr = read_value_right_of(page, lr, c.pick) + candidates.append((page, lr, text, vr)) + if not candidates: + results.append(r) + continue + + best = next((x for x in candidates if x[2] and same_value(x[2], c.expected, c.tolerance)), + candidates[0]) + page, lr, text, vr = best + r.page, r.found = page.number, text + r.label_bbox = tuple(round(v, 1) for v in lr) + + if text is None: + r.status = "value_not_found" + mark(page, lr, ORANGE, c.field, f"{c.field}: no value next to label; expected {c.expected}") + target = lr + else: + r.value_bbox = tuple(round(v, 1) for v in vr) + target = vr + if same_value(text, c.expected): + r.status = "match" + mark(page, vr, GREEN, c.field, f"{c.field} OK: {text}") + elif c.tolerance > 0: + if same_value(text, c.expected, c.tolerance): + r.status = "close match" + mark(page, vr, ORANGE, c.field, f"{c.field} PDF shows {text}, reference has {c.expected}, but within tolerance: {c.tolerance} ") + else: + r.status = "mismatch" + mark(page, vr, RED, c.field, f"{c.field}: PDF shows {text}, reference has {c.expected}, not within tolerance: {c.tolerance} ") + else: + r.status = "mismatch" + mark(page, vr, RED, c.field, f"{c.field}: PDF shows {text}, reference has {c.expected}") + + if snippet_prefix: + r.snippet = snippet(page, target | lr, f"{snippet_prefix}_{c.field}.png") + results.append(r) + + doc.save(out_pdf, garbage=3, deflate=True) + return results + + + def checks_from_csv(path: str, pick: str = "last") -> list[Check]: + """CSV/Excel export with columns: field,label,expected[,tolerance]""" + with open(path, newline="", encoding="utf-8-sig") as f: + return [Check(row["field"], row["label"], row["expected"], pick=pick, + tolerance=float(row.get("tolerance") or 0)) + for row in csv.DictReader(f)] + + + def checks_from_xlsx(path: str, sheet: str | None = None, pick: str = "last") -> list[Check]: + from openpyxl import load_workbook + ws = load_workbook(path, data_only=True)[sheet] if sheet else load_workbook(path, data_only=True).active + rows = ws.iter_rows(values_only=True) + header = [str(h).strip().lower() for h in next(rows)] + out = [] + for row in rows: + d = dict(zip(header, row)) + if d.get("label"): + out.append(Check(str(d["field"]), str(d["label"]), str(d["expected"]), pick=pick, + tolerance=float(d.get("tolerance") or 0))) + return out + + + # ========================================================================== + # B. PDF vs PDF (version comparison) + # ========================================================================== + @dataclass + class Change: + kind: str # replace | insert | delete + old_text: str + new_text: str + old_page: int | None = None + new_page: int | None = None + old_bbox: tuple | None = None + new_bbox: tuple | None = None + + + def doc_words(doc) -> list[tuple]: + """(page_no, rect, word, line_key) for every word in reading order.""" + out = [] + for page in doc: + for w in page.get_text("words", sort=True): + out.append((page.number, pymupdf.Rect(w[:4]), w[4], (page.number, w[5], w[6]))) + return out + + + def group_by_line(words: list[tuple]) -> list[tuple[int, pymupdf.Rect]]: + """Merge consecutive words on the same line into one rect per line.""" + groups = [] + for pno, rect, _, key in words: + if groups and groups[-1][0] == key: + groups[-1][2] |= rect + else: + groups.append([key, pno, pymupdf.Rect(rect)]) + return [(g[1], g[2]) for g in groups] + + + def compare_versions(old_doc, new_doc, old_out: str, new_out: str, + snippet_prefix: str | None = None) -> list[Change]: + old_w, new_w = doc_words(old_doc), doc_words(new_doc) + a = [normalise(w[2]) or w[2] for w in old_w] + b = [normalise(w[2]) or w[2] for w in new_w] + # For very large documents, diff page by page (or paragraph by paragraph) instead. + sm = difflib.SequenceMatcher(a=a, b=b, autojunk=False) + + changes = [] + for tag, i1, i2, j1, j2 in sm.get_opcodes(): + if tag == "equal": + continue + ow, nw = old_w[i1:i2], new_w[j1:j2] + ch = Change(kind=tag, + old_text=" ".join(w[2] for w in ow), + new_text=" ".join(w[2] for w in nw)) + + if ow: # struck out in the old version + ch.old_page = ow[0][0] + for pno, rect in group_by_line(ow): + page = old_doc[pno] + annot = page.add_strikeout_annot(rect) + annot.set_colors(stroke=RED) + annot.set_info(title="Removed/changed", + content=f"was: {ch.old_text}\nnow: {ch.new_text or '(deleted)'}") + annot.update() + ch.old_bbox = tuple(round(v, 1) for v in union([w[1] for w in ow if w[0] == ch.old_page])) + + if nw: # highlighted in the new version + ch.new_page = nw[0][0] + for pno, rect in group_by_line(nw): + page = new_doc[pno] + annot = page.add_highlight_annot(rect) + annot.set_colors(stroke=(1, 0.85, 0.2) if tag == "replace" else (0.6, 0.9, 0.6)) + annot.set_info(title="Changed" if tag == "replace" else "Inserted", + content=f"now: {ch.new_text}\nwas: {ch.old_text or '(new)'}") + annot.update() + ch.new_bbox = tuple(round(v, 1) for v in union([w[1] for w in nw if w[0] == ch.new_page])) + else: # pure deletion: point at where it was + anchor = new_w[j1] if j1 < len(new_w) else new_w[-1] if new_w else None + if anchor: + ch.new_page = anchor[0] + r = anchor[1] + pin = pymupdf.Rect(r.x0 - 3, r.y0, r.x0 - 1, r.y1) + annot = new_doc[anchor[0]].add_rect_annot(pin) + annot.set_colors(stroke=RED, fill=RED) + annot.set_info(title="Deleted", content=f"removed: {ch.old_text}") + annot.update() + ch.new_bbox = tuple(round(v, 1) for v in pin) + + changes.append(ch) + + if snippet_prefix: + for n, ch in enumerate(changes, 1): + if ch.new_page is not None: + snippet(new_doc[ch.new_page], pymupdf.Rect(ch.new_bbox), + f"{snippet_prefix}_change{n}.png") + + old_doc.save(old_out, garbage=3, deflate=True) + new_doc.save(new_out, garbage=3, deflate=True) + return changes + + + # -------------------------------------------------------------------------- + # Self-contained demo + # -------------------------------------------------------------------------- + def make_statement(path: str, rows: list[tuple[str, str]], footer: str): + doc = pymupdf.open() + page = doc.new_page() + page.insert_text((72, 72), "Quarterly Financial Statement", fontsize=16) + page.insert_text((72, 100), "Reporting period: Q2 2026", fontsize=11) + y = 140 + for label, amount in rows: + page.insert_text((72, y), label, fontsize=11) + page.insert_text((400, y), amount, fontsize=11) + y += 20 + page.insert_text((72, y + 20), footer, fontsize=10) + doc.save(path) + doc.close() + + + if __name__ == "__main__": + make_statement("statement_v1.pdf", + [("Revenue", "1,250,000.00"), ("Cost of sales", "(480,000.00)"), + ("Operating expenses", "(310,500.00)"), ("Net profit", "459,500.00")], + "Figures are unaudited and presented in EUR.") + make_statement("statement_v2.pdf", + [("Revenue", "1,265,000.00"), ("Cost of sales", "(480,000.00)"), + ("Operating expenses", "(310,500.00)"), ("Other income", "12,000.00"), + ("Net profit", "486,500.00")], + "Figures are presented in EUR.") + + # --- A. Verify v2 against reference data (e.g. an Excel export) -------- + with open("reference.csv", "w", newline="") as f: + f.write("field,label,expected,tolerance\n" + "period,Reporting period,Q2 2026,\n" + "revenue,Revenue,1265000,\n" + "cost_of_sales,Cost of sales,-480000,\n" + "operating_expenses,Operating expenses,-310500,\n" + "net_profit,Net profit,485700,1000\n" # wrong in the PDF, but with a tolerance value which puts it "in bounds" + "tax,Income tax,-95000,\n") # not in the PDF + + checks = checks_from_csv("reference.csv") + checks[0].pick = "first" + doc = pymupdf.open("statement_v2.pdf") + report_a = verify_against_reference(doc, checks, "statement_v2_verified.pdf", + snippet_prefix="check") + doc.close() + + # --- B. Compare v1 and v2 --------------------------------------------- + report_b = compare_versions(pymupdf.open("statement_v1.pdf"), + pymupdf.open("statement_v2.pdf"), + "statement_v1_marked.pdf", "statement_v2_marked.pdf", + snippet_prefix="diff") + + print(json.dumps({"reference_check": [asdict(r) for r in report_a], + "version_changes": [asdict(c) for c in report_b]}, indent=2)) + + +.. include:: footer.rst + \ No newline at end of file diff --git a/docs/header.rst b/docs/header.rst index 06bfe84a8..c22f09c50 100644 --- a/docs/header.rst +++ b/docs/header.rst @@ -24,9 +24,9 @@ PyMuPDF -.. |PyMuPDF Pro| raw:: html +.. |PyMuPDF Office| raw:: html - PyMuPDF Pro + PyMuPDF Office .. |PyMuPDF4LLM| raw:: html diff --git a/docs/how-to-open-a-file.rst b/docs/how-to-open-a-file.rst index 06ddf0039..5d3a5f973 100644 --- a/docs/how-to-open-a-file.rst +++ b/docs/how-to-open-a-file.rst @@ -30,10 +30,10 @@ The following file types are supported: ---- -PyMuPDF Pro +PyMuPDF Office """"""""""""""" -|PyMuPDF Pro| can open Office files. +|PyMuPDF Office| can open Office files. The following file types are supported: diff --git a/docs/images/grounding/citation.png b/docs/images/grounding/citation.png new file mode 100644 index 000000000..8538cf9d7 Binary files /dev/null and b/docs/images/grounding/citation.png differ diff --git a/docs/images/grounding/extraction.png b/docs/images/grounding/extraction.png new file mode 100644 index 000000000..30d3d398f Binary files /dev/null and b/docs/images/grounding/extraction.png differ diff --git a/docs/images/grounding/verification.png b/docs/images/grounding/verification.png new file mode 100644 index 000000000..42dbdc62d Binary files /dev/null and b/docs/images/grounding/verification.png differ diff --git a/docs/index.rst b/docs/index.rst index 511b6ed1a..99204ea3c 100644 --- a/docs/index.rst +++ b/docs/index.rst @@ -41,10 +41,9 @@ This documentation covers all versions up to |version|. :maxdepth: 1 about.rst - pymupdf4llm/index.rst - pymupdf-pro/index.rst - - + pymupdf-office/index.rst + rag.rst + grounding.rst .. toctree:: :caption: User Guide @@ -53,7 +52,7 @@ This documentation covers all versions up to |version|. installation.rst the-basics.rst tutorial.rst - rag.rst + ocr/index.rst resources.rst faq/index.rst @@ -71,7 +70,6 @@ This documentation covers all versions up to |version|. :maxdepth: 2 classes.rst - pymupdf4llm/api.rst algebra.rst lowlevel.rst glossary.rst diff --git a/docs/llms/llms-full.txt b/docs/llms/llms-full.txt index 1469b6cfc..4835e8286 100644 --- a/docs/llms/llms-full.txt +++ b/docs/llms/llms-full.txt @@ -10,7 +10,7 @@ PyMuPDF is hosted on [GitHub](https://github.com/pymupdf/PyMuPDF) and registered ``` pip install pymupdf -pip install pymupdf4llm # for LLM/RAG features +pip install pymupdf4llm ``` Import as: @@ -33,7 +33,7 @@ doc = pymupdf.open("a.pdf") # open a document `pymupdf.open(...)` is an alias for `pymupdf.Document(...)`. -Supported file types include: PDF, XPS, EPUB, MOBI, FB2, CBZ, SVG, TXT, and image formats (PNG, JPEG, BMP, GIF, TIFF, etc.). PyMuPDF Pro adds support for Office formats (DOCX, XLSX, PPTX, HWP, etc.). +Supported file types include: PDF, XPS, EPUB, MOBI, FB2, CBZ, SVG, TXT, and image formats (PNG, JPEG, BMP, GIF, TIFF, etc.). PyMuPDF Office adds support for Office formats (DOCX, XLSX, PPTX, HWP, etc.). ### Extract Text from a PDF @@ -105,7 +105,7 @@ PyMuPDF4LLM is a lightweight extension for PyMuPDF that converts documents into - Smart OCR — automatically OCRs only regions that need it, skipping clean text - Framework integrations — drop-in support for LlamaIndex and LangChain - Page chunking — chunk output by page with full metadata per chunk, ready for vector stores -- Office document support — works with PyMuPDF Pro for DOCX, XLSX, PPTX, etc. +- Office document support — works with PyMuPDF Office for DOCX, XLSX, PPTX, etc. ### Installation @@ -311,24 +311,22 @@ splitter = MarkdownTextSplitter(chunk_size=500, chunk_overlap=50) chunks = splitter.create_documents([md_text]) ``` -### Office Document Support (PyMuPDF Pro) +### Office Document Support (PyMuPDF Office) ```python -import pymupdf4llm -import pymupdf.pro +import pymupdf.office -pymupdf.pro.unlock() +pymupdf.office.unlock() # Now supports DOCX, XLSX, PPTX, DOC, HWP, etc. -md = pymupdf4llm.to_markdown("report.docx") -md = pymupdf4llm.to_markdown("spreadsheet.xlsx") +doc = pymupdf.open("report.docx") +md = doc.to_markdown() ``` -### to_markdown() Full Parameter Reference +### Document.to_markdown() Full Parameter Reference | Parameter | Type | Default | Description | |-----------|------|---------|-------------| -| `doc` | `Document` or `str` | required | File path or PyMuPDF Document | | `pages` | `list` or `None` | `None` | 0-based page numbers to process; `None` = all | | `page_chunks` | `bool` | `False` | Return list of per-page dicts instead of one string | | `write_images` | `bool` | `False` | Save images to disk; referenced in Markdown | @@ -406,6 +404,9 @@ md = pymupdf4llm.to_markdown("spreadsheet.xlsx") | `Document.get_ocgs()` | Get optional content groups (PDF layers) | | `Document.bake()` | Make annotations permanent | | `Document.journal_enable()` | Enable journalling (undo/redo) | +| `Document.to_json()` | Convert document to JSON +| `Document.to_markdown()` | Convert document to Markdown +| `Document.to_text()` | Convert document to TXT ### Key Attributes diff --git a/docs/locales/ja/LC_MESSAGES/how-to-open-a-file.po b/docs/locales/ja/LC_MESSAGES/how-to-open-a-file.po index b89457186..88a2be94c 100644 --- a/docs/locales/ja/LC_MESSAGES/how-to-open-a-file.po +++ b/docs/locales/ja/LC_MESSAGES/how-to-open-a-file.po @@ -61,7 +61,7 @@ msgid "PyMuPDF Pro" msgstr "" #: ../../how-to-open-a-file.rst:36 c4725d269a0e41b0867c5b6b16b8960b -msgid "|PyMuPDF Pro| can open Office files." +msgid "|PyMuPDF Office| can open Office files." msgstr "" #: ../../how-to-open-a-file.rst:43 d91d49bc138843188ab0077259cd82ae diff --git a/docs/locales/ja/LC_MESSAGES/pymupdf-layout/index.po b/docs/locales/ja/LC_MESSAGES/pymupdf-layout/index.po index 31efbb49e..d66c02919 100644 --- a/docs/locales/ja/LC_MESSAGES/pymupdf-layout/index.po +++ b/docs/locales/ja/LC_MESSAGES/pymupdf-layout/index.po @@ -172,16 +172,16 @@ msgid "Extending Capability" msgstr "機能の拡張" #: ../../pymupdf-layout/index.rst:107 f3bec49e1ee042caad2ba1f1f8d07230 -msgid "Using with Pro" -msgstr "PyMuPDF Pro との使用" +msgid "Using with Office" +msgstr "PyMuPDF Office との使用" #: ../../pymupdf-layout/index.rst:109 1a6a4c0ada2842d588768a849fe3c7f3 msgid "" -"We are able to extend |PyMuPDF Layout| to work with |PyMuPDF Pro| and " +"We are able to extend |PyMuPDF Layout| to work with |PyMuPDF Office| and " "thus increase our capability by allowing Office documents to be provided " "as input files. In this case all we have to do is to add the import for " -"|PyMuPDF Pro| and unlock it::" -msgstr "|PyMuPDF Layout| を |PyMuPDF Pro| と連携させることで機能を拡張し、Office ドキュメントを入力ファイルとして提供できるようにすることで能力を向上させることができます。この場合、|PyMuPDF Pro| のインポートを追加してロックを解除するだけです::" +"|PyMuPDF Office| and unlock it::" +msgstr "|PyMuPDF Layout| を |PyMuPDF Office| と連携させることで機能を拡張し、Office ドキュメントを入力ファイルとして提供できるようにすることで能力を向上させることができます。この場合、|PyMuPDF Office| のインポートを追加してロックを解除するだけです::" #: ../../pymupdf-layout/index.rst:116 695ebd9778304b3da362108ea2e9695e msgid "Now we can happily load Office files and convert them as follows::" diff --git a/docs/locales/ja/LC_MESSAGES/pymupdf-pro.po b/docs/locales/ja/LC_MESSAGES/pymupdf-pro.po index 06b3959d5..a94393fee 100644 --- a/docs/locales/ja/LC_MESSAGES/pymupdf-pro.po +++ b/docs/locales/ja/LC_MESSAGES/pymupdf-pro.po @@ -40,7 +40,7 @@ msgid "PyMuPDF Pro" msgstr "" #: ../../pymupdf-pro.rst:12 aa17b9e367774bb89e40e6d34d5ec0fe -msgid "|PyMuPDF Pro| is a set of *commercial extensions* for |PyMuPDF|." +msgid "|PyMuPDF Office| is a set of *commercial extensions* for |PyMuPDF|." msgstr "" #: ../../pymupdf-pro.rst:14 11f22babd0e94a46b70d2de6efd179bc @@ -71,7 +71,7 @@ msgstr "" #: ../../pymupdf-pro.rst:25 1321bc95b45541b48337e44f78b58020 msgid "" -"A licensed version of |PyMuPDF Pro| also gives you a licensed version of " +"A licensed version of |PyMuPDF Office| also gives you a licensed version of " "|PyMuPDF4LLM|. If you are interested in using the |PyMuPDF4LLM| package " "you should install it separately." msgstr "" @@ -107,7 +107,7 @@ msgstr "" #: ../../pymupdf-pro.rst:42 b8f3023035f74b3dacbf2e3210f2c12f msgid "" "In addition to the `standard file types supported by PyMuPDF " -"`, |PyMuPDF Pro| supports:" +"`, |PyMuPDF Office| supports:" msgstr "" #: ../../pymupdf-pro.rst:47 45349c74355f48c0abf3ae8f486843a6 @@ -144,7 +144,7 @@ msgstr "" #: ../../pymupdf-pro.rst:82 f5cf1df845294870b35e718fadf85b1f msgid "" -"Import |PyMuPDF Pro| and you can then reference **Office** documents " +"Import |PyMuPDF Office| and you can then reference **Office** documents " "directly, e.g.:" msgstr "" @@ -177,7 +177,7 @@ msgstr "" #: ../../pymupdf-pro.rst:123 2391a063f21c4de092b0beca7b94377e msgid "" -"|PyMuPDF Pro| functionality is restricted without a license key as " +"|PyMuPDF Office| functionality is restricted without a license key as " "follows:" msgstr "" @@ -207,13 +207,13 @@ msgid "Using a key" msgstr "" #: ../../pymupdf-pro.rst:142 25f50e294ef5429e86781bea1e1b1e29 -msgid "Initialize |PyMuPDF Pro| with a key as follows:" +msgid "Initialize |PyMuPDF Office| with a key as follows:" msgstr "" #: ../../pymupdf-pro.rst:150 b01044f9b9f34927892044fa1c007b7f msgid "" "This will allow you to evaluate the product for a limited time. If you " -"want to use |PyMuPDF Pro| after this time you should then `enquire about " +"want to use |PyMuPDF Office| after this time you should then `enquire about " "obtaining a commercial license `_." msgstr "" @@ -260,7 +260,7 @@ msgstr "" #~ msgstr "" #~ msgid "" -#~ "|PyMuPDF Pro| offers all the features" +#~ "|PyMuPDF Office| offers all the features" #~ " of |PyMuPDF|, plus enhanced functionality" #~ " to support **Office** documents." #~ msgstr "" diff --git a/docs/locales/ja/LC_MESSAGES/pymupdf-pro/index.po b/docs/locales/ja/LC_MESSAGES/pymupdf-pro/index.po index ddfb45f64..1dbdc6c2e 100644 --- a/docs/locales/ja/LC_MESSAGES/pymupdf-pro/index.po +++ b/docs/locales/ja/LC_MESSAGES/pymupdf-pro/index.po @@ -40,8 +40,8 @@ msgid "PyMuPDF Pro" msgstr "" #: ../../pymupdf-pro/index.rst:11 879666c817c94726a9df9e5880d5425a -msgid "|PyMuPDF Pro| is a set of *commercial extensions* for |PyMuPDF|." -msgstr "|PyMuPDF Pro| は、|PyMuPDF| 用の *商用拡張機能* セットです。" +msgid "|PyMuPDF Office| is a set of *commercial extensions* for |PyMuPDF|." +msgstr "|PyMuPDF Office| は、|PyMuPDF| 用の *商用拡張機能* セットです。" #: ../../pymupdf-pro/index.rst:13 ae90e3fd121f4e04a0a9ecd747514717 msgid "" @@ -71,10 +71,10 @@ msgstr "商用ライセンスの取得についてお問い合わせの場合は #: ../../pymupdf-pro/index.rst:24 f0a9c51213864c898871ca01453557e1 msgid "" -"A licensed version of |PyMuPDF Pro| also gives you a licensed version of " +"A licensed version of |PyMuPDF Office| also gives you a licensed version of " "|PyMuPDF4LLM|. If you are interested in using the |PyMuPDF4LLM| package " "you should install it separately." -msgstr "|PyMuPDF Pro| のライセンス版には、|PyMuPDF4LLM| のライセンス版も含まれています。|PyMuPDF4LLM| パッケージの使用に興味がある場合は、個別にインストールする必要があります。" +msgstr "|PyMuPDF Office| のライセンス版には、|PyMuPDF4LLM| のライセンス版も含まれています。|PyMuPDF4LLM| パッケージの使用に興味がある場合は、個別にインストールする必要があります。" #: ../../pymupdf-pro/index.rst:28 4a1d503fea744f4ba62d54e186f02d87 msgid "Platform support" @@ -107,8 +107,8 @@ msgstr "Office ファイルサポート" #: ../../pymupdf-pro/index.rst:41 7c38d85a3c164f3a90d869b52ef96eb1 msgid "" "In addition to the `standard file types supported by PyMuPDF " -"`, |PyMuPDF Pro| supports:" -msgstr "`PyMuPDF がサポートする標準ファイルタイプ ` に加えて、|PyMuPDF Pro| は以下をサポートします" +"`, |PyMuPDF Office| supports:" +msgstr "`PyMuPDF がサポートする標準ファイルタイプ ` に加えて、|PyMuPDF Office| は以下をサポートします" #: ../../pymupdf-pro/index.rst:46 67438064f36f474186aeb4fbfb093d31 msgid "**DOC/DOCX**" @@ -144,15 +144,15 @@ msgstr "**Office** ドキュメントを読み込む" #: ../../pymupdf-pro/index.rst:81 08e39b8bfc514665aafb5c1aa211bcfc msgid "" -"Import |PyMuPDF Pro| and you can then reference **Office** documents " +"Import |PyMuPDF Office| and you can then reference **Office** documents " "directly, e.g.:" -msgstr "|PyMuPDF Pro| をインポートすると、**Office** ドキュメントを直接参照できます。例えば:" +msgstr "|PyMuPDF Office| をインポートすると、**Office** ドキュメントを直接参照できます。例えば:" #: ../../pymupdf-pro/index.rst:92 e21e201731344245a88c7e6bdd9c2dc2 msgid "" "All standard |PyMuPDF| functionality is exposed as expected - |PyMuPDF " "Pro| handles the extended **Office** file types" -msgstr "すべての標準 |PyMuPDF| 機能が期待どおりに公開されます - |PyMuPDF Pro| は拡張された **Office** ファイルタイプを処理します" +msgstr "すべての標準 |PyMuPDF| 機能が期待どおりに公開されます - |PyMuPDF Office| は拡張された **Office** ファイルタイプを処理します" #: ../../pymupdf-pro/index.rst:95 32d84f02e19a40b5bb42c1d26a81afe5 msgid "" @@ -177,9 +177,9 @@ msgstr "制限事項" #: ../../pymupdf-pro/index.rst:122 badbfba0a2a0495290e9594585aaefc9 msgid "" -"|PyMuPDF Pro| functionality is restricted without a license key as " +"|PyMuPDF Office| functionality is restricted without a license key as " "follows:" -msgstr "|PyMuPDF Pro| 機能は、ライセンスキーがない場合、次のように制限されます:" +msgstr "|PyMuPDF Office| 機能は、ライセンスキーがない場合、次のように制限されます:" #: ../../pymupdf-pro/index.rst:124 0edbc072da504add956116e21d3eebde msgid "**Only the first 3 pages of any document will be available.**" @@ -207,16 +207,16 @@ msgid "Using a key" msgstr "キーの使用" #: ../../pymupdf-pro/index.rst:141 6766aa2a1d144b81a15c3ad28ce403da -msgid "Initialize |PyMuPDF Pro| with a key as follows:" -msgstr "次のようにキーで |PyMuPDF Pro| を初期化します" +msgid "Initialize |PyMuPDF Office| with a key as follows:" +msgstr "次のようにキーで |PyMuPDF Office| を初期化します" #: ../../pymupdf-pro/index.rst:149 92e495f1e7934e1e83d346f56ab9833d msgid "" "This will allow you to evaluate the product for a limited time. If you " -"want to use |PyMuPDF Pro| after this time you should then `enquire about " +"want to use |PyMuPDF Office| after this time you should then `enquire about " "obtaining a commercial license `_." -msgstr "これにより、限られた期間製品を評価できます。この期間後も |PyMuPDF Pro| を使用したい場合は、`商用ライセンスの取得についてお問い合わせ `_ ください。" +msgstr "これにより、限られた期間製品を評価できます。この期間後も |PyMuPDF Office| を使用したい場合は、`商用ライセンスの取得についてお問い合わせ `_ ください。" #: ../../pymupdf-pro/index.rst:153 a394162132a04de092aabb5d2fcbc814 msgid "Fonts" diff --git a/docs/locales/ja/LC_MESSAGES/pymupdf4llm/index.po b/docs/locales/ja/LC_MESSAGES/pymupdf4llm/index.po index cbe005bbd..62918b36c 100644 --- a/docs/locales/ja/LC_MESSAGES/pymupdf4llm/index.po +++ b/docs/locales/ja/LC_MESSAGES/pymupdf4llm/index.po @@ -63,9 +63,9 @@ msgstr "PyMuPDF4LLM を PyMuPDF-Layout と併用すると、ページレイア msgid "" "You can extend the supported file types to also include **Office** " "document formats (DOC/DOCX, XLS/XLSX, PPT/PPTX, HWP/HWPX) by :ref:`using " -"PyMuPDF Pro with PyMuPDF4LLM `." +"PyMuPDF Office with PyMuPDF4LLM `." msgstr "" -":ref:`PyMuPDF ProをPyMuPDF4LLMと併用することで " +":ref:`PyMuPDF OfficeをPyMuPDF4LLMと併用することで " "`、対応するファイル形式を拡張し、 **Office** " "ドキュメント形式(DOC/DOCX、XLS/XLSX、PPT/PPTX、HWP/HWPX)も含めることができます。" @@ -190,24 +190,24 @@ msgstr "" "**Markdown** 形式に変換され、その後、以下のように **LlamaIndex** ドキュメントとして返されます。" #: ../../pymupdf4llm/index.rst:107 e653025ecf1948759d96b33b7902c524 -msgid "Using with |PyMuPDF Pro|" +msgid "Using with |PyMuPDF Office|" msgstr "PyMuPDF Proとの使用 " #: ../../pymupdf4llm/index.rst:110 9e877b4bd26d43cbaf81821091d51a1b msgid "" "For **Office** document support, |PyMuPDF4LLM| works seamlessly with " -"|PyMuPDF Pro|. Assuming you have :doc:`../pymupdf-pro` installed you will" +"|PyMuPDF Office|. Assuming you have :doc:`../pymupdf-pro` installed you will" " be able to work with **Office** documents as expected:" msgstr "" -"**Office** ドキュメントのサポートのために、|PyMuPDF4LLM| は |PyMuPDF Pro| " +"**Office** ドキュメントのサポートのために、|PyMuPDF4LLM| は |PyMuPDF Office| " "とシームレスに動作します。:doc:`../pymupdf-pro` がインストールされている場合、期待通りに **Office** " "ドキュメントを操作できます。" #: ../../pymupdf4llm/index.rst:121 ff1a10504dbe4b5381a2f6232a21b346 msgid "" -"As you can see |PyMuPDF Pro| functionality will be available within the " +"As you can see |PyMuPDF Office| functionality will be available within the " "|PyMuPDF4LLM| context!" -msgstr "ご覧のとおり、|PyMuPDF Pro| の機能は |PyMuPDF4LLM| のコンテキスト内で利用可能になります!" +msgstr "ご覧のとおり、|PyMuPDF Office| の機能は |PyMuPDF4LLM| のコンテキスト内で利用可能になります!" #: ../../pymupdf4llm/index.rst:126 32a7f061e08e4a6ca0f4f3fddb070669 msgid "API" diff --git a/docs/locales/ja/LC_MESSAGES/resources.po b/docs/locales/ja/LC_MESSAGES/resources.po index 423030542..4d23fac72 100644 --- a/docs/locales/ja/LC_MESSAGES/resources.po +++ b/docs/locales/ja/LC_MESSAGES/resources.po @@ -43,8 +43,8 @@ msgid "**PyMuPDF Pro**" msgstr "" #: ../../resources.rst:12 08c45bb03e394054bea50f41ce8a06d0 -msgid "For **Office** file support `try PyMuPDF Pro `." -msgstr "**Office** ファイルのサポートには、`PyMuPDF Pro ` をお試しください。" +msgid "For **Office** file support `try PyMuPDF Office `." +msgstr "**Office** ファイルのサポートには、`PyMuPDF Office ` をお試しください。" #: ../../resources.rst:20 5af5e4a006674fbf9c2bb32995565bc2 msgid "Find out about **PyMuPDF Utilities**" diff --git a/docs/locales/ko/LC_MESSAGES/how-to-open-a-file.po b/docs/locales/ko/LC_MESSAGES/how-to-open-a-file.po index 18e76fb1c..6667914cb 100644 --- a/docs/locales/ko/LC_MESSAGES/how-to-open-a-file.po +++ b/docs/locales/ko/LC_MESSAGES/how-to-open-a-file.po @@ -57,12 +57,12 @@ msgid "The following file types are supported:" msgstr "다음 파일 유형이 지원됩니다:" #: ../../how-to-open-a-file.rst:34 fa9ae965eea142478505848892067bb0 -msgid "PyMuPDF Pro" -msgstr "PyMuPDF Pro" +msgid "PyMuPDF Office" +msgstr "PyMuPDF Office" #: ../../how-to-open-a-file.rst:36 c4725d269a0e41b0867c5b6b16b8960b -msgid "|PyMuPDF Pro| can open Office files." -msgstr "|PyMuPDF Pro| 는 Office 파일을 열 수 있습니다." +msgid "|PyMuPDF Office| can open Office files." +msgstr "|PyMuPDF Office| 는 Office 파일을 열 수 있습니다." #: ../../how-to-open-a-file.rst:43 d91d49bc138843188ab0077259cd82ae msgid "**DOC/DOCX**" diff --git a/docs/locales/ko/LC_MESSAGES/pymupdf-layout/index.po b/docs/locales/ko/LC_MESSAGES/pymupdf-layout/index.po index 8594dd299..2218b78f8 100644 --- a/docs/locales/ko/LC_MESSAGES/pymupdf-layout/index.po +++ b/docs/locales/ko/LC_MESSAGES/pymupdf-layout/index.po @@ -177,12 +177,12 @@ msgstr "Pro 와 함께 사용" #: ../../pymupdf-layout/index.rst:111 1a6a4c0ada2842d588768a849fe3c7f3 msgid "" -"We are able to extend |PyMuPDF Layout| to work with |PyMuPDF Pro| and " +"We are able to extend |PyMuPDF Layout| to work with |PyMuPDF Office| and " "thus increase our capability by allowing Office documents to be provided " "as input files. In this case all we have to do is to add the import for " -"|PyMuPDF Pro| and unlock it::" +"|PyMuPDF Office| and unlock it::" msgstr "" -"|PyMuPDF Layout| 를 |PyMuPDF Pro| 와 함께 작동하도록 확장하여 Office 문서를 입력 파일로 제공할 수 있게 하여 기능을 향상시킬 수 있습니다. 이 경우 |PyMuPDF Pro| 를 가져오고 잠금을 해제하기만 하면 됩니다::" +"|PyMuPDF Layout| 를 |PyMuPDF Office| 와 함께 작동하도록 확장하여 Office 문서를 입력 파일로 제공할 수 있게 하여 기능을 향상시킬 수 있습니다. 이 경우 |PyMuPDF Office| 를 가져오고 잠금을 해제하기만 하면 됩니다::" #: ../../pymupdf-layout/index.rst:118 695ebd9778304b3da362108ea2e9695e msgid "Now we can happily load Office files and convert them as follows::" @@ -278,12 +278,12 @@ msgstr "이 문서는 |version| 버전까지의 모든 버전을 다룹니다." #~ msgid "" #~ "We are able to extend |PyMuPDF " -#~ "Layout| to work with |PyMuPDF Pro| " +#~ "Layout| to work with |PyMuPDF Office| " #~ "and thus increase our capability by " #~ "allowing Office documents to be provided" #~ " as input files. In this case " #~ "all we have to do is to " -#~ "include the import for |PyMuPDF Pro| " +#~ "include the import for |PyMuPDF Office| " #~ "and unlock it before we import &" #~ " activate |PyMuPDF Layout|::" #~ msgstr "" diff --git a/docs/locales/ko/LC_MESSAGES/pymupdf-pro/index.po b/docs/locales/ko/LC_MESSAGES/pymupdf-pro/index.po index 25d27c0b5..947431cb8 100644 --- a/docs/locales/ko/LC_MESSAGES/pymupdf-pro/index.po +++ b/docs/locales/ko/LC_MESSAGES/pymupdf-pro/index.po @@ -36,12 +36,12 @@ msgid "" msgstr "PDF 텍스트 추출, PDF 이미지 추출, PDF 변환, PDF 테이블, PDF 분할, PDF 생성, Pyodide, PyScript" #: ../../pymupdf-pro/index.rst:8 07624f1e03dd463c99bf82f95b91a35e -msgid "PyMuPDF Pro" -msgstr "PyMuPDF Pro" +msgid "PyMuPDF Office" +msgstr "PyMuPDF Office" #: ../../pymupdf-pro/index.rst:11 879666c817c94726a9df9e5880d5425a -msgid "|PyMuPDF Pro| is a set of *commercial extensions* for |PyMuPDF|." -msgstr "|PyMuPDF Pro| 는 |PyMuPDF| 를 위한 *상용 확장* 세트입니다." +msgid "|PyMuPDF Office| is a set of *commercial extensions* for |PyMuPDF|." +msgstr "|PyMuPDF Office| 는 |PyMuPDF| 를 위한 *상용 확장* 세트입니다." #: ../../pymupdf-pro/index.rst:13 ae90e3fd121f4e04a0a9ecd747514717 msgid "" @@ -71,10 +71,10 @@ msgstr "상용 라이선스 취득에 대한 문의는 `이 연락처 페이지 #: ../../pymupdf-pro/index.rst:24 f0a9c51213864c898871ca01453557e1 msgid "" -"A licensed version of |PyMuPDF Pro| also gives you a licensed version of " +"A licensed version of |PyMuPDF Office| also gives you a licensed version of " "|PyMuPDF4LLM|. If you are interested in using the |PyMuPDF4LLM| package " "you should install it separately." -msgstr "|PyMuPDF Pro| 의 라이선스 버전은 |PyMuPDF4LLM| 의 라이선스 버전도 제공합니다. |PyMuPDF4LLM| 패키지를 사용하려면 별도로 설치해야 합니다." +msgstr "|PyMuPDF Office| 의 라이선스 버전은 |PyMuPDF4LLM| 의 라이선스 버전도 제공합니다. |PyMuPDF4LLM| 패키지를 사용하려면 별도로 설치해야 합니다." #: ../../pymupdf-pro/index.rst:28 4a1d503fea744f4ba62d54e186f02d87 msgid "Platform support" @@ -107,8 +107,8 @@ msgstr "Office 파일 지원" #: ../../pymupdf-pro/index.rst:41 7c38d85a3c164f3a90d869b52ef96eb1 msgid "" "In addition to the `standard file types supported by PyMuPDF " -"`, |PyMuPDF Pro| supports:" -msgstr "`PyMuPDF 가 지원하는 표준 파일 타입 ` 외에도 |PyMuPDF Pro| 는 다음을 지원합니다:" +"`, |PyMuPDF Office| supports:" +msgstr "`PyMuPDF 가 지원하는 표준 파일 타입 ` 외에도 |PyMuPDF Office| 는 다음을 지원합니다:" #: ../../pymupdf-pro/index.rst:46 67438064f36f474186aeb4fbfb093d31 msgid "**DOC/DOCX**" @@ -144,15 +144,15 @@ msgstr "**Office** 문서 로드" #: ../../pymupdf-pro/index.rst:81 08e39b8bfc514665aafb5c1aa211bcfc msgid "" -"Import |PyMuPDF Pro| and you can then reference **Office** documents " +"Import |PyMuPDF Office| and you can then reference **Office** documents " "directly, e.g.:" -msgstr "|PyMuPDF Pro| 를 가져오면 **Office** 문서를 직접 참조할 수 있습니다. 예:" +msgstr "|PyMuPDF Office| 를 가져오면 **Office** 문서를 직접 참조할 수 있습니다. 예:" #: ../../pymupdf-pro/index.rst:92 e21e201731344245a88c7e6bdd9c2dc2 msgid "" "All standard |PyMuPDF| functionality is exposed as expected - |PyMuPDF " "Pro| handles the extended **Office** file types" -msgstr "모든 표준 |PyMuPDF| 기능이 예상대로 노출됩니다 - |PyMuPDF Pro| 는 확장된 **Office** 파일 타입을 처리합니다" +msgstr "모든 표준 |PyMuPDF| 기능이 예상대로 노출됩니다 - |PyMuPDF Office| 는 확장된 **Office** 파일 타입을 처리합니다" #: ../../pymupdf-pro/index.rst:95 32d84f02e19a40b5bb42c1d26a81afe5 msgid "" @@ -177,9 +177,9 @@ msgstr "제한 사항" #: ../../pymupdf-pro/index.rst:122 badbfba0a2a0495290e9594585aaefc9 msgid "" -"|PyMuPDF Pro| functionality is restricted without a license key as " +"|PyMuPDF Office| functionality is restricted without a license key as " "follows:" -msgstr "라이선스 키 없이 |PyMuPDF Pro| 기능은 다음과 같이 제한됩니다:" +msgstr "라이선스 키 없이 |PyMuPDF Office| 기능은 다음과 같이 제한됩니다:" #: ../../pymupdf-pro/index.rst:124 0edbc072da504add956116e21d3eebde msgid "**Only the first 3 pages of any document will be available.**" @@ -207,16 +207,16 @@ msgid "Using a key" msgstr "키 사용" #: ../../pymupdf-pro/index.rst:141 6766aa2a1d144b81a15c3ad28ce403da -msgid "Initialize |PyMuPDF Pro| with a key as follows:" -msgstr "다음과 같이 키로 |PyMuPDF Pro| 를 초기화합니다:" +msgid "Initialize |PyMuPDF Office| with a key as follows:" +msgstr "다음과 같이 키로 |PyMuPDF Office| 를 초기화합니다:" #: ../../pymupdf-pro/index.rst:149 92e495f1e7934e1e83d346f56ab9833d msgid "" "This will allow you to evaluate the product for a limited time. If you " -"want to use |PyMuPDF Pro| after this time you should then `enquire about " +"want to use |PyMuPDF Office| after this time you should then `enquire about " "obtaining a commercial license `_." -msgstr "이것은 제한된 시간 동안 제품을 평가할 수 있게 해줍니다. 이 시간 이후에 |PyMuPDF Pro| 를 사용하려면 `상용 라이선스 취득에 대해 문의하세요 `_." +msgstr "이것은 제한된 시간 동안 제품을 평가할 수 있게 해줍니다. 이 시간 이후에 |PyMuPDF Office| 를 사용하려면 `상용 라이선스 취득에 대해 문의하세요 `_." #: ../../pymupdf-pro/index.rst:153 a394162132a04de092aabb5d2fcbc814 msgid "Fonts" diff --git a/docs/locales/ko/LC_MESSAGES/pymupdf4llm/index.po b/docs/locales/ko/LC_MESSAGES/pymupdf4llm/index.po index 442e2250f..7384fb409 100644 --- a/docs/locales/ko/LC_MESSAGES/pymupdf4llm/index.po +++ b/docs/locales/ko/LC_MESSAGES/pymupdf4llm/index.po @@ -61,8 +61,8 @@ msgstr "" msgid "" "You can extend the supported file types to also include **Office** " "document formats (DOC/DOCX, XLS/XLSX, PPT/PPTX, HWP/HWPX) by :ref:`using " -"PyMuPDF Pro with PyMuPDF4LLM `." -msgstr ":ref:`PyMuPDF Pro와 PyMuPDF4LLM 함께 사용 ` 하여 지원되는 파일 타입을 **Office** 문서 형식(DOC/DOCX, XLS/XLSX, PPT/PPTX, HWP/HWPX)도 포함하도록 확장할 수 있습니다." +"PyMuPDF Office with PyMuPDF4LLM `." +msgstr ":ref:`PyMuPDF Office PyMuPDF4LLM 함께 사용 ` 하여 지원되는 파일 타입을 **Office** 문서 형식(DOC/DOCX, XLS/XLSX, PPT/PPTX, HWP/HWPX)도 포함하도록 확장할 수 있습니다." #: ../../pymupdf4llm/index.rst:19 38428359cdaf4039b439454f5c712ec7 msgid "Features" @@ -185,21 +185,21 @@ msgid "" msgstr "|PyMuPDF4LLM| 는 **LlamaIndex** 문서로 직접 변환을 지원합니다. 문서는 먼저 **Markdown** 형식으로 변환된 다음 다음과 같이 **LlamaIndex** 문서가 반환됩니다:" #: ../../pymupdf4llm/index.rst:107 e653025ecf1948759d96b33b7902c524 -msgid "Using with |PyMuPDF Pro|" -msgstr "|PyMuPDF Pro| 와 함께 사용" +msgid "Using with |PyMuPDF Office|" +msgstr "|PyMuPDF Office| 와 함께 사용" #: ../../pymupdf4llm/index.rst:110 9e877b4bd26d43cbaf81821091d51a1b msgid "" "For **Office** document support, |PyMuPDF4LLM| works seamlessly with " -"|PyMuPDF Pro|. Assuming you have :doc:`../pymupdf-pro` installed you will" +"|PyMuPDF Office|. Assuming you have :doc:`../pymupdf-pro` installed you will" " be able to work with **Office** documents as expected:" -msgstr "**Office** 문서 지원을 위해 |PyMuPDF4LLM| 는 |PyMuPDF Pro| 와 원활하게 작동합니다. :doc:`../pymupdf-pro` 가 설치되어 있다고 가정하면 예상대로 **Office** 문서로 작업할 수 있습니다:" +msgstr "**Office** 문서 지원을 위해 |PyMuPDF4LLM| 는 |PyMuPDF Office| 와 원활하게 작동합니다. :doc:`../pymupdf-pro` 가 설치되어 있다고 가정하면 예상대로 **Office** 문서로 작업할 수 있습니다:" #: ../../pymupdf4llm/index.rst:121 ff1a10504dbe4b5381a2f6232a21b346 msgid "" -"As you can see |PyMuPDF Pro| functionality will be available within the " +"As you can see |PyMuPDF Office| functionality will be available within the " "|PyMuPDF4LLM| context!" -msgstr "보시다시피 |PyMuPDF Pro| 기능이 |PyMuPDF4LLM| 컨텍스트 내에서 사용 가능합니다!" +msgstr "보시다시피 |PyMuPDF Office| 기능이 |PyMuPDF4LLM| 컨텍스트 내에서 사용 가능합니다!" #: ../../pymupdf4llm/index.rst:126 32a7f061e08e4a6ca0f4f3fddb070669 msgid "API" diff --git a/docs/locales/ko/LC_MESSAGES/resources.po b/docs/locales/ko/LC_MESSAGES/resources.po index 2d28144ca..c3b146e80 100644 --- a/docs/locales/ko/LC_MESSAGES/resources.po +++ b/docs/locales/ko/LC_MESSAGES/resources.po @@ -40,12 +40,12 @@ msgid "Resources" msgstr "리소스" #: ../../resources.rst:9 226156d6aa574ae49ebbc29f7cc0ef0d -msgid "**PyMuPDF Pro**" -msgstr "**PyMuPDF Pro**" +msgid "**PyMuPDF Office**" +msgstr "**PyMuPDF Office**" #: ../../resources.rst:12 67c0d5d5d80c4878907a81eb3091f9c3 -msgid "For **Office** file support `try PyMuPDF Pro `." -msgstr "**Office** 파일 지원을 위해 `PyMuPDF Pro ` 를 사용해 보세요." +msgid "For **Office** file support `try PyMuPDF Office `." +msgstr "**Office** 파일 지원을 위해 `PyMuPDF Office ` 를 사용해 보세요." #: ../../resources.rst:20 28881f0ff7624205aac76bf0a54670e2 msgid "Find out about **PyMuPDF Utilities**" diff --git a/docs/ocr/index.rst b/docs/ocr/index.rst new file mode 100644 index 000000000..a9d18ec61 --- /dev/null +++ b/docs/ocr/index.rst @@ -0,0 +1,298 @@ + +.. include:: ../header.rst + +.. raw:: html + + + +.. _ocr-index: + +OCR +=== + +.. meta:: + :description: How automatic OCR works in PyMuPDF, when to force it, and how to swap in a different OCR engine. + +Overview +-------- + +|PyMuPDF| includes built-in OCR support for scanned documents and image-based PDFs. By default, OCR runs **automatically** when needed — you don't have to opt in. For more control, you can force OCR on specific pages, disable it entirely, or swap in a different OCR engine using the adaptor interface. + +.. note:: + + OCR requires a working Tesseract installation. See `OCR Installation ` for setup instructions. + +---- + +Hybrid OCR strategy +------------------------------------ + +|PyMuPDF| applies OCR only when it is genuinely required to obtain the complete text of a PDF page. If a page already contains sufficient extractable text, OCR is skipped entirely — avoiding unnecessary work and eliminating the risk of degrading high-quality digital text. + +When OCR is needed, |PyMuPDF| automatically selects the most suitable OCR plugin available in the runtime environment, balancing detection accuracy with processing speed. + +Its built-in OCR plugins implement a Hybrid OCR strategy: only those regions lacking extractable, legible text are passed to the OCR engine. This selective approach typically reduces OCR processing time by around 50% while improving recognition accuracy, since the engine focuses exclusively on the problematic regions. The recognized text is then merged back into the original page, enriching it without disturbing existing digital content. + + + + +Auto-OCR Behaviour +------------------ + +|PyMuPDF| inspects each page before extracting text. If a page contains **no selectable text** — meaning all content is rasterised into images — OCR is triggered automatically for that page. + +Pages that contain native text are never sent through OCR, even if they also contain embedded images. This keeps processing fast and avoids degrading already-clean text. + +.. code-block:: python + + import pymupdf + + # OCR runs automatically on any page with no selectable text + doc = pymupdf.open("scanned-document.pdf") + md_text = doc.to_markdown() + +The resulting Markdown is seamless — pages extracted via OCR and pages extracted natively are combined into a single output with no distinction between them. + +---- + +How OCR is Triggered +-------------------- + +There are two scenarios where OCR is applied automatically: + +**No text at all** — if a page contains roughly no text but is covered with images or many character-sized vectors, PyMuPDF uses `OpenCV `_ to check whether text is *probably* detectable on the page. This distinguishes image-based text (e.g. a scanned document) from ordinary pictures like photographs. + +**Garbled text** — if a page does contain text but too many characters are unreadable (e.g. ``"�����"``), OCR is applied **for the affected text areas only**, not the full page. This preserves already-readable text, images, and vectors while recovering only what is broken. + +.. note:: + + For these heuristics to work, both a `Tesseract installation ` and `OpenCV `_ must be available in your Python environment. If either is missing, no OCR is attempted. + +---- + +Decision Tree +~~~~~~~~~~~~~ + +OCR is applied only when all four of the following conditions are met: + +1. **PyMuPDF Layout is imported** — The Layout analysis module must be active. This is ``True`` by default. +2. **use_ocr is enabled** — The ``use_ocr`` parameter in the PyMuPDF API must be ``True``. This is the default value. +3. **Tesseract is installed** — Tesseract must be correctly installed on your system. See `OCR installation `. +4. **OpenCV is available** — The ``opencv-python`` package must be available in your Python environment. + +---- + +Forcing OCR +----------- + +In some cases you may want to force OCR even on pages that contain selectable text — for example, when the native text layer is corrupt, misencoded, or misaligned with the visual content. + +Use ``force_ocr=True`` to bypass the auto-detection check entirely: + +.. code-block:: python + + doc = pymupdf.open("document.pdf") + md_text = doc.to_markdown(force_ocr=True) + +.. warning:: + + Forcing OCR on clean, text-based PDFs will slow down processing significantly and may reduce output quality. Only use ``force_ocr=True`` when you have reason to distrust the native text layer. + +You can also force OCR on specific pages rather than the whole document: + +.. code-block:: python + + doc = pymupdf.open("document.pdf") + md_text = doc.to_markdown( + pages=[2, 3, 4], + force_ocr=True + ) + +---- + +Disabling OCR +------------- + +To prevent OCR from running at all — even on pages with no selectable text — set ``use_ocr=False``: + +.. code-block:: python + + doc = pymupdf.open("document.pdf") + md_text = doc.to_markdown(use_ocr=False) + +Pages with no selectable text will return empty strings in this mode. This is useful when you know your documents are always text-based, or when you want to handle OCR yourself in a downstream step. + +---- + +OCR Adaptors +------------ + +By default, PyMuPDF uses **Tesseract** and internal models for image pre-processing. If you need a different OCR engine you can plug in a custom adaptor. + +Built-in Adaptors +~~~~~~~~~~~~~~~~~ + +RapidOCR +^^^^^^^^ + +If `RapidOCR `_ and the RapidOCR ONNX Runtime are available, you can use a pre-made callable OCR function for it, which is provided in the ``PyMuPDF.ocr`` module as ``rapidocr_api.exec_ocr``. + +**Example** + +.. code-block:: python + + from pymupdf.ocr import rapidocr_api + + doc = pymupdf.open("document.pdf") + md = doc.to_markdown( + ocr_function=rapidocr_api.exec_ocr, + force_ocr=True + ) + +In this way RapidOCR can be used as an alternative OCR engine to Tesseract for all pages (if ``force_ocr=True``) or just for those pages which meet the default criteria for applying OCR (if ``force_ocr=False`` or omitted). + +RapidOCR & Tesseract Side-by-Side +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results. + +This pre-made callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``. + +**Example** + +.. code-block:: python + + from pymupdf.ocr import rapidtess_api + + doc = pymupdf.open("document.pdf") + md = doc.to_markdown( + ocr_function=rapidtess_api.exec_ocr, + force_ocr=True + ) + +.. list-table:: + :header-rows: 1 + :widths: 35 25 40 + + * - Adaptor + - Engines + - Notes + * - ``rapidocr_api.exec_ocr`` + - RapidOCR + - Requires RapidOCR and ONNX Runtime + * - ``rapidtess_api.exec_ocr`` + - RapidOCR & Tesseract + - Better accuracy for bounding box detection and text recognition + + +.. _ocr_custom_adaptor: + +Writing a Custom Adaptor +~~~~~~~~~~~~~~~~~~~~~~~~ + +An OCR adaptor can be defined by passing your own Python ``ocr_function`` method which follows a protocol as follows: + +.. code-block:: python + + def exec_ocr(page, dpi=300, pixmap=None): + """ + Custom OCR function to replace the default Tesseract-based implementation. + + Parameters: + - page: The PyMuPDF page object being processed. + - dpi: The resolution at which to render the page for OCR. + - pixmap: An optional pre-rendered pixmap of the page, if available. + + If a Pixmap is provided, the DPI parameter is ignored. Otherwise, an RGB + Pixmap is created from the page at the specified DPI. + """ + + # Your custom OCR logic here. + # The method should render the OCR'ed text onto the page + # so that PyMuPDF can extract it as usual. + + ... + +.. tip:: + + Custom adaptors receive a page or image object and must render what they "see" onto the source document. + PyMuPDF handles interpretation — your adaptor needs to handle the text recognition & rendering step. + +---- + +OCR Language Support +-------------------- + +When using the default Tesseract adaptor, you can specify one or more languages using Tesseract's language codes. + +Specify the language to be used by the Tesseract OCR engine. Default is ``"eng"`` (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign ``"+"``, for example ``"eng+deu"`` for English and German. + +.. code-block:: python + + doc = pymupdf.open("multilingual.pdf") + md_text = doc.to_markdown(ocr_language="eng+deu") + +Tesseract language packs must be installed separately on your system. For example, on Ubuntu: + +.. code-block:: bash + + sudo apt install tesseract-ocr-deu tesseract-ocr-fra + +See the page on :ref:`installing Tesseract language packs ` for further details. + +---- + + +How to OCR an Image +-------------------- +A supported image must first be converted to a :ref:`Pixmap`. The Pixmap can then be saved to a 1-page PDF. This page will look like the original image with the same width and height. It will contain a layer of text as recognized by Tesseract. + +The PDF can be generated via one of the methods :meth:`Pixmap.pdfocr_save` or :meth:`Pixmap.pdfocr_tobytes`, as a file on disk or as a PDF in memory. + +The text can be extracted and searched with the usual text extraction and search methods (:meth:`Page.get_text`, :meth:`Page.search_for`, etc.). Please also note the following important facts and prerequisites: + +* When converting the image to a Pixmap, please confirm that the color space is RGB and alpha is `False` (no transparency). Convert the original Pixmap if necessary. +* All text is written as "hidden" with Tesseract's own `GlyphLessFont`, a mono-spaced font with metrics comparable to Courier. +* All text has the properties regular and black (i.e. no bold, no italic, no information about the original fonts). +* Tesseract does not recognize vector graphics (i.e. no drawings / line-art). + +This approach is also recommended to OCR a complete scanned PDF: + +* Render each page to a :ref:`Pixmap` with desired resolution +* Append the resulting 1-page PDF to the output PDF + +How to OCR a Document Page +---------------------------- +Any supported document page can be OCR-ed -- either the complete page or only the image areas on it. + +Because optical character recognition is about one thousand times slower than standard text extraction, we make sure to do OCR only once per page and store the result in a :ref:`TextPage`. Using this TextPage for all subsequent extractions and text searches will then happen with |PyMuPDF|'s usual top speed. + +To OCR a document page, follow this approach: + +1. Determine whether OCR is needed / beneficial at all. A number of criteria can be used for this decision, like: + + * page is completely covered by an image + * no text exists on the page + * thousands of small vector graphics (indicating *simulated* text) + +2. OCR the page and store result in a :ref:`TextPage` object using an instruction like `tp = page.get_textpage_ocr(...)`. + +3. Refer to the produced :ref:`TextPage` in all subsequent text extractions and searches via the `textpage=tp` parameter. + +---- + +Performance Tips +---------------- + +OCR is the most compute-intensive part of the extraction pipeline. A few ways to keep it fast: + +- **Process only the pages you need** using the ``pages`` parameter to avoid running OCR on the entire document. +- **Cache results** — write the output to disk after the first run so you don't re-process the same file. +- **Use** ``force_ocr=False`` (the default) so clean pages skip OCR entirely. +- **Resize images before passing to OCR** — very high DPI scans can slow Tesseract down without improving accuracy. + +---- + +.. include:: ../footer.rst \ No newline at end of file diff --git a/docs/ocr/tesseract-language-packs.rst b/docs/ocr/tesseract-language-packs.rst index d57839bb8..34fd6ab1d 100644 --- a/docs/ocr/tesseract-language-packs.rst +++ b/docs/ocr/tesseract-language-packs.rst @@ -1,7 +1,7 @@ .. include:: ../header.rst -.. _pymupdf-pro: + .. raw:: html diff --git a/docs/pymupdf-pro/index.rst b/docs/pymupdf-office/index.rst similarity index 51% rename from docs/pymupdf-pro/index.rst rename to docs/pymupdf-office/index.rst index 2b77d2b9f..88cb5c51c 100644 --- a/docs/pymupdf-pro/index.rst +++ b/docs/pymupdf-office/index.rst @@ -1,7 +1,7 @@ .. include:: ../header.rst -.. _pymupdf-pro: +.. _pymupdf-office: .. raw:: html @@ -9,25 +9,18 @@ document.getElementById("headerSearchWidget").action = '../search.html'; -PyMuPDF Pro +PyMuPDF Office ============= +Enhance |PyMuPDF| capability with **Office** document support. -|PyMuPDF Pro| is a set of *commercial extensions* for |PyMuPDF|. - -Enhance |PyMuPDF| capability with **Office** document support & **RAG/LLM** integrations. - -- Enables Office document handling, including ``doc``, ``docx``, ``hwp``, ``hwpx``, ``ppt``, ``pptx``, ``xls``, ``xlsx``, and others. +- Enables Office document handling, including ``doc``, ``docx``, ``ppt``, ``pptx``, ``xls``, ``xlsx``, ``hwp``, ``hwpx``. - Supports text and table extraction, document conversion and more. -- Includes the commercial version of |PyMuPDF4LLM|. +- Includes everything from |PyMuPDF| To enquire about obtaining a commercial license, then `use this contact page `_. -Trial license keys are available for evaluation purposes. Please `fill out the form on this page `_ to obtain a trial key. - -.. note:: - - A licensed version of |PyMuPDF Pro| also gives you a licensed version of |PyMuPDF4LLM|. If you are interested in using the |PyMuPDF4LLM| package you should install it separately. +Trial license keys are available for evaluation purposes. Please `fill out the form on this page `_ to obtain a trial key. Platform support @@ -35,16 +28,16 @@ Platform support Available for these platforms only: -- Windows x86_64. -- Linux x86_64 (glibc). -- MacOS x86_64. -- MacOS arm64. +- Windows x86_64 +- Linux x86_64 (glibc) +- MacOS x86_64 +- MacOS arm64 Office file support ---------------------- -In addition to the `standard file types supported by PyMuPDF `, |PyMuPDF Pro| supports: +In addition to the `standard file types supported by PyMuPDF `, |PyMuPDF Office| supports: .. list-table:: :header-rows: 1 @@ -78,63 +71,62 @@ Install via pip with: .. code-block:: bash - pip install pymupdfpro - + pip install pymupdf-office Loading an **Office** document ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Import |PyMuPDF Pro| and you can then reference **Office** documents directly, e.g.: +Import |PyMuPDF Office| and you can then reference **Office** documents directly, e.g.: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock() - # PyMuPDF has now been extended with PyMuPDF Pro features, with some restrictions. + import pymupdf.office + pymupdf.office.unlock() + # PyMuPDF has now been extended with PyMuPDF Office features, with some restrictions. doc = pymupdf.open("my-office-doc.xls") .. note:: - All standard |PyMuPDF| functionality is exposed as expected - |PyMuPDF Pro| handles the extended **Office** file types + All standard |PyMuPDF| functionality is exposed as expected - |PyMuPDF Office| handles the extended **Office** file types -From then on you can work with document pages just as you would do normally, but with respect to the `restrictions `. +From then on you can work with document pages just as you would do normally, but with respect to the `restrictions `. -.. _PyMuPDFPro_Restrictions: +.. _PyMuPDFOffice_Restrictions: Restrictions ~~~~~~~~~~~~~~~~~~~~ -|PyMuPDF Pro| functionality is restricted without a license key as follows: +|PyMuPDF Office| functionality is restricted without a license key as follows: **Only the first 3 pages of any document will be available.** - To unlock full functionality you should `obtain a trial key `_. + To unlock full functionality you should `obtain a trial key `_. -.. _PyMuPDFPro_TrialKeys: +.. _PyMuPDFOffice_TrialKeys: Trial keys ----------------------- -To obtain a license key `please fill out the form on this page `_. You will then have the trial key emailled to the address you submitted. +To obtain a license key `please fill out the form on this page `_. You will then have the trial key emailled to the address you submitted. Using a key ~~~~~~~~~~~~~~~~ -Initialize |PyMuPDF Pro| with a key as follows: +Initialize |PyMuPDF Office| with a key as follows: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock(my_key) - # PyMuPDF has now been extended with PyMuPDF Pro features. + import pymupdf.office + pymupdf.office.unlock(my_key) + # PyMuPDF has now been extended with PyMuPDF Office features. -This will allow you to evaluate the product for a limited time. If you want to use |PyMuPDF Pro| after this time you should then `enquire about obtaining a commercial license `_. +This will allow you to evaluate the product for a limited time. If you want to use |PyMuPDF Office| after this time you should then `enquire about obtaining a commercial license `_. @@ -145,24 +137,24 @@ Converting **Office** document to |PDF| ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Use the :meth:`office_to_pdf()` method to convert an **Office** document to |PDF|, e.g.: +Use the :meth:`to_pdf()` method to convert an **Office** document to |PDF|, e.g.: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock() + import pymupdf.office + pymupdf.office.unlock(my_key) - pymupdf.pro.office_to_pdf("input.docx", "output.pdf") + pymupdf.office.to_pdf("input.docx", "output.pdf") -If you require a byte representation of the PDF data, you can use the :meth:`office_to_pdf()` without specifying an output file, e.g.: +If you require a byte representation of the PDF data, you can use the :meth:`to_pdf()` without specifying an output file, e.g.: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock() + import pymupdf.office + pymupdf.office.unlock(my_key) - pdfdata = pymupdf.pro.office_to_pdf("input.docx") + pdfdata = pymupdf.office.to_pdf("input.docx") **Office** document to Images @@ -183,36 +175,36 @@ In order to convert an **Office** document to images, you should iterate the doc **Office** document to |Markdown| ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Use the :meth:`office_to_markdown()` method to convert an **Office** document to |Markdown|, e.g.: +Use the :ref:`to_markdown() ` method to convert an **Office** document to |Markdown|, e.g.: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock() + import pymupdf.office + pymupdf.office.unlock(my_key) - pymupdf.pro.office_to_markdown("input.docx", "output.md") + pymupdf.office.to_markdown("input.docx", "output.md") -If you require a string representation of the Markdown data, you can use the :meth:`office_to_markdown()` without specifying an output file. +If you require a string representation of the Markdown data, you can use the :ref:`to_markdown() ` method without specifying an output file. **Office** document to |JSON| ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -Use the :meth:`office_to_json()` method to convert an **Office** document to |JSON|, e.g.: +Use the :ref:`to_json() ` method to convert an **Office** document to |JSON|, e.g.: .. code-block:: python - import pymupdf.pro - pymupdf.pro.unlock() + import pymupdf.office + pymupdf.office.unlock(my_key) - pymupdf.pro.office_to_json("input.docx", "output.json") + pymupdf.office.to_json("input.docx", "output.json") -If you require a string representation of the JSON data, you can use the :meth:`office_to_json()` without specifying an output file. +If you require a string representation of the JSON data, you can use the :ref:`to_json() ` method without specifying an output file. Fonts ----------------------- -By default `pymupdf.pro.unlock()` searches for all installed font directories. +By default `pymupdf.office.unlock()` searches for all installed font directories. This can be controlled with keyword-only args: @@ -222,14 +214,27 @@ This can be controlled with keyword-only args: If None (the default) we use true if `os.environ['PYMUPDFPRO_FONT_PATH_AUTO']` is '1'. If true we append all system font directories. -Function `pymupdf.pro.get_fontpath()` returns a tuple of all font directories used by `unlock()`. +Function `pymupdf.office.get_fontpath()` returns a tuple of all font directories used by `unlock()`. API ----------------------- +.. _pymupdf_office_api_unlock: + +.. method:: unlock(my_key: str = None, *, fontpath: str | list | tuple = None, fontpath_auto: bool = None) -.. method:: office_to_pdf(input_path: str, output_path:str = None) -> bytes | None + Unlocks the PyMuPDF Office functionality with the provided key. + + :arg str my_key: the license key for PyMuPDF Office. + :arg str | list | tuple fontpath: specific font directories, either as a list/tuple or `os.sep`-separated string. + :arg bool fontpath_auto: Whether to append system font directories. + + Grants access to the full functionality of PyMuPDF Office. For a free trial key please visit: `https://pymupdf.io/office/try/ `_. + +.. _pymupdf_office_api_to_pdf: + +.. method:: to_pdf(input_path: str, output_path:str = None) -> bytes | None Reads the input file and converts its contents into |PDF| format. @@ -240,7 +245,9 @@ API :returns: Either bytes of the PDF content, `None` if `output_path` is specified. -.. method:: office_to_markdown(input_path: str, output_path:str = None) -> str | None +.. _pymupdf_office_api_to_markdown: + +.. method:: to_markdown(input_path: str, output_path:str = None) -> str | None Reads the input file and outputs the text of its pages in |Markdown| format. @@ -251,7 +258,10 @@ API :returns: Either a string of the Markdown representation, `None` if `output_path` is specified. -.. method:: office_to_json(input_path: str, output_path:str = None) -> str | None + +.. _pymupdf_office_api_to_json: + +.. method:: to_json(input_path: str, output_path:str = None) -> str | None Reads the input file and outputs the text of its pages in |JSON| format. @@ -266,7 +276,7 @@ API .. raw:: html - + diff --git a/docs/pymupdf4llm/api.rst b/docs/pymupdf4llm/api.rst index edca4cc99..0f48a0750 100644 --- a/docs/pymupdf4llm/api.rst +++ b/docs/pymupdf4llm/api.rst @@ -177,6 +177,42 @@ The PyMuPDF4LLM API :returns: Either a string of the combined text of all selected document pages, or a list of dictionaries if `page_chunks=True`. + + +.. method:: to_json(doc: pymupdf.Document | str, *, **kwargs) -> str + + Parses the document and the specified pages and converts the result into a `JSON formatted string `_. + + :arg Document,str doc: the file, to be specified either as a file path string, or as a |PyMuPDF| :class:`Document` (created via `pymupdf.open`). In order to use `pathlib.Path` specifications, Python file-like objects, documents in memory etc. you **must** use a |PyMuPDF| :class:`Document`. + + :arg bool use_ocr: |PyMuPDFLayoutMode_Valid| use :ref:`OCR capability ` to help analyse the page. + + :arg str ocr_language: |PyMuPDFLayoutMode_Valid| specify the language to be used by the Tesseract OCR engine. Default is "eng" (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign "+", for example "eng+deu" for English and German. + + :arg int ocr_dpi: |PyMuPDFLayoutMode_Valid| specify the desired image resolution in dots per inch for applying OCR to the intermediate image of the page. Default value is 400. Only relevant if the page has been determined to profit from OCR (no or few text, most of the page covered by images or character-like vectors, etc.). Large values may increase the OCR precision but increase memory requirements and processing time. There also is a risk of over-sharpening the image which may decrease OCR precision. So the default value should probably be sufficiently high. + + :arg int image_dpi: specify the desired image resolution in dots per inch. Default value is 150. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + + :arg str image_format: specify the desired image format via its extension. Default is "png" (portable network graphics). Another popular format may be "jpg". Possible values are all :ref:`supported output formats `. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + + :arg str image_path: store images in this folder. Relevant if `write_images=True`. Default is the path of the script directory. Page areas classified as "picture" will be written as image files to the specified location. The image file names will be of the format `{image_path}/{filename}-pagenumber-image_number.{image_format}`. + + :arg bool force_text: generate text output for text that is written upon areas that are classified as "picture" by the layout module. This may be especially be useful when picture content is not stored. + + :arg bool show_progress: display a progress bar during processing. + + :arg bool embed_images: store image binaries for "picture" boundary boxes. Base64-encoded images are included in the JSON output. Ignores `image_path` if used. This may drastically increase the size of your JSON text. + + :arg bool write_images: store image files "picture" boundary boxes. When encountering images, image files will be created from the respective page area and stored in the specified folder. Any text contained in these areas will still be included in the text output. + + :arg list pages: optional, the pages to consider for output (caution: specify 0-based page numbers). If omitted (`None`) all pages are processed. Specify any valid Python sequence containing integers between `0` and `page_count - 1`. + + :rtype: str + + See `JSON Schema `_ for the structure of the output JSON string. + + + .. method:: to_text(doc: pymupdf.Document | str, *, **kwargs) -> str Reads the pages of the file and outputs the text of its pages in plain text (|TXT|) format. @@ -226,64 +262,191 @@ The PyMuPDF4LLM API See: :ref:`box classes ` -.. method:: to_json(doc: pymupdf.Document | str, *, **kwargs) -> str +.. _pymupdf4llm-api-to-chunks: - Parses the document and the specified pages and converts the result into a `JSON formatted string `_. +.. method:: to_chunks(doc: pymupdf.Document | str, **kwargs) -> ChunkedDocument - :arg Document,str doc: the file, to be specified either as a file path string, or as a |PyMuPDF| :class:`Document` (created via `pymupdf.open`). In order to use `pathlib.Path` specifications, Python file-like objects, documents in memory etc. you **must** use a |PyMuPDF| :class:`Document`. + Creates retrieval-oriented chunks from a document. - :arg bool use_ocr: |PyMuPDFLayoutMode_Valid| use :ref:`OCR capability ` to help analyse the page. + Chunk boundaries are determined from layout information, including box boundaries, page breaks, vertical gaps, font changes, and structural hints. Token limits guide chunk assembly but are not guarantees: indivisible content, such as a preserved table, may exceed ``max_tokens``. - :arg str ocr_language: |PyMuPDFLayoutMode_Valid| specify the language to be used by the Tesseract OCR engine. Default is "eng" (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign "+", for example "eng+deu" for English and German. + :arg int max_tokens: target maximum number of tokens per chunk. Default is ``400``. - :arg int ocr_dpi: |PyMuPDFLayoutMode_Valid| specify the desired image resolution in dots per inch for applying OCR to the intermediate image of the page. Default value is 400. Only relevant if the page has been determined to profit from OCR (no or few text, most of the page covered by images or character-like vectors, etc.). Large values may increase the OCR precision but increase memory requirements and processing time. There also is a risk of over-sharpening the image which may decrease OCR precision. So the default value should probably be sufficiently high. + :arg int min_tokens: minimum size used when merging small neighboring chunks. Default is ``120``. A value of ``0`` disables this minimum-size behavior. - :arg int image_dpi: specify the desired image resolution in dots per inch. Default value is 150. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + :arg float breakpoint_threshold: boundary score threshold used when splitting chunks. Default is ``0.5``. - :arg str image_format: specify the desired image format via its extension. Default is "png" (portable network graphics). Another popular format may be "jpg". Possible values are all :ref:`supported output formats `. Only relevant if one of the parameters `write_images=True` or `embed_images=True` is used. + :arg bool merge_small_chunks: whether to merge chunks that are below ``min_tokens`` with neighboring chunks. Default is ``True``. - :arg str image_path: store images in this folder. Relevant if `write_images=True`. Default is the path of the script directory. Page areas classified as "picture" will be written as image files to the specified location. The image file names will be of the format `{image_path}/{filename}-pagenumber-image_number.{image_format}`. + :arg str table_mode: ``"preserve"`` keeps table content together as one chunk; ``"isolate"`` prevents tables from being merged into neighboring chunks. Default is ``"preserve"``. - :arg bool force_text: generate text output for text that is written upon areas that are classified as "picture" by the layout module. This may be especially be useful when picture content is not stored. + :arg bool respect_section_starts: whether to keep a chunk starting a detected section separate from the preceding section when merging to meet token budgets. Default is ``True``. - :arg bool show_progress: display a progress bar during processing. + :arg str header_footer_mode: controls handling of page headers and footers. ``"exclude"`` omits them from chunk text, ``"auto"`` omits repeated headers and footers, and ``"include"`` retains them. Default is ``"exclude"``. The element registry in the result retains the parsed layout elements in all modes. - :arg bool embed_images: store image binaries for "picture" boundary boxes. Base64-encoded images are included in the JSON output. Ignores `image_path` if used. This may drastically increase the size of your JSON text. + :arg str sentence_splitter: sentence splitting mode: ``"default"`` or ``"multilingual"`` (for additional CJK support). Default is ``"default"``. - :arg bool write_images: store image files "picture" boundary boxes. When encountering images, image files will be created from the respective page area and stored in the specified folder. Any text contained in these areas will still be included in the text output. + :arg tokenizer: token-counting strategy. Use ``None`` for the default character-based estimate, a callable accepting text and returning an integer token count, or a ``tiktoken`` encoding name. Default is ``None``. - :arg list pages: optional, the pages to consider for output (caution: specify 0-based page numbers). If omitted (`None`) all pages are processed. Specify any valid Python sequence containing integers between `0` and `page_count - 1`. + :arg dict weights: optional overrides for the layout boundary-score weights. The default ``None`` uses the built-in weights for box boundaries, box classes, page breaks, vertical and horizontal gaps, font changes, headings, headers and footers, lists, tables, and captions. - :rtype: str + The enum-valued arguments accept only the values listed above. ``max_tokens`` must be a positive integer and ``min_tokens`` a non-negative integer; unknown keyword arguments raise :exc:`TypeError` and invalid values raise :exc:`ValueError`. - See `JSON Schema `_ for the structure of the output JSON string. + :returns: a :class:`ChunkedDocument`, a sequence of :class:`Chunk` objects with document-level element, table, figure, and section views. It also provides joined text, ID-based lookup, diagnostics, JSON-safe export, and reassembly with different chunk-budget parameters. -.. _pymupdf4llm-api-boxclasses: +.. class:: ChunkedDocument -.. note:: + Retrieval-ready chunks and document-level structure returned by :meth:`to_chunks`. This object implements the sequence protocol over its chunks: use ``len(cd)``, integer indexing, iteration, or slicing. Integer indexing returns a chunk; slicing returns a list of chunks. Chunk ids are ``c{n}`` in reading order, so ``cd.get("c3")`` addresses the same chunk as ``cd[3]``. - **About box classes** + The object also keeps layout elements and table, figure, and section views. These views are linked to their owning chunks through ids. Element ids have the form ``p{page}.b{box}`` (1-based page number and 0-based box index). The ``hierarchy`` property represents sections as a tree; its root has level 0. - If `page_chunks = True` the return objects for `to_markdown` & `to_text` contains a list of dictionaries representing the layout boundary boxes `page_boxes`, within that a key ``class`` indicates the type of box content therein. + .. attribute:: chunks - The return object for `to_json` contains a similar key called ``boxclass``. + The chunks as a tuple, equivalent to ``tuple(cd)``. - The possible string values are for this ``class`` / ``boxclass`` key are: + .. attribute:: text + + All chunk text joined in reading order with blank lines between chunks. The value is computed lazily and cached. + + .. attribute:: elements + + A list containing every parsed layout box, including headers and footers excluded from chunk text. Each element has an id, page and box indices, box class, bounding box, canonical text, and a flag indicating whether it is a header or footer. + + .. attribute:: tables + + A list of table views. Each table is linked to its owning chunk when one exists, and provides its id, source element id, page, bounding box, canonical text, optional caption and section id, and token count. The canonical text is Markdown or HTML according to the table output mode. + + .. attribute:: figures + + A list of figure and formula views, linked to their owning chunk when one exists. Views include source location, extracted text when available, caption and section links, and image data when it was extracted. + + .. attribute:: sections + + A list of section views with title, heading level, page range, path, child chunk ids, token count, and section body text. + + .. attribute:: hierarchy + + Sections arranged as a tree of section nodes. The root node has level 0 and no section id. + + .. attribute:: params + + A read-only mapping of the parameters used to create this object. + + .. attribute:: diagnostics + + A dictionary of extraction and ingestion facts: chunk, element, table, figure, section, and page counts; pages without chunks; reasons for an empty result; figures without extracted text; degenerate tables; and the number of excluded header/footer units. This reports facts only; consumers decide how to act on them. + + .. method:: get(id: str, default=...) + + Return a chunk, table, figure, section, or element by its public id: ``c{n}``, ``t{n}``, ``f{n}``, ``s{n}``, or ``p{page}.b{box}``. Raises :exc:`KeyError` if the id is unknown and no default is provided; otherwise returns ``default`` for an unknown id. + + .. method:: to_dicts(*, include_tagged: bool = True) -> list[dict] + + Return the chunks as flat, JSON-safe dictionaries. Each dictionary contains ``id``, ``text``, ``content_hash``, and ``metadata``; it also contains ``tagged_content`` unless ``include_tagged=False``. + + .. method:: to_json(*, include_tagged: bool = True, **json_kwargs) -> str + + Serialize the result of :meth:`to_dicts` as JSON. ``include_tagged`` controls whether tagged content is included; additional keyword arguments are passed to :func:`json.dumps`. + + .. method:: reassemble_chunks(**params) -> ChunkedDocument + + Create a new ``ChunkedDocument`` by assembling chunks again from the retained parsed units; the document does not need to be parsed again. Accepts only ``max_tokens``, ``min_tokens``, ``breakpoint_threshold``, ``table_mode``, ``merge_small_chunks``, and ``respect_section_starts``. Other chunking parameters and parse options raise :exc:`ValueError` because they require a new :meth:`to_chunks` call. The original object is unchanged. + + .. method:: to_langchain_documents(*, doc_id: str | None = None) + + Return the chunks as LangChain ``Document`` objects. Requires the optional ``langchain-core`` package. If ``doc_id`` is provided, it is prefixed to each exported chunk id to make ids unique across documents. + + .. method:: to_llama_nodes(*, doc_id: str | None = None) + + Return the chunks as LlamaIndex ``TextNode`` objects. Requires the optional ``llama-index-core`` package. If ``doc_id`` is provided, it is prefixed to each exported chunk id to make ids unique across documents; the original chunk id is retained in node metadata. + + +.. class:: Chunk + + A finalized, retrieval-ready portion of a document, returned as an item in :class:`ChunkedDocument`. Treat chunks returned by :meth:`to_chunks` as read-only. Mutating ``text`` or ``metadata`` after reading ``content_hash`` can leave that cached hash, and caches on the owning ``ChunkedDocument``, out of date. To change chunk boundaries, use :meth:`ChunkedDocument.reassemble_chunks`; keep application-specific data separately, keyed by the chunk id. + + .. attribute:: id + + The chunk id, formatted as ``c{n}``, where ``n`` is its position in this ``ChunkedDocument``. Chunk ids are local to the result and may change when chunks are reassembled with different parameters. + + .. attribute:: text + + The chunk's Markdown text, rendered from its layout content. + + .. attribute:: tagged_content + + Context-enriched text for embedding. It includes the chunk's section path, page, and non-paragraph content types where available, followed by its Markdown text. This is separate from ``text`` and is included by default in :meth:`ChunkedDocument.to_dicts` and :meth:`ChunkedDocument.to_json`. + + .. attribute:: metadata + + A :class:`ChunkMetadata` instance containing page, layout, structure, token-count, and provenance information for this chunk. + + .. attribute:: content_hash + + A lazily computed SHA-256 hash of ``text`` after runs of whitespace have been collapsed to one space and leading and trailing whitespace has been removed. The hash is cached. It is stable across rendering-only whitespace changes, but changes when the normalized text changes. + +.. class:: ChunkMetadata + + Payload-ready metadata associated with a :class:`Chunk`. Its fields are also included in the ``metadata`` object returned by :meth:`ChunkedDocument.to_dicts` and :meth:`ChunkedDocument.to_json`. + + .. attribute:: page_start + + 1-based number of the first page contributing content to the chunk. + + .. attribute:: page_end + + 1-based number of the last page contributing content to the chunk. + + .. attribute:: bboxes + + List of bounding boxes in ``(page, x0, y0, x1, y1)`` form. Page numbers are 1-based. + + .. attribute:: types + + Content types present in the chunk, in reading order, such as ``"heading"``, ``"paragraph"``, ``"table"``, ``"list"``, ``"figure"``, or ``"caption"``. + + .. attribute:: section_id + + Id of the innermost detected section containing the chunk, in ``s{n}`` form, or ``None`` when the chunk is not associated with a section. + + .. attribute:: section_path + + List of section titles from the document's section hierarchy to the innermost section. Empty when no section applies. + + .. attribute:: token_count + + Token count used by the chunk assembler. Depending on the tokenizer and the final rendered text, this can differ slightly from a separate tokenization of ``text``; token budgets are targets rather than hard limits. + + .. attribute:: element_ids + + Ids of the source layout elements contributing to the chunk, formatted as ``p{page}.b{box}`` (1-based page number, 0-based box index). + + .. attribute:: table_ids + + Ids of table views associated with the chunk, formatted as ``t{n}``. + + .. attribute:: figure_ids + + Ids of figure or formula views associated with the chunk, formatted as ``f{n}``. + + .. attribute:: lists + + List groups found in the chunk. Each group contains ``items`` with item text, page number, and bounding box, and ``bboxes`` for the group. + + .. attribute:: ocr + + ``True`` if content contributing to this chunk came from a page processed by OCR; otherwise ``False``. + + .. attribute:: file_path + + Source document path, or ``None`` when no path is available. + + .. attribute:: page_count + + Number of pages in the source document, or ``None`` when unavailable. - .. code-block:: bash - text - picture - table - caption - title - section-header - page-header - page-footer - list-item - footnote - formula .. _pymupdf4llm-api-layout: @@ -512,6 +675,36 @@ This is a version of previous **example 2** that uses :class:`TocHeaders` for he ----- + + +.. _pymupdf4llm-api-boxclasses: + +.. note:: + + **About box classes** + + If `page_chunks = True` the return objects for `to_markdown` & `to_text` contains a list of dictionaries representing the layout boundary boxes `page_boxes`, within that a key ``class`` indicates the type of box content therein. + + The return object for `to_json` contains a similar key called ``boxclass``. + + The possible string values are for this ``class`` / ``boxclass`` key are: + + .. code-block:: bash + + text + picture + table + caption + title + section-header + page-header + page-footer + list-item + footnote + formula + +----- + For a list of changes, please see file `CHANGES.md `_. .. rubric:: Footnotes @@ -560,3 +753,5 @@ For a list of changes, please see file `CHANGES.md + + diff --git a/docs/pymupdf4llm/index.rst b/docs/pymupdf4llm/index.rst index 46d63af06..07ee2ade3 100644 --- a/docs/pymupdf4llm/index.rst +++ b/docs/pymupdf4llm/index.rst @@ -20,326 +20,15 @@ PyMuPDF4LLM |PyMuPDF4LLM| makes it easy to extract document content in the format you need for **LLM** & **RAG** environments. It supports structured data extraction to :ref:`Markdown `, :ref:`JSON ` and :ref:`TXT `, as well as :ref:`LlamaIndex ` and :ref:`LangChain ` integration. -.. important:: - - You can also extend the supported file types to also include **Office** document formats (DOC/DOCX, XLS/XLSX, PPT/PPTX, HWP/HWPX) by :ref:`using PyMuPDF Pro with PyMuPDF4LLM `. - -Features -------------------------------- - - - Support for Markdown, JSON and plain text output formats. - - Support for multi-column pages. - - Support for image and vector graphics extraction. - - Layout analysis for better semantic understanding of document structure. - - Support for page chunking output. - - Automatic detection of pages which profit from OCR and support for various OCR engines. - - Integration with :ref:`LlamaIndex ` & :ref:`LangChain `. - API ------- See: :doc:`api`. -Installation ----------------- - - -Install the package via **pip** with: - - -.. code-block:: bash - - pip install pymupdf4llm - - -Extracting -------------------------------- - - -.. _extracting_as_md: - -As **Markdown** -~~~~~~~~~~~~~~~~~~~~~~~ - -To retrieve your document content in **Markdown** use the :meth:`to_markdown` method as follows: - -.. code-block:: python - - import pymupdf4llm - md = pymupdf4llm.to_markdown("input.pdf") - - - -.. _extracting_as_json: - -As **JSON** -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -To retrieve your document content in **JSON** use the :meth:`to_json` method as follows: - -.. code-block:: python - - import pymupdf4llm - json = pymupdf4llm.to_json("input.pdf") - -The JSON export will give you bounding box information and layout data for each element on the page. This can be used to create your own custom output formats or to simply have more detailed information about the document structure for RAG workflows & LLM integrations. - - -.. _extracting_as_txt: - -As **TXT** -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -To retrieve your document content in **TXT** use the :meth:`to_text` method as follows: - -.. code-block:: python - - import pymupdf4llm - txt = pymupdf4llm.to_text("input.pdf") - - - ----- - -.. note:: - Instead of using filename strings as above, one can also provide a :ref:`PyMuPDF Document `. - - Finally we can save the output to an external file as follows:: - - from pathlib import Path - suffix = ".md" # or ".json" or ".txt" - Path(doc.name).with_suffix(suffix).write_bytes(md.encode()) - -Headers & Footers -~~~~~~~~~~~~~~~~~~~~~~~ - - -Many documents will have header and footer information on each page of a PDF which you may or may not want to include. This information can be repetitive and simply not needed ( e.g. the same logo and document title or page number information is not always required when it comes to extracting the document content ). - -|PyMuPDF4LLM| is trained in detecting these typical document elements and able to omit them. - -So in this case we can adjust our API calls to ignore these elements as follows:: - - md = pymupdf4llm.to_markdown(doc, header=False, footer=False) - - -.. note:: - - Please note that page ``header`` / ``footer`` exclusion is not applicable to JSON output as it aims to always represent all data for the included pages. Please refer to :doc:`api` for more. - - -Integrations -------------------------------- - -.. _integration_with_llamaindex: - -With **LlamaIndex** -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -|PyMuPDF4LLM| supports direct conversion to a **LlamaIndex** document. A document is first converted into **Markdown** format and then a **LlamaIndex** document is returned as follows: - - -.. code-block:: python - - import pymupdf4llm - llama_reader = pymupdf4llm.LlamaMarkdownReader() - llama_docs = llama_reader.load_data("input.pdf") - - -.. _integration_with_langchain: - -With **LangChain** -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -|PyMuPDF4LLM| also supports **LangChain** integration, see the `PyMuPDF4LLM Document Loader`_ for more details. - - -.. _using_pymupdf4llm_with_pymupdfpro: - -Using with |PyMuPDF Pro| ---------------------------- - - -For **Office** document support, |PyMuPDF4LLM| works seamlessly with |PyMuPDF Pro|. Assuming you have :doc:`../pymupdf-pro/index` installed you will be able to work with **Office** documents as expected: - - -.. code-block:: python - - import pymupdf4llm - import pymupdf.pro - pymupdf.pro.unlock() - md = pymupdf4llm.to_markdown("sample.doc") - - -.. _pymupdf4llm_and_layout: - -PyMuPDF4LLM & PyMuPDF Layout ------------------------------------ - -By default |PyMuPDF4LLM| includes a `layout analysis module`_ to enhance output results. To disable this module you can do so by calling the :meth:`use_layout` method. - - - -OCR --------- - -PyMuPDF4LLM includes built-in OCR support for scanned documents and image-based PDFs. By default, OCR runs **automatically** when needed — you don't have to opt in. For more control, you can force OCR on specific pages, disable it entirely, or swap in a different OCR engine using the adaptor interface. - -.. note:: - - If you want to use an OCR engine other than Tesseract, see :ref:`OCR Engines ` for details. - - -Hybrid OCR strategy -~~~~~~~~~~~~~~~~~~~~~~~~~ - -PyMuPDF4LLM applies OCR only when it is genuinely required to obtain the complete text of a PDF page. If a page already contains sufficient extractable text, OCR is skipped entirely — avoiding unnecessary work and eliminating the risk of degrading high-quality digital text. - -When OCR is needed, PyMuPDF4LLM automatically selects the most suitable OCR plugin available in the runtime environment, balancing detection accuracy with processing speed. - -Its built-in OCR plugins implement a Hybrid OCR strategy: only those regions lacking extractable, legible text are passed to the OCR engine. This selective approach typically reduces OCR processing time by around 50% while improving recognition accuracy, since the engine focuses exclusively on the problematic regions. The recognized text is then merged back into the original page, enriching it without disturbing existing digital content. - - ----- - - -Auto-OCR Behaviour -~~~~~~~~~~~~~~~~~~~~~~~~~~ - -PyMuPDF4LLM inspects each page before extracting text. If a page contains **no selectable text** — meaning all content is rasterised into images — OCR is triggered automatically for that page. - -Pages that contain native text only are never sent through OCR. This keeps processing fast and avoids degrading already-clean text. - -.. code-block:: python - - import pymupdf4llm - - # OCR runs automatically on any page with no selectable text - md_text = pymupdf4llm.to_markdown("scanned-document.pdf") - -The resulting Markdown is seamless — pages extracted via OCR and pages extracted natively are combined into a single output with no distinction between them. - ----- - -How OCR is Triggered -~~~~~~~~~~~~~~~~~~~~~~~~~~ - -There are two scenarios where OCR is applied automatically: - -**No text at all** — if a page contains roughly no text but is covered with images or many character-sized vectors, PyMuPDF4LLM checks whether text is *probably* detectable on the page. This distinguishes image-based text (e.g. a scanned document) from ordinary pictures like photographs. - -**Garbled text** — if a page does contain text but too many characters are unreadable (e.g. ``"�����"``), OCR is applied **for the affected text areas only**, not the full page. This preserves already-readable text, images, and vectors while recovering only what is broken. - - ----- - -Forcing OCR -~~~~~~~~~~~~~ - -In some cases you may want to force OCR even on pages that contain selectable text — for example, when the native text layer is corrupt, misencoded, or misaligned with the visual content. - -Use ``force_ocr=True`` to bypass the auto-detection check entirely: - -.. code-block:: python - - md_text = pymupdf4llm.to_markdown("document.pdf", force_ocr=True) - -.. warning:: - - Forcing OCR on clean, text-based PDFs will slow down processing significantly and may reduce output quality. Only use ``force_ocr=True`` when you have reason to distrust the native text layer. - -You can also force OCR on specific pages rather than the whole document: - -.. code-block:: python - - md_text = pymupdf4llm.to_markdown( - "document.pdf", - pages=[2, 3, 4], - force_ocr=True - ) - ----- - -Disabling OCR -~~~~~~~~~~~~~ - -To prevent OCR from running at all — even on pages with no selectable text — set ``use_ocr=False``: - -.. code-block:: python - - md_text = pymupdf4llm.to_markdown("document.pdf", use_ocr=False) - -Pages with no selectable text will return empty strings in this mode. This is useful when you know your documents are always text-based, or when you want to handle OCR yourself in a downstream step. - ----- - -.. _ocr-adaptors: -.. _ocr-engines: -.. _ocr-plugins: - -OCR Engines -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - -Other OCR Engines (OCR Adaptors or Plugins) can be used with PyMuPDF4LLM. - -See :doc:`ocr-plugins` for details on how to use different OCR engines with PyMuPDF4LLM, including Tesseract, RapidOCR, and how to implement your own custom OCR function. - - - -OCR Language Support -~~~~~~~~~~~~~~~~~~~~~~~~~~ - -When using the default Tesseract adaptor, you can specify one or more languages using Tesseract's language codes. - -Specify the language to be used by the Tesseract OCR engine. Default is ``"eng"`` (English). Make sure that the respective language data files are installed. Remember to use correct Tesseract language codes. Multiple languages can be specified by concatenating the respective codes with a plus sign ``"+"``, for example ``"eng+deu"`` for English and German. - -.. code-block:: python - - md_text = pymupdf4llm.to_markdown("multilingual.pdf", - ocr_language="eng+deu") - -Tesseract language packs must be installed on your system. For example, on Ubuntu: - -.. code-block:: bash - - sudo apt install tesseract-ocr-deu tesseract-ocr-fra - -See the page on :ref:`installing Tesseract language packs ` for further details. - ----- - -Performance Tips -~~~~~~~~~~~~~~~~~~~~~~~~~~ - -OCR is the most compute-intensive part of the extraction pipeline. A few ways to keep it fast: - -- **Process only the pages you need** using the ``pages`` parameter to avoid running OCR on the entire document. -- **Cache results** — write the output to disk after the first run so you don't re-process the same file. -- **Use** ``force_ocr=False`` (the default) so clean pages skip OCR entirely. -- **Resize images before passing to OCR** — very high DPI scans can slow Tesseract down without improving accuracy. - ----- - - -Further Resources -------------------- - - -Sample code -~~~~~~~~~~~~~~~ - -- `Command line RAG Chatbot with PyMuPDF `_ -- `Example of a Browser Application using LangChain and PyMuPDF `_ -Blogs -~~~~~~~~~~~~~~ -- `RAG/LLM and PDF: Enhanced Text Extraction `_ -- `Creating a RAG Chatbot with ChatGPT and PyMuPDF `_ -- `Building a RAG Chatbot GUI with the ChatGPT API and PyMuPDF `_ -- `RAG/LLM and PDF: Conversion to Markdown Text with PyMuPDF `_ .. include:: ../footer.rst diff --git a/docs/rag.rst b/docs/rag.rst index 29082fa69..ba8ed6cfd 100644 --- a/docs/rag.rst +++ b/docs/rag.rst @@ -8,23 +8,56 @@ PyMuPDF, LLM & RAG Integrating |PyMuPDF| into your :title:`Large Language Model (LLM)` framework and overall :title:`RAG (Retrieval-Augmented Generation`) solution provides the fastest and most reliable way to deliver document data. -There are a few well known :title:`LLM` solutions which have their own interfaces with |PyMuPDF| - it is a fast growing area, so please let us know if you discover any more! +If you need to export to :title:`Markdown`, structured |JSON| or |TXT| formats, |PyMuPDF| provides the necessary tools to achieve this efficiently in the :class:`Document` object. -If you need to export to :title:`Markdown` or obtain a :title:`LlamaIndex` Document from a file: -.. raw:: html +Converting to |Markdown| +------------------------------------- + + +.. code-block:: python + + doc = pymupdf.open("input.pdf") + md = doc.to_markdown() + +See the API at: :meth:`Document.to_markdown`. + + +Converting to |JSON| +------------------------------------- - -

- +.. + .. raw:: html + + +

+ + Integration with :title:`LangChain` @@ -61,59 +94,60 @@ See `Building RAG from Scratch `_ is supported. +By using the :doc:`PyMuPDF4LLM module `, you can efficiently prepare your documents for chunking and subsequent processing with your :title:`LLM`. +Create chunked documents as follows: +.. code-block:: python -.. _rag_outputting_as_md: + import pymupdf4llm + + chunked_document = pymupdf4llm.to_chunks("input.pdf") -Outputting as :title:`Markdown` -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +Basic queries with chunked documents can be made as follows: -In order to export your document in :title:`Markdown` format you will need a separate helper. Package :doc:`pymupdf4llm/index` is a high-level wrapper of |PyMuPDF| functions which for each page outputs standard and table text in an integrated Markdown-formatted string across all document pages: +.. code-block:: python + len(chunked_document) # Count the number of Chunks + chunked_document[0], chunked_document[2:5] # a Chunk, a list of Chunks + chunked_document.index(chunked_document[3]) # 3 (Sequence protocol: iteration, slicing, index) + chunked_document.chunks # the same chunks as a tuple + chunked_document.text # all chunk text joined (lazy) + chunked_document.get("c0") # get chunk by any public id: c{n}, t{n}, f{n}, s{n}, p{page}.b{box} -.. code-block:: python +See :ref:`the full API documentation ` for PyMuPDF4LLM's chunking capabilities. - # convert the document to markdown - import pymupdf4llm - md_text = pymupdf4llm.to_markdown("input.pdf") - # Write the text to some file in UTF8-encoding - import pathlib - pathlib.Path("output.md").write_bytes(md_text.encode()) +.. _using_pymupdf_with_pymupdf_office: -For further information please refer to: :doc:`pymupdf4llm/index`. +Using with |PyMuPDF Office| +--------------------------- -How to use :title:`Markdown` output -~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ +For **Office** document support, |PyMuPDF| works seamlessly with |PyMuPDF Office|. Assuming you have :doc:`../pymupdf-office/index` installed you will be able to work with **Office** documents as expected: -Once you have your data in :title:`Markdown` format you are ready to chunk/split it and supply it to your :title:`LLM`, for example, if this is :title:`LangChain` then do the following: .. code-block:: python - import pymupdf4llm - from langchain.text_splitter import MarkdownTextSplitter + import pymupdf + import pymupdf.office + pymupdf.office.unlock() + md = pymupdf.office.to_markdown("sample.doc") - # Get the MD text - md_text = pymupdf4llm.to_markdown("input.pdf") # get markdown for all pages - splitter = MarkdownTextSplitter(chunk_size=40, chunk_overlap=0) +.. _pymupdf_and_layout: - splitter.create_documents([md_text]) +PyMuPDF & PyMuPDF Layout +----------------------------------- +By default |PyMuPDF| includes a `layout analysis module`_ to enhance output results. To disable this module you can do so by calling the :meth:`Document.use_layout` method. -For more see `5 Levels of Text Splitting `_ Related Blogs --------------------- - -To find out more about |PyMuPDF|, :title:`LLM` & :title:`RAG` check out our blogs for implementations & tutorials. - +----------------------------- Methodologies to Extract Text ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ @@ -123,7 +157,7 @@ Methodologies to Extract Text -Create a Chatbot to discuss your documents +Create a Chatbot to connect with your documents ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ - `Make a simple command line Chatbot `_ @@ -132,6 +166,10 @@ Create a Chatbot to discuss your documents +.. _PyMuPDF4LLM Document Loader: https://docs.langchain.com/oss/python/integrations/providers/pymupdf4llm/ + +.. _layout analysis module: https://pymupdf.io/use-cases/layout + diff --git a/docs/recipes-ocr.rst b/docs/recipes-ocr.rst deleted file mode 100644 index fca4455c6..000000000 --- a/docs/recipes-ocr.rst +++ /dev/null @@ -1,60 +0,0 @@ -.. include:: header.rst - -.. _RecipesOCR: - - -.. |toggleStart| raw:: html - -
- See code - -.. |toggleEnd| raw:: html - -
- -==================================== -OCR - Optical Character Recognition -==================================== - -|PyMuPDF| has integrated support for OCR (Optical Character Recognition). It is possible to use OCR for both, images (via the :ref:`Pixmap` class) and for document pages. - -The feature is currently based on Tesseract-OCR which must be installed as a separate application -- see the :ref:`installation_ocr`. - -How to OCR an Image --------------------- -A supported image must first be converted to a :ref:`Pixmap`. The Pixmap can then be saved to a 1-page PDF. This page will look like the original image with the same width and height. It will contain a layer of text as recognized by Tesseract. - -The PDF can be generated via one of the methods :meth:`Pixmap.pdfocr_save` or :meth:`Pixmap.pdfocr_tobytes`, as a file on disk or as a PDF in memory. - -The text can be extracted and searched with the usual text extraction and search methods (:meth:`Page.get_text`, :meth:`Page.search_for`, etc.). Please also note the following important facts and prerequisites: - -* When converting the image to a Pixmap, please confirm that the color space is RGB and alpha is `False` (no transparency). Convert the original Pixmap if necessary. -* All text is written as "hidden" with Tesseract's own `GlyphLessFont`, a mono-spaced font with metrics comparable to Courier. -* All text has the properties regular and black (i.e. no bold, no italic, no information about the original fonts). -* Tesseract does not recognize vector graphics (i.e. no drawings / line-art). - -This approach is also recommended to OCR a complete scanned PDF: - -* Render each page to a :ref:`Pixmap` with desired resolution -* Append the resulting 1-page PDF to the output PDF - -How to OCR a Document Page ----------------------------- -Any supported document page can be OCR-ed -- either the complete page or only the image areas on it. - -Because optical character recognition is about one thousand times slower than standard text extraction, we make sure to do OCR only once per page and store the result in a :ref:`TextPage`. Using this TextPage for all subsequent extractions and text searches will then happen with |PyMuPDF|'s usual top speed. - -To OCR a document page, follow this approach: - -1. Determine whether OCR is needed / beneficial at all. A number of criteria can be used for this decision, like: - - * page is completely covered by an image - * no text exists on the page - * thousands of small vector graphics (indicating *simulated* text) - -2. OCR the page and store result in a :ref:`TextPage` object using an instruction like `tp = page.get_textpage_ocr(...)`. - -3. Refer to the produced :ref:`TextPage` in all subsequent text extractions and searches via the `textpage=tp` parameter. - - -.. include:: footer.rst diff --git a/docs/recipes.rst b/docs/recipes.rst index afa81e863..39a0f9d77 100644 --- a/docs/recipes.rst +++ b/docs/recipes.rst @@ -24,12 +24,6 @@ ---- -.. toctree:: - - recipes-ocr.rst - ----- - .. toctree:: recipes-text.rst diff --git a/docs/resources.rst b/docs/resources.rst index 37be56dec..9e2b4f91e 100644 --- a/docs/resources.rst +++ b/docs/resources.rst @@ -5,11 +5,11 @@ Resources ============= -**PyMuPDF Pro** +**PyMuPDF Office** -------------------- -For **Office** file support `try PyMuPDF Pro `. +For **Office** file support `try PyMuPDF Office `. |