Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 10 additions & 39 deletions docs/about.rst
Original file line number Diff line number Diff line change
Expand Up @@ -60,29 +60,11 @@ The following table illustrates how |PyMuPDF| compares with other typical soluti

Therefore input files are mostly in a form that's useful for text extraction.

If faithful reproduction of layout is important, then consider using :ref:`PyMuPDF Pro <pymupdf-pro>`.
If faithful reproduction of layout is important, then consider using :ref:`PyMuPDF Office <pymupdf-office>`.


----

.. _About_PyMuPDF_Product_Suite:

PyMuPDF Product Suite
-----------------------------------------------

|PyMuPDF| is the standard version of the library, however there are a family of additional products each with different features and functionality.

**Additional products** in the |PyMuPDF| product suite are:

- |PyMuPDF Pro| adds support for Office document formats.
- |PyMuPDF4LLM| is optimized for large language model (LLM) applications, providing enhanced text extraction and processing capabilities.
It focuses on layout analysis and semantic understanding, ideal for document conversion and formatting tasks with enhanced results.

.. note::
All of the products above depend on the same core product - |PyMuPDF| and therefore have full access to all of its features.
These additional products can be seen as optional extras to the enhance the core |PyMuPDF| library.


.. _About_PyMuPDF_Products_Comparison:

PyMuPDF Products Comparison
Expand All @@ -91,51 +73,40 @@ PyMuPDF Products Comparison
The following table illustrates what features the products offer:

.. list-table:: PyMuPDF Products Comparison
:widths: 10 30 30 30
:widths: 10 45 45
:header-rows: 1

* -
- PyMuPDF
- PyMuPDF Pro
- PyMuPDF4LLM
- PyMuPDF Office
* - **Input Documents**
- `PDF`, `XPS`, `EPUB`, `CBZ`, `MOBI`, `FB2`, `SVG`, `TXT`, `MD`, Images (*standard document types*)
- *as PyMuPDF* and:
`DOC`/`DOCX`, `XLS`/`XLSX`, `PPT`/`PPTX`, `HWP`/`HWPX`
- *as PyMuPDF*
* - **Output Documents**
- Can convert any input document to `PDF`, `SVG` or Image
- Can convert any input document to `PDF`, Markdown (`MD`), `JSON`,`TXT`,`SVG`, Image
- *as PyMuPDF*
- *as PyMuPDF* and:
Markdown (`MD`), `JSON` or `TXT`
* - **Page Analysis**
- Basic page analysis to return document structure
- *as PyMuPDF*
- Advanced Page Analysis with trained data for enhanced results
* - **Data extraction**
- Basic data extraction with structured layout information and bounding box data
- *as PyMuPDF*
* - **Data extraction**
- Advanced data extraction including layout analysis with semantic understanding and enhanced bounding box data
* - **Table extraction**
- Basic table extraction as part of text extraction
- *as PyMuPDF*
* - **Table extraction**
- Advanced table extraction with cell structure, including support for merged cells and complex layouts
* - **Image extraction**
- Basic image extraction
- *as PyMuPDF*
* - **Image extraction**
- Advanced detection and rendering of image areas on page saving them to disk or embedding in MD output
* - **Vector extraction**
- Vector extraction and clustering
- *as PyMuPDF*
* - **Vector extraction**
- Superior detection of "picture" areas
- *as PyMuPDF*
* - **Popular RAG Integrations**
- Langchain, LlamaIndex
- *as PyMuPDF*
- *as PyMuPDF* and with some additional help methods for RAG workflows
* - **OCR**
- On-demand invocation of built-in Tesseract for text detection on pages or images
- Hybrid OCR based on page content analysis. OCR adapators for popular OCR engines available
- *as PyMuPDF*
- Automatic OCR based on page content analysis. OCR adapators for popular OCR engines available

----

Expand Down
13 changes: 9 additions & 4 deletions docs/converting-files.rst
Original file line number Diff line number Diff line change
Expand Up @@ -158,19 +158,24 @@ For example, assuming you have access to the source files for the "Comic Sans" f
PDF to Markdown
~~~~~~~~~~~~~~~~~

By utlilizing the :doc:`PyMuPDF4LLM API <pymupdf4llm/api>` we are able to convert PDF to a Markdown representation.
By utlilizing the :meth:`Document.to_markdown` method we are able to convert any of the following document types to a Markdown representation:

- PDFs
- Image documents (because they are internally converted to 1-page PDFs)
- Office documents opened using PyMuPDF Office (because they are converted to PDFs via :ref:`to_pdf() <pymupdf_office_api_to_pdf>` before processing them)

**Example**

.. code-block:: python

import pymupdf4llm
import pymupdf
import pathlib

md_text = pymupdf4llm.to_markdown("test.pdf")
doc = pymupdf.open("test.pdf")
md_text = doc.to_markdown()
print(md_text)

pathlib.Path("4llm-output.md").write_bytes(md_text.encode())
pathlib.Path("output.md").write_bytes(md_text.encode())


PDF to SVG
Expand Down
Loading
Loading