Skip to content

Docs update for 2.0 - #5146

Open
jamie-lemon wants to merge 13 commits into
mainfrom
docs-update-for-2.0
Open

jamie-lemon wants to merge 13 commits into
mainfrom
docs-update-for-2.0

Conversation

@jamie-lemon

Copy link
Copy Markdown
Collaborator

No description provided.

Comment thread docs/ocr/index.rst Outdated
OCR Adaptors
------------

By default, PyMuPDF uses **Tesseract** and **OpenCV** for image pre-processing. If you need a different OCR engine — for higher accuracy, language support, or cloud-based processing — you can plug in a custom adaptor.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We do not use OpenCV anymore since some. For image pre-processing, we use our own utility that is based on a LightGBM model (Gradient Boosting Decision Tree) that has been trained on 70 thousand documents of DocLayNet's ground truth data.

Comment thread docs/converting-files.rst Outdated
~~~~~~~~~~~~~~~~~

By utlilizing the :doc:`PyMuPDF4LLM API <pymupdf4llm/api>` we are able to convert PDF to a Markdown representation.
By utlilizing the :meth:`Document.to_markdown` method we are able to convert PDF to a Markdown representation.

@JorjMcKie JorjMcKie Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This conversion is not restricted to PDF:

By utlilizing the :meth:Document.to_markdown method we are able to convert any of the following document types to a Markdown representation:

  • PDFs
  • Image documents (because they are internally converted to 1-page PDFs)
  • Office documents opened using PyMuPDF-Office (because they are converted to PDFs via .to_pdf() before processing them).

Comment thread docs/document.rst Outdated
:meth:`Document.set_xml_metadata` PDF only: create or update document XML metadata
:meth:`Document.subset_fonts` PDF only: create font subsets
:meth:`Document.switch_layer` PDF only: activate OC configuration
:meth:`Document.to_json` PDF only: convert the document to JSON

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know I communicated it differently once. But the restriction to PDF is exclusively connected to the fact that we currently require PDF to write back OCRed text to pages that have been determined to need OCR.
We are investigating how to avoid this restriction.
But still, while this result is pending, the following document types are also eligible:

  • Image documents (because we internally convert them to 1-page PDFs)
  • Office documents handled with PyMuPDF-Office (because we internally deal with the .to_pdf() output)

@JorjMcKie JorjMcKie left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have made a handful of comments...

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants