Docs update for 2.0 - #5146
Docs update for 2.0#5146jamie-lemon wants to merge 13 commits into
Conversation
| OCR Adaptors | ||
| ------------ | ||
|
|
||
| By default, PyMuPDF uses **Tesseract** and **OpenCV** for image pre-processing. If you need a different OCR engine — for higher accuracy, language support, or cloud-based processing — you can plug in a custom adaptor. |
There was a problem hiding this comment.
We do not use OpenCV anymore since some. For image pre-processing, we use our own utility that is based on a LightGBM model (Gradient Boosting Decision Tree) that has been trained on 70 thousand documents of DocLayNet's ground truth data.
| ~~~~~~~~~~~~~~~~~ | ||
|
|
||
| By utlilizing the :doc:`PyMuPDF4LLM API <pymupdf4llm/api>` we are able to convert PDF to a Markdown representation. | ||
| By utlilizing the :meth:`Document.to_markdown` method we are able to convert PDF to a Markdown representation. |
There was a problem hiding this comment.
This conversion is not restricted to PDF:
By utlilizing the :meth:Document.to_markdown method we are able to convert any of the following document types to a Markdown representation:
- PDFs
- Image documents (because they are internally converted to 1-page PDFs)
- Office documents opened using PyMuPDF-Office (because they are converted to PDFs via
.to_pdf()before processing them).
| :meth:`Document.set_xml_metadata` PDF only: create or update document XML metadata | ||
| :meth:`Document.subset_fonts` PDF only: create font subsets | ||
| :meth:`Document.switch_layer` PDF only: activate OC configuration | ||
| :meth:`Document.to_json` PDF only: convert the document to JSON |
There was a problem hiding this comment.
I know I communicated it differently once. But the restriction to PDF is exclusively connected to the fact that we currently require PDF to write back OCRed text to pages that have been determined to need OCR.
We are investigating how to avoid this restriction.
But still, while this result is pending, the following document types are also eligible:
- Image documents (because we internally convert them to 1-page PDFs)
- Office documents handled with PyMuPDF-Office (because we internally deal with the
.to_pdf()output)
JorjMcKie
left a comment
There was a problem hiding this comment.
I have made a handful of comments...
No description provided.