Skip to content
thinkoidPublic

About

A modern C++ PDF parser library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

117 Commits

Folders and files

Repository files navigation

yPDF

A toy PDF document parser, written to exercise document traversal and modification — and to keep an older, larger PDF codebase honest.

Where it came from

ypdf began in 2020 as a minimal recursive-descent parser: small, composable parsing functions all the way down, an AST of plain values, boost::iostreams filter chains for stream decoding. Then it stalled, for the longest time, on a missing piece — a CCITT fax codec for the stream filters — and relaunched once libccittfax existed to fill the gap.

Since then development has gone back and forth between this project and xpdf, continuously adjusted in whatever direction appeared most natural at each step. ypdf has grown into a small library with a document model: cross-reference tables and streams, object streams, a linear-scan fallback for broken files, and a central filter dispatcher.

How it is built

Three layers, each usable without the one above it.

The parser is recursive descent, and deliberately granular: eol, skip, comment, lit, name, numeric, string, array, dict, ref, obj, iobj, stream, xref, xrefstm, xreftbl, trailer — one header each, one job each, composed upward. Every parser is a function template over an iterator pair, so the same code reads a memory-mapped file or anything else that produces characters. A failed match consumes nothing: the contract is enforced by the API rather than by each caller remembering to guard.

The AST is plain values, not a class hierarchy. obj_t is a boost::variant over null, bool, int, double, reference, string, name, array, dictionary and stream; dict_t is a vector of pairs, array_t a vector of objects. There is nothing to subclass and nothing to own — you inspect it with ast::is<T> and ast::as<T>, or visit it.

The model (include/ypdf/model/doc.hh) is where a pile of objects becomes a document. load() walks the cross-reference chain — classic tables, xref streams, hybrid files with /XRefStm, /Prev traversal hardened against cycles and bad offsets — and when that chain is unusable it falls back to sweeping the bytes for indirect objects and synthesizing a table from what it finds. fetch() is the on-demand deref, memoized, and reaches into compressed object streams, which the sweep cannot.

Stream decoding is a boost::iostreams filter chain, resolved up front from /Filter and /DecodeParms rather than negotiated as the bytes flow.

There is no decryption. ypdf does not implement the standard security handler, or any other; an encrypted document's strings and streams are read as the ciphertext they are. /Encrypt survives in the trailer and is otherwise ignored. This is a deliberate boundary, not a gap waiting to be filled.

How it is tested

Three kinds of check, because unit tests alone would not have found most of what has been fixed here.

Unit tests — twenty-four of them, Boost.Test, one per parser or filter. They pin the fiddly cases: token boundaries after keywords, non-destructive literal matching, integer and floating-point overflow rejected rather than clamped, escaped names, indirect /Length, LZW early-change.

Cross-parser oracle (tools/xparser-oracle.sh) — dump every object canonically and diff against xpdf's pdfobjdump over a shared corpus, three-way when pdfalo, a Bison LALR parser for the same grammar, is also built. xpdf is the reference; ypdf is scored by how often its dump matches.

Decode oracle (tools/xstream-oracle.sh) — the filter-level counterpart. For every in-use stream with a fully lossless filter chain, both tools emit <num> <gen> <len> fnv1a:<hash>, and ypdf is scored on byte-identical digests. Lossy image codecs are skipped by both sides; they are not bit-identical across decoders and belong to a different kind of oracle.

The corpus is pinned in xpdf's checkout — the veraPDF PDF/A-1b pass set and the PDF Association's PDF 2.0 examples for well-formed files, the Isartor suite and JHOVE's Cabinet of Horrors for deliberately broken ones, plus mupdf's test files. tools/sweep.cc is the triage companion: one file per process, so a hang or a crash under timeout cannot take a whole run down.

Where it stands

Both oracles over the 485 files of xpdf-corpus at c40e189, run 2026-08-15. The mupdf set is part of the default corpus and is not counted here:

oracle xpdf, the reference identical differing ypdf declined
decode 2118 decoded streams 2105 — 99% 0 13
parser 485 files parsed 461 — 96% of the 478 both parsed 17 7

The decode oracle is the one to read closely: not one stream ypdf decodes disagrees with xpdf. The thirteen it scores against are streams ypdf declined to decode at all, in nine files, mostly Isartor's deliberate breakage — none of them encrypted. The parser's seventeen differences are the categories the script expects and does not treat as bugs — object streams, incremental updates, dropped nulls — and seven more files xpdf parses that ypdf does not.

Read the percentages knowing what the corpus does not contain: not one of the 485 files is encrypted, so nothing here exercises the boundary above, and nothing here should be taken as evidence about encrypted documents.

Reproduce with either script; both take seconds:

tools/xstream-oracle.sh ~/build/xpdf-corpus/{solid,malformed}

The two projects keep each other honest — the old code checks the new design, the new design questions the old code. Several fixes in the list above began as an oracle disagreement rather than a failing test, and at least one ended with ypdf being right: an Isartor case with a dictionary integer past 2³¹−1, which ypdf rejects and xpdf clamps.

Building

git submodule update --init
meson setup build && ninja -C build && meson test -C build

Needs a C++20 compiler, Boost, range-v3, libjpeg/libturbojpeg, jbig2dec and OpenJPEG; libccittfax comes in as the submodule. Stream decoding covers ASCIIHex, ASCII85, LZW, RunLength, Flate, CCITT fax, JBIG2, JPX and DCT, with the PNG and TIFF predictors, and honours the abbreviated filter names as well as the long ones.

Writing something with it

Here is examples/objects.cc — which lists every indirect object in a document — being written. Five sittings, each one compiled and run, starting from an empty main that only proves the library links:

A terminal session typing a program in five stages, compiling and running it after each, as the listing it prints gains a column at a time

Nothing there is filmed and nothing is faked. doc/demo/demo.toml scripts the session and doc/demo/mkcast.py runs the commands for real, so what appears on screen is what the code printed. The five stages are real files under doc/demo/steps/; doc/demo/steps/check.sh compiles each of them against an installed ypdf exactly as shown, and the last of them is examples/objects.cc — checked with cmp, so the two cannot drift apart.

See doc/demo/ANIMATION.md for how the animation is made.

About

A modern C++ PDF parser library

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages