A toy PDF document parser, written to exercise document traversal and modification — and to keep an older, larger PDF codebase honest.
ypdf began in 2020 as a minimal recursive-descent parser: small, composable
parsing functions all the way down, an AST of plain values, boost::iostreams
filter chains for stream decoding. Then it stalled, for the longest time, on a
missing piece — a CCITT fax codec for the stream filters — and relaunched once
libccittfax existed to fill the gap.
Since then development has gone back and forth between this project and xpdf, continuously adjusted in whatever direction appeared most natural at each step. ypdf has grown into a small library with a document model: cross-reference tables and streams, object streams, a linear-scan fallback for broken files, and a central filter dispatcher.
Three layers, each usable without the one above it.
The parser is recursive descent, and deliberately granular: eol, skip,
comment, lit, name, numeric, string, array, dict, ref, obj,
iobj, stream, xref, xrefstm, xreftbl, trailer — one header each,
one job each, composed upward. Every parser is a function template over an
iterator pair, so the same code reads a memory-mapped file or anything else
that produces characters. A failed match consumes nothing: the contract is
enforced by the API rather than by each caller remembering to guard.
The AST is plain values, not a class hierarchy. obj_t is a
boost::variant over null, bool, int, double, reference, string, name, array,
dictionary and stream; dict_t is a vector of pairs, array_t a vector of
objects. There is nothing to subclass and nothing to own — you inspect it with
ast::is<T> and ast::as<T>, or visit it.
The model (include/ypdf/model/doc.hh) is
where a pile of objects becomes a document. load() walks the cross-reference
chain — classic tables, xref streams, hybrid files with /XRefStm, /Prev
traversal hardened against cycles and bad offsets — and when that chain is
unusable it falls back to sweeping the bytes for indirect objects and
synthesizing a table from what it finds. fetch() is the on-demand deref,
memoized, and reaches into compressed object streams, which the sweep cannot.
Stream decoding is a boost::iostreams filter chain, resolved up front from
/Filter and /DecodeParms rather than negotiated as the bytes flow.
There is no decryption. ypdf does not implement the standard security
handler, or any other; an encrypted document's strings and streams are read as
the ciphertext they are. /Encrypt survives in the trailer and is otherwise
ignored. This is a deliberate boundary, not a gap waiting to be filled.
Three kinds of check, because unit tests alone would not have found most of what has been fixed here.
Unit tests — twenty-four of them, Boost.Test, one per parser or filter.
They pin the fiddly cases: token boundaries after keywords, non-destructive
literal matching, integer and floating-point overflow rejected rather than
clamped, escaped names, indirect /Length, LZW early-change.
Cross-parser oracle (tools/xparser-oracle.sh)
— dump every object canonically and diff against xpdf's pdfobjdump over a
shared corpus, three-way when pdfalo, a Bison LALR parser for the same
grammar, is also built. xpdf is the reference; ypdf is scored by how often its
dump matches.
Decode oracle (tools/xstream-oracle.sh) — the
filter-level counterpart. For every in-use stream with a fully lossless filter
chain, both tools emit <num> <gen> <len> fnv1a:<hash>, and ypdf is scored on
byte-identical digests. Lossy image codecs are skipped by both sides; they are
not bit-identical across decoders and belong to a different kind of oracle.
The corpus is pinned in xpdf's checkout — the veraPDF PDF/A-1b pass set and the
PDF Association's PDF 2.0 examples for well-formed files, the Isartor suite and
JHOVE's Cabinet of Horrors for deliberately broken ones, plus mupdf's test
files. tools/sweep.cc is the triage companion: one file per
process, so a hang or a crash under timeout cannot take a whole run down.
Both oracles over the 485 files of xpdf-corpus at c40e189, run 2026-08-15.
The mupdf set is part of the default corpus and is not counted here:
| oracle | xpdf, the reference | identical | differing | ypdf declined |
|---|---|---|---|---|
| decode | 2118 decoded streams | 2105 — 99% | 0 | 13 |
| parser | 485 files parsed | 461 — 96% of the 478 both parsed | 17 | 7 |
The decode oracle is the one to read closely: not one stream ypdf decodes disagrees with xpdf. The thirteen it scores against are streams ypdf declined to decode at all, in nine files, mostly Isartor's deliberate breakage — none of them encrypted. The parser's seventeen differences are the categories the script expects and does not treat as bugs — object streams, incremental updates, dropped nulls — and seven more files xpdf parses that ypdf does not.
Read the percentages knowing what the corpus does not contain: not one of the 485 files is encrypted, so nothing here exercises the boundary above, and nothing here should be taken as evidence about encrypted documents.
Reproduce with either script; both take seconds:
tools/xstream-oracle.sh ~/build/xpdf-corpus/{solid,malformed}The two projects keep each other honest — the old code checks the new design, the new design questions the old code. Several fixes in the list above began as an oracle disagreement rather than a failing test, and at least one ended with ypdf being right: an Isartor case with a dictionary integer past 2³¹−1, which ypdf rejects and xpdf clamps.
git submodule update --init
meson setup build && ninja -C build && meson test -C buildNeeds a C++20 compiler, Boost, range-v3, libjpeg/libturbojpeg, jbig2dec and OpenJPEG; libccittfax comes in as the submodule. Stream decoding covers ASCIIHex, ASCII85, LZW, RunLength, Flate, CCITT fax, JBIG2, JPX and DCT, with the PNG and TIFF predictors, and honours the abbreviated filter names as well as the long ones.
Here is examples/objects.cc — which lists every
indirect object in a document — being written. Five sittings, each one compiled
and run, starting from an empty main that only proves the library links:
Nothing there is filmed and nothing is faked. doc/demo/demo.toml scripts the
session and doc/demo/mkcast.py runs the commands for real, so what appears on
screen is what the code printed. The five stages are real files under
doc/demo/steps/; doc/demo/steps/check.sh compiles each of
them against an installed ypdf exactly as shown, and the last of them is
examples/objects.cc — checked with cmp, so the two cannot drift apart.
See doc/demo/ANIMATION.md for how the animation is made.