PDF, the Portable Document Format, was designed to solve one problem: making a document look the same on every screen and printer. It does this by describing pages precisely, with each character, line and image placed at fixed coordinates, rather than describing content and leaving layout to the reader as HTML does.
Adobe introduced the format in 1993 and handed it to the International Organization for Standardization in 2008. The current standard is ISO 32000-2, known as PDF 2.0.
What is inside a PDF
Open a simple PDF in a text editor and much of it is readable. The file has four parts.
- Header. The first line gives the version, for example
%PDF-1.4. A second line of binary characters usually follows, to signal that the file is not plain text. - Body. A list of numbered objects: pages, fonts, images and streams of drawing instructions. Objects refer to each other by number.
- Cross-reference table. An index giving the byte offset of every object, so that a reader can jump straight to page 200 without reading pages 1 to 199.
- Trailer. A pointer to the cross-reference table and to the root object, ending with the marker
%%EOF.
A reader opens a PDF from the end. It finds %%EOF, follows the pointer to the cross-reference table, and from there locates everything else. This has a practical consequence for testing: a truncated PDF is almost always unreadable, because the index at the end is missing. If an upload is cut short, a PDF will reveal it where a text file would not.
Recognising a PDF
Every PDF starts with the five bytes 25 50 44 46 2D, which spell %PDF-. You can check from a terminal:
head -c 8 sample-pdf-10kb.pdfThe output is %PDF-1.4 or similar. The registered MIME type is application/pdf, and the extension is .pdf.
A robust upload check looks for the signature at the start and the %%EOF marker near the end. Passing the file to a PDF library and asking for the page count is stronger still.
Why PDF sizes vary so much
There is no relationship between page count and file size. What matters is what the pages contain.
- Text is compact. A page of text is two or three kilobytes, and fonts, if embedded, add a fixed cost.
- Vector graphics such as charts and diagrams are also small.
- Images dominate. A scanned page is a full-page image, typically 100 KB to 1 MB depending on resolution and compression.
So a 300-page novel can be smaller than a three-page scanned contract. The sized samples on this site reflect this: the 10 KB PDF is text only, while the 10 MB PDF reaches its size with full-page images, as a real document of that size would.
Versions and variants
- PDF 1.4 to 1.7 are what most files in circulation use. Version 1.4 is a safe baseline that every reader supports.
- PDF 2.0 is the current ISO standard. It clarifies the specification and removes some obsolete features.
- PDF/A is a restricted profile for long-term archiving. It requires fonts to be embedded and forbids encryption and external dependencies.
- PDF/UA defines requirements for accessible, tagged documents.
- Linearised PDF, sometimes called fast web view, arranges the file so that the first page can be shown before the rest has downloaded.
What a PDF can contain
Beyond text and images, the format allows interactive forms, digital signatures, file attachments, JavaScript, audio and video, and encryption with passwords. This matters if your application accepts PDFs from the public: a PDF is a container, and what is inside may not be benign.
The sample PDFs here use none of those features. They contain text in standard fonts and, in the larger sizes, plain RGB images.
Testing an application that accepts PDFs
A reasonable test plan covers four areas.
Acceptance. Upload the single-page invoice and the multi-page report. Check that page counts, previews and thumbnails are right.
Size. Walk your limit with the sized PDF samples, from 5 KB to 10 MB. See the guide to testing file size limits.
Validation. Try a non-PDF renamed to .pdf, and a PDF renamed to something else. Then truncate a valid PDF and upload the result:
head -c 50000 sample-pdf-100kb.pdf > truncated.pdfThe truncated file still begins with %PDF-, so a signature-only check accepts it. Anything that parses the document will fail. Make sure that failure reaches the user as a clear message and not as a server error.
Processing. If your application extracts text, generates thumbnails or counts pages, run the 10 MB sample through it and watch memory and time. PDF libraries often load whole documents into memory.
Useful commands
With the Poppler utilities installed, you can inspect any PDF from the command line:
pdfinfo sample-pdf-1mb.pdf # version, page count, page size
pdftotext sample-pdf-1mb.pdf - # extract the textThe qpdf --check command validates the internal structure and reports damage.
Further reading
The full specification, ISO 32000-2, is available at no cost from the PDF Association. The media type is defined in RFC 8118.