Sample files for parser and importer testing

Well-formed files and deliberate edge cases for testing code that reads CSV, JSON, XML, PDF and more.

A parser is tested by what it is fed. Clean input shows that the common path works; awkward but valid input shows whether the parser follows the format or merely an optimistic idea of it.

The curated library contains files built around the features that hand-written parsers most often mishandle: quoted CSV fields, byte order marks, XML namespaces and CDATA, deeply nested JSON and multipart email.

Recommended files

FileSizeWhy this oneDownload
CSV with quoted fields, embedded commas and line breakssample-csv-quoted-fields.csv324 B324 bytesRFC 4180 quoting rulesDownload CSV
CSV with a UTF-8 byte order mark and non-ASCII textsample-csv-utf8-bom.csv193 B193 bytesByte order markDownload CSV
Sample product catalogue XML with namespaces and CDATAsample-product-catalog.xml1.07 KB1,100 bytesNamespaces, CDATA and entitiesDownload XML
Nested JSON with every value typesample-nested-config.json829 B829 bytesAll JSON types and deep nestingDownload JSON
10 MB CSV samplesample-csv-10mb.csv10 MB10,485,760 bytesVolume testDownload CSV
5 MB JSONL samplesample-jsonl-5mb.jsonl5 MB5,242,880 bytesStreaming testDownload JSONL

Checklist

  1. Parse the clean file

    Confirm record counts and a few known values. Each download page states how many records a file holds.

  2. Parse the edge cases

    Quoted fields, embedded line breaks, a byte order mark, different delimiters, non-ASCII text.

  3. Scale up

    Move to 1 MB, then 10 MB. Watch memory as well as time: reading everything at once is the usual fault.

  4. Truncate deliberately

    Cut a file in half and check that the parser reports an error instead of returning partial data as if it were complete.

Guides

Related use cases

Frequently asked questions

How do I make a truncated file?
Use head -c with a byte count, for example head -c 5000 sample-json-10kb.json > truncated.json. The result is invalid JSON by construction.