The generator behind huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr — 32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR), each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription and structured KIE fields.
The dataset itself (4.3 GB, 30k train / 2k eval) lives on the Hub. This repo is the code that made it: content synthesis, rendering, degradation, validation, and the figure pipeline for the card.
Public receipt datasets are small (SROIE ~1k), single-locale, and their labels come from human OCR — with human errors. Here the ground truth is exact by construction: boxes are captured during rendering, never re-OCR'd; every number on every receipt adds up (line totals × quantities, per-class contained VAT, US sales tax, tendered and change). The contract: a field is non-null iff its value is printed on the image — nothing invisible to hallucinate.
| stage | file | what happens |
|---|---|---|
| content | generator/content.py |
Realistic merchants, addresses, tax ids (P.IVA / USt-IdNr / SIRET / VAT No), dates, payments; line items drawn from a 60k-title real product vocabulary (data/product_vocab.json, distilled from real Amazon listings) and 8k brands, abbreviated the way tills do it |
| render | generator/render.py |
Thermal-printer aesthetic, 2 font families (OFL fonts in fonts/), word boxes recorded as it draws |
| degrade | generator/degrade.py |
Homography onto a surface, uneven lighting, shadow bands, thermal fade, noise, blur, JPEG grunge — boxes mapped through the same 3×3 matrix |
| validate | generator/validate.py |
Arithmetic re-checked per receipt: anything that doesn't add up is rejected |
| build | generator/build.py |
Shards to parquet, splits with held-out fonts/vocab so eval isn't memorised |
samples/ holds real outputs (clean, degraded, with-boxes); card_assets/gen_figures.py rebuilds the card figures from actual dataset rows.
python3 generator/build.py --split train --n 30000
python3 generator/build.py --split eval --n 2000Dependencies: Pillow, NumPy, PyArrow — no cloud, no GPU. The published shards were pushed with push_hf_images.py.
Code: Apache-2.0. Fonts: SIL OFL (IBM Plex Mono, JetBrains Mono, Space Mono). The generated dataset is Apache-2.0 on the Hub.

