Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

synthetic-receipts-ocr

The generator behind huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr — 32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR), each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription and structured KIE fields.

The dataset itself (4.3 GB, 30k train / 2k eval) lives on the Hub. This repo is the code that made it: content synthesis, rendering, degradation, validation, and the figure pipeline for the card.

Hero — five degraded receipts, one per locale

Why it exists

Public receipt datasets are small (SROIE ~1k), single-locale, and their labels come from human OCR — with human errors. Here the ground truth is exact by construction: boxes are captured during rendering, never re-OCR'd; every number on every receipt adds up (line totals × quantities, per-class contained VAT, US sales tax, tendered and change). The contract: a field is non-null iff its value is printed on the image — nothing invisible to hallucinate.

Clean render vs photo-degraded twin

What the generator does

stage file what happens
content generator/content.py Realistic merchants, addresses, tax ids (P.IVA / USt-IdNr / SIRET / VAT No), dates, payments; line items drawn from a 60k-title real product vocabulary (data/product_vocab.json, distilled from real Amazon listings) and 8k brands, abbreviated the way tills do it
render generator/render.py Thermal-printer aesthetic, 2 font families (OFL fonts in fonts/), word boxes recorded as it draws
degrade generator/degrade.py Homography onto a surface, uneven lighting, shadow bands, thermal fade, noise, blur, JPEG grunge — boxes mapped through the same 3×3 matrix
validate generator/validate.py Arithmetic re-checked per receipt: anything that doesn't add up is rejected
build generator/build.py Shards to parquet, splits with held-out fonts/vocab so eval isn't memorised

samples/ holds real outputs (clean, degraded, with-boxes); card_assets/gen_figures.py rebuilds the card figures from actual dataset rows.

Regenerate

python3 generator/build.py --split train --n 30000
python3 generator/build.py --split eval  --n 2000

Dependencies: Pillow, NumPy, PyArrow — no cloud, no GPU. The published shards were pushed with push_hf_images.py.

License

Code: Apache-2.0. Fonts: SIL OFL (IBM Plex Mono, JetBrains Mono, Space Mono). The generated dataset is Apache-2.0 on the Hub.

About

32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR), each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription and structured KIE fields.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages