Jupyter notebooks testing different OCR models for document parsing (Dolphin, MonkeyOCR, Marker, Nanonets, ...)
-
Updated
Aug 2, 2026 - Jupyter Notebook
Jupyter notebooks testing different OCR models for document parsing (Dolphin, MonkeyOCR, Marker, Nanonets, ...)
collection of notebooks for finetuning donut model for various visual document understanding tasks, using huggingface Trainer.
Free, open-source Google Colab notebook that converts PDF & image files (JPG, PNG, BMP, TIFF, WEBP) into clean Markdown using Baidu's Unlimited-OCR (MIT license, 93%+ OmniDocBench). No API key, no cost — runs on free Colab GPU. PDF to Markdown, OCR, document parsing, image to text converter.
To associate your repository with the document-parsing topic, visit your repo's landing page and select "manage topics."