Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
-
Updated
Sep 29, 2026 - Python
Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
A system for agentic LLM-powered data processing and ETL
The Semantic Intelligence Layer for Ontologies
在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档
WAS-NS Reborn; Tools for image processing, filters, masking, text, logic, numbers, latents, files, 3D scenes, and animation.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
Transform unstructured documents into validated, rich and queryable knowledge graphs.
电子发票整理与报销准备工具:从邮箱批量收集 PDF/OFD/XML 发票,OCR 识别、分类归档并生成 Excel 汇总;提供 Windows/macOS 桌面版与 DSH 插件。
Open-source batch OCR workbench — a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCR工作台,纯本地运行,免费平替ABBYY,适合书籍文档数字化。
Generic framework for historical document processing
TWIX is an open-source data extraction tool that reconstructs structured data from documents at scale, accurately and at low cost, by inferring the shared underlying visual template across documents
Turn PDFs into clean, structured Markdown
Python, LlamaIndex, LangChain, 15 Property Graph, 4 RDF , 10 Vector, OpenSearch, Elasticsearch, Alfresco, Nuxeo DBs. 14 data sources (10 auto-sync), KG auto-building, Ontologies, LLMs, Docling, LlamaParse, LiteParse, GraphRAG, RAG, Hybrid Search, AI Chat. TypeScript React, Vue, Angular frontends, REST, MCP Server. Options: Langflow, CocoIndex
Open-source toolkit for reliable RAG pipelines: convert PDFs to Markdown, clean documents, inspect chunks, compare chunking strategies, and enrich metadata for LLM applications.
Claude Code and Codex SKILLs for PDF, Excel, Word, and PowerPoint manipulation — extraction, forms, formulas, tracked changes, adapted from Anthropic skills.
Local official-document review, formatting repair, and compliant export desktop assistant.
MCP server that lets Claude Code and other AI agents read and search large PDFs, one file or a whole folder: agentic RAG with hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts.
Python CLI to replace visible PDF text with font-aware fallback and cleaner word spacing.
To associate your repository with the document-processing topic, visit your repo's landing page and select "manage topics."