Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.
-
Updated
Sep 17, 2026 - Java
Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.
Open-source document management platform leveraging AWS managed services. RESTful API for document storage, processing, full-text search, and metadata management. Multi-tenant serverless architecture with auto-scaling... deployed entirely in your AWS account.
A lightweight, framework-agnostic Java library for adding watermarks to various file types, including PDFs and videos
Docling simplifies document processing, parsing diverse formats — including advanced PDF understanding — and providing seamless integrations with the gen AI ecosystem
基于 Spring AI Alibaba 架构的企业级智能客服系统 | RAG + 混合检索(向量+全文)| PostgreSQL pgvector + tsvector | 支持多格式文档处理与意图识别
Open-source RAG backend for document ingestion and AI-powered chat with on-premise LLMs
A RAG (Retrieval-Augmented Generation) backend that lets you ask natural-language questions about uploaded PDF documents and get cited, grounded answers.
Document & OCR text/metadata extraction for DuckDB (Java, Apache Tika)
Multi-language SDKs (TypeScript, Python, Go, Java, C#, Ruby, Rust, Swift, PHP) for AI-powered document processing
AI-powered Product Backlog Generator built with Spring Boot and Spring AI
Microsserviço de assistentes de IA com Spring Boot e Spring AI baseado em RAG. Integra OpenAI e pgvector para ingestão de documentos, busca vetorial e geração de respostas contextualizadas por domínio.
Batch digitization tool for handwritten historical documents. Draw a template once — the system crops fields, runs OCR, and applies LLM correction
LambDB, is a lightweight, command-line driven NoSQL database prototype built entirely in Java.
Event-driven file upload & search demo on OpenShift
State-machine driven Document Processing Orchestration Service built with Spring Boot. Orchestrates OCR, Document Classification, and Named Entity Recognition pipelines with async processing, retry support, logging, and Dockerized deployment.
Event-driven document processing platform powered by Change Data Capture: PostgreSQL, Debezium, Apache Kafka, Spring Boot and Angular.
This project is a document processing tool that converts HWP, PDF, DOCX, and other formats into HTML.
Distributed document-processing platform for asynchronous PDF generation using Spring Boot, RabbitMQ, PostgreSQL, MinIO, Docker, and scalable workers.
AI-powered enterprise knowledge assistant — upload documents, ask questions, and get context-aware answers using RAG, Spring Boot 3.5, Spring AI, and pgvector.
Enterprise Java document processing system with AI extraction, classification and workflow routing
To associate your repository with the document-processing topic, visit your repo's landing page and select "manage topics."