This guide explains how to build a Docker image of the Unstructured API with pre-loaded ML models (YOLO, PaddleOCR, Table Transformer) included in the image.
Pre-loading models in the Docker image provides several benefits:
- Faster Cold Starts: No need to download models on first request
- Offline Operation: Models available without internet connection
- Predictable Performance: No network delays for model downloads
- Production Ready: Consistent behavior across deployments
- Reduced Bandwidth: Models included in image, no runtime downloads
Building the image with preloaded models requires:
- Disk Space: Minimum 40GB free space (for build cache, models, and intermediate layers)
- RAM: At least 8GB RAM recommended
- Docker: Docker with buildx support enabled
- Network: Good internet connection for downloading dependencies and models (first time only)
- Time: 10-15 minutes for a complete build (depends on hardware and internet speed)
cd /path/to/unstructured-api
bash scripts/docker-build-with-models.shThis script will:
- Check if you have sufficient disk space
- Build the Docker image with all models preloaded
- Tag the image appropriately
- Show you how to run and push the image
cd /path/to/unstructured-api
DOCKER_REPOSITORY="your-registry" \
PIPELINE_PACKAGE="general" \
docker buildx build --load -f Dockerfile \
--build-arg PIPELINE_PACKAGE="general" \
--progress plain \
-t your-registry/unstructured-api:with-models \
.The pre-built image includes:
- Purpose: Object detection for high-resolution document processing
- Use Case: Identifying layout elements and sections in documents
- Size: ~250MB
- Environment Variable:
HI_RES_MODEL_NAME=yolox
- Purpose: Text extraction and OCR
- Languages: English (Spanish and other languages can be added)
- Size: ~500MB for English model
- Auto-downloads: Additional language models on first use if needed
- Model:
microsoft/table-transformer-structure-recognition - Purpose: Table structure recognition and extraction
- Size: ~100MB
- Framework: HuggingFace Transformers
- Purpose: Natural language processing utilities
- Size: ~100MB
- Included: Tokenizers and language patterns
The Dockerfile includes the following steps for model preloading:
- Base Setup: Installs system dependencies (Tesseract, LibreOffice, Poppler)
- Python Environment: Sets up Python 3.12 and project dependencies via
uv - NLTK Packages: Downloads NLTK language data
- Unstructured Models: Initializes general ML models
- Table Transformer: Pre-loads Microsoft's table structure recognition model
- YOLO Model: Pre-loads YOLO detection model
- PaddleOCR: Pre-loads PaddleOCR English language model
- Application: Copies application code and sets up entry point
After a successful build, you'll have an image with:
docker images | grep "unstructured-api"
The image size will be approximately 5-6GB due to the included models and dependencies.
docker run -p 8000:8000 your-registry/unstructured-api:with-modelscurl -X 'POST' \
'http://localhost:8000/general/v0/general' \
-H 'Content-Type: multipart/form-data' \
-F 'files=@document.pdf' \
-F 'strategy=hi_res' \
-F 'hi_res_model_name=yolox'docker run -p 8000:8000 \
-e API_ROOT_PATH="/api/v1" \
your-registry/unstructured-api:with-modelsdocker tag your-registry/unstructured-api:with-models username/unstructured-api:with-models
docker push username/unstructured-api:with-modelsdocker tag your-registry/unstructured-api:with-models quay.io/your-org/unstructured-api:with-models
docker login quay.io
docker push quay.io/your-org/unstructured-api:with-modelsSolution:
# Clean up Docker system
docker system prune -af --volumes
# Check available space
df -h /var/lib/docker
# Ensure at least 40GB free spaceSolution:
# Increase Docker daemon memory limit
# Edit Docker daemon configuration or restart with more memory
# Typically set to 4GB+ in Docker Desktop settingsSolution: Check that the build completed successfully:
# Test image
docker run --rm your-registry/unstructured-api:with-models \
python -c "from paddleocr import PaddleOCR; print('Models loaded!')"Causes and Solutions:
- Slow internet: Use a faster network or wait longer
- Slow disk: Use SSD storage for Docker
- Limited RAM: Allocate more memory to Docker daemon
- CPU: Slower CPUs take longer; wait or use faster hardware
Typical build times on modern hardware:
| Configuration | Time |
|---|---|
| On SSD with good internet | 10-15 minutes |
| On HDD with good internet | 20-30 minutes |
| On SSD with slow internet | 20-40 minutes |
| First build (no cache) | Add 5-10 minutes |
Models are cached in standard locations within the container:
~/.cache/ # YOLO and other models
~/.paddleocr/ # PaddleOCR models
~/.cache/huggingface/ # Table Transformer model
/usr/local/share/tessdata/ # Tesseract OCR data
Set these to customize model behavior:
# Use YOLO for high-resolution processing
HI_RES_MODEL_NAME=yolox
# Set number of OMP threads for parallel processing
OMP_NUM_THREADS=4
# Enable verbose logging during initialization
UNSTRUCTURED_LOG_LEVEL=DEBUGYou can modify the Dockerfile to include additional languages or models:
Edit the PaddleOCR loading section in Dockerfile:
RUN echo "Loading PaddleOCR models..." && \
${PYTHON} -c "
from paddleocr import PaddleOCR
import logging
logging.getLogger('ppocr').setLevel(logging.ERROR)
# Load multiple languages
for lang in ['en', 'es', 'fr', 'de']:
ocr = PaddleOCR(lang=lang, show_log=False)
print(f'PaddleOCR ({lang}) loaded')
"Add additional model loading steps after the existing model preloading sections.
For automated builds in CI/CD:
# GitHub Actions example
- name: Build Docker image with models
run: |
bash scripts/docker-build-with-models.sh
env:
DOCKER_REPOSITORY: ${{ secrets.REGISTRY }}/unstructured-api
PIPELINE_PACKAGE: general- First request: 30-60 seconds (includes model download)
- Subsequent requests: <2 seconds
- First request: <2 seconds (models already loaded)
- Subsequent requests: <2 seconds
- Improvement: 15-30x faster first response
For build issues:
- Check available disk space:
df -h /var/lib/docker - Review build logs for specific error messages
- Ensure Docker daemon has sufficient resources
- Try building on a different machine with more resources
- Check GitHub Issues for known issues
- PROXY_DEPLOYMENT.md - Proxy deployment configuration
- README.md - General API documentation
- docker-build-with-models.sh - Build script