Member-only story
Advanced Document Classification with LLMs
This article will cover in detail how to classify documents using LLMs. This is part of Document Intelligence when you intend to divide and group documents. Before the rise of LLMs, this used to be accomplished (and still is) with AI models training in-house for certain use cases. Service such as Azure Document Intelligence gives you this feature, but they are not dynamic and will set you up for “Vendor lock-in”.
LLMs may not be the most efficient for this task, but they are agnostic and near-perfect for it.
Classification with LLMs — Vanilla
When classifying documents, the process always involves extracting the content of the document (e.g., using pypdf) and adding it to the prompt with several possible classifications. It is strongly recommended to start with the instructor example, as it is straightforward to use.
from typing import Optional
from pydantic import BaseModel, Field
class ClassificationResponse(BaseModel):
name: str
confidence: Optional[int] = Field("From 1 to 10. 10 being the highest confidence. Always integer", ge=1, le=10)
