Sitemap
Level Up Coding

Coding tutorials and news. The developer homepage gitconnected.com && skilled.dev && levelup.dev

Member-only story

Advanced Document Classification with LLMs

--

Press enter or click to view image in full size

This article will cover in detail how to classify documents using LLMs. This is part of Document Intelligence when you intend to divide and group documents. Before the rise of LLMs, this used to be accomplished (and still is) with AI models training in-house for certain use cases. Service such as Azure Document Intelligence gives you this feature, but they are not dynamic and will set you up for “Vendor lock-in”.

LLMs may not be the most efficient for this task, but they are agnostic and near-perfect for it.

Classification with LLMs — Vanilla

Press enter or click to view image in full size
Example of Classification with LLMs

When classifying documents, the process always involves extracting the content of the document (e.g., using pypdf) and adding it to the prompt with several possible classifications. It is strongly recommended to start with the instructor example, as it is straightforward to use.

from typing import Optional
from pydantic import BaseModel, Field

class ClassificationResponse(BaseModel):
name: str
confidence: Optional[int] = Field("From 1 to 10. 10 being the highest confidence. Always integer", ge=1, le=10)

--

--

Júlio Almeida
Júlio Almeida

Written by Júlio Almeida

Creator of ExtractThinker | Contractor - Focused on extraction in enterprise | Fintech and Legal. https://www.linkedin.com/in/j%C3%BAlio-almeida-21772a125