An Agentic Approach to Textual Data Extraction Using LLMs and LangGraph
From unstructured Wikipedia text to structured JSON: A step-by-step guide.
Imagine you are tasked with building a clean and structured dataset that includes all the cities and travel locations of the world. This dataset might be used to make travel suggestions or plan trips for users. You know this information exists in Wikipedia and is constantly updated thanks to community effort, but it’s all plain text or semi-structured in a way that does not fit your need.
This type of problem, meaning missing high-quality data that fits a specific format, is one of the primary challenges that data scientists and machine learning engineers face when solving business problems. This post is an attempt at turning large text corpora into usable, relevant and structured data.
The base idea of the solution that we will be exploring is taking advantage of the zero-shot abilities of LLMs to build a text-to-structured-dataset pipeline. Multiple approaches exist, but we will explore an agent-based approach to perform JSON-like structured information parsing. This involves parsing complex information containing multiple, potentially linked entity types into a structured format.
Our input data consists of Wikipedia pages, many of which are noisy or irrelevant to our need. Using an agent to perform this task allows the LLM to focus on specific sub-tasks, one a at a time, while also giving it enough freedom to guide the overall parsing process.
We will use LangGraph, a low-level, graph-shaped agent builder. This allows us to create agents with predefined actions, reducing errors. This contrasts with more high level tools like CrewAI where it is easy to get started but can be more error prone.
This post is divided into two parts:
- An introduction to LangGraph and its three main components.
- A detailed implementation of our agent, capable of parsing Wikipedia pages and returning structured information from relevant pages.
How LangGraph Works
LangGraph has three main visible components:
- State: A memory that persists during graph execution, storing information relevant to execution (LLM messages, classification results, tool calls, human inputs, API results). This state can be represented as a TypedDict or Pydantic object.
- Nodes: Functions performing actions (using LLMs, requesting user feedback, using APIs).
- Edges: Define execution flow — static (e.g., from node A to node B) or dynamic (from node A to a node depending on results).
These components create an agent as a directed graph. Execution starts at the start node, follows edges until reaching an end node, and then returns the final execution state.
Key advantages of this approach include:
- Logic is split into easily understood, self-contained components. The modular approach of LangGraph allows us to build and debug individual parts of the agent before even putting them all together.
- Logging and tracing is easy and helps understanding the execution flow.
- Unit testing of individual components and the entire agent is simple to do.
- The agent is more easily interpretable than approaches solely relying on LLM handling all the logic and decision-making.
- The entire workflow is easily visualized.
A limitation:
Defining all nodes and possibilities is initially time-consuming. However, this upfront work allows for focusing on improving individual node performance and the overall agent. This leads to a more reliable and predictable workflow.
We will begin with a toy example of LangGraph using dummy functions and nodes to illustrate the graph’s operation. This simple example will compile, run, and execute on an input, demonstrating LangGraph agent creation.
from typing import Annotated, Optional
from langchain_core.messages import AnyMessage
from langgraph.graph.message import add_messages
from langchain_core.messages import HumanMessage
from langgraph.graph import END, START, StateGraph
from pydantic import BaseModel
class OverallState(BaseModel):
result: Optional[int] = None
messages: Annotated[list[AnyMessage], add_messages]
def step_1(state: OverallState):
return {"messages": ["step_1 message"]}
def step_2(state: OverallState):
return {"messages": ["step_2 message"]}
def step_3(state: OverallState):
return {"messages": ["step_3 message"], "result": len(state.messages)}
# Define the function that determines whether to continue or not
def which_step_next(state: OverallState):
if len(state.messages) < 20:
return "step_1"
return END
# Define a new graph
workflow = StateGraph(OverallState)
# Define the two nodes we will cycle between
workflow.add_node("step_1", step_1)
workflow.add_node("step_2", step_2)
workflow.add_node("step_3", step_3)
# Set the entrypoint as `agent`
# This means that this node is the first one called
workflow.add_edge(START, "step_1")
workflow.add_edge("step_1", "step_2")
workflow.add_edge("step_2", "step_3")
# We now add a conditional edge
workflow.add_conditional_edges(
"step_3",
which_step_next,
)
app = workflow.compile()
# Use the Runnable
final_state = app.invoke(
{"messages": [HumanMessage(content="Init input")]},
config={"configurable": {"thread_id": 42}}
)
print(final_state["result"])This code first imports necessary libraries and defines OverallState, a Pydantic model holding results and a message list (using add_messages for appending new messages). Next, it defines three nodes (step_1, step_2, step_3), each adding a specific message. A dynamic edge, which_step_next, conditionally directs execution after step_3.
The workflow is built by adding these nodes and edges to a StateGraph. Edges connect START to step_1, step_1 to step_2, step_2 to step_3, and step_3 conditionally to either step_1 or END based on which_step_next. The workflow is compiled and invoked with an initial state. The result, determined by the conditional edge from step_3, is then printed. This example can be used as a foundation for understanding the basic features and design of langGraph.
Now that we know how it works, let’s start building the real Wikipedia Parser agent.
To build a competent agent, we need an LLM API. We’ll use the Nebius AI Studio API (not sponsored, offering $100 credit upon signup — enough for a few hundred million tokens). This API works like OpenAI’s and offers multiple LLMs like LLaMA 3.1 8B.
Since our goal is to parse Wikipedia into structured data, we need to define that structure. We’ll focus on two entity types:
- City: Name, description, information, latitude, and longitude.
- Attraction: Museums, water parks, etc., located in or near a city.
Here is how they can be defined with Pydantic:
from pydantic import BaseModel, Field
from typing import List, Optional
class Attraction(BaseModel):
"""Model for an attraction"""
name: str = Field(..., description="Name of the attraction")
description: str = Field(..., description="Description of the attraction")
city: str = Field(..., description="City where the attraction is located or is closest to")
countryw: str = Field(..., description="Country where the attraction is or is closest to")
activity_types: List[ActivityType] = Field(..., description="List of activity types and attractions available at the attraction")
tags: List[str] = Field(..., description="List of tags describing the attraction (e.g., accessible, sustainable, sunny, cheap, pricey)", min_length=1)
class City(BaseModel):
"""Model for a tourist city"""
name: str = Field(..., description="Name of the city")
description: str = Field(..., description="Few sentences description of the city, can be long if there is enough relevant information. Includes what the city is famous for and why people might visit it.")
country: str = Field(..., description="Country")
continent: Optional[str] = Field(None, description="Continent if applicable, otherwise leave empty")
location_lat_long: Optional[List[float]] = Field([], description="Geographic coordinates [latitude, longitude] of the location, [] if unknown")
climate_type: ClimateType = Field(..., description="Type of climate at the city")
class Config:
use_enum_values = True
extra = "forbid"Next, we define the agent’s state, also a Pydantic class. The state will store the Wikipedia page title, the page content, the page type (city, attraction, country, or other/unknown), lists of cities and attractions, and the page URL. Some of these attributes, like the page content, are inputs and some others, like the list of attractions, will be populated by the agent.
class OverallState(BaseModel):
page_title: str
page_content: str
page_type: PageType = PageType.unknown
cities: Optional[Cities] = None
attractions: Optional[Attractions] = None
url: Optional[str] = None
Now, we’ll define our nodes. The first node classifies the Wikipedia page:
def predict_page_type(state: OverallState) -> dict:
logger.info(f"Entering predict_page_type function. State: {state}")
page_summary = f"{state.page_title} {state.page_content[:400]}"
messages = [
SystemMessage(
content=f"You are a helpful assistant that outputs in JSON. Follow this schema {Page.model_json_schema()}"
),
HumanMessage(content="Give me an example Json of such output"),
AIMessage(content=page_example.model_dump_json()),
HumanMessage(content=f"What is the type of this page?\n {page_summary}"),
]
local_client = client_medium.with_structured_output(Page)
try:
page = local_client.invoke(messages)
logger.info(f"Page type prediction successful: {page}")
return {"page_type": page.page_type}
except Exception as e:
logger.error(f"Error in predict_page_type: {e}")
return {"page_type": None}This node’s output dictates the next steps. If it’s a city page, we extract city and attraction entities; if its an attraction we extract attractions; otherwise (other or unknown), we stop.
Next is the node to parse city data:
def parse_city(state: OverallState) -> dict:
logger.info(f"Entering parse_city function. State: {state}")
messages = [
SystemMessage(
content=f"You are a helpful assistant that outputs in JSON. Follow this schema {Cities.model_json_schema()}. Only answer with information from the context. Keep missing information empty."
),
HumanMessage(content="Give me an example Json of such output"),
AIMessage(content=Cities(cities=[city_example]).model_dump_json()),
HumanMessage(content=f"What are the cities mentioned in this page?\n {state.page_content}"),
]
local_client = client_medium.with_structured_output(Cities)
try:
cities = local_client.invoke(messages)
logger.info(f"City parsing successful: {cities}")
return {"cities": cities}
except Exception as e:
logger.error(f"Error in parse_city: {e}")
return {"cities": None}This line local_client = client_medium.with_structured_output(Cities) ensure that the output will follow the same schema as the Cities object.
Similarity, we have a node to parse attraction data.
Now, we can build the agent:
workflow = StateGraph(OverallState)
# Add nodes
workflow.add_node(NodeNames.predict_page_type.value, predict_page_type)
workflow.add_node(NodeNames.parse_city.value, parse_city)
workflow.add_node(NodeNames.parse_attraction.value, parse_attractions)
# Add edges
workflow.add_edge(START, NodeNames.predict_page_type.value)
workflow.add_conditional_edges(NodeNames.predict_page_type.value, parse_city_edge)
workflow.add_conditional_edges(NodeNames.predict_page_type.value, parse_attraction_edge)
workflow.add_edge(NodeNames.parse_attraction.value, END)
workflow.add_edge(NodeNames.parse_city.value, END)
app = workflow.compile()As we saw with the toy example, we add the nodes to the agent, then we add the edges, finally we compile.
We can now run the agent on an example Wikipedia page:
docs = WikipediaLoader(query="Paris", load_max_docs=1).load()
page_title = docs[0].metadata["title"]
page_content = docs[0].page_content
initial_state = OverallState(page_title=page_title, page_content=page_content)
final_state_dict = app.invoke(initial_state)This outputs (example for Paris):
Cities(
cities=[
City(
name='Paris',
description=(
'The capital and largest city of France, known for its museums, architectural landmarks, and sustainable transportation system.'
),
country='France',
continent='Europe',
location_lat_long=[
48.8567,
2.2945,
],
climate_type='temperate',
),
],
)Or, with the page “Tourism in Paris”, it finds all attractions:
Attractions(
attractions=[
Attraction(
name='Notre Dame',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.historical_sites: 'historical sites'>,
<ActivityType.cultural_events: 'cultural events'>,
],
tags=[
'landmark',
],
),
Attraction(
name='Disneyland Paris',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.amusement_parks: 'amusement parks'>,
<ActivityType.children_activities: 'children activities'>,
],
tags=[
'theme park',
],
),
Attraction(
name='Sacre Cœur',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.historical_sites: 'historical sites'>,
<ActivityType.cultural_events: 'cultural events'>,
],
tags=[
'landmark',
],
),
Attraction(
name='Versailles Palace',
description='',
city='Versailles',
country='France',
activity_types=[
<ActivityType.historical_sites: 'historical sites'>,
<ActivityType.cultural_events: 'cultural events'>,
],
tags=['palace'],
),
Attraction(
name='Louvre Museum',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.museums: 'museums'>,
<ActivityType.art_galleries: 'art galleries'>,
],
tags=['art'],
),
Attraction(
name='Eiffel Tower',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.historical_sites: 'historical sites'>,
<ActivityType.cultural_events: 'cultural events'>,
],
tags=[
'landmark',
],
),
Attraction(
name='Centre Pompidou',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.museums: 'museums'>,
<ActivityType.art_galleries: 'art galleries'>,
],
tags=['art'],
),
Attraction(
name="Musée d'Orsay",
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.museums: 'museums'>,
<ActivityType.art_galleries: 'art galleries'>,
],
tags=['art'],
),
Attraction(
name='Arc de Triomphe',
description='',
city='Paris',
country='France',
activity_types=[
<ActivityType.historical_sites: 'historical sites'>,
<ActivityType.cultural_events: 'cultural events'>,
],
tags=[
'landmark',
],
),
],
)This workflow is a first version and works well in most cases. It adheres to the dataset schema we defined and is able to weed out all the irrelevant pages and information. However, very long pages cause the Nebius API to time out, likely due to slow LLM response times. To address this, I’ll likely move some nodes to the faster Gemini platform, leveraging LangGraph’s flexibility to mix and match LLM models and providers. Also, while the workflow’s structure is complete, processing the entire Wikipedia is another challenge, to be discussed in a later blog post.
Summary
This blog post details building a LangChain-based graph agent to extract structured travel data (cities and attractions) from Wikipedia. The agent leverages the zero-shot capabilities of LLMs (using the Nebius AI Studio API) within a LangGraph framework. LangGraph’s modular design, consisting of interconnected nodes (functions) and edges (execution flow), allows for creating easily understandable, debuggable, and testable agents.
This example use-case and implementation will hopefully give you new ideas on how to solve your own complex data extraction problems using LLM graph agents.

