From a Scraping Script to a Modular Web Crawler with Crawlee
Part 1: Separating page discovery, data extraction, logging, and storage with Crawlee for Python
A small web scraping script can be enough for a simple, one-time task.
You send a request, parse the HTML, extract a few values, and save the result. However, as the project grows, the script may need to handle different page types, discover new URLs, store structured data, and remain understandable when its behavior changes.
At that point, code organization becomes as important as data extraction.
In the first part of this series, I built a small product crawler using Crawlee for Python. The goal was not to create a large production system immediately, but to establish a modular foundation that could be extended in the next parts.
What This Crawler Does
The crawler starts from a product category page and performs a simple sequence of operations:
- It loads the initial category URL.
- It finds links that point to individual product pages.
- It labels those requests as
PRODUCT. - It sends each labeled request to a dedicated handler.
- It extracts the product page URL and title.
- It stores the extracted records in a Crawlee Dataset.
- It writes crawl information to both the terminal and a log file.
The extracted data is intentionally simple in this first part. The main focus is the structure of the crawler rather than the number of collected fields.
Why I Used Crawlee
It is possible to implement the same basic behavior with requests and BeautifulSoup alone. I chose Crawlee because it provides useful crawling components through a consistent interface, including request handling, link enqueueing, routing, storage, and crawler configuration.
For this part, I used BeautifulSoupCrawler because the required data can be obtained from the server-rendered HTML and does not require browser automation.
Later parts of the series use other Crawlee components when the target pages require more advanced navigation or JavaScript execution.
Separating Configuration from Routing
The project is divided into two main Python files:
main.pyconfigures and starts the crawler.routes.pydefines how different page types are processed.
This separation keeps the entry point small and moves page-specific behavior into dedicated handlers.
The crawler is initialized in main.py:
crawler = BeautifulSoupCrawler(
request_handler=router,
max_requests_per_crawl=50,
)The maximum request count keeps this example bounded while testing. The same file also defines the starting URL, configures Loguru, starts the crawler, and reports when the crawl is complete.
Discovering Product Pages
The default handler processes the initial category page.
Its job is not to extract the final product data. Instead, it discovers product links and adds them to the crawling queue:
@router.default_handler
async def default_handler(context):
logger.info(f"Processing URL: {context.request.url}")
await context.enqueue_links(
selector="a[href*='/product-detail/']",
label="PRODUCT",
unique=True,
)The CSS selector limits discovery to links containing /product-detail/.
Each discovered request receives the PRODUCT label. This label determines which handler will process that URL later.
Processing Product Pages
Product pages are handled separately:
@router.handler("PRODUCT")
async def product_handler(context):
logger.info(
f"Processing PRODUCT URL: {context.request.url}"
)
data = {
"url": context.request.url,
"title": (
context.soup.title.string.strip()
if context.soup.title
else "Unknown"
),
}
await context.push_data(data)This handler extracts two fields:
- The current product URL
- The HTML page title
The resulting dictionary is then passed to context.push_data(), which stores it in the crawler’s default Dataset.
The extraction logic is basic, but the important point is that discovery and extraction no longer share the same handler.
Why the Router Helps
Without routing, a crawler can quickly turn into a long function containing conditions for category pages, product pages, pagination, and other URL types.
The Router allows each request type to have its own handler.
In this example:
- Unlabeled requests use the default handler.
- Requests labeled
PRODUCTuse the product handler.
This makes the execution flow easier to follow and gives the project a clear place for adding new page types later.
For example, a future version could introduce separate labels for pagination pages, search results, or API endpoints without putting all the logic into one function.
Logging the Crawl
The project uses Loguru to write messages to both the terminal and a file located at:
logs/crawler.logThe logs record which URLs are being processed and when the crawl starts or finishes.
This is useful during development because it makes the crawler’s behavior visible without adding print statements throughout the code.
What This First Part Does Not Cover
This first implementation is deliberately limited.
It does not yet extract detailed product attributes, follow recursive pagination, control a browser, or collect variant-level price and stock information.
Those features are introduced gradually in the later parts of the series.
Keeping Part 1 small makes it easier to understand the basic responsibilities of the crawler before adding more complicated behavior.
Project Structure
The Part 1 directory contains:
Part-1-Foundation/
├── main.py
├── routes.py
└── README.mdThis is a small structure, but each file has a clear responsibility:
main.pyhandles configuration and execution.routes.pyhandles page-specific crawling logic.README.mddocuments the example.
Source Code
The complete source code for this part is available on GitHub:
https://github.com/Saeiii-d/crawlee-python-masterclass/tree/main/Part-1-Foundation
The main repository also contains the remaining parts of the crawling series:
What Comes Next
Part 1 establishes the routing and storage foundation.
In Part 2, I extend the crawler to work with pagination and extract more detailed product information while keeping the page-discovery and data-extraction responsibilities separated.
Connect
You can find my projects and future updates here:
- GitHub: https://github.com/Saeiii-d
- LinkedIn: https://www.linkedin.com/in/saeidkhazaei/
- Medium: https://medium.com/@saeiiid.khazaei/