Skip to content
Draft
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
Next Next commit
docs: restructure documentation to align with SDK and Client docs
- Add a Concepts section holding the component and abstraction pages moved
  from Guides, plus a new Logging page.
- Merge the Examples section into Concepts and Guides: each example is folded
  into its related page or combined into a new guide (Crawling links,
  Stopping and resuming crawlers), and redundant duplicates are dropped.
- Number the section directories and files to drive sidebar order, matching
  the apify-sdk-python and apify-client-python docs conventions.
- Enable @docusaurus/plugin-client-redirects with version-aware redirects for
  every moved URL.
- Update the navbar, footer, homepage, README, and lint config paths, and add
  frontmatter descriptions to pages that lacked them.
  • Loading branch information
vdusek committed Aug 24, 2026
commit 857bca19975d60f40b357dfacb4c9355d094b363
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,7 +167,7 @@ if __name__ == '__main__':

### More examples

Explore our [Examples](https://crawlee.dev/python/docs/examples) page in the Crawlee documentation for a wide range of additional use cases and demonstrations.
Explore the [Guides](https://crawlee.dev/python/docs/guides) section of the Crawlee documentation for a wide range of additional use cases and demonstrations.

## Features

Expand Down
7 changes: 4 additions & 3 deletions docs/quick-start/index.mdx → docs/01_quick-start/index.mdx
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: quick-start
title: Quick start
description: Build and run your first Crawlee crawler in a few minutes, and pick the crawler type that fits your project.
---

import ApiLink from '@site/src/components/ApiLink';
Expand All @@ -15,7 +16,7 @@ import PlaywrightCrawlerExample from '!!raw-loader!roa-loader!./code_examples/pl

import PlaywrightCrawlerHeadfulExample from '!!raw-loader!./code_examples/playwright_crawler_headful_example.py';

This short tutorial will help you start scraping with Crawlee in just a minute or two. For an in-depth understanding of how Crawlee works, check out the [Introduction](../introduction/index.mdx) section, which provides a comprehensive step-by-step guide to creating your first scraper.
This short tutorial will help you start scraping with Crawlee in just a minute or two. For an in-depth understanding of how Crawlee works, check out the [Introduction](../02_introduction/index.mdx) section, which provides a comprehensive step-by-step guide to creating your first scraper.

## Choose your crawler

Expand Down Expand Up @@ -61,7 +62,7 @@ If you plan to use the <ApiLink to="class/PlaywrightCrawler">`PlaywrightCrawler`
playwright install
```

For detailed installation instructions, see the [Setting up](../introduction/01_setting_up.mdx) documentation page.
For detailed installation instructions, see the [Setting up](../02_introduction/01_setting_up.mdx) documentation page.

## Crawling

Expand Down Expand Up @@ -128,6 +129,6 @@ If you want to change the storage directory, you can set the `CRAWLEE_STORAGE_DI

## Examples and further reading

For more examples showcasing various features of Crawlee, visit the [Examples](/docs/examples) section of the documentation. To get a deeper understanding of Crawlee and its components, read the step-by-step [Introduction](../introduction/index.mdx) guide.
For more examples showcasing various features of Crawlee, visit the [Guides](../guides) section of the documentation. To get a deeper understanding of Crawlee and its components, read the step-by-step [Introduction](../02_introduction/index.mdx) guide.

[//]: # (TODO: add related links once they are ready)
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: setting-up
title: Setting up
description: How to install Crawlee, set up your environment, and bootstrap a new project.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: first-crawler
title: First crawler
description: Build your first Crawlee crawler - set up a request queue, write a request handler, and crawl your first page.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: adding-more-urls
title: Adding more URLs
description: Grow the crawl by discovering and enqueuing new links, with filtering and deduplication handled for you.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: real-world-project
title: Real-world project
description: Plan a real scraping project - choose the data to collect and analyze the target website before writing code.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: crawling
title: Crawling
description: Crawl the example Warehouse store - visit the category listings and enqueue the product detail pages.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: scraping
title: Scraping
description: Extract structured data such as titles, prices, and stock information from the product detail pages.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: saving-data
title: Saving data
description: Persist the scraped results into a dataset and find them on disk.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down Expand Up @@ -88,19 +89,10 @@ A helper <ApiLink to="class/PushDataFunction">`context.push_data`</ApiLink> save

:::info Automatic dataset initialization

Each time you start Crawlee a default <ApiLink to="class/Dataset">`Dataset`</ApiLink> is automatically created, so there's no need to initialize it or create an instance first. You can create as many datasets as you want and even give them names. For more details see the <ApiLink to="class/Dataset#open">`Dataset.open`</ApiLink> function.
Each time you start Crawlee a default <ApiLink to="class/Dataset">`Dataset`</ApiLink> is automatically created, so there's no need to initialize it or create an instance first. You can create as many datasets as you want and even give them names. For more details see the [Storages](../concepts/storages#dataset) page and the <ApiLink to="class/Dataset#open">`Dataset.open`</ApiLink> function.

:::

{/* TODO: mention result storage guide once it's done

:::info Automatic dataset initialization

Each time you start Crawlee a default <ApiLink to="class/Dataset">`Dataset`</ApiLink> is automatically created, so there's no need to initialize it or create an instance first. You can create as many datasets as you want and even give them names. For more details see the [Result storage guide](../guides/result-storage#dataset) and the `Dataset.open()` function.

:::
*/}

## Finding saved data

Unless you changed the configuration that Crawlee uses locally, which would suggest that you knew what you were doing, and you didn't need this tutorial anyway, you'll find your data in the storage directory that Crawlee creates in the working directory of the running script:
Expand All @@ -111,16 +103,12 @@ Unless you changed the configuration that Crawlee uses locally, which would sugg

The above folder will hold all your saved data in numbered files, as they were pushed into the dataset. Each file represents one invocation of <ApiLink to="class/Dataset#push_data">`Dataset.push_data`</ApiLink> or one table row.

{/* TODO: add mention of "Result storage guide" once it's ready:

:::tip Single file data storage options

If you would like to store your data in a single big file, instead of many small ones, see the [Result storage guide](../guides/result-storage#key-value-store) for Key-value stores.
If you would like to store your data in a single big file, instead of many small ones, see how to [export the whole dataset](../concepts/storages#exporting-the-dataset) to JSON or CSV, or use a [key-value store](../concepts/storages#key-value-store).

:::

*/}

## Next steps

Next, you'll see some improvements that you can add to your crawler code that will make it more readable and maintainable in the long run.
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: refactoring
title: Refactoring
description: Clean up the crawler code with a router and separate handlers to keep the project maintainable.
---

import ApiLink from '@site/src/components/ApiLink';
Expand All @@ -21,7 +22,7 @@ You might be wondering about the **anti-blocking, bot-protection avoiding stealt

However, the default configuration, while powerful, may not cover every scenario.

If you want to learn more, browse the [Avoid getting blocked](../guides/avoid-blocking), [Proxy management](../guides/proxy-management) and [Session management](../guides/session-management) guides.
If you want to learn more, browse the [Avoid getting blocked](../guides/avoid-blocking), [Proxy management](../concepts/proxy-management) and [Session management](../concepts/session-management) guides.
*/}

To promote good coding practices, let's look at how you can use a <ApiLink to="class/Router">`Router`</ApiLink> class to better structure your crawler code.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
id: deployment
title: Running your crawler in the Cloud
sidebar_label: Running in the Cloud
description: Deploying Crawlee-python projects to the Apify platform
description: Deploy your Crawlee for Python project to the Apify platform and run it in the cloud.
---

import CodeBlock from '@theme/CodeBlock';
Expand Down
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
---
id: introduction
title: Introduction
description: A step-by-step tutorial that takes you from your first crawler to a production-ready scraper for a real website.
---

import ApiLink from '@site/src/components/ApiLink';
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ description: An overview of the core components of the Crawlee library and its a

import ApiLink from '@site/src/components/ApiLink';

Crawlee is a modern and modular web scraping framework. It is designed for both HTTP-only and browser-based scraping. In this guide, we will provide a high-level overview of its architecture and the main components that make up the system.
Crawlee is a modern and modular web scraping framework. It is designed for both HTTP-only and browser-based scraping. This page provides a high-level overview of its architecture and the main components that make up the system.

## Crawler

Expand Down Expand Up @@ -93,7 +93,7 @@ You can learn more about HTTP crawlers in the [HTTP crawlers guide](./http-crawl
Browser crawlers use a real browser to render pages, enabling scraping of sites that require JavaScript. They manage browser instances, pages, and context lifecycles. Crawlee provides two browser crawlers:

- <ApiLink to="class/PlaywrightCrawler">`PlaywrightCrawler`</ApiLink> utilizes the [Playwright](https://playwright.dev/) library and provides a high-level API for controlling and navigating browsers. You can learn more about it in the [Playwright crawler guide](./playwright-crawler).
- <ApiLink to="class/StagehandCrawler">`StagehandCrawler`</ApiLink> extends `PlaywrightCrawler` with AI-powered browser automation via [Stagehand](https://github.com/browserbase/stagehand). It adds natural-language methods (`act`, `extract`, `observe`, `execute`) directly on the page object. You can learn more about it in the [Stagehand crawler guide](./stagehand-crawler).
- <ApiLink to="class/StagehandCrawler">`StagehandCrawler`</ApiLink> extends `PlaywrightCrawler` with AI-powered browser automation via [Stagehand](https://github.com/browserbase/stagehand). It adds natural-language methods (`act`, `extract`, `observe`, `execute`) directly on the page object. You can learn more about it in the [Stagehand crawler guide](../guides/stagehand-crawler).

### Adaptive crawler

Expand Down Expand Up @@ -412,7 +412,7 @@ The core component of session management in Crawlee is <ApiLink to="class/Sessio

:::info

You can learn more about fingerprints and how to avoid getting blocked in the [Avoid blocking guide](./avoid-blocking).
You can learn more about fingerprints and how to avoid getting blocked in the [Avoid blocking guide](../guides/avoid-blocking).

:::

Expand Down Expand Up @@ -441,8 +441,3 @@ The system includes error tracking through the <ApiLink to="class/ErrorTracker">

Statistics are logged at configurable intervals in both table and inline formats, with final summary data returned from the `crawler.run` method available through <ApiLink to="class/FinalStatistics">`FinalStatistics`</ApiLink>.

## Conclusion

In this guide, we provided a high-level overview of the core components of the Crawlee library and its architecture. We covered the main components like crawlers, crawling contexts, storages, request routers, service locator, request loaders, event manager, session management, and statistics. Check out other guides, the [API reference](https://crawlee.dev/python/api), and [Examples](../examples) for more details on how to use these components in your own projects.

If you have questions or need assistance, feel free to reach out on our [GitHub](https://github.com/apify/crawlee-python) or join our [Discord community](https://discord.com/invite/jyEM2PRvMU). Happy scraping!
Original file line number Diff line number Diff line change
Expand Up @@ -20,13 +20,16 @@ import LexborParser from '!!raw-loader!roa-loader!./code_examples/http_crawlers/
import PyqueryParser from '!!raw-loader!roa-loader!./code_examples/http_crawlers/pyquery_parser.py';
import ScraplingParser from '!!raw-loader!roa-loader!./code_examples/http_crawlers/scrapling_parser.py';

import FileDownloadExample from '!!raw-loader!roa-loader!./code_examples/http_crawlers/file_download.py';
import FileDownloadStreamExample from '!!raw-loader!./code_examples/http_crawlers/file_download_stream.py';

import SelectolaxParserSource from '!!raw-loader!./code_examples/http_crawlers/selectolax_parser.py';
import SelectolaxContextSource from '!!raw-loader!./code_examples/http_crawlers/selectolax_context.py';
import SelectolaxCrawlerSource from '!!raw-loader!./code_examples/http_crawlers/selectolax_crawler.py';
import SelectolaxCrawlerRunSource from '!!raw-loader!./code_examples/http_crawlers/selectolax_crawler_run.py';
import AdaptiveCrawlerRunSource from '!!raw-loader!./code_examples/http_crawlers/selectolax_adaptive_run.py';

HTTP crawlers are ideal for extracting data from server-rendered websites that don't require JavaScript execution. These crawlers make requests via HTTP clients to fetch HTML content and then parse it using various parsing libraries. For client-side rendered content, where you need to execute JavaScript consider using [Playwright crawler](https://crawlee.dev/python/docs/guides/playwright-crawler) instead.
HTTP crawlers are ideal for extracting data from server-rendered websites that don't require JavaScript execution. These crawlers make requests via HTTP clients to fetch HTML content and then parse it using various parsing libraries. For client-side rendered content, where you need to execute JavaScript, consider using the [Playwright crawler](./playwright-crawler) instead.

## Overview

Expand Down Expand Up @@ -136,7 +139,19 @@ The following examples demonstrate how to integrate with several popular parsing

## FileDownloadCrawler

The <ApiLink to="class/FileDownloadCrawler">`FileDownloadCrawler`</ApiLink> downloads files instead of scraping pages. It accepts any content type without parsing and gives the request handler direct access to the response body. By default the whole file is buffered in memory. For large files, construct the crawler with `stream=True` and consume the body in chunks via <ApiLink to="class/HttpResponse#read_stream">`read_stream()`</ApiLink>. For usage, see the [Download files](../examples/file-download) example.
The <ApiLink to="class/FileDownloadCrawler">`FileDownloadCrawler`</ApiLink> downloads files such as PDFs, images or videos instead of scraping pages. It accepts any content type without parsing and gives the request handler direct access to the response body. Each downloaded file can be saved to the <ApiLink to="class/KeyValueStore">`KeyValueStore`</ApiLink> together with the content type reported by the server.

<RunnableCodeBlock className="language-python" language="python">
{FileDownloadExample}
</RunnableCodeBlock>

### Streaming large files

By default the whole file is buffered in memory, which doesn't scale to large downloads. Construct the crawler with `stream=True` and the request handler receives a response whose body hasn't been read yet. Consume it in chunks with <ApiLink to="class/HttpResponse#read_stream">`read_stream()`</ApiLink> and write each chunk to disk as it arrives.

<CodeBlock className="language-python" language="python">
{FileDownloadStreamExample}
</CodeBlock>

## Custom HTTP crawler

Expand Down Expand Up @@ -192,9 +207,3 @@ The custom crawler works like any built-in crawler. Request handlers receive you
</CodeBlock>
</TabItem>
</Tabs>

## Conclusion

This guide provided a comprehensive overview of HTTP crawlers in Crawlee. You learned about the three main crawler types - <ApiLink to="class/BeautifulSoupCrawler">`BeautifulSoupCrawler`</ApiLink> for fault-tolerant HTML parsing, <ApiLink to="class/ParselCrawler">`ParselCrawler`</ApiLink> for high-performance extraction with XPath and CSS selectors, and <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> for raw response processing. You also discovered how to integrate third-party parsing libraries with <ApiLink to="class/HttpCrawler">`HttpCrawler`</ApiLink> and how to create fully custom crawlers using <ApiLink to="class/AbstractHttpCrawler">`AbstractHttpCrawler`</ApiLink> for specialized parsing requirements.

If you have questions or need assistance, feel free to reach out on our [GitHub](https://github.com/apify/crawlee-python) or join our [Discord community](https://discord.com/invite/jyEM2PRvMU). Happy scraping!
Loading