Skip to content

Repository files navigation

Carbon.txt crawler

Purpose

This project uses Apache StormCrawler to measure the adoption of the carbon.txt protocol. It takes as a starting point a list of hostnames and generates candidate URLs that are then fetched and verified as being valid carbon.txt files. OpenSearch is used for storing the URL Frontier and the metrics generated by the crawler. Apache StormCrawler runs on an Apache Storm crawler. Please go to the Apache StormCrawler website for more information.

Prerequisites

  • Java 17
  • Maven 3.x
  • Docker and Docker Compose

An Ansible playbook playbook.yml is provided to install these on a remote server.

Compilation

First generate an uberjar:

mvn clean package

This packages the StormCrawler code and resources that are then deployed on the Apache Storm cluster.

Hostnames

We get the list of hostnames to process from the CommonCrawl Foundation. They maintain a list of the most popular hostnames from their graphs, see https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2026-mar-apr-may/index.html.

First retrieve the full list of sorted hostnames with

curl -O https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2026-mar-apr-may/host/cc-main-2026-mar-apr-may-host-ranks.txt.gz

The entries are reversed e.g. com.facebook.www.

The following command will keep the top 10M hostnames, normalise the entries and get them ready for injection.

zcat cc-main-2026-mar-apr-may-host-ranks.txt.gz |  head -n 10000000 | cut -f5 | awk -F. '{for (i=NF; i>0; i--) printf "%s%s", $i, (i>1 ? "." : "\n")}' | gzip > hostnames.gz

The variant below illustrates how to get a list of hostnames for the ones ranked between 50M and 100M.

zcat cc-main-2026-mar-apr-may-host-ranks.txt.gz | sed -n '50000001,100000000p' | cut -f5 | awk -F. '{for (i=NF; i>0; i--) printf "%s%s", $i, (i>1 ? "." : "\n")}' | gzip > 50_100M_hostnames.gz

The file containing the hostnames is mounted on the Storm supervisor container in the next step. Make sure you generate the list of hostnames before starting the containers (or restart them).

Start Docker containers

We provide a docker-compose.yaml file to launch OpenSearch, Zookeeper, Storm Nimbus, Storm Supervisor, and the Storm UI. Please note that this is designed to work on a single machine to keep things simple.

 docker compose up -d

then once everything is up and running call

 dashboards/importDashboards.sh

This populates the OpenSearch Dashboard running on http://localhost:5601/. Check that the Storm UI is up and running on http://localhost:8080/.

NOTE To keep the instructions simple, the OpenSearch setup used here is not secure: ONLY USE IN A SAFE ENVIRONMENT WHERE THE PORTS ARE NOT PUBLICLY AVAILABLE. Even with the security off, OpenSearch requires a OPENSEARCH_INITIAL_ADMIN_PASSWORD variable to be set, this is stored in the .env file.

URL injection

The next step consists in populating the status index in OpenSearch; it is used by StormCrawler to store the information it has about the URLs and their metadata. For each hostname in the file generated above, the injection topology puts 2 URLs in the status index.

xxx/carbon.txt
xxx/.well-known/carbon.txt

To start the injection, we need to create a temporary container to join the services managed by our docker-compose file in their network. It loads the JAR and additional resources and sends them to the Storm cluster.

export NETWORK=carbontxt-crawler_default
                                                                        
docker run --network $NETWORK -it --rm \
  -v "$(pwd)/injection-conf.yaml:/apache-storm/injection-conf.yaml" \
  -v "$(pwd)/crawler-conf.yaml:/apache-storm/crawler-conf.yaml" \
  -v "$(pwd)/opensearch-conf.yaml:/apache-storm/opensearch-conf.yaml" \
  -v "$(pwd)/injection.flux:/apache-storm/injection.flux" \
  -v "$(pwd)/target/carbontxt-crawler-1.0-SNAPSHOT.jar:/apache-storm/carbontxt-crawler-1.0-SNAPSHOT.jar" \
  storm:2.8.8 storm jar carbontxt-crawler-1.0-SNAPSHOT.jar \
  org.apache.storm.flux.Flux injection.flux

This container terminates as soon as the topology has been submitted to the Storm cluster. The topology itself runs on the Storm cluster and needs to be terminated explicitly.

You can list the topologies currently running with

docker run --network $NETWORK --rm storm:2.8.8 storm list

and kill one with

docker run --network $NETWORK --rm storm:2.8.8 storm kill crawler -w 0

To see whether the injection has finished, you can use the Storm UI at [http://localhost:8080], query OpenSearch for the number of documents in the status index or look at the OpenSearch dashboard.

Remember: there should be 2 URLs injected per hostname.

Running the crawl

The crawl fetches the URLs that have been injected by the previous step and checks whether they are valid carbon.txt files. If so, they are indexed in a separate index called content. In addition, the crawl topology checks the http headers and DNS records for pointers to carbon.txt files. If the values found are valid URLs, these are added to the status index, however if they are hostnames, we generate 2 candidates URLs for them, just like in the injection step.

The crawl is triggered like so:

export NETWORK=carbontxt-crawler_default

docker run --network $NETWORK -it --rm \
  -v "$(pwd)/crawler-conf.yaml:/apache-storm/crawler-conf.yaml" \
  -v "$(pwd)/opensearch-conf.yaml:/apache-storm/opensearch-conf.yaml" \
  -v "$(pwd)/crawler.flux:/apache-storm/crawler.flux" \
  -v "$(pwd)/target/carbontxt-crawler-1.0-SNAPSHOT.jar:/apache-storm/carbontxt-crawler-1.0-SNAPSHOT.jar" \
  storm:2.8.8 storm jar carbontxt-crawler-1.0-SNAPSHOT.jar \
  org.apache.storm.flux.Flux crawler.flux

The Storm UI can be used in combination with the Metrics dashboard to monitor the progress of the crawl.

The crawl keeps track of the method used to find the carbon.txt file (root, well-known,dns,http), in the event where a URL is found via more than one path, only the first one is tracked.

Hardware and configuration

Due to the large number of hostnames and the fact that the crawl queries for DNS records, it is important to use a fast and reliable DNS server (i.e. probably not the one from your internet provider). Google DNS (8.8.8.8 and 8.8.4.4) are free, robust and quite fast; the docker compose file uses them for the container supervisor but this can be changed if needed.

Apart from a fast internet connection, you need a machine with at least 16GB RAM and 100GB SSD storage. The status index on OpenSearch takes 45.5GB for 200M URLs (from 100M hostnames), i.e. an average of 227.5 bytes per doc.

Finally, OpenSearch recommends that the swap is turned off, this is done with sudo swapoff -a on Linux.

Index reset

The script OS_IndexInit.sh can be used to wipe out the OpenSearch indices.

NDJSON export

The script ./export_content.sh extracts the carbon.txt files found during the crawl and exports it as a ndjson file called carbontxt.ndjson.

The content of the carbon.txt files are represented in Base64. Here is an example:

{"url":"https://thegreenwebfoundation.org/carbon.txt","hostname":"thegreenwebfoundation.org","method":"root","fetch_date":"20260617","content":"dmVyc2lvbj0iMC40IgpsYXN0X3VwZGF0ZWQ9IjIwMjUtMTItMjIiCgpbb3JnXQpkaXNjbG9zdXJlcyA9IFsKCXsgZG9jX3R5cGU9J3dlYi1wYWdlJywgdXJsPSdodHRwczovL3d3dy50ZWNoY2FyYm9uc3RhbmRhcmQub3JnL2Nhc2Utc3R1ZGllcy9ncmVlbi13ZWItZm91bmRhdGlvbi9vdmVydmlldycsIHRpdGxlPSdNZXRob2RvbG9neSBvZiBhcHBseWluZyB0aGUgVGVjaCBDYXJib24gU3RhbmRhcmQgdG8gR3JlZW4gV2ViIEZvdW5kYXRpb24nIH0sCiAgICB7IGRvY190eXBlPSdvdGhlcicsIHVybD0naHR0cHM6Ly93d3cudGhlZ3JlZW53ZWJmb3VuZGF0aW9uLm9yZy8ud2VsbC1rbm93bi90Y3MuanNvbicsIHZhbGlkX3VudGlsPScyMDIzLTEyLTMxJywgdGl0bGU9JzIwMjMgZGlnaXRhbCBlc3RhdGUgZW1pc3Npb24gZXN0aW1hdGVzIGZvciBHcmVlbiBXZWIgRm91bmRhdGlvbicgfQpdCgpbdXBzdHJlYW1dCnNlcnZpY2VzID0gWwoJeyBkb21haW49J3d3dy5oZXR6bmVyLmNvbScsIHNlcnZpY2VfdHlwZT0nYmxvY2stc3RvcmFnZScgfSwKICAgIHsgZG9tYWluPSd3d3cuaGV0em5lci5jb20nLCBzZXJ2aWNlX3R5cGU9J3ZpcnR1YWwtcHJpdmF0ZS1zZXJ2ZXJzJyB9LAogICAgeyBkb21haW49J3d3dy5zY2FsZXdheS5jb20nLCBzZXJ2aWNlX3R5cGU9J29iamVjdC1zdG9yYWdlJyB9LAogICAgeyBkb21haW49J3d3dy4zNHNwLmNvbScsIHNlcnZpY2VfdHlwZT0nbWFuYWdlZC13b3JkcHJlc3MtaG9zdGluZycgfSwKICAgIHsgZG9tYWluPSd3d3cuY2xvdWRmbGFyZS5jb20nLCBzZXJ2aWNlX3R5cGU9J2NvbnRlbnQtZGVsaXZlcnktbmV0d29yaycgfQpd"}

The file resulting from the initial crawl of the top 100M hostnames in June/July 2026 is at carbontxt.ndjson.

License

Licensed under the Apache License, Version 2.0: http://www.apache.org/licenses/LICENSE-2.0

About

Crawling pipeline to estimate the adoption of the carbon.txt protocol

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages