This project uses Apache StormCrawler to measure the adoption of the carbon.txt protocol. It takes as a starting point a list of hostnames and generates candidate URLs that are then fetched and verified as being valid carbon.txt files. OpenSearch is used for storing the URL Frontier and the metrics generated by the crawler. Apache StormCrawler runs on an Apache Storm crawler. Please go to the Apache StormCrawler website for more information.
- Java 17
- Maven 3.x
- Docker and Docker Compose
An Ansible playbook playbook.yml is provided to install these on a remote server.
First generate an uberjar:
mvn clean packageThis packages the StormCrawler code and resources that are then deployed on the Apache Storm cluster.
We get the list of hostnames to process from the CommonCrawl Foundation. They maintain a list of the most popular hostnames from their graphs, see https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2026-mar-apr-may/index.html.
First retrieve the full list of sorted hostnames with
curl -O https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main-2026-mar-apr-may/host/cc-main-2026-mar-apr-may-host-ranks.txt.gz
The entries are reversed e.g. com.facebook.www.
The following command will keep the top 10M hostnames, normalise the entries and get them ready for injection.
zcat cc-main-2026-mar-apr-may-host-ranks.txt.gz | head -n 10000000 | cut -f5 | awk -F. '{for (i=NF; i>0; i--) printf "%s%s", $i, (i>1 ? "." : "\n")}' | gzip > hostnames.gz
The variant below illustrates how to get a list of hostnames for the ones ranked between 50M and 100M.
zcat cc-main-2026-mar-apr-may-host-ranks.txt.gz | sed -n '50000001,100000000p' | cut -f5 | awk -F. '{for (i=NF; i>0; i--) printf "%s%s", $i, (i>1 ? "." : "\n")}' | gzip > 50_100M_hostnames.gz
The file containing the hostnames is mounted on the Storm supervisor container in the next step. Make sure you generate the list of hostnames before starting the containers (or restart them).
We provide a docker-compose.yaml file to launch OpenSearch, Zookeeper, Storm Nimbus, Storm Supervisor, and the Storm UI.
Please note that this is designed to work on a single machine to keep things simple.
docker compose up -dthen once everything is up and running call
dashboards/importDashboards.sh
This populates the OpenSearch Dashboard running on http://localhost:5601/. Check that the Storm UI is up and running on http://localhost:8080/.
NOTE
To keep the instructions simple, the OpenSearch setup used here is not secure: ONLY USE IN A SAFE ENVIRONMENT WHERE THE PORTS ARE NOT PUBLICLY AVAILABLE.
Even with the security off, OpenSearch requires a OPENSEARCH_INITIAL_ADMIN_PASSWORD variable to be set, this is stored in the .env file.
The next step consists in populating the status index in OpenSearch; it is used by StormCrawler to store the information it has about the URLs and their metadata.
For each hostname in the file generated above, the injection topology puts 2 URLs in the status index.
xxx/carbon.txt
xxx/.well-known/carbon.txt
To start the injection, we need to create a temporary container to join the services managed by our docker-compose file in their network. It loads the JAR and additional resources and sends them to the Storm cluster.
export NETWORK=carbontxt-crawler_default
docker run --network $NETWORK -it --rm \
-v "$(pwd)/injection-conf.yaml:/apache-storm/injection-conf.yaml" \
-v "$(pwd)/crawler-conf.yaml:/apache-storm/crawler-conf.yaml" \
-v "$(pwd)/opensearch-conf.yaml:/apache-storm/opensearch-conf.yaml" \
-v "$(pwd)/injection.flux:/apache-storm/injection.flux" \
-v "$(pwd)/target/carbontxt-crawler-1.0-SNAPSHOT.jar:/apache-storm/carbontxt-crawler-1.0-SNAPSHOT.jar" \
storm:2.8.8 storm jar carbontxt-crawler-1.0-SNAPSHOT.jar \
org.apache.storm.flux.Flux injection.flux
This container terminates as soon as the topology has been submitted to the Storm cluster. The topology itself runs on the Storm cluster and needs to be terminated explicitly.
You can list the topologies currently running with
docker run --network $NETWORK --rm storm:2.8.8 storm list
and kill one with
docker run --network $NETWORK --rm storm:2.8.8 storm kill crawler -w 0
To see whether the injection has finished, you can use the Storm UI at [http://localhost:8080], query OpenSearch for the number of documents in the status index or look at the OpenSearch dashboard.
Remember: there should be 2 URLs injected per hostname.
The crawl fetches the URLs that have been injected by the previous step and checks whether they are valid carbon.txt files. If so, they are indexed in a separate index called content.
In addition, the crawl topology checks the http headers and DNS records for pointers to carbon.txt files. If the values found are valid URLs, these are added to the status index, however if they are
hostnames, we generate 2 candidates URLs for them, just like in the injection step.
The crawl is triggered like so:
export NETWORK=carbontxt-crawler_default
docker run --network $NETWORK -it --rm \
-v "$(pwd)/crawler-conf.yaml:/apache-storm/crawler-conf.yaml" \
-v "$(pwd)/opensearch-conf.yaml:/apache-storm/opensearch-conf.yaml" \
-v "$(pwd)/crawler.flux:/apache-storm/crawler.flux" \
-v "$(pwd)/target/carbontxt-crawler-1.0-SNAPSHOT.jar:/apache-storm/carbontxt-crawler-1.0-SNAPSHOT.jar" \
storm:2.8.8 storm jar carbontxt-crawler-1.0-SNAPSHOT.jar \
org.apache.storm.flux.Flux crawler.fluxThe Storm UI can be used in combination with the Metrics dashboard to monitor the progress of the crawl.
The crawl keeps track of the method used to find the carbon.txt file (root, well-known,dns,http), in the event where a URL is found via more than one path, only the first one is tracked.
Due to the large number of hostnames and the fact that the crawl queries for DNS records, it is important to use a fast and reliable DNS server (i.e. probably not the one from your internet provider).
Google DNS (8.8.8.8 and 8.8.4.4) are free, robust and quite fast; the docker compose file uses them for the container supervisor but this can be changed if needed.
Apart from a fast internet connection, you need a machine with at least 16GB RAM and 100GB SSD storage.
The status index on OpenSearch takes 45.5GB for 200M URLs (from 100M hostnames), i.e. an average of 227.5 bytes per doc.
Finally, OpenSearch recommends that the swap is turned off, this is done with sudo swapoff -a on Linux.
The script OS_IndexInit.sh can be used to wipe out the OpenSearch indices.
The script ./export_content.sh extracts the carbon.txt files found during the crawl and exports it as a ndjson file called carbontxt.ndjson.
The content of the carbon.txt files are represented in Base64. Here is an example:
{"url":"https://thegreenwebfoundation.org/carbon.txt","hostname":"thegreenwebfoundation.org","method":"root","fetch_date":"20260617","content":"dmVyc2lvbj0iMC40IgpsYXN0X3VwZGF0ZWQ9IjIwMjUtMTItMjIiCgpbb3JnXQpkaXNjbG9zdXJlcyA9IFsKCXsgZG9jX3R5cGU9J3dlYi1wYWdlJywgdXJsPSdodHRwczovL3d3dy50ZWNoY2FyYm9uc3RhbmRhcmQub3JnL2Nhc2Utc3R1ZGllcy9ncmVlbi13ZWItZm91bmRhdGlvbi9vdmVydmlldycsIHRpdGxlPSdNZXRob2RvbG9neSBvZiBhcHBseWluZyB0aGUgVGVjaCBDYXJib24gU3RhbmRhcmQgdG8gR3JlZW4gV2ViIEZvdW5kYXRpb24nIH0sCiAgICB7IGRvY190eXBlPSdvdGhlcicsIHVybD0naHR0cHM6Ly93d3cudGhlZ3JlZW53ZWJmb3VuZGF0aW9uLm9yZy8ud2VsbC1rbm93bi90Y3MuanNvbicsIHZhbGlkX3VudGlsPScyMDIzLTEyLTMxJywgdGl0bGU9JzIwMjMgZGlnaXRhbCBlc3RhdGUgZW1pc3Npb24gZXN0aW1hdGVzIGZvciBHcmVlbiBXZWIgRm91bmRhdGlvbicgfQpdCgpbdXBzdHJlYW1dCnNlcnZpY2VzID0gWwoJeyBkb21haW49J3d3dy5oZXR6bmVyLmNvbScsIHNlcnZpY2VfdHlwZT0nYmxvY2stc3RvcmFnZScgfSwKICAgIHsgZG9tYWluPSd3d3cuaGV0em5lci5jb20nLCBzZXJ2aWNlX3R5cGU9J3ZpcnR1YWwtcHJpdmF0ZS1zZXJ2ZXJzJyB9LAogICAgeyBkb21haW49J3d3dy5zY2FsZXdheS5jb20nLCBzZXJ2aWNlX3R5cGU9J29iamVjdC1zdG9yYWdlJyB9LAogICAgeyBkb21haW49J3d3dy4zNHNwLmNvbScsIHNlcnZpY2VfdHlwZT0nbWFuYWdlZC13b3JkcHJlc3MtaG9zdGluZycgfSwKICAgIHsgZG9tYWluPSd3d3cuY2xvdWRmbGFyZS5jb20nLCBzZXJ2aWNlX3R5cGU9J2NvbnRlbnQtZGVsaXZlcnktbmV0d29yaycgfQpd"}The file resulting from the initial crawl of the top 100M hostnames in June/July 2026 is at carbontxt.ndjson.
Licensed under the Apache License, Version 2.0: http://www.apache.org/licenses/LICENSE-2.0