Skip to content

Instructions for Lab Members Performing Crawls

KQ edited this page Oct 7, 2026 · 94 revisions

This page provides guidance for running and saving data from a crawl of our complete 11,708-site dataset. Performing a full crawl involves three main stages:

  • Crawling crawl-set-pt1.csv through crawl-set-pt8.csv (our crawl set divided into 8 batches)
  • Creating and crawling redo-sites.csv (which you'll generate based on the initial crawl results)
  • Parsing the crawled data and saving it to Google Drive

Below, we outline the steps involved in each stage.

Crawling the First 8 Batches

  1. Ensure that the XPI file is up to date. Each time the crawler code is changed you must re-package the Extension in XPI Format. Check commit history of this repo against the date on the XPI file to make sure the XPI was last repackaged with the most recent crawler code changes. The updated XPI file must be pushed to the GitHub along with the crawler code changes.

  2. Set the Mullvad VPN to the appropriate location you wish to crawl for:

  • For CA crawls we set the location to Los Angeles, California.
  • For CO crawls we set the location to Denver, Colorado.
  • For NJ_NY crawls we set the location to Secaucus, New Jersey. NOTE: Mullvad's NJ VPN uses an NY IP address, see Mullvad's New Jersey VPN for more information.
  • For CT crawls we perform the crawl from Wesleyan University's campus (if we are not on campus, we use the Wesleyan VPN; however, if possible, we try to not use a VPN for CT).
  1. Before starting the crawl, ensure you've set up Docker and cloned the crawler repository.

    • If you haven't already, follow steps 1–4 in the README to install and initialize Docker and set up the gpc-web-crawler repository:

      • Install Docker Desktop.
      • Authenticate Docker.
      • Clone the repository.
  2. In your Terminal, navigate to the gpc-web-crawler directory.

  3. Remove previous crawl outputs if present:

    rm -rf crawl_results
  4. Run the following command to verify that the Docker compose stack (gpc-web-crawler) isn't already running:

    make check-if-up
    • If the command prints true (the stack is running), run make stop to shut down the compose stack.
    • If it prints false (the stack is not running), proceed to the next step.
  5. For each batch number n (where n is from 1 to 8), repeat these steps:

    1. If using a Mac to perform the crawl, open a new terminal window and run the command: caffeinate This prevents your Mac from falling asleep during the crawl while it is running without altering its settings. Command C to end the caffeinated session. Otherwise adjust your settings to prevent your laptop from sleeping/shutting down during the crawl.

    2. Run:

      make start-debug
    3. When prompted for a number between 1 and 8, enter your chosen batch number (n).

    4. Once the crawl is complete, shut down the compose stack:

      make stop
    5. Inside the crawl_results folder, find the corresponding batch folder for the crawl you just performed. Locate analysis.json and debug.json, and affix the batch number to their name. ie, analysis.json should be renamed to analysis-pt1.json, and debug.json to debug-pt1.json for batch 1.

    6. Save the terminal output as a text file, named similarly to Terminal Saved Output-pt1.txt, to the batch folder, affixing the batch number as before.

    7. Open the batch folder, and ensure well-known-data.csv and well-known-errors.json are inside the folder entitled Extra. Sometimes, these files don't appear. If this is the case, you should redo the crawl. We also have a dedicated well-known/gpc.json Python Script to only perform a .well-known crawl. However, for data consistency, we have done a full redo crawl in the past. A possible remedy for repeated non-generation issues is restarting your computer and using a new terminal. Please also note that well-known-errors.json is not generated immediately after the crawler finishes gathering data and it can take up to 10 minutes to appear in the batch folder.

    8. Inside Docker, locate the crawl_driver1 container. Copy the output and save it as an .rtf file named similarly to logs-pt1.rtf inside of the Extra folder, affixing the batch number as before.

    9. Navigate to the localhost, and export the entries table as a .csv file. Save this to the Extra folder, naming it similarly to entries-pt1.csv, affixing the batch number as before.

  6. With all 8 baches completed, generate the merged well-known-data.csv file by running the following command, and rename its folder from well-known to Extra

./merge_well_known_data.sh
  1. Rename crawl_results to St_Crawl_Data_Mon_Year, where St is the state name, Mon is the month the crawl took place, and Year is the year the crawl took place.

  2. Upload the renamed crawl_results directory to the Web Crawler folder in Google drive.

    Alternatively, you may create a Crawl_Data_Mon_Year folder in the Web Crawler folder and upload each batch as you complete it, instead of uploading all of them at once.

  3. Upon finishing the crawl, open an issue in this repo "Back up crawl data Mon Year", where Mon is the month the crawl took place, and Year is the year the crawl took place. Assign it to @SebastianZimmeck.

Identifying and Crawling Redo Sites

After completing the initial crawl of the 8 batches, the next stage involves identifying and crawling the "redo sites." Redo sites are those that failed during the initial crawl due to issues with their subdomains. These sites will be recrawled without their subdomains.

To complete this stage:

  1. Go into your google drive and access MyDrive. Verify that that the GPC folder is accessible from MyDrive. If it isn't, the Colab notebook will not work and you must first perform the following steps:

    • Find the GPC folder either by clicking the hyperlink or looking inside Shared with me
    • Right click on the GPC folder, click on Organize, and then click on Add shortcut
    • Navigate to All locations, click on MyDrive, and finally click Add.
  2. Open the following Colab notebook.

  3. Update the variable path1 in the notebook to reference your most recent crawl dataset.

  4. Execute the notebook to produce the following files:

    • redo-sites.csv: Lists sites requiring a second crawl (with subdomains removed).
    • redo-original.csv: Stores the original URLs (including subdomains) for reference.
  5. Upload both redo-sites.csv and redo-original.csv into a folder named similarly to Crawl_Data_Mon_Year (corresponding to the month and year of your crawl) within the designated Google Drive directory.

  6. In your Terminal, navigate to the gpc-web-crawler directory.

  7. Run the following command to verify that the Docker compose stack (gpc-web-crawler) isn't already running:

    make check-if-up
    • If the command prints true (the stack is running), run make stop to shut down the compose stack.
    • If it prints false (the stack is not running), proceed to the next step.
  8. Update the contents of selenium-optmeowt-crawler/crawl-sets/sites.csv with the contents of redo-sites.csv

  9. Run a custom batch on the list of redo-sites:

    1. Run:
      make custom
    2. Once the crawl is complete, shut down the compose stack:
      make stop
  10. In the newly created crawl_results folder, find the folder prefixed CUSTOMCRAWL and rename it to redo.

  11. As for the other batches 1-8, create the Extra folder, move files around, save outputs, and rename files appropriately. Instead of affixing the batch number, affix -redo. ie, rename analysis.json to analysis-redo.json

  12. Upload the newly renamed folder redo to Crawl_Data_Mon_Year.

Post Processing the Crawl Data:

After completing all crawl batches including the redo crawls, you may now move on to post processing the data.

  1. Make sure that the GPC folder is accessible from MyDrive. If it isn't, follow the instructions under Step 1 of Identifying and Crawling Redo Sites.

  2. Begin by parsing the crawl data into the appropriate state google sheet (for example see Crawl_Data_CA) by using the Parse_Data_For_Google_Sheet colab. Edit the crawl_time, sheet_name, and data_path to define the month and state of the new crawl. Also create a new empty tab in appropriate state google sheet with the name of the monthyear of the crawl (This must match the sheet_name variable). Then run the colab and the sheet should be populated with all the crawl data.

  3. Inside of the Crawl_Data_<state> sheet, populate the BlankCellsCount and Counts tabs, and then copy these values into the <state>_Figures_Over_Time tabs.

    1. For BlankCellsCount, add an extra row, and populate it following the previous pattern. You should use the equations =COUNTBLANK('<monthYear>'!1:11709) for column B, =COUNTBLANK('<monthYear>'!M1:M11710) for column C, and =COUNTBLANK(<monthYear>!N1:N11709) for column D, where <monthYear> is the tab name for the data you're adding. Then, make sure to update the equations in the totals row such that they include this new data in their calculation.
    2. For the Counts tab, fill in a new row at the bottom, continuing the previous pattern. You'll use the following general format of equation =COUNTIF('NonCompliantSites<monthYear>'!reasonsColumn>, "*<applicable reason>s*"). To be clear, 'NonCompliantSites<monthYear>' refers to a tab inside the sheet. For an example of how this equation is used, column D, corresponding to Invalid_uspapi, should read =COUNTIF(NonCompliantSites<monthYear>!P:P, "*Invalid_uspapi*"). Small note: after the August 2025 crawl, the Reasons_Non_Compliant column has changed from O to P, since Tranco ranks were added into the data at column O. This is why the column referenced in the equation is different for crawls pre-Jan2026.
    3. Once you have all the new data filled out, navigate to the corresponding <state>_Figures_Over_Time sheet, and locate the similarly named tabs in there entitled BlankCellsCount and Counts. Manually copy the new data over into these tabs.

In event that changes are made to the Parse_Data_For_Google_Sheet colab, the post processing for all previous datasets must be redone. This can be done by simply re-running the colab on each dataset EXCEPT FOR the February 2024 CA crawl, which cannot be processed by the colab due to its raw data being lost (see #314 for more information). Instead, the February 2024 CA crawl can only be updated by creating specialized scripts to convert the old results to a newer version (e.g., the script made for issue #329).

Note: The Processing_Analysis_Data colab serves as a library for the other colabs.

Creating a new release

After a full crawl, we also want to publish a new image to the container repository and create a new release for the crawl. To do so, see the following steps:

  1. If you haven't already, create a Personal Access Token (PAT) with write:packages permission here: https://github.com/settings/tokens/new?type=classic. Make sure to save the token when you create it, you can only view it once.

  2. Open a new terminal.

  3. In the gpc-web-crawler root directory, run the following command, replacing YOUR_PAT with the access token from step 1 and YOUR_GITHUB_USERNAME with your Github username:

    echo YOUR_PAT | docker login ghcr.io -u YOUR_GITHUB_USERNAME --password-stdin
    
  4. Follow the instructions to Pack Extension in XPI Format

  5. Run the following commands:

    docker compose build
    docker compose push
    

    This will push the images to the Github container registry with a tag value of "latest".

  6. In your browser, navigate to https://github.com/orgs/privacy-tech-lab/packages.

  7. For every package, click the three buttons circled in the following image: Screenshot 2025-04-17 184758

  8. Copy the value of sha256 and save it somewhere.

  9. Write the changelog of the release normally. At the end, include the following lines:

    To pull the exact image versions used in this release:
    
    docker pull ghcr.io/privacy-tech-lab/crawl-driver@sha256:<SHA256 PLACEHOLDER>
    docker pull ghcr.io/privacy-tech-lab/well-known-crawl@sha256:<SHA256 PLACEHOLDER>
    docker pull ghcr.io/privacy-tech-lab/rest-api@sha256:<SHA256 PLACEHOLDER>
    docker pull ghcr.io/privacy-tech-lab/mariadb-custom@sha256:<SHA256 PLACEHOLDER>
    

    Make sure to replace the SHA256 placeholder with the value you found in step 7 for the specific image.

  10. Publish the release.

How to see the crawler browser

  1. Delete lines 115 and 116 in compose.yaml. These lines start the crawl_browser with VNC mode off, effectively mirroring a headless crawler.
  2. Start the crawler.
  3. In your browser, navigate to the Selenium Grid Hub located at localhost:4444. You should see a page that looks like this: image
  4. Hit sessions on the left of the screen.
  5. On the single running session, hit the camera all the way on the left.
  6. When prompted for a password, enter secret.

Accuracy Check Protocol

IMPORTANT NOTE: The accuracy check compares what the crawler reports vs your manual observation (i.e., ground truth) of the SAME DATA INSTANCE. You manually check whether the result the crawler reports for a site is the same you see in the data. The accuracy check does NOT involve any existing results from our regular crawls because over time a website may change its third party integrations. There may also be non-determinism, e.g., some site elements appear in one site load but not in another. That is why we have to perform the accuracy check FOR THE SAME SITE LOAD. See issue #214 for more details.

High-level accuracy check system:

We will be conducting accuracy check once every few months alternating between the three different locations (CO, CT, CA, NJ/NY, ...). Since we are crawling different locations using the same crawler and with the same methodology, one location chosen for an accuracy check would suffice as our general goal is to confirm whether our crawler is working as expected. To decide which state to use, check previous states used for accuracy check in Web_Crawl_Accuracy_Overtime, and choose one that has not been used recently.

Additional Note: While our goal is to confirm the validity and consistency of the data acquired from both the crawler and manual check at the same time, should we find an error that could fall under any of the category here, we should note down the error statement and include it in our codebase so that it will be flagged as an error appropriately in the following automated crawl.

Random Selection Sample of Sites

We choose random sample of sites from our batches using Google Apps Script. The script focuses on relevant columns of interest (uspapi, usp cookies, OptanonConsent, gpp, usps, Well-known, gpp_version, etc...) and checks if a row has valid non-null data in at least one of these columns. For each column of interest, 5 random rows with valid data are selected and added to an output sheet titled something like Crawl_Data_<state>_<monYear>_ReviewSites for manual ground truth analysis. In order to generate a sample of random sites, follow the steps below:

  1. Navigate to the relevant Google Sheet in the Overall Crawl Results Folder of the Google Drive for the desired crawl location.
  2. Select the sheet tab corresponding to the month you want to evaluate.
  3. Run the attached Google App Script function called selectRandomSites for that month's tab to generate a random sample of sites.
    • To run the App Script, navigate to Extensions/AppScript in the Tool Bar at the top of the sheet and then select run.
  4. After running the script, it will output a test list of sampled sites to a tab titled Crawl_Data_<state>_<monYear>_ReviewSites inside the Web Crawler Accuracy Overtime spreadsheet. Use the sites in this list for the accuracy check, and colour the blank cells using the legend according to the following rules. When done, rename the tab to follow the naming scheme of the other tabs.

Setting up VNC for accuracy check:

  1. To ensure the possibility of a manual check while the crawl is running, we need a VNC interface that would be smooth. While Selenium Grid is an accessible and friendly VNC for the viewing of the crawler, We found TigerVNC to offer a better option for manual verification involving accessing the cookie storage and putting commands on the console.
  2. Download the self-contained binaries for TigerVNC appropriate to your local device.
  3. Update the compose.yaml file in the codebase to add the port 5900:5900 for the crawl_browser and delete the line environment: - SE_START_VNC=false. These changes ensure that the VNC would be accessible and visible; it will look like this in the end: Screenshot 2025-03-21 at 6 12 37 PM

Accuracy Check Methodology

You will need the following commands to copy within TigerVNC: To do so, open a new tab and visit https://privacytechlab.org/gpccode.html to copy the commands over easily (it can alternatively be included as the first site on the custom crawl list). If on mac, make sure to use "control" instead of "command" when copy and pasting. Alternatively, the text can be highlighted and right-clicked the achieve the same effect. Ensure you type 'allow pasting' into console after you try pasting for the first time)

__uspapi('getUSPData', 1, (data) => { console.log("USP Data: ", data); });
__gpp('ping', (data, success) => { console.log(data, success); });
__gpp('getGPPData', (data, success) => { console.log(data, success); });

Pressing the up arrow while typing into the console will cycle through previous commands. This is the recommended method of interacting with the console after the above commands have all been pasted. Additionally, adjusting the filters in the top right of the console tab to only show "Logs" addresses most situations where a website's debug messages obscure the output of the pasted commands.

While adjusting the filters as described above is advised, please do not use the "filter output" textbox at the top of the console. While filtering for "__" will successfully filter out all irrelevant debug messages, it also occasionally hides the output of the console commands, resulting in an incorrectly-assigned NULL value.

Given the short time for gathering data before the gpc signal + refresh, it is recommended to screen recording or use a physical log to quickly record values. In the event that not all of the data was collected in a single attempt, the missing data must be obtained by crawling the website again. The following directions should be followed for every crawl:

  1. Increase Timeout in line 63 of webcrawler.js to 120000 or 180000 to allow time for manually writing the commands in the console before and after gpc signal is detected
  2. Load subset of custom sites to crawl to sites.csv. IMPORTANT: do the accuracy check in batches of ~5 sites given the relatively slow interface of the VNC.
  3. Start the crawler following the steps outlined in the ReadMe. Importantly, ensure before every crawl you run
make stop && make clean
make custom
  1. Open TigerVNC app on your local device and follow the steps to see the crawler browser. Please note that instead of localhost:4444, the VNC server inputted in TigerVNC should be localhost:5900. Connect with the password 'secret' if prompted.
  2. Determine the value of the US Privacy String Value by (1) checking the site's cookies via the Inspect site Storage Tab and (2) calling the USPAPI from the Web Console
  3. Determine the GPP String value by calling the GPP CMPAPI from the Web Console
  4. Determine OneTrust’s OptanonConsent cookie value by checking the site’s cookies via the Storage Tab
  5. After a few seconds, the crawler will send a gpc signal and the site will refresh. Repeat steps 4-7 and log these values under their "after gpc" variants. If any of the console commands return undefined or Uncaught ReferenceError: "command" is not defined, that is to be logged as a null value. Similarly, Cookies are given a null value if the website doesn't list them in the storage tab. Otherwise, the output should be recorded as it is presented.
  6. Determine .Well-known value by appending /.well-known/gpc.json to the URL path
  7. After the crawler has completely visited the sites in sites.csv, check that the data from steps 5-8 for each website matches the data in crawl_results/analysis.json (the comparison should be on live crawl data – not on crawl data from when the original crawl was conducted, so you should be comparing the results that appear in this analysis.json versus the manual values you got in 5-8)

If the manually collected values differ from what the crawler reports, we must determine whether that discrepancy arose from user error or from the crawler itself. While the methodology described above produces accurate results for most sites, there are cases where it proves insufficient. Issue #269, for example, documents a website for which the above console commands were ineffective, and required looking at the website's local storage. In such edge cases, a thorough ad-hoc investigation beyond the standard procedure is necessary for assessing the web crawler's accuracy.

Google Drive directories:

For detailed information about the folders and files, please check out the Google Drive ReadMe

General info about data analysis in the colabs:

GPP String decoding:

The GPP String encoding/decoding process is described by the IAB here. The IAB has a website to decode and encode GPP strings. This is helpful for spot checking and is the quickest way to encode/decode single GPP strings. They also have a JS library to encode and decode GPP strings on websites. Because we cannot directly use this library to decode GPP strings in Python, we converted the JS library to Python and use that for decoding (Python library found here). The Python library will need to be updated when the IAB adds more sections to the GPP string. More information on updating the Python library and why we use it can be found in issue 89. GPP strings are automatically decoded using the Python library in the colabs.

Updating the Crawler:

If you make a change to the crawler code, you must re-package the Extension in XPI Format. The updated XPI file must be pushed along with the crawler code changes.