Sign inSign up

johnfmorton/scraper

By johnfmorton

•Updated almost 4 years ago

WIP: This image takes a CSV file of URLs and scans each URL for the presence of a string.

Image
0

61

johnfmorton/scraper repository overview

⁠Overview of "johnfmorton/scraper"

I had a CSV file of URLs that I needed to scrape, looking for the presence or absence of a defined string. This repo automated that process.

This image wasn't made for public consumption, but feel free to explore it.

⁠What does it do?

This is a page scanner for scraping the URLs for the presence of a string in the HTML code.

It is a node program that runs in a Docker container.

I've built the image on my Mac and tagged it as scraper.

docker image build -t scraper .

I've made an alias in my .zshrc file to make running it easier.

alias scraper='docker container run --rm -it -v "$PWD":/root scraper'

⁠What the program expects

  1. landingPageExport.csv - this is a CSV of all the URLs. It should be on the root level of the directory.
  2. In the index.js file, update line 5, const landingPagesCsv = 'landingPageExport.csv' as needed to point to a different location.
  3. The script is looking for a string of text to have been included on each page. It can be updated on line 6 of index.js, const oneTrustString = 'REPLACE_WITH_YOUR_STRING'.

⁠Running the script

From the command line, scraper dumps me into a CLI with node installed. From there, I can run the index.js file because I have a start script in my package.json file.

npm run start

⁠Expected output

Successfully output will generate results.csv, which will include the results.

Tag summary

Content type

Image

Digest

sha256:c380d5bfe…

Size

156.7 MB

Last updated

almost 4 years ago

docker pull johnfmorton/scraper:dev