WIP: This image takes a CSV file of URLs and scans each URL for the presence of a string.
61
I had a CSV file of URLs that I needed to scrape, looking for the presence or absence of a defined string. This repo automated that process.
This image wasn't made for public consumption, but feel free to explore it.
This is a page scanner for scraping the URLs for the presence of a string in the HTML code.
It is a node program that runs in a Docker container.
I've built the image on my Mac and tagged it as scraper.
docker image build -t scraper .
I've made an alias in my .zshrc file to make running it easier.
alias scraper='docker container run --rm -it -v "$PWD":/root scraper'
landingPageExport.csv - this is a CSV of all the URLs. It should be on the root level of the directory.index.js file, update line 5, const landingPagesCsv = 'landingPageExport.csv' as needed to point to a different location.index.js, const oneTrustString = 'REPLACE_WITH_YOUR_STRING'.From the command line, scraper dumps me into a CLI with node installed. From there, I can run the index.js file because I have a start script in my package.json file.
npm run start
Successfully output will generate results.csv, which will include the results.
Content type
Image
Digest
sha256:c380d5bfe…
Size
156.7 MB
Last updated
almost 4 years ago
docker pull johnfmorton/scraper:dev