domain crawler determine urls for a website asynchronously
342
crawls websites to gather all possible pages
Make sure to have Rust installed.
cargo runBuild and run the service.
docker build -t crawler . && docker run -dp 8000:8000 crawler
Build and run the service with compose.
docker-compose up
You can use program as a docker image.
you can use the crate to setup a tcp server to run on the machine.
curl --location --request POST 'http://0.0.0.0:8000/crawl' \
--header 'Content-Type: application/json' \
--data-raw '{"url": "http://www.drake.com", "id": 0 }'
// results
{
"pages": [
"http://www.drake.com/",
"http://www.drake.com/?hsLang=en"
],
"user_id": 0,
"domain": "http://www.drake.com"
}
ROCKET_ENV=dev
CRAWL_URL="http://api:8080/api/website-crawl-background"
check the license file in the root of the project.
Content type
Image
Digest
Size
496 MB
Last updated
over 5 years ago
docker pull jeffmendez19/crawler