Sign inSign up

unocha/hdx-scraper-dhs

By unocha

•Updated about 5 years ago

DHS HDX Scraper

Image
0

2.8K

unocha/hdx-scraper-dhs repository overview

⁠Collector for DHS's Datasets

Build Status Coverage Status

This script connects to the DHS API⁠ and extracts data country by country creating two datasets per country in HDX (national and subnational). The scraper takes around 10 hours to run. It makes in the order of 200 reads from DHS and 1000 read/writes (API calls) to HDX in total. It creates around 7000 temporary files of at most 1Mb in size and uploads them into HDX. It will be run monthly.

⁠Usage
python run.py

For the script to run, you will need to have a file called .hdx_configuration.yml in your home directory containing your HDX key eg.

hdx_key: "XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX"
hdx_read_only: false
hdx_site: prod

You will also need to supply the universal .useragents.yml file in your home directory as specified in the parameter user_agent_config_yaml passed to facade in run.py. The collector reads the key hdx-scraper-dhs as specified in the parameter user_agent_lookup.

Alternatively, you can set up environment variables: USER_AGENT, HDX_KEY, HDX_SITE, EXTRA_PARAMS, TEMP_DIR, LOG_FILE_ONLY

Tag summary

Content type

Image

Digest

Size

65.5 MB

Last updated

about 5 years ago

docker pull unocha/hdx-scraper-dhs