A dockerized version of OHDSI's WhiteRabbit tool.
555
A Dockerized build of OHDSI WhiteRabbit with:
docker pull edence/whiterabbit-container:main
Create a WhiteRabbit .ini file and pass this to the container:. The .ini file content can vary depending on what the source database/file format is, but generally has the following format:
WORKING_FOLDER = /data # Path to the folder where all output will be written
DATA_TYPE = PostgreSQL # "Delimited text files", "MySQL", "Oracle", "SQL Server", "PostgreSQL", "MS Access", "Redshift", "BigQuery", "Azure", "Teradata", "SAS7bdat"
SERVER_LOCATION = 127.0.0.1/db_name # Name or address of the server. For Postgres, add the database name
USER_NAME = postgres # User name for the database
PASSWORD = supersecret # Password for the database
DATABASE_NAME = schema_name # Name of the data schema used
DELIMITER = , # The delimiter that separates values
TABLES_TO_SCAN = * # Comma-delimited list of table names to scan. Use "*" (asterix) to include all tables in the database
SCAN_FIELD_VALUES = yes # Include the frequency of field values in the scan report? "yes" or "no"
MIN_CELL_COUNT = 5 # Minimum frequency for a field value to be included in the report
MAX_DISTINCT_VALUES = 1000 # Maximum number of distinct values per field to be reported
ROWS_PER_TABLE = 100000 # Maximum number of rows per table to be scanned for field values
CALCULATE_NUMERIC_STATS = no # Include average, standard deviation and quartiles in the scan report? "yes" or "no"
NUMERIC_STATS_SAMPLER_SIZE = 500 # Maximum number of rows used to calculate numeric statistics
See https://ohdsi.github.io/WhiteRabbit/WhiteRabbit.html#Source_Data for some more information about how to format the parameters for the different sources.
In the following example, there is a data sub-directory with a collection of .csv files, and an .ini file csv-source.ini that defines the parameters:
csv-source.ini
└── data
├── allergies.csv
├── careplans.csv
├── conditions.csv
├── encounters.csv
├── imaging_studies.csv
├── immunizations.csv
├── medications.csv
├── observations.csv
├── organizations.csv
├── patients.csv
├── procedures.csv
└── providers.csv
The content of csv-source.ini:
WORKING_FOLDER = /data
DATA_TYPE = Delimited text files
DELIMITER = ,
TABLES_TO_SCAN = *
SCAN_FIELD_VALUES = yes
MIN_CELL_COUNT = 5
MAX_DISTINCT_VALUES = 1000
ROWS_PER_TABLE = 100000
CALCULATE_NUMERIC_STATS = no
NUMERIC_STATS_SAMPLER_SIZE = 500
macOS / Linux:
docker run --rm \
-v "$(pwd)/data:/data:rw" \
-v "$(pwd)/csv-source.ini:/config/csv-source.ini:ro" \
edence/whiterabbit-container:main \
gui -ini /config/csv-source.ini
Windows (PowerShell):
docker run --rm `
-v "${PWD}\data:/data:rw" `
-v "${PWD}\csv-source.ini:/config/csv-source.ini:ro" `
edence/whiterabbit-container:main `
gui -ini /config/csv-source.ini
Windows (CMD):
docker run --rm ^
-v "%cd%\data:/data:rw" ^
-v "%cd%\csv-source.ini:/config/csv-source.ini:ro" ^
edence/whiterabbit-container:main ^
gui -ini /config/csv-source.ini
The resulting ScanReport.xlsx file will be written to the same directory as the .csv files, so make sure you have write access to that directory.
If you only want to include a sub-set of the .csv files, change the line for TABLES_TO_SCAN to the following format, listing the file names to include separated by comma:
TABLES_TO_SCAN = patients.csv,encounters.csv,conditions.csv,procedures.csv,medications.csv
In the following example, there is a data sub-directory available for writing the resulting ScanReport.xlsx file, and an .ini file postgres-source.ini that defines the parameters:
postgres-source.ini
└── data
The content of postgres-source.ini:
WORKING_FOLDER = /data
DATA_TYPE = PostgreSQL
SERVER_LOCATION = 192.168.1.123:5432/postgres
USER_NAME = postgres
PASSWORD = somestrongpassword
DATABASE_NAME = source_schema
DELIMITER = ,
TABLES_TO_SCAN = *
SCAN_FIELD_VALUES = yes
MIN_CELL_COUNT = 5
MAX_DISTINCT_VALUES = 1000
ROWS_PER_TABLE = 100000
CALCULATE_NUMERIC_STATS = no
NUMERIC_STATS_SAMPLER_SIZE = 500
macOS / Linux:
docker run --rm \
-v "$(pwd)/data:/data:rw" \
-v "$(pwd)/postgres-source.ini:/config/postgres-source.ini:ro" \
edence/whiterabbit-container:main \
gui -ini /config/postgres-source.ini
Windows (PowerShell):
docker run --rm `
-v "${PWD}\data:/data:rw" `
-v "${PWD}\postgres-source.ini:/config/postgres-source.ini:ro" `
edence/whiterabbit-container:main `
gui -ini /config/postgres-source.ini
Windows (CMD):
docker run --rm ^
-v "%cd%\data:/data:rw" ^
-v "%cd%\postgres-source.ini:/config/postgres-source.ini:ro" ^
edence/whiterabbit-container:main ^
gui -ini /config/postgres-source.ini
The resulting ScanReport.xlsx file will be written directory specified (data in this example), so make sure you have write access to that directory.
If you only want to include a sub-set of the tables, change the line for TABLES_TO_SCAN to the following format, listing the table names to include separated by comma:
TABLES_TO_SCAN = patient,blood_samples,comorb
There are a few other examples of .ini files in the WhiteRabbit GitHub repop: https://github.com/OHDSI/WhiteRabbit/tree/master/iniFileExamples
GUI mode runs WhiteRabbit in a headless X server with optional VNC.
macOS / Linux:
docker run --rm -p 5900:5900 \
-e ENABLE_VNC=1 \
-e VNC_PASSWORD=ohdsi \
edence/whiterabbit-container:main gui
Windows (PowerShell):
docker run --rm -p 5900:5900 `
-e ENABLE_VNC=1 `
-e VNC_PASSWORD=ohdsi `
edence/whiterabbit-container:main gui
Windows (CMD):
docker run --rm -p 5900:5900 ^
-e ENABLE_VNC=1 ^
-e VNC_PASSWORD=ohdsi ^
edence/whiterabbit-container:main gui
Then connect using any VNC client:
Host: localhost:5900
Password: ohdsi
WhiteRabbit is an OHDSI tool.
Please follow the OHDSI project’s licensing and attribution requirements.
Content type
Image
Digest
sha256:09b3550ed…
Size
502.2 MB
Last updated
7 months ago
docker pull edence/whiterabbit-container:sha-d0640c6