Assigns WHO ICF domains and levels to Dutch clinical notes
485
This docker image is built from: https://github.com/cltl/aproof-icf-classifier
This repository contains a machine learning pipeline that reads a clinical note in Dutch and assigns the functioning level of the patient based on the textual description.
We focus on 9 WHO-ICF domains, which were chosen due to their relevance to recovery from COVID-19:
| ICF code | Domain | name in repo |
|---|---|---|
| b1300 | Energy level | ENR |
| b140 | Attention functions | ATT |
| b152 | Emotional functions | STM |
| b440 | Respiration functions | ADM |
| b455 | Exercise tolerance functions | INS |
| b530 | Weight maintenance functions | MBW |
| d450 | Walking | FAC |
| d550 | Eating | ETN |
| d840-d859 | Work and employment | BER |
The input is a csv file with at least one column containing the text (one clinical note per row).
The csv must follow the following specifications:
See example in example/input.csv.
The output file is saved in the same location as the input; it has 'output' added to the original file name.
The output file contains the same columns as the input + 9 new columns with the functioning levels per domain.
The functioning levels are generated per row. If a cell is empty, it means that this domain is not discussed in this note (according to the algorithm).
See example in example/input_output.csv.
The pipeline includes a multi-label classification model that detects the domains mentioned in a sentence, and 9 regression models that assign a level to sentences in which a specific domain was detected. All models were created by fine-tuning a pre-trained Dutch medical language model.
The pipeline includes the following steps:

docker pull piekvossen/a-proof-icf-classifier
docker run piekvossen/a-proof-icf-classifier
This will download all the required models from https://huggingface.co/CLTL and store them in the Docker's .cache, so that in subsequent runs cached models can be used. In total, 10 transformers models are downloaded, each between 500MB and 1GB.
To run the pipeline on your own data (i.e. a csv file on your local machine), you need to mount the local directory where the file is stored to the docker container. This is done with the -v flag and then <local_dir>:<docker_dir>. In addition, you need to pass the following arguments:
--in_csv: path to the input csv file--text_col: name of the text column in the csv (when running on Windows, do not use quotes or use double quotes for text values, e.g. --text_col “text”, instead of single quotes)--encoding (optional): use if input csv is not utf-8For example, if your csv file is in C:\Users\User\Desktop, it is called myfile.csv and the text is in the column note; you need to run the following command:
docker run -v C:\Users\User\Desktop:/root piekvossen/a-proof-icf-classifier --in_csv /root/myfile.csv --text_col note
To save the cached models on the local file system, or use them in a different container in a follow-up run, mount the Huggingface cache dir to a local directory. For example:
docker run -v <local_path_to_cache>:/root/.cache/huggingface/transformers/ piekvossen/a-proof-icf-classifier --in_csv example/input.csv --text_col text --sep ;
To use the cached models in an environment without internet connection, set TRANSFORMERS_OFFLINE=1 as environment variable (see Huggingface documentation). For example:
docker run -v <local_path_to_cache>:/root/.cache/huggingface/transformers/ -e TRANSFORMERS_OFFLINE=1 piekvossen/a-proof-icf-classifier --in_csv example/input.csv --text_col text --sep ;
The code runs faster if GPU is available on your machine; it is used automatically if it's available, no need to configure anything.
On some machines, you might run into issues when generating domains predictions (this function is applied to each sentence in the input file). If this is the case, split the input into smaller batches.
Content type
Image
Digest
sha256:34a7d4005…
Size
4.2 GB
Last updated
over 3 years ago
docker pull piekvossen/a-proof-icf-classifier