Assigns the functioning level of the patient based on the Dutch textual description.
411
This repository is an update of the original repository: https://github.com/cltl/aproof-icf-classifier in hwich the classification is extended to 17 ICF categories. It contains a machine learning pipeline that reads a clinical note in Dutch and assigns the functioning level of the patient based on the textual description. For a global overview of the A-PROOF project see our website: https://cltl.github.io/a-proof-project/. For a more detailed description checkout the technical report in doc.
We focus on 17 WHO-ICF domains, which were chosen due to their relevance to recovery from COVID-19:
| ICF code | Domain | name in repo |
|---|---|---|
| b1300 | Energy level | ENR |
| b140 | Attention functions | ATT |
| b152 | Emotional functions | STM |
| b440 | Respiration functions | ADM |
| b455 | Exercise tolerance functions | INS |
| b530 | Weight maintenance functions | MBW |
| d450 | Walking | FAC |
| d550 | Eating | ETN |
| d840-d859 | Work and employment | BER |
| B280 | Sensations of pain | SOP |
| B134 | Sleep functions | SLP |
| D760 | Family relationships | FML |
| B164 | Higher-level cognitive functions | HLC |
| D465 | Moving around using equipment | MAE |
| D410 | Changing basic body position | CBP |
| B230 | Hearing functions | HRN |
| D240 | Handling stress and other psychological demands | HSP |
The input is a csv file with at least one column containing the text (one clinical note per row).
The csv must follow the following specifications:
See example in example/input.csv.
The output file is saved in the same location as the input; it has 'output' added to the original file name.
The output file contains the same columns as the input + 17 new columns with the functioning levels per domain.
The functioning levels are generated per row. If a cell is empty, it means that this domain is not discussed in this note (according to the algorithm).
See example in example/input_output.csv.
The pipeline includes a multi-label classification model that detects the domains mentioned in a sentence, and a single regression model that assign a level to sentences in which a specific domain was detected. All models were created by fine-tuning a pre-trained Dutch medical language model.
The pipeline includes the following steps:

docker pull piekvossen/a-proof-icf17-classifier
docker run --shm-size=1g piekvossen/a-proof-icf17-classifier
Note: The models are pre-cached inside the Docker image, so it is fully ready to run offline without downloading models at runtime.
To save the docker container, use:
docker save piekvossen/a-proof-icf17-classifier > aproof_image.tar
To run the pipeline on your own data (i.e. a csv file on your local machine), you need to mount the local directory where the file is stored to the docker container. This is done with the -v flag and then <local_dir>:<docker_dir>.
CRITICAL: You must include --shm-size=1g (or more) when running the docker container. The script uses multiprocessing and passes large chunks of data between processes via shared memory. The default Docker shared memory limit (64MB) will cause the container to crash with a Bus error. To use GPU acceleration, also add --gpus all.
For example, if your csv file is in C:\Users\User\Desktop, it is called myfile.csv and the text is in the column note where columns are separated with ";" you need to run the following command:
docker run --gpus all --shm-size=1g -v C:\Users\User\Desktop:/data piekvossen/a-proof-icf17-classifier --in_csv /data/myfile.csv --text_col note --sep ';'
All arguments are passed directly after the image name in the docker run command. Below is a full overview of all available parameters:
| Parameter | Default | Description |
|---|---|---|
--in_csv | ./example/input.csv | Path to the input CSV file (inside the container). |
--text_col | text | Name of the column containing the clinical notes. |
--encoding | utf-8 | File encoding of the input CSV. |
--sep | ; | Column separator of the input CSV. |
--out_csv | (next to input) | Path to write the output CSV. Defaults to <input_name>_output.csv in the same directory as the input. |
--out_sep | (same as --sep) | Column separator for the output CSV. |
--chunk_size | 1000 | Number of rows processed per chunk. Reduce if you run into memory issues. |
--prediction_batch_size | 32 | Batch size for transformer model inference. Reduce if GPU memory is insufficient. |
--spacy_batch_size | 128 | Batch size for spaCy sentence splitting. |
--spacy_n_process | 1 | Number of parallel spaCy processes. Keep at 1 inside Docker. |
--cuda_device | 0 | GPU device index to use. Ignored if no GPU is available. |
--domain_token | True | Whether to prepend a domain token before level prediction (recommended). |
--resume | False | Resume a previously interrupted run using saved checkpoints. See section below. |
--checkpoint_dir | (next to output) | Directory where chunk checkpoints and the manifest are stored. |
--snapshot_every_n_chunks | 5 | Write a partial assembled output file every N completed chunks. |
--stop_on_chunk_error | False | Stop the entire run if a single chunk fails. By default, failed chunks are skipped and retried on --resume. |
--chunk_timeout_seconds | 0 (disabled) | Kill and skip a chunk if it takes longer than this many seconds. |
--worker_startup_timeout_seconds | 1800 | Maximum time (in seconds) to wait for the worker process to load models on startup. |
--progress_every | 250 | Log a progress message every N sentences during spaCy preprocessing. |
--log_level | INFO | Logging verbosity: DEBUG, INFO, WARNING, ERROR. |
The pipeline saves a checkpoint after each successfully processed chunk. If a run is interrupted (e.g. the container is stopped, the server reboots, or a chunk times out), you can resume exactly where it left off.
For this to work with Docker, you must mount the checkpoint directory to a folder on the host machine, so that the checkpoints survive when the container stops. The simplest way is to mount your data directory and let the script create the checkpoint folder there automatically:
# First run
docker run --gpus all --shm-size=1g \
-v /your/data:/data \
piekvossen/a-proof-icf17-classifier \
--in_csv /data/myfile.csv --text_col note --sep ';'
# If it gets interrupted, resume with the exact same command + --resume
docker run --gpus all --shm-size=1g \
-v /your/data:/data \
piekvossen/a-proof-icf17-classifier \
--in_csv /data/myfile.csv --text_col note --sep ';' --resume
The checkpoint folder (myfile_output__checkpoints/) will be created automatically inside /your/data on your host machine. On --resume, the script reads the manifest from that folder and skips all chunks that were already completed, then picks up from the first unprocessed chunk.
Note: If you do not mount a volume, the checkpoint folder will be created inside the container and will be lost when the container exits. Always use
-vwhen you want resumable runs.
Because the models are now pre-downloaded and baked into the Docker image itself, you no longer need to mount a local cache directory or specify TRANSFORMERS_OFFLINE=1. The container works completely offline by default.
The code runs faster if GPU is available on your machine; it is used automatically if it's available, no need to configure anything.
On some machines, you might run into memory issues when generating predictions. In this case, reduce --chunk_size and/or --prediction_batch_size.
The a-proof-icf-classifier is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See MIT license for more details.
When using this repository please cite:
J. Kim, S. Verkijk, E. Geleijn, M. van der Leeden, C. Meskers, C. Meskers, S. van der Veen, P. Vossen, and G. Widdershoven, Modeling dutch medical texts for detecting functional categories and levels of covid-19 patients, 2022. In: Proceedings of the 13th Language Resources and Evaluation Conference, Marseille, June, 2022.
@proceedings{kim-etal-lrec2022, author={Jenia Kim and Stella Verkijk and Edwin Geleijn and Marieke van der Leeden and Carel Meskers and Caroline Meskers and Sabina van der Veen and Piek Vossen and Guy Widdershoven}, title={Modeling Dutch Medical Texts for Detecting Functional Categories and Levels of COVID-19 Patients}, booktitle={Proceedings of the 13th Language Resources and Evaluation Conference, Marseille, June, 2022}, year={2022} }
Content type
Image
Digest
sha256:5f042f8d6…
Size
5.7 GB
Last updated
9 days ago
docker pull piekvossen/a-proof-icf17-classifier