Sign inSign up

piekvossen/a-proof-icf17-classifier

By piekvossen

•Updated 9 days ago

Assigns the functioning level of the patient based on the Dutch textual description.

Image
Machine learning & AI
0

411

piekvossen/a-proof-icf17-classifier repository overview

⁠a-proof-icf-classifier

⁠Contents

  1. Description⁠
  2. Input File⁠
  3. Output File⁠
  4. Machine Learning Pipeline⁠
  5. How to use?⁠
  6. Cached models⁠
  7. Runtime and File Size⁠

⁠Description

This repository is an update of the original repository: https://github.com/cltl/aproof-icf-classifier⁠ in hwich the classification is extended to 17 ICF categories. It contains a machine learning pipeline that reads a clinical note in Dutch and assigns the functioning level of the patient based on the textual description. For a global overview of the A-PROOF project see our website: https://cltl.github.io/a-proof-project/⁠. For a more detailed description checkout the technical report in doc.

We focus on 17 WHO-ICF⁠ domains, which were chosen due to their relevance to recovery from COVID-19:

ICF codeDomainname in repo
b1300Energy levelENR
b140Attention functionsATT
b152Emotional functionsSTM
b440Respiration functionsADM
b455Exercise tolerance functionsINS
b530Weight maintenance functionsMBW
d450WalkingFAC
d550EatingETN
d840-d859Work and employmentBER
B280Sensations of painSOP
B134Sleep functionsSLP
D760Family relationshipsFML
B164Higher-level cognitive functionsHLC
D465Moving around using equipmentMAE
D410Changing basic body positionCBP
B230Hearing functionsHRN
D240Handling stress and other psychological demandsHSP
⁠Functioning Levels
  • FAC and INS have a scale of 0-5, where 5 means there is no functioning problem.
  • The rest of the domains have a scale of 0-4, where 4 means there is no functioning problem.
  • For more information about the levels, refer to the annotation guidelines⁠.
  • NOTE: the values generated by the machine learning pipeline might sometimes be outside of the scale (e.g. 4.2 for ENR); this is normal in a regression model.

⁠Input file

The input is a csv file with at least one column containing the text (one clinical note per row).

The csv must follow the following specifications:

  • sep = ;
  • quotechar = "
  • the first row is the header (column names)

See example in example/input.csv⁠.

⁠Output file

The output file is saved in the same location as the input; it has 'output' added to the original file name.

The output file contains the same columns as the input + 17 new columns with the functioning levels per domain.

The functioning levels are generated per row. If a cell is empty, it means that this domain is not discussed in this note (according to the algorithm).

See example in example/input_output.csv⁠.

⁠Machine Learning Pipeline

The pipeline includes a multi-label classification model that detects the domains mentioned in a sentence, and a single regression model that assign a level to sentences in which a specific domain was detected. All models were created by fine-tuning a pre-trained Dutch medical language model⁠.

The pipeline includes the following steps:

ml_pipe drawio

⁠How to use?

⁠Step 1: Setting up Docker

  1. Install Docker Desktop: see here⁠ for Windows and here⁠ for macOS.
  2. Pull the docker image from DockerHub⁠ by typing in your command line:
docker pull piekvossen/a-proof-icf17-classifier
  1. Run the docker on the example/input.csv⁠ file (it is already in the docker image and is given as the default argument to the main.py⁠ script):
docker run --shm-size=1g piekvossen/a-proof-icf17-classifier

Note: The models are pre-cached inside the Docker image, so it is fully ready to run offline without downloading models at runtime.

To save the docker container, use: docker save piekvossen/a-proof-icf17-classifier > aproof_image.tar

⁠Step 2: Running the pipeline on your data

To run the pipeline on your own data (i.e. a csv file on your local machine), you need to mount the local directory where the file is stored to the docker container. This is done with the -v flag and then <local_dir>:<docker_dir>.

CRITICAL: You must include --shm-size=1g (or more) when running the docker container. The script uses multiprocessing and passes large chunks of data between processes via shared memory. The default Docker shared memory limit (64MB) will cause the container to crash with a Bus error. To use GPU acceleration, also add --gpus all.

For example, if your csv file is in C:\Users\User\Desktop, it is called myfile.csv and the text is in the column note where columns are separated with ";" you need to run the following command:

docker run --gpus all --shm-size=1g -v C:\Users\User\Desktop:/data piekvossen/a-proof-icf17-classifier --in_csv /data/myfile.csv --text_col note --sep ';'

⁠All Parameters

All arguments are passed directly after the image name in the docker run command. Below is a full overview of all available parameters:

ParameterDefaultDescription
--in_csv./example/input.csvPath to the input CSV file (inside the container).
--text_coltextName of the column containing the clinical notes.
--encodingutf-8File encoding of the input CSV.
--sep;Column separator of the input CSV.
--out_csv(next to input)Path to write the output CSV. Defaults to <input_name>_output.csv in the same directory as the input.
--out_sep(same as --sep)Column separator for the output CSV.
--chunk_size1000Number of rows processed per chunk. Reduce if you run into memory issues.
--prediction_batch_size32Batch size for transformer model inference. Reduce if GPU memory is insufficient.
--spacy_batch_size128Batch size for spaCy sentence splitting.
--spacy_n_process1Number of parallel spaCy processes. Keep at 1 inside Docker.
--cuda_device0GPU device index to use. Ignored if no GPU is available.
--domain_tokenTrueWhether to prepend a domain token before level prediction (recommended).
--resumeFalseResume a previously interrupted run using saved checkpoints. See section below.
--checkpoint_dir(next to output)Directory where chunk checkpoints and the manifest are stored.
--snapshot_every_n_chunks5Write a partial assembled output file every N completed chunks.
--stop_on_chunk_errorFalseStop the entire run if a single chunk fails. By default, failed chunks are skipped and retried on --resume.
--chunk_timeout_seconds0 (disabled)Kill and skip a chunk if it takes longer than this many seconds.
--worker_startup_timeout_seconds1800Maximum time (in seconds) to wait for the worker process to load models on startup.
--progress_every250Log a progress message every N sentences during spaCy preprocessing.
--log_levelINFOLogging verbosity: DEBUG, INFO, WARNING, ERROR.

⁠Resuming interrupted runs

The pipeline saves a checkpoint after each successfully processed chunk. If a run is interrupted (e.g. the container is stopped, the server reboots, or a chunk times out), you can resume exactly where it left off.

For this to work with Docker, you must mount the checkpoint directory to a folder on the host machine, so that the checkpoints survive when the container stops. The simplest way is to mount your data directory and let the script create the checkpoint folder there automatically:

# First run
docker run --gpus all --shm-size=1g \
  -v /your/data:/data \
  piekvossen/a-proof-icf17-classifier \
  --in_csv /data/myfile.csv --text_col note --sep ';'

# If it gets interrupted, resume with the exact same command + --resume
docker run --gpus all --shm-size=1g \
  -v /your/data:/data \
  piekvossen/a-proof-icf17-classifier \
  --in_csv /data/myfile.csv --text_col note --sep ';' --resume

The checkpoint folder (myfile_output__checkpoints/) will be created automatically inside /your/data on your host machine. On --resume, the script reads the manifest from that folder and skips all chunks that were already completed, then picks up from the first unprocessed chunk.

Note: If you do not mount a volume, the checkpoint folder will be created inside the container and will be lost when the container exits. Always use -v when you want resumable runs.

⁠Cached models

Because the models are now pre-downloaded and baked into the Docker image itself, you no longer need to mount a local cache directory or specify TRANSFORMERS_OFFLINE=1. The container works completely offline by default.

⁠Runtime and File Size

The code runs faster if GPU is available on your machine; it is used automatically if it's available, no need to configure anything.

On some machines, you might run into memory issues when generating predictions. In this case, reduce --chunk_size and/or --prediction_batch_size.

⁠License:

The a-proof-icf-classifier is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See MIT license⁠ for more details.

⁠Reference

When using this repository please cite:

J. Kim, S. Verkijk, E. Geleijn, M. van der Leeden, C. Meskers, C. Meskers, S. van der Veen, P. Vossen, and G. Widdershoven, Modeling dutch medical texts for detecting functional categories and levels of covid-19 patients, 2022. In: Proceedings of the 13th Language Resources and Evaluation Conference, Marseille, June, 2022.

⁠Bibtext:

@proceedings{kim-etal-lrec2022, author={Jenia Kim and Stella Verkijk and Edwin Geleijn and Marieke van der Leeden and Carel Meskers and Caroline Meskers and Sabina van der Veen and Piek Vossen and Guy Widdershoven}, title={Modeling Dutch Medical Texts for Detecting Functional Categories and Levels of COVID-19 Patients}, booktitle={Proceedings of the 13th Language Resources and Evaluation Conference, Marseille, June, 2022}, year={2022} }

Tag summary

Content type

Image

Digest

sha256:5f042f8d6…

Size

5.7 GB

Last updated

9 days ago

docker pull piekvossen/a-proof-icf17-classifier