Default preprocessing pipeline for MSI data based on Gaussian Mixture Modeling (GMM)
3.2K
Default preprocessing pipeline for MSI data in raw ASCII format, as used by Data Mining Group in Silesian University of Technology.
The packaged pipeline consists of the following steps:
The preferred installation is via Docker.
Having Docker installed, you can just pull the image:
docker pull gmrukwa/msi-preprocessing
You need to prepare your data for processing:
/mydata/mydata/raw - this is where pipeline expects your
original data/mydata
|- raw
|- my-dataset1
|- my-dataset2
|- ...
/mydata
|- raw
|- my-dataset1
|- my-dataset1_0_R00X309Y111_1.txt
|- my-dataset1_0_R00X309Y112_1.txt
|- my-dataset1_0_R00X309Y113_1.txt
|- my-dataset1_0_R00X309Y114_1.txt
|- my-dataset1_0_R00X309Y115_1.txt
|- ...
Note: File names are important, since R, X and Y is parsed as
metadata! If you put there broken values, spatial dependencies between
spectra will be lost.
700,043096125457 2
700,051503297599 2
700,059910520559 1
700,068317794335 0
...
<another-mz-value> <another-ions-count>
...
3496,66186447226 1
3496,68071341296 3
3496,69956240485 2
Both . and , are supported as decimal separator.
Example of the expected structure can be found in
sample-data.
You can launch preprocessing via:
docker run -v /mydata:/data -p 8082:8082 gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]'
Results will appear in the /mydata directory as soon as they are
available. You can track the progress live at
localhost:8082.
If you need output data also in the format of .csv files (not a binary
numpy-related .npy), you can simply add a switch --export-csv:
docker run -v /mydata:/data -p 8082:8082 gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]' --export-csv
Note: There is no space between dataset names.
Note: The --export-csv switch must appear right after the datasets
(due to the way Docker handles arguments).
If you want to review time needed for each task to process, you can prevent
scheduler from being stopped with --keep-alive switch:
docker run -v /mydata:/data -p 8082:8082 gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]' --keep-alive
Note: --keep-alive switch must always come last.
sample-data directorydocker run -v sample-data:/data -p 8082:8082 gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]'localhost:8082Building GMM model takes longer time (at least 1 hour), so be patient.
You can simply add e-mail notifications to your configuration. They will provide you with failure messages and notification, when the pipeline completes successfully. Two methods are supported: via SendGrid and via SMTP server.
template.env as .env.env file, set following values (rest of the content preserve
intact):LUIGI_EMAIL_METHOD=sendgrid
LUIGI_EMAIL_RECIPIENT=<your-email-here>
LUIGI_SENDGRID_APIKEY=<your-api-key-here>
--env-file .env:docker run -v /mydata:/data --env-file .env gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]'
template.env](https://github.com/gmrukwa/msi-preprocessing-pipeline/template.env) as .env`.env file, set following values (rest of the content preserve
intact):LUIGI_EMAIL_METHOD=smtp
LUIGI_EMAIL_RECIPIENT=<your-email-here>
LUIGI_EMAIL_SENDER=<your-email-here>
LUIGI_SMTP_HOST=<smtp-host-of-your-provider>
LUIGI_SMTP_PORT=<smtp-port-of-your-provider>
LUIGI_SMTP_NO_TLS=<False-if-your-provider-uses-TLS-True-otherwise>
LUIGI_SMTP_SSL=<False-if-your-provider-uses-TLS-True-otherwise>
LUIGI_SMTP_PASSWORD=<password-to-your-email-account>
LUIGI_SMTP_USERNAME=<login-to-your-email-account>
--env-file .env:docker run -v /mydata:/data --env-file .env gmrukwa/msi-preprocessing '["my-dataset1","my-dataset2"]'
Task history is collected to SQLite database. If you want to persist the
database, you need to mount the /luigi directory. This can be done via:
docker run -v tasks-history:/luigi -v /mydata:/data gmrukwa/msi-preprocessing
Content type
Image
Digest
Size
1.3 GB
Last updated
almost 7 years ago
docker pull gmrukwa/msi-preprocessing