MS1-Tools annotates unknown features in LC-MS metabolomics data using probabilistic inference. This image can be
used two ways: a browser GUI (no commands needed after starting the container)
or the classic command-line interface. Full technical details, including the
underlying method, are in docs/main.tex in the source repository.
This repository contains docker images implementing MS1-Tools, a machine learning method described in:
Martin Stražar, Ahmed Mohamed, Eivgeni Mashin, Rachele Invernizzi, Sarah Jeanfavre, Julian Avila-Pacheco, Sam Zimmerman, Motohiko Kadoki, Jeongho Lee, Zhihan Nan, Chenhao Li, Yeun-Hyeok Shin, Menglei Shuai, Marie-Madlen Pust, Panhasith Ung, Zher Yin Tan, Gleb Pishchany, Mohammad Seyedsayamdost, Daniel B. Graham, Damian Plichta, Clary B. Clish, Ramnik J. Xavier: Probabilistic inference of metabolites across inflammation states and microbiomes (in preparation, 2026).
Pull the image for your architecture (x86-64 for Intel/AMD, arm64 for Apple Silicon) and give it a short tag:
docker pull ms1tools/ms1tools-x86-64
docker tag ms1tools/ms1tools-x86-64 ms1tools
or
docker pull ms1tools/ms1tools-arm64
docker tag ms1tools/ms1tools-arm64 ms1tools
The container picks a mode at docker run time — no rebuild needed to switch:
| Run with... | You get |
|---|---|
| no command | the browser GUI, on port 8080 |
an explicit command (e.g. ms1-main.py ...) | the classic command-line interface |
Start the container and expose port 8080. No volume mount is needed — inputs and outputs are exchanged through the browser itself:
docker run -p 8080:8080 ms1tools
Open http://localhost:8080.
To try it immediately without preparing your own data, copy out the bundled example dataset (bacterial isolates, 329 samples / 351 annotated metabolites):
docker run -v ~/ms_data/:/data/ -it ms1tools cp -r /app/examples/isolates /data
This produces ~/ms_data/isolates/mbx_metadata_all.csv and mbx_counts_all.csv —
upload those two files directly in the browser (or use the Use example dataset
button in the GUI, which fetches them for you automatically).
From the page you can:
config.yml) and model (models.yml) parameter files;config.yml
(one described subpane per pipeline stage) and models.yml (only the
currently selected model's hyperparameters are shown), and limit the number of
CPU cores used (default: all cores minus one);ms1-main.py's progress
stream live, line by line; pause/resume or stop the run at any point; re-running
resumes from the last incomplete step, re-executing only steps whose settings
changed;ms1.search.search function as the Jupyter notebook below. The queried
structure is drawn alongside a results table (ID, method, m/z rounded to four
decimals, retention time rounded to two, predicted identity, query, probability
and match level), sorted by match level first (exact, then mass, then
fingerprint matches) and descending probability within each level. Downloadable
as a CSV with every underlying column at full precision.Closing the browser tab stops the pipeline (after a short grace period, so a page refresh is safe); the page warns before closing while a job is running.
The two main inputs are a peak intensity table and peak metadata table. Copy the
bundled example dataset to a local working directory — note the mapping
-v ~/ms_data/:/data/ of your local directory to /data inside the container:
mkdir ~/ms_data/
docker run -v ~/ms_data/:/data/ -it ms1tools cp -r /app/examples/isolates /data
Input/output paths must be given relative to the container's view of your
directory, i.e. under /data/:
IN_META=/data/isolates/mbx_metadata_all.csv
IN_COUNT=/data/isolates/mbx_counts_all.csv
OUT_DIR=/data/isolates/
Run the pipeline, again mounting your working directory:
docker run -it -v ~/ms_data/:/data ms1tools ms1-main.py -i $IN_META -c $IN_COUNT -o $OUT_DIR
This takes 10–20 minutes on most laptops. The default config.yml uses the
lmogp_mini model (15 epochs, for quick debugging); change config: lmogp in
config.yml for the full, more accurate 100-iteration model. Results appear under
~/ms_data/isolates/output/, including 7a_formatted_output/ with the final
annotated tables.
To search the results interactively, start a Jupyter Notebook server:
docker run -it -v ~/ms_data/:/data -w /data/isolates/ -p 8888:8888 ms1tools jupyter notebook
Open the printed http://127.0.0.1:8888/tree?token=... URL, select
search_notebook.ipynb, and search for compounds by structure.
Commands are analogous, using forward slashes in paths for Docker's parser. Note that
the variables below are set without quotes — unlike Bash, cmd.exe's set stores
any quote characters as a literal part of the value, so quoting here and again at
each use below (as these commands do) would double up and break the command line:
set DATA_DIR=C:/ms_data/
mkdir "%DATA_DIR%"
docker run -v "%DATA_DIR%":"/data" -it ms1tools sh -c "cp -r /app/examples/isolates /data"
set IN_META=/data/isolates/mbx_metadata_all.csv
set IN_COUNT=/data/isolates/mbx_counts_all.csv
set OUT_DIR=/data/isolates/
docker run -v "%DATA_DIR%":"/data" -it ms1tools sh -c "ms1-main.py -i %IN_META% -c %IN_COUNT% -o %OUT_DIR%"
docker run -v "%DATA_DIR%":"/data" -w /data/isolates/ -p 8888:8888 -it ms1tools jupyter notebook
Results appear under C:\ms_data\isolates\, same layout as above.
The 1a_nominal_mass.db setting (or the GUI's database dropdown) accepts: HMDB,
Norman, COCONUT, BUDDY, PubChem, MONA, LIPID_MAPS — or a custom CSV
database (see docs/main.tex for the required columns).
Content type
Image
Digest
sha256:4b51da753…
Size
6 GB
Last updated
about 1 month ago
docker pull ms1tools/ms1tools-x86-64