Sign inSign up

ms1tools/ms1tools-x86-64

By ms1tools

•Updated about 1 month ago

Image
0

1.1K

ms1tools/ms1tools-x86-64 repository overview

⁠MS1-Tools — Docker Quick Start

MS1-Tools annotates unknown features in LC-MS metabolomics data using probabilistic inference. This image can be used two ways: a browser GUI (no commands needed after starting the container) or the classic command-line interface. Full technical details, including the underlying method, are in docs/main.tex in the source repository.

This repository contains docker images implementing MS1-Tools, a machine learning method described in:

Martin Stražar, Ahmed Mohamed, Eivgeni Mashin, Rachele Invernizzi, Sarah Jeanfavre, Julian Avila-Pacheco, Sam Zimmerman, Motohiko Kadoki, Jeongho Lee, Zhihan Nan, Chenhao Li, Yeun-Hyeok Shin, Menglei Shuai, Marie-Madlen Pust, Panhasith Ung, Zher Yin Tan, Gleb Pishchany, Mohammad Seyedsayamdost, Daniel B. Graham, Damian Plichta, Clary B. Clish, Ramnik J. Xavier: Probabilistic inference of metabolites across inflammation states and microbiomes (in preparation, 2026).

⁠1. Get the image

Pull the image for your architecture (x86-64 for Intel/AMD, arm64 for Apple Silicon) and give it a short tag:

docker pull ms1tools/ms1tools-x86-64
docker tag ms1tools/ms1tools-x86-64 ms1tools

or

docker pull ms1tools/ms1tools-arm64
docker tag ms1tools/ms1tools-arm64 ms1tools

⁠2. Choose a mode

The container picks a mode at docker run time — no rebuild needed to switch:

Run with...You get
no commandthe browser GUI, on port 8080
an explicit command (e.g. ms1-main.py ...)the classic command-line interface

⁠3. Browser GUI

Start the container and expose port 8080. No volume mount is needed — inputs and outputs are exchanged through the browser itself:

docker run -p 8080:8080 ms1tools

Open http://localhost:8080⁠.

To try it immediately without preparing your own data, copy out the bundled example dataset (bacterial isolates, 329 samples / 351 annotated metabolites):

docker run -v ~/ms_data/:/data/ -it ms1tools cp -r /app/examples/isolates /data

This produces ~/ms_data/isolates/mbx_metadata_all.csv and mbx_counts_all.csv — upload those two files directly in the browser (or use the Use example dataset button in the GUI, which fetches them for you automatically).

From the page you can:

  • upload inputs — the metabolite metadata and counts tables, plus optional pipeline (config.yml) and model (models.yml) parameter files;
  • adjust settings — togglable panes generated from the default config.yml (one described subpane per pipeline stage) and models.yml (only the currently selected model's hyperparameters are shown), and limit the number of CPU cores used (default: all cores minus one);
  • run, pause, stop and re-run the pipeline — watch ms1-main.py's progress stream live, line by line; pause/resume or stop the run at any point; re-running resumes from the last incomplete step, re-executing only steps whose settings changed;
  • download results once the run finishes, as a single zip archive of the entire output directory; and
  • search by structure — query a completed run with a SMILES string, using the same ms1.search.search function as the Jupyter notebook below. The queried structure is drawn alongside a results table (ID, method, m/z rounded to four decimals, retention time rounded to two, predicted identity, query, probability and match level), sorted by match level first (exact, then mass, then fingerprint matches) and descending probability within each level. Downloadable as a CSV with every underlying column at full precision.

Closing the browser tab stops the pipeline (after a short grace period, so a page refresh is safe); the page warns before closing while a job is running.

⁠4. Command-line interface

⁠macOS / Linux

The two main inputs are a peak intensity table and peak metadata table. Copy the bundled example dataset to a local working directory — note the mapping -v ~/ms_data/:/data/ of your local directory to /data inside the container:

mkdir ~/ms_data/
docker run -v ~/ms_data/:/data/ -it ms1tools cp -r /app/examples/isolates /data

Input/output paths must be given relative to the container's view of your directory, i.e. under /data/:

IN_META=/data/isolates/mbx_metadata_all.csv
IN_COUNT=/data/isolates/mbx_counts_all.csv
OUT_DIR=/data/isolates/

Run the pipeline, again mounting your working directory:

docker run -it -v ~/ms_data/:/data ms1tools ms1-main.py -i $IN_META -c $IN_COUNT -o $OUT_DIR

This takes 10–20 minutes on most laptops. The default config.yml uses the lmogp_mini model (15 epochs, for quick debugging); change config: lmogp in config.yml for the full, more accurate 100-iteration model. Results appear under ~/ms_data/isolates/output/, including 7a_formatted_output/ with the final annotated tables.

To search the results interactively, start a Jupyter Notebook server:

docker run -it -v ~/ms_data/:/data -w /data/isolates/ -p 8888:8888 ms1tools jupyter notebook

Open the printed http://127.0.0.1:8888/tree?token=... URL, select search_notebook.ipynb, and search for compounds by structure.

⁠Windows (Command Prompt)

Commands are analogous, using forward slashes in paths for Docker's parser. Note that the variables below are set without quotes — unlike Bash, cmd.exe's set stores any quote characters as a literal part of the value, so quoting here and again at each use below (as these commands do) would double up and break the command line:

set DATA_DIR=C:/ms_data/
mkdir "%DATA_DIR%"
docker run -v "%DATA_DIR%":"/data" -it ms1tools sh -c "cp -r /app/examples/isolates /data"
set IN_META=/data/isolates/mbx_metadata_all.csv
set IN_COUNT=/data/isolates/mbx_counts_all.csv
set OUT_DIR=/data/isolates/
docker run -v "%DATA_DIR%":"/data" -it ms1tools sh -c "ms1-main.py -i %IN_META% -c %IN_COUNT% -o %OUT_DIR%"
docker run -v "%DATA_DIR%":"/data" -w /data/isolates/ -p 8888:8888 -it ms1tools jupyter notebook

Results appear under C:\ms_data\isolates\, same layout as above.

⁠Reference databases

The 1a_nominal_mass.db setting (or the GUI's database dropdown) accepts: HMDB, Norman, COCONUT, BUDDY, PubChem, MONA, LIPID_MAPS — or a custom CSV database (see docs/main.tex for the required columns).

Tag summary

Content type

Image

Digest

sha256:4b51da753…

Size

6 GB

Last updated

about 1 month ago

docker pull ms1tools/ms1tools-x86-64