Sign inSign up

rochoa85/panmhc-parce

By rochoa85

•Updated about 5 years ago

PanMHC-PARCE: Protocol to engineer variants in epitopes for MHC class II pan-allele binding

Image
0

912

rochoa85/panmhc-parce repository overview

⁠PanMHC-PARCE

⁠Protocol to engineer variants in epitopes for MHC class II pan-allele binding

⁠Purpose

We present PanMHC-PARCE, a docker container for engineering new modifications to improve the binding of epitopes toward different MHC class II alleles. The protocol performs a random mutation in the binder sequence, then samples the bound conformations using molecular dynamics simulations, and evaluates the MHC-peptide interactions using multiple scoring functions. This is done in parallel for multiple alleles. The mutation is accepted if a consensus criterion (based on the binding scores) is favorable for the majority of alleles. The procedure is iterated with the aim to explore efficiently novel sequences with potential better affinities toward the specified MHC II alleles. The methodology has been prepared for of pan-allele binding but it has a design option to use a single allele. Below we provide detailed descriptions about the docker container, running the protocol, input files and tutorials.

⁠Docker setup

To run the docker image (https://hub.docker.com/r/rochoa85/panmhc-parce⁠), first you require to install Docker in the operating system of interest. A guide of the distributions for Docker Desktop (Windows and Mac) and Docker server (Linux) can be found here: https://docs.docker.com/engine/install/⁠

After verifying the correct installation of Docker, check if sudo is required to execute it. If this is the case, please add sudo to the docker commands described below.

The PanMHC-PARCE image can be downloaded by:

docker pull rochoa85/panmhc-parce:latest

To create the container execute the following command:

docker run -it rochoa85/panmhc-parce /bin/bash

After that, you can find the code and tutorial material in the folder: /home/PanMHC-PARCE (see below for details on running the protocol).

The previous docker run command is used only the first time. After that, please exit by entering the command exit. Then, you can check into the container created. With the command sudo docker ps -a, you find the container-id being in the first column. This id will be used to access the docker container later, so that the changes made to the container will be stored.

To check into the container again:

docker start container-id

then, open the bash environment with the following command

docker exec -it container-id /bin/bash

The previous commands enable reentering to the container bash environment. The start and exec commands will be used to keep accessing the container (with previous installed programs) and maintain the modifications and output files generated by the protocol.

Note: For instructions about how to use docker in different OS, follow these tutorials:

⁠Running the protocol

The PanMHC-PARCE protocol has been created and tested using Python3.5. The code utilizes several preinstalled packages and dependencies in the docker container (see details below).

The basic command line to run the script is:

python3.5 run_protocol.py [-h] -m MODE_DESIGN -s MUTATION_STRATEGY -p PEPTIDE_SOURCE -a ALLELE

where the arguments are:

arguments:

  -h, --help            show this help message and exit

  -m MODE_DESIGN        Choose a mode to run the script from two options: 1) single, 2) multiple. 
                        If single, only one allele will be used. If multiple, the four available
                        alleles (0101, 0301, 0401, 0501) will be selected. 
                        NOTE: The protocol is currently created for using either one or four alleles.

  -s MUTATION_STRATEGY  Mutation strategy to change the amino acids: 
                        1) random, 2) filters, 3) motif.

  -p PEPTIDE_SOURCE     Keyword that references the peptide name, and it is used for reading the input 
                        configuration files (see below). In the tutorials below, we use "reference" 
                        for the single allele example  and "vivax" for the pan-allele design.
                        NOTE: For a new peptide, the keyword name can be provided in this option.

  -a ALLELE            Allele(s) of the MHC class II DRB1 that will be selected from five options: 
                       1) 0101, 2) 0301, 3) 0401,  4) 1501 (if using the single mode) and 5) multi 
                       (if using the multiple mode).

This script has been configured to provide the main design strategies, select the appropriate inputs and configuration files based on these choices.

The protocol enables selecting different mutation strategies. The first one is random, which means that the mutations are chosen randomly in any of the positions selected. The second is filters, where a set of bioinformatics predictions are used to pre-select amino acids that do not increase the peptide hydrophobicity, or avoid the presence of patterns associated to solubility or synthesis issues. The third strategy is motif, where a set of scoring-matrices are used to select mutations with higher probabilities of being part of the core region. For any of the strategies, two different programs are available to do the single-point mutations (see Configuration file section).

In addition to the above keywords, a configuration file is necessary with the description of the parameters to run the protocol. This include providing the positions that will be mutated. A detailed explanation is provided in the next section.

⁠Configuration file

The configuration file should be located in the directory config_files. Its name should be config_<ALLELE>_<PEPTIDE_SOURCE>.txt where ALLELE and PEPTIDE_SOURCE are the same as the keywords used above.

The configuration file requires the following parameters, which can be changed by the user depending on the design requirements:

  • folder: Name of the folder where the output files of the protocol will be stored. This folder is created inside the design_output folder, which is generated in the selected workspace after running the protocol. NOTE: If the same design is run multiple times, change the folder name to avoid losing previous outputs.
  • src_route: Route of the PanMHC-PARCE folder where the src folder is located. This enables running the protocol in any location.
  • mode: The design mode, which has three possible options: start (start the protocol from zero), restart (start from a particular iteration of a previous run) and nothing (just run without modifying existing files for debugging purposes).
  • peptide_reference: The sequence of the peptide, or protein fragment that will be modified.
  • pdbID: Name of the structure that is used as input, which contains the protein, the peptide/protein binder and the solvent molecules.
  • chain: Chain id of the peptide/protein binder in the complex.
  • sim_time: Time in nanoseconds that will be used to sample the complex after each mutation. Recommended a minimum value of 5 ns.
  • num_mutations: Number of mutations that will be attempted.
  • try_mutations: Number of mutations tried after having minimization or equilibration problems. These issues might be encountered depending on the system and the previous sampling of the complex before starting the protocol.
  • residues_mod: These are the specific positions of the residues that will be modified. This depends on the peptide/protein length and the numbering in the PDB file. It is recommended to renumber the input structure to associate the first position to residue number 1, for avoiding errors in posterior stages related with the renumbering of the coordinate files.
  • md_route: Path to the folder containing the input files, which are the files used during the previous MD sampling of the system.
  • md_original: Name of the system PDB file located in the folder containing the previous MD sampling.
  • score_list: List of the scoring functions that will be used to calculate the consensus. Currently the package has available: BACH, Pisa, ZRANK, IRAD, BMF-BLUUES and FireDock. At least two should be selected.
  • half_flag: Flag that controls which part of the trajectory is used to obtain the average score. If 0, the full trajectory is used, if 1, only the last half.
  • threshold: Threshold used for the consensus. If the number of scoring functions in agreement are equal or greater than the threshold, then the mutation is accepted.
  • mutation_method: Protocol to perform the single-point mutations. There are 2 options: faspr (included within the code), and scwrl4 (requires external installation).

Optional arguments:

  • scwrl_path: Provide the path to Scwrl4 in case it is not installed in a PATH folder. By default the system will use the system path to call the program.
  • gmxrc_path: Provide the path to GMXRC in case Gromacs was not included previously in the system path.

For pan-allele design, the code contains the routes to search for the P. vivax allele input files. In the config_files folder, we provide a set of configurations files as examples for the tutorial systems described below.

⁠Input files

The input files required to run the protocol are:

  • Equilibrated structure of starting complex from long MD simulations. For each system, there should be a PDB file containing the initial complex: protein, peptide and solvent; the topology files of the chains; a GRO file of the PDB template structure with the same name of the PDB file. The folder with all the files should be located inside the design_input folder. The folder should be named <ALLELE>_<PEPTIDE_SOURCE>, the structures should be named <ALLELE>_<PEPTIDE_SOURCE>.pdb and <ALLELE>_<PEPTIDE_SOURCE>.gro, and the topology file should be topol.top.
  • Configuration file. Named: config_<ALLELE>_<PEPTIDE_SOURCE>.txt (described above). The file should be located in the config_files folder.

As starting-complexes examples, we provide 5 complexes that were subjected to long MD simulations in Gromacs. These files are located in the folder design_input.

For the single-allele tutorial:

  • 0101_reference: MHC II allele DRB1*01:01 (PDB id 1dlh) bound to reference peptide (YPKYVKQNTLKLAT)

For the multiple-allele tutorial:

  • 0101_vivax: MHC II allele DRB1*01:01 (PDB id 1dlh) bound to P. vivax peptide (DYDVVYLKPLAGMYK)
  • 0301_vivax: MHC II allele DRB1*03:01 (PDB id 1a6a) bound to P. vivax peptide (DYDVVYLKPLAGMYK)
  • 0401_vivax: MHC II allele DRB1*04:01 (PDB id 1j8h) bound to P. vivax peptide (DYDVVYLKPLAGMYK)
  • 1501_vivax: MHC II allele DRB1*15:01 (PDB id 1bx2) bound to P. vivax peptide (DYDVVYLKPLAGMYK)

⁠Changing the input files

Currently, the protocol is configured for running a single or 4 MHC II alleles with specific PDB id files. The user has the option to modify the starting peptide bound to the alleles (keyword PEPTIDE_SOURCE). To model new peptides in complex with structures of reference alleles, we recommend following the protocol: https://github.com/rochoa85/MHC_class_II_matrices⁠

Please note, after modeling the new peptide, the complexes have to be sampled with MDs in Gromacs and the equilibrated PDB, GRO and topology files should be provided to start the design protocol. NOTE: be careful with the naming of the starting complexes following the above guidelines.

⁠MDP files for Gromacs MD simulations

The protocol has a set of mdp files (with fixed names) to run the MD simulations at each mutation step. These include the minimization steps, the NVT equilibration and the NPT production stages. The parameters have been optimized to improve the protocol's efficiency and accuracy. However, these parameters can be modified by the user, the source files can be found in the folder src/start/mdp. The default temperature of the system is 310K.

⁠Tutorial examples

We present two examples, one using a single allele with a reference peptide, and a second using four alleles with the P. vivax epitope.

1. Single-allele design

This example involves the design an Influenza peptide bound to allele DRB1*01:01. The command to reproduce it is:

python3.5 run_protocol.py -m single -s random -p reference -a 0101

A design run, using the 0101 allele, is performed starting from reference peptide YPKYVKQNTLKLAT (PDB id 1dlh) with a random mutation strategy. The code calls the configuration file config_0101_reference.txt located in the config_files folder. The parameters are:

folder: design_reference_0101
src_route: /home/PanMHC-PARCE
mode: start
peptide_reference: YPKYVKQNTLKLAT
pdbID: 0101_reference
chain: C
sim_time: 5
num_mutations: 10
try_mutations: 10
half_flag: 1
residues_mod: 1,2,3,4,5,6,7,8,9,10,11,12,13,14
md_route: /home/PanMHC-PARCE/design_input/0101_reference
md_original: 0101_reference
score_list: bach,pisa,zrank,irad,bmf-bluues,firedock
threshold: 3
mutation_method: scwrl
scwrl_path: /usr/local/bin/scwrl4/Scwrl4
gmxrc_path: /usr/local/gromacs/bin/GMXRC

The output folder is created inside the design_output folder, with the designed peptides and the files generated during the MD simulations that are created during the run. Details of the outputs are provided below.

2. Multiple-allele design

For this example, the python script can be called as follows:

python3.5 run_protocol.py -m multiple -s motif -p vivax -a multi

where the mode is multiple (pan-allele), the mutation strategy uses a MHC II matrix motif to guide the most frequent substitutions, the peptide to be optimized is the P. vivax epitope toward the four alleles (0101, 0301, 0401, 0501). The code reads a file located in the config_files folder called (config_multi_vivax.txt) with the following information:

src_route: /home/PanMHC-PARCE
mode: start
peptide_reference: DYDVVYLKPLAGMYK
chain: C
sim_time: 5
num_mutations: 10
try_mutations: 10
half_flag: 1
residues_mod: 1,2,3,4,5,6,7,8,9,10,11,12,13,14,15
score_list: bach,pisa,zrank,irad,bmf-bluues,firedock
threshold: 3
mutation_method: scwrl
scwrl_path: /usr/local/bin/scwrl4/Scwrl4
gmxrc_path: /usr/local/gromacs/bin/GMXRC

If any of these parameters are missing, the protocol stops and prints a warning message to the user. This example has been configured to run the pan-allele design with the same P. vivax modelled epitope for the four MHC II alleles. The routes are internally provided in the code to find the input folders. If a new peptide wants to be run, the input folders and configuration files should be named and located according to the P. vivax example. For this example, the mpi version version of Gromacs uses all the possible cores in the machine. Due to that, we recommend to run the pan-allele design in a server with a sufficient number of cores (ideally more than 24 cores).

For each allele an output folder called design_vivax_X is generated with the output information about the design run, where X is 1 for allele 0101, 2 for 0301, 3 for 0401 and 4 for 1501. If a new pan-allele design is run with another peptide, the keyword in the middle of the folder names will change.

⁠Outputs

When a design run starts, an initial folder is created with the required input, as well as the folders that will store the outputs step-by-step. The following is a list of the folders created, and their specific content:

  • binder: Stores the peptide/protein structure after each mutation attempt.
  • target: Stores the target structure after each mutation attempt.
  • iterations: Stores the average scores calculated per each iteration.
  • mdp: Stores the mdp files used for the molecular dynamics simulations.
  • complexP: Stores the target-peptide/protein structure after each mutation attempt.
  • solvent: Stores the solvent box after each mutation attempt.
  • system: Stores the complete target-peptide/protein-solvent complex after each mutation attempt.
  • trajectory: Stores the MD trajectory of the previous mutations.
  • score_trajectory: Stores the average scores for each snapshot from the trajectories. The file is split into four columns. The first column is the score of the complex. The second and third are the scores for the receptor and peptide alone. The fourth column is the total score after doing the difference between the complex and each component. For all of them it is possible to score the complex only. However the scoring function BACH can calculate the energies per component and subsequently apply the difference.
  • log_npt: Stores the log file from each npt run to verify possible errors.
  • log_nvt: Stores the log file from each nvt run to verify possible errors.

The design results are summarized in the output file called mutation_report.txt, which contains details at each mutation step, such as the type of mutation, the average scores, the binder sequence and if the mutation was accepted or not. In the case of pan-allele design, an extra column is reported mentioning if the mutation was Favourable or Unfavourable for the particular allele, given the consensus acceptance for 3 of the 4 alleles. Regarding the type of mutation, it is defined with the nomenclature: [old amino acid][binder chain][position][new amino acid]. An example of a mutation is AB2P, which means that an alanine located in the position number 2 of the chain B is replaced by a proline.

The report includes failed attempts based on minimization or equilibration problems. The latest can happen depending on the MD result. To overcome these issues, the protocol automatically attempts a number of mutations using the immediately accepted structure. If the simulation keeps failing after a certain number of attempts (defined by the user in the input file with the keyword try_mutations), a new mutation will be attempted but using the complex that was accepted previous to the current one. If the problem persist during a number of mutations, the run will stop.

In addition, a file named gromacs.log stores the logging of all the Gromacs commands, and the file gromacs_general_output.txt stores the latest Gromacs command used, just to track the progress during the run.

With the final mutation_report.txt, it is possible to check and select the accepted sequences, and plot the scores to verify that they are minimizing after the mutation steps. The results are numbered per iteration step, and the folder's content facilitates locating the information.

⁠Preinstalled third-party tools (in the docker container):

In the following, we provide a brief description of the third-party tools, preinstalled in the docker container used in the design protocol:

If this program is already installed, its path to executable can be provided in the configuration file. The protocol uses Gromacs 5.1.4.

  • The scoring functions (BACH, Firedock, BMF-BLUUES, Pisa, ZRANK and IRAD) are provided in the src folder and configured to run the analysis.
  • An open source method to perform the single-point mutations, named FASPR (https://zhanglab.ccmb.med.umich.edu/FASPR/⁠) is available in the docker container. The executable is included within the src folder.
  • Scwrl4: Scwrl4 can be also used to perform the single-point mutations in replacement of FASPR. However, Scwrl4 is not installed in the docker container because it requires an academic license. If you want to use this option, please visit http://dunbrack.fccc.edu/retro/scwrl4/license/index.html⁠ to download it, accept the license and install it inside the container. The path to executable can be provided in the configuration file.

NOTE 1: The mutation method can be changed in the configuration files (see sections above). The path to the executable can be provided in the same configuration files. NOTE 2: All the required tools are installed in the docker container under Ubuntu OS.

⁠Preinstalled dependencies (in the docker container):

The following BioPython and additional python modules (minimum python3.5) are preinstalled in the container

pdb2pqr
python3-biopython
python3-pip
python3-tk
python3-yaml

NOTE: If the user wants to install these dependencies, an install_dependencies.sh file is provided to automatize the installation of dependencies in the Linux (Ubuntu) operating system.

⁠References

These are the references of the external programs used within the container:

  • Gromacs: Hess et al., J. Chem. Theory Comput., 4, 435–447, 2008
  • Scwrl4: Krivov et al., Proteins: Struct. Funct. Bioinf., 77 (4), 778–795, 2009.
  • FASPR: Huang, Pearce and Zhang. Bioinformatics, 36, 3758-3765, 2020.
  • Bach: Cossio et. al., Sci. Rep., 2, 1–8. 2012. Sarti et al., Comput. Phys. Commun., 184 (12), 2860–2865. 2013. Sarti et. al, Proteins: Struct. Funct. Bioinf., 83 (4), 621–630, 2015.
  • Pisa: Krissinel and Henrick, J. Mol. Biol., 372 (3), 774–797, 2007
  • Firedock: Andrusier et al., Proteins: Struct. Funct. Bioinf., 69 (1), 139–159, 2007.
  • Irad: Vreven et al., Protein Sci., 20 (9), 1576–1586, 2011.
  • Zrank: Pierce and Weng, Proteins: Struct. Funct. Bioinf., 67 (4), 1078–1086 2007.
  • BMF-Bluues: The programs bluues (Fogolari et al., BMC Bioinformatics, 13 (S4), S18, 2012) and score_bmf (Berrera et al., BMC Bioinformatics, 4, 8, 2003) were kindly provided by Prof. F. Fogolari, University of Udine.

⁠Support

In case the protocol is useful for other research projects and require some advice, please contact us to the email: [email protected]⁠

Tag summary

Content type

Image

Digest

Size

1.4 GB

Last updated

about 5 years ago

docker pull rochoa85/panmhc-parce