Sign inSign up

eeadcsiccompbio/tfmodeller

By eeadcsiccompbio

•Updated almost 5 years ago

Scans a library of protein-DNA complexes and builds comparative models of proteins bound to DNA.

Image
0

911

eeadcsiccompbio/tfmodeller repository overview

⁠tfcompare

This repository distributes a Docker image with TFmodeller ready to use. The TFmodeller program takes a protein sequence P and a template protein-DNA complex in PDB format and builds comparative models. These models are used to get an idea of the P-DNA interface, its evolution and the putative recognised DNA sequences.

This legacy container replaces the web server at http://www.ccg.unam.mx/tfmodeller⁠ , which originally allowed users to search for optimal templates; instead, the Docker container assumes users have an idea of what template is best to build this model.

TFmodeller exploits the current knowledge about protein-DNA interfaces contained in the Protein Data Bank and uses it to model similar interfaces related by homology. Results include an evolutionary contact matrix, a schematic representation of the putative binding interface and atomic coordinates of the modelled complex.

flowchart

Read more at:

B Contreras-Moreira, PA Branger, J Collado-Vides (2007) TFmodeller: comparative modelling of protein-DNA complexes. https://doi.org/10.1093/bioinformatics/btm148⁠

⁠File formats

Template files are expected to be in PDB format. Should you have other protein-DNA complexes in other formats you can convert them with other tools such as MAXIT⁠ or CIFTr⁠.

⁠Installing Docker

See instructions here⁠.

⁠Examples of use

Once Docker is set up you can see the options or run the demo in the terminal as follows:


docker run --rm -it eeadcsiccompbio/tfmodeller TFmodeller.pl

docker run --rm -it eeadcsiccompbio/tfmodeller TFmodeller.pl -i example/example.fas -t example/1cgp.pdb 

Note: This takes longer the first time as the Docker image first needs to be downloaded.

In order to analyze your own, local files you can do:


cd your_pdb_folder/
docker run --rm -v "$PWD:$PWD" -w "$PWD" -u $UID:$GROUPS -it eeadcsiccompbio/tfmodeller TFmodeller.pl -i query.fasta -t file.pdb

Note: The results will be saved to your local folder.

⁠User template + alignment mode (-a)

TFmodeller allows you to use custom alignments of the amino acid sequence of query and template. This might be useful when you are not satisfied by the automatic alignment. The alignment must be in FASTA format as well, with the input sequence on top, followed by the template's sequence. Headers will be ignored:


>sp|P0A9E5|FNR_ECOLI monomer
MIPEKRIIRRIQSGGCAIHCQDCSISQLCIPFTLNEHELDQLDNIIERKKPIQKGQTLFK
AGDELKSLYAIRSGTIKSYTITEQGDEQITGFHLAGDLVGFDAIG--SGHHPSFAQALET
SMVCEIPFETLDDLSGKMPNLRQQMMRLMSGEIKGDQDMILLLSKKNAEERLAAFIYNLS
RRFAQRGFSPREFRLTMTRGDIGNYLGLTVETISRLLGRFQKSGMLAVKGKYITIENNDA
LAQLAGHTRNVA
>template 1CGP chain A
---------------------------PTLEWFLSHCHIHKYP----------SKSTLIH
QGEKAETLYYIVKGSVAVLIKDEEGKEMILSYLNQGDFIGELGLFEEGQERSAWVRAKTA
CEVAEISYKKFRQLIQVNPDILMRLSAQMARRLQVTSEKVGNLAFLDVTGRIAQTLLNLA
K-QPDAMTHPDGMQIKITRQEIGQIVGCSRETVGRILKMLEDQNLISAHGKTIVV-----
------------

It is important to note that the software requires that the amino acid sequence of the aligned template exactly matches the sequence in the PDB file with the coordinates.

⁠Modelling a multimeric complex

It is possible to take advantage of TFmodeller to build multimeric models, in which two or more protein chains bind to the same DNA molecule. Of course it is necessary to use a multimeric template to do this, extracted from the PDB or generated with symmetry matrices. Here I will illustrate how to model a FNR dimer, the protein introduced earlier, which is known to be functional as a dimer. We will first obtain the sequence of the FNR dimer by concatenating two copies of the sequence; note that with heterodimers we should concatenate two different sequences:


>sp|P0A9E5|FNR_ECOLI monomer
MIPEKRIIRRIQSGGCAIHCQDCSISQLCIPFTLNEHELDQLDNIIERKKPIQKGQTLFK
AGDELKSLYAIRSGTIKSYTITEQGDEQITGFHLAGDLVGFDAIG--SGHHPSFAQALET
SMVCEIPFETLDDLSGKMPNLRQQMMRLMSGEIKGDQDMILLLSKKNAEERLAAFIYNLS
RRFAQRGFSPREFRLTMTRGDIGNYLGLTVETISRLLGRFQKSGMLAVKGKYITIENNDA
LAQLAGHTRNVA
MIPEKRIIRRIQSGGCAIHCQDCSISQLCIPFTLNEHELDQLDNIIERKKPIQKGQTLFK
AGDELKSLYAIRSGTIKSYTITEQGDEQITGFHLAGDLVGFDAIG--SGHHPSFAQALET
SMVCEIPFETLDDLSGKMPNLRQQMMRLMSGEIKGDQDMILLLSKKNAEERLAAFIYNLS 
RRFAQRGFSPREFRLTMTRGDIGNYLGLTVETISRLLGRFQKSGMLAVKGKYITIENNDA
LAQLAGHTRNVA

Then we need to align the FNR dimer to the dimeric PDB template, with two concatenated protein chains, A and B, and put the alignment in a FASTA formatted text file:


>sp|P0A9E5|FNR_ECOLI dimer
MIPEKRIIRRIQSGGCAIHCQDCSISQLCIPFTLNEHELDQLDNIIERKKPIQKGQTLFK
AGDELKSLYAIRSGTIKSYTITEQGDEQITGFHLAGDLVGFDAIG--SGHHPSFAQALET
SMVCEIPFETLDDLSGKMPNLRQQMMRLMSGEIKGDQDMILLLSKKNAEERLAAFIYNLS
RRFAQRGFSPREFRLTMTRGDIGNYLGLTVETISRLLGRFQKSGMLAVKGKYITIENNDA
LAQLAGHTRNVAMIPEKRIIRRIQSGGCAIHCQDCSISQLCIPFTLNEHELDQLDNIIER
KKPIQKGQTLFKAGDELKSLYAIRSGTIKSYTITEQGDEQITGFHLAGDLVGFDAIG--S
GHHPSFAQALETSMVCEIPFETLDDLSGKMPNLRQQMMRLMSGEIKGDQDMILLLSKKNA
EERLAAFIYNLSRRFAQRGFSPREFRLTMTRGDIGNYLGLTVETISRLLGRFQKSGMLAV
KGKYITIENNDALAQLAGHTRNVA
>template 1CGP chains A,B
---------------------------PTLEWFLSHCHIHKYP----------SKSTLIH
QGEKAETLYYIVKGSVAVLIKDEEGKEMILSYLNQGDFIGELGLFEEGQERSAWVRAKTA
CEVAEISYKKFRQLIQVNPDILMRLSAQMARRLQVTSEKVGNLAFLDVTGRIAQTLLNLA
K-QPDAMTHPDGMQIKITRQEIGQIVGCSRETVGRILKMLEDQNLISAHGKTIVV-----
--------PTLEWFLSHCHIHKYPSK----------------------------------
-------STLIHQGEKAETLYYIVKGSVAVLIKDEEGKEMILSYLNQGDFIGELGLFEEG
QERSAWVRAKTACEVAEISYKKFRQLIQVNPDILMRLSAQMARRLQVTSEKVGNLAFLDV
TGRIAQTLLNLAK-QPDAMTHPDGMQIKITRQEIGQIVGCSRETVGRILKMLEDQNLISA
HGKTIVV-----------------

⁠Output

Results include:

  • a matrix of homologous interface contacts, that shows a multiple interface alignment of the user's input sequence to one or more structurally related protein-DNA complexes. Note this is a vertical alignment with residues from the query on the left and then one column per related complex. Each column shows the aligned equivalent protein residue and the contacted nucleotide. Only N-ring (purine/pyrimidine) contacts are considered here. Residues marked with * are supposed to be interface residues in the query sequence. For instance

0208 E* ECRG------SCKG--HG----RG--EC 0.91

represents residue E(Glu) 208 from the query aligned to 7 equivalent residues, two of which are E that contact C nucleotides. The stats line


_ stats: contacts=13 Nring=5 specif=0.38

shows the evolutionary proportion of sequence-specific contacts for this complex (0.38), a number that can be related to the number of different DNA sequences potentially bound by this protein.

  • one or more comparative models of the input sequence in complex with DNA. Sequence alignments are printed in order to highlight interface contact residues and its degree of conservation with respect to the template complex. In addition, a schematic representation of the interface is printed to help in the task of identifying key residues and the recognised DNA motifs:

_1.00          0.23   1.00                        
_R0212A        T0206A E0208A                      
_G      A      t      C      G      c      a      
_:      :      :      :      :      :      :      
_       T      a      G      C      g      t      
_                            E0208A V0207A V0207A 
_                            1.00   0.00   0.00   

In this example we learn that E(Glu) 208, chain A, from the query protein sequence is very likely to be contacting two C nucleotides, as this interaction is conserved from the template, with a contact probability of 1.00. However, V(Val) 207 most likely will not contact G or T nucleotides, since this residue actually mutated with respect to the original R in the template and no base contacts can be identified for it. Nucleotides in lower case show parts of the putative DNA motif that have probably changed, as a result of mutations in contacting residues. A special case would be T(Thr) 206, chain A, found to be contacting a T nucleotide with a probability of 0.23. However, the matrix of homologous interface contacts provides further support for this contact, as there is a similar contact (TT) in the PDB:


0206 T* ST--------------SG----HC--TT 0.89

Each comparative model is attached in PDB format to the results email in compressed form (.tgz format). Programs such as Rasmol, PyMOL or DeepView can be used to display them.

⁠Performance

It is important to acknowledge key observations that affect the value of results generated by TFmodeller:

  • Generally, structurally related protein-DNA complexes have similar interfaces.
  • Interface similarity decreases as the protein sequences being compared diverge. However, not all interfaces are equally conserved. TFmodeller was mostly benchmarked using 3-helical bundle proteins, that include well known folds or motifs such as HTH or homeodomains, but not Zn-finger proteins.
  • Proteins with similar interfaces recognise similar DNA motifs, but the predictive value of modelled interfaces decreases linearly as templates diverge.
  • The accuracy of modelled interface side-chains also depends on the sequence divergence between the sequence being modelled and the template used. We actually benchmarked this effect by modelling 2193 interface hydrogen-binding residues. By classifying complexes in terms of %sequence identity between template and target, we were able to compile the following tables, that summarize the reliability (REL, frequency of conserved contacts after modelling) and deviations (RMSD, in angstrom) observed for each residue type in four %sequence identity intervals: 20-39 (0.2), 40-59 (0.4), 60-79 (0.6) and 80-100 (0.8):

REL  0.2  0.4  0.6  0.8  RMSD 0.2  0.4  0.6  0.8  obs  0.2 0.4 0.6 0.8  
ASP 0.50 0.57 0.00 ----  ASP 1.78 1.35 1.60 ----  ASP  028 021 003 ---
ASN 0.63 0.75 0.86 0.86  ASN 1.11 0.70 1.24 0.28  ASN  252 061 014 022
LYS 0.52 0.65 0.90 0.70  LYS 1.47 1.07 0.66 0.68  LYS  204 175 010 010
TYR 0.40 0.38 0.50 ----  TYR 2.44 2.00 3.54 ----  TYR  015 008 002 ---
GLU 0.50 0.70 0.60 0.80  GLU 1.51 1.10 0.76 1.18  GLU  115 122 005 005
ARG 0.48 0.73 0.63 0.71  ARG 2.13 1.41 1.20 1.16  ARG  507 283 054 021
CYS 1.00 1.00 ---- ----  CYS 0.90 0.99 ---- ----  CYS  002 001 --- ---
THR 0.23 0.40 0.40 0.80  THR 1.28 0.90 0.70 0.39  THR  013 005 005 010
HIS 0.42 0.18 0.33 ----  HIS 1.88 2.10 2.00 ----  HIS  036 011 006 ---
GLN 0.31 0.67 0.90 0.71  GLN 1.93 1.20 1.00 0.39  GLN  051 039 010 007
SER 0.39 0.18 0.56 ----  SER 2.05 1.57 1.56 ----  SER  038 011 009 ---
TRP ---- ---- 1.00 ----  TRP ---- ---- 0.46 ----  TRP  --- --- 002 ---
Total observations = 2193

Tag summary

Content type

Image

Digest

Size

302.3 MB

Last updated

almost 5 years ago

docker pull eeadcsiccompbio/tfmodeller