Sign inSign up

eeadcsiccompbio/dnaprot

By eeadcsiccompbio

•Updated over 4 years ago

Structure-based sequence motifs of transcription factors captured in protein-DNA PDB complexes.

Image
Data science
0

976

eeadcsiccompbio/dnaprot repository overview

⁠dnaprot

This repository distributes a Docker image with dnaprot ready to use with your own PDB file. DNAPROT is an algorithm used to produce structure-based sequence motifs that predict the specificity of transcription factors captured in protein-DNA complexes from the Protein Data Bank⁠.

Binding specificities are estimated as position weight matrices (PWMs) of two types: contact and readout.

Contact PWMs are calculated by adding contacts between side-chains and nitrogen bases, assuming that the DNA molecule in the complex is the cognate sequence, after Morozov's⁠ approach.

Readout PWMs are instead derived both from a linear combination of

  • i) the array of scored atomic interactions at the interface (direct readout)
  • ii) the set of sequence-dependent deformations inferred from the DNA coordinates (indirect readout)

In general, both approaches provide similar PWMs (see the DNAPROT⁠ paper), and mean PWMs and sequence logos are provided. However, readout PWMs are evaluated as unreliable when the number of interface atomic contacts is less than 5 or when the cognate DNA sequence has an associated PWM score below the top 80%.

The Web database https://3dfootprint.eead.csic.es⁠ contains precomputed reports for such complexes and it is updated yearly. With this Docker container you can analyze your own complexes in PDB format.

Read more at:

Contreras-Moreira (2010) https://doi.org/10.1093/nar/gkp781⁠

Espinosa-Angarica et al (2008) https://doi.org/10.1186/1471-2105-9-436⁠

Morozov & Siggia (2007) https://doi.org/10.1073/pnas.0701356104⁠

⁠File formats

Input files are expected to be in PDB format. Should you have other protein-DNA complexes in in other formats you can convert them with other tools such as gemmi⁠, MAXIT⁠ or CIFTr⁠.

⁠Installing Docker

See instructions here⁠.

⁠Examples of use

Once Docker is set up you can see the options or run a dnaprot example in the terminal as follows:


docker run --rm -it eeadcsiccompbio/dnaprot dnaprot.pl 

docker run --rm -it eeadcsiccompbio/dnaprot dnaprot.pl -i example/1je8_AB.pdb

Note: This takes longer the first time as the Docker image first needs to be downloaded.

In order to analyze your own files and save the output in a persistent, local folder you can do as follows:


docker run --rm -v "$PWD:$PWD" -w "$PWD" -u $UID:$GROUPS -it eeadcsiccompbio/dnaprot dnaprot.pl -i file1.pdb -P local_folder/

Option -d can be used to pass comma-separated options to the underlying DNAPROT algorithm, eg: -d -D,0.8. You can check the manual here⁠.

⁠Output

Results include a text summary of the protein-DNA interface, with interactions classified as hydrogen bonds (H), water-mediated hydrogen bonds (w) or hydrophobic interactions (V) (read more at 3d-footprint⁠):

: w : LYS   NZ  A0188 <- 5.54 ->  DT   O4 b0013 : score 2.199  
: w : LYS   NZ  A0188 <- 5.40 ->  DA   N7 b0012 : score 1.727  
: H : LYS   NZ  A0192 <- 2.86 ->  DG   N7 b0015 : score 5.24907  
: H : LYS   NZ  A0192 <- 2.91 ->  DG   O6 b0015 : score 5.53573  
: H : LYS   NZ  A0192 <- 3.05 ->  DG   O6 b0016 : score 5.37724  
: w : LYS   NZ  B0188 <- 5.74 ->  DT   O4 a0033 : score 2.199  
: w : LYS   NZ  B0188 <- 5.41 ->  DA   N7 a0032 : score 1.727  
: H : LYS   NZ  B0192 <- 3.22 ->  DG   N7 a0035 : score 4.86654  
: H : LYS   NZ  B0192 <- 3.28 ->  DG   O6 a0036 : score 5.11687  
: V : VAL  CG1  A0189 <- 3.65 ->  DT   C7 a0023 : score 5.16515  
: V : LYS   CE  B0188 <- 3.85 ->  DT   C7 a0033 : score 2.34671  
: V : VAL  CG1  B0189 <- 3.80 ->  DT   C7 b0003 : score 4.97846  
: V : LYS   CE  A0188 <- 3.83 ->  DT   C7 b0013 : score 2.35859  

A readout PWM:

> PW = (PS * (1-w)) + (PD * weight) | weight = 0.50 | PMscale = 100
A |  23  19   3  29  17   0  21  75  12  25  29  34   4  26  14   4  15  84  19  21
C |  29  24   5  24  70  94  27   7  36  19  23  11   7  17   7   3  28   4  30  25
G |  25  30   5  27   5   1  23   7  11  21  18  37   7  34  75  67  18   5  24  32
T |  19  23  83  16   4   1  25   7  37  31  26  14  78  19   0  22  35   3  23  18

A cumulative contact PWM:

> PW = (PS * (1-w)) + (PD * weight) | weight = -1.00 | PMscale = 100
A |  24  24  13  56   9  14  23  38  22  24  24  30  19  19  14  10  14  60  24  24
C |  24  24  14  14  69  52  31  20  22  24  24  22  19  19  16  10  14  12  24  24
G |  24  24  13  13   9  14  21  19  25  24  24  22  19  39  52  66  14  12  24  24
T |  24  24  56  13   9  16  21  19  27  24  24  22  39  19  14  10  54  12  24  24

And the resulting mean PWM:

> PW = (PS * (1-w)) + (PD * weight) | weight = 0.50 | PMscale = 100
A |  25  22   8  42  13   8  22  56  17  24  26  33  11  23  14   7  14  72  21  23
C |  26  24  10  19  69  73  29  14  29  23  24  16  13  18  12   6  22   9  28  24
G |  24  27   9  21   8   7  22  13  18  22  21  29  13  36  63  67  16   8  24  28
T |  21  23  69  14   6   8  23  13  32  27  25  18  59  19   7  16  44   7  23  21

In addition, the script computes the best motif after selecting the top 50 scoring sites. This motif is now based on occurrences of nucleotides from top sites, and it is re-scaled in the arbitrary range [0-96], just like in 3d-footprint⁠:

# IC=12.477 IC/col=0.780
A |   0  75   0   0  24  96  24  24  24  56   0  11   0   0   5  96
C |   0   5  96  96  24   0  24  24  24  11   0  11   0   0   5   0
G |   0  11   0   0  24   0  24  24  24  16   0  63  96  96  13   0
T |  96   5   0   0  24   0  24  24  24  13  96  11   0   0  73   0

Finally, sequence logos and two types of plots (graph, matrix) dissecting the protein-DNA interface are also produced (see examples at 3d-footprint⁠).

Tag summary

Content type

Image

Digest

Size

166.4 MB

Last updated

over 4 years ago

docker pull eeadcsiccompbio/dnaprot