Structure-based sequence motifs of transcription factors captured in protein-DNA PDB complexes.
976
This repository distributes a Docker image with dnaprot ready to use with your own PDB file. DNAPROT is an algorithm used to produce structure-based sequence motifs that predict the specificity of transcription factors captured in protein-DNA complexes from the Protein Data Bank.
Binding specificities are estimated as position weight matrices (PWMs) of two types: contact and readout.
Contact PWMs are calculated by adding contacts between side-chains and nitrogen bases, assuming that the DNA molecule in the complex is the cognate sequence, after Morozov's approach.
Readout PWMs are instead derived both from a linear combination of
In general, both approaches provide similar PWMs (see the DNAPROT paper), and mean PWMs and sequence logos are provided. However, readout PWMs are evaluated as unreliable when the number of interface atomic contacts is less than 5 or when the cognate DNA sequence has an associated PWM score below the top 80%.
The Web database https://3dfootprint.eead.csic.es contains precomputed reports for such complexes and it is updated yearly. With this Docker container you can analyze your own complexes in PDB format.
Read more at:
Contreras-Moreira (2010) https://doi.org/10.1093/nar/gkp781
Espinosa-Angarica et al (2008) https://doi.org/10.1186/1471-2105-9-436
Morozov & Siggia (2007) https://doi.org/10.1073/pnas.0701356104
Input files are expected to be in PDB format. Should you have other protein-DNA complexes in in other formats you can convert them with other tools such as gemmi, MAXIT or CIFTr.
See instructions here.
Once Docker is set up you can see the options or run a dnaprot example in the terminal as follows:
docker run --rm -it eeadcsiccompbio/dnaprot dnaprot.pl
docker run --rm -it eeadcsiccompbio/dnaprot dnaprot.pl -i example/1je8_AB.pdb
Note: This takes longer the first time as the Docker image first needs to be downloaded.
In order to analyze your own files and save the output in a persistent, local folder you can do as follows:
docker run --rm -v "$PWD:$PWD" -w "$PWD" -u $UID:$GROUPS -it eeadcsiccompbio/dnaprot dnaprot.pl -i file1.pdb -P local_folder/
Option -d can be used to pass comma-separated options to the underlying DNAPROT algorithm, eg: -d -D,0.8. You can check the manual here.
Results include a text summary of the protein-DNA interface, with interactions classified as hydrogen bonds (H), water-mediated hydrogen bonds (w) or hydrophobic interactions (V) (read more at 3d-footprint):
: w : LYS NZ A0188 <- 5.54 -> DT O4 b0013 : score 2.199
: w : LYS NZ A0188 <- 5.40 -> DA N7 b0012 : score 1.727
: H : LYS NZ A0192 <- 2.86 -> DG N7 b0015 : score 5.24907
: H : LYS NZ A0192 <- 2.91 -> DG O6 b0015 : score 5.53573
: H : LYS NZ A0192 <- 3.05 -> DG O6 b0016 : score 5.37724
: w : LYS NZ B0188 <- 5.74 -> DT O4 a0033 : score 2.199
: w : LYS NZ B0188 <- 5.41 -> DA N7 a0032 : score 1.727
: H : LYS NZ B0192 <- 3.22 -> DG N7 a0035 : score 4.86654
: H : LYS NZ B0192 <- 3.28 -> DG O6 a0036 : score 5.11687
: V : VAL CG1 A0189 <- 3.65 -> DT C7 a0023 : score 5.16515
: V : LYS CE B0188 <- 3.85 -> DT C7 a0033 : score 2.34671
: V : VAL CG1 B0189 <- 3.80 -> DT C7 b0003 : score 4.97846
: V : LYS CE A0188 <- 3.83 -> DT C7 b0013 : score 2.35859
A readout PWM:
> PW = (PS * (1-w)) + (PD * weight) | weight = 0.50 | PMscale = 100
A | 23 19 3 29 17 0 21 75 12 25 29 34 4 26 14 4 15 84 19 21
C | 29 24 5 24 70 94 27 7 36 19 23 11 7 17 7 3 28 4 30 25
G | 25 30 5 27 5 1 23 7 11 21 18 37 7 34 75 67 18 5 24 32
T | 19 23 83 16 4 1 25 7 37 31 26 14 78 19 0 22 35 3 23 18
A cumulative contact PWM:
> PW = (PS * (1-w)) + (PD * weight) | weight = -1.00 | PMscale = 100
A | 24 24 13 56 9 14 23 38 22 24 24 30 19 19 14 10 14 60 24 24
C | 24 24 14 14 69 52 31 20 22 24 24 22 19 19 16 10 14 12 24 24
G | 24 24 13 13 9 14 21 19 25 24 24 22 19 39 52 66 14 12 24 24
T | 24 24 56 13 9 16 21 19 27 24 24 22 39 19 14 10 54 12 24 24
And the resulting mean PWM:
> PW = (PS * (1-w)) + (PD * weight) | weight = 0.50 | PMscale = 100
A | 25 22 8 42 13 8 22 56 17 24 26 33 11 23 14 7 14 72 21 23
C | 26 24 10 19 69 73 29 14 29 23 24 16 13 18 12 6 22 9 28 24
G | 24 27 9 21 8 7 22 13 18 22 21 29 13 36 63 67 16 8 24 28
T | 21 23 69 14 6 8 23 13 32 27 25 18 59 19 7 16 44 7 23 21
In addition, the script computes the best motif after selecting the top 50 scoring sites. This motif is now based on occurrences of nucleotides from top sites, and it is re-scaled in the arbitrary range [0-96], just like in 3d-footprint:
# IC=12.477 IC/col=0.780
A | 0 75 0 0 24 96 24 24 24 56 0 11 0 0 5 96
C | 0 5 96 96 24 0 24 24 24 11 0 11 0 0 5 0
G | 0 11 0 0 24 0 24 24 24 16 0 63 96 96 13 0
T | 96 5 0 0 24 0 24 24 24 13 96 11 0 0 73 0
Finally, sequence logos and two types of plots (graph, matrix) dissecting the protein-DNA interface are also produced (see examples at 3d-footprint).
Content type
Image
Digest
Size
166.4 MB
Last updated
over 4 years ago
docker pull eeadcsiccompbio/dnaprot