Sign inSign up

filippis/phyrepower-docker

By filippis

•Updated about 10 years ago

PhyrePower is a software package for distant homology detection based on contact threading.

Artifact
Image
0

345

filippis/phyrepower-docker repository overview

⁠PhyrePower

PhyrePower is a software package for distant homology detection based on contact threading i.e. pairwise alignment of eigendecomposed contact maps.

Phyrepower-docker is the docker image containing the PhyrePower binaries.

⁠Requirements

Only Docker needs to be installed.

⁠Installation

Download image from public Docker Hub Registry.

$ docker pull filippis/phyrepower-docker

⁠General Use
  • Run a container, an instance of the image, without specifying a program. $ docker run --rm -it -e DOCKER_UID=`id -u` -v ~/PhyrePowerData:/input -v ~/output:/output filippis/phyrepower-docker

This will drop you into the command prompt of a bash shell in the home directory of phyrepower user within the container. Input data are mounted as /input. Data generated must be written into a mounted volume (e.g. /output), else they will be deleted along with the container (--rm automatically cleans up the container when the container exits). An alternative option is to run docker without --rm and mounted volume for output data and manually transfer the data using docker cp. Read below regarding the environmental variable DOCKER_UID.

  • Run a container specifying a PhyrePower program.
    • To generate an optimal alignment betwen a query and a template protein. $ docker run --rm -e DOCKER_UID=`id -u` -v ~/PhyrePowerData:/input -v ~/output:/output filippis/phyrepower-docker PhyrePower -q /input/queries/d1a5ta1.pp -d /input/templates_samples/d2ccya_.pp -o /output/d1a5ta1_d2ccya -m 1
    • To generate a PhyrePower file for a query protein. $ docker run --rm -e DOCKER_UID=`id -u` -v ~/PhyrePowerData:/input -v ~/output:/output filippis/phyrepower-docker PhyrePowerInput -f /input/<fasta-file> -s /input/<psipred-file> -c /input/<metapsicov-file> -o /output/<output-prefix> -t 1.19
    • To generate a PhyrePower file for a template protein. $ docker run --rm -e DOCKER_UID=`id -u` -v ~/PhyrePowerData:/input -v ~/output:/output filippis/phyrepower-docker PhyrePowerInput -f /input/<fasta-file> -s /input/<secondary-structure-file> -c /input/<pdb-file> -o /output/<output-prefix> -n

The above are general suggestions. Docker containers can be used in various ways.

⁠Usage

You can get a detailed list of options for PhyrePower binaries by running them with the -h option.

$ docker run --rm filippis/phyrepower-docker PhyrePower -h $ docker run --rm filippis/phyrepower-docker PhyrePowerInput -h

⁠User, Permissions and Security

The image is currently run as phyrepower user. This user will have uid (user id) as specified by the value of the environmetal variable DOCKER_UID passed in the commands above. In these commands DOCKER_UID is set to the uid of the user running the command. If DOCKER_UID is not passed, then phyrepower will have uid 9001. The gid (group id) of phyrepower is set equal to its uid.

This configuration serves two purposes:

  • The files written from within the container will be owned by the same user running the container.
  • Best security practise is to run applications within a container as a non root user.
⁠PhyrePower Output Data Format

PhyrePower program generates two output files: a txt one that contains info regarding the alignment(s) and a fasta one that contains the alignment(s) itself. Depending on the output mode these files have either only the best alignment or all alignments.

The txt file has one row per alignment. In case it contains info for the best alignment, it has all 7 following columns, while for all alignments only the first 5 columns are reported:

  • Column 1: query protein id.
  • Column 2: template protein id.
  • Column 3: alignment id. This corresponds to how many eigenvectors have been used and whether their product has been multiplied by -1.
  • Column 4: alignment description. This is a string of 0-1 digits. The number of digits corresponds to how many eigenvectors are used. Each digit corresponds to a specific number of eigenvectors with the rightmost digit corresponding to using 1 eigenvector, the second from the right to using 2 eigenvectors and so on. Digit value of 1 means that the product of eigenvectors has been multiplied by -1.
  • Column 5: Contact Map Overlap (CMO) value for the specific alignment.
  • Column 6: sum of the maximum CMO values, one for the best alignment for each number of eigenvectors i.e. 1, 2 and so on up to the max number of eigenvectors used.
  • Column 7: range of aligned residues from the query protein.

The fasta file has 5 lines per alignment. The first line is an id #<query-id>_<template-id>_<alignment-id> while the rest follow a standard alignment fasta format.

⁠PhyrePower Input Data Format

PhyrePower program accepts input data in a specific format. Such data can be generated using the PhyrePowerInput program. The input format of PhyrePower is an 11-line plain text format per protein:

  • Line 1: protein id (as within the fasta file) and length space-separated.
  • Line 2: sequence.
  • Line 3: 3-state secondary structure assignment or prediction.
  • Line 4: native or predicted contact map. Each contact is given as a pair of residue numbers space-separated. If the contact is predicted, then the confidence is included as well. All contacts are comma-separated. Zero-based numbering is used for residues.
  • Lines 5-11: top 7 eigenvectors of the contact map weighted by the square root of the corresponding eigenvalue. Values are space-separated. 1st line corresponds to 1st eigenvector, 2nd line to 2nd eigenvector and so on.

Files based on this format are named with suffix .pp.

⁠Download PhyrePower Input Data

Download current data from our server⁠.

$ curl -sf http://www.sbg.bio.ic.ac.uk/~phyrepower/public_data/PhyrePowerData.tgz | tar xz

The data folder contains 5 folders:

  • queries. It contains all pp files for 151 query proteins, proteins with predicted contact maps.
  • templates_db. It contains a merged version of all pp files for ~23k template proteins, proteins with native contact maps from the ASTRAL SCOPe 2.04. The individual pp files are concatenated and provided in 6 equal size splitted files.
  • templates_samples. It contains individual pp files for the 151 template proteins chosen.
  • models. It contains all 320 models generated by Tinker for each query protein (tgz files) as well as the final PhyrePower and Tinker models as selected by CMO (pdb files). Tinker model filenames follow this format <query-id>_<contacts-threshold>.<model-id>.pdb, while PhyrePower ones follow this <query-id>_<template-id>_<alignment-id>_<gap-opening-penalty>.pdb.
  • structures. It contains native structures for all 151 query proteins (cleaned up pdb files).
⁠License

PhyrePower is available for free for researchers at academic and non profit-making institutions. Commercial users please contact Prof. Sternberg⁠.

⁠Contact

For comments, bug reports and suggestions for improvement please contact us at this address⁠.

Tag summary

Content type

Unrecognized

Digest

Size

4.7 MB

Last updated

about 10 years ago

docker pull filippis/phyrepower-docker