PPaxe is an easy-to-use protein-protein interaction extraction tool from scientific literature.
349
PPaxe?PPaxe is an easy to use Python app that scans MEDLINE records from PUBMED
and PMC , to extract sentences describing protein-protein interactions. Upon a list of PUBMED identifiers obtained by a keyword search at NCBI-Entrez (e.g. a set of gene names or a disease), it provides with a summary of putative protein-protein interactions and the sentences of the PUBMED records that describe them. This docker container provides all the tools and libraries required by PPaxe to run the command-line out of the box; just follow these simple steps:
1.- get the container image:
docker pull compgenlabub/ppaxe
2.- retrieve information about its usage with the help switch:
docker run compgenlabub/ppaxe -h
The output from PPaxe is saved into the container folder /ppaxe/output, in order to access to those files you should mount a local folder to that container folder via the volume switch, as in:
cd /your/folder
cat > search.pmids <<'EOF'
25211495
25196150
EOF
# remind to use absolute paths for the volume mounting
docker run -v /your/folder:/ppaxe/output:rw \
compgenlabub/ppaxe -v -p search.pmids \
-o interactions_found.tbl \
-r summary_report;
The above command will save in /your/folder the interactions found as a tabular file (interactions_found.tbl), and will provide with a summary report in html using the prefix of the -r switch (summary_report.html). This file will display several tables (sentences, interactions, and identifier frequencies), a simple interactive graph for the interactions found, as well as some plots.
Further usage details can be found on the PPaxe GitHub site.
Here you can find a short description of the available customizable environment vars to fit PPaxe-app to your system:
CORENLP_THREADS: set up the number of threads for the Stanford-CoreNLP annotator (defaults to 4).CORENLP_MAXMEM: set up the maximum amount of system memory for the Stanford-CoreNLP annotator (defaults to 4g).PPaxe project is licensed under the GNU GPL3 license - see the LICENSE file for details. PPaxe uses Stanford-CoreNLP as POS-tagger (aka Parse-Of-Speech tagger), which is also under the GNU GPL3 license. This software requires Java, we have chosen openjdk provided by the Ubuntu Linux distribution that serves as base of the container.
If you find this software tool useful for your research, we will appreciate if you cite it as follows:
PPaxe: easy extraction of protein occurrence and interactions from the scientific literature.
S. Castillo-Lara, J.F. Abril.
Bioinformatics, AOP November 2018, doi:bty988.
Content type
Image
Digest
Size
2.3 GB
Last updated
almost 8 years ago
docker pull compgenlabub/ppaxe