bio-nlp — gene alias mapping & triplet extraction module (Jupyter)
This project is a small research workspace for biomedical text mining. It focuses on two things:
What it does
Relationship extraction module
-This module runs biomedical relation extraction on PubMed-style texts. It uses two spaCy NER models (BioNLP13CG on raw text and BC5CDR on normalized text) to detect chemicals, diseases, genes and related biomedical entities, then merges these mentions. Using curated lists of “trigger” verbs (increase, decrease, protect, cause, etc.), it scans each sentence and builds simple triplets of the form (chemical, verb, affected object, sentence context) that describe how a chemical influences a process, phenotype or disease. A helper routine can apply this extractor to all articles in a PubMed XML ZIP archive and export the resulting relations to TSV/CSV files.
Notes & limitations
Content type
Image
Digest
sha256:566d9feef…
Size
402 MB
Last updated
10 months ago
docker pull asava388/bio-nlp:0.3