Sign inSign up

asava388/bio-nlp

By asava388

•Updated 10 months ago

Image
0

183

asava388/bio-nlp repository overview

bio-nlp — gene alias mapping & triplet extraction module (Jupyter)

This project is a small research workspace for biomedical text mining. It focuses on two things:

  1. Mapping gene mentions to official symbols using a large alias dictionary.
  2. Extracting simple cause–effect “triplets” from sentences (e.g., "Valproic acid —increase→ CD36"). It’s packaged as a Docker image with JupyterLab so anyone can reproduce the environment quickly.

What it does

  • Loads a gene dictionary (gene_dict.tsv, ~16 MB) from a GitHub RAW URL. Only Homo sapiens entries (taxid 9606) are used.
  • Normalizes text and builds an alias -> official map, also keeping identity mappings so official names match as-is.
  • Scans free text and prints matches as human-readable lines, e.g.: Valproic acid --increase--> CD36 (...) Exendin-4 --increase--> PRKAA1 (as AMPK) (...) [no gene targets]

Relationship extraction module

-This module runs biomedical relation extraction on PubMed-style texts. It uses two spaCy NER models (BioNLP13CG on raw text and BC5CDR on normalized text) to detect chemicals, diseases, genes and related biomedical entities, then merges these mentions. Using curated lists of “trigger” verbs (increase, decrease, protect, cause, etc.), it scans each sentence and builds simple triplets of the form (chemical, verb, affected object, sentence context) that describe how a chemical influences a process, phenotype or disease. A helper routine can apply this extractor to all articles in a PubMed XML ZIP archive and export the resulting relations to TSV/CSV files.

Notes & limitations

  • Dictionary-based matching; no NER model required.
  • Human genes only (NCBI Taxonomy 9606).
  • Very short/ambiguous tokens are filtered out on purpose.
  • Designed for exploratory analysis; not a production pipeline

Tag summary

Content type

Image

Digest

sha256:566d9feef…

Size

402 MB

Last updated

10 months ago

docker pull asava388/bio-nlp:0.3