Sign inSign up

frinve/graph-search

By frinve

•Updated about 3 years ago

GRAPH-SEARCH is a graphical abstract search engine over a semantic network.

Image
Data science
0

47

frinve/graph-search repository overview

⁠GRAPH-SEARCH Data Pipeline

================ Visit us at http://gmql.eu/graph-search⁠ ================

These are the instructions to reproduce the semantic network described in the related paper by Francesco Invernici, Prof. Anna Bernasconi, and Prof. Stefano Ceri, from the Department of Electronic, Information, and Bioengineering, Politecnico di Milano, Italy.

⁠System Requirements

The GRAPH-SEARCH data pipeline is developed in Python 3.10 and it is available as a docker (version >= 18.03.0-ce) image on DockerHub with the tag "pipeline".

The mining process has been tested on Ubuntu 18.04.5 LTS with Linux 4.15.0-197-generic, with Docker 18.03.0-ce. Our machine is powered by an Intel Xeon CPU E5-2660 with 56 cores and 378GB of RAM. We strongly suggest running the data pipeline scripts on a similar, or more powerful, machine. Please, be aware that, even on this machine, the mining process requires several days of computation.

⁠Dependencies

All the software dependencies and the assets required to run the pipeline are already included in the Docker image. Clearly, a working Docker installation is necessary. We also suggest running the pipeline container and the databases container with the provided docker-compose.yaml configuration file.

In addition, installations of neo4j Graph DB and MariaDB are necessary to process and store the semantic network. In these instructions, we suggest making these systems available through the docker-compose.yaml file reported below.

⁠Using the pipeline

⁠Preparation
  • Create the following directory hierarchy to share the files among the different containers. In this way, the semantic network will be preserved after removing the container.
# Initialize neo4j directories
mkdir -p graph-search/gs_neo4j/{conf,data,import,logs,plugins}
# Initialize mariaDB directories
mkdir -p graph-search/gs_mariadb/{entrypoint,data,conf}
# Initialize the export directories for the data pipeline
mkdir -p graph-search/pipeline_products/{bigrams,bigrams_cido,citedby,clean,curator,curator_cido,export,mine,mine_cido,raw,neo4j}
  • Copy the following init script for MariaDB into graph-search/gs_mariadb/entrypoint/00_init_tables.sql
CREATE DATABASE IF NOT EXISTS agave;
USE agave;

CREATE TABLE IF NOT EXISTS metadata
(
    cord_uid VARCHAR(10) PRIMARY KEY,
    sha VARCHAR(5000),
    source_x VARCHAR(35),
    title VARCHAR(2000),
    doi VARCHAR(100),
    pmcid VARCHAR(20),
    pubmed_id VARCHAR(50),
    license VARCHAR(20),
    abstract TEXT,
    publish_time DATE,
    authors TEXT,
    journal VARCHAR(600),
    mag_id VARCHAR(5),
    who_covidence_id VARCHAR(26),
    arxiv_id VARCHAR(20),
    pdf_json_files TEXT,
    pmc_json_files TEXT,
    url TEXT,
    s2_id VARCHAR(15),
    citedby_count INT
);
  • Copy the following MariaDB configuration file to graph-search/gs_mariadb/conf/custom.cnf
[mariadb]

innodb_buffer_pool_size=16GB
  • Copy the following docker-compose.yaml to graph-search/docker-compose.yaml
version: '3.1'

services:
  mariadb:
    image: mariadb:10.8.2
    environment:
      MYSQL_ROOT_PASSWORD: agave_mariadb_0*
      MYSQL_USER: agave
      MYSQL_PASSWORD: agave_password
      MYSQL_DATABASE: agave
    volumes:
      - ./gs_mariadb/entrypoint:/docker-entrypoint-initdb.d
      - ./gs_mariadb/data:/var/lib/mysql
      - ./gs_mariadb/conf:/etc/mysql/conf.d
    container_name: graph-search-db

  neo4j:
    image: neo4j:4.4-community
    environment:
      NEO4J_AUTH: neo4j/agave
      NEO4JLABS_PLUGINS: '["graph-data-science", "apoc"]'
    ports:
      - "7474:7474"
      - "7687:7687"
    volumes:
      - ./gs_neo4j/logs:/logs
      - ./gs_neo4j/data:/data
      - ./gs_neo4j/plugins:/plugins
      - ./gs_neo4j/import:/var/lib/neo4j/import
      - ./gs_neo4j/conf:/conf
    container_name: graph-search-neo4j

  pipeline:
	image: frinve/graph-search:pipeline
	volumes:
	  - ./pipeline_products:/pipeline/products
	container_name: graph-search-pipeline

A few additional lines for neo4j might be necessary for some older docker installations. First, identify your username id with

id -u

Add the following lines after image: neo4j:4.4-community

    user: 'XXXX:XXXX'
    privileged: true

where, with XXXX, we intend your username id found before.

⁠Creation of Semantic Network

The pipeline is managed with ploomber⁠, a Python package for data pipeline orchestration. Multiple steps compose the data pipeline in the pipeline.yaml config file, which is used by ploomber to orchestrate the execution of the tasks. The pipeline execution is automatically started when the container is run in the following steps.

⁠Semantic Network mining
  • Move inside graph-search directory
cd graph-search
  • Start the pipeline and the databases with:
docker-compose up -d --rm
  • You can check if everything is working (all containers should be "Up"):
docker-compose ps
  • Wait for the pipeline to complete. You can check the progression with:
docker-compose logs pipeline
  • Copy the files with the relationships of the semantic network to the import folder of neo4j
cp pipeline_products/bigrams.csv gs_neo4j/import/
cp pipeline_products/bigrams_cido.csv gs_neo4j/import/
  • Connect to neo4j Browser at the localhost address https://127.0.0.1:7474⁠ and execute the following instructions.
  • Create an index for all the entities
MATCH (n) SET n :ENTITY RETURN count(n);  
CREATE INDEX idx_umls_id IF NOT EXISTS FOR (n:ENTITY) ON (n.umls_id);  
CREATE CONSTRAINT ON (n:ENTITY) assert n.umls_id IS UNIQUE;  
  • Import the relationships. Please, wait for these files to be loaded, it might take up to some minutes.
:auto USING PERIODIC COMMIT  
LOAD CSV WITH HEADERS FROM 'file:///bigrams.csv' AS line FIELDTERMINATOR ';'  
MATCH (first:ENTITY), (second:ENTITY)  
WHERE first.umls_id = line.first_entity_id AND second.umls_id = line.second_entity_id  
CREATE (first)-[r:co_occurrence {name: line.bigram_name, frequency: toInteger(line.frequency), cramers_v: toFloat(line.cramers_v), pmi: toFloat(line.pmi), npmi: toFloat(line.npmi)}]->(second); 
:auto USING PERIODIC COMMIT  
LOAD CSV WITH HEADERS FROM 'file:///bigrams_cido.csv' AS line FIELDTERMINATOR ';'  
MATCH (first:ENTITY), (second:ENTITY)  
WHERE first.umls_id = line.first_entity_id AND second.umls_id = line.second_entity_id  
CREATE (first)-[r:co_occurrence {name: line.bigram_name, frequency: toInteger(line.frequency), cramers_v: toFloat(line.cramers_v), pmi: toFloat(line.pmi), npmi: toFloat(line.npmi)}]->(second); 
⁠Explore the Semantic Network

Now the Semantic Network is ready to be explored. Check the number of nodes:

MATCH (n) RETURN COUNT (n)

Tag summary

Content type

Image

Digest

sha256:68b83ddd7…

Size

942.8 MB

Last updated

about 3 years ago

docker pull frinve/graph-search:pipeline