GRAPH-SEARCH is a graphical abstract search engine over a semantic network.
47
================ Visit us at http://gmql.eu/graph-search ================
These are the instructions to reproduce the semantic network described in the related paper by Francesco Invernici, Prof. Anna Bernasconi, and Prof. Stefano Ceri, from the Department of Electronic, Information, and Bioengineering, Politecnico di Milano, Italy.
The GRAPH-SEARCH data pipeline is developed in Python 3.10 and it is available as a docker (version >= 18.03.0-ce) image on DockerHub with the tag "pipeline".
The mining process has been tested on Ubuntu 18.04.5 LTS with Linux 4.15.0-197-generic, with Docker 18.03.0-ce. Our machine is powered by an Intel Xeon CPU E5-2660 with 56 cores and 378GB of RAM. We strongly suggest running the data pipeline scripts on a similar, or more powerful, machine. Please, be aware that, even on this machine, the mining process requires several days of computation.
All the software dependencies and the assets required to run the pipeline are already included in the Docker image. Clearly, a working Docker installation is necessary.
We also suggest running the pipeline container and the databases container with the provided docker-compose.yaml configuration file.
In addition, installations of neo4j Graph DB and MariaDB are necessary to process and store the semantic network. In these instructions, we suggest making these systems available through the docker-compose.yaml file reported below.
# Initialize neo4j directories
mkdir -p graph-search/gs_neo4j/{conf,data,import,logs,plugins}
# Initialize mariaDB directories
mkdir -p graph-search/gs_mariadb/{entrypoint,data,conf}
# Initialize the export directories for the data pipeline
mkdir -p graph-search/pipeline_products/{bigrams,bigrams_cido,citedby,clean,curator,curator_cido,export,mine,mine_cido,raw,neo4j}
graph-search/gs_mariadb/entrypoint/00_init_tables.sqlCREATE DATABASE IF NOT EXISTS agave;
USE agave;
CREATE TABLE IF NOT EXISTS metadata
(
cord_uid VARCHAR(10) PRIMARY KEY,
sha VARCHAR(5000),
source_x VARCHAR(35),
title VARCHAR(2000),
doi VARCHAR(100),
pmcid VARCHAR(20),
pubmed_id VARCHAR(50),
license VARCHAR(20),
abstract TEXT,
publish_time DATE,
authors TEXT,
journal VARCHAR(600),
mag_id VARCHAR(5),
who_covidence_id VARCHAR(26),
arxiv_id VARCHAR(20),
pdf_json_files TEXT,
pmc_json_files TEXT,
url TEXT,
s2_id VARCHAR(15),
citedby_count INT
);
graph-search/gs_mariadb/conf/custom.cnf[mariadb]
innodb_buffer_pool_size=16GB
docker-compose.yaml to graph-search/docker-compose.yamlversion: '3.1'
services:
mariadb:
image: mariadb:10.8.2
environment:
MYSQL_ROOT_PASSWORD: agave_mariadb_0*
MYSQL_USER: agave
MYSQL_PASSWORD: agave_password
MYSQL_DATABASE: agave
volumes:
- ./gs_mariadb/entrypoint:/docker-entrypoint-initdb.d
- ./gs_mariadb/data:/var/lib/mysql
- ./gs_mariadb/conf:/etc/mysql/conf.d
container_name: graph-search-db
neo4j:
image: neo4j:4.4-community
environment:
NEO4J_AUTH: neo4j/agave
NEO4JLABS_PLUGINS: '["graph-data-science", "apoc"]'
ports:
- "7474:7474"
- "7687:7687"
volumes:
- ./gs_neo4j/logs:/logs
- ./gs_neo4j/data:/data
- ./gs_neo4j/plugins:/plugins
- ./gs_neo4j/import:/var/lib/neo4j/import
- ./gs_neo4j/conf:/conf
container_name: graph-search-neo4j
pipeline:
image: frinve/graph-search:pipeline
volumes:
- ./pipeline_products:/pipeline/products
container_name: graph-search-pipeline
A few additional lines for neo4j might be necessary for some older docker installations. First, identify your username id with
id -u
Add the following lines after image: neo4j:4.4-community
user: 'XXXX:XXXX'
privileged: true
where, with XXXX, we intend your username id found before.
The pipeline is managed with ploomber, a Python package for data pipeline orchestration. Multiple steps compose the data pipeline in the pipeline.yaml config file, which is used by ploomber to orchestrate the execution of the tasks. The pipeline execution is automatically started when the container is run in the following steps.
graph-search directorycd graph-search
docker-compose up -d --rm
docker-compose ps
docker-compose logs pipeline
cp pipeline_products/bigrams.csv gs_neo4j/import/
cp pipeline_products/bigrams_cido.csv gs_neo4j/import/
MATCH (n) SET n :ENTITY RETURN count(n);
CREATE INDEX idx_umls_id IF NOT EXISTS FOR (n:ENTITY) ON (n.umls_id);
CREATE CONSTRAINT ON (n:ENTITY) assert n.umls_id IS UNIQUE;
:auto USING PERIODIC COMMIT
LOAD CSV WITH HEADERS FROM 'file:///bigrams.csv' AS line FIELDTERMINATOR ';'
MATCH (first:ENTITY), (second:ENTITY)
WHERE first.umls_id = line.first_entity_id AND second.umls_id = line.second_entity_id
CREATE (first)-[r:co_occurrence {name: line.bigram_name, frequency: toInteger(line.frequency), cramers_v: toFloat(line.cramers_v), pmi: toFloat(line.pmi), npmi: toFloat(line.npmi)}]->(second);
:auto USING PERIODIC COMMIT
LOAD CSV WITH HEADERS FROM 'file:///bigrams_cido.csv' AS line FIELDTERMINATOR ';'
MATCH (first:ENTITY), (second:ENTITY)
WHERE first.umls_id = line.first_entity_id AND second.umls_id = line.second_entity_id
CREATE (first)-[r:co_occurrence {name: line.bigram_name, frequency: toInteger(line.frequency), cramers_v: toFloat(line.cramers_v), pmi: toFloat(line.pmi), npmi: toFloat(line.npmi)}]->(second);
Now the Semantic Network is ready to be explored. Check the number of nodes:
MATCH (n) RETURN COUNT (n)
Content type
Image
Digest
sha256:68b83ddd7…
Size
942.8 MB
Last updated
about 3 years ago
docker pull frinve/graph-search:pipeline