Sign inSign up

ahujalab/evoage-project

By ahujalab

โ€ขUpdated about 2 months ago

Image
0

438

ahujalab/evoage-project repository overview

โ EvoAGE Project โ€“ Docker Image

This guide explains how to run the evoage-project Docker image. Before running the EvoAGE container, a Neo4j graph database and a Redis server must be installed and properly configured on the host. These are the following steps:

โ ๐Ÿ“š Table of Contents

  1. Neo4j Setupโ 
  2. Redis Setupโ 
  3. Choose Your LLM Backend โ€” Gemini or MedGemmaโ 
  4. Running the EvoAGE Containerโ 
  5. Environment Variable Referenceโ 

โ 1. Neo4j Setup (Required Before Running EvoAGE Backend)

The EvoAGE backend uses Neo4j as the primary graph database. Follow the steps below to start Neo4j, configure it, restore a database from a dump, and enable the APOC plugin.

โ 1.1 Install, Start & Configure Neo4j

Install Neo4j

# Install Java (Neo4j requires Java 17)
sudo apt update
sudo apt install -y openjdk-17-jdk

# Add Neo4j repository
wget -O - https://debian.neo4j.com/neotechnology.gpg.key | sudo apt-key add -
echo "deb https://debian.neo4j.com stable 5" | sudo tee /etc/apt/sources.list.d/neo4j.list

# Install Neo4j
sudo apt install -y neo4j

Check installed Neo4j version

neo4j --version

Set initial password BEFORE first start

sudo neo4j-admin dbms set-initial-password <YOUR_NEO4J_PASSWORD>

Start Neo4j

sudo systemctl start neo4j
sudo systemctl status neo4j

Test login

cypher-shell -u neo4j -p '<YOUR_NEO4J_PASSWORD>' "SHOW DATABASES;"
โ 1.2 Restore Database from a .dump File

Neo4j must be stopped before restoring. You can get the EvoAge_neo4j.dump file from https://zenodo.org/records/17711174โ 

sudo systemctl stop neo4j
sudo cp neo4j.dump /var/lib/neo4j/import/
sudo neo4j-admin database load neo4j \
  --from-path=/var/lib/neo4j/import/ \
  --overwrite-destination=true

Check the graph is built by getting the total node count

cypher-shell -u neo4j -p '<YOUR_NEO4J_PASSWORD>' "MATCH (n) RETURN count(n) AS nodeCount;"

Open the config file

sudo nano /etc/neo4j/neo4j.conf

Add or un-comment these lines:

# Enable APOC Core
dbms.security.procedures.unrestricted=apoc.*
dbms.security.procedures.allowlist=apoc.*

# Allow file imports (optional)
server.directories.import=import

Note on the Bolt port. A default Neo4j install listens on 7687. If you change it (for example to 7333/3333 via server.bolt.listen_address), use that same port in NEO4J_URI below. The container connects over the network, so Neo4j must listen on an address reachable from Docker โ€” set server.bolt.listen_address=0.0.0.0:7687 and use the host's LAN IP (or host.docker.internal) in NEO4J_URI, not localhost.

Start Neo4j after restoration

sudo systemctl enable neo4j
sudo systemctl start neo4j

# This will show the working status of neo4j
sudo systemctl status neo4j
โ 1.3 Install APOC Plugin (Required)

Stop Neo4j before adding plugins

sudo systemctl stop neo4j

Go to the Neo4j plugin directory and check existing plugins

cd /var/lib/neo4j/plugins
ls -l

Download APOC (example for Neo4j 5.x)

sudo wget https://github.com/neo4j/apoc/releases/download/5.26.14/apoc-5.26.14-core.jar

Set correct permissions

sudo chown neo4j:neo4j apoc-5.26.14-core.jar

Enable APOC in neo4j.conf

sudo nano /etc/neo4j/neo4j.conf

Ensure this line exists:

dbms.security.procedures.unrestricted=apoc.*

Restart Neo4j

sudo systemctl restart neo4j

Neo4j + APOC is now ready for the EvoAGE backend! ๐ŸŽ‰


โ 2. Redis Setup (Required Before Running the Container)

The backend uses Redis for caching and job state. The image does bundle a local redis-server, but the recommended setup is a host Redis with a password, which survives container restarts and can be shared across deployments.

The following steps install Redis, enable it as a service, set a password, and verify the setup.

Update and install Redis

sudo apt update
sudo apt install redis-server -y

Enable Redis to start automatically

sudo systemctl enable redis-server

Configure the Redis password

REDIS_PASSWORD="<YOUR_REDIS_PASSWORD>"
sudo sed -i "s/^# requirepass .*/requirepass $REDIS_PASSWORD/" /etc/redis/redis.conf

Allow connections from the container

The container reaches Redis over the Docker bridge, so Redis must not be bound to loopback only. In /etc/redis/redis.conf set:

bind 0.0.0.0
protected-mode no

Only do this on a trusted/private network, and keep requirepass set.

Restart Redis to apply changes

sudo systemctl restart redis-server

Test Redis authentication

redis-cli
127.0.0.1:6379> AUTH default <YOUR_REDIS_PASSWORD>
OK
127.0.0.1:6379> PING
PONG

Check Redis service status

systemctl status redis-server --no-pager

Redis setup completed successfully! ๐ŸŽ‰


โ 3. Choose Your LLM Backend โ€” Gemini or MedGemma

The hypothesis pipeline routes every filter / agent / judge call to a single LLM backend, selected by the USE variable. Pick one before starting the container.

Gemini (hosted)MedGemma (local)
USEgeminimedgemma
Runs onGoogle's serversYour own GPU, via SGLang
NeedsA Gemini API key~60 GB VRAM, an HF read token, extra setup
DataPrompts leave your machinePrompts stay on your machine
Setup effortNone โ€” skip to ยง4โ Complete ยง3.2โ  first
โ 3.1 Option A โ€” Gemini (hosted)

No extra setup. Set these in the run command:

USE=gemini
GEMINI_API_KEY=<YOUR_GEMINI_API_KEY>      # comma-separate several keys to rotate them
GEMINI_MODEL=gemini-2.5-flash-lite

If you use Gemini, skip ยง3.2 and continue directly to ยง4โ .

โ 3.2 Option B โ€” MedGemma (local)

Set these in the run command:

USE=medgemma
MEDGEMMA_BASE_URL=http://host.docker.internal:30001/v1
MEDGEMMA_MODEL=medgemma-27b-local

โš ๏ธ Use host.docker.internal, not localhost, in MEDGEMMA_BASE_URL. The SGLang server runs on the host; localhost inside the container refers to the container itself. This requires --add-host=host.docker.internal:host-gateway, which is already in the run command in ยง4.

โš ๏ธ GEMINI_API_KEY must still be set even in MedGemma mode โ€” the configuration object validates it at startup and the container will not boot without it. Any non-empty value works if you never fall back to Gemini.

โ MedGemma Model Setup (SGLang)

These steps run on the host, not inside the container, using the setup scripts from the EvoAGE project repository. Run them before starting the container.

1. Create the separate SGLang environment

bash scripts/setup.sh --medgemma

2. Download the MedGemma model (~54 GB; you need a Hugging Face read token and accepted model licence)

conda run -n sglang hf download google/medgemma-27b-text-it \
  --local-dir ./scripts/medgemma-27b-local \
  --token <YOUR_HF_READ_TOKEN> \
  --max-workers 4

3. Start the local MedGemma server

bash scripts/setup_medgemma.sh

This launches SGLang in the background on port 30001, bound to 0.0.0.0 so the container can reach it, and logs to scripts/logs/medgemma_port30001.log.

What setup_medgemma.sh does (click to expand)
#!/bin/bash

# Exit on errors
set -e

# ==========================================================
# Environment Setup
# ==========================================================
source ~/miniconda3/etc/profile.d/conda.sh
export TMPDIR=$HOME/tmp
export TEMP=$HOME/tmp
export TMP=$HOME/tmp
export TRITON_CACHE_DIR=$HOME/triton_cache
export TORCHINDUCTOR_CACHE_DIR=$HOME/torch_cache
export HF_HOME=$HOME/hf_cache
export HUGGINGFACE_HUB_CACHE=$HOME/hf_cache

conda activate sglang

export LIBRARY_PATH=/usr/lib/x86_64-linux-gnu:$LIBRARY_PATH
export LD_LIBRARY_PATH=/usr/lib/x86_64-linux-gnu:$LD_LIBRARY_PATH

PORT=30001

# Get script's directory
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
cd "$SCRIPT_DIR"

# Create logs directory inside script's directory
mkdir -p logs

MODEL_PATH="medgemma-27b-local/"
LOG_PREFIX="medgemma"

echo "=========================================================="
echo " Starting server for: $MODEL_PATH"
echo " Port: $PORT"
echo "=========================================================="

# 1. Launch SGLang Server in background
nohup python3 \
  -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --port $PORT \
  --host 0.0.0.0 \
  --mem-fraction-static 0.9 \
  --context-length 32000 \
  --schedule-policy lpm \
  --chunked-prefill-size 2048 \
  > "logs/${LOG_PREFIX}_port${PORT}.log" 2>&1 &

SERVER_PID=$!
echo "Server started with PID: $SERVER_PID. Logging to logs/${LOG_PREFIX}_port${PORT}.log"

4. Verify the server is up before starting the container โ€” model loading takes several minutes:

tail -f scripts/logs/medgemma_port30001.log     # wait for the server-ready line
curl http://localhost:30001/v1/models           # should list medgemma-27b-local

GPU note. SGLang is started with --mem-fraction-static 0.9, so it reserves ~90% of its GPU's memory. The EvoAGE container also loads a DGL-KE model onto the GPU named by DGLKE_DEVICE. On a multi-GPU host, point DGLKE_DEVICE at a different GPU than the one SGLang uses (set CUDA_VISIBLE_DEVICES before running setup_medgemma.sh), or both will contend for memory.


โ 4. Run the EvoAGE Container

Before running the container:

  1. Ensure Redis and Neo4j are running and reachable from Docker.
  2. Retrieve an App Password for Gmail (Google Account โ†’ Security โ†’ App Passwords).
  3. Complete ยง3โ  โ€” if you chose MedGemma, its SGLang server must already be running on the host.
  4. Pull the EvoAGE Docker image.
โ ๐Ÿ“ฅ Pull the EvoAGE Docker Image
docker pull ahujalab/evoage-project:latest
โ ๐Ÿ”Œ Ports

The container runs three processes under supervisord:

ServiceContainer portPurpose
Streamlit frontend8502Web UI
FastAPI backend1027REST API
Redis (bundled)6379Only used if you point REDIS_HOST at the container itself

Both 8502 and 1027 must be published. The frontend calls the backend from the user's browser via API_BASE_URL, so that URL has to be an address the browser can reach โ€” the host's IP or hostname, never localhost (which would resolve inside the container).

โ โ–ถ๏ธ Run the EvoAGE Container
docker rm -f evoage-container 2>/dev/null || true

docker run -d \
    --name evoage-container \
    --gpus all \
    --runtime=nvidia \
    --add-host=host.docker.internal:host-gateway \
    -p 8502:8502 \
    -p 1027:1027 \
    \
    -e API_BASE_URL="http://<YOUR_HOST_IP>:1027" \
    -e API_BASE="http://<YOUR_HOST_IP>:1027" \
    -e FRONTEND_URL="http://<YOUR_HOST_IP>:8502" \
    \
    -e JWT_SECRET_KEY="<YOUR_SECRET_KEY>" \
    -e JWT_ALGORITHM="HS256" \
    -e JWT_ACCESS_TOKEN_EXPIRE_MINUTES="30" \
    \
    -e NEO4J_URI="neo4j://<YOUR_NEO4J_HOST>:7687" \
    -e NEO4J_USERNAME="neo4j" \
    -e NEO4J_PASSWORD="<YOUR_NEO4J_PASSWORD>" \
    \
    -e REDIS_HOST="<YOUR_REDIS_HOST>" \
    -e REDIS_PORT="6379" \
    -e REDIS_USERNAME="default" \
    -e REDIS_PASSWORD="<YOUR_REDIS_PASSWORD>" \
    \
    -e MAIL_USERNAME="<MAIL_USERNAME>" \
    -e MAIL_PASSWORD="<MAIL_APP_PASSWORD>" \
    -e MAIL_FROM="<MAIL_FROM_ADDRESS>" \
    -e MAIL_FROM_NAME="Arushi Sharma [Project Lead - EvoAGE]" \
    -e MAIL_ADMIN_EMAIL="<ADMIN_EMAIL>" \
    -e MAIL_SERVER="smtp.gmail.com" \
    -e MAIL_PORT="587" \
    -e MAIL_STARTTLS="True" \
    -e MAIL_SSL_TLS="False" \
    \
    -e GEMINI_API_KEY="<YOUR_GEMINI_API_KEYS_COMMA_SEPARATED>" \
    -e GEMINI_MODEL="gemini-2.5-flash-lite" \
    -e USE="gemini" \
    \
    -e DGLKE_DEVICE="0" \
    -e DGLKE_SFUNC="logsigmoid" \
    -e DEFAULT_ENTITY_PROP="id" \
    ahujalab/evoage-project:latest

๐Ÿ” Never commit a filled-in version of this command to a public repository โ€” it contains your JWT secret, database passwords, mail app password, and API keys. Prefer --env-file .env with the file kept out of version control.

โ ๐Ÿง  Running with MedGemma instead of Gemini

The command above uses the Gemini backend. To use MedGemma โ€” after completing ยง3.2โ  and confirming the SGLang server is up โ€” replace the single USE line with these three:

    -e USE="medgemma" \
    -e MEDGEMMA_BASE_URL="http://host.docker.internal:30001/v1" \
    -e MEDGEMMA_MODEL="medgemma-27b-local" \

Leave everything else, including GEMINI_API_KEY, unchanged.

โ ๐Ÿ“„ Check Container Logs
sudo docker logs -f evoage-container
โ โณ Notes on Startup Behaviour

The backend performs an initialization routine where the trained model is loaded onto the GPU. This GPU warm-up step takes around 5 minutes, during which the service is not yet ready.

Monitor progress with:

sudo docker logs -f evoage-container

You may begin using the frontend only after the logs show:

Application startup complete.

Then open http://<YOUR_HOST_IP>:8502 in a browser.

โ ๐Ÿ”ฌ Hypothesis Testing (First-Time Delay)
  • The first hypothesis testing request may take 3โ€“5 minutes due to GPU priming.
  • Subsequent requests are optimized and typically complete in under 1 minute.

โ 5. Environment Variable Reference

VariableRequiredDefaultDescription
API_BASE_URLโœ…http://localhost:8000Backend URL used by the Streamlit frontend. Must be browser-reachable.
API_BASEโœ…โ€“Backend URL used internally by the hypothesis pipeline.
FRONTEND_URLโœ…http://localhost:8501Used to build links in verification / notification emails.
JWT_SECRET_KEYโœ…โ€“Signing key for auth tokens. Startup fails if unset.
JWT_ALGORITHMโ€“HS256JWT signing algorithm.
JWT_ACCESS_TOKEN_EXPIRE_MINUTESโ€“30Access-token lifetime.
NEO4J_URIโœ…bolt://localhost:3333Bolt URI of the Neo4j instance.
NEO4J_USERNAMEโœ…โ€“Neo4j user.
NEO4J_PASSWORDโœ…โ€“Neo4j password.
REDIS_HOSTโ€“localhostRedis host. Use the host IP for the external Redis from ยง2.
REDIS_PORTโ€“6379Redis port.
REDIS_USERNAMEโ€“โ€“Usually default.
REDIS_PASSWORDโ€“โ€“Redis password (requirepass).
MAIL_USERNAME / MAIL_PASSWORDโ€“โ€“SMTP credentials; use a Gmail App Password.
MAIL_FROM / MAIL_FROM_NAMEโ€“โ€“Sender identity on outgoing mail.
MAIL_SERVER / MAIL_PORTโ€“587SMTP server and port.
MAIL_STARTTLS / MAIL_SSL_TLSโ€“True / FalseTLS mode for SMTP.
MAIL_ADMIN_EMAILโ€“โ€“Address that receives admin notifications.
GEMINI_API_KEYโœ…โ€“One or more Gemini keys, comma-separated (rotated across calls).
GEMINI_MODELโ€“gemini-1.5-flashGemini model id.
USEโ€“geminiLLM backend for the hypothesis pipeline: gemini or medgemma.
MEDGEMMA_BASE_URLโ€“http://localhost:30001/v1OpenAI-compatible MedGemma endpoint.
MEDGEMMA_MODELโ€“medgemma-27b-localModel name sent to that endpoint.
DGLKE_DEVICEโ€“โ€“GPU index used for DGL-KE inference.
DGLKE_SFUNCโ€“โ€“Scoring function, e.g. logsigmoid.
DEFAULT_ENTITY_PROPโ€“โ€“Neo4j node property used for entity lookup, e.g. id.

โ ๐Ÿ“ Support

For issues, open a ticket in the repository or contact the project maintainers.

Tag summary

Content type

Image

Digest

sha256:77c039ffaโ€ฆ

Size

20.9 GB

Last updated

about 2 months ago

docker pull ahujalab/evoage-project