Sign inSign up

jonathanquai/alpha-inference-miner

By jonathanquai

•Updated 3 months ago

Turnkey RTX 5090 LLM server that mines Pearl proof-of-useful-work when idle (int7 Qwen3.6-27B).

Image
0

763

jonathanquai/alpha-inference-miner repository overview

⁠alpha-inference-miner

A turnkey Pearl merge-mining miner in a single container. It serves a chat LLM and mines the Pearl⁠ chain whenever the GPU is idle, and it also merge-mines each chat's prompt (the prefill) - so mining is not limited to idle time, and the same card does useful inference and proof-of-useful-work. The model's weight matmuls are int7 GEMMs in Pearl's mining shape, so a chat request and a mining share come from the same arithmetic.

⚠️ Only the NVIDIA RTX 5090 (sm_120, 32 GB) is supported. The mining kernels and the serve sizing are tuned and validated exclusively for the 5090; other GPUs (including the 5080) are not supported.

⁠Quick start

Requires Docker, the NVIDIA Container Toolkit⁠, and one RTX 5090.

Recommended - run the prebuilt image:

docker run --rm -it --gpus all --ipc=host --shm-size=2g \
  -v "$HOME/.cache/alpha-inference-miner/model:/root/models/qwen36-27b-pearl-e8" \
  jonathanquai/alpha-inference-miner:1.1.0 \
    -a <payout-address> -p <worker-password> [-w <rig-name>] [-d <difficulty>] [-o <pool-host:port>]

Or, from a clone of the source repo⁠ (run.sh just wraps the same docker run):

git clone https://github.com/jdowning100/alpha-inference-miner
cd alpha-inference-miner
./run.sh -a <payout-address> -p <worker-password> [-w <rig-name>] [-d <difficulty>] [-o <pool-host:port>]

The model (~26 GB) is pulled from Hugging Face on first run by the container entrypoint into the mounted volume - with a progress bar - and reused after that. It is not baked into the image, so the image stays small and the download happens once on the host.

Disk: plan for ~52 GB total - the ~25.7 GB image plus the ~26 GB model (the image goes to Docker's storage; the model goes to the mounted volume). Allow ~60 GB free for headroom.

Flags (passed to either command): -a payout address (required), -p worker password (required - gates the chat tunnel and the pool worker), -w rig name (default: hostname), -d static pool difficulty (default 500000; -d 0 for vardiff), -o pool host:port (default us1.alphapool.tech:5566), --no-chat to disable the chat tunnel. Each flag also has a RELAY_* env-var equivalent - see GPU access & cloud⁠ below.

Difficulty is static 500000 by default (tuned for a 5090 - this miner does not use vardiff). Override with -d <n> (or RELAY_DIFFICULTY), or -d 0 to fall back to the pool's vardiff. Advanced: you can instead put it straight in the password (-p "mysecret;d=100000") - the pool reads the ;d=NNN, and the chat tunnel uses only the secret part before the ;, so the difficulty never reaches chat.

⁠GPU access & cloud (vast.ai, RunPod)

A container can't see the GPU on its own. The NVIDIA Container Toolkit⁠ is the host package that lets docker run --gpus all expose the GPU - and the host's NVIDIA driver - to the container. Install it on whatever machine runs Docker, then --gpus all works. CUDA itself is bundled inside this image, so the host only needs the NVIDIA driver (R570 or newer for the RTX 5090) plus the toolkit, not a CUDA install.

On managed container hosts (vast.ai, RunPod, Lambda, ...) the driver + toolkit are already set up and the GPU is passed into your image for you - you don't install anything or pass --gpus all. Just select this image and configure it with environment variables instead of flags (the entrypoint accepts either):

env var= flagmeaning
RELAY_ADDRESS-ayour prl1... payout address (required)
RELAY_PASSWORD-psecret; gates chat + the pool worker
RELAY_WORKER-wrig / worker name (default: hostname)
RELAY_DIFFICULTY-dstatic pool difficulty (default 500000; 0 = vardiff)
RELAY_POOL-opool host:port (default us1.alphapool.tech:5566)

On vast.ai specifically: pick an RTX 5090 offer whose host driver is R570 or newer; give the instance ~55+ GB disk (the ~25.7 GB image + the ~26 GB model, downloaded on first run - use a persistent volume so the model isn't re-fetched if the instance is recreated); and make sure /dev/shm is a couple GB (add --ipc=host / --shm-size=2g in the Docker options) so vLLM has shared memory. The relay and chat tunnel are outbound-only, so they work behind the host's NAT with no port mapping.

⁠What it does

One entrypoint supervises the whole stack - pool connection, share submission, the lazy-load vLLM serve, and (optionally) a chat tunnel - and keeps the card mining the entire time it is up:

statewhat minesrate (RTX 5090)
idle > 90 s (model asleep)random 131072²~385 TMAC/s
between requests (model loaded)16384² co-resident~240 TMAC/s
chat prefillthe prompt matmul (merge-mined)~130 TMAC/s
chat decode (default)-0 mining, ~50 tok/s

When idle the model sleeps and the card mines the full profile; between requests it stays loaded and mines a smaller co-resident profile; a chat prefill merge-mines and decode streams tokens. After 90 s idle it re-sleeps to the full profile. The card never stops mining while the container is up. Context window: 102400 tokens (text); the chat endpoint is OpenAI-compatible at http://localhost:8000/v1.

⁠Remote chat (tunnel)

By default the container opens an outbound-only chat tunnel so you can talk to your miner's model from anywhere, with no port-forwarding or public IP. The miner dials out to alphapool's tunnel broker (wss://alphapool.tech/tunnel-ws), which proxies chat requests back to the local vLLM serve on :8000. Open the chat UI at pearl.alphapool.tech/chat⁠ and enter your worker name + password to reach your model.

Access is gated by the worker password you pass with -p (use a real secret, not x). Pass --no-chat to disable the tunnel; the local http://localhost:8000/v1 endpoint still works either way.

⁠Configuration

Opt-in knobs (set with -e on docker run; defaults shown):

  • PEARL_LAZY_AWAKE_IDLE_MN=16384 - co-resident "awake-idle" profile (0 disables).
  • PEARL_LAZY_IDLE_TIMEOUT_S=90 - awake→sleep idle timeout (seconds).
  • PEARL_VLLM_IDLE_MINE_VRAM_RESERVE_BYTES=134217728 - miner VRAM floor (128 MiB).
  • PEARL_LAZY_MINE_DURING_DECODE=0 - set 1 to also mine during decode: +~100-130 TMAC/s of shares while the response streams, at the cost of decode speed (~50 → ~34 tok/s). For operators who value always-on hashrate over chat speed.
  • PEARL_VLLM_VISION=0 - set 1 (or pass --vision) to enable image input via the mmproj/vision encoder. The encoder reserves activation VRAM + a larger cudagraph, so the context drops from 102400 to 16384. Default off = maximum text context.
  • PREFILL_MINE_PAD_M=8192 - prompt rows the prefill is padded to for merge-mining (on by default; ~130 TMAC/s). 0 disables prefill mining.

Launcher overrides for run.sh:

  • ALPHA_MINER_IMAGE - image ref (default jonathanquai/alpha-inference-miner:latest).
  • ALPHA_MINER_MODEL_DIR - host model cache (default ~/.cache/alpha-inference-miner/model).
  • ALPHA_MINER_DETACH=1 - run detached instead of foreground.

⁠Model

dominant-strategies/Qwen3.6-27B-heretic-pearl⁠

  • an int7-quantized Qwen3.6-27B. Quantizing the weight matmuls to int7 is what makes inference be Pearl proof-of-useful-work; the model card has the full quantization scheme.

⁠Dev fee

The inference miner contributes 1% of its mining time - 10 s of every 1000 s - to the project's dev address. It uses a brief separate pool connection for that window; the operator's connection and shares are otherwise untouched.

⁠Versioning

Releases are pinned with matching tags across GitHub and Docker Hub:

  • Git tag vMAJOR.MINOR.PATCH (e.g. v1.1.0) on the source commit.
  • Docker image tag MAJOR.MINOR.PATCH (e.g. 1.1.0) - the same bytes as latest at release time. cu128-sm120 is the build tag.

Pin a version for reproducibility (jonathanquai/alpha-inference-miner:1.1.0); use latest to always pull the newest build. Current release: 1.1.0 (adds the -d static-difficulty flag with a 500000 default).

⁠What's in this repo

  • run.sh - the one-command launcher (pull + run, with the first-run model download).
  • deploy/ - the Dockerfile and build.sh used to assemble the image.
  • src/ - the Python the image runs: the vLLM mining plugin (vllm-miner), the Pearl gateway (pearl-gateway), the reference mining base (miner-base), shared utilities (miner-utils), and the launchers.

The GPU kernels (the fused mining/inference GEMM and the in-context solvers) ship as compiled artifacts inside the published image and are intentionally not in this repository. Building the image yourself therefore requires those artifacts placed in deploy/dist/ (see deploy/build.sh); most users should just use the prebuilt image via run.sh.

Tag summary

Content type

Image

Digest

sha256:f1277fceb…

Size

7.8 GB

Last updated

3 months ago

docker pull jonathanquai/alpha-inference-miner