Turnkey RTX 5090 LLM server that mines Pearl proof-of-useful-work when idle (int7 Qwen3.6-27B).
763
A turnkey Pearl merge-mining miner in a single container. It serves a chat LLM and mines the Pearl chain whenever the GPU is idle, and it also merge-mines each chat's prompt (the prefill) - so mining is not limited to idle time, and the same card does useful inference and proof-of-useful-work. The model's weight matmuls are int7 GEMMs in Pearl's mining shape, so a chat request and a mining share come from the same arithmetic.
⚠️ Only the NVIDIA RTX 5090 (sm_120, 32 GB) is supported. The mining kernels and the serve sizing are tuned and validated exclusively for the 5090; other GPUs (including the 5080) are not supported.
Requires Docker, the NVIDIA Container Toolkit, and one RTX 5090.
Recommended - run the prebuilt image:
docker run --rm -it --gpus all --ipc=host --shm-size=2g \
-v "$HOME/.cache/alpha-inference-miner/model:/root/models/qwen36-27b-pearl-e8" \
jonathanquai/alpha-inference-miner:1.1.0 \
-a <payout-address> -p <worker-password> [-w <rig-name>] [-d <difficulty>] [-o <pool-host:port>]
Or, from a clone of the source repo
(run.sh just wraps the same docker run):
git clone https://github.com/jdowning100/alpha-inference-miner
cd alpha-inference-miner
./run.sh -a <payout-address> -p <worker-password> [-w <rig-name>] [-d <difficulty>] [-o <pool-host:port>]
The model (~26 GB) is pulled from Hugging Face on first run by the container entrypoint into the mounted volume - with a progress bar - and reused after that. It is not baked into the image, so the image stays small and the download happens once on the host.
Disk: plan for ~52 GB total - the ~25.7 GB image plus the ~26 GB model (the image goes to Docker's storage; the model goes to the mounted volume). Allow ~60 GB free for headroom.
Flags (passed to either command): -a payout address (required), -p worker
password (required - gates the chat tunnel and the pool worker), -w rig name
(default: hostname), -d static pool difficulty (default 500000; -d 0 for
vardiff), -o pool host:port (default us1.alphapool.tech:5566), --no-chat to
disable the chat tunnel. Each flag also has a RELAY_* env-var equivalent - see
GPU access & cloud below.
Difficulty is static 500000 by default (tuned for a 5090 - this miner does
not use vardiff). Override with -d <n> (or RELAY_DIFFICULTY), or -d 0 to fall
back to the pool's vardiff. Advanced: you can instead put it straight in the password
(-p "mysecret;d=100000") - the pool reads the ;d=NNN, and the chat tunnel uses only
the secret part before the ;, so the difficulty never reaches chat.
A container can't see the GPU on its own. The
NVIDIA Container Toolkit
is the host package that lets docker run --gpus all expose the GPU - and the host's
NVIDIA driver - to the container. Install it on whatever machine runs Docker, then
--gpus all works. CUDA itself is bundled inside this image, so the host only needs
the NVIDIA driver (R570 or newer for the RTX 5090) plus the toolkit, not a CUDA install.
On managed container hosts (vast.ai, RunPod, Lambda, ...) the driver + toolkit are
already set up and the GPU is passed into your image for you - you don't install
anything or pass --gpus all. Just select this image and configure it with
environment variables instead of flags (the entrypoint accepts either):
| env var | = flag | meaning |
|---|---|---|
RELAY_ADDRESS | -a | your prl1... payout address (required) |
RELAY_PASSWORD | -p | secret; gates chat + the pool worker |
RELAY_WORKER | -w | rig / worker name (default: hostname) |
RELAY_DIFFICULTY | -d | static pool difficulty (default 500000; 0 = vardiff) |
RELAY_POOL | -o | pool host:port (default us1.alphapool.tech:5566) |
On vast.ai specifically: pick an RTX 5090 offer whose host driver is R570 or
newer; give the instance ~55+ GB disk (the ~25.7 GB image + the ~26 GB model,
downloaded on first run - use a persistent volume so the model isn't re-fetched if the
instance is recreated); and make sure /dev/shm is a couple GB (add --ipc=host /
--shm-size=2g in the Docker options) so vLLM has shared memory. The relay and chat
tunnel are outbound-only, so they work behind the host's NAT with no port mapping.
One entrypoint supervises the whole stack - pool connection, share submission, the lazy-load vLLM serve, and (optionally) a chat tunnel - and keeps the card mining the entire time it is up:
| state | what mines | rate (RTX 5090) |
|---|---|---|
| idle > 90 s (model asleep) | random 131072² | ~385 TMAC/s |
| between requests (model loaded) | 16384² co-resident | ~240 TMAC/s |
| chat prefill | the prompt matmul (merge-mined) | ~130 TMAC/s |
| chat decode (default) | - | 0 mining, ~50 tok/s |
When idle the model sleeps and the card mines the full profile; between requests it
stays loaded and mines a smaller co-resident profile; a chat prefill merge-mines and
decode streams tokens. After 90 s idle it re-sleeps to the full profile. The card
never stops mining while the container is up. Context window: 102400 tokens
(text); the chat endpoint is OpenAI-compatible at http://localhost:8000/v1.
By default the container opens an outbound-only chat tunnel so you can talk to
your miner's model from anywhere, with no port-forwarding or public IP. The miner
dials out to alphapool's tunnel broker (wss://alphapool.tech/tunnel-ws), which
proxies chat requests back to the local vLLM serve on :8000. Open the chat UI at
pearl.alphapool.tech/chat and enter your
worker name + password to reach your model.
Access is gated by the worker password you pass with -p (use a real secret, not
x). Pass --no-chat to disable the tunnel; the local http://localhost:8000/v1
endpoint still works either way.
Opt-in knobs (set with -e on docker run; defaults shown):
PEARL_LAZY_AWAKE_IDLE_MN=16384 - co-resident "awake-idle" profile (0 disables).PEARL_LAZY_IDLE_TIMEOUT_S=90 - awake→sleep idle timeout (seconds).PEARL_VLLM_IDLE_MINE_VRAM_RESERVE_BYTES=134217728 - miner VRAM floor (128 MiB).PEARL_LAZY_MINE_DURING_DECODE=0 - set 1 to also mine during decode:
+~100-130 TMAC/s of shares while the response streams, at the cost of decode
speed (~50 → ~34 tok/s). For operators who value always-on hashrate over chat speed.PEARL_VLLM_VISION=0 - set 1 (or pass --vision) to enable image input via the
mmproj/vision encoder. The encoder reserves activation VRAM + a larger cudagraph, so
the context drops from 102400 to 16384. Default off = maximum text context.PREFILL_MINE_PAD_M=8192 - prompt rows the prefill is padded to for merge-mining
(on by default; ~130 TMAC/s). 0 disables prefill mining.Launcher overrides for run.sh:
ALPHA_MINER_IMAGE - image ref (default jonathanquai/alpha-inference-miner:latest).ALPHA_MINER_MODEL_DIR - host model cache (default ~/.cache/alpha-inference-miner/model).ALPHA_MINER_DETACH=1 - run detached instead of foreground.dominant-strategies/Qwen3.6-27B-heretic-pearl
The inference miner contributes 1% of its mining time - 10 s of every 1000 s - to the project's dev address. It uses a brief separate pool connection for that window; the operator's connection and shares are otherwise untouched.
Releases are pinned with matching tags across GitHub and Docker Hub:
vMAJOR.MINOR.PATCH (e.g. v1.1.0) on the source commit.MAJOR.MINOR.PATCH (e.g. 1.1.0) - the same bytes as latest at
release time. cu128-sm120 is the build tag.Pin a version for reproducibility (jonathanquai/alpha-inference-miner:1.1.0); use
latest to always pull the newest build. Current release: 1.1.0 (adds the -d
static-difficulty flag with a 500000 default).
run.sh - the one-command launcher (pull + run, with the first-run model download).deploy/ - the Dockerfile and build.sh used to assemble the image.src/ - the Python the image runs: the vLLM mining plugin (vllm-miner), the
Pearl gateway (pearl-gateway), the reference mining base (miner-base), shared
utilities (miner-utils), and the launchers.The GPU kernels (the fused mining/inference GEMM and the in-context solvers) ship as
compiled artifacts inside the published image and are intentionally not in this
repository. Building the image yourself therefore requires those artifacts placed in
deploy/dist/ (see deploy/build.sh); most users should just use the prebuilt image
via run.sh.
Content type
Image
Digest
sha256:f1277fceb…
Size
7.8 GB
Last updated
3 months ago
docker pull jonathanquai/alpha-inference-miner