This Docker container builds FastFlowLM that talks directly to the NPU.
1.3K
This Docker image provides a high-performance, ready-to-use environment for FastFlowLM, specifically optimized for AMD Ryzen AI (XDNA 2) NPUs (e.g., Strix Halo / Ryzen AI 9 HX 395).
It allows you to offload background AI tasks—such as Document Embedding (RAG), Summarization, and Translation—from your GPU to the NPU, achieving extreme power efficiency and low latency without interrupting your main GPU workloads.
To run the OpenAI-compatible API server on port 52625:
docker run -d --name fastflowlm \
--device=/dev/accel/accel0:/dev/accel/accel0 \
--ulimit memlock=-1:-1 \
-v /your/local/path/to/models:/root/.config/flm \
-p 52625:52625 \
maharajah/fastflowlm:latest serve --host 0.0.0.0 --port 52625
--ulimit memlock=-1:-1 is mandatory for DMA buffer allocation.You can interact with the running container to manage your NPU models:
docker exec -it fastflowlm flm listdocker exec -it fastflowlm flm pull qwen3.5-9b-npu2docker exec -it fastflowlm flm validate这是专为 AMD Strix Halo (Ryzen AI 9 HX 395) 优化的 FastFlowLM NPU 推理镜像。
核心作用: 将文档检索(Embedding)、网页总结、后台翻译等低功耗 AI 任务从 GPU 卸载到 NPU 执行。在不占用显存的情况下,实现“AI 助理”常驻运行。
使用关键:
/dev/accel/accel0。ulimit memlock=-1,否则驱动无法申请 DMA 缓冲区。Content type
Image
Digest
sha256:5ab6273ed…
Size
135.5 MB
Last updated
6 months ago
docker pull maharajah/fastflowlm