Sign inSign up

maharajah/fastflowlm

By maharajah

•Updated 6 months ago

This Docker container builds FastFlowLM that talks directly to the NPU.

Image
Machine learning & AI
0

1.3K

maharajah/fastflowlm repository overview

⁠FastFlowLM - AMD Ryzen AI (NPU) Inference Engine

⁠🚀 Overview

This Docker image provides a high-performance, ready-to-use environment for FastFlowLM, specifically optimized for AMD Ryzen AI (XDNA 2) NPUs (e.g., Strix Halo / Ryzen AI 9 HX 395).

It allows you to offload background AI tasks—such as Document Embedding (RAG), Summarization, and Translation—from your GPU to the NPU, achieving extreme power efficiency and low latency without interrupting your main GPU workloads.


⁠🌟 Key Features
  • Hardware Native: Full support for XDNA 2 architecture via AMD XRT.
  • Optimized Models: Supports FP4/INT4 hardware-quantized models (NPU2/FLM formats).
  • Standard API: Provides an OpenAI-compatible API server.
  • Persistence: Model storage is decoupled from the container for easy management.

⁠🛠️ Quick Start (Usage)

To run the OpenAI-compatible API server on port 52625:

docker run -d --name fastflowlm \
  --device=/dev/accel/accel0:/dev/accel/accel0 \
  --ulimit memlock=-1:-1 \
  -v /your/local/path/to/models:/root/.config/flm \
  -p 52625:52625 \
  maharajah/fastflowlm:latest serve --host 0.0.0.0 --port 52625
⁠Core Requirements:
  1. Hardware: AMD Ryzen AI NPU (Strix Halo / Hawk Point).
  2. Host Driver: XDNA 2 drivers (XRT) must be installed on the host machine.
  3. Memory Lock: --ulimit memlock=-1:-1 is mandatory for DMA buffer allocation.

⁠📦 Model Management

You can interact with the running container to manage your NPU models:

  • List Models: docker exec -it fastflowlm flm list
  • Pull a Model (e.g., Qwen 3.5 9B NPU2): docker exec -it fastflowlm flm pull qwen3.5-9b-npu2
  • Validate Hardware Setup: docker exec -it fastflowlm flm validate

⁠🏗️ Build Info
  • Base Image: Ubuntu 22.04/24.04 (LTS)
  • Driver Support: XRT (Xilinx Runtime) latest
  • Inspiration: Based on the excellent work from hpenedones/fastflowlm-docker⁠.

⁠🇨🇳 中文说明 (Chinese Summary)

这是专为 AMD Strix Halo (Ryzen AI 9 HX 395) 优化的 FastFlowLM NPU 推理镜像。

核心作用: 将文档检索(Embedding)、网页总结、后台翻译等低功耗 AI 任务从 GPU 卸载到 NPU 执行。在不占用显存的情况下,实现“AI 助理”常驻运行。

使用关键:

  1. 设备映射: 必须映射 /dev/accel/accel0。
  2. 内存锁定: 必须设置 ulimit memlock=-1,否则驱动无法申请 DMA 缓冲区。
  3. 路径挂载: 建议将模型目录挂载到极速 SSD(如 RAID0 阵列)以提升加载速度。

Tag summary

Content type

Image

Digest

sha256:5ab6273ed…

Size

135.5 MB

Last updated

6 months ago

docker pull maharajah/fastflowlm