Production-ready local speech recognition API service powered by FunASR and Qwen3-ASR.
docker run -d --name funasr-api \
--gpus all \
-p 17003:8000 \
-e ENABLED_MODELS=auto \
-e API_KEY=your_api_key \
-v ./models/modelscope:/root/.cache/modelscope \
-v ./models/huggingface:/root/.cache/huggingface \
-v ./logs:/app/logs \
-v ./temp:/app/temp \
quantatrisk/funasr-api:gpu-latest
docker run -d --name funasr-api \
-p 17003:8000 \
-e ENABLED_MODELS=paraformer-large \
-e API_KEY=your_api_key \
-v ./models/modelscope:/root/.cache/modelscope \
-v ./logs:/app/logs \
-v ./temp:/app/temp \
quantatrisk/funasr-api:cpu-latest
| Tag | Description |
|---|---|
gpu-latest | GPU version with CUDA 12.6, auto model selection |
cpu-latest | CPU-only version, Paraformer model only |
/v1/audio/transcriptions endpoint| Variable | Default | Description |
|---|---|---|
ENABLED_MODELS | auto | Models to load: auto, all, or comma-separated list |
API_KEY | - | API authentication key (optional) |
LOG_LEVEL | INFO | Log level: DEBUG, INFO, WARNING, ERROR |
MAX_AUDIO_SIZE | 2048 | Max audio file size in MB |
ASR_BATCH_SIZE | 4 | Batch size for inference (GPU: 4, CPU: 2) |
MAX_SEGMENT_SEC | 90 | Max audio segment duration in seconds |
DEVICE | auto | Device: auto, cpu, cuda:0 |
qwen3-asr-1.7b + paraformer-largeqwen3-asr-0.6b + paraformer-largeparaformer-large (Qwen3 requires GPU)POST /v1/audio/transcriptionsPOST /stream/v1/asr/ws/v1/asr, /ws/v1/asr/qwenGET /stream/v1/asr/healthhttp://localhost:17003/docs# Health check
curl http://localhost:17003/stream/v1/asr/health
# Transcription with OpenAI API
curl -X POST "http://localhost:17003/v1/audio/transcriptions" \
-H "Authorization: Bearer your_api_key" \
-F "[email protected]" \
-F "model=qwen3-asr-1.7b" \
-F "response_format=verbose_json"
# Transcription with Alibaba Cloud API
curl -X POST "http://localhost:17003/stream/v1/asr" \
-H "Content-Type: application/octet-stream" \
--data-binary @audio.wav
Models are cached in Docker volumes for offline use:
# ModelScope models (Paraformer, VAD, CAM++)
./models/modelscope:/root/.cache/modelscope
# HuggingFace models (Qwen3-ASR, GPU only)
./models/huggingface:/root/.cache/huggingface
First run downloads models automatically. Pre-download for offline deployment:
# Use the helper script (recommended)
./scripts/prepare-models.sh
# Or manually with Docker
docker run --rm \
-v ./models/modelscope:/root/.cache/modelscope \
-v ./models/huggingface:/root/.cache/huggingface \
quantatrisk/funasr-api:gpu-latest \
python -c "from app.utils.download_models import download_models; download_models()"
Minimum (CPU):
Recommended (GPU):
MIT License
Content type
Image
Digest
sha256:f72394f10…
Size
8.6 GB
Last updated
6 months ago
docker pull quantatrisk/funasr-api