Sign inSign up

quantatrisk/funasr-api

By quantatrisk

•Updated 6 months ago

Image
0

10K+

quantatrisk/funasr-api repository overview

⁠FunASR-API

Production-ready local speech recognition API service powered by FunASR⁠ and Qwen3-ASR⁠.

⁠Quick Start

docker run -d --name funasr-api \
  --gpus all \
  -p 17003:8000 \
  -e ENABLED_MODELS=auto \
  -e API_KEY=your_api_key \
  -v ./models/modelscope:/root/.cache/modelscope \
  -v ./models/huggingface:/root/.cache/huggingface \
  -v ./logs:/app/logs \
  -v ./temp:/app/temp \
  quantatrisk/funasr-api:gpu-latest
⁠CPU Version
docker run -d --name funasr-api \
  -p 17003:8000 \
  -e ENABLED_MODELS=paraformer-large \
  -e API_KEY=your_api_key \
  -v ./models/modelscope:/root/.cache/modelscope \
  -v ./logs:/app/logs \
  -v ./temp:/app/temp \
  quantatrisk/funasr-api:cpu-latest

⁠Supported Tags

TagDescription
gpu-latestGPU version with CUDA 12.6, auto model selection
cpu-latestCPU-only version, Paraformer model only

⁠Features

  • Multi-Model Support: Qwen3-ASR (1.7B/0.6B) + Paraformer Large
  • OpenAI API Compatible: /v1/audio/transcriptions endpoint
  • Alibaba Cloud Compatible: RESTful and WebSocket streaming API
  • Speaker Diarization: Automatic multi-speaker identification
  • Word-Level Timestamps: Qwen3-ASR supports precise timestamps
  • Smart Far-Field Filtering: Reduces ambient noise in streaming
  • GPU Batch Processing: 2-3x faster with batch inference

⁠Environment Variables

VariableDefaultDescription
ENABLED_MODELSautoModels to load: auto, all, or comma-separated list
API_KEY-API authentication key (optional)
LOG_LEVELINFOLog level: DEBUG, INFO, WARNING, ERROR
MAX_AUDIO_SIZE2048Max audio file size in MB
ASR_BATCH_SIZE4Batch size for inference (GPU: 4, CPU: 2)
MAX_SEGMENT_SEC90Max audio segment duration in seconds
DEVICEautoDevice: auto, cpu, cuda:0

⁠Auto Mode Behavior

  • VRAM >= 32GB: Auto-load qwen3-asr-1.7b + paraformer-large
  • VRAM < 32GB: Auto-load qwen3-asr-0.6b + paraformer-large
  • No CUDA: Only paraformer-large (Qwen3 requires GPU)

⁠API Endpoints

  • OpenAI Compatible: POST /v1/audio/transcriptions
  • Alibaba Cloud: POST /stream/v1/asr
  • WebSocket: /ws/v1/asr, /ws/v1/asr/qwen
  • Health Check: GET /stream/v1/asr/health
  • API Docs: http://localhost:17003/docs

⁠Quick Test

# Health check
curl http://localhost:17003/stream/v1/asr/health

# Transcription with OpenAI API
curl -X POST "http://localhost:17003/v1/audio/transcriptions" \
  -H "Authorization: Bearer your_api_key" \
  -F "[email protected]" \
  -F "model=qwen3-asr-1.7b" \
  -F "response_format=verbose_json"

# Transcription with Alibaba Cloud API
curl -X POST "http://localhost:17003/stream/v1/asr" \
  -H "Content-Type: application/octet-stream" \
  --data-binary @audio.wav

⁠Model Storage

Models are cached in Docker volumes for offline use:

# ModelScope models (Paraformer, VAD, CAM++)
./models/modelscope:/root/.cache/modelscope

# HuggingFace models (Qwen3-ASR, GPU only)
./models/huggingface:/root/.cache/huggingface

First run downloads models automatically. Pre-download for offline deployment:

# Use the helper script (recommended)
./scripts/prepare-models.sh

# Or manually with Docker
docker run --rm \
  -v ./models/modelscope:/root/.cache/modelscope \
  -v ./models/huggingface:/root/.cache/huggingface \
  quantatrisk/funasr-api:gpu-latest \
  python -c "from app.utils.download_models import download_models; download_models()"

⁠Resource Requirements

Minimum (CPU):

  • CPU: 4 cores
  • Memory: 16GB
  • Disk: 20GB

Recommended (GPU):

  • CPU: 4 cores
  • Memory: 16GB
  • GPU: NVIDIA GPU (16GB+ VRAM)
  • Disk: 20GB

⁠License

MIT License

Tag summary

Content type

Image

Digest

sha256:f72394f10…

Size

8.6 GB

Last updated

6 months ago

docker pull quantatrisk/funasr-api