ScribeCrate: self-hosted Whisper transcription API. OpenAI-compatible, powered by faster-whisper.
100K+
Open-source, self-hosted transcription API.
GitHub: https://github.com/hwdsl2/scribecrateβ
ScribeCrate is a self-hosted speech-to-text API for transcribing audio, generating subtitles, and translating speech into English. It provides OpenAI-compatible transcription and translation endpoints using Whisperβ models, powered by faster-whisperβ . Deploy with Docker on CPU or an NVIDIA GPU.
Previously known as docker-whisper, maintained by hwdsl2β . The Docker image remains
hwdsl2/whisper-server; existing configuration, API endpoints, and persistent data remain compatible.
Features:
POST /v1/audio/transcriptions and POST /v1/audio/translations endpoints for integration with clients that support the OpenAI Whisper API.stream=true to receive transcription segments via Server-Sent Events as they are decoded, without waiting for the entire uploaded file to finish processing.tiny, base, small, medium, large-v3, large-v3-turbo, and more.:cuda image.WHISPER_LOCAL_ONLY.Also available as part of the Self-Hosted AI Stackβ , which deploys a complete self-hosted AI stack with a single command.
π The Self-Hosted AI Builderβs Guideβ is a practical guide to building, securing, and operating your own private AI stack.
Also available:
Use this command to start ScribeCrate:
docker run \
--name whisper \
--restart=always \
-v whisper-data:/var/lib/whisper \
-p 9000:9000 \
-d hwdsl2/whisper-server
If you have an NVIDIA GPU, use the :cuda image for hardware-accelerated inference:
docker run \
--name whisper \
--restart=always \
--gpus=all \
-v whisper-data:/var/lib/whisper \
-p 9000:9000 \
-d hwdsl2/whisper-server:cuda
Requirements: NVIDIA GPU, NVIDIA driverβ 575.57.08+ (Linux) or 576.57+ (Windows), and the NVIDIA Container Toolkitβ installed on the host. The :cuda image is linux/amd64 only.
Important: This image requires at least 700 MB of available RAM for the default base model. Systems with 512 MB or less of RAM are not supported.
Note: For internet-facing deployments, use a reverse proxyβ to add HTTPS. Also replace -p 9000:9000 with -p 127.0.0.1:9000:9000 in the docker run command above, to prevent direct access to the unencrypted port.
The Whisper base model (~145 MB) is downloaded and cached on first start. Check the logs to confirm the server is ready:
docker logs whisper
Once you see "ScribeCrate transcription server is ready", retrieve the API key generated for a fresh install with the persistent volume shown above:
scribecrate_api_key="$(docker exec whisper whisper_manage --getkey)"
Transcribe your first audio file, replacing your_server_ip with your server address and audio.mp3 with your audio file:
curl http://your_server_ip:9000/v1/audio/transcriptions \
-H "Authorization: Bearer $scribecrate_api_key" \
-F [email protected] \
-F model=whisper-1
Response:
{"text": "Your transcribed text appears here."}
Tip: Need a sample audio file to test? Download this English speech sample (WAV, MIT License) from the Azure Samplesβ repository:
curl -L -o sample_speech.wav \
"https://github.com/Azure-Samples/cognitive-services-speech-sdk/raw/master/sampledata/audiofiles/katiesteve.wav"
curl http://your_server_ip:9000/v1/audio/transcriptions \
-H "Authorization: Bearer $scribecrate_api_key" \
-F file=@sample_speech.wav \
-F model=whisper-1
Alternatively, you may set up Whisper without Dockerβ . To learn more about how to use this image, read the sections below.
ScribeCrate Live is a separately deployed server powered by WhisperLiveβ .
| ScribeCrate | ScribeCrate Liveβ | |
|---|---|---|
| Use case | Transcribe complete audio files | Live microphone / real-time audio streaming |
| Protocol | HTTP REST | WebSocket (streaming) + HTTP REST |
| Latency | JSON after processing; segments via SSE | Progressive segment updates |
| Best for | Meeting recordings, uploaded audio | Browser capture, RTSP streams, live captions |
| Image size | ~190 MB (~3.1 GB for :cuda) | ~750 MB (~4.5 GB for :cuda) |
amd64 (x86_64), arm64 (e.g. Raspberry Pi 4/5, AWS Graviton)base model (see model tableβ )WHISPER_LOCAL_ONLY=true with pre-cached models.For GPU acceleration (:cuda image):
:cuda image supports linux/amd64 onlyFor internet-facing deployments, see Using a reverse proxyβ to add HTTPS.
Get the trusted build from the Docker Hub registryβ :
docker pull hwdsl2/whisper-server
For NVIDIA GPU acceleration, pull the :cuda tag instead:
docker pull hwdsl2/whisper-server:cuda
Alternatively, you may download from Quay.ioβ :
docker pull quay.io/hwdsl2/whisper-server
docker image tag quay.io/hwdsl2/whisper-server hwdsl2/whisper-server
CPU images support linux/amd64 and linux/arm64. The :cuda tag supports linux/amd64 only.
All variables are optional. Fresh installs with a mounted /var/lib/whisper volume auto-generate a Bearer token. Existing installs without a key remain open for backward compatibility.
This Docker image uses the following variables, that can be declared in an env file (see exampleβ ):
| Variable | Description | Default |
|---|---|---|
WHISPER_MODEL | Whisper model to use. See model tableβ for options. | base |
WHISPER_LANGUAGE | Default transcription language. BCP-47 code (e.g. en, fr, de, zh, ja) or auto to autodetect. | auto |
WHISPER_PORT | HTTP port for the API (1β65535). | 9000 |
WHISPER_DEVICE | Compute device: cpu, cuda, or auto. Use cuda with the :cuda image for GPU acceleration. auto detects GPU and falls back to CPU. | cpu |
WHISPER_COMPUTE_TYPE | Quantization / compute type. int8 is recommended for CPU; float16 is recommended for CUDA. | int8 (CPU) / float16 (CUDA) |
WHISPER_THREADS | CPU threads for inference. Set to the number of physical cores for best latency. | 2 |
WHISPER_API_KEY | Optional Bearer token. Fresh persistent installs auto-generate one. If set, all API requests must include Authorization: Bearer <key>. Set explicitly empty to disable authentication. | Auto-generated for fresh persistent installs |
WHISPER_LOG_LEVEL | Log level: DEBUG, INFO, WARNING, ERROR, CRITICAL. | INFO |
WHISPER_BEAM | Beam size for transcription and translation decoding. Higher values may improve accuracy at the cost of speed. Use 1 for fastest (greedy) decoding. | 5 |
WHISPER_MAX_REQUEST_BEAM | Maximum beam size allowed for the per-request beam override. Set to 0 to disable this limit. | 10 |
WHISPER_MAX_UPLOAD_MB | Maximum uploaded audio file size in MB. Requests above this limit return HTTP 413. Set to 0 to disable the limit. | 1024 |
WHISPER_LOCAL_ONLY | When set to any non-empty value (e.g. true), disables all HuggingFace model downloads. For offline or air-gapped deployments with pre-cached models. | (not set) |
WHISPER_WORD_TIMESTAMPS | When set to true, enables word-level timestamps globally for all requests. The verbose_json output will include a top-level words array with per-word timing and confidence. Can also be enabled per-request via timestamp_granularities[]=word. | (not set) |
WHISPER_DIARIZATION | Set to true to enable speaker diarization. Identifies who is speaking in each segment. Uses sherpa-onnxβ with pyannote segmentation-3.0 ONNX models (~45 MB, auto-downloaded on first use). Not supported in streaming mode. | (not set) |
WHISPER_DIARIZE_NUM_SPEAKERS | Exact number of speakers (if known). Improves clustering accuracy. Set to -1 or leave unset for auto-detection. | -1 |
WHISPER_DIARIZE_THRESHOLD | Clustering threshold for auto-detection. Lower = more speakers detected, higher = fewer. Ignored when exact speaker count is set. | 0.5 |
WHISPER_DISABLE_USAGE_COUNTS | Set to 1 to disable anonymous aggregate usage counts. | (not set) |
Note: In your env file, you may enclose values in single quotes, e.g. VAR='value'. Do not add spaces around =. If you change WHISPER_PORT, update the -p flag in the docker run command accordingly.
Example using an env file:
cp whisper.env.example whisper.env
# Edit whisper.env with your settings, then:
docker run \
--name whisper \
--restart=always \
-v whisper-data:/var/lib/whisper \
-v ./whisper.env:/whisper.env:ro \
-p 9000:9000 \
-d hwdsl2/whisper-server
The env file is bind-mounted into the container, so changes are picked up on every restart without recreating the container.
--env-filedocker run \
--name whisper \
--restart=always \
-v whisper-data:/var/lib/whisper \
-p 9000:9000 \
--env-file=whisper.env \
-d hwdsl2/whisper-server
cp whisper.env.example whisper.env
# Edit whisper.env as needed, then:
docker compose up -d
docker logs whisper
Example docker-compose.yml (already included):
services:
whisper:
image: hwdsl2/whisper-server
container_name: whisper
restart: always
ports:
- "9000:9000/tcp" # For a host-based reverse proxy, change to "127.0.0.1:9000:9000/tcp"
volumes:
- whisper-data:/var/lib/whisper
- ./whisper.env:/whisper.env:ro
volumes:
whisper-data:
name: whisper-data
Note: For internet-facing deployments, use a reverse proxyβ to add HTTPS. Also change "9000:9000/tcp" to "127.0.0.1:9000:9000/tcp" in docker-compose.yml, to prevent direct access to the unencrypted port.
A separate docker-compose.cuda.yml is provided for GPU deployments:
cp whisper.env.example whisper.env
# Edit whisper.env as needed, then:
docker compose -f docker-compose.cuda.yml up -d
docker logs whisper
Example docker-compose.cuda.yml (already included):
services:
whisper:
image: hwdsl2/whisper-server:cuda
container_name: whisper
restart: always
ports:
- "9000:9000/tcp" # For a host-based reverse proxy, change to "127.0.0.1:9000:9000/tcp"
volumes:
- whisper-data:/var/lib/whisper
- ./whisper.env:/whisper.env:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
volumes:
whisper-data:
name: whisper-data
ScribeCrate provides endpoints compatible with OpenAI's audio transcriptionβ and audio translationβ interfaces. For clients using the OpenAI SDK, configure the base URL and use your ScribeCrate API key:
Authentication: Fresh persistent installs require an API key. Retrieve it for the SDK configuration and curl examples below:
scribecrate_api_key="$(docker exec whisper whisper_manage --getkey)"
export OPENAI_BASE_URL="http://your_server_ip:9000/v1"
export OPENAI_API_KEY="$scribecrate_api_key"
If API key authentication is disabled, omit the Authorization header in curl examples. OpenAI SDK clients still require a nonempty key; set OPENAI_API_KEY=unused.
Speaker diarization, when enabled, is a local sherpa-onnx extension and is not equivalent to OpenAI diarization models. OpenAI-only transcription options such as gpt-4o-transcribe-diarize, response_format=diarized_json, include=logprobs, chunking_strategy, known_speaker_names, and known_speaker_references are not supported and return 400.
POST /v1/audio/transcriptions
Content-Type: multipart/form-data
Parameters:
| Parameter | Type | Required | Description |
|---|---|---|---|
file | file | β | Audio file. Supported formats: mp3, mp4, m4a, wav, webm, ogg, flac and all other formats supported by ffmpeg. |
model | string | β | Pass whisper-1 (value is accepted but the active model is always used). |
language | string | β | BCP-47 language code. Overrides WHISPER_LANGUAGE for this request. |
prompt | string | β | Optional text to guide the model's style or continue a previous segment. |
response_format | string | β | Output format. Default: json. See response formatsβ . Ignored when stream=true. OpenAI-only diarized_json is not supported. |
temperature | float | β | Sampling temperature (0β1). Default: 0. |
stream | boolean | β | Enable SSE streaming. When true, segments are returned as text/event-stream events as they are decoded. Default: false. |
timestamp_granularities[] | array | β | Timestamp granularities to populate. Values: word, segment. When word is included, verbose_json output includes a top-level words array. Default: ["segment"]. |
Local faster-whisper extension: You can set beam to override WHISPER_BEAM for a single transcription or translation request. This is not part of the OpenAI API schema, so do not send it to the hosted OpenAI API or strict OpenAI-compatible gateways. The default per-request cap is 10 (WHISPER_MAX_REQUEST_BEAM); set that variable to 0 to disable the cap. Beam search mainly affects deterministic decoding when temperature=0.
Example:
curl http://your_server_ip:9000/v1/audio/transcriptions \
-H "Authorization: Bearer $scribecrate_api_key" \
-F [email protected] \
-F model=whisper-1 \
-F language=en
With API key authentication:
curl http://your_server_ip:9000/v1/audio/transcriptions \
-H "Authorization: Bearer your_api_key" \
-F [email protected] \
-F model=whisper-1
response_format | Description |
|---|---|
json | {"text": "..."} β default, matches OpenAI's basic response |
text | Plain text, no JSON wrapper |
verbose_json | Full JSON with language, duration, per-segment timestamps, log-probabilities |
srt | SubRip subtitle format (.srt) |
vtt | WebVTT subtitle format (.vtt) |
See Response formatsβ for more details.
POST /v1/audio/translations
Content-Type: multipart/form-data
Translates audio in any language to English text. Compatible with OpenAI's audio translation endpointβ . Accepts the common translation parameters. The output is always in English.
Note: Translation is not supported with English-only (
.en) models. Use a multilingual model (e.g.base,small,large-v3-turbo).
Example:
curl http://your_server_ip:9000/v1/audio/translations \
-H "Authorization: Bearer $scribecrate_api_key" \
-F file=@french_audio.mp3 \
-F model=whisper-1
GET /v1/models
Returns the active model in OpenAI-compatible format.
curl http://your_server_ip:9000/v1/models -H "Authorization: Bearer $scribecrate_api_key"
An interactive Swagger UI is available at:
http://your_server_ip:9000/docs
All server data is stored in the Docker volume (/var/lib/whisper inside the container):
/var/lib/whisper/
βββ models--Systran--faster-whisper-*/ # Cached Whisper model files (downloaded from HuggingFace)
βββ .port # Active port (used by whisper_manage)
βββ .model # Active model name (used by whisper_manage)
βββ .server_addr # Cached server IP (used by whisper_manage)
Back up the Docker volume to preserve downloaded models. Models are large (145 MB β 3 GB) and can take several minutes to download on first start; preserving the volume avoids re-downloading on container recreation.
Tip: The /var/lib/whisper volume uses the same HuggingFace cache layout as docker-whisper-live's /var/lib/whisper-live volume. If you have already downloaded a model with docker-whisper-live, you can bind-mount the same volume directory to avoid re-downloading.
Use whisper_manage inside the running container to inspect and manage the server.
Show server info:
docker exec whisper whisper_manage --showinfo
List available models:
docker exec whisper whisper_manage --listmodels
Pre-download a model:
docker exec whisper whisper_manage --downloadmodel large-v3-turbo
See Switching modelsβ .
See Securing your serverβ .
For internet-facing deployments, place a reverse proxy in front of ScribeCrate to handle HTTPS termination. The server works without HTTPS on a local or trusted network, but HTTPS is recommended when the API endpoint is exposed to the internet.
Use one of the following addresses to reach the ScribeCrate container from your reverse proxy:
whisper:9000 β if your reverse proxy runs as a container in the same Docker network as ScribeCrate (e.g. defined in the same docker-compose.yml).127.0.0.1:9000 β if your reverse proxy runs on the host and port 9000 is published (the default docker-compose.yml publishes it).Example with Caddyβ (Docker imageβ ) (automatic TLS via Let's Encrypt, reverse proxy in the same Docker network):
Caddyfile:
whisper.example.com {
reverse_proxy whisper:9000
}
Example with nginx (reverse proxy on the host):
See Using a reverse proxyβ .
Fresh persistent installs auto-generate a WHISPER_API_KEY. Display it with docker exec whisper whisper_manage --showkey, or use docker exec whisper whisper_manage --getkey in scripts. For existing installs without a key, set WHISPER_API_KEY in your env file to enable authentication.
See Update Docker imageβ .
See Using with other AI servicesβ .
Speaker diarization identifies who is speaking in each transcribed segment. It is a local extension powered by sherpa-onnx.
See Speaker diarizationβ for setup, examples, output formats, and notes.
See Usage countsβ .
See Technical detailsβ .
Note: The software components inside the pre-built image (such as faster-whisper and its dependencies) are under the licenses chosen by their copyright holders. As for any pre-built image usage, it is the image user's responsibility to ensure that any use of this image complies with any relevant licenses for all software contained within.
Copyright (C) 2026 Lin Song
This work is licensed under the MIT Licenseβ .
faster-whisper is Copyright (C) SYSTRAN, and is distributed under the MIT Licenseβ .
ScribeCrate is an independent server using Whisper models and is not affiliated with, endorsed by, or sponsored by OpenAI or SYSTRAN.
Content type
Image
Digest
sha256:184cd6a14β¦
Size
196.9 MB
Last updated
3 days ago
docker pull hwdsl2/whisper-server