vLLM OpenAI API Endpoint Server for OSS LLMs
1.8K
docker run wasertech/vllm-inference-api:latest --help
usage: api_server.py [-h] [--host HOST] [--port PORT] [--allow-credentials] [--allowed-origins ALLOWED_ORIGINS] [--allowed-methods ALLOWED_METHODS]
[--allowed-headers ALLOWED_HEADERS] [--served-model-name SERVED_MODEL_NAME] [--model MODEL] [--tokenizer TOKENIZER]
[--revision REVISION] [--tokenizer-revision TOKENIZER_REVISION] [--tokenizer-mode {auto,slow}] [--trust-remote-code]
[--download-dir DOWNLOAD_DIR] [--load-format {auto,pt,safetensors,npcache,dummy}]
[--dtype {auto,half,float16,bfloat16,float,float32}] [--max-model-len MAX_MODEL_LEN] [--worker-use-ray]
[--pipeline-parallel-size PIPELINE_PARALLEL_SIZE] [--tensor-parallel-size TENSOR_PARALLEL_SIZE] [--block-size {8,16,32}]
[--seed SEED] [--swap-space SWAP_SPACE] [--gpu-memory-utilization GPU_MEMORY_UTILIZATION]
[--max-num-batched-tokens MAX_NUM_BATCHED_TOKENS] [--max-num-seqs MAX_NUM_SEQS] [--max-paddings MAX_PADDINGS] [--disable-log-stats]
[--quantization {awq,squeezellm,None}] [--engine-use-ray] [--disable-log-requests] [--max-log-len MAX_LOG_LEN]
vLLM OpenAI-Compatible RESTful API server.
options:
-h, --help show this help message and exit
--host HOST host name
--port PORT port number
--allow-credentials allow credentials
--allowed-origins ALLOWED_ORIGINS
allowed origins
--allowed-methods ALLOWED_METHODS
allowed methods
--allowed-headers ALLOWED_HEADERS
allowed headers
--served-model-name SERVED_MODEL_NAME
The model name used in the API. If not specified, the model name will be the same as the huggingface name.
--model MODEL name or path of the huggingface model to use
--tokenizer TOKENIZER
name or path of the huggingface tokenizer to use
--revision REVISION the specific model version to use. It can be a branch name, a tag name, or a commit id. If unspecified, will use the default
version.
--tokenizer-revision TOKENIZER_REVISION
the specific tokenizer version to use. It can be a branch name, a tag name, or a commit id. If unspecified, will use the default
version.
--tokenizer-mode {auto,slow}
tokenizer mode. "auto" will use the fast tokenizer if available, and "slow" will always use the slow tokenizer.
--trust-remote-code trust remote code from huggingface
--download-dir DOWNLOAD_DIR
directory to download and load the weights, default to the default cache dir of huggingface
--load-format {auto,pt,safetensors,npcache,dummy}
The format of the model weights to load. "auto" will try to load the weights in the safetensors format and fall back to the
pytorch bin format if safetensors format is not available. "pt" will load the weights in the pytorch bin format. "safetensors"
will load the weights in the safetensors format. "npcache" will load the weights in pytorch format and store a numpy cache to
speed up the loading. "dummy" will initialize the weights with random values, which is mainly for profiling.
--dtype {auto,half,float16,bfloat16,float,float32}
data type for model weights and activations. The "auto" option will use FP16 precision for FP32 and FP16 models, and BF16
precision for BF16 models.
--max-model-len MAX_MODEL_LEN
model context length. If unspecified, will be automatically derived from the model.
--worker-use-ray use Ray for distributed serving, will be automatically set when using more than 1 GPU
--pipeline-parallel-size PIPELINE_PARALLEL_SIZE, -pp PIPELINE_PARALLEL_SIZE
number of pipeline stages
--tensor-parallel-size TENSOR_PARALLEL_SIZE, -tp TENSOR_PARALLEL_SIZE
number of tensor parallel replicas
--block-size {8,16,32}
token block size
--seed SEED random seed
--swap-space SWAP_SPACE
CPU swap space size (GiB) per GPU
--gpu-memory-utilization GPU_MEMORY_UTILIZATION
the percentage of GPU memory to be used forthe model executor
--max-num-batched-tokens MAX_NUM_BATCHED_TOKENS
maximum number of batched tokens per iteration
--max-num-seqs MAX_NUM_SEQS
maximum number of sequences per iteration
--max-paddings MAX_PADDINGS
maximum number of paddings in a batch
--disable-log-stats disable logging statistics
--quantization {awq,squeezellm,None}, -q {awq,squeezellm,None}
Method used to quantize the weights
--engine-use-ray use Ray to start the LLM engine in a separate process as the server process.
--disable-log-requests
disable logging requests
--max-log-len MAX_LOG_LEN
max number of prompt characters or prompt ID numbers being printed in log. Default: unlimited.
ehartford/dolphin-2.2.1-mistral-7bdocker run \
-it \
-p 5085:5085 \
--gpus=all \
--privileged \
--shm-size=8g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--ipc=host \
--mount type=bind,src=`echo ~/.cache/huggingface/hub/`,dst=/root/.cache/huggingface/hub/ \
wasertech/vllm-inference-api:latest --host 0.0.0.0 --port 5085 --model ehartford/dolphin-2.2.1-mistral-7b --tokenizer ehartford/dolphin-2.2.1-mistral-7b --dtype half
Content type
Image
Digest
sha256:a04dbd12d…
Size
12.7 GB
Last updated
over 2 years ago
docker pull wasertech/vllm-inference-api