Sign inSign up

dengcao/vllm-openai

By dengcao

•Updated 12 months ago

vLLM is a fast and easy-to-use library for LLM inference and serving. Originally developed...

Image
Machine learning & AI
3

10K+

dengcao/vllm-openai repository overview

采用vllm最新源码制作的Docker镜像:dengcao/vllm-openai:latest 经测试正常,可放心使用。

使用示例:

支持Qwen3-Reranker-0.6B,Qwen3-Reranker-4B,Qwen3-Reranker-8B等模型。

详见:https://modelscope.cn/models/dengcao/Qwen3-Reranker-0.6B⁠

https://modelscope.cn/models/dengcao/Qwen3-Reranker-4B⁠

https://modelscope.cn/models/dengcao/Qwen3-Reranker-8B⁠

Support models such as Qwen3-Embedding,Qwen3-Reranker.

docker-compose.yaml

services:

Qwen3-Reranker-0.6B:

container_name: Qwen3-Reranker-0.6B

restart: no

image: dengcao/vllm-openai:v0.9.2-dev #采用vllm最新的开发版制作的镜像,经在NVIDIA RTX3060平台主机上测试正常,可放心使用。

ipc: host

volumes:

  - ./models:/models

command: ['--model', '/models/Qwen3-Reranker-0.6B',  '--served-model-name', 'Qwen3-Reranker-0.6B',  '--gpu-memory-utilization', '0.90', '--hf_overrides','{"architectures": ["Qwen3ForSequenceClassification"],"classifier_from_token": ["no", "yes"],"is_original_qwen3_reranker": true}']

ports:

  - 8010:8000

deploy:

  resources:

    reservations:

      devices:

        - driver: nvidia

          count: all

          capabilities: [gpu]

`

将以上代码另存为docker-compose.yaml,执行:docker compose up -d

调用模型API接口:

========================

Docker内的容器APP调用:

API请求地址:http://host.docker.internal:8011/v1/rerank⁠

请求Key:NOT_NEED

模型名称:Qwen3-Reranker-0.6B

========================

Docker外部的APP调用:

API请求地址:http://localhost:8011/v1/rerank⁠

请求Key:NOT_NEED

模型名称:Qwen3-Reranker-0.6B

此方法已经在FastGPT上测试通过,可正常排序。

Tag summary

Content type

Image

Digest

sha256:d8d39b59e…

Size

11.6 GB

Last updated

12 months ago

docker pull dengcao/vllm-openai