AMD Inference server
765
The AMD Inference Server is an open-source tool to deploy your machine learning models and make them accessible to clients for inference. Out-of-the-box, the server can support selected models that run on AMD CPUs, GPUs or FPGAs by leveraging existing libraries. For all these models and hardware accelerators, the server presents a common user interface based on community standards so clients can make requests to any using the same API. The server provides HTTP/REST and gRPC interfaces for clients to submit requests. For both, there are C++ and Python bindings to simplify writing client programs. You can also use the server backend directly using the native C++ API to write local applications.
To use the inference server, you need a Docker image for it, which you can get by using a prebuilt image or building one from the inference server repository on GitHub.
You can pull the appropriate deployment Docker image(s) from DockerHub using:
docker pull amdih/serve:uif1.2_zendnn_amdinfer_0.4.0
docker pull amdih/serve:uif1.2_migraphx_amdinfer_0.4.0
docker pull amdih/serve:uif1.2_vai_amdinfer_0.4.0
The instructions provided here are an overview, but you can see more complete information about the AMD Inference Server in the documentation.
Content type
Image
Digest
sha256:420c5c754…
Size
132.1 MB
Last updated
about 3 years ago
docker pull amdih/serve:uif1.2_vai_amdinfer_0.4.0