Sign inSign up

adimyth/serverless-stt-deployment

By adimyth

•Updated about 2 years ago

Deploying AI4Bharat ASR using runpod. Refer https://github.com/adimyth/serverless-s2t-indicwav2vec

Image
Machine learning & AI
0

959

adimyth/serverless-stt-deployment repository overview

⁠Deploying AI4Bharat S2T Models on RunPod

⁠Model Information

The models used in this project come from AI4Bharat's IndicWav2Vec⁠ project. These models are designed for Automatic Speech Recognition (ASR) for Indic languages.

To use these models with the Hugging Face transformers "automatic-speech-recognition" pipeline, additional steps were required to convert the ASR models into a HuggingFace-compatible format. An iPython notebook⁠ detailing this conversion process is available in this repository.

⁠Building the Docker Image

Before running the project, you need to build a Docker image that includes the necessary dependencies and the pre-downloaded model.

  1. Ensure you have Docker installed on your system.

  2. Set your HuggingFace API key as an environment variable:

    export HF_API_KEY=your_huggingface_api_key
    
  3. Build the Docker image:

      docker build --build-arg HF_API_KEY=$HF_API_KEY -t ai4bharat-s2t-runpod .
    

This command builds the Docker image with the tag ai4bharat-s2t-runpod. The --build-arg flag passes your HuggingFace API key to the build process, allowing it to download the private model during the build.

This is needed because I had pushed the weights to my HF account.

Why build a custom Docker image?

  • It includes all necessary dependencies.
  • The model is pre-downloaded, reducing startup time when deploying to serverless GPUs.
  • It ensures consistency across all deployments.
  • It optimizes for the RunPod environment.
⁠Running the project locally
  1. Clone the repository
  2. Install the requirements
pip install -r builder/requirements.txt
  1. Run the project. Refer the docs⁠ for more options. This will start the FastAPI server on the specified host at port 8000.
python3 src/handler.py --rp_serve_api --rp_api_host 0.0.0.0 --rp_log_level DEBUG
  1. Test the project
curl --location 'http://0.0.0.0:8000/runsync' \
--header 'accept: application/json' \
--header 'Content-Type: application/json' \
--data '{"audioURL": "https://www.tuttlepublishing.com/content/docs/9780804844383/06-18%20Part2%20Car%20Trouble.mp3", "language": "hi"}'
⁠Running with Docker

After building the image, you can run the container:

docker run -p 8000:8000 ai4bharat-s2t-runpod

This command runs the container and maps port 8000 from the container to port 8000 on your host machine.

Note

The weights were openly available. I just pushed the weights to my HF account and used the HF API as well as the pipeline to load & infer the model making it a whole lot easier

Tag summary

Content type

Image

Digest

sha256:386db936a…

Size

7.1 GB

Last updated

about 2 years ago

docker pull adimyth/serverless-stt-deployment:v1.5.0