Sign inSign up

drvpn/runpod_serverless_openvoice_worker

By drvpn

•Updated about 2 years ago

RunPod Serverless OpenVoice Worker

Image
Machine learning & AI
1

3.5K

drvpn/runpod_serverless_openvoice_worker repository overview

⁠RunPod Serverless OpenVoice Worker

⁠Explore RunPod⁠

This project allows users to install OpenVoice, an AI speech model that can clone voices, on RunPod serverless platform.

⁠Suggested version

I suggest you use the v2.0.0 version as the older versions have a URL issue when loading the OpenVoice model. v2.0.0 uses the latest model(v2)

⁠Environment Variables

To run this application on RunPod serverless, you need to set the following environment variables:

  • BUCKET_ENDPOINT_URL: The endpoint URL of your S3-compatible storage.
  • BUCKET_ACCESS_KEY_ID: The access key ID for your S3-compatible storage.
  • BUCKET_SECRET_ACCESS_KEY: The secret access key for your S3-compatible storage.

These variables are required to store and host the generated WAV files.

⁠Running on RunPod Serverless

  1. Deploy on RunPod

    • Go to RunPod's dashboard and create a new serverless function.
    • Use drvpn/runpod_serverless_openvoice_worker:latest for image.
    • Set the environment variables: BUCKET_ENDPOINT_URL, BUCKET_ACCESS_KEY_ID, BUCKET_SECRET_ACCESS_KEY.
  2. Invoke the Function

You can invoke the function with a JSON payload specifying the text, language, and voice URL. Here is an example:

{
    "input": {
        "text": "Hello, world!",
        "voice_url": "https://example.com/path/to/voice.mp3",
        "language": "EN-NEWEST",
        "speed": 1.0
    }
}

Use RunPod's interface or an HTTPS client (i.e. Postman) to send this payload to the deployed function.

⁠Input

  • text: The text the AI will transcribe
  • voice_url: A URL to a wav file. This file should contain spoken words recorded in a quite environment. This will become the voice of the speaker.
  • language: The language the speaker will use when transcribing your text. Choose on of the following ['EN', 'EN-AU', 'EN-BR', 'EN-INDIA', 'EN-US', 'EN-DEFAULT', 'EN-NEWEST', 'ES', 'FR', 'ZH', 'JP', 'KR']
  • speed: Speed is the pace the speaker will use when speaking.

⁠Default values

  • text: required no default
  • voice_url: required no default
  • language: default value is EN-NEWEST
  • speed: default value is 1.0

To override default values, you can set the following (optional) environment variables:

  • DEFAULT_TEXT: sets new default for text
  • DEFAULT_LANGUAGE: sets new default for language
  • DEFAULT_VOICE_URL: Sets new default for voice_url
  • DEFAULT_SPEED: Sets new default for speed

⁠Sample return value

{
  "delayTime": 789,
  "executionTime": 16608,
  "id": "your-unique-id-will-be-here",
  "output": {
    "output_audio_url": "https://mybucket.nyc3.digitaloceanspaces.com/OpenVoice/OpenVoice_20240613_213640_i7bzrf_32f210.wav"
  },
  "status": "COMPLETED"
}

⁠License

This project is licensed under the MIT License.

⁠Source

Source on GitHub⁠

Tag summary

Content type

Image

Digest

sha256:85c172240…

Size

8.8 GB

Last updated

about 2 years ago

docker pull drvpn/runpod_serverless_openvoice_worker