RunPod Serverless OpenVoice Worker
3.5K
This project allows users to install OpenVoice, an AI speech model that can clone voices, on RunPod serverless platform.
I suggest you use the v2.0.0 version as the older versions have a URL issue when loading the OpenVoice model. v2.0.0 uses the latest model(v2)
To run this application on RunPod serverless, you need to set the following environment variables:
BUCKET_ENDPOINT_URL: The endpoint URL of your S3-compatible storage.BUCKET_ACCESS_KEY_ID: The access key ID for your S3-compatible storage.BUCKET_SECRET_ACCESS_KEY: The secret access key for your S3-compatible storage.These variables are required to store and host the generated WAV files.
Deploy on RunPod
drvpn/runpod_serverless_openvoice_worker:latest for image.BUCKET_ENDPOINT_URL, BUCKET_ACCESS_KEY_ID, BUCKET_SECRET_ACCESS_KEY.Invoke the Function
You can invoke the function with a JSON payload specifying the text, language, and voice URL. Here is an example:
{
"input": {
"text": "Hello, world!",
"voice_url": "https://example.com/path/to/voice.mp3",
"language": "EN-NEWEST",
"speed": 1.0
}
}
Use RunPod's interface or an HTTPS client (i.e. Postman) to send this payload to the deployed function.
text: The text the AI will transcribevoice_url: A URL to a wav file. This file should contain spoken words recorded in a quite environment. This will become the voice of the speaker.language: The language the speaker will use when transcribing your text. Choose on of the following ['EN', 'EN-AU', 'EN-BR', 'EN-INDIA', 'EN-US', 'EN-DEFAULT', 'EN-NEWEST', 'ES', 'FR', 'ZH', 'JP', 'KR']speed: Speed is the pace the speaker will use when speaking.text: required no defaultvoice_url: required no defaultlanguage: default value is EN-NEWESTspeed: default value is 1.0To override default values, you can set the following (optional) environment variables:
DEFAULT_TEXT: sets new default for textDEFAULT_LANGUAGE: sets new default for languageDEFAULT_VOICE_URL: Sets new default for voice_urlDEFAULT_SPEED: Sets new default for speed{
"delayTime": 789,
"executionTime": 16608,
"id": "your-unique-id-will-be-here",
"output": {
"output_audio_url": "https://mybucket.nyc3.digitaloceanspaces.com/OpenVoice/OpenVoice_20240613_213640_i7bzrf_32f210.wav"
},
"status": "COMPLETED"
}
This project is licensed under the MIT License.
Content type
Image
Digest
sha256:85c172240…
Size
8.8 GB
Last updated
about 2 years ago
docker pull drvpn/runpod_serverless_openvoice_worker