Sign inSign up

gcareri/asterisk-voicebot-rt

By gcareri

•Updated almost 2 years ago

It's an application that utilizes Asterisk for real-time audio transmission and processing.

Image
Networking
Integration & delivery
Machine learning & AI
2

1.6K

gcareri/asterisk-voicebot-rt repository overview

⚠️ DEPRECATED PROJECT
This project is no longer maintained.
The new official version of this VoiceBot application is now available at agentvoiceresponse.com⁠
➡️ GitHub: github.com/agentvoiceresponse⁠
➕ Docker Hub: hub.docker.com/u/agentvoiceresponse⁠


⁠Asterisk VoiceBot Realtime Application

Asterisk VoiceBot Realtime is a Node.js application that utilizes Asterisk for real-time audio transmission and processing. This application leverages Google Cloud Speech-to-Text, Google Cloud Text-to-Speech, and OpenAI for transcription, speech synthesis, and intelligent response generation.

⁠How It Works

Asterisk VoiceBot Realtime connects to a TCP server and handles audio data packets sent by Asterisk through the AudioSocket module. Upon receiving audio data, the application transcribes it using Google Cloud Speech-to-Text, sends the transcription to OpenAI for generating a response, and finally synthesizes the response into audio using Google Cloud Text-to-Speech. The generated audio is then sent back to the client through the socket.

⁠Example Usage in Asterisk extensions.conf

To use this application, you need to have a SIP client configured on your Asterisk server. You can use a softphone to make a call to extension 570.

exten => 570,1,Answer()
    same => n,AudioSocket(66b7feb0-8938-11ee-abd7-0242ac150002,YOUR_IP:5001)
⁠Making a Call
  1. Configure a SIP client on your Asterisk server.
  2. Use a softphone to call the number 570.

⁠Configuration

Before running the application, ensure to configure the required environment variables.

GOOGLE_APPLICATION_CREDENTIALS=path_to_your_google_credentials.json
OPENAI_API_KEY=your_openai_api_key

Ecco la sezione convertita nel formato Markdown:

⁠Optional Variables (with default values)

If you want to modify the parameters for the Speech Recognition, use these variables:

SPEECH_RECOGNITION_LANGUAGE=en-US
  • SPEECH_RECOGNITION_LANGUAGE (default: en-US): Specifies the language for Google Cloud Speech-to-Text to understand the voice input. Change this to the appropriate language code if needed. For more information about the supported languages for Google Cloud Speech-to-Text, visit the official page here⁠.
SPEECH_RECOGNITION_MODEL=phone_call
  • SPEECH_RECOGNITION_MODEL (default: phone_call): Specifies the model to be used for speech recognition. Different models are optimized for various use cases. For more information, refer to the documentation here⁠.
SPEECH_RECOGNITION_ALTERNATIVE_LANGUAGES=null
  • SPEECH_RECOGNITION_ALTERNATIVE_LANGUAGES (default: null): Allows the addition of alternative languages for speech recognition. This should follow the syntax with languages separated by commas (e.g., it-IT,en-US). Google will prioritize the language configured under SPEECH_RECOGNITION_LANGUAGE and use the alternatives as secondary options. However, based on my experience, I do not recommend this for production environments. For further details on usage, refer to this documentation here⁠.

These configurations provide flexibility in handling various voice input scenarios and can be adjusted based on your application's requirements.

If you want to modify the parameters for the Text to Speech use these variables:

TEXT_TO_SPEECH_LANGUAGE=en-AU
TEXT_TO_SPEECH_GENDER=FEMALE
TEXT_TO_SPEECH_NAME=en-AU-Neural2-C
  • TEXT_TO_SPEECH_LANGUAGE (default: en-AU): Specifies the language for Google Cloud Text-to-Speech to generate the voice output.
  • TEXT_TO_SPEECH_GENDER (default: FEMALE): Specifies the gender of the voice used by Google Cloud Text-to-Speech.
  • TEXT_TO_SPEECH_NAME (default: en-AU-Neural2-C): Specifies the specific voice name used by Google Cloud Text-to-Speech.

For more information about the supported voices for Google Cloud Text-to-Speech, visit the official page here⁠.

If you want to use a general OpenAI Assistant use these variables:

OPENAI_MODEL=gpt-3.5-turbo
SYSTEM_PROMPT="You are a helpful assistant."
  • OPENAI_MODEL (default: gpt-3.5-turbo): The OpenAI model used for generating responses. For more information about the supported OpenAI models, visit the official page here⁠.
  • SYSTEM_PROMPT (default: "You are a helpful assistant."): The system prompt used by OpenAI for generating responses.

But if you want to use a specific agent you have to set the assistant ID that you find in your OpenAI account under the section Accounts OpenAI Assistants⁠:

OPENAI_ASSISTANT_ID=asst_1234
  • OPENAI_ASSISTANT_ID: The specific assistant ID for using a pre-configured OpenAI Assistant.

If you use the assistant, the variables OPENAI_MODEL and SYSTEM_PROMPT will not be considered, but the application will use your OpenAI Assistant Configuration that you created in OpenAI.

This setup allows for flexible configurations depending on your requirements for speech recognition, text-to-speech synthesis, and intelligent response generation.

⁠Docker Setup

To run the application using Docker, you can use the following docker-compose.yml example:

services:
  asterisk-voicebot-rt:
    image: gcareri/asterisk-voicebot-rt
    container_name: asterisk-voicebot-rt
    ports:
      - 5001:5001
    environment:
      - GOOGLE_APPLICATION_CREDENTIALS=/usr/src/app/service-account-key.json
      - OPENAI_API_KEY=sk-proj-1234
      - OPENAI_MODEL=gpt-3.5-turbo
      - SYSTEM_PROMPT="You are a helpful assistant."
      - SPEECH_RECOGNITION_LANGUAGE=en-US
      - TEXT_TO_SPEECH_LANGUAGE=en-AU
      - TEXT_TO_SPEECH_GENDER=FEMALE
      - TEXT_TO_SPEECH_NAME=en-AU-Neural2-C
    volumes:
      - ./my-local-service-account-key.json:/usr/src/app/service-account-key.json
⁠Example Configuration for Italian Language

To configure the application to use Italian for speech recognition and text-to-speech, you can use the following docker-compose.yml example:

services:
  asterisk-voicebot-rt:
    image: gcareri/asterisk-voicebot-rt
    container_name: asterisk-voicebot-rt
    ports:
      - 5002:5001
    environment:
      - GOOGLE_APPLICATION_CREDENTIALS=/usr/src/app/service-account-key.json
      - OPENAI_API_KEY=sk-proj-1234
      - OPENAI_MODEL=gpt-3.5-turbo
      - SYSTEM_PROMPT="Sei un operatore di contact center"
      - SPEECH_RECOGNITION_LANGUAGE=it-IT
      - TEXT_TO_SPEECH_LANGUAGE=it-IT
      - TEXT_TO_SPEECH_GENDER=FEMALE
      - TEXT_TO_SPEECH_NAME=it-IT-Neural2-A
    volumes:
      - ./my-local-service-account-key.json:/usr/src/app/service-account-key.json
⁠Running with Docker Compose
  1. Ensure you have Docker and Docker Compose installed.
  2. Create a my-local-service-account-key.json file with your Google Cloud service account credentials.
  3. Create a docker-compose.yml file with the content provided above.
  4. Run docker-compose up -d to start the application.

If you want to debug your application, run docker-compose logs -f and check if you set the variables correctly:

Server v1.0.2 listening on port 5001

Variables for Speech Recognition:
   SPEECH_RECOGNITION_LANGUAGE: en-US

Variables for OpenAI:
   OPENAI_MODEL: gpt-3.5-turbo
   SYSTEM_PROMPT: You are a helpful assistant.

Variables for Text-To-Speech:
   TEXT_TO_SPEECH_LANGUAGE: en-AU
   TEXT_TO_SPEECH_GENDER: FEMALE
   TEXT_TO_SPEECH_NAME: en-AU-Neural2-C

⁠Contact

For more information, please contact [email protected]⁠.

Tag summary

Content type

Image

Digest

sha256:15e597cdc…

Size

91.3 MB

Last updated

almost 2 years ago

docker pull gcareri/asterisk-voicebot-rt