ElevenLabs

ElevenLabs

Official ElevenLabs Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech and audio processing APIs.

10K+

27 Tools

Packaged by
Requires Configuration
Requires Secrets
Add to Docker Desktop

Version 4.43 or later needs to be installed to add the server automatically

Tools

NameDescription
add_knowledge_base_to_agentAdd a knowledge base to ElevenLabs workspace. Allowed types are epub, pdf, docx, txt, html. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
check_subscriptionCheck the current subscription status. Could be used to measure the usage of the API.
compose_musicConvert a prompt to music and save the output audio file to a given directory. Directory is optional, if not provided, the output file will be saved to $HOME/Desktop. Saves output file to directory (default: $HOME/Desktop). Two models are supported: - music_v2 (default): latest model. Composition plans use a `chunks` array where each chunk is either a `GenerationChunk` (text, duration_ms, positive_styles, negative_styles, context_adherence, optional conditioning_ref + condition_strength) or an `AudioRefChunk` ({song_id, range: {start_ms, end_ms}}) for inpainting. Inpainting also requires the source song to have been stored — call this tool with store_for_inpainting=True or use upload_music_for_inpainting first to get a song_id. - music_v1: legacy model. Composition plans use positive_global_styles, negative_global_styles, sections.
create_agentCreate a conversational AI agent with custom configuration. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
create_composition_planCreate a composition plan for music generation. Usage of this endpoint does not cost any credits but is subject to rate limiting depending on your tier. Composition plans can be used when generating music with the compose_music tool. The returned plan shape depends on model_id: - music_v2 (default): `{"chunks": [GenerationChunk | AudioRefChunk, ...]}`. Each GenerationChunk has `text`, `duration_ms`, `positive_styles`, `negative_styles`, `context_adherence` and optional `conditioning_ref` + `condition_strength`. AudioRefChunks reference a stored song via `song_id` and `range: {start_ms, end_ms}` for inpainting. - music_v1: `{"positive_global_styles": [...], "negative_global_styles": [...], "sections": [...]}`.
create_voice_from_previewAdd a generated voice to the voice library. Uses the voice ID from the `text_to_voice` tool. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
get_agentGet details about a specific conversational AI agent
get_conversationGets conversation with transcript. Returns: conversation details and full transcript. Use when: analyzing completed agent conversations.
get_voiceGet details of a specific voice
isolate_audioIsolate audio from a file. Saves output file to directory (default: $HOME/Desktop). ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
list_agentsList all available conversational AI agents
list_conversationsLists agent conversations. Returns: conversation list with metadata. Use when: asked about conversation history.
list_modelsList all available models
list_phone_numbersList all phone numbers associated with the ElevenLabs account
make_outbound_callMake an outbound call using an ElevenLabs agent. Automatically detects provider type (Twilio or SIP trunk) and uses the appropriate API. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
play_audioPlay an audio file. Supports WAV and MP3 formats.
search_voice_librarySearch for a voice across the entire ElevenLabs voice library.
search_voicesSearch for existing voices, a voice that has already been added to the user's ElevenLabs voice library. Searches in name, description, labels and category.
simulate_conversationSimulate a text conversation between a conversational AI agent and a simulated user. Runs the full conversation and returns the transcript plus analysis. Use this to test agent behaviour, evaluate prompts, and catch failure modes without a live call. The simulated user follows the persona you describe. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
speech_to_speechTransform audio from one voice to another using provided audio files. Saves output file to directory (default: $HOME/Desktop). ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
speech_to_textTranscribe speech from an audio file. When save_transcript_to_file=True: Saves output file to directory (default: $HOME/Desktop). When return_transcript_to_client_directly=True, always returns text directly regardless of output mode. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
text_to_sound_effectsConvert text description of a sound effect to sound effect with a given duration. Saves output file to directory (default: $HOME/Desktop). Duration must be between 0.5 and 5 seconds. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
text_to_speechConvert text to speech with a given voice. Saves output file to directory (default: $HOME/Desktop). Only one of voice_id or voice_name can be provided. If none are provided, the default voice will be used. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
text_to_voiceCreate voice previews from a text prompt. Creates three previews with slight variations. Saves output file to directory (default: $HOME/Desktop). If no text is provided, the tool will auto-generate text. Voice preview files are saved as: voice_design_(generated_voice_id)_(timestamp).mp3 Example file name: voice_design_Ya2J5uIa5Pq14DNPsbC1_20250403_164949.mp3 ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
upload_music_for_inpaintingUpload an existing audio file to ElevenLabs so it can be referenced in music_v2 inpainting workflows. Returns a song_id you can plug into a composition plan's AudioRefChunks (or a generation chunk's conditioning_ref) to edit or extend the track via the compose_music tool. Optionally extracts a composition plan from the uploaded audio so you have a starting point to mutate. Note: this endpoint is gated to enterprise customers with inpainting access.
video_to_musicGenerate background music for one or more video files. Saves output file to directory (default: $HOME/Desktop). The videos are concatenated server-side in the order provided; the generated score targets the combined duration. Constraints: 1-10 videos per call, combined size <= 200 MB, combined duration <= 600 seconds.
voice_cloneCreate an instant voice clone of a voice using provided audio files. ⚠️ COST WARNING: This tool makes an API call to ElevenLabs which may incur costs. Only use when explicitly requested by the user.
Related servers