ElevenLabs capabilities
15 mapped capabilities, each graded and dated. The map shows what ElevenLabs can do; the audit shows whether it’s worth consolidating — and a guide shows how to move.
Capabilities
AI Dubbing
provisionalverified 25 days agoTranslates and re-voices audio or video content into 90+ languages, preserving speaker identity. Offers automatic dubbing (Dubbing v2, currently in alpha) and manual Dubbing Studio for fine-grained editing.
AI Music Generation
provisionalverified 25 days agoGenerates original songs with vocals and instrumentals from text prompts using the Music v2 model. Supports genre, mood, and structural customization including mid-track transitions.
AI Sound Effects Generation
provisionalverified ~2 months agoGenerates custom royalty-free sound effects and ambient audio from text prompts using a dedicated AI model. Returns multiple distinct samples per generation.
Ads Engine: automated multi-market ad localization in ElevenCreative
provisionalverified 25 days agoAds Engine is a product within ElevenCreative that connects directly to Google Ads and Meta Ads accounts to pull existing ad creatives and performance data, then automates localizing those campaigns across 50+ languages including text translation, image adaptation, and video dubbing.
Conversational AI Agents (ElevenAgents)
provisionalverified ~2 months agoPlatform for building and deploying real-time voice and chat agents that combine speech recognition, configurable LLMs, and low-latency TTS. Formerly called Conversational AI.
Developer API and SDKs
provisionalverified ~2 months agoREST API exposing all ElevenLabs capabilities (TTS, STT, voice cloning, sound effects, music, dubbing, voice agents, speech-to-speech) with official SDKs for Python, TypeScript/JavaScript, Flutter, Swift, and Kotlin.
Flows Agent: conversational AI builder for multi-model creative pipelines
provisionalverified 25 days agoFlows Agent is a conversational AI assistant embedded in ElevenCreative's Flows canvas that builds audio/video production pipelines from natural-language prompts instead of manual node wiring, spanning ElevenLabs audio models plus 50+ third-party image and video models.
Pricing Plans and Commercial Rights
provisionalverified ~2 months agoSix self-serve subscription tiers plus Enterprise, governing monthly credit allowances, commercial use rights, voice clone slots, audio quality, and workspace seats.
Procedures: structured SOP builder for ElevenAgents conversational agents
provisionalverified 25 days agoProcedures let teams define standard-operating-procedure-style instruction sets that govern how a conversational agent handles specific scenarios (refunds, billing, identity verification, troubleshooting), moving scenario-specific logic out of a single monolithic system prompt so agents respond faster and more consistently.
Speech-to-Text (Scribe)
provisionalverified ~2 months agoAI transcription system (Scribe) converting audio and video to text with speaker diarization, word-level timestamps, and non-speech event tagging across 90+ languages.
Studio (Long-Form Audio/Video Editor)
provisionalverified ~2 months agoTimeline-based end-to-end production environment for creating audiobooks, narrated videos, and long-form audio with AI voice, music, sound effects, and captions.
Text-to-Speech
provisionalverified ~2 months agoConverts text into lifelike, emotionally expressive speech using multiple AI models. Supports multilingual synthesis, emotion control via context and audio tags (Eleven v3), and real-time streaming.
Voice Cloning
provisionalverified ~2 months agoReplicates a speaker's voice from audio samples using two tiers: Instant Voice Cloning (IVC) for rapid prototyping from short samples, and Professional Voice Cloning (PVC) for near-indistinguishable results via fine-tuned model training.
Voice Design (AI-generated voices)
provisionalverified ~2 months agoGenerates entirely new synthetic voices from a text description, with no audio sample required. You write a prompt describing the voice you want (age, gender, accent, tone, pacing, emotion, and audio quality) and the model returns three candidate voice previews to choose from. It fills the gap when the exact voice you need is not available in the Voice Library, and the saved voice can then be used across Text to Speech, Studio, and the API.