ElevenLabs Launches v4 Voice Models with 90+ Languages and 10-Second Cloning
US-based AI voice company ElevenLabs has officially launched two new voice models: ElevenLabs v4 and v4 Turbo. The new models use a fresh architecture to improve emotional expression control, reduce response latency for voice agents, and support over 90 languages, up from 70 in the previous version.
The v4 series offers finer voice control and faster voice cloning. Users need only 10 seconds of audio to clone a voice. When reading long texts, the model better maintains the speaker's tone and adjusts based on surrounding context. Inline tags have been expanded to allow stacking multiple tags executed in order. Japanese, Brazilian Portuguese, Mandarin, and Cantonese show the most notable quality improvements.
Over 55% of ElevenLabs' revenue comes from large enterprises. The new models suit voice agent scenarios, generating audio while the backend LLM begins outputting answers. Annualized revenue has climbed from about $330 million to over $600 million, with staff exceeding 800. The company plans to IPO within the next few years.