In brief

Recognizing a human voice, transcribing it, synthesizing it, responding to it in real time: these four operations now belong to distinct pipelines or unified models depending on the architecture. The field covers three model families — STT (speech-to-text), TTS (text-to-speech), and multimodal voice-text models known as speech-to-speech. Each family has its own players, its own benchmarks, its own trade-offs. And its own risks.


Recognizing speech: speech-to-text

Whisper, the open-source reference

Whisper (OpenAI, 2022) established itself as the reference model for automatic speech recognition. Its architecture is a classic Transformer encoder-decoder: the encoder transforms the audio signal into a latent representation, the decoder generates text tokens. The large-v3 version has 1.55 billion parameters, trained on more than 5 million hours of labeled audio.

In October 2024, OpenAI released Whisper large-v3-turbo: the decoder was reduced from 32 to 4 layers. The result — 8× faster processing with marginal quality degradation. A useful trade-off for production deployments where latency matters.

In March 2025, OpenAI launched gpt-4o-transcribe and its mini variant. The WER (Word Error Rate) dropped by approximately 35% compared to Whisper. Combined with the Realtime API (generally available since August 2025), this model family enables low-latency speech-to-speech interactions.

Commercial alternatives

Deepgram Nova-3 targets production use cases: multilingual transcription, real-time processing, noise robustness. February 2026 benchmark: 8.1% WER in English, 6.8% in multilingual.

AssemblyAI released Universal-2 in 2024 (21% gain in alphanumeric precision), followed by Slam-1 in October 2025, with multilingual streaming across 6 languages and built-in safety guardrails.

According to a 2026 benchmark published by AssemblyAI, gpt-4o-transcribe leads on English accuracy (6.5% WER), followed by Deepgram Nova-3 and Universal-2. Notable caveat: the majority of available benchmarks are produced by the industry players themselves — no standardized, independent third-party benchmark yet exists covering languages beyond English.

In short: a WER (Word Error Rate) of 6.5% means roughly 6 mistranscribed words per 100. That is acceptable for note-taking; insufficient for consumer-grade automatic subtitling or legal transcription. Low-resource languages (Lingala, Swahili, Breton…) remain at 15-30% WER depending on the model.


Synthesizing speech: text-to-speech

ElevenLabs and commercial voice cloning

ElevenLabs has positioned itself as the leader in commercial TTS: 10,000 voices available, 70 languages, voice cloning from a few seconds of reference audio. The company raised $80 million in November 2025 and expanded its offering to synchronized video generation.

PlayHT is gaining ground. In blind tests on the TTS Leaderboard, 65.77% of users preferred PlayHT over ElevenLabs on certain benchmarks — illustrating both the challengers’ progress and the inherent subjectivity of evaluating vocal naturalness.

Open-source advances

Several open models are reaching quality levels comparable to commercial offerings:

  • Coqui XTTS v2.5: voice cloning from 6 seconds of reference audio.
  • Bark: emotional expressiveness superior to ElevenLabs in certain tests.
  • Piper: optimized for edge and real-time deployment (mobile, embedded).
  • StyleTTS2: naturalness close to commercial models.
  • Chatterbox (2025): voice cloning from 5 seconds, 17 languages, Apache 2.0 license. Beats ElevenLabs in blind tests according to its developers.

An academic milestone: VALL-E 2

VALL-E 2 (Microsoft Research, June 2024) is presented as the first zero-shot TTS model achieving human parity on the LibriSpeech and VCTK benchmarks. Two technical innovations: Repetition Aware Sampling (stabilizing autoregressive decoding) and Grouped Code Modeling (compressing audio codec sequences). This model is research-only — no public deployment is planned.

Limits of expressive TTS

TTS performance drops as soon as one moves beyond neutral registers. Emotional fidelity — sarcasm, urgency, sadness — remains difficult to evaluate: no standardized public benchmark exists for prosody. Comparisons remain subjective. The scarcity of prosody-annotated training data partly explains this limitation.


Real-time voice: speech-to-speech models

GPT-4o and native voice conversation

GPT-4o voice mode (OpenAI) processes audio signals directly on input and output, without going through intermediate transcription. The model detects prosody (sarcasm, urgency, hesitation), handles interruptions smoothly, and maintains low latency. The Realtime API has been generally available since August 2025.

The end-to-end approach — a single model for incoming and outgoing audio — involves a trade-off: it improves latency but can degrade transcription quality compared to a separate STT + LLM + TTS pipeline.

Gemini Live

Gemini Live (Google, updated in 2025–2026) integrates real-time speech rate adjustment, adaptive emotional tone, and accent-changing options. The architecture is optimized for accessibility. Notable limitation: API integration is less open than GPT-4o Realtime. Gemini Live remains primarily mobile-first.

Moshi: the open-source full-duplex architecture

Moshi (Kyutai, France, July 2024) is an open-source speech-text foundation model designed for full-duplex conversation — meaning the model can speak and listen simultaneously, as in a real conversation.

Its architecture rests on two distinct components: a Temporal Transformer (7 billion parameters) that captures long temporal dependencies, and a Depth Transformer that manages dependencies between the different layers of the audio codec. The neural codec Mimi (developed by Kyutai) encodes audio at 12.5 Hz and 1.1 kbps, combining semantic and acoustic information in a compact representation.

The Inner Monologue method is particularly notable: before generating audio tokens, the model predicts temporally aligned text tokens. This intermediate step improves the factual accuracy of vocal responses.

Theoretical latency: 160 ms (80 ms for the Mimi frame + 80 ms acoustic delay). In practice, ~200 ms on an L4 GPU.

Sesame CSM-1B: conversational naturalness

Sesame CSM-1B (Sesame AI, March 2025, Apache 2.0) adopts a two-level architecture: a Llama backbone (1 to 8 billion parameters depending on the variant) and a Mimi audio decoder (100 to 300 million parameters). Trained on 1 million hours of English audio, it generates natural pauses, hesitations, laughter, and tonal variations — naturalness markers that classical TTS systems struggle to reproduce.


Which model to choose for which need

ContextRecommendationWhy
Transcription where accuracy is critical (legal, medical)gpt-4o-transcribe or Universal-2WER below 7% in English, mature commercial multilingual support
On-premise / private deploymentWhisper large-v3 (self-hosted)Open source, no data leaving for an external cloud
Multilingual synthetic brand voiceElevenLabs or PlayHT70+ languages, complete production pipeline, commercial expressiveness
Embedded / edge TTS (mobile, IoT)PiperOptimised for latency and footprint, runs without a GPU
Real-time voice conversationGPT-4o Realtime API or MoshiFull-duplex architecture, latency under 200 ms, handles interruptions
Apache 2.0 voice cloningChatterbox or Coqui XTTS v2.5Permissive licence, 5-6 seconds of reference audio are enough

Ethical stakes

Voice as a vector for fraud

Voice cloning from a few seconds of audio exposes a concrete threat: the generation of synthetic voices indistinguishable from the original, used to deceive interlocutors or steal identities. Detection tools (Resemble Detect, ElevenLabs Speech Classifier) lag behind recent generative models. The arms race between generation and detection is accelerating.

The Tennessee ELVIS Act (2024) is the first American law explicitly protecting voices cloned by AI. The EU AI Act classifies voice cloning as a high-risk AI use. But no international standard yet exists — jurisprudence on voice as a personality attribute remains fragmentary.

The question of training data provenance remains opaque for commercial models. Neither ElevenLabs nor PlayHT publicly details the composition of their corpora. This makes any ethical and legal assessment of provenance currently impossible.

Voice cloning without the explicit consent of the person being imitated lacks uniform regulation. The right to one’s voice as a personality attribute exists in certain jurisdictions, but its application to generative AI cases is still being constructed through the courts.


Key takeaways

  • STT is dominated by Whisper (open-source) and gpt-4o-transcribe (commercial), with English WER dropping below 7% for the best models — but low-resource languages remain significantly behind.
  • Commercial TTS is concentrated around ElevenLabs and PlayHT; open-source (Chatterbox, Coqui XTTS) is advancing and making self-hosted deployment viable.
  • Speech-to-speech models (GPT-4o, Moshi, Sesame CSM) unify the audio chain in a single model, at the cost of a latency/quality trade-off that is still poorly calibrated empirically.
  • Voice cloning is technically accessible from a few seconds of audio. The legal framework exists but remains fragmented. Detection tools are outpaced by recent generative models.
  • Available benchmarks are predominantly published by the industry players themselves. No standardized third-party benchmark yet exists for non-English languages or emotional expressiveness.