Text-to-Speech (TTS)
in Today’s World
From Concatenative to Flow Matching and Hybrid AR+LLM models; a comprehensive review of architectures, top 2025–2026 models, Persian challenges, and the future path of this technology.
What is TTS and why is it important?
Text-to-Speech or speech synthesis is the technology that converts written text into natural-sounding speech. However, “natural” has a deeper meaning here — not just correct pronunciation, but accurate prosody, appropriate emotion, natural pauses, and a tone that makes the listener forget they are talking to a machine.
If just a few years ago TTS was merely an accessibility tool, today it is the core engine of voice agents, audiobooks, automated dubbing, and conversational AI. The difference between a voice agent that attracts a user and one that drives them away often lies in the quality of the TTS.
In 2026, for the first time, open-source models surpassed ElevenLabs — the best commercial service — in blind tests. Chatterbox Turbo received a 65.3% listener preference in a blind test. This was the threshold that changed everything once crossed.
The Evolution of TTS Architectures
TTS has evolved through five distinct eras — each era representing a qualitative leap, not just an improvement:
The computer extracted small audio units (diphones or larger units) from a recorded corpus and spliced them together. The voice sounded artificial and robotic, and the fragment boundaries were audible. But for decades, it was the only practical option.
Unit Selection · Diphone · FestivalStatistical models predicted acoustic parameters (pitch, duration, spectral), and then a vocoder converted these parameters into a sound wave. It offered more flexibility, but the voice still sounded “muffled” and lifeless.
HMM · MERLin · DNN · SPSSGoogle’s WaveNet (2016) and Tacotron (2017) created a true qualitative leap. For the first time, synthesized voices could fool humans. Two-part models (Acoustic model + Vocoder) became the standard: FastSpeech2 and HiFi-GAN brought this era to its peak.
WaveNet · Tacotron2 · FastSpeech2 · HiFi-GAN · VITSModels like VALL-E, YourTTS, and NaturalSpeech 2 showed that anyone’s voice could be cloned from a few seconds of audio. Diffusion and Flow Matching replaced older architectures. XTTS v2 brought this capability to open-source.
VALL-E · NaturalSpeech2 · XTTS-v2 · Diffusion · Flow MatchingThe new generation uses LLMs as a backbone (CosyVoice2, Fish Speech, Chatterbox) and is fine-tuned with Reinforcement Learning to offer more natural prosody, emotion control, and rapid zero-shot voice cloning. Voxtral TTS from Mistral introduced the hybrid AR + Flow Matching architecture. F5R-TTS uses GRPO to simultaneously improve WER and speaker similarity.
LLM-TTS · CosyVoice2 · Flow Matching + RL · GRPO · Chatterbox · VoxtralTop 2025 Models — Foundations of the New Generation
These models became the industry standard in 2025 and many are still used in production today:
The first model with a complete audio tag system — you can control emotions with tags like [laughs], [sighs], or [excited]. Text to Dialogue blends multiple voices into a single output. ElevenLabs is no longer just TTS; it’s a complete audio platform.
The best option for real-time voice agents where latency is crucial. Sonic Turbo with 40ms is designed for conversational apps. Sonic 3.5 arrived in May 2026 with better prosody and a broader emotional range.
Highest ELO among multilingual open-source models. Built on an LLM-based architecture, offering high-quality zero-shot voice cloning. Ideal for multilingual content generation applications.
The smallest model delivering commercial-grade quality. 82M parameters running 96× real-time on a standard GPU. No encoders or diffusion — direct and fast. Best choice for edge deployment and latency-sensitive services.
Best streaming latency among open-source models. Uses supervised semantic tokens (from ASR) for better quality. Ideal for real-time voice agents requiring self-hosting.
The main challenge of automated dubbing is fitting audio perfectly into timelines. IndexTTS-2 separates duration control from emotion — allowing independent adjustments. A core tool for post-production teams.
New 2026 Models NEW
March 2026 was a historic day for TTS — three major models were released within hours, sparking a new wave of competition:
The first open-source model to surpass ElevenLabs in blind tests — 65.3% of listeners preferred it. Trained on 500k+ hours of audio. Offers emotion control, zero-shot voice cloning, and sub-150ms streaming. Also fine-tuned for Persian (by the community).
A unique hybrid architecture: AR for natural prosody + Flow Matching for speed. Preferred by 62.8% in blind tests and competes with the ElevenLabs v3 tier in expressiveness. Pricing is nearly one-fifth of ElevenLabs.
The first voice model with GPT-5 reasoning capabilities. Supports tool calls, interruption management, and real-time correction. GPT-Realtime-Translate for live translation and GPT-Realtime-Whisper for real-time transcription were also added. Ideal for voice agents requiring reasoning.
Just as ASR improved significantly with RL, TTS is following the same path. F5R-TTS utilizes GRPO with a dual reward system (WER + Speaker Similarity). Result: Simultaneous improvement in intelligibility and voice cloning quality.
Highest ELO on the TTS Arena Leaderboard for accessible models. Voice cloning is included in the base price — unlike ElevenLabs which charges separately. Inworld Mini is designed for real-time gaming with sub-130ms latency.
The most affordable model in the top 10 leaderboard. Designed for teams with high-volume audio generation needs and tight budgets — offering the lowest price for decent quality.
March 26, 2026 saw the release of Voxtral TTS from Mistral, Transcribe from Cohere, and CoVo-Audio from Tencent all within a few hours. A user on Reddit wrote: “The on-prem voice stack is here.” — a perfect summary of that day.
Quantitative Comparison — MOS, ELO, and Latency
Evaluating TTS differs fundamentally from STT. In TTS, the primary metric is MOS (Mean Opinion Score) — human judgment — not a computational figure. In 2026, both TTS Arena Leaderboard and Artificial Analysis utilize blind test-based ELO:
| Model | ELO / MOS | Latency | Voice Cloning | Type | Approximate Price |
|---|---|---|---|---|---|
| Inworld TTS MAX 2026 | ELO 1594 | — | Free (included) | Commercial | Medium |
| ElevenLabs v3 | Top Tier | 75–150ms | Extra Cost | Commercial | $5–$1300/mo |
| Chatterbox Turbo 2026 | 65.3% blind win | <150ms | Free (MIT) | Open (MIT) | Free |
| Voxtral TTS 2026 | 62.8% blind win | — | Zero-shot | Open-weight | ~1/5 ElevenLabs |
| Cartesia Sonic 3.5 2026 | Temp #1 AA | 40ms (Turbo) | Limited | Commercial | Medium |
| Fish Speech V1.5 | ELO 1339 | — | Zero-shot | CC-BY-NC | Free (non-com) |
| CosyVoice2-0.5B | Good | 150ms streaming | Zero-shot | Open (Apache 2) | Free |
| Kokoro (82M) | Great for size | 96× RT | None | Open (Apache 2) | Free |
| Speechify SIMBA 3.0 2026 | ELO ~1159 | — | Limited | Commercial | $10/1M char |
ELO and MOS heavily depend on the test data and the evaluator demographic. A model that excels at audiobook narration might be mediocre in a conversational voice agent. Always test on your specific use case with your actual listeners.
Voice Cloning in 2026 — Current Status
Voice cloning (generating a specific person’s voice from a few seconds of audio) became a mainstream capability in 2026. But with this power, crucial issues have emerged:
Models like Chatterbox, Fish Speech, and CosyVoice2 clone voices with just 3–10 seconds of audio. Fine-tuning is no longer required.
You can clone a Persian speaker’s voice and make it speak English or Japanese — without a Persian accent. This is revolutionary for international dubbing.
Resemble AI, ElevenLabs, and Google have audio watermarking systems. Detecting audio deepfakes has become an independent research field.
The EU and several US states have passed restrictive laws. Cloning a voice without the owner’s consent is banned on most platforms.
New models can separate identity from emotion — creating the same voice with happiness or sadness without altering its core identity.
ElevenLabs Professional Voice Cloning requires 30 minutes of high-quality recording for production-grade results. The quality is significantly better than zero-shot.
Why is Persian TTS a Unique Problem?
Persian TTS has different challenges compared to Persian STT — and in some ways it’s harder, because TTS requires much cleaner data:
Everyday written Persian lacks vowels (diacritics). “Madraseh” (School) or “Modarres” (Teacher)? The model must deduce it from context — mispronunciations make the output sound unnatural.
Numbers (Persian/Arabic/Latin), dates, units, abbreviations, and emojis — each must be read correctly. The TTS Frontend for Persian is far more complex than for English.
TTS data must be much cleaner than ASR data: no noise, perfect pronunciation, high sample rate. ManaTTS and ParsVoice have made strides, but it’s still not enough.
The stress and intonation patterns of Persian are completely different from English. A model trained only on English will produce incorrect intonations when generating Persian.
Everyday Persian is full of English, French, and Arabic words, each with different pronunciation rules. The model needs to know when to read “Software” with a Persian accent vs English.
Most Persian TTS datasets feature single or limited speakers. A model trained on them will exhibit lower quality during voice cloning.
The best available open-source Persian TTS model (based on Tacotron2 + ManaTTS) has reached a MOS of around 3.76 — compared to the 4.5+ MOS ElevenLabs claims for English. This gap shows that truly high-quality Persian TTS remains an unsolved problem.
The State of Persian TTS in 2026
| Model / Approach | Type | Status | Note |
|---|---|---|---|
| Chatterbox Persian Fine-tuned 2025 | Open-source + Community | Available (HuggingFace) | Community fine-tune on Persian is available. Quality is acceptable for general use but not production-grade. |
| Tacotron2 + ManaTTS | Custom Open-source | Best Open-source | ManaTTS: 114 hours of high-quality single-speaker data. Best baseline for Persian TTS. MOS 3.76 |
| Persian VITS (ZabanZad.ai) | Open-source | Research | SAIL Lab project aiming for production-grade Persian TTS. Currently in development. |
| ParsVoice (Multi-speaker) 2025 | Dataset | Published | First large-scale multi-speaker Persian corpus. Extracted from IranSeda. Foundation for multi-speaker Persian models. |
| XTTS v2 + Persian Fine-tuning | Open-source + Custom | Needs Work | Can be fine-tuned with 6 minutes of Persian data, but Persian text normalization must be implemented independently. |
| Iranian Commercial Services | Commercial | Limited | Some Iranian companies have Persian TTS (like Balad), but lack open access and quality varies. |
Practical Recommended Approach for Persian TTS
- Short-term: XTTS v2 or Chatterbox fine-tuned on your own custom dataset — at least 30 minutes of clean, labeled audio.
- Frontend is Critical: Build a robust Persian text normalization module — numbers, dates, abbreviations, English words. This is more important than model selection.
- Grapheme-to-Phoneme: A precise Persian G2P is essential for models that ingest phonemes — especially for Arabic and English words within Persian text.
- Medium-term: Train on ParsVoice for multi-speaker capabilities. It provides multi-speaker data, yielding better voice cloning.
- Long-term: Fine-tune Voxtral TTS (once generalized to Persian) with a domain-specific dataset — this path will likely yield the best quality.
A natural-sounding, multi-speaker, clonable Persian TTS — with a MOS above 4 — does not yet exist. Any team that builds this and offers it via API will capture the markets of Iran, Afghanistan, and Tajikistan. The first step is creating a clean, multi-speaker dataset.
The Future Path of TTS — Where Are We Heading?
Just as Qwen3-ASR with GSPO-RL revolutionized ASR, F5R-TTS and GLM-TTS showed that RL with GRPO can simultaneously improve intelligibility and speaker similarity. By 2027, most leading models will use RL.
GPT-Realtime-2 paved the way: Not STT → LLM → TTS, but an integrated model that directly understands audio and replies in audio. This slashes latency from 800ms to under 300ms, fundamentally upgrading conversation quality.
ElevenLabs v3 started with audio tags, but the path is leading to models that infer these states autonomously from textual context. A TTS that understands “this sentence should be spoken with hesitation” without a programmer writing tags.
Kokoro with 82M parameters boasts commercial-grade quality. The trajectory aims at models that run real-time on smartphones. This is crucial for privacy, offline modes, and markets with limited connectivity.
The combination of zero-shot voice cloning and RL fine-tuning allows every user to have a personal voice profile. Your AI assistant will speak with your own voice — or one explicitly designed to your preferences.
Industrial Applications of TTS
Highest Quality + Emotion: ElevenLabs v3 · Lowest Latency (40ms): Cartesia Sonic Turbo · Best Blind Open-Source: Chatterbox Turbo · Multilingual Open: Fish Speech V1.5 · Edge/On-device: Kokoro 82M · Streaming Open: CosyVoice2-0.5B · Dubbing: IndexTTS-2 · Persian: Tacotron2 + ManaTTS + Custom Normalization
Conclusion
In 2026, TTS has entered a phase where “good enough” is no longer acceptable. When Chatterbox Turbo with an MIT license beats ElevenLabs in blind tests, it means any team with reasonable resources can build a commercially viable TTS.
Three key takeaways for teams active in this field:
- The AR+Flow Matching+RL architecture is the future: Models combining all three elements — like Voxtral — offer the best balance of quality, speed, and controllability. This architecture will become the standard over the next two years.
- For Persian, Frontend is more critical than the model: A robust Persian text normalization module — which accurately converts numbers, dates, English terms, and specialized jargon — has a greater impact than choosing the best base model.
- Multi-speaker Persian data is the main missing piece: ParsVoice was a huge step, but we still need larger, more diverse datasets with broader domain coverage. The team that builds this data will win the Persian TTS market.