AI Voices Sound Entirely Human. What's Next? — Inworld TTS-2
In 2021 the Blizzard Challenge saturated — human listeners could no longer tell AI speech from a real actor. This deep dive covers how voice AI labs define, measure, and train for the new frontier: expressiveness. Covers the Mean Opinion Score (MOS) and its ITU-standardization history, LLM-style arenas with Elo rankings (Artificial Analysis), AI judges like UTMOS/DNSMOS (and why speech language models can't yet replace human raters), Inworld's deterministic DSP proxy (pitch range + variation + loudness variation), and how DPO/GRPO shape a voice — including a Goodhart's-law case where optimizing for intelligibility flattened prosody, and a reward-hacking case where a model rambled forever. Closes with the Stephen Hawking lesson: sounding human isn't always the goal.








