Voice
Select speech models and prepare scripts for reliable automated narration
Automated voice work has two separate failure surfaces: choosing a model that fits the delivery and compute budget, and preparing text that the model can pronounce reliably. Clean narration may not expose meaningful differences between capable models, while numbers, identifiers and terminal commands can fail even on larger systems. Expression, multilingual delivery and commercial licensing create additional selection gates. Keep the written source and pronunciation-oriented rendering copy separate so audio fixes do not corrupt reusable content.
Start here
- Selecting an Open-Source TTS Model for Content Automation — Choose candidates by narration, expression, hardware and license constraints.
- Normalizing TTS Scripts Before Synthesis — Add a pronunciation and review boundary for difficult tokens before rendering.
Pages
- Selecting an Open-Source TTS Model for Content Automation — Sam Hottman's five-model Korean comparison under one RTX 5090 setup
- Normalizing TTS Scripts Before Synthesis — Sam Hottman's mitigation for dates, money, versions, identifiers and commands
Gaps
- Independent, blinded listening tests across Korean voices and content formats.
- Reproducible latency, VRAM, CPU throughput and long-run stability measurements by model revision.
- Verified code, weight and generated-output license terms for commercial deployment.
- A tested normalization rule set with error rates for Korean numbers and technical notation.