Benchmarking models for voice

Benchmarking models for voice is hard. Most published numbers tell you if a model is fast. They don’t tell you if it actually sounds good or performs tasks well in real production environments. The bar is high: humanlike. We’re learning exactly what’s required to clear it.

Today, LiveKit is releasing the first iteration of benchmarks that test real agent scenarios (with a 100-scenario corpus, tool-heavy, adversarial callers) and tracks both how well LLMs complete the task at hand and how they perform under real conditions: time to first sentence, worst-case behavior, consistency over days instead of a single run. You can also compare two models turn by turn in the same conversation and see exactly where they diverge.

Resources

2 Likes