Why sub-500ms latency matters for AI avatars
Most AI avatar demos look great in a recorded video and feel broken the moment you talk to one. The reason is almost always latency.
The 500ms wall
Human conversation has a measurable rhythm. Across cultures and languages, the gap between one person finishing a sentence and another responding averages around 200ms. Push past 500ms and the listener starts wondering whether the other person heard them.
What we measure
// Approximate latency budget per stage
const budget = {
stt: 150, // Deepgram streaming
llm_ttft: 200, // Grok first token
tts_first: 80, // Cartesia first audio chunk
lipsync: 40, // SyncTalk_2D @ 250 FPS on Triton
webrtc: 30 // LiveKit transport
};
The dominant cost is the LLM. Everything else has to be aggressive about streaming and chunking to stay under budget.
Why most competitors miss this
Commercial avatar APIs typically batch instead of stream. Fine for async use cases, fatal for real-time.