Why Orca Streaming Text-to-Speech?
Natural-sounding confirmation at 41 MB peak memory.
106 ms
First-token-to-speech latency (3.1x faster than ElevenLabs Streaming at 335 ms)
41 MB
Peak memory (11x less than the lightest on-device alternative)
3.1x
Less CPU than the most compute-efficient neural on-device TTS
Orca Streaming Text-to-Speech reads each dialer response aloud, such as "Calling Sarah Chen on mobile," "I found Sarah Chen and Sarah Khan. Which one?", or "There is no work number for Sarah Chen. Should I try mobile?". This way the user never has to look at the screen while driving, walking, or operating equipment. Most high-quality TTS engines require hundreds of megabytes of RAM. Orca uses 41 MB peak memory, 10–50x less than any natural-sounding on-device alternative, which fits easily inside a car head unit, a Bluetooth headset firmware image, or a pair of smart glasses. First-token latency is 106 ms, fast enough that spoken confirmations feel conversational.
ElevenLabs TTS Streaming335 ms