🎯 On-Device Voice AI for Enterprises
Get dedicated support to ensure your specific needs are met.
Talk to Sales

Multilingual speech-to-text converts spoken audio into text across languages with a single engine. Cheetah Streaming Speech-to-Text does it in real time and entirely on-device, in English, French, German, Italian, Portuguese, and Spanish, so transcripts appear as words are spoken without audio leaving the device.

From AI agents and real-time coaching to meeting transcription, speech-to-text brings new ideas to life. Cheetah Streaming Speech-to-Text has created even more opportunities with cloud-level accuracy on the device in real time, combining the best of the cloud and on-device processing.

Now, Cheetah Streaming Speech-to-Text officially supports five new languages: French, German, Italian, Portuguese, and Spanish. More developers can build private and cost-effective AI applications with zero network latency using accurate, production-ready, cross-platform Cheetah Streaming Speech-to-Text.

On-device transcription with cloud-level accuracy within your web browser

Picovoice’s web demos leverage its web SDKs. They run within your web browser, meaning the audio is processed locally without using 3rd party cloud services.

Real-time transcription with Cheetah Streaming Speech-to-Text

Test Cheetah Streaming Speech-to-Text in English or change the language to French, German, Italian, Portuguese, and Spanish to test out the new languages.

Cheetah accuracy across languages

Word error rate (WER, lower is better) from the open-source real-time transcription benchmark shows that Cheetah runs ahead of Google's streaming engine in all six languages and within range of the strongest cloud engines, from a 34 MB model that runs on-device.

Streaming word error rate by language for Cheetah, Amazon Streaming, Azure Real-time, and Google Streaming.
LanguageCheetahAmazon StreamingAzure Real-timeGoogle Streaming
English10.1%5.6%8.2%11.9%
French13.6%9.3%n.a.18.5%
German11.9%8.4%10.0%16.1%
Spanish8.6%6.4%9.4%11.6%
Italian14.3%11.5%18.5%18.0%
Portuguese12.3%8.0%9.7%12.8%

On word emission latency, Cheetah delivers words in 590 ms versus 830 ms for Google streaming and 920 ms for Amazon streaming (Azure: 530 ms). Cloud engines lead on raw WER in several languages; the on-device case is real-time latency without a network, privacy, offline operation, and cost-effectiveness at scale.

Real-time translation with on-device speech-to-text

The same on-device stack can translate as it transcribes. Zebra Translate can run alongside Cheetah Streaming Speech-to-Text, translating transcripts into another language in real time without a cloud round-trip. Combining the two into a single pipeline turns live captions into multilingual subtitles.

Train Use-case Specific Speech Models

Cheetah Streaming Speech-to-Text offers cloud-level accuracy out-of-the-box. However, some use cases and industries, such as healthcare, finance, or legal, have special terminology that cannot be accurately predicted by generic speech models. Custom Vocabulary & Keyword Boosting features allow developers to customize speech-to-text models specific to their application on the no-code Picovoice Console.

Automated Transcription in 6 Languages: Automated transcription (English), Transcription automatisée (French), Automatisierte Transkription (German), Transcripción automatizada (Spanish), Transcrição automatizada (Portuguese), Trascrizione automatizzata (Italian)

Start Building Now

Building with intuitive and cross-platform speech-to-text SDKs doesn’t require any experience in Machine Learning. Anyone can start transcribing with a few lines of code.

Frequently Asked Questions

+
Can real-time multilingual speech-to-text run fully on-device?

Yes, as long as your choice of speech-to-text can be stored and run on your target environment. For example, Cheetah's lightweight 34 MB model doesn't take up too much space even on resource-constrained devices to run multiple languages together. So, you can add as many language models as you want while enjoying low latency and full privacy.

+
How accurate is on-device speech-to-text compared with cloud?

In Picovoice's open-source benchmark, Cheetah's English WER is 10.1% versus 5.6% for Amazon streaming, 8.2% for Azure, and 11.9% for Google streaming. While cloud speech-to-text APIs from Amazon and Azure lead on raw accuracy in several languages; Cheetah runs ahead of Google streaming in all six supported languages while running entirely on-device.

+
Can I customize speech-to-text for industry terminology?

Yes. Custom Vocabulary and Keyword Boosting on the Picovoice Console or via Cheetah Model Training API adapt Cheetah and Leopard models to domain terms in healthcare, finance, legal, and other fields, with no machine learning experience required.