How local transcription works
SideSense turns speech into text on your own machine with one built-in model that runs on any CPU, faster than live speech. No GPU, no presets, nothing to tune.
One built-in model, on your CPU
Transcription uses a single built-in model that runs on your processor, faster than live speech, on any supported Windows PC. There are no presets to compare and no GPU settings to think about: the Transcription page in Settings is down to the model, its download, and your language. Captions and answer drafts keep pace with the call, and nothing about the audio leaves your machine.
Languages: English, Hindi, and more to come
English is the default. Hindi is available under Settings, Transcription as a one-time download of about 523 MB, with a model purpose-trained on Indian voices; after the download it runs offline like everything else. During a live session you can switch the transcription language from the top bar, so a call that drifts between English and Hindi keeps a usable transcript. More languages will unlock the same way.
The hosted lane (coming soon)
Paid plans will also be able to send transcription to a hosted lane that does the work off your machine entirely. Hosted minutes will count against your plan quota once the lane is live; the local lane stays unmetered on every tier. See pricing for the numbers.
On an older build?
Builds before 0.1.14 shipped separate GPU and CPU transcription presets, and their in-app banners may have sent you here. Those presets are retired: updating moves you to the built-in model automatically, with nothing to reconfigure, and tells you so. The same applies to the short-lived custom fine-tuned model option from those builds. Update from inside the app, or install the current version from the download page.
Questions, answered.
Do I need a GPU for transcription?
No. Since version 0.1.14, transcription uses one built-in model that runs on your CPU faster than live speech, on any supported Windows PC. A GPU is only useful if you also run larger local AI models through Ollama.
Which languages are supported?
English is the default. Hindi arrived in 0.1.17 as a one-time download of about 523 MB, with a model purpose-trained on Indian voices; after the download it runs fully offline like everything else. More languages will unlock the same way.
Does any audio leave my machine?
No. Local transcription runs entirely on your PC, offline. Paid plans will also get an optional hosted transcription lane (coming soon) that does the work off your machine; hosted minutes will count against your monthly quota once that lane is live, and the local lane stays unmetered on every tier.
My app talks about GPU and CPU presets. Where are they?
You are on a build older than 0.1.14. Those builds shipped separate GPU and CPU transcription presets; newer builds replace them with the single built-in CPU model, and the app migrates your setting automatically when you update. Update from inside the app, or grab the current installer from the download page.
Transcription that just works.
One built-in model, any CPU, fully on your machine.