Notes · Senary Launcher 1.5
Voice without a recogniser
Opening apps by voice on a phone with no speech recognition, and getting it from 28 seconds at worst to about 0.9.
1.5 lets you open an app by saying its name: hold the search bar, speak, let go. Most phones have speech recognition that runs on the phone, and Senary uses it when it's there, asking Android for the on-device kind only. (That last part is checked in the code, not by watching it run; the privacy policy says the same.)
One of my test phones, a Pixel 10a on GrapheneOS, has no speech recognition at all. So Senary needed its own: small enough for the phone, and good at app names.
Fourteen names
I recorded myself saying 14 apps I use, including Revolut, Todoist, Trading 212 and Pret, and counted how often each engine opened the right one.
| Engine | Opened the right app |
|---|---|
| Vosk, small English model | 6 of 14 |
| Whisper tiny | 7 of 14 |
| Whisper tiny, with the app names as its prompt | 12 of 14 |
| Whisper base, with the app names as its prompt | 10 of 14 |
Why a dictionary loses
Vosk works from a dictionary. You can give it a list of words to listen for, but only words it already knows; anything else is skipped, and adding a word means rebuilding the model. Most app names aren't in any dictionary.
Whisper, OpenAI's speech model, builds words from pieces, so it can spell names it's never heard. On its own it spells them how they sound: Revolut comes out as “Revolute”. The fix is the prompt. Whisper can be given text to carry on from, and Senary gives it the names of your installed apps. With that, “Revolute” became Revolut and “to do it” became Todoist.
Nothing is trained on the phone. A new app is in the prompt the next time you hold the bar. And bigger wasn't better: the next model up, base, got 10 rather than 12.
From 28 seconds to 0.9
The first working build took three to five seconds from letting go to the app opening. Once, on a one-second clip, it took 28: Whisper got stuck repeating itself, ran out of room, then retried up to five times. An app name is a few words, so Senary now stops it at 16 pieces and never retries. That fixed the worst case.
The rest came from benchmarking on the phone, each setting run in turn so heat hit them all equally:
- The less compressed model was faster: 0.91 seconds a name against 1.43, same accuracy.
- Four threads, not six. The 10a has four fast cores and four slow ones; six threads end up waiting on the slow ones.
- The full audio window, not a shorter one. Cutting it got 2 of 14, with WhatsApp heard as “WorldSupport”.
- Five guesses at once, not one. With the length capped, one guess wasn't faster, and it was less accurate.
Never on a guess
Opening the wrong app is worse than opening none. Senary only opens an app when it heard a whole app name, or something that matches exactly one app. Otherwise the words go in the search field with the matches below. Silence isn't sent to Whisper, and the phrases it makes up from nothing, “you” and “thanks for watching”, are thrown away, because “you” is the start of YouTube.
Where the model comes from
Whisper tiny is 44 MB, too big to put in every install when most phones won't need it. It's downloaded once from Google Play the first time it's needed, after the app asks. After that it works with no connection; I checked in aeroplane mode on a copy installed from Play. Your speech is turned into text on the phone.
Still open
12 of 14 isn't 14. And voice after a twist misheard “Maps” twice in the first four tries, where holding the bar never did. Watching both.
More notes: all of them. What changed in each version: What's new.