Notes · Senary Launcher 1.5
Voice without a recogniser
Opening an app by name on a phone with no speech recognition of its own, and the settings that took it from a 28-second worst case to about 0.9 seconds a name.
Senary Launcher 1.5 opens an app when you say its name: hold the search bar in the dock, speak, let go. Most phones come with speech recognition that runs on the phone, and where there is one, Senary uses it, asking Android for the on-device kind only. (That last part is read in the code rather than watched running, which the privacy policy also says.)
One of my test phones, a Pixel 10a running GrapheneOS, has no speech recognition at all. So Senary needed its own, small enough to run on the phone and good at the one thing it would be asked: the names of apps.
Fourteen names
I recorded myself saying fourteen apps I actually use, Revolut, Todoist, Trading 212 and Pret among them, and scored each engine by whether the right app would open.
| Engine | Opened the right app |
|---|---|
| Vosk, small English model | 6 of 14 |
| Whisper tiny | 7 of 14 |
| Whisper tiny, with the app names as its prompt | 12 of 14 |
| Whisper base, with the app names as its prompt | 10 of 14 |
Why a dictionary loses
Vosk recognises words from a dictionary. It can be handed a list of words to listen for, but only words the model already knows: a word it does not know is skipped, and adding one means rebuilding the model. App names are mostly words no dictionary has.
Whisper, OpenAI's speech model, writes text from pieces of words, so it can spell a name it has never heard. On its own it spells them the way they sound, which is how Revolut comes out as “Revolute”. What fixed that is its prompt: Whisper can be given some text to carry on from, and Senary gives it the names of the apps on the phone. With them, “Revolute” became Revolut and “to do it” became Todoist.
Nothing is trained on the phone. An app installed today is in the prompt the next time you hold the bar. And bigger was not better here: the next model up, base, opened 10 rather than 12.
From 28 seconds to 0.9
The first working build took three to five seconds from letting go to the app opening, and once, on a one-second recording, 28. Whisper had started repeating itself, written until it ran out of room, then retried at a higher temperature, up to five times. An app name is a few words, so Senary now stops it at 16 pieces and never retries. That took out the worst case.
The rest came from a bench on the phone, every setting run in turn, so that heat slowed them all equally:
- The less compressed of two versions of the model was faster: 0.91 seconds a name against 1.43, with the same accuracy.
- Four threads, not six. The 10a has four fast cores and four slow ones, and six threads spend their time waiting on the slow ones.
- Listening to the whole window Whisper expects, not a shorter one. Cutting it opened 2 names of 14, one of them WhatsApp heard as “WorldSupport”.
- Searching five ways at once rather than taking the first guess. Once the length was capped, the first guess was no faster, and it was less accurate.
Never on a guess
Opening the wrong app is worse than opening none. Senary opens one only when what it heard is a whole app name or matches exactly one app; otherwise the words go into the search field with the candidates under them. Silence is not sent to Whisper at all, and the phrases it is known to invent from nothing, “you” and “thanks for watching”, are thrown away, because “you” is the start of YouTube.
Where the model comes from
Whisper tiny is 44 MB, too much to put in every install for the phones that will never need it. It is downloaded once from Google Play the first time it is needed, after the app asks, and after that it works with no connection at all: that was checked in aeroplane mode on a copy installed from Play. Your speech is turned into text on the phone.
Still open
Twelve of fourteen is not fourteen. And speaking after a twist, with no finger on the screen, misheard “Maps” twice in the first four tries, where holding the bar never did. Both are being watched.
More notes: all of them. What changed in each version: What's new.