SENARYSTUDIOS

Notes · Senary Launcher 1.5

Voice without a recogniser

Opening an app by name on a phone with no speech recognition of its own, and the settings that took it from a 28-second worst case to about 0.9 seconds a name.

Senary Launcher 1.5 opens an app when you say its name: hold the search bar in the dock, speak, let go. Most phones come with speech recognition that runs on the phone, and where there is one, Senary uses it, asking Android for the on-device kind only. (That last part is read in the code rather than watched running, which the privacy policy also says.)

One of my test phones, a Pixel 10a running GrapheneOS, has no speech recognition at all. So Senary needed its own, small enough to run on the phone and good at the one thing it would be asked: the names of apps.

Fourteen names

I recorded myself saying fourteen apps I actually use, Revolut, Todoist, Trading 212 and Pret among them, and scored each engine by whether the right app would open.

EngineOpened the right app
Vosk, small English model6 of 14
Whisper tiny7 of 14
Whisper tiny, with the app names as its prompt12 of 14
Whisper base, with the app names as its prompt10 of 14

Why a dictionary loses

Vosk recognises words from a dictionary. It can be handed a list of words to listen for, but only words the model already knows: a word it does not know is skipped, and adding one means rebuilding the model. App names are mostly words no dictionary has.

Whisper, OpenAI's speech model, writes text from pieces of words, so it can spell a name it has never heard. On its own it spells them the way they sound, which is how Revolut comes out as “Revolute”. What fixed that is its prompt: Whisper can be given some text to carry on from, and Senary gives it the names of the apps on the phone. With them, “Revolute” became Revolut and “to do it” became Todoist.

Nothing is trained on the phone. An app installed today is in the prompt the next time you hold the bar. And bigger was not better here: the next model up, base, opened 10 rather than 12.

From 28 seconds to 0.9

The first working build took three to five seconds from letting go to the app opening, and once, on a one-second recording, 28. Whisper had started repeating itself, written until it ran out of room, then retried at a higher temperature, up to five times. An app name is a few words, so Senary now stops it at 16 pieces and never retries. That took out the worst case.

The rest came from a bench on the phone, every setting run in turn, so that heat slowed them all equally:

Never on a guess

Opening the wrong app is worse than opening none. Senary opens one only when what it heard is a whole app name or matches exactly one app; otherwise the words go into the search field with the candidates under them. Silence is not sent to Whisper at all, and the phrases it is known to invent from nothing, “you” and “thanks for watching”, are thrown away, because “you” is the start of YouTube.

Where the model comes from

Whisper tiny is 44 MB, too much to put in every install for the phones that will never need it. It is downloaded once from Google Play the first time it is needed, after the app asks, and after that it works with no connection at all: that was checked in aeroplane mode on a copy installed from Play. Your speech is turned into text on the phone.

Still open

Twelve of fourteen is not fourteen. And speaking after a twist, with no finger on the screen, misheard “Maps” twice in the first four tries, where holding the bar never did. Both are being watched.

More notes: all of them. What changed in each version: What's new.