Hands-free cooking: a wink for the next step, a voice for everything else
Voice control is the wrong instrument for turning a recipe page. A look at what you actually ask a recipe mid-cook, and the cheaper controls that fit each ask.

Hands-free cooking has meant voice control for about a decade, and voice is genuinely good at the hard part: answering a question you did not plan to ask, in the middle of a step. It is a poor fit for the easy part. Nine times out of ten what you need from a recipe is one bit of information passed the other way: forward. Saying a whole sentence out loud, over an extractor fan, to communicate one bit is an odd trade, and it is why most people give up on voice after a week and go back to wiping their hands.
The useful principle is that the cost of the control should match the size of the request. A wink is enough to turn a page. Holding your mouth open is enough to scroll. Speaking should be reserved for the things only speech can express. This post is about what that looks like when you take it seriously.
What you actually ask a recipe
Stand at a hob with a recipe on a phone and count the interactions. They fall into four buckets, and they are wildly uneven in both frequency and size.
Forward
By far the most common, and the smallest possible request. One bit. It happens ten to twenty times a meal.
Back
Rarer, equally small. Usually because you half-read a quantity and want to check it.
More of this step
A long step runs off the screen. Not a jump: a continuous, held request. You want it to move while you look, then stop.
A real question
"How much of the stock goes in now?" "Can I use crème fraîche instead?" Open-ended, unpredictable, genuinely needs language.
Only the fourth one needs a sentence. The first three are the overwhelming majority of what happens, and a voice assistant charges the same price for all four: get its attention, form a phrase, say it clearly enough to be transcribed, wait.
Voice is a general-purpose tool billed at a flat rate
The flat rate is the problem, and it gets worse in the exact conditions where you need hands-free most. Kitchens are loud in a way that is unusually hostile to speech recognition: an extractor fan, a running tap and something frying are all broadband noise sitting on top of the voice band. Wake words are tuned to be strict, because one that fires on a television is unusable, and strictness costs you misses. Three failed attempts to say "next step" over a fan is worse than just wiping your hands, and everyone works this out quickly.
Then there is the part nobody writes about: cooking is frequently social. Narrating instructions to a phone while someone else is in the kitchen talking to you is not a neutral act. People stop doing it, not because it fails, but because it is faintly embarrassing.
None of this makes voice bad. It makes it the wrong tier for the small stuff.
Voice, for everything
Fixed cost per request. Needs a quiet-ish room, a full phrase, and a moment of your attention to compose it.
Degrades exactly when the kitchen is busiest.
Tiered controls
A gesture for the one-bit asks, a held gesture for the continuous one, speech for the questions that need language.
The common case gets cheap; the rare case keeps its full power.
A blink is not a wink
The obvious cheap control is the face, and the obvious first idea is to watch the eyes. This is harder than it sounds, and the reason is worth understanding, because it is the thing that makes most face-gesture features feel broken.
A person blinks fifteen to twenty times a minute without deciding to. So "an eye is closed" carries no intent whatsoever. What separates a wink from a blink is not that the eye is shut. It is asymmetry. A natural blink closes both lids together, within twenty to forty milliseconds of each other. A wink closes one.
That single observation changes what you build. It is not a detector with a filter bolted on afterwards; it is an asymmetry detector where symmetric closure is the explicit reject case. Once you frame it that way, a wink can be confirmed the instant a blink has been ruled out (roughly a tenth of a second) rather than waiting for anything.
One 900 ms window, played back slowed
0 ms
A blink
A wink
The tempting mistake is to require a hold: keep the eye shut for a moment, so we know you meant it. It sounds safer and it is the wrong mechanism. A quick, natural wink closes and reopens in eighty to a hundred and twenty milliseconds, which is over before any sensible hold threshold elapses. So the gesture gets thrown away the moment the eye reopens, and the cook has to learn to squint deliberately at their dinner. Duration was never what distinguished a wink from a blink. Asking for it buys nothing and makes the whole thing feel sluggish in a way no amount of tuning fixes.
Eyes are worse pointers than you think
Scrolling is a different shape of request (held rather than fired) and the elegant-sounding answer is to use gaze. Look down, the text moves. Look up, it stops.
It measures badly. On a real device, the entire usable range of head-relative gaze is about 0.13 in normalised units: sweeping your eyes from the top of a step to a button at the bottom of the screen moves the signal by roughly 0.10. Meanwhile, simply reading the screen wanders by about ±0.065. The noise floor and the signal are the same size. No threshold separates them, and that is a property of the sensor rather than something you can tune your way out of.
The jaw, unglamorously, is much better. Mouth-open has a genuine zero-to-one range, needs no adaptive baseline, and is unambiguous. You either did it or you did not. Its one real false-positive risk is speech, which is solved by a short hold: ordinary talking opens the jaw repeatedly but briefly, and rarely holds it wide past a third of a second.
Step 4 of 9
ClosedTurn the heat right down and add the sliced onions with a good pinch of salt. Cook them gently for thirty to forty minutes, stirring every five or so, until they have collapsed to about a third of their volume and turned a deep amber. Do not rush this with a higher flame. They will colour at the edges before the middle has softened, and you will end up with burnt fragments in something that still tastes raw. If the pan looks dry, a spoonful of water lifts the browned bits off the bottom.
Gaze still earns its place, just not as a pointer. It answers a much coarser question (is this person looking at the phone at all?) and that question turns out to be the one that matters, because a wink aimed at someone across the kitchen is not a navigation request.
What this looks like in Sugo
Cook mode in Sugo puts three controls on the same face, and each one is sized to its job.
Wink right for the next step, wink left for the previous one. Right and left are yours, not the camera's, which is why setup asks you to perform each gesture once before switching anything on. A normal blink closes both eyes, so it never triggers anything. After a gesture fires there is a second of deafness, so one enthusiastic wink cannot page through three steps.
Hold your mouth open to scroll the current step. It moves while your mouth is open and stops when you close it, which is the right shape for a long paragraph you are reading at arm's length.
Say "hey Sugo" for anything that needs words. The wake word runs on your device, and the assistant that answers already has the recipe in front of it, so "how much butter was it?" is a complete question.
Crush the San Marzano by hand, straight into the pan. Don't be tidy.
Flour on your hands · nothing tapped
Three details are worth calling out because they are the ones that stop this feeling like a camera staring at you. Everything happens on your device. No photos or video are saved or sent anywhere. The camera indicator stays lit the whole time cook mode is open, deliberately. And a forward wink does nothing on the last step: finishing a recipe stays a deliberate tap, because it is the one action you cannot casually undo.
If you want the rest of the setup (the propping, the auto-lock, the reading-ahead) that is covered in cooking from a recipe without touching your phone. And the whole thing only works if the recipe is already in a form worth cooking from, which is what saving recipes properly is for.
What it doesn't do
Winking one eye independently is a learned motor skill, and a real share of people simply cannot do it. That is why the setup makes you perform the gesture rather than just explaining it: if it does not work for you, you find out in ten seconds instead of mid-recipe. It also means this is genuinely not a feature for everyone, and the tap targets are not going anywhere.
The camera has to see your face, which means the phone needs to be propped up, facing you, within arm's reach, in enough light. A phone lying flat on the counter defeats the entire thing, though it was already the worst place for it.
Eating in front of the phone looks a lot like holding your mouth open. Attention gating helps, because it only counts when you are looking at the screen, but if you are tasting and reading at the same time, the step will scroll.
Running the front camera costs battery, and the camera indicator stays visible throughout. Some people will not want either of those for a forty-minute recipe, which is reasonable. And while the wake word listens on-device, the assistant that answers needs a connection: the gestures keep working when it drops; the questions do not.
The five-minute version
Prop the phone facing you, within arm's reach
Every face control depends on this one physical fact.Run the setup once and actually perform the gestures
It confirms left and right on your device, and tells you in seconds whether winking works for you.Use winks for forward and back
The most frequent request should be the cheapest one to make.Hold your mouth open for long steps
A held control for a continuous request, rather than a jump.Save your voice for the real questions
Substitutions, quantities, "what did that step say", the things only language can ask.
The point is not that gestures are clever. It is that a recipe should cost you as little attention as possible while something is on the heat, and most of what you ask it is too small to deserve a sentence.
Sugo is in open beta on iPhone. If you have ever shouted "next step" at a phone over an extractor fan, join the beta.