Silent Speech at 26 Percent Word Error Rate
SoniSpeech drops the first open-vocabulary silent speech dataset for wearable eyewear - the missing piece for conversational coaching.
Meta gave the coaching wearable silent typing last month - Neural Handwriting on Ray-Ban Display lets you trace letters via EMG without opening your mouth, and it is a genuine milestone. But typing is not talking, and the coaching wearable needs you to talk.
TL;DR
- SoniSpeech, published August 2026, is the first large-scale open-vocabulary dataset for silent speech recognition on wearable eyewear - 34 hours of tri-modal data across 18,000 utterances with 5,356 unique words, captured via ultrasound acoustic sensing built into glasses frames
- EMG-based silent typing lets you input text to an AI coach, but coaching is conversational - text input strips the emotional nuance that makes interventions work
- Just-in-time adaptive interventions (JITAIs) - coaching nudges timed to moments of vulnerability - require rich dialogue about emotional state and context rather than keyword commands
- The gap between silent typing and silent speech is the gap between a search bar and a therapist, and the dataset to close it now exists
Why Does Silent Typing Fall Short for Coaching?
Silent text input was the right first step. The coaching wearable needs to accept user responses during emotionally charged moments without breaking social context, and EMG handwriting solves that input problem - at the wrong level of abstraction.
When you are mid-argument with your partner - stress climbing, tone shifting, the wearable detecting the escalation pattern it has seen before - and the coach nudges you to pause, what do you need to say back? Not "help." Not "breathing exercise." Something like: "she just brought up the money thing again and I can feel myself shutting down like last Tuesday."
That sentence carries emotional texture, temporal reference, and relationship context that only natural language can encode. Just-in-time adaptive interventions - coaching nudges delivered at the precise moment of vulnerability and receptivity - require exactly this kind of input to close the feedback loop, because the nudge is only half the exchange: the JITAI reads your response, adjusts its model of your state, and calibrates the next intervention. Typing "stressed" gives it a keyword. Speaking your actual experience gives it something to coach with.
The clinical evidence keeps landing in the same place. AI coaching without emotional attunement performs no better than a control group. What separates human coaching from AI coaching is the ability to read nuance in how someone describes their experience, and that read requires speech-level input.
What Is SoniSpeech and Why Does It Change the Constraint?
SoniSpeech is the first large-scale, open-vocabulary, tri-modal dataset for wearable silent speech interfaces using acoustic-sensing eyewear. Published in August 2026, it provides 34 hours of synchronized data across 18,000 utterances.
The sensing modality is ultrasound echo profiling: the glasses emit inaudible sound pulses and capture the echoes reflected by facial muscle movements during silent articulation. There are no EMG electrodes on skin and no camera watching lips - the ultrasound transducers sit in the frame of eyewear you are already wearing. The dataset captures three synchronized modalities - ultrasound echo profiles, voiced audio, and frontal video - in both voiced and silent modes, with full phoneme coverage across 5,356 unique words drawn from conversational English dialogue.
A baseline CTC-ResNet-34 model achieves 26.3% word error rate on open-vocabulary silent speech recognition. For context, production voiced ASR sits below 5% WER. The gap is real, but it is now a measurable engineering problem with a public benchmark instead of an open research question with no training data.
Two weeks ago I said silent input would expand from letters to gestures to sub-vocal speech, and I assumed the path ran through Meta's EMG programs. The first public benchmark for the last step arrived faster than that, and through ultrasound in the frame rather than electrodes on the wrist.
When I worked on Ray-Ban Meta smart glasses, the hardest interaction design problem was finding input modalities that did not break social context. Voice worked for commands. It failed for anything private. EMG handwriting solves part of that problem - you can type without speaking. Ultrasound silent speech solves the rest - you can talk without sound. For coaching, which is a conversation, that distinction is everything.
Why Does the Coaching Loop Need Conversation?
The ambient coaching wearable runs a continuous feedback loop: sense context, detect a coaching moment, deliver an intervention, capture user response, update the model. Every element of this loop has a viable technical path today except one - capturing a nuanced user response during the moments that matter.
Motivational interviewing - the gold-standard therapeutic framework for behavior change - is built on open-ended questions, reflective listening, and change talk that emerges from dialogue. The therapist does not issue directives; the client articulates ambivalence, the therapist reflects it back, and the client's own language drives the shift. An AI coach running motivational interviewing needs to hear how you describe your experience, and a search query does not carry that.
Silent speech on eyewear closes this loop without the social cost. The wearable senses context through its always-on microphone and physiological sensors, identifies the coaching moment, and delivers a nudge through the lens or bone conduction. The user responds in full natural language - silently, by mouthing words while wearing glasses nobody else notices.
The 26.3% WER is nowhere near production today. But research baselines compress fast when public benchmarks and large datasets exist, and the first team to push wearable silent speech below 10% WER on conversational dialogue will own the input layer of the coaching wearable.
What This Means for Investors
The coaching wearable market has been stuck on a last-mile interaction problem. Sensing is solved (always-on mics, physiological sensors, speaker diarization at conversation speed), intelligence is solved (on-device models running on wearable NPUs), and nudge delivery is solved (lens displays, bone conduction, haptics). The user response channel - the one that lets the AI actually coach rather than broadcast - has been missing.
EMG typing was half an answer. Silent speech is the full answer: it lets the user have a conversation with their AI coach while sitting across from their partner, their boss, or their teenager.
SoniSpeech is the training data, and the dataset itself is public. The moat is the model trained on it, refined with proprietary coaching dialogue data, and deployed on hardware with the ultrasound sensor stack already in the frame. That combination - silent speech model plus coaching intelligence plus sensor-integrated eyewear - is the product the ambient wearable market has been circling without being able to ship. Silent speech now has a public number attached, and public numbers get beaten.