Skip to content
Back to Blog

Lip-Sync Debt

Meta's avatar model draws 320 ms chunks that an agent's turn can hide. Driven by a live caller's voice, a chunk would put lips at least 340 ms late.

Evyatar Bluzer4 min read

So far the press has printed both answers. Engadget and Digital Trends say Hologram, the voice-driven face Meta is bringing to WhatsApp calls from Meta Ray-Ban Display, runs on Muse Realtime Avatar, and techmymoney says it does not. None of them cites Meta, which has named no model.

Part 5 of six on the Connect 2026 documents; part 4 traced the glasses' memory to the key that unlocks it.

A face generated in chunks from a live voice must trail or delay it; at Meta's published Muse Realtime Avatar chunk size, either costs at least a third of a second (derived).

Meta's research post on Muse Realtime Avatar (opens in a new tab) is exact: 25 fps video in causal chunks from Muse Realtime Voice's speech tokens, eight frames or 320 ms each, drawn in 20 ms on a GB200. It measures about 870 ms from the end of your turn to the first byte of synced voice and video. An agent's speech tokens exist before its audio plays, so the draw can hide in that wait (derived).

Why can't a chunked face keep up with a live caller?

Because a chunked generator has to hear each chunk before it can draw it. One 20 ms step draws all eight frames once the eighth frame's audio is in, so the first frame's lips trail the caller by 340 ms before any network, or the voice waits as long (derived: 320 + 20, assuming no lookahead).

Muse Charm and its avatar timingPhoto of a hand holding the Muse Charm keychain device with a character on its screen; the screen is outlined in the accent color and joined by a leader line to the first of three text notes in the left margin, and the other two notes point at no part of the device. The agent, drawn on Charm's screen Per TechCrunch, Zuckerberg said Charm packs the real-time voice and avatar stack into a keychain Driven by the agent's own voice Per Meta: 8 frames = 320 ms, drawn in 20 ms; about 870 ms from the end of your turn to the first byte. Derived: a chunk can be drawn ahead in that wait If a live caller's voice drove it (derived) 320 ms of audio must arrive, then 20 ms to draw: at least 340 ms of lip lag, or of held audio
Muse Charm with the Muse agent on its screen; Meta has published no Charm spec sheet, the timing numbers come from Meta's research post on Muse Realtime Avatar, and the 340 ms floor at that published chunk size is derived from them. Photo: Meta.

Late lips mean sound leads picture, the stricter direction: ITU-R BT.1359-1 (opens in a new tab) accepts only 90 ms of it (1998 television tests, not calls). A 340 ms lag is almost four times that. Even half a chunk, one latent frame in Meta's architecture figure, is 160 ms, or 180 with the draw (derived), twice the line. Holding the voice 340 - 90 = 250 ms so lips trail by 90 leaves 150 ms of ITU-T G.114's 400 ms mouth-to-ear planning ceiling for everything else (derived).

Meta's Hologram announcement (opens in a new tab) gives the model one line: it "infers your expression from what you're saying and how you're saying it." A voice-inferred face is arguably emotion recognition, so a correction: the August 1 post on emotion inference put the EU's high-risk obligations at August 2, 2026. The AI Omnibus, Regulation (EU) 2026/1744, had moved them to December 2, 2027. It entered into force July 27, five days before I wrote it.

With Seoul National University, Meta's Codec Avatars Lab published a frame-by-frame audio-to-face model (opens in a new tab) (arXiv:2510.01176) that uses no future audio and takes 10 ms a frame on an A100, though Meta has never linked it to Hologram. Add one frame of audio, 33 ms at the model's 30 fps, and the model floor is 33 + 10 = 43 ms (derived). That is a floor, not a call latency: the paper reports no decode time for its 3D headset avatar and defers on-device quantization.

My position: an agent's face can borrow time from its own turn; a caller's cannot. On June 30, 2027, Display Hologram will still be in Early Access on WhatsApp, and Meta will not have published an end-to-end Hologram call latency. What would change my mind is a Hologram call with lips within 90 ms of a voice arriving no later than on a plain WhatsApp call. Frame-by-frame models are the calling asset; chunked agent avatars lose that market.

Comments