So far the press has printed both answers. Engadget and Digital Trends say Hologram, the voice-driven face Meta is bringing to WhatsApp calls from Meta Ray-Ban Display, runs on Muse Realtime Avatar, and techmymoney says it does not. None of them cites Meta, which has named no model.
Part 5 of six on the Connect 2026 documents; part 4 traced the glasses' memory to the key that unlocks it.
A face generated in chunks from a live voice must trail or delay it; at Meta's published Muse Realtime Avatar chunk size, either costs at least a third of a second (derived).
Meta's research post on Muse Realtime Avatar (opens in a new tab) is exact: 25 fps video in causal chunks from Muse Realtime Voice's speech tokens, eight frames or 320 ms each, drawn in 20 ms on a GB200. It measures about 870 ms from the end of your turn to the first byte of synced voice and video. An agent's speech tokens exist before its audio plays, so the draw can hide in that wait (derived).
Why can't a chunked face keep up with a live caller?
Because a chunked generator has to hear each chunk before it can draw it. One 20 ms step draws all eight frames once the eighth frame's audio is in, so the first frame's lips trail the caller by 340 ms before any network, or the voice waits as long (derived: 320 + 20, assuming no lookahead).
Late lips mean sound leads picture, the stricter direction: ITU-R BT.1359-1 (opens in a new tab) accepts only 90 ms of it (1998 television tests, not calls). A 340 ms lag is almost four times that. Even half a chunk, one latent frame in Meta's architecture figure, is 160 ms, or 180 with the draw (derived), twice the line. Holding the voice 340 - 90 = 250 ms so lips trail by 90 leaves 150 ms of ITU-T G.114's 400 ms mouth-to-ear planning ceiling for everything else (derived).
Meta's Hologram announcement (opens in a new tab) gives the model one line: it "infers your expression from what you're saying and how you're saying it." A voice-inferred face is arguably emotion recognition, so a correction: the August 1 post on emotion inference put the EU's high-risk obligations at August 2, 2026. The AI Omnibus, Regulation (EU) 2026/1744, had moved them to December 2, 2027. It entered into force July 27, five days before I wrote it.
With Seoul National University, Meta's Codec Avatars Lab published a frame-by-frame audio-to-face model (opens in a new tab) (arXiv:2510.01176) that uses no future audio and takes 10 ms a frame on an A100, though Meta has never linked it to Hologram. Add one frame of audio, 33 ms at the model's 30 fps, and the model floor is 33 + 10 = 43 ms (derived). That is a floor, not a call latency: the paper reports no decode time for its 3D headset avatar and defers on-device quantization.
My position: an agent's face can borrow time from its own turn; a caller's cannot. On June 30, 2027, Display Hologram will still be in Early Access on WhatsApp, and Meta will not have published an end-to-end Hologram call latency. What would change my mind is a Hologram call with lips within 90 ms of a voice arriving no later than on a plain WhatsApp call. Frame-by-frame models are the calling asset; chunked agent avatars lose that market.