What the Puck Can Afford to Think
Qualcomm's claim of a 3B model at 45 tokens a second on the VR Glasses chip needs the puck's entire memory bus, by my arithmetic.
Qualcomm's overview graphic for Snapdragon Reality Elite, the chip in the puck of Meta's VR Glasses, prints two numbers under "On-device AI": "3B Parameter on-device" and "45 Tokens-per-second." Read them as bytes. At 4 bits a weight, a 3-billion-parameter model is 1.5 GB, and a model serving one conversation reads every weight for each token, so 45 tokens a second is 45 x 1.5 = 67.5 GB/s of memory reads. By my arithmetic below, the chip's memory bus carries 67.2 GB/s.
Part 2 of six on the Connect 2026 documents; part 1 argued that the tether's pixel rate forces Meta's own link silicon onto the face.
The short version: Qualcomm's claim of a 3B model at 45 tokens per second on Snapdragon Reality Elite needs 67.5 GB/s of 4-bit weight reads, and by my arithmetic the chip's bus peaks at 67.2 GB/s. Scaled to the VR Glasses puck's bus, Meta's MobileMoE phone timings put a dense 3B model at 13 to 18 tokens per second with the displays idle, by my extrapolation. By my arithmetic, the VR Glasses puck has the bandwidth for a fast command model or a slow 3B one, not an agent that plans before it speaks, which leaves Meta AI's reasoning to Meta's cloud.
How fast is the puck's memory bus?
About 67 GB/s by my arithmetic, reading the brief's memory clock the only way Qualcomm's own phone numbers add up. The Reality Elite product brief lists "4x16 LP-DDR5 memory, up to 4.2 GHz," a 64-bit bus, and prints no bandwidth. Qualcomm's January 2024 generative-AI whitepaper gives the Snapdragon 8 Gen 3 "LPDDR5x at 4.8GHz and 77GB/s," which only adds up if the quoted clock is half the per-pin data rate: 64 bits x 9.6 Gbps / 8 = 76.8 GB/s. Read the same way, Reality Elite is 64 x 8.4 / 8 = 67.2 GB/s.
Qualcomm's 45 is 100.4 percent of the bus. A model of exactly 3 billion weights tops out at 67.2 / 1.5 = 44.8 tokens a second, and if the "Llama 3B" in Fonearena's June AWE report is Llama 3.2 3B, its 3.21 billion parameters put the ceiling at 41.9, under the claim.
That does not make it false. Speculative decoding, where a small draft model proposes tokens and the big one checks several at once, reads the big weights once per several tokens; sub-4-bit layers or a "3B" lighter than its name would also get there. Qualcomm has not said which.
Meta's phone timings, scaled to the puck
Meta's MobileMoE paper (arXiv:2605.27358) times small models on a Galaxy S25, one conversation at a time, and names the limit itself:
Decode is memory-bandwidth-bound: per-step weight reads from RAM also scale with active parameters, so MoE transfers fewer bytes per token, yielding higher throughput.
Its dense baseline, MobileLLM-Pro, holds 0.55 GB of 4-bit weights and decodes 45.8 tokens a second at a 1k-token context: 0.55 x 45.8 = 25 GB/s, 30 percent of the 84.8 GB/s the same half-clock reading gives the phone's Snapdragon 8 Elite (40 percent at 256 tokens). Those are four-thread CPU runs, and the NPU does no better: Meta's MobileLLM-Pro report has a Galaxy S24 decoding 31.6 tokens a second from a 0.72 GB model at 2k context, at most 0.72 x 31.6 = 22.8 GB/s, 30 percent of its bus. At that efficiency, by my extrapolation, a dense 3B on the puck decodes 67.2 x 0.3 / 1.5 = 13.4 tokens a second, or 17.9 at 40 percent, with nothing else reading memory.
Something always is. Meta's compare-devices page lists two 2412x2288 panels at 120 Hz, so scanout alone reads 2 x 2412 x 2288 x 120 x 4 bytes = 5.3 GB/s. Unless the panel scans straight from the app's frames, that is one of up to four passes of that size (the app writes each eye's frame, the compositor reads it and writes the warped result, the panel reads that): about 21 GB/s, near a third of the bus by my count, before any frame-buffer compression the chip applies. Qualcomm's passthrough ceiling, two 12 MP streams at 90 frames a second, writes up to 2 x 12 MP x 90 x 1.25 bytes = 2.7 GB/s more at 10 bits a pixel (Meta publishes no camera resolution). At Qualcomm's token rate, demand reaches 67.5 + 5.3 + 2.7 = 75.5 GB/s, 112 percent of the bus, before the GPU draws anything. So either Qualcomm timed its 45 with nothing else on the bus, or its model read the weights less than once a token, and the graphic says neither.
MobileMoE's catch is its own headline. The 2.2 to 3.4x decode speedup over MobileLLM-Pro belongs to the smallest model, MobileMoE-S: 272 million active parameters of 1.3 billion, 0.68 GB, 112 tokens a second at 1k context. MobileMoE-L, 922 million active of 5.3 billion, decodes 36.6, slower than the dense 1B, though its 0.46 GB of active weights would allow 84.8 / 0.46 = 184 tokens a second. At a fifth of that, the quoted sentence fails on the paper's own table: something other than the bus sets its pace. On this evidence MoE makes the puck's small model faster and does nothing yet for a bigger one.
So where does Meta AI run?
Mostly in Meta's datacenters, by this arithmetic, though Meta has not said. Its announcement lists the agent's jobs: play a movie, open an app, adjust your workspace, search. An agent writes before it speaks: a 100-token tool call, my round number, takes 100 / 17.9 to 100 / 13.4 = 5.6 to 7.5 seconds on the puck's 3B. Those numbers leave room for a router: "open an app" is intent classification with a slot or two, a job for a 0.68 GB model like MobileMoE-S.
I read the puck as a display computer with a command model in it: anything Meta AI does past routing a command goes to Meta's servers. That makes the agent a serving bill, and Meta can pay it. Qualcomm's June release named XREAL and Play for Dream as building on the chip; a maker without its own cloud agent rents one or ships a 3B model at 13 to 18 tokens a second, less with the displays lit, and either way loses the premium local-AI pitch Qualcomm's graphic was drawn for.
That pitch was mine in June, when I repeated the 45 tokens a second as a local capability and predicted a premium tier on Reality Elite with "full on-device inference." That call stays open until late 2027, and I would now bet against its on-device half. The check I skipped applies to any wearable with a local agent: decode rate is bus bandwidth, minus what its sensors and display read, over the weight bytes each token needs.
I am putting dates on two calls. The first, checked when the VR Glasses go on sale (spring 2027, Meta says) or on June 30, 2027, whichever is later: with the network off, Meta AI (or Muse, which the VR Glasses announcement never names) carries out device commands such as opening an app and declines anything past them, and no offline assistant mode ships in 2027. The second, by December 31, 2027: no Reality Elite device ships a built-in assistant whose default model, 3 billion parameters or larger, runs on the device.
Qualcomm's 45 tokens a second describes a chip with nothing else to do. Meta's puck has two 120 Hz panels to feed, so the arithmetic leaves it the commands and sends the thinking to Meta's datacenters.