Repository navigation
feat(demos): add on-device gemma 4 assistant demo - #634
Conversation
b9761cc to
d740865
Compare
|
Tested it on Galaxy XR. The model loads and works. |
Thanks for testing! The freeze is the GPU prefilling the prompt at the start of each reply, fighting with rendering. I pushed a change to cut that down, basically the system prompt gets prefilled while the model loads now, and the scene info only gets re-sent when something changed, so follow-ups are way shorter. Seems a bit quicker now. The first question in a chat still sends the full scene, so there might still be a smaller hitch there. From what I've tested it seems not fully possible to totally get rid of it from my test on my phone, didn't retry the test on a meta quest but can do later today. I also tried feeding the prompt to the GPU in small chunks instead of one big call. On a really long prompt it took the worst frame from tests on M$ from 1.1 s to 0.1 s, but everything else got slower: time to first text went from about 0.65 s to 1.3 s on a scene question and from 0.25 s to about 0.55s on a follow-up, replies dropped from about 30 to 21 tok/s, and on the long prompt first text went from 1.8 s to 5.2 s. It could be worth trying that behind a flag to see if the trade off of average time being longer and lag is worth it, but leaving it out of the pr for now. I could do a follow up PR for you to test it? |
|
Thanks. I'll retest this shortly on Galaxy XR. |
gemma_test_20261005.mp4The performance is still pretty bad on Galaxy XR. |
|
My antigravity found a way to get it working without stuttering. I'll update this PR and submit it shortly. |
oh wow that's awesome! thank you David! |
Keep the large model download opt-in and reuse complete cached weights on later visits.
Keep model initialization and generation off the UI thread while supporting streaming, cancellation, and reset.
Coordinate validated prompts and worker requests without accepting stale replies after cancellation or disposal.
Let users select and move objects, then discuss their metadata through streamed local chat.
Expose the spatial assistant in the browser with pinned SDK peer dependencies and XR entry controls.
Make the model download, offline behavior, licensing, and observed responsiveness explicit.
Make the runnable demo discoverable from the samples navigation.
Name the selected shape directly and round useful metadata so the small model does not have to resolve opaque IDs.
Avoid rebuilding the whole spatial card for each chat entry, which caused repeatable main-thread stalls.
Keep typed questions useful without forcing every answer to describe scene objects or adding keyword routing.
Move initialization out of the immersive session and reuse the loaded worker on entry without another download.
Record measured improvements without claiming freeze-free headset behavior or verified phone compatibility.
Keep replies readable without rebuilding the spatial card or interpreting HTML. Cache formatted messages and preserve streaming indentation in the existing text node.
Avoid a separate shadow shader on first keyboard open using the existing demo-local style options. Keep shared keyboard behavior unchanged.
Keep loading and device guidance short, reflect confirmed phone use, and document the stable Markdown presentation without repeated warnings.
A Quest report had Gemma describe a different object than the one the user believed was selected. The scene context and prompt matched the selection in every desktop and simulator repro, so make the selection the model receives visible: tag each user turn with the selected object's name, read from the same context sent to the worker, and highlight the selected object in its own colour.
The pause when a reply starts is GPU prefill in the worker competing with rendering, not main-thread work. Prefill the system prompt when the conversation is created, and send the scene block again only when it changed since the last completed turn; otherwise send a short "unchanged" marker with the selected object's name. On an M4 in desktop Chrome, the first query's prefill went from 322 to ~195 tokens (first text 848 -> ~650 ms, worst frame 83 -> ~50 ms), and repeat queries from ~195 to ~30 tokens (first text ~480 -> ~250 ms, worst frame <= 34 ms). Answers stayed grounded across reselection.
…ence Split large WebGPU compute passes and command encoders in the Gemma worker into smaller command buffer chunks, replay pipeline and snapshotted bind-group state across split passes, and pace submission with adaptive onSubmittedWorkDone synchronization so the WebXR compositor is not starved during prefill or decode without stalling fast desktop GPUs. Also avoid synchronous WebGL readPixels readback on the main thread when getProfileSummary is unavailable on GpuArtisan.
7607cde to
11c9f48
Compare
Description
for #629. new demo at
demos/gemma_on_device/: Gemma 4 E2B running fully on the device, through LiteRT-LM in the browser. ask it anything, or select and drag the three objects and use the presets to ask about them. the model gets the objects' names, shapes and rounded positions from the scene context API. no camera, no cloud, and the model never runs code.the ~2 GB model only downloads when you press Download, then it's cached in the browser so later visits just load it. you can load it from the 2D page before entering XR, so the load happens outside the headset. inference runs in a dedicated worker. replies render markdown through a pinned
marked(headings, lists, code) into one stable text node, because rebuilding UI blocks per reply caused 250-300 ms frames.phones can enter XR now that #636 made hand tracking optional (it used to fail with NotSupportedError), and this is rebased on top of it.
on an M4 in desktop Chrome: first text in about 0.9 s, ~15 tokens/s, and the longest frame during a reply was 34 ms. works with the network blocked once the model is loaded. tested on an Android phone and it works there. an earlier build ran on Quest 3 but paused a lot, the latest one still needs a Quest retest. rendering and inference share the GPU, so short pauses while loading or starting a reply are possible on standalone headsets.
markedis also added as a pinned devDependency so the markdown tests can import it.Type of Change
Media / Screen Recordings & Screenshots (If Applicable)
Simulator Recording:
markdown reply in the spatial card:
follow-up with the network blocked after loading:
Device Recording:
tested on an Android phone, works. no recording attached yet.
Checklist