Skip to content

feat(demos): add on-device gemma 4 assistant demo - #634

Merged
dli7319 merged 18 commits into
google:mainfrom
salmanmkc:salmanmkc-gemma-4-on-device-demo
Oct 6, 2026
Merged

dli7319 merged 18 commits into
google:mainfrom
salmanmkc:salmanmkc-gemma-4-on-device-demo

Conversation

@salmanmkc

@salmanmkc salmanmkc commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Description

for #629. new demo at demos/gemma_on_device/: Gemma 4 E2B running fully on the device, through LiteRT-LM in the browser. ask it anything, or select and drag the three objects and use the presets to ask about them. the model gets the objects' names, shapes and rounded positions from the scene context API. no camera, no cloud, and the model never runs code.

the ~2 GB model only downloads when you press Download, then it's cached in the browser so later visits just load it. you can load it from the 2D page before entering XR, so the load happens outside the headset. inference runs in a dedicated worker. replies render markdown through a pinned marked (headings, lists, code) into one stable text node, because rebuilding UI blocks per reply caused 250-300 ms frames.

phones can enter XR now that #636 made hand tracking optional (it used to fail with NotSupportedError), and this is rebased on top of it.

on an M4 in desktop Chrome: first text in about 0.9 s, ~15 tokens/s, and the longest frame during a reply was 34 ms. works with the network blocked once the model is loaded. tested on an Android phone and it works there. an earlier build ran on Quest 3 but paused a lot, the latest one still needs a Quest retest. rendering and inference share the GPU, so short pauses while loading or starting a reply are possible on standalone headsets.

marked is also added as a pinned devDependency so the markdown tests can import it.

Type of Change

  • Bug fix
  • New feature / enhancement
  • New demo or sample
  • Documentation update

Media / Screen Recordings & Screenshots (If Applicable)

  • Simulator Recording:

    markdown reply in the spatial card:

    Gemma markdown reply

    follow-up with the network blocked after loading:

    Gemma reply with the network blocked

  • Device Recording:

    tested on an Android phone, works. no recording attached yet.

Checklist

  • Tested in simulator & device: Verified functionality in desktop simulator and/or physical hardware (where applicable).
  • Large Assets ($\ge$ 1MB): Submitted separately to xrblocks/proprietary-assets via jsdelivr CDN. (n/a, no weights committed. the Apache-2.0 model downloads from HF at a pinned commit)
  • SDK Dynamic Dependencies: All new SDK dependencies are dynamically loaded at runtime.
  • Security: Confirmed no hardcoded API keys or secrets are committed.

@salmanmkc
salmanmkc force-pushed the salmanmkc-gemma-4-on-device-demo branch 2 times, most recently from b9761cc to d740865 Compare September 26, 2026 10:04
@dli7319

dli7319 commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Tested it on Galaxy XR. The model loads and works.
There's a ~1 second freeze and sequence of frame drops when we start any query tho.

@salmanmkc

salmanmkc commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor Author

Tested it on Galaxy XR. The model loads and works. There's a ~1 second freeze and sequence of frame drops when we start any query tho.

Thanks for testing! The freeze is the GPU prefilling the prompt at the start of each reply, fighting with rendering. I pushed a change to cut that down, basically the system prompt gets prefilled while the model loads now, and the scene info only gets re-sent when something changed, so follow-ups are way shorter. Seems a bit quicker now. The first question in a chat still sends the full scene, so there might still be a smaller hitch there. From what I've tested it seems not fully possible to totally get rid of it from my test on my phone, didn't retry the test on a meta quest but can do later today.

I also tried feeding the prompt to the GPU in small chunks instead of one big call. On a really long prompt it took the worst frame from tests on M$ from 1.1 s to 0.1 s, but everything else got slower: time to first text went from about 0.65 s to 1.3 s on a scene question and from 0.25 s to about 0.55s on a follow-up, replies dropped from about 30 to 21 tok/s, and on the long prompt first text went from 1.8 s to 5.2 s. It could be worth trying that behind a flag to see if the trade off of average time being longer and lag is worth it, but leaving it out of the pr for now. I could do a follow up PR for you to test it?

@dli7319

dli7319 commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

Thanks. I'll retest this shortly on Galaxy XR.

@dli7319

dli7319 commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator
gemma_test_20261005.mp4

The performance is still pretty bad on Galaxy XR.
I'm not sure if this is something that's ready to merge yet.
@salmanmkc Are you seeing the same stutters on Quest?

@dli7319

dli7319 commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

My antigravity found a way to get it working without stuttering. I'll update this PR and submit it shortly.

@salmanmkc

salmanmkc commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

My antigravity found a way to get it working without stuttering. I'll update this PR and submit it shortly.

oh wow that's awesome! thank you David!

salmanmkc and others added 18 commits October 6, 2026 13:48
Keep the large model download opt-in and reuse complete cached weights on later visits.
Keep model initialization and generation off the UI thread while supporting streaming, cancellation, and reset.
Coordinate validated prompts and worker requests without accepting stale replies after cancellation or disposal.
Let users select and move objects, then discuss their metadata through streamed local chat.
Expose the spatial assistant in the browser with pinned SDK peer dependencies and XR entry controls.
Make the model download, offline behavior, licensing, and observed responsiveness explicit.
Make the runnable demo discoverable from the samples navigation.
Name the selected shape directly and round useful metadata so the small model does not have to resolve opaque IDs.
Avoid rebuilding the whole spatial card for each chat entry, which caused repeatable main-thread stalls.
Keep typed questions useful without forcing every answer to describe scene objects or adding keyword routing.
Move initialization out of the immersive session and reuse the loaded worker on entry without another download.
Record measured improvements without claiming freeze-free headset behavior or verified phone compatibility.
Keep replies readable without rebuilding the spatial card or interpreting HTML. Cache formatted messages and preserve streaming indentation in the existing text node.
Avoid a separate shadow shader on first keyboard open using the existing demo-local style options. Keep shared keyboard behavior unchanged.
Keep loading and device guidance short, reflect confirmed phone use, and document the stable Markdown presentation without repeated warnings.
A Quest report had Gemma describe a different object than the one the
user believed was selected. The scene context and prompt matched the
selection in every desktop and simulator repro, so make the selection
the model receives visible: tag each user turn with the selected
object's name, read from the same context sent to the worker, and
highlight the selected object in its own colour.
The pause when a reply starts is GPU prefill in the worker competing
with rendering, not main-thread work. Prefill the system prompt when the
conversation is created, and send the scene block again only when it
changed since the last completed turn; otherwise send a short
"unchanged" marker with the selected object's name.

On an M4 in desktop Chrome, the first query's prefill went from 322 to
~195 tokens (first text 848 -> ~650 ms, worst frame 83 -> ~50 ms), and
repeat queries from ~195 to ~30 tokens (first text ~480 -> ~250 ms,
worst frame <= 34 ms). Answers stayed grounded across reselection.
…ence

Split large WebGPU compute passes and command encoders in the Gemma
worker into smaller command buffer chunks, replay pipeline and snapshotted
bind-group state across split passes, and pace submission with adaptive
onSubmittedWorkDone synchronization so the WebXR compositor is not
starved during prefill or decode without stalling fast desktop GPUs. Also
avoid synchronous WebGL readPixels readback on the main thread when
getProfileSummary is unavailable on GpuArtisan.
@dli7319
dli7319 force-pushed the salmanmkc-gemma-4-on-device-demo branch from 7607cde to 11c9f48 Compare October 6, 2026 20:49
@dli7319
dli7319 merged commit a9c28bc into google:main Oct 6, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants