Skip to content

Generative objects: summon AI-generated objects onto the surfaces around you - #405

Merged
dli7319 merged 55 commits into
google:mainfrom
salmanmkc:feat/generative-object
Oct 3, 2026
Merged

dli7319 merged 55 commits into
google:mainfrom
salmanmkc:feat/generative-object

Conversation

@salmanmkc

@salmanmkc salmanmkc commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

Description

opening demos/generative_object/ against the current SDK stopped before startup. Clear could bring a cancelled object back when Gemini responded, removing the demo left objects and controls alive, and entering a key reloaded the page with that key in the request URL.

ports the demo to built-in spatial UI and xb.manipulation, aligns three.js with 0.186.0, and keeps prompt-to-object generation demo-owned. buttons or voice request an image, remove its background and place the cutout on a nearby depth surface, with an in-front fallback. the SDK additions are pure image-processing and placement helpers only, not a Core/Options generative subsystem. optional 2.5D relief still uses brightness as an approximation, not measured depth.

Clear and teardown now invalidate work before the AI request and after image decoding. old completions cannot place objects or overwrite a replacement demo's status. both scripts release owned objects, speech handlers, DOM controls and lights; UI is released through the normal lifecycle. each cleanup step runs independently if another release throws.

the global depth-raycast override was deleted, not restored during teardown. current pointer filtering already ignores the depth mesh, while direct placement raycasts still hit it. occlusion requires both depth and occlusion to be enabled, including when a disabled Depth is registered. each object tracks every compiled shader variant and removes all registrations on disposal.

key entry now passes through options.ai.gemini.apiKey, without navigation or adding it to the URL. legacy URL input is consumed and stripped before resource loading, and referrers only carry the origin, so website-restricted keys still work. the overlay now says plainly that the initial URL-key request may already have been logged. demo-local and repository-root keys.json files use the SDK's {gemini: {apiKey: "..."}} shape, and key resolution finishes before initialization. localStorage and client-side keys remain local-prototype conveniences only; production apps must proxy AI calls through a server they control.

Coverage

red-first evidence is against #405 after merging c8d86793, not against unmodified main, because main has no generative demo. generation/ownership tests exposed the old behavior with a temporary legacy-enum shim to get past its removed import. UI teardown and stale-status tests were red after the API port but before the lifecycle fixes. real Chromium also reproduced the unchanged overlay ignoring a valid SDK keys file and navigating with the typed key.

preservation coverage is separate: the existing image-processing, sizing and facing helpers, already-correct cancellation during image decoding, distinct-texture disposal, current semantic UI composition, and direct depth placement while shared pointer hits stay disabled. there were 19 helper tests before the repair, not 45; 31 demo tests were added, for 50 focused tests at the time. the voice, page and startup checks since bring it to 94 (20 helper, 74 demo). the SDK build, full suite, lint, formatting and demo/test type checks passed locally.

Browser checks and limits

Chromium loaded the freshly built SDK and compiled demo. manual, stored, local-file, root-file and legacy URL keys reached the expected options without another navigation or subsequent key-bearing request/referrer. fixture-image decoding and generation, Clear during an AI wait, and removal/recreation of both scripts passed without module exceptions. live Gemini requests were blocked during the smoke test.

re-checked after merging main at e2c6ae64. only change needed was the importmap bump to three 0.186.0. the latest SDK calls Object3D.dispose(), which 0.185.0 doesn't have, so removing the demo threw from every spatial UI script's dispose. 0.186.0 is also what the SDK and every other demo load now. also dropped updateFullResolutionGeometry: placement already raycasts the SDK's downsampled depth mesh and occlusion reads the depth texture, so the hidden full-res mesh was being rebuilt (~23.7k vertices) every depth frame for nothing. with a fixture Gemini response in the simulator: summon from the DOM and spatial buttons, background keying, grounding on the simulated table, occlusion shader registration, Clear during a pending request, relief, mouse-drag translate, and removal/recreation of both scripts all work, no page errors. SDK build, full suite, lint, formatting and demo/test type checks pass again.

voice now goes through Gemini like roomcraft: Speak records one clip (30 s / 4 MB cap), a second tap, or the 30 s limit, sends it to the same key for a transcript, then summons it. the Gemini client is pinned to 2.7.0, like roomcraft. the SDK's browser speech recognizer is off, since Web Speech doesn't work in Quest Browser and the old button sat on "listening..." forever. Clear also drops any recording or transcription in flight.

spatial text renders cleanly in this Chromium run, and so does untouched templates/01_spatial_ui/, so the earlier fragmentation doesn't reproduce anymore. no unrelated SDK UI changes were made.

summon and Speak were checked on a Meta Quest over a local HTTPS preview. real hand/controller dragging and live AI output quality weren't checked beyond that.

Type of Change

  • Bug fix
  • New feature / enhancement
  • New demo or sample
  • Documentation update

Media / Screen Recordings & Screenshots (If Applicable)

  • Simulator Recording: Not attached. Chromium smoke results are described above.
  • Device Recording: Not attached. summon and Speak were checked by hand on a Meta Quest.

Checklist

  • Tested in simulator & device: Verified functionality in desktop simulator and/or physical hardware (where applicable). Chromium desktop simulator with fixture AI, plus a manual check of summon and Speak on a Meta Quest over a local HTTPS preview, with the limitations above.
  • Large Assets ($\ge$ 1MB): Submitted separately to xrblocks/proprietary-assets via jsdelivr CDN. N/A: no assets of this size added.
  • SDK Dynamic Dependencies: All new SDK dependencies are dynamically loaded at runtime. N/A: no new SDK dependencies; three.js remains external.
  • Security: Confirmed no hardcoded API keys or secrets are committed.

salmanmkc added 23 commits June 23, 2026 12:45
A prompt becomes a placed, draggable object: GenerativeObjects.imagine()
asks the AI image model to generate an image, decodes it into a texture,
and drops a billboard into the scene in front of the user, occludable by
real-world depth. Pure helpers (aspect-preserving scale, place-in-front
pose) and the orchestration are unit-tested with mocked AI + texture
source. Wired into Core/Options via enableGenerativeObjects() and exported
from xrblocks.ts.
Add keyOutBackground (pure, tested) and a browser CanvasBackgroundTextureSource
that decodes the generated image, keys out the plain background, and returns a
CanvasTexture so the subject reads as a cutout rather than a flat card. Gated by
GenerativeOptions.removeBackground (on by default); Core swaps in the canvas
source when enabled.
Speak or pinch to summon an AI-generated object into your space via
xb.core.generative.imagine(); generated subjects are keyed to cutouts,
placed in front of you, draggable, and occluded by real depth. Voice
trigger via SpeechRecognizer; pinch cycles preset prompts so it works
without a mic.
Add enableGenerativeObjects() to the options list, an xb.core.generative.imagine
usage snippet, and a generative/ directory-map entry.
Add @google/genai to the importmap (the demo failed to init Gemini without
it) and a key-entry overlay that resolves the key from ?key= > localStorage >
keys.json > a prompt, matching the world_companion/objects_3d demos.
Every pinch/click summoned a new object, which fought with grabbing an
existing one. Track whether a select started on an existing generative
object and, if so, let DragManager move it instead of summoning. Add a
keyboard 'G' summon for desktop where dragging uses the mouse.
Add a quaternionFacingCamera helper and a GenerativeObjects.update() that
turns tracked objects to face the user each frame, gated by the new
GenerativeOptions.billboard flag (on by default). Keeps the flat cutout from
looking paper-thin from the side. Pure helper + billboard behavior unit-tested.
Add a netblocks-styled 🎙️ push-to-talk button (speech was undiscoverable
before) and move the status HUD to the top-left so it no longer collides
with the simulator's settings gear.
Add an opt-in GenerativeOptions.relief that builds the object as a densely
subdivided plane displaced by the generated image's brightness (three.js
displacementMap + bumpMap on a lit standard material) instead of a flat
cutout, giving real shaded surface relief. Approximate (brightness is not
true depth) and needs a light in the scene; default off. Structure unit-tested.
Press R to switch subsequently summoned objects between flat cutout and
2.5D relief (pausing billboarding so you can orbit the relief), and add
ambient + directional lights so the lit relief material shows shading.
Raycast the camera forward against the depth mesh and place the object
there: stand it on horizontal surfaces, float it off vertical ones so it
doesn't blend into walls, falling back to in-front-of-camera with no hit.
Also opt the material into the occlusion shader (the layer alone only
builds the mask) so it's hidden behind real geometry.
Add a draggable uiblocks control panel (summon/speak/relief/clear) that
head-leashes to follow the user, plus a top-right on-screen button bar and a
push-to-talk voice button. Summoning is now via the controls/voice only
(removed click-to-spawn), enable spatial UI + the depth texture for
occlusion, and use the 'flare' icon for summon.
draggable=true alone wasn't enough: DragManager.beginDragging bails when
there's no draggingMode, so grabbing never started. Set draggingMode to
TRANSLATING.
@dli7319
dli7319 self-requested a review June 26, 2026 22:19
@ruofeidu ruofeidu self-assigned this Jun 26, 2026
@ruofeidu ruofeidu added the demo New demo for XR Blocks demonstrating novel interactivity or perception features. label Jun 26, 2026
@ruofeidu

Copy link
Copy Markdown
Collaborator

Hi Salman,

Thank you for your contribution in this!!! I would like to request to switch to a demo.
Also the demo ignores keys.json file I used to debug locally.

I won't say this is ready to be put inside the SDK (for now).
Even for inside SDK, it should be under ai.generateBillboard(image), etc.

We need to carefully think of the high-level picture of generativeAssets:

abstract --------> photorealistic
photo vs. mesh vs. 3D Gaussians etc.
LLM / local model vs. cloud models

Internally, we have a demo like this, but with better quality & confidential tech :)

@ruofeidu
ruofeidu marked this pull request as draft June 26, 2026 22:31
@ruofeidu
ruofeidu self-requested a review June 26, 2026 22:31
@ruofeidu ruofeidu added the algorithm spatial algorithm label Jun 26, 2026

@ruofeidu ruofeidu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

upgrade this to the latest SDK and see if it still works, thanks for the update

the latest SDK calls Object3D.dispose(), which three 0.185.0 doesn't
have, so removing the demo threw from every UICard, UIText, UIButton,
UIPanel, FollowHead and FaceCamera dispose. 0.186.0 is also what the
SDK and every other demo on main load.
…jects

placement raycasts already go through the SDK's 40x40 downsampled
depth mesh and occlusion reads the depth texture, so the hidden
full-resolution mesh was being unprojected (~23.7k vertices) on the
main thread every depth frame for nothing. also corrects the comment
that said placement hit the full-resolution surface.
typescript port of roomcraft's GeminiVoice.js: one bounded recording
(30 s / 4 MB) sent to the configured Gemini key for a JSON transcript,
with the mic released on finish, cancel or dispose. drops roomcraft's
review and reconnect checks since this demo sets Gemini up once at
startup. keeps its 4096 output tokens, since thinking tokens count
toward that limit, and calls isAvailable() before reading the client
because Gemini only creates it there.
the Speak button used the browser's speech-recognition service, which
doesn't work in Quest Browser, so the demo sat on "listening..."
forever. Speak now starts a recording, a second tap sends it to Gemini
and summons the transcript, and the status says what's happening or
why the mic couldn't start. turns the SDK recognizer off, like
roomcraft.
Clear already discards late images, but a recording or transcription
in flight could still summon afterwards, and a live recording kept the
mic on until it sent itself at the 30 s limit.
the key screen only mentioned image generation. it now says Speak
records what you say and sends it to Gemini to transcribe.
@dli7319
dli7319 removed their request for review September 24, 2026 18:17
no-referrer also stripped the Referer from Gemini calls, so a key
limited to this site's address got rejected. strict-origin sends only
the origin, which still never carries a ?key= value.
the page loaded whatever @google/genai was newest while holding a key
and mic recordings. 2.7.0 is what the SDK is built against and what
roomcraft pins.
runs the real key-stripping script from index.html against a stub
window and checks it comes before anything loads, pins the referrer
policy and the three/genai versions to package.json, and asserts the
reticle, depth, depth texture and occlusion options start() depends
on.
poseInFrontOfCamera and quaternionFacingCamera produce world-space
results, which only line up for direct scene children. says so, and
adds a test with the camera under a moved, rotated rig.
@salmanmkc
salmanmkc marked this pull request as ready for review September 25, 2026 07:23
@salmanmkc

Copy link
Copy Markdown
Contributor Author

upgrade this to the latest SDK and see if it still works, thanks for the update

merged latest main and it still works since I fixed latest apis here already, only real fix needed was bumping three to 0.186, since the sdk now calls Object3D.dispose() and 0.185 doesn't have it.

Tried on a quest, speaking now goes through gemini like roomcraft does, since web speech doesn't work in the quest browser. Ready for another review.

@dli7319 dli7319 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tested it on my Galaxy XR and it works. Occlusion is working as well!
There are still perf issues on Galaxy XR, but likely not specific to this demo.
Can you check these few nits?

Comment thread src/generative/GenerativeObjectUtils.ts Outdated
cameraPosition: THREE.Vector3,
target = new THREE.Quaternion()
): THREE.Quaternion {
const awayFromCamera = new THREE.Vector3().subVectors(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this a temporary vector to avoid allocations?
Or maybe just use target as a temporary vector?

Comment thread src/generative/GenerativeObjectUtils.ts Outdated
position = new THREE.Vector3(),
quaternion = new THREE.Quaternion()
): {position: THREE.Vector3; quaternion: THREE.Quaternion} {
const forward = new THREE.Vector3();

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you make this a temporary vector to avoid allocations?
Or maybe just use target as a temporary vector?

Comment thread demos/generative_object/src/main.ts Outdated
const card = new xb.UICard({
size: {width: 0.62, height: 0.24},
manipulation: {
actions: {translate: {faceCamera: false}},

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this should be true.
Right now, it rotates only after we finish dragging it.
And then you can drop the FaceCamera component.

…gging

turns on faceCamera for the card's translate action so it rotates during the drag instead of after, and drops the separate FaceCamera component.
@salmanmkc

Copy link
Copy Markdown
Contributor Author

I tested it on my Galaxy XR and it works. Occlusion is working as well! There are still perf issues on Galaxy XR, but likely not specific to this demo. Can you check these few nits?

thanks for testing! fixed all three nits, temp vectors + faceCamera: true with FaceCamera dropped.

@dli7319

dli7319 commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Thank you!

@dli7319
dli7319 merged commit 863004a into google:main Oct 3, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

demo New demo for XR Blocks demonstrating novel interactivity or perception features.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants