← Back

Workshop

Spatial intelligence is not a property of a model. It's a property of a system.

Sole DeveloperApr 2026 – PresentResearch prototypeGitHub

I built this as a personal fun project to explore building 3D objects like rockets using only voice and hand gestures, similar to how Tony Stark talks to JARVIS in the movies. It starts with only basic building blocks like a cone. The complexity is in teaching the system what more complex 3D objects mean, like a nozzle, and then composing the nozzle into a rocket.

A multi-agent spatial perception framework. You say "build me a rocket" and three agents assemble and self-correct a 3D structure: Builder creates, Owl evaluates what it looks like, Mechanic corrects the geometry, and the loop runs until it converges. No keyboard, no mouse, and no model trained on 3D data.

"No model in the system is trained on 3D data. Spatial accuracy comes from the interaction between semantic, visual, and geometric agents, which is close to how embodied cognition describes it in people."
Three parts, a nose cone, a fuselage and an engine nozzle, appear apart on a 3D grid and then snap together into a rocket
Port-based assembly in the real app: parts are created apart, then each connection snaps the target part so the two ports meet. This clip replays the canvas calls directly, without voice or the agents.

Multi-Agent Architecture

iterate ≤5×Voice InputDeepgram STTBuilder AgentParametric PartsOwl (Vision)Mechanic (Geo)

Multi-agent cycle: no single agent understands geometry — accuracy emerges from interaction. Issue convergence: 3→1 in 2 iterations for well-formed meshes.

Three Agents, Separate Senses

The loop

  • Builder decomposes the spoken request into parts and says which connector port joins which. It never computes coordinates; the canvas aligns the ports
  • Owl looks at front, side, and top renders with Claude Vision, alongside scene-graph measurements and the original request, and evaluates every part before giving a verdict
  • Mechanic receives the Owl's evaluation, exact scene-graph geometry, OpenCV three-view estimates, and the history of fixes already tried, and corrects rotation, then scale, then position

The result

  • Correction loop went from 3 open issues to 1 in 2 iterations
  • Voice in through Deepgram, hand tracking through MediaPipe
  • 23 parametric component types (cone, cylinder, fin, nozzle, gear, bearing and more), each with typed connector ports
  • Up to five capture, measure, evaluate, correct iterations per build

What Broke, and What It Taught Me

  • Mixed perceptual channels caused the convergence failures. Oscillating rotations and false approvals traced back to agents sharing inputs: when the correcting agent is handed screenshots and numbers together, the image tends to win even when the numbers are exact. The Owl now owns the visual judgment, the Mechanic is told to trust measurements and CV estimates over its own impression, and repeated corrections are filtered so it cannot loop. Cutting the Mechanic down to numeric input only is the next experiment.
  • Generated meshes were the wrong foundation. The first version asked a text-to-3D service for each part. Pivoting to parametric primitives with typed connector ports made assembly a geometry problem the system could reason about, and the rocket came out right on the first build.
  • The evaluators never saw the request. Owl and Mechanic were judging the scene without the user's original words, so they approved things that were well-formed but wrong.
  • A naming mismatch silently disabled triangulation for two days. Frames were saved as frame_0/1/2 while the OpenCV step looked for front/side/top. Nothing errored; the geometry channel was simply empty.
  • Titles vs. IDs. Builder passed node titles while the connection code looked up by ID, so auto-align never ran. One resolver function fixed a bug that took two days to see.

Stack

Next.jsTypeScriptReact Three FiberThree.jsZustandClaude APIClaude VisionOpenCV (Python)MediaPipe HandsDeepgram