Workshop
Spatial intelligence is not a property of a model. It's a property of a system.
I built this as a personal fun project to explore building 3D objects like rockets using only voice and hand gestures, similar to how Tony Stark talks to JARVIS in the movies. It starts with only basic building blocks like a cone. The complexity is in teaching the system what more complex 3D objects mean, like a nozzle, and then composing the nozzle into a rocket.
A multi-agent spatial perception framework. You say "build me a rocket" and three agents assemble and self-correct a 3D structure: Builder creates, Owl evaluates what it looks like, Mechanic corrects the geometry, and the loop runs until it converges. No keyboard, no mouse, and no model trained on 3D data.
"No model in the system is trained on 3D data. Spatial accuracy comes from the interaction between semantic, visual, and geometric agents, which is close to how embodied cognition describes it in people."

Multi-Agent Architecture
Multi-agent cycle: no single agent understands geometry — accuracy emerges from interaction. Issue convergence: 3→1 in 2 iterations for well-formed meshes.
Three Agents, Separate Senses
The loop
- Builder decomposes the spoken request into parts and says which connector port joins which. It never computes coordinates; the canvas aligns the ports
- Owl looks at front, side, and top renders with Claude Vision, alongside scene-graph measurements and the original request, and evaluates every part before giving a verdict
- Mechanic receives the Owl's evaluation, exact scene-graph geometry, OpenCV three-view estimates, and the history of fixes already tried, and corrects rotation, then scale, then position
The result
- Correction loop went from 3 open issues to 1 in 2 iterations
- Voice in through Deepgram, hand tracking through MediaPipe
- 23 parametric component types (cone, cylinder, fin, nozzle, gear, bearing and more), each with typed connector ports
- Up to five capture, measure, evaluate, correct iterations per build
What Broke, and What It Taught Me
- Mixed perceptual channels caused the convergence failures. Oscillating rotations and false approvals traced back to agents sharing inputs: when the correcting agent is handed screenshots and numbers together, the image tends to win even when the numbers are exact. The Owl now owns the visual judgment, the Mechanic is told to trust measurements and CV estimates over its own impression, and repeated corrections are filtered so it cannot loop. Cutting the Mechanic down to numeric input only is the next experiment.
- Generated meshes were the wrong foundation. The first version asked a text-to-3D service for each part. Pivoting to parametric primitives with typed connector ports made assembly a geometry problem the system could reason about, and the rocket came out right on the first build.
- The evaluators never saw the request. Owl and Mechanic were judging the scene without the user's original words, so they approved things that were well-formed but wrong.
- A naming mismatch silently disabled triangulation for two days. Frames were saved as frame_0/1/2 while the OpenCV step looked for front/side/top. Nothing errored; the geometry channel was simply empty.
- Titles vs. IDs. Builder passed node titles while the connection code looked up by ID, so auto-align never ran. One resolver function fixed a bug that took two days to see.