← deckwright.ai
Multi Modal Capture

We don't record what you typed.
We record what you meant.

MMC is deckwright's intent-capture layer. While you work on the Bench - our endless spatial canvas, on desktop or iPad - it records the whole living act of creation: what you draw, how your hand moves, what you say and exactly when you say it, and how the thing you're building actually behaves. Those streams fuse into a single, durable record of intent the crew can read, replay, and build against.

Most tools take your click as a coordinate and throw the rest away. MMC keeps the rest. The rest is where the meaning is.

Why we built it this way

Artists don't think in specifications. They think responsively, in real time - gesturing, pointing, reacting to what's in front of them. "More like this. No, here. Bigger. Softer. That one." The intent lives in the motion, the glance, the timing, and the word said at the moment the hand arrives - not in a paragraph written after the fact.

Prompt-based tools force that fluid, embodied way of thinking through the keyhole of a text box. You stop, translate yourself into prose, and hand the machine a flattened description of something you were showing it a second ago. Everything expressive - the hesitation, the decisiveness, the pointing, the pace - is lost in the translation.

MMC refuses that trade. It meets you in your native mode: you gesture and speak and arrange the way you always have, and the system treats that - the actual performance of creating - as the real specification. It's the difference between describing a melody and humming it.

How it works

Four moves: capture, align, decode, keep.

No screenshots. Nothing inferred that can be recorded.

Deckwright does not work from screenshots. A screenshot is a single flattened surface with no time axis: no motion, no order of operations, no hesitation, no voice - everything a model reads from it is inference, and inference is lossy. Multi-Modal Capture never flattens. Each channel of intent is recorded at the source as its own discrete, first-class, timestamped stream, and the streams are fused on one shared clock. Signal-native, not pixel-derived. The result is not a picture of the work to be interpreted after the fact - it is the work's own instrumentation: exact, replayable, and made of what actually happened.

And no vision model in the loop. MMC does not rely on pre-baked vision modalities to understand the work: no screenshots are taken, and no screen-recorded video is parsed by a model's eyes. The capture arrives as structured, machine-native state - exact coordinates, exact timings, exact canvas, DOM, and console events - so nothing is estimated from pixels that the system knew precisely at the source. Three consequences follow: precision, because a cursor position is a number rather than a guess; economy, because structured deltas cost a fraction of the tokens of video frames; and independence, because the record is legible to any model, with or without vision.

What it captures

Far more than the picture. MMC listens on every channel a moment of creation actually happens on:

Why it matters

The paired gesture-and-voice idea is 45 years old. Everyone who ever built it built it as a command - fire it, execute it, forget it. We're the ones who keep it. That one decision - treating the living act of creation as a durable record of intent instead of a disposable input - is what turns a canvas into a context engine, and a sketch into a spec.

Read: metaphor as harness → See our production pipeline →