The Bench walkthrough
What this demonstrates
This is the operational Deckwright The Bench, not a concept animation or a design mockup. Synchronized voice and spatial capture on a live canvas: speech, gesture, and timing parsed into structured intent, image and shape generation directed by pointing and talking, live HTML loaded onto the canvas with behavioral and dev-console capture, and attachment of the results to governed work.
Capabilities demonstrated, with timestamps
- [00:47] A single control starts a full-canvas multi-modal capture session: voice, gesture, and timing recorded together.
- [01:00-02:39] Natural speech plus pointing produces placed, styled canvas elements: an image in a circular mask, shapes indicated by gesture rather than name, and an arc of seven size-graded triangles.
- [01:51-02:23] The system parses the intent streams - gestural content, semantic content, hesitation patterns, captured motions - temporally coordinated and packaged into a persistent payload delivered to an officer for execution.
- [02:39-03:04] Generated elements remain live: draggable, recolorable, animatable; HTML widgets can be created and moved for prototyping.
- [03:04-03:44] Any element, or an entire recorded multi-modal capture, can be attached to the Board or a card, where it becomes a template for building that piece.
- [03:44-04:37] A live HTML page is loaded onto the Bench and captured behaviorally: DOM highlights, inputs, scrolls, hovers, mutations, dev-console errors, and network fetches, all timestamped against user behavior.
- [04:37-04:57] Bugs, buttons, interactions, and scroll positions captured this way attach directly to the production pipeline, ready to be fixed or implemented.
- [05:06-08:58] Image generation directed multi-modally: a box-cover composition described by voice while gesturing regions of the canvas, fused with previously developed project context, produces a usable one-shot result.
- [09:05] The generated result, whole or in pieces, attaches to the production pipeline.
What this walkthrough does not establish
- It does not show a captured artifact carried end-to-end through the pipeline to deployed code within this video (the Board walkthrough shows the pipeline side).
- It does not demonstrate multi-user collaboration on one canvas.
- It shows one image-generation example; it does not establish generation quality across styles or domains.
- It does not demonstrate performance under load or at scale.
Timestamped transcript
[00:00] So, this is the bench. This is the surface where we collect and develop gestural, spatial,
[00:08] interactive context for the project. The really important part of this is not, in fact, the
[00:17] chat interface, but the canvas behind it so that we can minimize the size of the chat
[00:24] interface or completely get rid of it so that we have space to work with the agents
[00:30] here. But of course, if we don't have the chat interface, how are we supposed to work
[00:36] with the agents? This is how it works. Okay, so what I'm going to do is click this red
[00:47] button and that's going to start a full screen, full canvas, multimodal capture. Okay, Mapmaker,
[00:55] here we are. We're going to do some work on this canvas. The first thing I want you to do is
[01:00] generate an image of a dog down here, about that size, in a circular mask. And then up here, I want
[01:10] you to give me a shape like this. Normally, I'd call the name out, but we're going to do,
[01:15] we're going to test you. So give me a shape like that and make it dark blue. And then over here,
[01:22] give me a shape like this and make that one dark red. And then right here, I want starting about
[01:32] here and swooping down like that on an arc. I want a series of seven triangles. They can all be green
[01:41] and I want them to be sized larger to smaller. And the largest one can be, let's say, that big
[01:51] and the smallest one can be that big. Okay. Now the system is parsing the users intent streams,
[02:03] all the modalities, temporarily coordinated. So it's gestural content, semantic content,
[02:08] we use hesitation patterns, captured motions, this is a lot more. It's pretty complex. We package it
[02:16] all up into a persistent payload and we deliver it to the officer to execute. So obviously, usually
[02:23] you just draw or, you know, name those shapes that I named on, or sort of gestured on the canvas,
[02:31] but we really do want to kind of stress the system a little bit like that's easy stuff.
[02:39] Excellent. All right. So these are all individual elements that can be dragged and mapmaker can
[02:49] continue to work on them, change the color position, etc. There's animations that can be done,
[02:55] you can create HTML widgets on this page, move them around, do prototyping. It's super powerful.
[03:04] This is just sort of scratching the surface of it. And of course, any of these things from colors
[03:11] to shapes to, you know, gestural, intense easing animations, it can all be attached either individually
[03:22] or as a single kind of recorded MMC, which would be this piece to the board to a card to anything. So
[03:33] if you have a speed of something approaching something else that you get to be just perfect,
[03:40] you throw it over to the board and that becomes a template for building that piece of your
[03:44] interactive masterpiece. But another thing that you can do is you can load an HTML page onto the
[03:54] bench. And if you do, then you get all the benefits of multimodal capture to debug and design and
[04:02] build. So if you're working on something and in in Deckwright, and it's not working, just load it
[04:08] into the bench and start looking around so you can highlight the DOMs and etc. All of it gets
[04:17] recorded if you start a multi modal capture. So semantic const context, and I'm instantly
[04:25] collecting inputs and scrolls and DOM hovers and mutations, all the data from the dev console,
[04:32] errors, network fetches, everything bundled together and time stamped against user behavior.
[04:37] And again, I can attach any of that. So any bugs, any buttons, any interactivity, any scroll
[04:47] positions, any features can be attached to the production pipeline. And they're, at that point,
[04:57] ready to be fixed or changed or implemented. And another thing you can do with MMC is any image
[05:06] generation. So let's get a new page and get this out of the way. Go to Mapmaker. So one of the
[05:18] things that the officers and I have kind of figured out in terms of making a historically bad game is
[05:24] that a really important piece of it would be to have historically unrelatable player characters or
[05:30] groups of player characters. So we went back and forth on a few and we decided on two groups,
[05:37] weasels and leprechauns, because you know, weasels are chaotic, backstabbing, sneaks,
[05:45] nobody likes them. And leprechauns are vindictive, greedy, baby stealing bitches,
[05:50] or at least that's according to my officers research. So now let's develop a box cover image
[05:57] for our new game. Mapmaker has been listening to all of this, but he already knows everything
[06:02] anyway. So Mapmaker, I'm going to tell you what I want. So over here on this side, we're going to
[06:08] have the weasels just in a general sense. And over here on this side, we're going to have the
[06:13] leprechauns, generally speaking. Between the two of them, there is a table. So this is the edge of
[06:20] the table and then it's seen from below a little bit, right? So you maybe see pieces of it going
[06:25] back. There's a bench here for the weasels, a bench here for the leprechauns. The environment itself
[06:32] is a bar vague, maybe up in this area or around behind, you can see elements of the bar. But the
[06:41] main focus is obviously the weasels and the leprechauns. On the top of the table, let's put some
[06:47] beer. And the weasels, let's say they're each about maybe this tall like that. And the thing
[06:57] with the weasels is they're just, it's just a, it's a chaos machine. They can't cooperate. They're
[07:01] just constantly betraying each other. But they do know that they need to get up here to the beer
[07:06] because that's part of the game. And they need to do it before the leprechauns. And the leprechauns,
[07:12] you know, they're a little better at cooperating, but none of them like each other. You know,
[07:16] they're addictive backstabbers. They do cooperate slightly, but they also trick each other and steal
[07:25] each other's babies. It's very confusing. And they're all kind of, you know, getting up on these
[07:30] benches and trying to get up on the table. The weasels are trying to get up on these benches and
[07:35] trying to get up on this table. And then right at the top, we'll put our, the title of our
[07:43] historically bad game, which is Beer Bust, leprechauns versus weasels. All right, let's see what you got.
[07:55] So we're going to take all of that descriptive context, all the time stamped ideas, the coordinates,
[08:04] and bundle it up with some other slightly more subtle data streams and pump it into
[08:13] the LLMs image generating module.
[08:17] Yeah, that's great. One shot at it, as they say, you know, and that's part of the power of this
[08:33] multimodal input system. The, all the gestural stuff, all the semantic stuff plus the model is
[08:42] drawing on the existing context. So all the previous work that we've done describing the game,
[08:48] breaking down the schemes, discussions about the various unrelatable NPCs, it all
[08:58] comes into play here in addition to the instructions that I give it at the moment.
[09:05] Now, obviously, we can attach any of this to the production pipeline. So the whole thing on its own
[09:18] or any, you know, individual pieces of it, etc. So now let's take a look at the production pipeline.