Current AI writing tools, which rely on text prompts, poorly support the spatial and interactive nature of storytelling where ideas emerge from direct manipulation and play. We present PlayWrite, a mixed-reality system where users author stories by directly manipulating virtual characters and props. A multi-agent AI pipeline interprets these actions into Intent Frames -structured narrative beats visualized as rearrangeable story marbles on a timeline. A large language model then transforms the user's assembled sequence into a final narrative. A user study (N=13) with writers from varying domains found that PlayWrite fosters a highly improvisational and playful process. Users treated the AI as a collaborative partner, using its unexpected responses to spark new ideas and overcome creative blocks. PlayWrite demonstrates an approach for co-creative systems that move beyond text to embrace direct manipulation and play as core interaction modalities.
Alex: Welcome to another episode of ResearchPod. Sam, we've been talking about how AI is changing creative work—what's this paper we're diving into today?
Sam: This is about PlayWrite, a mixed-reality system designed to help people create stories. Instead of just typing words into a computer, users move virtual characters and objects around in their real space, like playing with toys on a table, and talk to them—the AI watches and turns those actions into story pieces.
Alex: So the core problem here is that regular AI story tools force everything through text prompts? And that doesn't match how people naturally think up stories, especially if they're more visual or hands-on?
Sam: Exactly. Text prompts work for some, but they limit storytellers who build ideas by arranging things in space or acting them out—like blocking a scene in theater or setting up action figures. PlayWrite shifts to XR environments, where virtual elements blend with your physical room so you can grab and move characters directly, making the process feel more like play than writing.
Alex: Right, and the paper says current tools poorly support that spatial side of storytelling. How does PlayWrite actually capture those playful actions into something the AI can use for a full story?
Sam: It uses a multi-agent setup—think of it as a team of AI watchers. One tracks the environment, like where characters stand near props; another notes social cues, such as characters facing each other closely; a third handles the overall story flow. Together, they boil down your moves and spoken lines into structured summaries called Intent Frames—basic story beats, like a moment of tension or characters meeting, shown as draggable marbles you can rearrange on a timeline.
Alex: Okay, so these Intent Frames are like editable cards that capture the intent behind your play? Without typing, the AI gets the big picture from gestures and space, then a language model weaves them into a script.
Sam: That's the key idea. A study with 13 writers found users felt the AI was a real partner, sparking ideas through surprises in play—helping beat creative blocks. The paper suggests this embodied approach makes co-creation more intuitive for spatial thinkers.
Alex: Makes sense why text feels rigid... This opens up storytelling in a more natural way.
Alex: So if PlayWrite makes storytelling feel more natural through play, how does the actual interface in XR handle that—without forcing users back to typing or rigid controls?
Sam: The interface layers virtual story elements right onto the user's real room using passthrough XR—like seeing digital characters and props mixed with your actual table or floor, so you can grab them as if they're toys. It has a story playground, a marked area in front of you for placing characters; a dialogue history panel to review past talk; an intent frame view showing those summary cards; and an assembly space for rearranging marbles on a timeline.
Alex: Okay, so everything's right there in your physical space. But how do you interact—does grabbing a character do different things depending on how you do it?
Sam: Yes—direct manipulation keeps it intuitive. Grab the character's body to talk, which starts a smooth animation and listens to your voice, hiding the move handle to avoid mix-ups. Or drag a small handle near the character to reposition it, with the system calculating paths and stopping at safe bounds; props snap to hand zones when close, like handing over an item naturally. These let you stage scenes multimodally, matching whatever feels right in the moment.
Alex: That sounds like it gives real control over positioning and talk. What keeps the AI from just taking over—does it respond only when you prompt?
Sam: The dialogue feels like a live improv chat. You grab and speak through a character anytime; AI ones respond reactively—after your line, targeting by name or who you're facing—or proactively during pauses over ten seconds, suggesting next steps to keep momentum without railroading. Users override AI lines easily by grabbing, balancing your lead with helpful extensions while preserving your intent.
Alex: Huh—so the design invites play without locking you in... It prioritizes curiosity, expression, and turning moves into story beats.
Sam: Precisely. The creators set three goals: spark open-ended improvisation free of failure fear; support speech, grabs, and space as equal inputs; and interpret actions like moving foes near a prop as meaningful plot turns, not just motions—the AI elaborates while staying true to your direction. This fosters agency, where you steer but the system collaborates meaningfully.
Alex: Those design goals sound solid for keeping the user in charge... But how do the agents actually turn all those grabs, moves, and words into coherent Intent Frames without missing the point of what you're improvising?
Sam: The three agents watch the scene constantly, each picking up different clues—like one noting where a character moves a prop close to another, another spotting if characters face off tensely, and the third linking it to the overall story flow. They tag these as simple observations with details like who did it, what they touched, where it happened, and how sure they are. These tags get filtered and combined by a central piece that drops shaky or repeated stuff, highlights what's fresh or plot-relevant, and blends them into a single summary, say, a character approaching angrily.
Alex: So it's like a team of spotters at a game, each calling out key plays, then a referee merges the best ones into official highlights? What decides which observations make the cut—does it just average them?
Sam: Close—each merged summary gets a score blending how much space changed, if it's a new social dynamic, and if it pushes the story forward, with weights shifting based on what's happened lately to avoid repeats. For instance, moving foes near a gun might score high on tension, tagged as a threat buildup. The paper describes this as fusing low-level actions into narrative beats, like zones where hiding matters or trails showing chase direction.
Alex: Huh—that explains why play feels meaningful; the system spots intent without you spelling it out. And users can drag those as marbles to reorder later?
Sam: Yes, once committed, they become visual marbles on a timeline, each holding the summary, characters, tension level from one to ten, and story role like rising action. Rearranging them feeds straight to a language model for a synopsis or script, keeping your improvisational choices intact while smoothing transitions. This setup balances your control with AI help, ensuring the output stays true to the play.
Alex: So the fusion makes editing intuitive... A notable way to capture spatial storytelling.
Alex: You've described how the marbles let users reorder their play into a story... But did real users find that editing process gave them enough control, or did the AI override their vision?
Sam: The paper reports an exploratory study with 13 writers of varying experience—from published authors to hobbyists and game designers—using structured and open tasks in three scenes, like a tutorial interview, a guided Aladdin betrayal, and free Robin Hood improvisation. After play, participants reviewed and rearranged marbles, then exported to synopses or screenplays via a web tool. Ratings on a creativity support index were notably high for expressiveness at around 6.4 out of 7 and enjoyment near 6.2, but more mixed for control at about 3.9—suggesting users felt free to explore but sometimes wanted finer narrative grip.
Alex: So expressiveness won out over tight control... Like, the play sparked ideas, but reshaping the final story wasn't always precise?
Sam: Precisely—that aligns with the design goals. It fostered playful improvisation without failure pressure, like directing actors on a set where surprises from AI responses opened new paths, turning a planned horror into comedy for one user. Grabs and moves steered plots directly, such as positioning foes near props to signal threats.
Alex: Huh—that director feel preserves agency amid collaboration. In the Robin Hood example, how did reordering marbles actually shape the output?
Sam: One user shifted a confrontation marble earlier, grouping tension beats, then exported a JSON log that generated a screenplay preserving vocal tones and authorship details—like user moves plus speech triggering "Mary pleads with Robin Hood." Review showed outputs faithful to play, letting users refine further, which underscores transparent agency: you see summaries, replay moments, and reorder without losing improvisational essence.
Alex: A balanced way to co-create... Notable for blending control with surprise.
Alex: What did the writers actually say about steering the story through grabs and space—did it really give them the agency they wanted?
Sam: Writers noted that moving characters near props signaled plot turns naturally, like placing foes by a gun to hint at threat, without typing prompts. One said it picked up subtext on moral conflict from pointing the gun, giving a sense of direction through scene-based actions rather than scripting every detail. The study suggests this built control by turning physical setup into narrative nudges, though not always perfectly precise.
Alex: So space and grabs inspired ideas on the fly... Like putting characters behind a rock to hide—did the AI read that as intended?
Sam: Yes—users praised how rare it is for systems to process such details into the story, sparking character development through grabs and speech that put them in the role. Dialogue complemented this, letting them test lines live and get character voices back, reducing pressure for perfect wording and aiding spontaneous flow.
Alex: Huh—that turn-taking mimics improv... But with AI, weren't there hiccups in timing or when it jumped ahead?
Sam: Participants mentioned friction, like waiting for AI to finish before interrupting or lag making grabs ignored, sometimes shifting the story unexpectedly. Some adapted these as new material, reframing mismatches into plot, but it highlighted a trade-off: valued improvisation yet desired better cues for precise steering.
Alex: Makes sense—the play invites emergence, not rigid plans... Users saw AI as a bouncing partner, not just a tool.
Sam: Exactly—they felt collaboration through context-aware responses, like a playmate in doll play where you negotiate direction. This aligned with design aims: playful curiosity via surprises, multimodal expression blending moves and talk, and intent-preserving events from actions—achieving fluid co-creation, though refining control remains key.
Alex: That balance of control and surprise seems central... Different writers approached it in their own ways, right?
Sam: Yes—the study showed variety. Some came with a clear plot in mind and used the system to test it out, valuing when the AI stuck to their direction. Others discovered the story through surprises, treating unexpected AI lines as new ideas to build on. Writers with improv or game backgrounds liked it most, as it fit their unscripted style.
Alex: Huh—so it suits explorers more than strict planners... And looking ahead, what design lessons come from that?
Sam: The paper outlines a few key ones. Systems should make room for spontaneous play over tight planning, treat speech and moves as equal ways to express ideas, and let the AI respond to your hints while adding its own touches. Play itself emerges as a useful way to think about these tools—not just fun, but a real guide for building them.
Alex: Makes sense as a fresh angle... Though with only 13 participants, mostly experienced writers, how general is this?
Sam: Fair point—the small group and short sessions limit broader claims, especially for beginners or long-term use. Extreme changes to marble order could also glitch story flow, as the language model doesn't enforce strict timelines. Still, it points to play empowering non-writers like kids or teachers to shape stories through hands-on trials, with potential in education or early game design.
Alex: A solid step toward intuitive co-creation. Thanks for listening to ResearchPod.