K2

K² · Artificial intelligence

AI generates whole conversations — messy room sounds and all — from a written scene

MF
DS

Michael Finkelson, Daniel Segal, Eitan Richardson et al.

10 authors · cs.SD, cs.AI, cs.CV

arXiv preprintArtificial intelligenceJun 2026 · ~70s read

Like explaining it at the dinner table.

Most AI that fakes a conversation between people works like a scriptwriting tool. You hand it a tidy script: this line is Speaker A, that line is Speaker B, tagged turn by turn. The result sounds clean — and clinical. No background hum, no overlapping voices, none of the texture of a real room.

ScenA throws out the script format. You give it a few sample clips of each person's voice, plus a plain-English description of the scene you want — "two friends arguing in a busy cafe," say — and it generates the whole soundscape: the voices, the chatter over each other, the laughs and sighs, the room echo. It can do this because it's built on a model already trained on huge amounts of real-world audio, not studio recordings, so messiness comes baked in.

The clever part is a trap they had to dodge. The model builds audio by cleaning up noise step by step. Early on, it figured out which voice sample matched which speaker just by comparing sounds — ignoring the written description entirely. So they trained it on extra-noisy versions where that acoustic shortcut doesn't work, forcing it to actually read the prompt to decide who speaks.

On one benchmark, it pinned the right voice to the right speaker better than existing systems. Still, that's a single test set — how it handles many distinct speakers or long exchanges isn't shown.

Why you should care: Generating overlapping dialogue with ambient sound from a written description — not a tagged script — means audiobooks, game scenes, and dubbing could be authored in plain sentences instead of engineered line by line.

arXiv preprint — these findings haven’t been peer-reviewed yet. Treat them as early results, not settled science.