article · 2026-06-24
Can AI Write a Game? Dialogue, Quests, and the Level-Design Wall
The narrative kitchen scores 2.78 in our buildability model — but that average hides a hard split between the words AI writes for free and the design it still can't.
Can AI write a game's narrative? Yes — but 'narrative' is three jobs, not one
If you ask whether AI can write the story in your game, the honest answer is: it depends entirely on which part of the story you mean. People hear the word narrative and picture a single craft — someone writing the script. But a game's narrative is really a stack of very different jobs, and AI is brilliant at the bottom of that stack and helpless at the top. Lumping them together is how you end up either over-promising or dismissing the tools entirely.
In our AI Game-Buildability study, we broke the medium into roughly 45 reusable ingredients spread across five kitchens — Art, Audio, Narrative, Code, and Polish — and scored each ingredient on how much of the work a generalist can hand to AI tooling today. The Narrative kitchen averages 2.78 in our buildability model, which puts it squarely in the middle of the pack. But the average is the least interesting thing about it. The interesting thing is the spread underneath: this is the most internally split kitchen of the five.
Narrative runs from a flat 1 — flavour text, the single most automatable thing in the kitchen — all the way up to a flat 4 at level, puzzle, and encounter design. That is a three-point range inside a single kitchen. Knowing where on that range your project actually lives is worth more than any headline average, because it tells you exactly which parts you can generate this afternoon and which parts will still take a human a month.
The easy shelf: flavour text, dialogue, worldbuilding, and localization
Start at the bottom of the kitchen, where AI does almost all of the lifting. Flavour text — item descriptions, signposts, ambient barks, the little lines nobody writes a design doc for — scores a 1 in our assessment. It is text generation with low stakes and forgiving constraints, which is precisely the shape of work large language models are best at. You curate and trim; the machine supplies the volume.
One rung up, at 2, sit three load-bearing crafts: branching dialogue, worldbuilding and lore, and localization. Branching dialogue is the conversation tree behind every quest-giver and shopkeeper — the lines, the player options, the conditional responses. Worldbuilding is the connective tissue of names, factions, histories, and place. Localization is taking all of that and rendering it in another language. All three are content-generation problems where the output can be read and judged on its own, which is exactly why our model rates them automatable: an AI can draft, vary, and translate them at a scale no single writer matches.
The catch — and it is the whole catch — is that automatable does not mean finished. AI can write a thousand serviceable tavern lines; it cannot tell you which fifty have a voice, or whether your bard actually sounds like a bard rather than a generic fantasy NPC. The performance layer matters too. Written lines are not the same as delivered ones, which is why drop-in voiced libraries like our Bard Dialogue Pack — 112 minutes of professionally delivered bardic, quest-giving, tavern-tale dialogue — exist alongside generated text: the words may be cheap to draft, but a consistent, characterful spoken voice across an entire cast is a different problem from generating the script.
The middle shelf: quest design and the economy creep up to 3
Move up the kitchen and the work changes character. Quest design and economy-and-progression both score a 3 in our buildability model — the boundary between content AI generates cleanly and design AI can only assist with. The line between a quest's dialogue (a 2) and a quest's design (a 3) is subtle but decisive, and it is where a lot of AI game-writing demos quietly fall apart.
Writing the words a quest-giver says is content. Deciding what the quest does — where it sends the player, what it gates, how it threads into the wider critical path, whether it teaches something the next quest assumes you now know — is design. AI can draft a perfectly plausible 'fetch the amulet from the cave' quest in seconds. Whether that quest lands at the right point in your difficulty curve, respects your pacing, and doesn't hand the player a reward that breaks the next three hours is a judgement about the whole system, not the single artefact. The same is true of an economy: AI will happily generate prices, drop tables, and crafting costs, but whether they hold together so the game is neither trivial nor grindy is a balance question, and balance is where automation thins out.
Practically, this is the rung where you stop treating AI as an author and start treating it as a fast, tireless first-draft generator that you supervise. Let it propose ten quest structures and twenty economy tunings; keep the human in the seat that decides which one actually fits the curve. The leverage is real — but it is leverage on a craft you still have to direct, not a craft you can outsource.
The wall at 4: level, puzzle, and encounter design
Then you hit the wall. Level design, puzzle design, and encounter balance all score a flat 4 in our assessment — the top of the Narrative kitchen and one of the four hard ceilings that recur across our entire study. This is where the story everyone repeats about AI building games quietly stops being true.
The trap is that these crafts look automatable. AI can generate a plausible dungeon layout, a plausible puzzle, a plausible mix of enemies, and on a screenshot every one of them passes. The problem is that a level, a puzzle, and an encounter are not artefacts you judge by looking at them — they are experiences you judge by moving through them. A layout that reads fine on paper has to teach the player its own rules, escalate at the right rate, reward exploration without trivialising it, and stay fair while still surprising. A puzzle has to be solvable for the reason you intended and not by accident. An encounter has to be tense, winnable, and tuned to the exact moment in the game it appears. Generating something that looks like these is easy. Generating one that is tuned to teach and escalate is the hard part, and it is the part AI barely touches.
This is the same fault line that runs through the whole medium: generation of things you can judge on their own is largely solved; synthesis under constraints you can only judge in motion is not. Plausible is cheap. Tuned-to-teach is expensive. Everything above a 3 in this kitchen is on the expensive side of that line, and no current tool moves it cleanly.
Why a 'writing' genre can still be hard to build
This split inside the Narrative kitchen is what decides whether a story-led genre is actually buildable. The genres at the very top of our buildability ranking are the ones whose narrative work lives entirely on the easy shelf. Interactive Fiction scores 1.83 and Visual Novel 1.80 — pure writing, branching dialogue, and worldbuilding, with no level-design or encounter-balance burden dragging them up. Those are the genres where AI's strongest domain and the game's core are the same thing.
Contrast that with a genre like Puzzle, which scores 3.17 in our model and sits in the systems-craft band — despite being, on the surface, a 'content' genre. The reason is that a puzzle game's soul is puzzle design, a 4, and no amount of cheap AI-generated flavour text rescues it. Roguelikes (3.25) and Metroidvanias (3.12) carry the same problem from the other direction: their identity rests on encounter and level design, the hard ceilings, even though their writing is easy. The narrative kitchen's average of 2.78 tells you almost nothing about these games; the position of their core ingredient tells you everything.
So the rule for reading any story-led project is: find the highest-scoring ingredient in its core and that, not the average, is the real difficulty. A game whose narrative work tops out at a 2 is buildable by a generalist today. A game whose narrative work reaches a 4 is a design project wearing a writing costume, and the writing being easy will not save it.
The opportunity: let AI write the words, spend your budget on the design
Pulling the threads together gives a clean strategy. The bottom of the Narrative kitchen — flavour text, branching dialogue, worldbuilding, localization — is now effectively a commodity. You can fill a world with serviceable text faster than ever, which means filling a world with text is no longer where the value is. Everyone can do it. The defensible work has migrated upward, into the design rungs AI cannot finish: the quest that lands at the right point in the curve, the level that teaches, the encounter that is tuned rather than merely plausible.
That points to a specific way to spend a small team's budget. Let AI draft the dialogue, the lore, and the item descriptions; keep your scarce human hours for the parts our study marks as 4s. If you are building a story-heavy RPG, the corollary is that the spoken-voice layer is worth buying rather than rebuilding, because consistent delivery across a whole cast is its own craft. That is the niche our dialogue libraries fill — the Bard Dialogue Pack for quest-givers and minstrels, the Deity Dialogue Pack's 92 minutes of divine proclamations for oracles and gods, the free Assassin Dialogue Lore Pack for rogues, and the Fantasy NPC Voices megabundle's 33 hours and 13,668 voiced lines to cover an entire cast at once.
The most defensible position, in narrative as everywhere else in the study, is the overlap: a project whose words AI generates for nothing and whose design AI cannot touch. Build that, hand the cheap half to the machine, and spend every human hour on the level, puzzle, and quest design that no tool has learned to do. That is not a limitation to lament. On a board where production is collapsing in price, it is exactly where the remaining advantage lives.
How we scored this
Two different kinds of judgement sit in this piece, and it is worth keeping them apart. The AI-build difficulty numbers — the 1 on flavour text, the 4 on level design, the 2.78 kitchen average — are our editorial assessment of what AI tooling can do today, drawn from our AI Game-Buildability study and its buildability model. They are our judgement, not a measurement, and they score the automatability of the parts, not whether the resulting game is any good. Buildable is not the same as fun, and nothing in this kitchen changes that.
The genre figures — the buildability bands and the share of titles released recently — are relative popularity and momentum counts drawn from the MythicLemon Games Catalogue, our own compiled dataset of roughly 82,000 released games. They are a proxy for how crowded and how current a genre is, not a measure of sales or revenue, and we never present them as either. As always, the load-bearing caveat: these scores tell you where the labour collapses, not where the craft does.
The Narrative kitchen, ingredient by ingredient
| Narrative ingredient | AI-build difficulty | What it really is |
|---|---|---|
| Flavour text | 1 | Item lines, signposts, ambient barks — high volume, low stakes |
| Branching dialogue | 2 | The conversation tree behind quest-givers and shopkeepers |
| Worldbuilding / lore | 2 | Names, factions, histories — the connective tissue |
| Localization | 2 | Rendering the writing into other languages |
| Quest design | 3 | What a quest does, gates, and threads into the path |
| Economy / progression | 3 | Prices, drops, costs tuned so the game holds together |
| Level design | 4 | Layouts that teach, escalate, and stay fair in motion |
| Puzzle design | 4 | Solvable for the intended reason, not by accident |
| Encounter balance | 4 | Tense, winnable, tuned to the exact moment it appears |
AI-build difficulty is our editorial assessment of what AI tooling can do today (1 = AI does it near end-to-end; 4 = expert, design-bound craft). It measures automatability of the parts, not whether the game is good.
Same words, different difficulty: story genres by their core
| Genre | Buildability | Band | Why |
|---|---|---|---|
| Visual Novel | 1.80 | AI-buildable | Core is pure writing and dialogue — the easy shelf |
| Interactive Fiction | 1.83 | AI-buildable | Pure text, AI's single strongest domain |
| JRPG | 2.77 | AI-assisted | A content/art genre in a hard-RPG costume |
| Metroidvania | 3.12 | Systems-craft | Identity rests on level design, a 4 |
| Puzzle | 3.17 | Systems-craft | Soul is puzzle design — flavour text can't rescue it |
| Roguelike | 3.25 | Systems-craft | Encounter and level design carry the genre |
Genre buildability is our editorial assessment (lower = more automatable). A genre's difficulty is set by its hardest core ingredient, not by its writing.
FAQ
Can AI write the dialogue for my RPG?
Mostly yes. In our buildability model, branching dialogue, worldbuilding, and flavour text are the easy shelf of the Narrative kitchen — flavour text scores a 1 and dialogue a 2. AI can draft, vary, and translate them at scale. The human job shifts to curation (which lines have a voice) and to the spoken-delivery layer, which is a separate craft from writing the words.
What part of game narrative can AI not do?
Level design, puzzle design, and encounter balance — all 4s in our assessment, and three of the four hard ceilings in the whole study. AI can generate layouts and puzzles that look plausible on a screenshot, but it can't reliably tune them to teach, escalate, and stay fair in motion, because those are judged by playing them, not by reading them.
Why does quest design score harder than quest dialogue?
Because they are different jobs. Writing the words a quest-giver says is content, which AI handles (a 2). Deciding what the quest does — where it sends the player, what it gates, where it lands on the difficulty curve — is design, a judgement about the whole system rather than a single artefact, which is why our model rates it a 3.
If a game is mostly writing, is it easy to build with AI?
Often yes — Interactive Fiction (1.83) and Visual Novel (1.80) top our buildability ranking precisely because their core is pure writing. But a 'writing' genre can still be hard if its real core is design: Puzzle scores 3.17 and Roguelike 3.25 because their identity rests on puzzle and level design, the AI-hard ceilings, no matter how easy their text is.
Is the buildability score a measurement of how good these games are?
No. The difficulty scores are our editorial assessment of how much of each ingredient AI can produce today — automatability of the parts, not quality of the whole. You can generate every word of a game and it can still be unfun without the human design and balance pass that no tool automates. Buildable is not good.
Bard Dialogue Pack
Quest-givers, tavern tales and travelling-minstrel banter — 112 minutes of professionally delivered bard dialogue for medieval and fantasy worlds. Plug the cues straight into your dialogue system.