Local SEO blog

AI Em Dash Examples and Why Your AI Drafts Keep Reaching for Them

Bright clean 3D render of an abstract document with repeated long dash marks floating above it in cyan, mauve, blush pink, white, and light gray

The em dash is probably the most misunderstood punctuation mark in English right now. AI drafts use it so heavily that people have started treating the mark itself like proof that a machine wrote the page.

That misses the bigger picture. The mark has been around for a long time, writers have used it on purpose for a long time, and AI reaches for it because AI repeats patterns it has seen over and over. That makes a lot more sense when you look at literary history, editing rules, training data, and the way these tools generate text.

I can usually spot a machine-written draft in about three seconds without actually reading it. I look at the shape of the page first. When the same mark shows up in clusters, five to ten times across a short post, the rhythm gives the draft away before the words even register. That pattern is real. So is the mark’s long history in human writing.

The em dash extravaganza

The easiest way to see the problem is to compare rough AI copy with a human edit pass.

A lot of machine-written SEO copy comes back with one opening paragraph that does everything wrong at once. It packs in half a dozen of the marks, every banned phrase in the book, a conclusion that concludes nothing, and emoji confetti on top. It opens in today’s fast-paced digital landscape, calls SEO a journey, reaches for unlock the power of, calls the content game-changing, drops in synergistic best practices, then says at the end of the day and closes by saying absolutely nothing at all. Somewhere in there it also mentions the wrong year in a line about 2024 and throws a rocket emoji over the wall on the way out.

Screenshot of a machine-written SEO paragraph packed with em dashes, banned phrases, a wrong year, and rocket emoji

My edit pass fixes that by making the copy sound normal again. A cleaned up version says what the page is for, who it helps, and what to do next. It uses plain sentences, keeps the point in view, and sounds like somebody trying to communicate instead of somebody trying to sound impressive.

Meme photo of Randi with the caption “My Face When” above and “I see dashes everywhere” below

A mark with a long literary life

The name comes from typography. An em dash is one em wide, traditionally the width of the letter M at a given type size.

,

An en dash is half that width.

The terminology reaches back to metal type, where printers used quadrats measured in ems. Joseph Moxon recorded the terms “m quadrat” and “n quadrat” in his 1683 book on printing. The history of the typographic quad and the dash makes clear that this name came from the printer’s shop. The mark itself is older. The 99% Invisible episode on the em dash traces it back to eleventh-century Italy.

Writers adopted it because it does useful work. It can signal an interruption, a sharp pause, an aside, or a thought that changes direction in mid-sentence.

Some scholars read Emily Dickinson’s dashes as pauses, hitches, and spaces where ideas remain in tension. Critics have debated their meaning and even how editors should represent them in print. That debate itself tells you how much expressive weight readers find in her punctuation. John Hay’s overview of Dickinson’s dashes lays out several of those critical readings.

Poe used the mark to heighten a narrator’s rising alarm in “The Black Cat.” You can see the pressure build when the dashes start clustering. William Goldman uses the mark for restless dialogue that interrupts its own rhythm in Marathon Man. I have always liked that effect.

Two reading app screenshots showing examples of the em dash in Marathon Man by William Goldman and The Black Cat by Edgar Allan Poe

Other writers worked the mark hard. The 99% Invisible episode on the em dash counts one every 129 words in Moby-Dick, one every 224 in Oliver Twist, and one every 90 in Jane Eyre. The styles differ, and the punctuation serves different purposes in each. The point is simple. The mark was already a working literary tool long before people started blaming it on machines.

One of the screenshots above came from a post in r/WritingWithAI by u/Gold_Concentrate9249. That poster wrote:

“Just to avoid any misunderstandings: The page on the left is from Marathon Man by William Goldman. Page on the right, The Black Cat by Edgar Allan Poe. These two American authors were using “em dashes” effectively years before AI, Goldman to great effect. I grew up on a steady diet of writers like this and that’s a primary reason I appreciate the many uses of the em dash. I find it frustrating that use of them today is discouraged simply because their inclusion triggers a lot of people.”

I grew up on that kind of prose too. I knew the mark from books long before it became something people used to call out machine writing.

For a wider cultural history, I recommend the 99% Invisible episode about the em dash. It follows the mark through printing, literature, and the current AI conversation.

Editors use it too

Professional editors use the em dash all the time. Different style guides just give it different house rules.

The Associated Press Stylebook calls for spaces on both sides of the mark. Chicago style uses it without spaces, as described in this Chicago Manual of Style discussion.

There is also a useful old newsroom detail here. AP historically avoided the en dash, including in number ranges, and preferred hyphens. The reason was transmission: the en dash did not travel cleanly through early teletype machines used to send news copy. Old technology still leaves marks on style rules.

You will find em dashes in edited journalism, magazines, features, books, and essays. Editors decide where the mark fits and how it should look. That writing also becomes part of the material language models learn from.

How the model picks its next mark

People sometimes talk about an AI model’s “brain” as everything people put into it. That image is close enough to be useful, with one important correction. The training text gets turned into numerical weights that capture patterns. When the model generates a response, it uses those learned patterns to estimate which token is most likely to come next.

That matters here because the model has seen a lot of sentences where an em dash introduces an aside, adds emphasis, or shifts direction. Those examples help build the pattern. Then, when the model is asked for polished prose, the mark is an easy choice because it fits a lot of contexts.

The model is not sitting there thinking about Poe or Goldman. It is repeating habits from the text patterns it learned, plus whatever instructions and examples are shaping the current response. The exact data behind commercial models is often private, so I cannot point to one author or one dataset and say that is the source. The general mechanism is still pretty easy to understand. The model leans toward patterns it has seen a lot.

What happens when you tell the AI to stop

A direct instruction helps, but it helps in a limited way.

A chat instruction can steer the next response. It does not retrain the model on the spot. The habit comes from patterns the model already learned, so your instruction is competing with a pattern that is already strong.

Saved preferences help too. A standing instruction rides along with future requests instead of disappearing between chats. In my own use, that lowers how often the mark shows up. It still comes back sometimes when the context shifts, the request gets long, or the model falls back into its default polished rhythm.

So I use a simple system.

The brief comes first. I write punctuation and style restrictions into every content brief I hand off or work from. If I care about cadence, I say so before the draft starts.

Then I repeat the prompt. I keep the exact instruction saved as a desktop sticky note and paste it more than once, sometimes in the same chat, because one style instruction can fade as the conversation gets longer. My wording stays blunt because blunt works: stop using em dashes, never under any circumstances use them, and update the memory so it does not happen again.

Then I do the edit pass. I paste the problem paragraphs back into the tool and ask for a rewrite with the restriction stated again.

On the client side, I handle it in the edit pass. I swap the marks out, smooth the rhythm, and move on. Grammar checkers underline the mark, so it is easy to spot. The point is clean copy.

The instruction helps. The edit pass finishes the job.

Screenshot of a note containing the instruction to stop using em dashes

Stop using EM dashes , , never ever under any circumstances use them. Update the memory to make sure it never happens again.

The instruction block I keep on every tool

I keep one block of style rules and paste it into every tool I use. Claude, Magica, Perplexity, Gemini, and whatever ships next week. Using one version everywhere keeps the rules consistent.

The exclusions live in one block. They are marked non negotiable, and they apply to every chat on the account.

  • The em dash, banned outright.
  • The transition that opens with the paired customer construction.
  • The opener that announces a point before making it.
  • The adverb for something happening without anyone noticing.
  • The word that introduces a choice between two options.
  • The rhetorical reframe, the move where a writer negates one idea to set up another.

Screenshot of the instructions for Claude settings panel holding a block of banned writing rules for an account

One saved block is easier to manage than a pile of reminders spread across different tools. Every tool gets the same rules. I stop chasing five different saved settings, three half updated notes, and one mystery preference I forgot I turned on six months ago.

That setup helps. It lowers how often the habits show up. It still leaves room for them to come back, because a saved instruction rides along with the request instead of changing what the model learned. My edit pass still closes the gap.

The other patterns give a draft away

The mark gets most of the attention, but it usually shows up with a few other habits.

I see the same kind of overuse in stacked bullet points, repeated sentence structures, and predictable transitions that keep doing the same job. A draft starts sounding too even. Every paragraph lands the same way. Every explanation takes the same turn.

A lot more people are using these tools now, which makes those habits easier to spot in the wild. The Elon University Imagining the Digital Future Center reported that 52% of US adults use large language models in a national survey run January 21 to 23, 2025 through the SSRS Opinion Panel with 939 US adults and a margin of error near three percentage points, and that figure measures tool use rather than how many people publish AI tells (source).

That is why I keep coming back to process. I want a working system for spotting repeated patterns, cleaning them up quickly, and keeping the useful parts of the tool. The draft can still save me time. My edit pass makes it sound like somebody actually meant to write it.

Why older books can carry extra weight

Copyright is one practical reason older writing can stand out in openly licensed or rights-conscious datasets. Public-domain books can be digitized and reused without the permissions required for many modern titles. Licensing contemporary books can involve expense and legal risk. A UCL paper on the value of books in generative AI training data examines the role books play in this discussion.

Project Gutenberg provides free digital books, with a focus on older works whose US copyright has expired. Common Crawl makes large web archives available. That archive contains varied material, and a dataset needs to consider the rights and licenses attached to the pages it uses.

Newer openly licensed collections are also being built. The Common Pile research paper describes a large dataset assembled from public-domain and openly licensed material, including books, academic work, government material, and other text. An announcement on Mozilla’s builders site describes Common Corpus, a Pleias project, as a collection that includes books, newspapers, scientific articles, government and legal documents, code, and more.

For US training datasets, Jane Austen and Poe are practical public-domain sources. Last year’s bestseller generally still needs rights clearance. That gives the older canon a simpler path into openly usable collections, while newer books need more careful permissions work.

That is one reason Dickinson’s and Poe’s punctuation sits inside a corpus that can skew old. Plenty of other influences shape a model’s habits, and no single dataset explains any particular model on its own.

Contemporary journalism adds another layer. Training corpora can include large web collections and licensed or openly licensed news material. Newsrooms, magazines, and feature writing use the em dash as part of edited prose. This gives models many modern examples of the mark alongside the older literary ones. The Common Pile paper, for example, documents a news component as well as book collections.

A layered library of older books and paper pages feeding into a clean model inspired device

How I keep the mark in proportion

The em dash has real uses. It can make a strong break, set off an aside, show a sudden interruption, or help a punch line land harder than a comma would.

Comic illustration of Randi in armor facing robots under a sky full of flying dash marks

I have a standing rule for my own work. The mark stays out of my drafts, out of my site, and out of the client content I write. That is my own house-style choice, and everyone else gets to set their own.

Overuse is the part worth noticing. I count the marks in a draft when the rhythm starts to feel repetitive. As a working habit, a draft with one in every paragraph reads machine-like to me. One or two across a thousand words often feels deliberate. I use that as my editing guide, and your own style guide should tell you what it expects.

I also read the sentences around each mark. Does it create a meaningful pause? Does the aside help? Would a period make the point clearer? That is where human review matters. I review AI-assisted copy before publishing, every time.

That workflow connects right back to the literary history. The mark is worth keeping because it is worth using deliberately.

That same editorial care matters in my broader work on building topical authority, a content SEO strategy for AI search, and what ChatGPT reads from a page. Clear writing starts with useful answers and thoughtful editing. My SEO content strategy services follow that same principle.

The goal is proportion. The em dash belongs to writers. AI picked up a useful tool from the language it learned. Use it well, edit the draft, and let the mark do its job when the sentence calls for it.

Make Randi One of your Preferred Sources

Sources