Put a ChatGPT paragraph next to a human one on the same topic and the differences are rarely about quality. The AI version is often cleaner. What separates them is texture: what the writer knows, what they are willing to commit to, and the small irregularities that come from a person thinking on the page.
This is a side-by-side look at how the two actually read, what research says about our ability to tell them apart, and why the gap is narrowing.
A Direct Comparison
Same prompt, same topic — advice on remote team management.
Typical AI version: "Effective remote team management requires a multifaceted approach. Clear communication channels are essential, as is establishing regular check-ins. Additionally, fostering a culture of trust and autonomy allows team members to thrive. It's important to note that different teams may require different strategies."
Typical human version: "We tried daily standups for four months and everyone hated them. Switched to a written update in Slack by 10am and async video for anything complex. Standups only survive if your team is in three timezones or fewer — past that they become a tax on whoever drew the worst hours."
The AI passage is well-formed and says almost nothing actionable. The human one contains a duration, a specific failure, a replacement, and a rule with a stated boundary condition. That contrast — category versus incident — is the most consistent difference between the two.
The Five Patterns That Give AI Away
1. Consistent formality
Human writers drift. A formal paragraph is followed by an aside, a contraction slips in, register shifts as attention moves. Models hold a chosen register with unusual discipline from first sentence to last. That evenness is itself the tell — not the formality, but the absence of variation in it.
2. Perfect structure
Generated text arrives pre-organised: an intro that previews, body sections of comparable length, a conclusion that restates. Human writing is lumpier. One point gets four paragraphs because the writer found it interesting; another gets a sentence. Sections are uneven because attention is uneven.
3. Generic examples
Models produce "a marketing team might use this to improve campaign performance." Humans produce "our client's Black Friday email went out with a broken discount code and we spent the next six hours issuing manual refunds." A model has no incidents to draw on, so it generates categories. When it does produce specifics, they tend to be round, plausible and unverifiable.
4. No contradictions or self-corrections
People argue with themselves in writing. "I used to think X — actually, that's too strong. What I mean is..." Models rarely reverse a position mid-piece, because coherence is what they optimise for. The absence of any visible thinking-in-progress across a long piece is notable.
5. Emotional flatness
Generated text can describe frustration accurately without conveying it. Human writing about a genuinely annoying problem tends to carry some heat — a sharp aside, an impatient sentence fragment, an opinion that goes slightly further than strictly necessary. Models are trained toward measured neutrality and it shows.
What the Research Shows
The consistent finding across studies of human detection is that people perform close to chance — and importantly, that confidence does not track accuracy. People who feel certain are not much more likely to be right.
Two patterns show up repeatedly. First, we over-flag: careful, well-edited human writing gets called AI, particularly formal or academic prose and writing by non-native English speakers. Second, we under-flag conversational AI text: when a model is prompted to write casually, our intuitions largely stop working, because most people's mental model of "AI writing" is the formal default register.
Training helps somewhat — knowing the specific patterns above measurably beats going on instinct — but the honest summary is that unaided human judgement is not reliable enough to act on alone.
Why the Gap Is Narrowing
Most of the tells above are properties of *default* model output rather than hard limits.
Prompting changes them substantially. A model asked to write in a specific voice, vary sentence length, take a position and include concrete detail produces text that defeats most of the patterns on this list. And when a person edits the output — adding one real anecdote, cutting the hedges — the remaining signals largely disappear.
What has not changed is the underlying asymmetry: a model has no experiences to report. It can be given them in a prompt, and it can invent plausible ones, but it cannot draw on a memory of the Black Friday discount code. Specific, verifiable, checkable detail remains the hardest thing to fake — which is why it stays the strongest signal even as the surface tells fade.
Test Yourself
Reading about the patterns is different from applying them under uncertainty. Our AI vs Human Quiz presents passages with no labels and asks you to judge. Most people score meaningfully worse than they expect, which is the useful part — it recalibrates how much weight your instinct deserves.
For a statistical second opinion on a specific piece of text, our AI Content Detector measures the same properties numerically. For a deeper walkthrough of the individual signals, see our guide to the signs that text was written by AI, and for a comparison of the available tools, the best free AI content detectors.
Why This Matters
The practical stakes are mostly about disclosure rather than detection. Readers generally care whether the person publishing stands behind the content and whether the claims are accurate — not which software was involved in drafting.
Where it does matter, act carefully. Detector scores and personal intuition are both unreliable enough that neither should decide an academic misconduct case or a hiring outcome. Drafts, version history and a conversation about the material establish authorship far more reliably than statistics or a gut feeling.
Frequently Asked Questions
Can most people tell AI writing from human writing?
Not reliably. Studies consistently place unaided performance close to chance, and self-reported confidence correlates poorly with being correct. Knowing specific patterns improves accuracy, but not to a level worth acting on alone.
Is AI writing worse than human writing?
Usually not on grammar, structure or clarity — often better. It is weaker on specificity, genuine opinion, and anything requiring lived experience. Fluency and usefulness are different things.
Will AI writing become undetectable?
The surface tells are already largely removable with good prompting and light editing. The durable difference is access to real, specific, verifiable experience, which is not a stylistic feature and cannot be prompted into existence.
Does it matter if content was written by AI?
It depends on what was promised. If original human writing was commissioned or implied, it matters as a question of honesty. If the content is accurate, useful and someone stands behind it, most readers care more about that than about the drafting process.
Are AI detectors better than human judgement?
Somewhat, on long unedited text. Both degrade sharply on short or edited passages, and both produce false positives concentrated on formal writers and non-native English speakers. Neither is dependable enough to carry a consequential decision on its own.


