AI-generated text has become difficult to identify by eye alone. Detection tools exist to help, and they are genuinely useful — but only if you understand what they actually measure and where they fail. This guide covers how detection works, which free tools are worth using, and how to read a result without over-trusting it.
How AI Detectors Work
Detectors do not recognise text as "written by ChatGPT" in any direct sense. They measure statistical properties of the writing and compare them against patterns typical of language-model output.
Perplexity
Perplexity measures how surprising each word is given the words before it. Language models are trained to predict likely next words, so their output tends to be highly predictable — low perplexity. Human writing wanders more: unusual word choices, tangents, idiosyncratic phrasing. Consistently low perplexity across a passage is the single strongest signal most detectors rely on.
Burstiness
Burstiness describes variation in sentence length and complexity. Humans naturally mix long sentences with short ones. A three-word sentence lands after a forty-word one. Models trained to produce fluent prose tend toward uniformity, so a passage where nearly every sentence runs 15–25 words with similar structure reads as machine-generated to a detector.
Token probability distribution
More sophisticated detectors examine the probability distribution across the whole passage rather than word by word. Human writing contains occasional low-probability choices that a model optimising for fluency would rarely make. The absence of those outliers is itself evidence.
Best Free AI Detection Tools
Anonymiz AI Content Detector
Our AI Content Detector analyses text for the patterns above and returns a confidence score with a breakdown of the signals behind it. Free, no signup, no stored text. Best for quick checks on shorter passages where you want to see the reasoning rather than a bare verdict.
GPTZero
One of the earliest dedicated detectors, built around perplexity and burstiness scoring. It offers sentence-level highlighting, which is useful for spotting mixed documents where only part of the text was generated. The free tier limits word count per check.
Originality.ai
A paid tool aimed at agencies and publishers, generally reported as accurate on unedited model output. Worth mentioning because it sets a realistic ceiling: even commercial tools built for this purpose do not claim certainty, which tells you something about how hard the problem is.
Specialised checkers
For specific contexts, narrower tools do better. Our AI Essay Detector is tuned for academic writing, and the AI Email Detector is calibrated for shorter business correspondence, where general-purpose detectors struggle most.
Limitations of AI Detectors
This is the part most guides skip, and it matters more than the tool comparison.
False positives are real and unevenly distributed
Clear, well-structured, grammatically consistent human writing looks statistically similar to model output. Technical documentation, formal reports and academic prose are flagged more often than casual writing. Non-native English speakers are flagged disproportionately, because writing learned through formal instruction tends toward the regular sentence structures detectors treat as suspicious. Any process that penalises people based on a detector score will penalise these groups hardest.
Short text is close to unmeasurable
Statistical signals need volume. Under roughly 150–200 words there is not enough data for a meaningful reading, and confidence scores on a paragraph or two should be treated as noise. Most tools will still return a number, which is misleading.
Light editing defeats detection
Rewriting a handful of sentences, varying lengths, and swapping predictable word choices will move most text below detection thresholds. This means a low score does not establish human authorship — it establishes that the text does not currently exhibit the measured patterns.
Detectors cannot identify which model was used
A tool may report "likely AI-generated," but no detector can reliably distinguish GPT output from Claude or Gemini output. Claims to the contrary should be treated sceptically.
How to Read a Detection Score
Treat the output as evidence, not a verdict. A useful way to interpret results:
High score on a long, unedited passage — reasonable evidence of generated text, worth following up on. High score on short or heavily formal writing — weak evidence, plausibly a false positive. Low score — establishes very little, since editing suppresses the signal. Mid-range score — essentially uninformative; most tools are least reliable in the middle of their range.
Where the stakes are high — academic misconduct, employment decisions, publication — a detector score should never be the deciding factor. Ask for drafts, version history, or a conversation about the content instead. Those establish authorship in a way statistics cannot.
When to Use AI Detection
Editorial and content workflows
Checking freelance submissions before publication is a sensible screening step, particularly if your agreement specifies original writing. Use it to open a conversation, not to reject work outright.
Reviewing your own writing
If you used AI assistance and want the result to read naturally, a detector shows which passages are most uniform. Our Humanize AI Text tool helps rework those sections. This is the least controversial use case — you already know the provenance, and you are checking quality rather than making an accusation.
Understanding the technology
Testing detectors against text you know the origin of is the fastest way to calibrate your own expectations. Try our AI vs Human Quiz to see how well you do unaided — most people score close to chance, which explains why tools exist at all.
Where not to use it
Automated enforcement is the clear misuse: any system that fails a student, rejects an application or removes content based solely on a detector score will produce unjust outcomes at a predictable rate. The tools are not accurate enough to carry that weight, and their errors are not randomly distributed.
Frequently Asked Questions
How accurate are free AI detectors?
On long, unedited model output, well-built detectors perform reasonably. Accuracy drops sharply on short text, edited text, and formal human writing. No tool publishes a false-positive rate that would justify using it as sole evidence in a consequential decision.
Can AI detectors be wrong about human writing?
Yes, and predictably so. Formal, well-structured prose and writing by non-native English speakers are flagged more often than average. This is a known limitation of measuring statistical regularity rather than authorship.
Does editing AI text make it undetectable?
Largely, yes. Varying sentence lengths, replacing predictable phrasing and restructuring paragraphs will move most text below detection thresholds. This is why a low score cannot be read as proof of human authorship.
Can a detector prove someone used ChatGPT?
No. Detectors report a statistical likelihood based on observable patterns. They cannot identify which model produced a text, and they cannot establish authorship. Drafts and version history are far stronger evidence.
How much text do I need for a reliable check?
Aim for at least 300 words. Below roughly 150–200 words there is insufficient signal, and any score returned should be treated as unreliable regardless of how confident the tool appears.
Related Reading
Check any text instantly with our free AI Content Detector — no signup required. For website-level detection rather than text, see the AI Website Detector, which identifies sites built with tools like Lovable, Bolt and Framer AI.


