Skip to content

Tools

What AI Text Detectors Can and Cannot Tell You

A detector score looks like a verdict and is really a guess about style. Knowing how the guess is made, and where published research shows it goes wrong, is the difference between using a detector sensibly and letting it decide for you.

A browser window showing a passage with a score beside it and several sentences marked for a closer readexample.org/article#:~:text=the%20one%20linethe browser scrolls here and marks the text
Four highlighter pens lined up beside a paper bookmark, two paper clips and a card with one yellow stroke

Teachers, editors and hiring managers now meet AI text detectors every week, often built into tools they already use. The tools report a percentage or a label, and it is tempting to read that number as a finding. This page explains what the number is, summarises what independent research has found about how reliable it is, and sets out a fair way to use a result. It does not rate any product; the principles apply to all of them.

How AI text detectors work

Most detectors are classifiers: statistical models trained on large collections of text labelled as human-written or machine-generated. Given a new passage, they estimate how closely it resembles each collection. Many rely on the observation that language models tend to choose likely words, so machine text is often more predictable, word by word, than human text, and more uniform from sentence to sentence. Some newer research systems instead look for a watermark: a hidden statistical pattern that a model deliberately builds into its word choices, which a matching detector can later test for. Watermarks only exist if the model that produced the text was built to add one.

Three families of method; none of them can see who typed the words.
ApproachWhat it measuresMain weakness
Trained classifierResemblance to labelled human and machine textHuman writing that is plain, formulaic or non-native can resemble machine text
Predictability scoringHow expected each word is to a language modelEditing and paraphrasing change the score easily
Watermark testA deliberate pattern the generating model addedOnly works for text from a model that watermarks, and heavy rewriting weakens it

What independent research has found

The published evidence is consistent on three points. First, detectors misclassify human writing, and they do so unevenly: a Stanford study found that essays by non-native English writers were flagged as machine-written far more often than essays by native speakers, largely because simpler vocabulary and structure look predictable. Second, paraphrasing defeats them: researchers showed that running generated text through a rewording step sharply lowers detection rates, and argued that reliable detection becomes harder as models improve. Third, people are not much better: in one widely reported test, scientists reading abstracts could not reliably tell generated ones from real ones.

None of this means a detector result is meaningless. It means a result is one weak signal, with known biases, that needs other evidence before anyone acts on it.

Using a detector result fairly

  1. Treat the score as a reason to look, not a finding

    A high score justifies reading the work more closely. It does not justify a conclusion, a grade penalty or a rejection on its own.

  2. Read the passage yourself

    Mark the features that drew the score, using the approach in close reading for machine-writing habits, and look just as hard for specific detail only the writer could know.

  3. Ask for the process

    Drafts, notes, version history and sources settle most questions quickly. The page on reviewing authorship through drafts and sources describes the conversation.

  4. Record what you relied on

    If a decision follows, write down the evidence other than the score. If the only evidence is the score, there is no decision to make.

If a detector flags your own work

Stay calm and bring your process: dated drafts, the sources you read with your highlights, your notes, and your document's version history. Ask which passages were flagged and explain how you wrote them. Point out, politely, that published research documents false positives, especially for non-native writers. A clear record of your process is the strongest answer to any score.

Where this fits on the site

This site builds highlighting tools, not detectors, and the reading skills here are the part of the problem a detector cannot do for you. The tools and extensions section covers the highlighters and annotation apps that make close reading and process records practical, including annotation layers for shared reading and highlight links you can send as evidence.

Common questions

How accurate are AI text detectors?

Independent studies show they misclassify some human writing, especially by non-native English writers, and that paraphrased machine text often passes. A score is a probability about style, so treat it as a reason to look closer, not as proof.

How do AI detectors work?

Most are trained classifiers that compare a passage with examples of human and machine text, often using how predictable the wording is. Watermark tests look for a pattern deliberately added by the generating model.

Can a detector prove someone used AI?

No. It cannot see who wrote the words. Drafts, sources, version history and a conversation with the writer are the evidence that matters.

Why do detectors flag non-native English writers?

Simpler vocabulary and regular sentence structure make text more predictable, which is one of the features detectors associate with machine writing.

What should I do if my writing is flagged?

Share your drafts, notes, sources and version history, explain how you wrote the flagged passages, and ask what other evidence exists.