SiteTell
// comparison

AI detectors and AI writing checkers answer different questions

An AI detector estimates whether a machine wrote your text. A writing checker tells you whether it reads as generic, and names the rule. Here is where each one is the right tool, and the data showing why the distinction matters.

data measured September 14, 2026ruleset 2026-09-12
About these figures
They were produced under ruleset 2026-09-12, which changed on September 16, 2026. A scan you run today uses the newer one, so your score may differ from what these figures imply. See the current dataset.

Two tools that look almost identical from the outside are answering completely different questions, and picking the wrong one wastes a week.

An AI detector estimates whether a machine produced a piece of text. It takes what you paste in and returns a likelihood, usually a percentage. A writing checker makes no claim about authorship at all. It looks at the text and tells you which specific patterns in it read as generic, and why.

If you are a teacher deciding whether a student wrote their own essay, you want the first kind. If you are a founder wondering why your pricing page reads like every other pricing page, the first kind cannot help you, because "68% likely AI" is not an edit you can make.

The short version

What five writing tools output, what they accept as input, and whether they cover a whole site
GPTZeroOriginality.aiPangramHemingwaySiteTell
What you get backA percentage likely to be AI, with the phrases highlightedAn AI score and a plagiarism result, with sentence highlightsA classification: human, AI-generated, or AI-assistedA readability grade, plus style highlightsThe rule that fired, quoted on the exact sentence
What it takes as inputPasted text or an uploaded filePasted text, an upload, or one URLPasted text or uploads, up to 100 filesPasted or typed textA domain
Covers a whole siteNoOne page at a timeNoNoYes, up to 400 pages per scan
Makes a claim about who wrote itYesYesYesNoNo
What a wrong answer costs youAn accusation the writer has to disproveA rejected draft, or a flagged freelancerAn accusation, at a stated rate of 1 in 10,000A suggestion you ignoreA flag you disagree with, on a sentence you can see
Free tierUnder 10,000 characters without an accountThree scans a day, 2,000 wordsTwenty checks a day with an accountYes, with a paid tier above itThe whole scan, every page, no signup
Who it is built forEducators and institutionsPublishers, SEO teams, agenciesEducatorsAnyone editing proseTeams auditing a live marketing site
Claims checked September 21, 2026

Why does the distinction matter if both flag the same text?

Because they disagree far more often than people expect, and their disagreement carries the useful information.

Detectors ask about origin. Writing checkers ask about effect. A paragraph can be entirely human and still read as interchangeable, and a paragraph can be machine-drafted, heavily edited, and land sharp and particular. Those two facts are what keep the categories separate.

We have a number for the first half of that. Across 97 domains and 10,126 pages, 51% of sites carried at least one word from our own blacklist. That is a majority of sites, and much of the writing underneath those flags is unambiguously human: copy written in 2019, brochure language inherited from a rebrand, a founder's own draft that reached for the nearest available phrase. Marketing prose converged on this register long before models learned to imitate it.

So a detector and a checker flagging the same page tells you almost nothing about why. Only one of them will tell you what to change.

When is a detector the right tool?

When authorship is the actual question and consequences hang on the answer. Three situations where nothing else substitutes:

Academic integrity work, where somebody has made a claim about who produced an essay and an institution must assess it. Freelance and contract verification, where you paid for original work and want to know whether you received it. Compliance and provenance, where a policy or client agreement requires disclosure of generated material.

Across all three, pick whichever vendor publishes the lowest false-positive rate, then treat its output as one input to a conversation rather than a verdict. Pangram states a rate of 1 in 10,000 and frames a detection as "an invitation to talk, never as a presumption of guilt", which is the correct posture for this entire category.

When is a detector the wrong tool?

When you want to improve copy. Detection software answers a question whose answer you already hold.

You wrote that landing page, or commissioned it, or watched it emerge from a generator. A probability estimate about its origin tells you nothing new. What you lack is an inventory of the sentences costing you money, and probability output cannot supply one. A percentage passes judgement on a whole body of prose; judgement does not decompose into edits.

A second failure matters more commercially. Polished human writing gets flagged at meaningful rates, which explains why every honest roundup carries a false-positive warning. Run your own marketing copy through one, watch it return 80%, and you have gained nothing actionable. Worse, you may have picked up something actively misleading: a belief that origin, rather than craft, is your problem.

What does a rule-based checker do instead?

It names the pattern. Rather than scoring the page, it points at a sentence and says which rule fired on it.

That difference has a practical consequence. A rule is arguable. When SiteTell flags a sentence for , you can look at the sentence, decide the rule is wrong in this case, and move on. You cannot argue with 68%, because there is nothing in it to disagree with.

Flagged — negative_parallelism
Suggested rewrite
Illustrative. Nothing here is quoted from a scanned site.

It also means the score decomposes. Every point deducted traces to a specific flag on a specific sentence, so a score of 62 is a list of things to do rather than a grade to feel bad about.

Does Hemingway do the same job?

Closer than anything else here, and it is the honest comparison. Hemingway reads for readability and style: sprawling sentences, passive voice, adverbs, dense paragraphs. Authorship never enters into it, which puts Hemingway on our side of the line.

Scope and specificity separate them. Hemingway operates on prose you paste in, one piece at a time, and its categories cover general writing quality rather than the particular register marketing language converges toward. It will never mention that your closing paragraph restates the three above it, or that four separate sections borrowed a three-item list purely for rhythm.

Editing a single essay? Hemingway is an excellent choice and has been for a decade. Trying to discover which among 200 live URLs reads worst? Neither Hemingway nor any detection vendor was built for that, because none of them crawl.

So which should you use?

Choose by whichever question you genuinely have.

Asking "did a person write this"? Use a detector, preferably one publishing its false-positive rate. Asking "is this written well"? Hemingway, or a human editor. Asking "which URLs across my site read as interchangeable, and what precisely is wrong inside them"? You want something that crawls and names rules.

The single combination worth avoiding: running your own marketing copy through detection software and treating the resulting number as a quality score. That number measures something else entirely, and the something else is not available for you to fix.

Questions people actually ask

Is SiteTell an AI detector?

No. It makes no claim about whether a person or a model wrote your copy, and it does not produce a probability. It flags specific sentences against a deterministic ruleset and names the rule that fired on each one.

How accurate are AI detectors?

Accuracy varies widely by tool and by test set, and independent benchmarks disagree with each other and with vendors' own figures. The consistent finding across them is that false positives on polished human writing are common enough that no detector output should be treated as proof on its own.

Why did my own writing get flagged as AI?

Usually because it is written in a register that models also produce. Marketing and business prose converged on a particular set of phrases and rhythms well before generative models existed, so a detector trained on that output will flag human text written the same way.

Can I check a whole website rather than one page?

Not with a detector. Detectors take pasted text or file uploads, and Originality.ai additionally accepts a single URL. A full-site crawl is a different kind of tool: SiteTell scans up to 400 pages per site and orders results worst page first.

Does generic copy actually hurt search rankings?

Search guidance has moved toward rewarding content that shows first-hand experience and specificity, and away from content that reads as interchangeable. The more reliable argument is commercial rather than algorithmic: copy that reads like everyone else's gives a reader no reason to pick you.

Want to know which of these your site does?

SiteTell crawls every page and flags the exact sentences, with the rule that fired next to each one. Free for the whole site.

Scan your site free
SiteTell

The read a careful editor would give your site, run in twenty seconds. Built by Rally Digital.

PRODUCT
COMPANY
LEGAL
© 2026 Rally Digital