Two tools that look almost identical from the outside are answering completely different questions, and picking the wrong one wastes a week.
An AI detector estimates whether a machine produced a piece of text. It takes what you paste in and returns a likelihood, usually a percentage. A writing checker makes no claim about authorship at all. It looks at the text and tells you which specific patterns in it read as generic, and why.
If you are a teacher deciding whether a student wrote their own essay, you want the first kind. If you are a founder wondering why your pricing page reads like every other pricing page, the first kind cannot help you, because "68% likely AI" is not an edit you can make.
The short version
| GPTZero | Originality.ai | Pangram | Hemingway | SiteTell | |
|---|---|---|---|---|---|
| What you get back | A percentage likely to be AI, with the phrases highlighted | An AI score and a plagiarism result, with sentence highlights | A classification: human, AI-generated, or AI-assisted | A readability grade, plus style highlights | The rule that fired, quoted on the exact sentence |
| What it takes as input | Pasted text or an uploaded file | Pasted text, an upload, or one URL | Pasted text or uploads, up to 100 files | Pasted or typed text | A domain |
| Covers a whole site | No | One page at a time | No | No | Yes, up to 400 pages per scan |
| Makes a claim about who wrote it | Yes | Yes | Yes | No | No |
| What a wrong answer costs you | An accusation the writer has to disprove | A rejected draft, or a flagged freelancer | An accusation, at a stated rate of 1 in 10,000 | A suggestion you ignore | A flag you disagree with, on a sentence you can see |
| Free tier | Under 10,000 characters without an account | Three scans a day, 2,000 words | Twenty checks a day with an account | Yes, with a paid tier above it | The whole scan, every page, no signup |
| Who it is built for | Educators and institutions | Publishers, SEO teams, agencies | Educators | Anyone editing prose | Teams auditing a live marketing site |
Why does the distinction matter if both flag the same text?
Because they disagree far more often than people expect, and their disagreement carries the useful information.
Detectors ask about origin. Writing checkers ask about effect. A paragraph can be entirely human and still read as interchangeable, and a paragraph can be machine-drafted, heavily edited, and land sharp and particular. Those two facts are what keep the categories separate.
We have a number for the first half of that. Across 97 domains and 10,126 pages, 51% of sites carried at least one word from our own blacklist. That is a majority of sites, and much of the writing underneath those flags is unambiguously human: copy written in 2019, brochure language inherited from a rebrand, a founder's own draft that reached for the nearest available phrase. Marketing prose converged on this register long before models learned to imitate it.
So a detector and a checker flagging the same page tells you almost nothing about why. Only one of them will tell you what to change.
When is a detector the right tool?
When authorship is the actual question and consequences hang on the answer. Three situations where nothing else substitutes:
Academic integrity work, where somebody has made a claim about who produced an essay and an institution must assess it. Freelance and contract verification, where you paid for original work and want to know whether you received it. Compliance and provenance, where a policy or client agreement requires disclosure of generated material.
Across all three, pick whichever vendor publishes the lowest false-positive rate, then treat its output as one input to a conversation rather than a verdict. Pangram states a rate of 1 in 10,000 and frames a detection as "an invitation to talk, never as a presumption of guilt", which is the correct posture for this entire category.
When is a detector the wrong tool?
When you want to improve copy. Detection software answers a question whose answer you already hold.
You wrote that landing page, or commissioned it, or watched it emerge from a generator. A probability estimate about its origin tells you nothing new. What you lack is an inventory of the sentences costing you money, and probability output cannot supply one. A percentage passes judgement on a whole body of prose; judgement does not decompose into edits.
A second failure matters more commercially. Polished human writing gets flagged at meaningful rates, which explains why every honest roundup carries a false-positive warning. Run your own marketing copy through one, watch it return 80%, and you have gained nothing actionable. Worse, you may have picked up something actively misleading: a belief that origin, rather than craft, is your problem.
What does a rule-based checker do instead?
It names the pattern. Rather than scoring the page, it points at a sentence and says which rule fired on it.
That difference has a practical consequence. A rule is arguable. When SiteTell flags a sentence for , you can look at the sentence, decide the rule is wrong in this case, and move on. You cannot argue with 68%, because there is nothing in it to disagree with.
It also means the score decomposes. Every point deducted traces to a specific flag on a specific sentence, so a score of 62 is a list of things to do rather than a grade to feel bad about.
Does Hemingway do the same job?
Closer than anything else here, and it is the honest comparison. Hemingway reads for readability and style: sprawling sentences, passive voice, adverbs, dense paragraphs. Authorship never enters into it, which puts Hemingway on our side of the line.
Scope and specificity separate them. Hemingway operates on prose you paste in, one piece at a time, and its categories cover general writing quality rather than the particular register marketing language converges toward. It will never mention that your closing paragraph restates the three above it, or that four separate sections borrowed a three-item list purely for rhythm.
Editing a single essay? Hemingway is an excellent choice and has been for a decade. Trying to discover which among 200 live URLs reads worst? Neither Hemingway nor any detection vendor was built for that, because none of them crawl.
So which should you use?
Choose by whichever question you genuinely have.
Asking "did a person write this"? Use a detector, preferably one publishing its false-positive rate. Asking "is this written well"? Hemingway, or a human editor. Asking "which URLs across my site read as interchangeable, and what precisely is wrong inside them"? You want something that crawls and names rules.
The single combination worth avoiding: running your own marketing copy through detection software and treating the resulting number as a quality score. That number measures something else entirely, and the something else is not available for you to fix.