SiteTell
// comparison

GPTZero alternatives: four detectors, and what to use when authorship is not the question

Four AI detectors compared on what each one outputs, what it accepts, and the false positive rate it publishes. Plus what to reach for when you already know who wrote the copy and want to know why it reads generic.

data measured September 14, 2026ruleset 2026-09-12
About these figures
They were produced under ruleset 2026-09-12, which changed on September 16, 2026. A scan you run today uses the newer one, so your score may differ from what these figures imply. See the current dataset.

People search for a GPTZero alternative for two very different reasons, and the reasons want opposite tools.

Some have a genuine authorship problem. A teacher suspects an essay was generated, a publisher wants assurance before paying an invoice, a hiring manager received a cover letter that reads oddly. Those readers want a better detector, and several exist.

Others arrived from a different direction entirely. They pasted their own homepage into GPTZero, watched it come back at 82%, and felt confused, because they wrote that page themselves. Nothing in the detector category will help them, though the reason why turns out to be interesting.

This page covers both. The first half compares four detectors on what each one reports; the second half explains what to reach for when the percentage was never the thing you needed.

Which detectors are worth comparing against GPTZero?

Four, on the evidence each publishes about itself.

What four AI detectors output, what they accept as input, the false positive rate each one publishes, and their free tiers
GPTZeroOriginality.aiWinston AIPangram
What you get backA percentage, with sentence-level highlightsAn AI score plus a plagiarism result, and a shareable PDF reportA score from 0 to 100, with a colour-coded prediction mapA classification, plus a breakdown of which segments were AI-assisted
What it takes as inputPasted text or an uploaded filePasted text, uploads, or a URL, with bulk scanningPasted text or uploads, plus images read by OCR including handwritingPasted text, or PDF, DOCX and RTF files, up to 100 at once
Stated false positive rate1% for English-as-a-second-language writers, after de-biasingPublished per model in their own accuracy studyNot published; states 99.98% accuracy1 in 10,000
Covers a whole siteNoOne page at a timeNoNo
Free tierUp to 10,000 characters without an accountThree scans a day, 2,000 words, no account neededA trial, with no card requiredTwenty checks a day once you sign up
Who it is built forEducators and institutions, plus publishers and recruitersPublishers, SEO teams and agenciesEducation, publishing and SEO teamsUniversities, publishers, law firms and HR teams
Claims checked September 22, 2026

Read the third row carefully, since it sells most of these products while being the least suited to comparison. Every figure there was measured by whoever quotes it, on a corpus they assembled, under a definition they chose. Pangram's 1 in 10,000 and GPTZero's 1% describe different populations; neither firm audited the other. Treat each as a public commitment a company is willing to stand behind, genuinely valuable, rather than a benchmark you can rank.

Why do the published accuracy figures disagree so much?

Because performance here depends almost entirely on which texts get tested, and every vendor picks its own.

A classifier scoring 99% against essays produced by one model family, at default settings, in a single sitting, may behave quite differently on a paragraph somebody wrote and then revised twice. Independent evaluations repeatedly report lower numbers than vendors claim, and find the gap widening on revised passages, on very short extracts, and on prose by authors whose first language is other than English. GPTZero's site addresses the last category head-on, quoting a 1% false positive rate for second-language writers following retraining.

An honest summary, then: detection behaves respectably in easy conditions, degrades in hard ones, and the hard ones are precisely where somebody's grade or invoice hangs in the balance.

Which alternative suits which job?

For publishing and SEO teams, Originality.ai owns the most operationally practical shape of the four. It swallows a URL rather than demanding pasted text, runs bulk jobs, bundles plagiarism checking alongside classification, and emits a shareable PDF settling any dispute with a freelancer without resorting to screenshots.

For academic integrity work, Pangram publishes the most precise false positive claim anywhere on this page, at 1 in 10,000, and plugs into the learning management platforms institutions already operate. Its framing counts for as much as its arithmetic: a hit becomes grounds for a conversation, never proof of misconduct.

For document-heavy workflows, Winston AI reads scanned sheets and photographs via OCR, handwriting included, which nobody else here attempts. Should your material arrive as pictures of paper, that lone capability outweighs every other consideration.

For occasional spot checks, GPTZero stays hard to beat on sheer convenience, since 10,000 characters sail through without any account whatsoever.

How do they handle text a person has edited?

Worse than they handle raw model output, which is awkward, because edited text describes most of what anybody actually publishes.

A paragraph drafted by a model and then reworked by a person carries fewer of the statistical regularities detectors were trained to recognise. Several vendors now acknowledge this in the shape of their output rather than in a disclaimer. Pangram reports which segments appear assisted instead of returning a single verdict for a whole document, and classifies material as AI-assisted as a category of its own. Originality.ai highlights at sentence level, letting you see where confidence concentrates.

Treat any tool returning one number for a long document with more caution than one that shows you where the number came from. The second kind at least lets you check whether the evidence sits in the passage you care about.

How should you test one before trusting it?

With your own material, before anything is at stake.

Take four or five pieces of writing you know the provenance of, including something a colleague wrote years ago and something deliberately unremarkable, then run all of them through whichever tool you are evaluating. What you are looking for is the false positive rate on your own content, in your own register, which is the only figure that predicts how the tool will behave in your workflow.

Vendors cannot run that test for you, since they do not have your corpus. Doing it yourself takes perhaps twenty minutes and will tell you more than any published benchmark.

What about Copyleaks and Turnitin?

Both belong in this category, and we excluded both deliberately.

Turnitin sells through institutions rather than individuals, making it seldom a live option for anybody weighing up software on a Tuesday afternoon. Copyleaks enjoys wide adoption and frequent recommendation, yet its servers refuse automated requests, so nothing could be read off its own pages while we wrote this. The principle behind the omission governs every claim above: specifications come straight from whoever built the product, never from somebody else's roundup. Roundups decay silently, and one wrong number sitting beside a correct one poisons both.

What if authorship was never the question?

Then no detector on this page can help, and the reason is worth understanding before you spend another evening testing them.

Consider running your own marketing copy through detection software. You already know who produced it. A probability adds nothing you lacked beforehand, and should it return high, you now hold a figure with no action attached. Percentages cannot be edited.

A second problem exists, and our own corpus speaks to it. Across 97 domains and 10,126 pages measured on 14 September 2026, 51% of sites carried at least one blacklisted word. A majority. Plenty of the underlying prose predates generative models entirely: copy drafted in 2018, brochure phrasing inherited through a rebrand, a founder grabbing whichever expression came to hand. Business writing drifted toward this register long before software learned to imitate it.

So a high score on your own domain frequently measures convergence rather than origin. Your pages sound like everybody else's pages. Such a diagnosis is useful and directly actionable, yet stays wholly invisible inside a percentage.

What reads copy without guessing who wrote it?

Tools that report patterns instead of probabilities.

How two tools that make no authorship claim differ in what they read and what they report
HemingwaySiteTell
What you get backA readability grade, plus style highlightsThe rule that fired, quoted on the exact sentence
What it takes as inputPasted or typed textA domain
How much it reads at onceOne piece of writingUp to 400 pages per scan
Makes a claim about who wrote itNoNo
Free tierThe online editor, with paid tiers above itThe whole scan, every page, no signup
Claims checked September 22, 2026

Hemingway has done this well for over a decade, grading readability and marking passages that run long or lean on passive constructions. It reads whatever you paste into it.

SiteTell works from the other end. You give it a domain, it crawls the site, and for every sentence that trips a rule it reports which rule and why. No probability appears anywhere in the output, and no claim about authorship is made or implied.

The practical difference lies in what you can argue with. A rule is contestable in a way a score never is.

Flagged — rule_of_three_pattern
Suggested rewrite
Illustrative. Nothing here is quoted from a scanned site.

Look at that flag and you can decide we are wrong about your sentence, keep it, and carry on. Try doing that with 82%. There is nothing inside the number to disagree with, which is precisely why it cannot be turned into an edit.

So how should you choose?

Start from the question you actually hold, because it determines the category before it determines the vendor.

Do you need to know whether a person wrote something, with consequences attached to the answer? Pick a detector, favour whichever publishes its false positive rate most specifically, and treat what it returns as one input among several. Do you want to know why your own pages read like everybody else's? Detection software answers a question you already have the answer to; reach for something that names patterns instead.

The one move worth avoiding is treating a detection percentage as a quality grade on your own copy. It measures something else, and that something else cannot be edited.

Questions people actually ask

What is the best free GPTZero alternative?

Originality.ai has the most generous no-account tier of the four compared here, at three scans a day and 2,000 words, and it accepts a URL rather than only pasted text. Pangram allows twenty checks a day once you create an account, and publishes the most specific false positive rate of any vendor on this page.

Is GPTZero accurate?

GPTZero reports 99% accuracy and a 1% false positive rate for second-language writers after retraining. Independent evaluations generally find lower figures, particularly on edited text and short passages. No detector should be treated as proof on its own, which most vendors in this category now say themselves.

Why does my own writing get flagged as AI?

Usually because it sits in a register that models also produce. Marketing and business prose converged on a recognisable set of phrases and rhythms well before generative models existed, so a detector trained on model output will flag human text written the same way. In our own scan of 97 domains, 51% carried at least one blacklisted word, and much of that copy predates the models entirely.

Can any AI detector scan a whole website?

No. Detectors take pasted text or file uploads, and Originality.ai additionally accepts a single URL, which checks one page rather than crawling a site. Auditing every page is a different kind of tool: SiteTell crawls up to 400 pages from a domain and orders the results worst page first.

Is SiteTell an AI detector?

No, and it is listed on this page as the adjacent option rather than a competitor to the four detectors above. It makes no claim about whether a person or a model produced your copy and returns no probability. It flags sentences against a deterministic ruleset and names the rule that fired on each one.

What should I do about a false positive on my own content?

Nothing, if you know you wrote it. A false positive tells you the writing resembles model output statistically, which says more about the register than about the author. If several pages come back high, the useful follow-up is to find out which specific patterns are causing it, since those are editable in a way a score is not.

Do detectors work on text that has been edited by a human?

Less reliably. Accuracy on mixed or revised text is consistently lower than on unedited model output, which is why some vendors now report an assisted or mixed classification instead of a binary one. Pangram breaks results down by segment for this reason.

Does generic copy hurt search rankings?

Search guidance has moved toward rewarding pages that demonstrate first-hand experience and specificity. The stronger argument is commercial rather than algorithmic: copy indistinguishable from a competitor's gives a reader no particular reason to choose you, whatever a ranking algorithm makes of it.

Want to know which of these your site does?

SiteTell crawls every page and flags the exact sentences, with the rule that fired next to each one. Free for the whole site.

Scan your site free
SiteTell

The read a careful editor would give your site, run in twenty seconds. Built by Rally Digital.

PRODUCT
COMPANY
LEGAL
© 2026 Rally Digital