People who ask whether their website sounds AI-written are rarely asking who wrote it. They already know. What they want to find out is whether visitors will feel the copy was produced rather than written, and if so, which sentences cause that.
An AI detector cannot answer that. It returns a probability about authorship, and a page can score as human while reading like every competitor, or score as machine while being perfectly clear. The real question concerns patterns a reader notices, and patterns can be named, counted and fixed.
SiteTell's ruleset names 114 of them. Sixteen describe shape and rhythm; the other 98 are words and phrases. Below is every family, what each rule looks for, and how often it turned up. Of the 80 sites whose stored text we rescored, 74 tripped at least one structural rule and 54 used at least one listed word.
How these numbers were measured
Each figure comes from the stored text of each site's most recent scan, rescored on 2026-09-28 by the same scoring code the live product runs, under ruleset 2026-09-16. Wherever we could compare, the rescore matched the original scans flag for flag. Sites scanned before we began storing text are excluded, hence 80 rather than 103.
Three formatting rules depend on HTML measurements taken mid-crawl and never kept, so their figures come from the original scans, across all 103 sites, and are marked below.
The sample is self-selected: owners (or curious visitors) chose to run these scans, so rates are probably higher than across the web generally. No site is named.
Family one: the words
The most familiar family, though not the one with the widest reach. Its 98 entries split into inflated verbs, unprovable praise, borrowed metaphors, filler openers, puffed-up phrases, unsourced claims and chatbot leftovers.
Dull ones matter most. appears on 38% of sites, on 35% and on 33%, while the famous reaches just 13%. Each entry, with its prevalence and a plainer alternative, is listed on our page of AI words to avoid.
Family two: sentence shapes
These rules read one sentence at a time, looking for a construction whatever its vocabulary. Because every flag quotes its sentence, they are the easiest group to fix, and several rank among the strongest signals we have.
Deny-then-inflate
A figure denying a small claim to assert a bigger one, as in . Half of all rescored sites carried it, despite production scans missing it entirely until 2026-09-16. Severity is high, and our full page on negative parallelism covers its four shapes.
Lists of three, on repeat
A trio joined by "and" sounds finished, so generated prose reaches for three whether or not three things exist. One trio is harmless. Four or more on a page trips the rule, and 63% of sites crossed that line. Often the honest fix is two items, or four, since there usually were four.
Stock openers
Preambles carrying no information, such as or , found on 28% of sites. Deleting the opener and starting on the point is nearly always enough.
A tacked-on "-ing" clause
A plain fact followed by a comma and a participle praising it: . Significance gets added without content. One site in five showed it, at high severity; stop at the fact.
Buried verbs
An action turned into a noun and propped up with a carrier, so "review" becomes . Sixteen percent of sites had one. Restoring the verb shortens the sentence by two words.
Formal words where plain ones fit
A short list of formal terms, among them and , counted against a threshold of six per 500 words. That density is uncommon; 14% reached it. "Use" and "help" nearly always work.
Claim, then hedge
A sentence staking out a position, then withdrawing it mid-breath through a qualifier. This check is live, yet only one of those 80 sites triggered it, too few for any rate worth quoting.
Family three: the rhythm of a whole page
Statistics over an entire page, so no lone sentence trips them and no lone edit clears them. They are hard to see in your own writing and, taken together, the most widespread family.
A closing paragraph that repeats the page
Our most common finding, on 68% of sites. It triggers when over 60% of a final paragraph's meaningful words appeared earlier, which is precisely what a summary does. Models are trained to round things off. Anyone who reached the end does not need the beginning again, so deletion usually wins.
Low variety of words
The share of distinct words, flagged below 35% across at least 150. Sixty-one percent of sites fell short. Thin variety seldom signals a small vocabulary; more often, prose circled one point and restated its noun for lack of anything new.
Too little punctuation
Runs opposite to everything else here, tripping below four commas, semicolons or parentheses per 100 words, with 56% of sites under. Clipped sentences strung together with "and" dodge the commitment subordination demands, making them a model's safe output when asked for punch. Combine related thoughts; let clauses hang off each other.
Sentences all the same length
Human writing is lumpy. It runs long, then stops. Generated prose settles into a comfortable middle, and this check flags too little spread across six or more sentences, which 41% of sites showed.
Paragraphs all the same length
That measurement again, at paragraph scale, across four or more. A tenth of sites had it, typically where every section came from an identical brief.
Family four: formatting
Habits of layout and typography rather than wording.
Em dashes as a crutch
Judged by rate, never presence: counting begins near one dash per 500 words. Sixty-three percent of sites crossed it, and homepages ran about 5.6 times the blog rate. Our em dash page covers thresholds, exemptions and why we kept this rule after weighing its retirement.
Pages made of bullet points
Tripped when list items outnumber paragraphs among a page's text blocks. Original scans put it on 51% of all 103 sites. Bullets are fine; what goes missing is the argument connecting them.
Bullets led by emoji or symbols
Most list items opening with an emoji or decorative glyph instead of a word. Original scans found this on 6% of sites.
Bold used for emphasis everywhere
Over 15% of words set bold. Live, but zero of 103 sites triggered it, which we report rather than omit.
Why one hit proves nothing, and several do
Each rule here matches things people write unaided. What makes copy read as generated is accumulation: several patterns from different families arriving together.
Co-occurrence shows this plainly. Pages carrying are about 20 times likelier than average to carry . Deny-then-inflate pages are about 6.6 times likelier to contain a buried verb too. Such habits travel in groups because they share an origin, and readers react to the group long before they could name any member.
That is why SiteTell scores each page and lists every flag, rather than issuing one verdict. A page with a single hit is almost certainly fine. One with a dozen, spread over three families, deserves rewriting first.
Where to start
Order the work by payoff, not by fame.
- Cut closing paragraphs that restate what came before. Most common pattern, one delete key.
- Rewrite every deny-then-inflate sentence. High severity, and each rewrite forces a vague claim to become specific.
- Swap inflated verbs for plain ones plus a number. Most vocabulary hits live here.
- Only then tackle rhythm, where fixes take genuine rewriting, and dashes, which cost least.