SiteTell
// guide

What makes a website sound AI-written: every pattern, and how often each appears

Sixteen structural patterns and 98 words, each with how often it turned up across 80 real sites and what fixes it. Why a single hit proves nothing, and where to start.

data measured September 28, 2026ruleset 2026-09-16

People who ask whether their website sounds AI-written are rarely asking who wrote it. They already know. What they want to find out is whether visitors will feel the copy was produced rather than written, and if so, which sentences cause that.

An AI detector cannot answer that. It returns a probability about authorship, and a page can score as human while reading like every competitor, or score as machine while being perfectly clear. The real question concerns patterns a reader notices, and patterns can be named, counted and fixed.

SiteTell's ruleset names 114 of them. Sixteen describe shape and rhythm; the other 98 are words and phrases. Below is every family, what each rule looks for, and how often it turned up. Of the 80 sites whose stored text we rescored, 74 tripped at least one structural rule and 54 used at least one listed word.

How these numbers were measured

Each figure comes from the stored text of each site's most recent scan, rescored on 2026-09-28 by the same scoring code the live product runs, under ruleset 2026-09-16. Wherever we could compare, the rescore matched the original scans flag for flag. Sites scanned before we began storing text are excluded, hence 80 rather than 103.

Three formatting rules depend on HTML measurements taken mid-crawl and never kept, so their figures come from the original scans, across all 103 sites, and are marked below.

The sample is self-selected: owners (or curious visitors) chose to run these scans, so rates are probably higher than across the web generally. No site is named.

Family one: the words

The most familiar family, though not the one with the widest reach. Its 98 entries split into inflated verbs, unprovable praise, borrowed metaphors, filler openers, puffed-up phrases, unsourced claims and chatbot leftovers.

Dull ones matter most. appears on 38% of sites, on 35% and on 33%, while the famous reaches just 13%. Each entry, with its prevalence and a plainer alternative, is listed on our page of AI words to avoid.

Family two: sentence shapes

These rules read one sentence at a time, looking for a construction whatever its vocabulary. Because every flag quotes its sentence, they are the easiest group to fix, and several rank among the strongest signals we have.

Deny-then-inflate

A figure denying a small claim to assert a bigger one, as in . Half of all rescored sites carried it, despite production scans missing it entirely until 2026-09-16. Severity is high, and our full page on negative parallelism covers its four shapes.

Lists of three, on repeat

A trio joined by "and" sounds finished, so generated prose reaches for three whether or not three things exist. One trio is harmless. Four or more on a page trips the rule, and 63% of sites crossed that line. Often the honest fix is two items, or four, since there usually were four.

Stock openers

Preambles carrying no information, such as or , found on 28% of sites. Deleting the opener and starting on the point is nearly always enough.

A tacked-on "-ing" clause

A plain fact followed by a comma and a participle praising it: . Significance gets added without content. One site in five showed it, at high severity; stop at the fact.

Buried verbs

An action turned into a noun and propped up with a carrier, so "review" becomes . Sixteen percent of sites had one. Restoring the verb shortens the sentence by two words.

Formal words where plain ones fit

A short list of formal terms, among them and , counted against a threshold of six per 500 words. That density is uncommon; 14% reached it. "Use" and "help" nearly always work.

Claim, then hedge

A sentence staking out a position, then withdrawing it mid-breath through a qualifier. This check is live, yet only one of those 80 sites triggered it, too few for any rate worth quoting.

Family three: the rhythm of a whole page

Statistics over an entire page, so no lone sentence trips them and no lone edit clears them. They are hard to see in your own writing and, taken together, the most widespread family.

A closing paragraph that repeats the page

Our most common finding, on 68% of sites. It triggers when over 60% of a final paragraph's meaningful words appeared earlier, which is precisely what a summary does. Models are trained to round things off. Anyone who reached the end does not need the beginning again, so deletion usually wins.

Low variety of words

The share of distinct words, flagged below 35% across at least 150. Sixty-one percent of sites fell short. Thin variety seldom signals a small vocabulary; more often, prose circled one point and restated its noun for lack of anything new.

Too little punctuation

Runs opposite to everything else here, tripping below four commas, semicolons or parentheses per 100 words, with 56% of sites under. Clipped sentences strung together with "and" dodge the commitment subordination demands, making them a model's safe output when asked for punch. Combine related thoughts; let clauses hang off each other.

Sentences all the same length

Human writing is lumpy. It runs long, then stops. Generated prose settles into a comfortable middle, and this check flags too little spread across six or more sentences, which 41% of sites showed.

Paragraphs all the same length

That measurement again, at paragraph scale, across four or more. A tenth of sites had it, typically where every section came from an identical brief.

Family four: formatting

Habits of layout and typography rather than wording.

Em dashes as a crutch

Judged by rate, never presence: counting begins near one dash per 500 words. Sixty-three percent of sites crossed it, and homepages ran about 5.6 times the blog rate. Our em dash page covers thresholds, exemptions and why we kept this rule after weighing its retirement.

Pages made of bullet points

Tripped when list items outnumber paragraphs among a page's text blocks. Original scans put it on 51% of all 103 sites. Bullets are fine; what goes missing is the argument connecting them.

Bullets led by emoji or symbols

Most list items opening with an emoji or decorative glyph instead of a word. Original scans found this on 6% of sites.

Bold used for emphasis everywhere

Over 15% of words set bold. Live, but zero of 103 sites triggered it, which we report rather than omit.

Why one hit proves nothing, and several do

Each rule here matches things people write unaided. What makes copy read as generated is accumulation: several patterns from different families arriving together.

Co-occurrence shows this plainly. Pages carrying are about 20 times likelier than average to carry . Deny-then-inflate pages are about 6.6 times likelier to contain a buried verb too. Such habits travel in groups because they share an origin, and readers react to the group long before they could name any member.

That is why SiteTell scores each page and lists every flag, rather than issuing one verdict. A page with a single hit is almost certainly fine. One with a dozen, spread over three families, deserves rewriting first.

Run the whole ruleset over your site and see which pages carry the most.
Scan your site free →

Where to start

Order the work by payoff, not by fame.

  1. Cut closing paragraphs that restate what came before. Most common pattern, one delete key.
  2. Rewrite every deny-then-inflate sentence. High severity, and each rewrite forces a vague claim to become specific.
  3. Swap inflated verbs for plain ones plus a number. Most vocabulary hits live here.
  4. Only then tackle rhythm, where fixes take genuine rewriting, and dashes, which cost least.

Questions people actually ask

How can I tell if my website sounds AI-written?

Look for patterns rather than single words: closing paragraphs that restate the page, the same sentence length throughout, lists of three on repeat, and the deny-then-inflate figure. A scan checks all 114 rules across up to 400 pages and quotes each flagged sentence.

Is this the same as an AI detector?

No. A detector estimates who or what wrote a text. SiteTell names patterns that make copy read as generic, whoever wrote it, and every flag points at a sentence you can change. The two answer different questions.

What is the most common sign of AI writing on websites?

In our rescored data, a closing paragraph that mostly repeats words already on the page, found on 68% of sites. The em dash rule and repeated lists of three follow, at 63% each.

Can human-written copy trip these rules?

Yes, often. Every one of these patterns predates language models. The ruleset measures whether copy reads as generic and trustworthy, so a flagged human sentence is still worth a second look.

Do search engines penalise AI-sounding copy?

We know of no reliable evidence that any engine penalises individual patterns like these, and we would distrust claims that it does. The case for fixing them rests on readers: vague, formulaic copy gives a visitor less reason to stay.

How many rules does SiteTell check?

One hundred and fourteen active rules as of ruleset 2026-09-16: sixteen structural rules measuring shape and rhythm, and 98 vocabulary rules matching words and phrases.

Why are some figures from the original scans instead of rescored?

Three formatting rules depend on measurements taken from the page's HTML during the crawl, and those measurements are never stored. Everything else was rescored from stored text under one ruleset so the numbers are comparable.

Does fixing one page change my overall site score?

Yes, and your weakest page moves it most. The site figure blends two halves: an average of every page weighted by length, and the single lowest-scoring page on its own. Lifting that bottom page shifts both halves at once.

Want to know which of these your site does?

SiteTell crawls every page and flags the exact sentences, with the rule that fired next to each one. Free for the whole site.

Scan your site free →
SiteTell

The read a careful editor would give your site, run in twenty seconds. Built by Rally Digital.

PRODUCT
COMPANY
LEGAL
© 2026 Rally Digital