All guides

AI content

How to humanize AI content, and why detector scores are the wrong target

Seven detectors flagged 61% of human written essays as AI. Rewriting for a detector score is a trap. Here is the edit that stops a ChatGPT draft reading like one.

Bar chart of the share of essays seven AI detectors flagged as AI generated: human non native writers 61.2%, human US eighth graders 5.2%, AI written before a rewrite 100%, AI written after one rewrite 13%

Two different jobs are hiding under the phrase humanize AI content. One is making a draft read as though a person with something to say wrote it. The other is making a detector return a lower score. They are not the same work, and only the first one is worth doing. The evidence on the second is worse than most people writing about it admit.

What the detector evidence actually says

Seven widely used GPT detectors were tested against 91 TOEFL essays written by real people whose first language was not English, and 88 essays written by US eighth graders.

Sample, all human written Average share flagged as AI
TOEFL essays, non native English writers 61.22%
US eighth grade essays 5.19%

Across the TOEFL set, 89 of the 91 essays were flagged by at least one detector and 18 were flagged by all seven. The pattern is not random. Detectors key on the statistical evenness of text, and someone writing carefully in a second language produces evenly built sentences for entirely human reasons.

The other half of the study is the part that ends the argument. When the researchers asked ChatGPT to rewrite genuinely AI generated college essays in more literary language, detection fell from 100% to 13%. On scientific abstracts it fell from as much as 68% to as much as 28%.

So the same tool fails honest writers most of the time and clears deliberate AI text almost entirely, after one prompt. A score from it is not evidence of anything, in either direction, and building an editing process around raising or lowering that number is optimising a broken instrument.

Google’s position, in Google’s words

This matters because most humanizing advice is sold on an implied threat of a penalty that the documentation does not describe.

Google published its guidance on 8 February 2023 under the heading Rewarding high-quality content, however it is produced, and stated that its focus is on “the quality of content, rather than how content is produced”. The FAQ on the same post answers the question directly: appropriate use of AI or automation is not against the guidelines.

The line is drawn somewhere narrower. Google’s spam policy says that if you use automation, including AI generation, to produce content for the primary purpose of manipulating search rankings, that is a violation. Thin pages produced at volume to catch queries are the target, and they were the target when humans were producing them a decade ago.

Nobody is coming to fine you for using a model. The reason to edit an AI draft is that an unedited one does not earn attention, links or citations, which is a slower and more expensive problem than a penalty.

The tells worth removing

Read a raw draft looking for these six things. They are what makes AI prose recognisable to a reader long before any tool is involved.

Claims with no specifics attached. A model will write that a strategy improves performance significantly. It cannot write that it moved a client from 4,100 to 6,800 sessions in nine weeks, because it does not know that. Every sentence in the draft that makes a claim with no number, date, name or source is a sentence only you can finish.

Hedging that carries no information. It is important to note, in today’s landscape, can play a crucial role. Cut every one of them. The paragraph survives.

The negation flourish. It is not just about X, it is about Y. Once per article is a rhythm. Four times is a signature.

Uniform rhythm. Models produce sentences of similar length in paragraphs of similar length. Human writing is lumpy. A three word sentence after a long one is not a style trick, it is what happens when a person is emphasising something.

Symmetrical lists. Five bullets, each one sentence, each the same shape. Real advice is unbalanced, because the second item genuinely deserves a paragraph and the fifth deserves half a line.

The closing summary. The paragraph that restates what the article said. It exists because the model was trained to close, not because a reader needed it. Delete it and end on the last real point.

Making ChatGPT specifically stop sounding like ChatGPT

The six tells above are general. If your drafts come from ChatGPT in particular, three of them dominate, and knowing which lets you edit faster.

The hedge stack is the loudest. ChatGPT opens paragraphs with a qualifying clause more persistently than the other assistants, and it is the first thing a reader registers. Delete the clause before the comma and check whether the sentence lost anything. It usually has not.

It closes everything. Sections get a summarising last line, articles get a summarising last paragraph, and lists get a sentence explaining the list. This is the single highest yield deletion in a ChatGPT draft because it is mechanical: find the last paragraph of each section and ask whether it introduced anything.

It is relentlessly balanced. Every list comes out the same length, every argument gets its counterargument, every section gets roughly the same word count. Real writing is lopsided because real knowledge is lopsided. The fix is not to add asymmetry as a style trick; it is to expand the part you actually know about and cut the part you were padding.

Two things that will not help, both of which are widely recommended. Telling ChatGPT to “write like a human” or “avoid AI language” in the prompt shifts the vocabulary and leaves the structure, which is what a reader is actually noticing. And asking it to vary sentence length produces varied sentences with the same empty content. The problem was never the prose. It is that the model wrote around the specifics it does not have, and only you can put those back, which is what the rest of this page is about.

If you would rather have this as one paste, the AI humanizer prompt does the diagnosis and marks every gap it cannot fill rather than inventing something to fill it with.

The edit that does the work

Cutting the tells makes a draft inoffensive. It does not make it worth publishing. That takes an addition, and there are only four kinds of thing you can add that a model could not have produced.

  1. A number you measured. Yours, with the date it was measured and how. One real figure carries more weight than a page of adjectives, and it is the single most quoted unit in this whole category.
  2. A thing that went wrong. What you tried that failed, and what it cost. Models are trained toward the consensus recommendation, so the failure case is almost always missing, and it is the part practitioners actually read.
  3. A named specific. The tool, the version, the client sector, the exact setting. Generic examples are the loudest tell in the draft and the easiest to fix.
  4. A judgement with a reason. Not both options are valid. Which one you use, and what would have to be true for you to switch.

If you would like the structure of that pass written out as a runnable prompt, including the constraint that stops the model inventing the specifics it does not have, we published it as the AI humanizer prompt.

Before and after, on one paragraph

The draft:

In today’s digital landscape, page speed is more important than ever. Studies have shown that slow loading websites can significantly impact user experience and conversion rates. It’s not just about speed, it’s about delivering a seamless experience that keeps visitors engaged. By optimising your site’s performance, you can improve both rankings and revenue.

The edit, with the figures standing in for the ones you would supply from your own project:

Page speed is a conversion problem before it is a ranking one. On a checkout page where largest contentful paint went from 4.1s to 1.9s, completed orders moved 11% in the following month with no other change shipped in that window, and rankings did not move at all. That is the usual shape of it: the search benefit is real and slow, and the revenue benefit arrives immediately, which is why the work is easier to fund as a revenue project than as an SEO one.

Same subject, same length. The second one contains four things that had to come from a person: a measurement, its before and after, the window it was observed in, and a judgement about how to fund the work. It is also quotable by an assistant answering does page speed affect conversions, because it answers that question in one self contained passage with a figure in it.

The numbers in that rewrite are illustrative and they are the one part of this you cannot outsource. Writing them from your own data is the job. Inventing them, or letting a model invent them, produces something that reads human and is worse than the draft it replaced, because it is now confidently wrong rather than merely empty.

Why this is the same work as getting cited

Assistants do not quote articles. They quote passages, and they favour passages that answer one question, stand alone, and contain something checkable.

Generic AI prose fails that test for a structural reason rather than an aesthetic one. There is nothing in the sentence to lift. “Page speed is more important than ever” is not an answer to anything, so no retrieval system has a use for it, and no human has a reason to link to the page it sits on.

The rewrite above is quotable because it has a fact, a magnitude and a boundary in one place. That is the entire mechanism behind what gets cited in AI Overviews, and it means the editing pass that removes the AI smell and the editing pass that makes a page citable are one pass.

You can check the result mechanically. The AI content readiness checker reads a live URL and reports whether a clean passage can actually be extracted from it, which is the closest thing to a detector worth paying attention to, because it measures something that has a consequence.

What to skip

Humanizer tools that rewrite for a score. Most substitute synonyms and shuffle clauses. The output says the same unsupported thing in slightly worse English, and the specifics are still missing because no tool can supply them.

Anything that inserts invisible characters or homoglyphs. This is text designed to read one way to a machine and another to a person. It breaks copy and paste, it breaks screen readers, and it is a manipulation attempt on a system that has been detecting exactly this class of trick for twenty years.

Rewriting until a detector goes green. You now know what that number is worth. You will also be editing away the evenness that makes technical writing clear, which is a real cost paid for a fake benefit.

Disclosure theatre. Adding a line saying parts of this were AI assisted changes nothing about ranking or citation. Publish it if it is true and you want to, but do not treat it as a fix for a thin page.

The order to do it in

  1. Read the draft once and mark every claim with no evidence attached. That marked list is your actual writing task.
  2. Fill those in, or cut the claim. A claim you cannot support is a liability whoever wrote it.
  3. Cut the hedges, the closing summary and the symmetry.
  4. Read one paragraph aloud. Rhythm problems are audible and invisible.
  5. Check that at least three passages answer one question each, on their own, in about forty to sixty words.
  6. Run the page through a free audit to confirm the structure holds up: headings that match the questions, extractable chunks, and nothing blocking the crawlers that would make the whole exercise moot.

The draft was never the deliverable. What you know is the deliverable, and the model is a way of getting a structure onto the page fast enough that you still have the afternoon to put yourself into it.

Sources

All read on 11 August 2026.

Questions people ask

How do you humanize AI generated content?

Put in what only you could have written, then cut what any model could have written. In practice that is four edits: add a specific number, date, name or first hand result to every claim that has none, delete hedging and transition sentences that carry no information, break the uniform sentence rhythm a model produces, and remove the closing paragraph that restates the article. A draft that survives those four edits reads as human because a human contributed something to it.

Do AI humanizer tools work?

They change detector scores and they do not reliably change quality. Most work by swapping words for synonyms and rearranging clauses, which leaves the text saying the same generic thing less clearly, and some insert invisible characters, which is a manipulation a search engine can see and you cannot. Neither adds the specifics that make writing worth reading.

Are AI detectors accurate?

They are unreliable in both directions and the published evidence is blunt about it. A Stanford study of seven detectors misclassified 61.22% of TOEFL essays written by real people as AI generated, while flagging only 5.19% of essays by US eighth graders. In the same study a single rewrite prompt dropped detection of genuinely AI written college essays from 100% to 13%. A test that fails honest writers and passes anyone who asks for a rewrite is not a test to build a process on.

Will Google penalise AI generated content?

Not for being AI generated. Google published its position on 8 February 2023 under the heading "Rewarding high-quality content, however it is produced", and its spam policy is narrower than most people assume: using automation to produce content primarily to manipulate search rankings is the violation. Mass produced thin content is the problem whether a person or a model typed it.

Does humanizing AI content help it get cited by ChatGPT or AI Overviews?

Indirectly, and it is the better reason to do the work. Assistants quote passages that answer one question and stand on their own, and generic AI prose is made of sentences that cannot be lifted because they contain no fact to lift. Adding the number, the date and the source that a model cannot invent is simultaneously what makes a draft read like a person and what makes it quotable.

How much of an AI draft should be rewritten?

Judge it by what is in the draft rather than by percentage. The parts a model handled well, structure, ordering, obvious explanations, can usually stand. Every sentence that makes a claim needs a human to supply the evidence, the caveat or the example. In most drafts that is a quarter of the text and it is the quarter that carries the whole piece.

Run your first audit
in about a minute

Free account, no card. Paste your URL and get a real, scored report of your AI and search visibility.

Measuring rankings in