on page SEO
On page SEO: the checklist barely changed, and the weighting did
Titles and meta descriptions are mostly fine on professional sites. What fails is heading order, extractable passages and attribution. Measured on 88 homepages.

On page SEO is everything on the page itself that affects how it is found and understood. The list of items has been broadly stable for a decade. What has moved is which of them are actually broken on real sites, and one item that was not on the list at all now decides whether a page gets quoted.
We can say which are broken rather than guess, because we measured it. Our audit of 98 agency homepages ran 36 checks against every site we could fetch, and the on page half of that result is the argument of this article.
The classic list is mostly done
On the 88 homepages we could fetch, titles, meta descriptions and image alt coverage each came back clean on roughly five sites in six. These are professional marketing sites, so the sample is flattering, but the direction is not in doubt: the items every checklist leads with are the items everybody has already done.
That does not make them unimportant. It makes them a poor place to spend an afternoon. Confirm them and move on:
- Title. One per page, front loaded with what the page is about, and short enough to survive truncation. Our title tag checker reports the rendered length.
- Meta description. Not a ranking factor and still worth writing, because it is ad copy for a result somebody is deciding whether to click. Checker here, and worked examples with scores here.
- Alt text. Every image needs the attribute. For a decorative one the correct value is empty, which is not the same as missing. The long version covers when to write nothing.
- URL shape. Readable, stable, lowercase, no session junk. Every site in our sample passed this. It is solved.
What actually fails
Three things, and none of them are on the front of a typical on page checklist.
Heading order, on nearly two thirds of pages
56 of the 88 homepages had a heading problem: a level that skips a rank, an H2 followed by an H4, or heading tags with no text in them. That is 64%, by far the most common on page defect we found.
This used to be an accessibility footnote. It is now structural, because the heading outline is what a model chunks a page on. A skipped level tells a machine that a subsection belongs to a parent that does not exist, and an empty heading tag is a section boundary with no label on it.
The fix is genuinely trivial. Headings descend one rank at a time, never skip, and never get used because of how big the text looks. Heading structure checker.
Extractability, on a third of pages
28 of 88 could not offer a passage worth lifting. This is the check that did not exist on the old list, and it is the one that decides whether you are cited when an assistant answers instead of linking.
Our engine scores it on four structural properties, and all four are things you can edit today:
- At least three sections bounded by headings. A page of unbroken prose is one undifferentiated block to a model.
- Sections that open with the answer. A short, direct first sentence under the heading, not a preamble that arrives at the point in the third paragraph. This carries the most weight of the four.
- Some headings phrased as questions, because that is the shape of a real prompt.
- Structured content present, a table or a real list. Models lift structured comparisons far more readily than they lift paragraphs.
Notice what is absent: word count, keyword density, and any measure of how well written it is. Extractability is a structural property. Excellent prose organised as one long argument scores badly, and that is not a flaw in the scoring.
The practical test is one question per section. If you cut this section out and showed it to somebody who had not read the page, would it still answer something? If it opens with “this” or “that” or “as we saw above”, it would not.
Run a page through AI content readiness to see the four properties scored on your own page.
Attribution, on five sites in six
74 of 88 had no clear author in the markup, or something too thin to resolve. Dates were only slightly better.
This is the cheapest fix in the whole set and the most neglected. A byline that exists as styled text but not as markup is invisible to anything reading the page mechanically, which now includes the systems deciding whether your page is a source worth naming.
The order to do it in
Technical first, always. The best on page work on a page that cannot be fetched or rendered is worth exactly nothing, and in our separate crawl of 1,048 of the most visited sites, 28.3% were blocking at least one major AI crawler in robots.txt. Check that floor before optimising anything above it.
Then, in this order, because it is ordered by what is actually broken rather than by tradition:
- Heading levels. An hour, sitewide, and it fixes the most common defect we measured.
- Author and date in the markup. A template change, once.
- Section openings. Rewrite the first sentence under each heading so it answers. This is the highest value writing work on the list and it is usually a one line edit per section.
- A table or a real list wherever you are comparing things in prose.
- Titles and descriptions, confirmed rather than agonised over.
- Internal links that point somewhere useful, with link text that says what is at the other end.
The technical SEO checklist covers the layer underneath this, and the free tools run each individual check against a single URL.
What is not on page work, despite being sold as it
Keyword density. It was never a thing you should tune, and nothing in any current guidance rewards it.
Word count targets. Length follows coverage. A target inverts that and produces padding, which our own content depth and readability checks tend to punish rather than reward.
Adding schema to a page that does not show what the schema claims. Structured data has to describe the visible page. Mismatched markup is a policy problem, not a shortcut, and it was the most common hard failure in the whole audit.
Any score out of 100, including ours. Google published guidance on third party SEO tools in June 2026 stating plainly that no third party tool has access to its internal ranking systems. Every difficulty score and visibility index in this category, ours included, models the outside of a system. That is useful, and it is not a reading from inside it.
The short version
The on page checklist did not change. The failure distribution did. Titles and descriptions are done on most professional sites; heading structure, extractable passages and basic attribution are not, and those three are precisely the ones that decide whether a machine reading your page can find a section, lift an answer out of it, and say who wrote it.
Sources
Read 16 August 2026. Measurements are our own and the raw data is published.
- Our agency homepage audit, 98 sites, 36 checks, collected 6 August 2026. Every percentage above is computed over the 88 sites that returned content, because the ten that answered 403 have nothing to check.
- Our AI crawler blocking study, 1,048 of the most visited websites, collected 11 August 2026, for the 28.3% crawler blocking figure.
- Google Search Central, third party SEO tools, services, and advice, for the statement about internal ranking systems.
Questions people ask
What is on page SEO?
Everything on the page itself that affects how it is found and understood: the title, the meta description, the headings, the body copy, the images and their alt text, the URL, the internal links and the structured data. It is the half of SEO you fully control, as distinct from links and other signals earned off the page.
What are the most important on page SEO factors in 2026?
The classic list has not changed much. What has changed is where the failures are. On the professional sites we measured, titles and descriptions were almost always fine, and the things that failed were heading order, whether any passage on the page could be lifted as an answer, and whether an author and a date existed in the markup at all.
What is the difference between on page and technical SEO?
On page work is about the content and markup of a single page. Technical work is about whether that page can be reached, rendered and indexed at all, which is crawling, robots rules, canonicals, status codes, rendering and site speed. They overlap in the head of the document, and the practical order is technical first, because the best on page work on a page nothing can fetch is worth nothing.
How many words should a page be for on page SEO?
There is no target and never has been. Length is a consequence of covering the question, not an input. What our own scoring actually rewards is structure: enough sections bounded by headings that a model can chunk on, and a short direct answer under each one. A tight page with five well bounded sections outperforms a long one written as unbroken prose.
Does on page SEO still work if AI answers the query?
It works differently. When an assistant answers in place, the value moves from the click to the citation, and being citable is largely an on page property: a bounded passage that answers in its first line, carries its own figures with units, and does not depend on the paragraph above it to make sense. That is on page work that no on page checklist was listing five years ago.
What should I fix first?
Whatever is broken rather than whatever is first on a list. Run the page, look at what actually fails, and start with the items that are cheap and structural: heading levels that skip, sections that open with a preamble instead of an answer, and a missing author or date in the markup. Those are edits, not projects.
Related guides
- Google's SEO starter guide is right, and it is answering one questionGoogle publishes a list of things it says not to focus on, and every item is correct about ranking. Three of them still decide whether you get quoted.
- SEO tips ranked by how often the thing is actually brokenMost SEO tips are opinions in no order. These are ordered by measured failure rates across our own audits of professional marketing sites.
- How to humanize AI content, and why detector scores are the wrong targetSeven detectors flagged 61% of human written essays as AI. Rewriting for a detector score is a trap. Here is the edit that stops a ChatGPT draft reading like one.