Skip to main content
USA-Based Digital Agency

Your Content Isn't the Problem. Your Index Coverage Is.

What five days of Search Console forensics across nine sites actually showed — and why “publish more” is usually the wrong prescription.

By George Shvaya · August 10, 2026 · 9 min read

Here is the finding, before the argument.

Across nine sites we audited this month, the sites that had stopped growing were not short on content. They were short on indexed content. One property had 43% of its submitted URLs indexed — 688 pages sitting in a sitemap, invisible. Another had 16 URL families indexed under more than one variant, and those duplicated families carried 85.1% of the site’s clicks. Its homepage existed twice in Google’s index — one variant at average position 3.0, the other at 15.9.

Figure 1

43% indexed. The other 688 pages were never seen.

Indexed
Eligible to rank or be cited
Submitted, never indexed
688 pages, invisible to search

One audited property: 43% of submitted URLs indexed, 57% — 688 pages — submitted and ignored. Source: Search Console Pages report, sitemap-filtered.

None of that is a content problem. No amount of publishing fixes it. And every one of those sites had been advised, at some point, to publish more.

That gap — between the pages you have and the pages a search engine can actually see, resolve, and attribute — is what we call index engineering. It is the least glamorous work in search, and in 2026 it is where most of the recoverable value sits.

The reflex that costs the most money

When traffic flattens, the default diagnosis is content. It is the easiest thing to sell, the easiest thing to brief, and the easiest thing to measure activity against. Fifty new pages feels like progress in a way that a canonical audit never will.

But content volume only works if the marginal page gets indexed, gets attributed to a query, and doesn’t cannibalize a page you already have. On four of the nine sites we looked at, all three of those assumptions were false at the same time.

What we foundWhat it actually means
43% sitemap indexationOver half the site is submitted and ignored. New pages join the ignored pile.
16 self-competing URL familiesThe site's authority is split across variants of itself.
Homepage at position 3.0 and 15.9Two versions of one page, one of them wasting every impression it earns.
14 URLs competing for one query clusterNobody wins. The intended page ranks worst.
67% of URLs never surfaced at allThe content exists. Search does not know it does.

Publishing into that is like adding water to a bucket you haven’t checked for holes.

Figure 2

85.1% of clicks went to pages competing with themselves.

Clean, single-variant URLs
Full value of each impression
Self-competing URL families
85.1% of clicks, split across variants

16 URL families were indexed under more than one variant. Those families carried 85.1% of the site's total clicks — performance split across duplicates of the same page.

Figure 3

One homepage. Two entries in the index. Twelve positions apart.

151015203.015.9

Average search position — lower is better

The same homepage existed twice in Google's index — one variant averaging position 3.0, the other 15.9. Every impression the second variant earned was wasted.

How to check yours in about thirty minutes

This is the part worth keeping. You can run most of this without a tool purchase.

  1. Compare submitted vs. indexed. In Search Console, Pages → open your sitemap-filtered view. Divide indexed by submitted. If you’re below roughly 70%, stop every content plan you have until you know why. That ratio is the single most diagnostic number on the property, and almost nobody looks at it.
  2. Group your Pages export by canonical path. Export Pages.csv, strip the protocol and the trailing slash, and group. Every group with more than one row is a page competing with itself. Sum the clicks in those groups — that’s the share of your performance that’s being split. On the site above it was 85.1%.
  3. Check the four homepage variants. http://, https://, with www, without. All four should terminate in one 301 to one canonical host. A 307, a redirect chain, or two live variants means you have been running two websites.
  4. Map queries to the page you intended to rank. For your top 20 queries, note the URL Google actually serves. Every mismatch is a page that is either miscanonicalized, mislinked, or redundant.
  5. Respect the export cap. This one separates careful analysis from confident nonsense. The GSC query export is capped at 1,000 rows. On one property those 1,000 rows accounted for 45% of total impressions — the rest sat in the anonymized long tail, invisible. Every query-level percentage you calculate is a share of the visible portion only. Page-level data doesn’t carry that limitation, which is exactly why the conclusions above are drawn at page level.

Work it out

What’s your indexation rate?

Search Console → Pages → filter to your sitemap. Take the two numbers and put them here. Nothing is sent anywhere; this runs in your browser.

Enter both numbers to see your rate and what it means.

If you take one thing from this article, take step 5. Most audits I read are built on the capped export and presented as if they were complete.

Figure 4

Your query export shows 45% of the picture.

Visible in the 1,000-row export
What most audits analyse
Anonymized long tail
Never appears in the export

The Search Console query export is capped at 1,000 rows. On one property those rows accounted for 45% of total impressions; the remainder sat in the anonymized long tail. Any query-level percentage describes the visible slice only.

The AI-search half nobody connects

There’s a reason this matters more now than it did two years ago, and it isn’t nostalgia for technical SEO.

An unindexed page cannot be cited. Generative engines — AI Overviews, ChatGPT search, Perplexity — retrieve from indexes. If your page isn’t in one, it does not exist to the model, no matter how well written it is. Every dollar spent on generative engine optimization sits on top of index coverage as a prerequisite, not beside it.

And the retrieval layer is less forgiving than the ranking layer used to be:

  • Duplicate entities confuse attribution. We found structured data on roughly 60 pages of one site naming a provider node that was never emitted — a dangling @id reference. A validator flags that instantly. A crawler just resolves nothing.
  • Unverifiable claims are a liability, not a differentiator. On one property, inherited copy claimed “over 20 years.” The business was founded in 2013. That wasn’t puffery; it was a factual error being fed to systems that increasingly cross-check.
  • Your AI-facing files drift silently. llms.txt — the file assistants read to describe your business — went stale twice on one site inside a month, quietly describing pages that no longer existed. It now fails a build check when it drifts.

The practice that came out of this week is unromantic: a claim registry. Every factual assertion on the site — years in business, certifications, capacity, ratings — recorded with its source and its verification date, enforced by a check that fails the build when unverified copy ships. It is bookkeeping. It is also the only thing that reliably stops a marketing site from lying by accident.

What we got wrong, since it’s instructive

Two failures from the same week, both mine, both worth publishing.

A guard tuned too tight blocks true copy. We gated the phrase “ASE-certified” as unverified. The owner then verified it. The check would have blocked accurate copy — and a check that punishes truth teaches a team to write around it, which produces worse writing than no check at all. Guards need an unblocking path, not just a blocking one.

Automated claim removal is not a find-and-replace. A mechanical token swap across a content corpus produced sentences like “cared for by genuinely good at this work.” Ten of them shipped before review. Prose needs sentence-level rewrites. There is no regex for meaning.

I’d rather publish those than another case study.

What to do Monday

  1. Calculate your indexation rate. Not your ranking, not your traffic — the ratio.
  2. Pause the content calendar if it’s under 70%. The pages aren’t the constraint.
  3. Resolve every duplicate URL family to one canonical, one host, one 301.
  4. Fix crawl and canonical hygiene before spending anything on GEO or AEO. Retrieval requires an index entry.
  5. Write down every factual claim on your site with its source. Then enforce it.

None of this produces a screenshot that looks good in a monthly report. All of it determines whether the next hundred pages are an asset or a liability.

Frequently asked questions

What is a good sitemap indexation rate?

Divide indexed URLs by submitted URLs in the Search Console Pages report. Below roughly 70% is the point at which adding content stops being the useful lever, because new pages join the ignored pile rather than the indexed one. Across the nine sites in this audit, one property sat at 43% — 688 submitted pages that were never indexed.

Why does the Google Search Console 1,000-row export cap matter?

The query export is capped at 1,000 rows. On one property in this audit those 1,000 rows accounted for only 45% of total impressions; the rest sat in the anonymized long tail. Any query-level percentage calculated from that export describes the visible portion only, not the site. Page-level data does not carry the same limitation, which is why conclusions should be drawn at page level.

Does index coverage affect AI search visibility?

Yes, and it is a prerequisite rather than a parallel task. Generative engines including AI Overviews, ChatGPT search and Perplexity retrieve from indexes. An unindexed page cannot be cited, however well written it is, so spending on generative engine optimization sits on top of index coverage rather than beside it.

How do I know if my pages are competing with each other?

Export the Pages report, strip the protocol and trailing slash, and group by canonical path. Any group with more than one row is a page competing with itself. Summing the clicks in those groups gives the share of performance being split — on one site in this audit that was 85.1%.

About the work

Webvello is an AI search optimization and index engineering company founded in 2024, based in Roseville, California. We’ve delivered 50+ projects across 37+ cities, with traffic growth typically in the 3–5x range on properties where the technical foundation was the binding constraint. Founder George Shvaya came to search from cybersecurity and systems engineering — which is roughly the reason this article is about index coverage and verification layers rather than content calendars.

Methodology: nine client properties audited over one month using Google Search Console page-level and query-level exports. Figures are reported anonymously and are drawn at page level except where stated. The 1,000-row query export cap applies to query-level figures and is called out where relevant.

Want to know your own indexation rate before you approve another content budget?

Book a call
Get Free Growth Plan