Submitted is not indexed, and indexed is not ranking
The most expensive assumption in SEO is that a page you published is a page Google has. Between 'it exists' and 'people find it' there's a funnel with several places to leak — and you have to measure each one, not the top.
By Andrew Pyle
The costliest mental shortcut in SEO is a quiet syllogism that feels obviously true: I published the page, therefore Google has the page, therefore people can find the page. Every step in that chain can fail independently, and when your traffic is inexplicably flat despite a growing site, it's almost always because one of those steps is failing silently while you assume the whole chain holds.
Between a page existing and a person finding it, there's a funnel: it has to exist, be submitted, be crawled, be indexed, and finally earn impressions for real queries. Each stage is a separate thing that can leak, and the only way to know your true state is to measure each stage on its own — because the top of the funnel tells you almost nothing about the bottom.
01Each stage leaks
Every stage can fail on its own
Walk the funnel and notice how each step is a distinct gate. A page can exist in your database but not be reachable by a crawler. It can be reachable but not submitted in your sitemap, so the engine never learns it's there. It can be submitted but never crawled, because crawlers budget their attention. It can be crawled but not indexed, because the engine decided it wasn't worth keeping. And it can be indexed but never surface for anything, because it's not competitive for any real query. Five gates, five independent ways to leak.
The reason this matters is that the failures are invisible if you only look at the top. Your publishing pipeline reports success — the page exists, the sitemap updated — and everything downstream could be quietly broken while your dashboard stays green. The page count goes up and to the right; the indexed count flatlines; and unless you're measuring the gap between those two numbers, you won't see the leak until the missing traffic forces you to.
Publishing reports success at the top of the funnel while every stage below it fails silently. The green dashboard is measuring the wrong end.
02Five numbers
The funnel is five real numbers, not a metaphor
In my own stack the funnel isn't an abstraction, it's five numbers I can pull and line up next to each other. How many posts exist in the database. How many actually render on the live site. How many appear in the sitemap the build generates. How many Search Console says it has submitted. And how many earn even one impression for a real query. Stack those five, and the leak stops being a mystery — it's whichever number drops off a cliff from the one above it.
The third number is the one people misread, because they assume the sitemap lists everything. Mine is generated from the source of truth and deliberately excludes noindexed and redirected URLs, so the sitemap count is already smaller than the raw page count, on purpose. A page that vanished between the database and the sitemap didn't leak; it was excluded by design. Separating the deliberate exclusions from the real leaks is half the work of reading the funnel honestly — a drop that looks alarming is often a rule you wrote doing exactly its job.
For the stages Google owns, I don't inspect every URL — I sample. A stratified sample run through the URL Inspection results, snapshotted so I can compare this week against last, tells me the indexed rate without tens of thousands of API calls. The snapshot is the point: a single reading is a number, but two readings are a trend, and a trend is the only thing that tells me whether a fix actually moved the indexed rate or just moved my mood. One measurement flatters you; the second one keeps you honest.
03Submitted vs indexed
The gap that hurts most: submitted versus indexed
The single most important comparison is between how many pages you submitted and how many the engine actually indexed. A large gap there is the tell that something in the middle of the funnel is broken — pages are being crawled and rejected, or not crawled at all, and either way the engine has looked at your site and declined to keep a big chunk of it. That gap is the number that predicts your traffic ceiling, and it's the one nobody looks at because the submitted number is so satisfying to watch grow.
When that gap is wide, the diagnosis is usually quality or crawlability, not volume. The engine crawled the pages and judged many of them not worth indexing — which is exactly the signal that you have thin or duplicative content, or that something structural is making the pages hard to crawl. Adding more pages when the indexed rate is already poor doesn't help; it just widens the gap. You have to fix why the existing pages aren't being kept before more of them will be.
04Measure, not command
You can measure it, but you mostly can't command it
There's a hard truth sitting in the middle of all this: you can measure the funnel far more easily than you can command it. The obvious wish is an API that says 'index this page now,' and Google does publish one — but the Indexing API officially only supports job postings, events, and live streams. For a portfolio of essays and project pages, calling it is not the intended path, and treating it as a force-index button is a way to lie to yourself about control you don't actually have.
So the tools I actually run are observers, not levers. A sitemap ping tells Google the map changed; the URL Inspection API tells me what Google currently thinks about a specific URL. Neither one makes a page get indexed — they make the state visible. That reframes the whole job: the lever isn't the submission, it's the page. If the engine crawled a URL and declined to keep it, re-submitting it a hundred times changes nothing; you change the page, and then the measurement tells you whether the engine changed its mind.
05Go look
Go look; don't trust the flag
The discipline the funnel forces is one I hold everywhere: don't trust the declared status, go observe the real one. "Is this page in Google?" has an easy wrong answer — my system says I published it — and a real answer that requires actually checking what the engine reports about that specific URL. Reconciling my side of the funnel against the engine's side, page by page or in stratified samples, is the only way to know where the leak actually is, as opposed to where I assume it is.
This is slower and less satisfying than watching a page-count go up, and it's the only measurement that tells the truth. The publishing pipeline's success is a claim; the engine's index is the fact; and the whole job is to keep checking the claim against the fact rather than letting the claim stand in for it. A status column that says 'done' is a convenience, not evidence — and indexing is the domain where believing the convenience costs you the most, because the gap between what you shipped and what got kept is invisible until traffic goes looking for pages that were never really there.
06Measure, then fix
Measure the funnel, fix the leak
The practical upshot is that you should track the funnel as a series of numbers, not a single one: how many pages exist, how many are submitted, how many are indexed, how many earn impressions. The shape of those numbers tells you exactly where your problem is. A leak between submitted and indexed is a quality or crawlability problem. A leak between indexed and impressions is a competitiveness problem. Each leak has a different fix, and you can't choose the right fix until you know which stage is actually failing.
The north star for a lot of what I do is getting more of what I publish actually indexed and found — and the only way to move that number is to stop assuming the funnel and start measuring it. The page you published is not the page Google has, and the page Google has is not the page people find. Treat those as three separate facts to be verified, and the mysterious flat traffic stops being mysterious: it's a leak at a specific stage, and now you can see which one.
07
Keep reading
This piece is part of my series on search, answer engines, and content quality. The anchor is AEO and GEO: optimizing for answer engines.
Related
writing
A sitemap should only list pages you actually want found
A sitemap is a set of recommendations you make to a search engine. Listing pages you've told it not to index, or that redirect elsewhere, is contradicting yourself — and a sitemap that contradicts itself teaches the crawler to trust you less.
writing
When four pages should have been one
I wrote four pages chasing the same topic from slightly different angles, and they spent a year competing with each other instead of ranking. The fix wasn't better pages. It was fewer of them.
writing
I let an AI pipeline write my blog. Google demoted the whole site.
A first-person autopsy of a self-inflicted Helpful-Content demotion — the junk I shipped, the code that shipped it, and the reversible way I dug out.
writing
Programmatic content is a quality gate, not a volume play
One template and a dataset can become ten thousand pages overnight — and that same leverage can quietly sink a whole domain. I learned it by getting it wrong. The craft isn't the volume; it's the quality gate, and it has to clear three surfaces now: SEO, AEO, and GEO.
writing
Naming a child is the least scalable thing there is
I built a hundred thousand pages for a decision that only ever comes down to a name or two, chosen slowly and with enormous care. NameBayBay taught me — the hard way, with a lot of pages Google quietly ignored — that the things that matter aren't made in bulk.