Catchy Clouds Icon
You Published 1,000 Posts and Google Indexed 200. Here Is What Happened to the Rest.
Technical SEO

You Published 1,000 Posts and Google Indexed 200. Here Is What Happened to the Rest.

An unindexed page cannot rank, cannot be found, and cannot be cited by an AI assistant. It is invisible, no matter how good it is. On sites with large content libraries, this is common: Google discovers pages it never crawls, crawls pages it never indexes, and quietly drops pages it once kept, usually because the site as a whole is asking for more crawling than its authority justifies or offering more pages than it offers value. Search Console tells you exactly which of those is happening to you, in a report most content teams have never opened.


First, find your real number

Open Google Search Console, go to Indexing, then Pages. You will see two totals: indexed pages, and pages not indexed. Compare the first number to how many pages you actually published.

The gap is your invisible library. Many content-heavy sites find that half or more of what they published is sitting outside the index, which means half the content budget produced nothing. Do not use the site: search operator for this; it is an approximation and it will mislead you. Search Console is Google reporting on your site directly.

Then click into "Why pages aren't indexed" and read the statuses, because each one names a different problem with a different fix.


Decoding what Search Console tells you


"Discovered, currently not indexed"

Meaning: Google knows the URL exists but has not crawled it yet. It found the link, put it in the queue, and never got there.

What it usually means in practice: your site is requesting more crawling than Google is willing to spend on it, or the pages look low-priority. This is the classic large-site symptom, and publishing more pages makes it worse rather than better, because the queue grows while the crawl allocation does not.


"Crawled, currently not indexed"

Meaning: Google visited the page, read it, and decided not to include it.

What it usually means: a quality or redundancy judgment. The page adds little that is not already available, overlaps heavily with your other pages, or reads as generic. This status is the single most useful signal in SEO, because it is Google telling you plainly that a page was not worth keeping, which is a verdict most content teams never learn about.


"Duplicate, Google chose a different canonical"

Meaning: Google decided another URL is the real version of this content, and it may be your own page or someone else's.

What it usually means: tag pages, category archives, parameter URLs, printer-friendly versions, or near-identical articles competing with each other. The fix is consolidation and correct canonical tags rather than new content.


"Excluded by noindex tag" or "Blocked by robots.txt"

Meaning: you told Google to stay away, whether you meant to or not.

What it usually means: a plugin setting, a staging configuration that shipped, or a template applying a rule far more widely than intended. Check this first, because it is the fastest possible fix and it is covered fully in our guide to why a website does not show up on Google.


Why publishing more can make this worse

Three mechanisms turn a big library into an invisible one:

  • Crawl allocation is finite. Google crawls each site at a rate its systems consider appropriate, based partly on how well the server responds and partly on how much value the site has demonstrated. Small sites rarely hit this ceiling. Sites with thousands of URLs do, and every low-value page in the queue is spending crawl capacity that a valuable page needed.
  • Quality is assessed site-wide, not only page by page. A library where most pages are thin teaches Google to expect thin pages from that domain, which lowers the enthusiasm for crawling and indexing the next one, including the genuinely good one you published this morning.
  • Internal link equity gets diluted. Attention and authority flow through your internal links. Spread across two thousand pages, each page receives a trickle. Pages that nothing links to, orphans buried in pagination, often never get crawled at all.

This is the mechanical version of a point we made in our myths article: more pages does not equal more traffic. Here is exactly why not.


The fix, in the order that works


1. Remove the blockers

Clear any accidental noindex tags and robots.txt disallows first. Nothing else matters while these are in place.


2. Cut the crawl waste

Identify what is consuming crawl capacity without earning anything: parameter URLs, endless faceted filters, thin tag and archive pages, duplicate category paths, and old paginated series. Consolidate them, noindex the ones that exist purely for navigation, and keep your sitemap listing only pages you genuinely want indexed. A sitemap full of URLs Google keeps rejecting is a credibility problem in itself.


3. Prune or merge the thin pages

Take everything sitting in "Crawled, currently not indexed" and be honest about each one. Most fall into three groups: merge into a stronger page and redirect, substantially improve, or remove and redirect to the closest relevant page. Publishers who consolidate a bloated library frequently find the remaining pages perform better afterwards, because the site as a whole finally looks like what it claims to be.

Deleting content feels wrong after paying to produce it. The money is already spent either way; the only question is whether those pages continue to hold back the rest of the library.


4. Rebuild internal linking toward what matters

Every page you want indexed should be reachable in a few clicks from your homepage and linked from related articles in context. Hub pages that gather a topic cluster do this naturally, and they are the cheapest indexing fix available, because they simultaneously tell Google what a page is about and give it a path to reach it. Our guide to why technical SEO comes first covers the structural side.


5. Then publish less, better

Once the library is clean, protect it. One substantial piece answering a real question outperforms five that restate what already ranks, and it does not add to the crawl burden that caused the problem. The keyword approach in our free keyword research guide is built for exactly this: one question, one page, answered properly.


The AI consequence nobody mentions

Indexing is now the gate to two channels rather than one. AI assistants find their sources by searching, then fetch and read the pages they select, which means an unindexed page is not merely unranked, it is ineligible to be cited in an AI answer at all. Every page sitting in "Crawled, currently not indexed" is invisible in both places simultaneously, as explained in where AI answers come from.

Which reframes an indexing audit as one of the highest-leverage projects available to any content-heavy business: the content is already written and already paid for. Getting it seen costs a fraction of producing it again.


Frequently asked questions


How do I know if crawl budget is my problem?

Crawl capacity rarely limits small sites. If you have a few hundred URLs, look at quality and duplication instead. If you have thousands, check the Crawl Stats report in Search Console: if Google is crawling far fewer pages per day than your library size, and "Discovered, currently not indexed" is large, allocation is your constraint.


Should I delete old blog posts?

Delete or merge the ones that are thin, redundant, and unindexed, and redirect them to the closest relevant page. Never delete pages that have backlinks or still earn traffic; improve those instead. When unsure, merge rather than remove, since merging preserves whatever value existed.


Does requesting indexing in Search Console help?

For a handful of important pages, yes, and it is worth doing after a fix. It is not a solution at scale, because it treats the symptom while the underlying quality or capacity problem keeps producing new cases.


How long does reindexing take after I fix things?

Individual pages can be reprocessed within days of a request. Site-wide changes in how Google treats your domain, particularly after a large consolidation, typically show over several weeks as crawling patterns adjust. Patience is part of the method, as with the timelines in our guide to how long SEO takes.


The bottom line

A content library is only as valuable as the portion Google actually keeps. Open the Pages report, compare indexed against published, and read the statuses honestly, because Google is already telling you which pages it judged not worth keeping. Clear the blockers, cut the waste, merge the thin material, link deliberately, and then publish less and better. If you would like that audit run properly on your own library, our free SEO and AEO audit reports exactly how much of your content is indexed, what is blocking the rest, and which pages are worth rescuing first.

Was this article helpful?
Search ranking growth illustration

Ready to Rank Higher and Get Cited in AI Answers?

Specialist SEO and AEO that builds lasting visibility

Scan Your Website for Free

No credit card, no signup to see your results

Icon
Shape