How search engines crawl and index pages (and why your best page sometimes never shows up)
You publish a page. You check it in an incognito window. Nothing. Not on page 8, not on page 30. Nowhere. You assume the page is bad. Usually it's not bad at all — it simply never made it through the two filters every search engine runs before ranking anything: crawling and indexing. Two different jobs, done by two different systems, and most site owners treat them as one.
I learned this the hard way on a client project where I moved an ecommerce category to a new URL and "fixed" it with a redirect. Search Console showed the new page as discovered. It stayed that way for eleven days. Not crawled, not indexed, invisible. The reason had nothing to do with content quality. It was a crawl queue decision I didn't understand yet.
Key Takeaways
- Crawling is discovery and reading. Indexing is storing and organising. A page can be crawled without ever being indexed.
- Googlebot decides what to crawl using a queue, a crawl budget, and signals like XML sitemaps and internal links.
noindexandrobots.txtdo opposite things: one blocks indexing, the other blocks crawling. Confusing them is the most common technical mistake I see.- A
noindexrule only works if the crawler can actually reach the page. Block it in robots.txt and the rule is never read. - Modern search engines render JavaScript before extracting content. If rendering fails, indexing fails.
- Duplicate content, thin pages, and wrong canonical tags are the usual reasons a crawled page never gets indexed.
What crawling in a search engine actually is
Crawling is an automated bot fetching a URL, reading the raw response, and extracting every link it finds. That's it. No ranking happens at this stage. No judgement of quality either. A crawler is closer to a mail carrier than an editor: it delivers the page to the next step and moves on.
The part almost nobody explains is how the bot chooses which URL to fetch next. It doesn't wander randomly. Every major engine maintains a URL frontier — a queue of known addresses waiting their turn — and assigns each one a priority and a revisit frequency.
What determines how often a page gets revisited
Pages that change often get crawled more often. Pages that never change drift to the back of the queue. I watched this happen on a blog where I published twice weekly for two years: new posts were picked up within hours, while an old "About" page took weeks to reflect a single edit. Same site, same domain authority, wildly different crawl rates.
- How frequently the content genuinely changes
- How many internal links point to the page
- Whether the URL appears in your XML sitemap
- How deep it sits in the site architecture (three clicks from the homepage is a very different situation from nine)
- Server response speed — slow pages get fewer requests, not more patience
The internal linking point is the one people underestimate. On a site with 40,000 pages and a shallow category structure, I cut average crawl latency from roughly three weeks to under four days just by linking new content from high-traffic pages instead of leaving it buried in an archive. No content changes. No new backlinks. Pure link architecture.
What do search engines use to index web pages?
Once a page is fetched, it enters the indexing pipeline — and the first stop is rendering. Modern engines don't index raw HTML anymore. They load the page, execute its JavaScript, build the final DOM, and index what a real visitor would see. If your content only appears after a client-side API call that the renderer can't complete, the engine indexes an empty shell.
From there, the system analyses the rendered page: text, headings, structured data, images and their alt attributes, internal and external links, and the entities the content is actually about. Then it decides whether the page deserves a slot in the index at all.
The indexing decision is not automatic
Being crawled does not guarantee being indexed. Roughly a third of the pages I audit on mid-sized sites are crawled regularly and still excluded. The usual culprits:
- Duplicate content — the same product description on 200 variants, with the canonical tag pointing at the wrong one
- Thin pages — tag archives, paginated listings, and auto-generated category pages with almost no unique text
- Wrong canonical signals — a page telling the engine "the real version of me is over there"
- Rendering failures — content that never materialises for the crawler
Here's the thing: crawl budget isn't really about how many pages you have. It's about how many of them are worth spending requests on. A site with 5,000 genuinely distinct pages will get indexed far more completely than one with 50,000 near-identical ones.
What is the difference between crawling and indexing in search engines?
Crawling is discovering and reading. Indexing is storing, parsing, and organising what was read so it can be retrieved later. Crawling is a live action against your server. Indexing is a state — your page either sits in the database or it doesn't.
That distinction matters because the fixes are completely different. A page that isn't crawled has a discovery or access problem: broken links, a robots.txt block, an orphaned URL. A page that's crawled but not indexed has a quality, duplication, or directive problem. Pulling the wrong lever wastes months.
| Aspect | Crawling | Indexing |
|---|---|---|
| What happens | Bot fetches the URL and extracts links | Content is parsed, stored, and associated with entities |
| Blocked by | robots.txt, server errors, firewall rules | noindex, canonical tags, quality filters |
| How you check it | Crawl stats, server logs | Coverage report, site: query |
| Visible symptom | Page never appears as "discovered" | Page shows "Crawled – currently not indexed" |
| Typical fix | Internal links, sitemap, faster responses | Unique content, correct canonical, remove noindex |
Does Google crawl noindex pages?
Yes. And this trips up almost everyone. noindex is a rule set with either a meta tag or an HTTP response header, used to prevent indexing content by search engines that support the rule, such as Google. When Googlebot crawls that page and extracts the tag or header, Google will drop that page entirely from Google Search results, regardless of whether other sites link to it.
Read that again: the bot has to crawl the page for the rule to take effect. The crawler cannot obey an instruction it never sees. So a noindex page is crawled, the directive is read, and only then is the page removed from the index.
The robots.txt trap that keeps pages indexed
For the rule to be effective, the page or resource must not be blocked by a robots.txt file and has to be otherwise accessible to the crawler. If the page is blocked by robots.txt, or the crawler can't reach it, the crawler will never see the noindex rule — and the page can still appear in search results, for example if other pages link to it.
I've seen this exact combination on a staging environment that leaked into production: Disallow: / in robots.txt plus a noindex tag, applied together for "extra safety". The result was the worst possible outcome. Google couldn't crawl the pages, so it never read the noindex directive, and a handful of staging URLs stayed in the index for months because other sites linked to them. Removing the robots.txt block was what finally let the noindex do its job.
Using noindex is useful when you don't have root access to your server, since it lets you control access on a page-by-page basis. The two implementation methods — meta tag and HTTP response header — have the same effect, so pick whichever fits your setup. Just don't stack them with a crawl block and expect the page to disappear.
How to tell which stage is failing
Stop guessing. Open the coverage report and read the status label, because it tells you exactly where the pipeline broke.
- Discovered – currently not indexed: the URL is known but the crawl queue hasn't reached it. This is a crawl-side problem. Add internal links, check server speed, verify the sitemap.
- Crawled – currently not indexed: the bot paid a visit and walked away unimpressed. Indexing-side problem. Look at duplication, thin content, canonical conflicts.
- Excluded by noindex tag: working as intended. If it shouldn't be excluded, remove the directive.
- Blocked by robots.txt: crawl-side. The bot isn't even trying.
One more tool worth knowing: running a site: query for a specific URL is a quick sanity check on whether a page sits in the index at all. It's imprecise and it's not a ranking signal, but as a yes/no answer it beats opening five reports.
What I've come to believe after years of this: indexing problems are almost never mysterious. They're boring. A canonical tag pointing the wrong way. A sitemap nobody updated. A render that times out on mobile connections. The glamorous explanations — "Google penalised me" — are rare. The dull structural ones are everywhere, and they're all fixable if you know which of the two stages you're actually fighting. Next time a page vanishes, don't rewrite the copy. Find out first whether anyone ever read it.