Web Crawlers: How Search Engines Discover Pages
Before a search engine can rank a page, it has to know the page exists. That job belongs to crawlers — the tireless robots that walk the link graph of the web.
A crawler — also called a spider or a bot — is a program with one loop: fetch a URL, parse the response, extract the links, add them to a queue, repeat. Google's crawler is called Googlebot; Bing's is Bingbot; every serious engine has its own. What separates a good crawler from a bad one is not the loop but the judgement applied inside it: which URL to fetch next, how often, and how politely.
What a crawler actually does
- Fetch. Send an HTTP request, usually with a conditional header so an unchanged page returns a cheap "not modified" response instead of the full body.
- Parse. Read the HTML, separate content from navigation and adverts, and collect links, text, metadata and structured data.
- Queue. Add newly found URLs to a frontier — a prioritised backlog. Priorities depend on how popular the URL's host is, how often the page seems to change, and how valuable the content appears to be.
- Deduplicate. Skip URLs already seen, and group near-identical pages so the index does not fill with copies.
- Hand off. Pass the processed document to the indexing pipeline, which decides whether and how to store it.
Five ways crawlers discover URLs
- Links from known pages. The classic route. Every hyperlink on a crawled page is a candidate.
- XML sitemaps. A file listing URLs you want crawled, with optional last-modified dates. The most reliable way to surface pages that few sites link to.
- Historical knowledge. An engine remembers URLs it has seen before and revisits them to check for changes, even without a new link.
- Manual submission. Search consoles accept individual URLs and sitemaps for faster discovery.
- Feeds and third-party sources. RSS, news feeds and other cooperative sources point at fresh content.
Note the implication: a page nothing links to and no sitemap mentions is discovered far more slowly, if at all. Internal links from your own site count. Orphan pages are the most common cause of "why is my page not showing up?"
Robots.txt, meta tags and other rules
Crawlers obey a small set of instructions, each answering a different question.
- robots.txt — a file at
/robots.txt. It controls fetching: which paths a given crawler may request. It is advisory, and it is public, so it is a tool for managing crawl load rather than a security mechanism. - Meta robots — a tag inside the page, such as
<meta name="robots" content="noindex">. It controls indexing: the page may be crawled, but not stored. - X-Robots-Tag — the same directives sent as an HTTP header. Useful for non-HTML files such as PDFs.
- rel="canonical" — nominates the preferred URL when several addresses serve the same content.
- rel="nofollow" — tells crawlers not to treat a link as an endorsement, and historically not to follow it at all.
A common mistake is disallowing a URL in robots.txt and adding noindex to it. If a crawler is forbidden from fetching the page, it never sees the noindex instruction, and the URL can still appear in results without a reliable snippet. Pick one technique and use it deliberately.
Crawl budget and crawl rate
Large sites have more URLs than any crawler cares to visit often, so engines allocate attention. That allocation is loosely called crawl budget, and it has two halves: how many URLs the engine is willing to crawl (crawl capacity limit) and how many it wants to (crawl demand, driven by popularity and freshness).
Signals that tell an engine your site deserves more attention:
- Fast server responses and few error responses.
- Fresh, genuinely changing content in the sections you care about.
- Inbound links and traffic to specific pages.
- Clean, unique URLs without endless near-duplicate parameter combinations.
What burns crawl budget: soft 404s, redirect chains, infinite calendar or faceted-navigation spaces, and thousands of thin pages that never change. If those consume a crawler's visit, the pages you care about wait.
Rendering JavaScript pages
Most crawlers first read raw HTML. If the content only appears after client-side JavaScript runs, many engines perform a second, more expensive rendering pass, sometimes delayed by hours or days. Server-side rendering, static HTML or pre-rendering removes that dependency and is almost always faster to be indexed.
How to verify a crawler
Anyone can put "Googlebot" in a User-Agent string, so look for proof. Reputable crawlers verify themselves by reverse DNS: the requesting IP resolves to a hostname in the engine's own domain, and that hostname resolves back to the same IP. Google and Bing both publish the IP ranges they crawl from. If a request claims to be a crawler and fails the lookup, treat it as a scraper.
Practical advice
- Serve a sitemap, reference it from robots.txt, and keep it accurate.
- Keep important pages within a few clicks of the homepage, linked with plain
<a>tags. - Return real status codes — 200, 404, 301 — instead of friendly-looking error pages.
- Do not block CSS and JavaScript files: crawlers use them to understand layout and rendering.
- Watch crawl statistics in a search console, and fix the errors it reports first.
Once a page is crawled and stored, the next stage begins. Continue with The Search Index: How Engines Store the Entire Web.
Published 21 September 2026 · Last reviewed 21 September 2026