How Search Engines Work: Crawling, Indexing and Ranking

Every search engine, from the largest to the smallest, performs the same three jobs in the same order. Understand them and the rest of search stops feeling like magic.

A search engine looks like a box that reads your mind. Under the surface it is a pipeline: find web pages, store what they contain, then decide which stored pages best answer a query. Crawling, indexing and ranking are the three stages of that pipeline, and they run continuously — long before you ever type a word.

What a search engine actually is

There is no single program called "the search engine". There are dozens of cooperating systems: crawlers that fetch pages, parsers that clean up the HTML they receive, indexers that turn text into lookup structures, retrieval services that find candidates for a query, ranking models that score those candidates, and the front end that draws the results page.

The scale explains most of the design. A general engine may know about hundreds of billions of URLs and answer tens of thousands of queries every second, each in well under a second. Nothing can be recomputed on the fly at that volume, so the expensive work happens ahead of time: crawling and indexing are preparation, and a query only pays for retrieval and scoring.

Stage 1 — Crawling: finding pages

Crawling starts from a list of seed URLs and a queue of links waiting to be visited. A crawler takes the next URL, fetches it, and extracts every link it finds. Those links become new queue entries. Follow enough links and the graph of the web emerges, one edge at a time.

Three rules keep this from becoming a denial-of-service attack on the web:

  • Politeness. Crawlers limit how often they hit a host, so a small site is never flattened by a bot.
  • robots.txt. A plain text file on the host says which areas may be fetched and by whom. It is a request, respected by reputable crawlers and ignored by bad actors.
  • Sitemaps. An XML list of important URLs, which lets crawlers skip the guesswork — particularly useful for pages that few other sites link to.

Crawling also involves decisions that sound bureaucratic but matter enormously: which of several near-identical URLs is the canonical one, whether a page has already been seen under another address, and how much crawl budget a large site deserves. Pages rendered with JavaScript need a second pass through a rendering service before their content can be read properly.

This stage is covered in more depth in Web Crawlers: How Search Engines Discover Pages.

Stage 2 — Indexing: turning pages into data

Fetching a page is the easy part. Indexing is where a raw document becomes something a machine can search. The engine parses the HTML, strips navigation and boilerplate, and keeps the parts that carry meaning: the visible text, the title, headings, links, image alt text, and structured data such as JSON-LD.

That text is then broken into tokens — roughly, words — normalised so that running, runs and ran can be related, and stored in an inverted index: a mapping from each term to the list of documents containing it. That structure is what makes a query fast. Instead of reading billions of pages, the engine looks up a handful of terms and intersects the lists.

The engine also keeps a compressed copy of the original page for building snippets, and records signals alongside the content: when the page was published, how it is structured, which pages link to it, and whether it is a duplicate of something already indexed. A page can be crawled and still never be indexed if it is a duplicate, blocked by a noindex directive, or judged too thin to be useful.

More on this in The Search Index: How Engines Store the Entire Web.

Stage 3 — Ranking: choosing an order

Ranking answers a deceptively simple question: of the thousands of pages that mention your words, which ones should occupy the top ten? The engine retrieves a candidate set — usually far more documents than it will ever show — and scores each one.

Historically, relevance was measured almost purely statistically: how often your terms appear in a document, weighted by how rare those terms are across the collection. That measure, best known as TF-IDF and its successor BM25, is still the first filter. Modern engines layer several more signal families on top:

  • Links and authority. A link is a vote, and votes from trusted pages count for more than votes from obscure ones.
  • Quality and helpfulness. Whether the page demonstrates first-hand experience, expertise and care — and whether real users seem satisfied by it.
  • Freshness. Vital for news and prices, nearly irrelevant for a recipe.
  • Experience and speed. Usability signals such as mobile friendliness and page load behaviour.
  • Context. Language, country and — for some engines and some queries — personalisation.

Those signals are combined by machine-learned models that are trained on relevance judgements and query logs, then reordered into the final list. The order can differ between engines for the same query, and can shift over time as content changes. That is why no fixed list of "ranking factors" exists, only families of signals with different weights.

The full picture is in Ranking Algorithms: How Search Engines Decide What Comes First.

One query, end to end

  1. You type a few words. The query is normalised, corrected, expanded with synonyms, and classified by likely intent.
  2. The engine looks up the query's terms in the inverted index and pulls together a candidate set of matching documents.
  3. Each candidate is scored by the ranking models, using stored signals — no live crawling of the wider web happens at this point.
  4. Results are assembled into a page with titles, snippets, sitelinks and any extra features the query deserves, then delivered back to your browser.

The whole sequence usually takes a few hundred milliseconds. Everything expensive — discovery, parsing, storage, signal gathering — was done days or minutes earlier.

What it means for site owners

If the pipeline is crawl → index → rank, then a page can fail at any of the three stages and the failure looks identical from the outside: nothing is visible in search. The diagnostic questions follow the pipeline in order.

  • Crawlable? Can robots reach it, can crawlers follow a path of links to it, does the server answer quickly?
  • Indexable? Is it unique, substantial, and free of blocking directives or accidental duplicates?
  • Rankable? Does it answer a real query better than the pages already there, and does it earn links or engagement?

Those three questions are the skeleton of SEO Basics, and they are worth internalising before touching any tactic.

Misconceptions worth dropping

  • "The engine reads my page like a human would." It extracts features and models them. Meaning is approximated, sometimes well, sometimes badly.
  • "There is a fixed list of ranking factors." Signals overlap, interact and change. Chasing a checklist is not the same as building something worth ranking.
  • "Paying for ads raises organic rank." It does not. Ads and organic results are ordered by different systems.
  • "Once indexed, always indexed." Index entries are refreshed, demoted or dropped as pages and the web around them change.

Published 28 September 2026 · Last reviewed 28 September 2026