The Search Index: How Engines Store the Entire Web
A crawler fetches a page. An index is what makes that page findable ten seconds later, for a query nobody has typed yet.
Indexing is the step that turns the open web into a database. A crawler can only hold the pages it has just fetched; the index holds everything the engine has ever decided was worth keeping, and it must do so in a form that answers a query in milliseconds rather than hours.
From page to tokens
The first job is turning a document into numbers a machine can compare. The engine parses the HTML, drops navigation, cookie banners and adverts, and keeps the text that carries meaning. Titles, headings, bold text, anchor text, image alt attributes and structured data are treated as more informative than body prose, so they get extra weight.
That text is then normalised:
- Tokenisation splits a string into terms. "Mugurdy's search-box" becomes something like
mugurdy,s,search,box. - Case folding makes Search and search the same term.
- Stemming reduces inflected words to a common root, so crawling, crawls and crawled all match crawl. Lemmatisation does the same job with grammatical rules.
- Stop-word handling decides what to do with high-frequency words such as the and of. Most engines keep them in phrases but ignore them when scoring single terms.
Normalisation is lossy on purpose. It trades a little precision for a much better chance of matching how people actually type.
The inverted index
Suppose you want to find every document containing the word ranking. A naive approach scans every page. That is hopeless at web scale, so search engines invert the mapping: instead of storing "document → words", they store "word → documents".
The result is a dictionary of terms, each pointing at a postings list: the documents containing that term, plus per-document details such as term frequency and where in the document the term appears. A query becomes an intersection of a few postings lists — cheap, because they are sorted, compressed, and often already cached in memory.
ranking → [doc17: tf 4, doc42: tf 2, doc91: tf 9, ...]
search → [doc17: tf 7, doc42: tf 5, doc88: tf 3, ...]
# "search ranking" = intersection of the two lists → doc17, doc42
Because both lists are ordered by document ID, the intersection is a linear merge rather than a search. Positional data inside each posting also makes phrase queries such as "search ranking factors" possible without a second pass over the text.
Signals stored next to the text
The index is not just words. Alongside each document the engine keeps the signals its ranking models need later:
- Document statistics such as length, language and the number of unique terms.
- Link data — which pages link here, with what anchor text, and how authoritative those sources are.
- Freshness: first-seen, last-modified and last-crawled dates.
- Quality and policy flags: whether the page is spam, a duplicate, malware-flagged, or paywalled.
- Canonical grouping, so duplicate URLs collapse into one representative document.
Collecting all of this before a query arrives is exactly why retrieval can be fast. The query only reads; it does not compute the world from scratch.
Shards, tiers and updates
A single machine cannot hold a web-scale index, so it is partitioned into shards and spread across many machines. A query fans out to the relevant shards, results are merged, and the top candidates are ranked. Engines also keep the index in tiers: highly popular, frequently queried documents live in fast memory, while long-tail pages sit in cheaper storage and are fetched less often. This keeps latency low without discarding the long tail.
Updating is a constant background process. Rebuilding the whole index every time a page changes is impossible, so engines merge small updates with large, mostly static segments, reordering and rewriting storage in the background — much like a log-structured database. From the outside this appears as the "freshness" of results: breaking news can enter the index in minutes, while a quiet page may wait days to be refreshed.
Why some pages are never indexed
- Blocked by a directive —
noindex, a robots.txt rule that prevents fetching, or an HTTP status that discourages storage. - Duplicate content — a near-copy of a page already indexed, or a URL variant that the engine treats as the same document.
- Thin or empty — pages with almost no unique text: tag archives, empty search results, placeholder pages.
- Low quality or spam — automatically detected and excluded or demoted.
- Never discovered — no links, no sitemap entry, nothing pointing at the URL at all.
Being crawled is therefore not the same as being indexed, and being indexed is not the same as being ranked. Each stage has its own gate.
Snippets need the original page
Results pages show a short excerpt around the matched terms. To produce that, engines store a compressed copy of each document's text alongside the postings lists, so a snippet can be assembled at query time without re-fetching the page. Cached copies — the "cached" link many engines still offer — come from the same storage layer.
That is the whole trick: fetch once, analyse thoroughly, store cleverly, and then answer every future query by reading rather than searching. Next, see how those stored candidates are ordered in Ranking Algorithms.
Published 14 September 2026 · Last reviewed 14 September 2026