Ranking Algorithms: How Search Engines Decide What Comes First

Two pages can mention exactly the same words and still be treated very differently. Ranking is where a search engine stops being a lookup table and starts making a judgement.

Ranking is a sorting problem with a philosophical edge: the engine must guess, from a handful of typed words, which of thousands of candidate pages a stranger would find most useful. Every engine solves it by combining signals into a score. The families of signals are stable; their weights and the models that combine them are not.

Retrieval first, ranking second

The index rarely gives you a short list. A single common word can match millions of documents, so the engine first retrieves a manageable candidate set using fast, cheap matching, then applies expensive ranking models to the survivors. This two-phase design is why relevance and importance are separable: one decides who gets considered, the other decides who wins.

Relevance: matching the words

The oldest and still most important family of signals asks how well a document matches the query text. Two ideas do most of the work:

  • Term frequency — a page that uses your words more often is usually more on-topic, but with diminishing returns, so a page repeating a term fifty times does not outrank a page using it five times by a factor of ten.
  • Inverse document frequency — rare words carry more meaning. Matching BM25 algorithm is far stronger evidence than matching the.

Refinements layer on top: which fields the term appears in (title beats body), how close the terms are to each other, whether they appear as an exact phrase, and whether the page uses obvious synonyms. The combined measure, commonly implemented as BM25 or a descendant, is the backbone of first-pass ranking in most engines, including open-source ones.

Text matching alone cannot tell a careful explanation from a careless one, so engines use the structure of the web itself. A hyperlink is a recommendation: someone chose to point readers here. Counting links naively fails immediately — links are easy to buy — so engines weight them by the authority of the linking page, and discount the ones that look manufactured.

The classic formalisation is PageRank, which treats the web as a graph and asks: if a reader wandered from page to page at random, how much of the time would they spend here? Pages linked from many respected pages score highly; links from link farms count for very little. Anchor text — the clickable words — also carries meaning and can describe a page better than the page describes itself.

Quality, experience and helpfulness

Since around 2010 the industry has moved from "does this page mention the query" towards "does this page satisfy the person behind the query". The signals here are softer and partly inferred:

  • Demonstrated expertise. Author credentials, citations, and whether the site is known for the subject.
  • First-hand experience. Reviews, tests and photographs that only come from actually doing the thing.
  • Trustworthiness. Clear ownership, contact information, and content that is transparent about its sources.
  • Page experience. Usability signals such as mobile friendliness, intrusive interstitials, and core web vitals measuring loading, interactivity and visual stability.
  • Engagement. Query logs suggest when users abandon a result quickly, or refine their query, both of which feed back into future ranking.

None of these signals is a switch. They move a page up or down within a band of comparably relevant results.

Freshness and time

Time matters differently depending on the query. For breaking news, a product price or a sports score, recency is a hard requirement. For a definition or a recipe, an older page is often the better one. Engines classify queries by how time-sensitive they appear to be, then weight document age and change history accordingly.

Machine-learned ranking

Somewhere around a hundred signals are combined by trained models rather than hand-written rules. The standard technique is learning to rank: engineers describe each candidate with features, collect human relevance judgements and click data, then train a model to reproduce the ordering that satisfies users most often.

Modern engines also score semantically, using language models that turn text into vectors and compare meaning rather than exact words. That is how a query like "why is my phone battery draining" can match a page titled "common causes of rapid battery loss" — no shared vocabulary, same intent. Neural re-rankers run only on the top slice of candidates because they are expensive.

Context, personalisation and local

The same query in two places can reasonably return different results. Engines use language, country, device type, and — depending on the product and the user's settings — location, search history or previous interactions. Some engines keep this minimal by design; others lean on it heavily. Where personalisation is applied, it usually reorders a broadly similar set rather than inventing a new one.

Why there is no fixed factor list

Search engines publish principles, not weights. Signals interact and can substitute for each other: a page with exceptional content can rank without strong links, and a page with excellent links will not rank for a query it does not answer. Models are retrained and updated constantly, and the same page can rank differently for the same query in two engines, or in one engine in two languages.

The practical conclusion is not to reverse-engineer the algorithm but to satisfy the intent behind the query better than the pages already there — which is much of what SEO Basics is about. When ranking goes wrong, the corrections fall to a different set of systems entirely: see Spam, Quality and Trust.

Published 7 September 2026 · Last reviewed 7 September 2026