Query Processing: How Search Engines Understand What You Typed

You type three clumsy words. The engine spends a few milliseconds deciding that you meant something more precise, then searches for that instead.

Queries are short, misspelled, ambiguous and often unfinished. "cheap flights rome march" is not a grammatical request; it is a compressed intent. Query processing is the pipeline that decompresses it into something the index can answer.

Before you press Enter

Autocomplete runs while you type. Each keystroke triggers a lookup of popular completions, filtered for policy problems and ranked by how often people searching from your region chose them. Suggestions are a feedback loop, not a dictionary: they reflect what previous searchers typed, which is why suggestions differ by country, by time and occasionally by language.

Autocomplete influences behaviour significantly. People accept a suggestion rather than finishing a thought, so popular queries grow more popular over time — a small structural bias baked into the interface itself.

Normalising the query

  1. Tokenisation splits the query into terms, handling punctuation, apostrophes and script-specific rules (Chinese, Japanese and Thai do not use spaces, so segmentation is a real problem).
  2. Case and accent folding makes Nairo match nairó.
  3. Spelling correction compares your terms against a language model of common queries and offers or silently applies a correction. Silent correction is why "recieve" still finds pages about receive.
  4. Stemming or lemmatisation reduces words to a base form so indexing matches index.
  5. Stop-word handling marks words such as the or near as low value for scoring — but not for phrases, and not when they change meaning. "The Who" is a band, not a question about a band.

Expanding the query

Matching literal strings is brittle, so engines quietly add related terms. A query for laptop battery life should also find pages about notebook runtime. Expansion happens in three ways:

  • Stem families — morphological variants of the words you typed.
  • Synonym sets and knowledge graphs — curated equivalences and entity relationships, including abbreviations and product names.
  • Embeddings — vector representations of meaning, so documents that share no vocabulary but express the same thing can still be retrieved.

Expansion is a trade-off. Too little and relevant pages are missed; too much and the results drift. The ranking stage then pulls the drift back by favouring documents that also match your original terms.

Classifying intent

Engines group queries by what the searcher probably wants, because different intents need different results pages.

  • Informational — "how does indexing work". Wants explanation, guides, sometimes a direct answer.
  • Navigational — "niguro search". Wants one specific destination, not a comparison.
  • Commercial investigation — "best search engines 2026". Wants reviews, comparisons and shortlists.
  • Transactional — "download browser". Wants to act.
  • Local — "coffee near me". Wants places, hours and a map.

Classification is statistical, based on the words, the pattern of previous searches and how users behaved on past results. It shapes the layout more than the ranking: the same ten documents would be wasted if the query wanted a map.

Natural language and entities

Modern engines parse questions rather than matching keywords. They identify entities — people, places, products, dates — and the relationships between them, so "who wrote the paper about ranking that inspired the search engine founded in 1998" can be resolved step by step. This is the difference between keyword matching and question answering, and it is the part of search that language models improved most in the last decade.

Choosing which results page to build

Once the engine knows the intent, it assembles more than ten blue links. Depending on the query it may add a direct answer panel, image or video results, shopping listings, a map, a "people also ask" block, news boxes, or clearly labelled ads. Each feature is triggered by the intent classification, and each one can remove a click from the page entirely.

Zero-click results and why they happen

When the engine can answer a query directly — a date, a definition, a conversion, a phone number — it does so above the links. That is convenient for the searcher and awkward for publishers, because those queries never reach a website. The practical response is not to fight it but to target queries that genuinely require reading, judgement or a purchase, and to make pages that would still be worth clicking even if the headline fact were already visible above them.

Query processing decides what to look for. The next guide covers how the candidates are ordered: Ranking Algorithms.

Published 31 August 2026 · Last reviewed 31 August 2026