Explain how search works
Walk any colleague through the four stages a page passes before it can appear: crawling, indexing, ranking, and serving.
Learning Hub · Start Here
A search engine serves as an automated retrieval system that returns web pages for a query (organic listings, map packs, featured snippets, and AI overviews)
A search engine serves as an automated retrieval system that returns web pages for a query (organic listings, map packs, featured snippets, and AI overviews). Search Engine Optimization (SEO) starts when a crawler discovers a URL, the index stores and interprets the page, and the ranking system scores it against each search. The engine serves the results from its stored copy of the web, ensuring answers in milliseconds rather than a live scan of every site. Opt for the URL Inspection tool in Google Search Console to confirm the indexed status if a page never appears in results. Visibility is influenced by crawlability, relevance, and trust, with eight beginner chapters, a 28-term glossary, and a knowledge check available below for readers on their first pass. Business owners and marketers utilize these basics to judge their own site, creating a shared vocabulary while briefing an SEO agency or running the work themselves.
Learning Outcomes
Four capabilities separate a reader who has studied SEO from a practitioner who can act on it. Each chapter below builds one of them, and the knowledge check at the end confirms all four.
Walk any colleague through the four stages a page passes before it can appear: crawling, indexing, ranking, and serving.
Tell organic listings, sponsored ads, map packs, featured snippets, and AI overviews apart at a glance — and know which one you can win.
Sort a keyword into informational, navigational, commercial, or transactional, then choose the page type that satisfies it.
Check whether a URL is indexed, find the queries it already earns impressions for, and rank the fixes worth doing first.
The Whole Model in One Diagram
Every query passes through the same four stages (crawl, index, rank, and serve). Almost every SEO problem traces back to a failure at one specific stage, so naming the stage delivers most of the fix.
STAGE 01
Bots discover the URL.
Googlebot follows links, reads XML sitemaps, and revisits pages it already knows to find what has changed.
STAGE 02
The page is stored and understood.
The engine renders the crawled page, splits it into words and entities, deduplicates it against near-copies, and files it in the index.
STAGE 03
Candidates are scored per query.
At the moment of the search, hundreds of signals score every indexed page that could answer the query, and an order falls out.
STAGE 04
A results page is assembled.
Ten blue links, a map pack, a snippet, an AI overview — the layout is built around what the query appears to need.
A search engine combines three systems: a crawler that discovers pages, an index that stores them, and a ranking system that orders the index against a query.
Google, Bing, and DuckDuckGo all run the same three-part structure. The size of the index, the frequency of the crawler's return, and the signals the ranking system trusts separate one engine from another. A reader who learns the structure once can read every engine, including the newer AI answer engines.
The engine searches nothing live. A typed query never sends the engine out to read the web; the engine reads its own copy of the web, assembled in advance. Every technique in search engine optimization exists either to get a page into that copy or to give the ranking system reasons to place the page near the top.
The stored copy also explains the lag beginners notice. A page published this morning might rank this afternoon or six weeks from now, and the timing depends entirely on when a crawler reaches the page and how readily the index accepts it.
Key concepts
A program such as Googlebot requests pages the way a browser does, then follows the links it finds to reach the next URL.
A distributed database of processed pages, broken into words, entities, and signals so a query can be matched in milliseconds.
Scoring happens at query time. The ranking system recalculates the order for every search, device, and location — no page holds a permanent position.
Takeaway
Search engines answer from a stored copy of the web, not the live web. Crawlability and indexability come before any attempt to rank.
Crawling covers discovery. A bot such as Googlebot fetches a URL, reads the HTML, extracts every link on the page, and queues those links to fetch next.
Links form the road network of a site. A page with no internal link pointing at it and no entry in a sitemap leaves the crawler without a route, however good the content is. Orphan pages quietly earn nothing for months for exactly that reason.
Two files shape the journey. An XML sitemap hands the crawler a list of URLs worth fetching, which matters most on large sites and brand-new domains. A robots.txt file does the opposite and tells compliant bots which paths to leave alone. A directory blocked in robots.txt ranks among the most common and expensive accidents: the pages inside stop being crawled, so their content never reaches the index.
Crawlers also budget their time. Google allocates each site a rough crawl capacity based on server response speed and the demand the site's pages generate. Slow responses, endless faceted URLs, and long redirect chains burn that crawl budget on pages nobody searches for.
Key concepts
The most reliable discovery path. Anything reachable in a few clicks from the homepage gets found early and re-crawled often.
A machine-readable list of the URLs a site wants crawled, submitted through Google Search Console and Bing Webmaster Tools.
A plain text file at the domain root that permits or blocks crawler access per path. The file controls crawling, never indexing.
The practical ceiling on how many URLs a bot fetches per visit. Fast servers and tidy URL structures stretch it further.
Takeaway
A URL the bot cannot reach earns nothing downstream. Link internally, publish an XML sitemap, and check robots.txt before blaming the content.
Indexing covers comprehension. The engine renders a crawled page, parses it into words and entities, compares it against near-duplicates, and files it — or discards it.
Crawled does not mean indexed. A filter sits between the two stages and asks whether the page is worth keeping. The index drops thin pages, near-identical variants, and content that merely restates what is already indexed, and Google Search Console reports them under coverage states such as Crawled — currently not indexed.
Rendering matters at this stage. Many sites build their content with JavaScript in the browser, so the raw HTML a crawler receives arrives close to empty. Google renders JavaScript, but on a second pass and after a delay. Content that exists in the initial HTML reaches the index faster and more reliably.
Canonicalization settles duplicates. When several URLs carry the same content — a printer version, a tracking parameter, an HTTP and HTTPS pair — the engine picks one URL to represent the set. A rel=canonical tag states a preference; the engine treats the tag as a strong hint rather than an order and can overrule it.
Key concepts
The engine executes the page the way a browser would run it, so scripts, templates, and lazily-loaded text resolve into final content.
The engine chooses one URL to represent a set of duplicates. Signals earned by the others consolidate onto the chosen one.
A meta robots directive that permits crawling but forbids inclusion. Useful for thank-you pages, filters, and staging content.
Schema markup labels what a page contains — a product, a review, an FAQ — making the content easier to classify and eligible for rich results.
Takeaway
The URL Inspection tool in Google Search Console answers the only question that matters at this stage: is this exact URL in the index, and if not, why not?
Ranking happens at query time. The ranking system scores every indexed page that could plausibly answer the search against hundreds of signals, and the order falls out of that scoring.
Google describes its systems as weighing many signals rather than following one master rule, and the weighting shifts by query. A medical question leans hard on expertise and source reputation; a query for pizza leans on proximity and opening hours. No single leaderboard exists — position 1 for a searcher in Tampa on a phone differs from position 1 for a desktop searcher in Dublin.
Four groups make the signals manageable. Relevance asks whether the page addresses the query. Quality asks whether the source deserves trust, largely through links and reputation. Usability asks whether the page works once opened. Context adjusts the result for who is searching, from where, and on which device.
Two habits protect beginners at this stage. Treat every published ranking-factor list as a rough map rather than a specification, and measure movement over weeks, because results fluctuate daily as tests run and competitors publish.
Key concepts
Does the page address the query and the intent behind it? Language, topic coverage, and related entities all feed this judgment.
Links and citations from reputable sites act as endorsements. Google's raters use the E-E-A-T framework — experience, expertise, authoritativeness, trustworthiness — to describe the same idea in human terms.
Mobile rendering, HTTPS, layout stability, and load speed decide whether an otherwise relevant page is pleasant to actually read.
Location, device, language, and recent activity reshape results. The engine rewrites local intent queries around proximity almost entirely.
Takeaway
Ranking runs as a per-query calculation, not a fixed table. Compete on relevance and trust for the specific searches your customers run.
A modern search engine results page (SERP) assembles a layout rather than a list. The blocks a query triggers tell you what winning that query would even look like.
Sponsored results sit at the top, and advertisers buy them through Google Ads. No budget buys the organic results below them — that separation counts as the single most useful fact on this page. Features drawn from the index sit between and around the two: map packs, featured snippets, People Also Ask panels, image and video carousels, and increasingly an AI-generated overview.
SERP features change the prize. When a query returns a map pack, the businesses inside it collect most of the clicks, and a Google Business Profile, not the website, earns that place. When a query returns a featured snippet, the answer often satisfies the searcher outright — a zero-click result — so the value shifts toward being the cited source.
AI overviews extend that logic rather than replacing it. An AI overview summarises and cites pages the engine has already crawled, indexed, and judged trustworthy, so the groundwork stays unchanged: be in the index, be clear, be citable.
Blocks you'll see on a results page
Earned listings ordered by the ranking systems.
Paid placements, labelled, sold by auction.
Three local businesses plus a map, driven by proximity and profile quality.
An extracted answer promoted above the first organic result.
Expandable related questions, each citing a source.
A generated summary that cites indexed pages beneath it.
Takeaway
Run your target query before writing anything. The layout that comes back tells you which asset you actually need to build.
A keyword names the phrase typed into the search box, and search intent names the reason behind it. A page ranks well when it answers the intent, not merely when it contains the keyword.
Keywords fall on a spectrum. Head terms run short, enormous in volume, and brutally competitive. Long-tail keywords run longer, quieter, and far more specific — and because specificity implies a decision nearly made, long-tail keywords convert better. Most new sites earn their first traffic on the tail.
Intent decides page type. A searcher typing how to fix a leaking tap wants instructions, and a service page ranking there gets skipped. A searcher typing emergency plumber near me wants a phone number now. An instruction manual served to that second searcher triggers an immediate bounce, no matter how well optimized the page is.
Practical keyword mapping follows a simple discipline: one primary intent per URL, one page per intent, and no two pages chasing the same phrase. When two of your own pages target one query, the engine picks between them — usually the weaker one — and both underperform. SEO calls that self-inflicted problem keyword cannibalization.
| Intent | What the searcher wants | Example query | Page that wins |
|---|---|---|---|
| Informational | To learn or understand something | how do search engines work | Guide, explainer, or hub page |
| Navigational | To reach a specific site or brand | keever seo pricing | The brand's own page |
| Commercial | To compare before deciding | best seo agency for dentists | Comparison, review, or case study |
| Transactional | To act, buy, book, or call | seo audit near me | Service or product page with a clear next step |
Takeaway
Pick the keyword, then read the top ten results already ranking for it. Their shared format shows the format the engine has decided satisfies that intent.
Every optimization task belongs to one of three pillars: on-page SEO covers what sits on the page, off-page SEO covers what the rest of the web says about it, and technical SEO covers whether the machine can process it at all.
On-page SEO covers everything a visitor and a parser both read — the title tag, the headings, the body copy, image alt text, internal links, and the structured data describing the page. The on-page pillar carries the fastest feedback loop, because changes ship the moment the page is re-crawled.
Off-page SEO covers signals earned elsewhere: backlinks from relevant sites, brand mentions, reviews, and citations of your name, address, and phone number across directories. Off-page signals resist faking and, unsurprisingly, carry the most weight for competitive queries.
Technical SEO provides the substrate: crawlability, indexation, site speed, mobile rendering, HTTPS, clean redirects, and a URL structure that mirrors the site's logic. None of it wins a ranking on its own, and all of it can prevent one.
Key concepts
Titles, headings, body copy, internal linking, and schema markup — the parts of relevance under direct editorial control.
Backlinks, unlinked brand mentions, reviews, and local citations that establish how the wider web regards the site.
Crawl paths, indexation rules, Core Web Vitals, mobile rendering, and redirect hygiene that let the other two pillars count.
Takeaway
Fix technical blockers first, then sharpen on-page relevance, then earn off-page authority. Reversing that order wastes the most money.
Two free tools answer nearly every beginner question. Google Search Console reports how search sees the site; Google Analytics reports what visitors do once they arrive.
Start with impressions rather than rankings. An impression records one appearance of a page in someone's results, and impressions climb before positions do, which makes impressions the earliest honest signal that the work is landing. A single keyword position, watched instead, misleads you for weeks.
Then read clicks and click-through rate (CTR) together. High impressions with a low click-through rate usually mean searchers see the listing and skip it, so the title tag and meta description need the fix rather than more content. Low impressions mean the relevance or authority problem still sits upstream.
Finish with conversions. Traffic that never enquires counts as a vanity number, and a page ranking third for a term that converts beats a page ranking first for one that never does.
| Metric | What it tells you | Where to find it |
|---|---|---|
| Impressions | How often your pages surfaced in results | Search Console — Performance |
| Clicks and CTR | Whether the listing persuades people to open it | Search Console — Performance |
| Average position | Roughly where a query places, averaged over time | Search Console — Performance |
| Indexed pages | How much of the site is eligible to rank at all | Search Console — Pages |
| Core Web Vitals | Whether real visitors experience the page as fast and stable | Search Console — Core Web Vitals |
| Conversions | Enquiries, calls, and sales attributable to organic search | Google Analytics 4 |
Takeaway
Verify the site in Google Search Console before anything else. Without Search Console, every judgment about search performance rests on guesswork.
Unlearn These First
Bad SEO advice survives because it sounds plausible. Each claim below circulated widely, and each one misdirects a budget in a specific way — the correction beside it names what actually holds.
Myth 01
What's true
Google confirmed in 2009 that it ignores the meta keywords tag. The tags that still earn their place are the title tag, which is a ranking and click factor, and the meta description, which shapes clicks alone.
Myth 02
What's true
Ad spend buys placement in the sponsored block, clearly labelled and separate. The ranking systems award organic positions, and no one sells them — not Google, and not anyone offering guarantees.
Myth 03
What's true
Competitors publish, algorithms update, and content ages. Positions decay without maintenance, which is why search performance behaves like a garden rather than a build.
Myth 04
What's true
Indexes reward pages that resolve a question. Thin, overlapping pages compete with each other, dilute internal linking, and often leave a site with more URLs and fewer visits.
Myth 05
What's true
Qualified enquiries are the goal. First place for a term nobody buys from is worth less than third place for the query your best customers actually type.
Myth 06
What's true
AI overviews and chat assistants summarise pages that were crawled, indexed, and judged trustworthy first. The groundwork did not change — the reward simply moved toward being the cited source.
Reference
The glossary defines the 28 terms that recur in every SEO conversation, one sentence per term (crawler, index, canonical tag, featured snippet, and more). Keep the filter open while you read the chapters.
No term matches that filter. Try a shorter word — or ask us directly.
Check Yourself
The knowledge check draws five questions from the chapters above and shows them one at a time. Pick an answer and the reasoning appears straight away, then the next question replaces it.
Your score
0 / 5
Nothing answered yet.
Question 1 of 5
Why: Results are drawn from an index assembled in advance. Pages that were never crawled and indexed cannot appear, however good they are.
Why: robots.txt controls crawler access only. To keep a crawlable page out of the index you need a noindex directive instead.
Why: The intent is informational. Serving a service page to an informational query loses the click, no matter how well optimized that page is.
Why: Impressions prove searchers see the listing. When searchers see it and skip it, the headline and summary in the results page become the lever.
Why: Google sells ads by auction and labels them as sponsored. The ranking systems award organic positions and features such as snippets.
Result
After the Basics
The fundamentals carry a site a long way. The guides below pick up where the chapters stop — one per pillar (technical, on-page, and off-page SEO), plus the strategy, the mistakes to avoid, and the paid-versus-organic comparison that decide where to start.
Thirteen tips cover a flat architecture, topical clusters, canonical tags, faceted navigation controls, and how structure affects rankings.
Read the guideA seven-step workflow maps priority pages, sets anchor text rules, and audits the internal links that spread authority through a site.
Read the guideAn eight-step method builds authority through canonical consolidation, intent-aligned content, linkable assets, editorial backlinks, and link reclamation.
Read the guideA nine-step workflow runs from auditing current positions and matching keywords to intent through on-page, technical, and backlink work to tracking results.
Read the guideThe eight most common SEO mistakes, from intent mismatch and thin content to broken architecture and spammy backlinks, with the fix order.
Read the guideSEO, SEM, and PPC differ in placement, cost, speed, longevity, and control; the guide shows when a business should choose each and how to combine them.
Read the guideFAQ
Beginners ask these questions most often once the chapters land. The groups below follow the stages (search basics, ranking, and getting started), so you can jump to the one you are stuck on.
A search engine is software that discovers pages on the web with automated crawlers, stores what it finds in a searchable index, and orders that index against whatever someone types. Google, Bing, and DuckDuckGo all follow that three-part pattern.
Google reaches new sites through links from pages it already crawls and through XML sitemaps submitted in Google Search Console. A site with no inbound links and no sitemap can wait weeks for discovery, so submitting the sitemap on launch day is the fastest reliable route.
The usual causes, in order of frequency: the page has not been crawled yet, it is blocked by robots.txt or a noindex tag, it duplicates another URL and was consolidated away, or it is indexed but ranking too low to notice. The URL Inspection tool in Search Console identifies which of the four applies.
Paid search buys labelled placements in the sponsored block and stops the moment the budget stops. SEO earns organic positions that cannot be bought, take longer to build, and keep returning traffic after the work is finished.
Google describes its systems as weighing hundreds of signals, and the exact list and weighting are not published. Treating any circulated factor list as a specification is a mistake; the durable groupings are relevance, quality and trust, usability, and searcher context.
Links from relevant, reputable sites remain among the strongest signals of trust, particularly for competitive queries. Volume alone has lost its power — a handful of genuine editorial links outperforms hundreds of low-quality ones.
Core Web Vitals measure speed and stability, and the ranking systems weigh them, most visibly as a tiebreaker between pages of similar relevance. The larger commercial effect is behavioural: slow pages lose visitors before the content is ever read.
Technical fixes can show up within days of a re-crawl. Content and authority work typically takes three to six months to move competitive positions, and longer in crowded markets or on new domains with no established trust.
Verify the site in Google Search Console, submit an XML sitemap, and confirm the pages that matter are indexed. Those three steps cost nothing, take under an hour, and turn every later decision from guesswork into measurement.
Google Search Console for impressions, clicks, indexing, and Core Web Vitals. Google Analytics 4 for what visitors do after arriving. Google Business Profile for any business with a physical location or service area. Together they cover almost everything a beginner needs.
Beginners can genuinely learn the fundamentals in this guide, and small sites often get a long way alone. Agencies earn their fee where the work becomes specialist or unbounded — technical migrations, competitive link acquisition, multi-location visibility, and penalty recovery.
Track impressions first, since they move before positions do, then clicks and click-through rate, then enquiries from organic search. Reviewing those three in Search Console and Analytics each month shows the trend far more honestly than a daily rank check.
Still have questions? Get in touch — we're happy to help.
Keever
Typically replies within 1 hour
Opens our live calendar — pick a slot in under a minute.