Learning Hub · Start Here

Search Engine Basics

A search engine serves as an automated retrieval system that returns web pages for a query (organic listings, map packs, featured snippets, and AI overviews)

A search engine serves as an automated retrieval system that returns web pages for a query (organic listings, map packs, featured snippets, and AI overviews). Search Engine Optimization (SEO) starts when a crawler discovers a URL, the index stores and interprets the page, and the ranking system scores it against each search. The engine serves the results from its stored copy of the web, ensuring answers in milliseconds rather than a live scan of every site. Opt for the URL Inspection tool in Google Search Console to confirm the indexed status if a page never appears in results. Visibility is influenced by crawlability, relevance, and trust, with eight beginner chapters, a 28-term glossary, and a knowledge check available below for readers on their first pass. Business owners and marketers utilize these basics to judge their own site, creating a shared vocabulary while briefing an SEO agency or running the work themselves.

  • 8 chapters
  • 38 min read
  • Beginner level
  • Updated August 2026

Learning Outcomes

What Search Engine Basics Teach You

Four capabilities separate a reader who has studied SEO from a practitioner who can act on it. Each chapter below builds one of them, and the knowledge check at the end confirms all four.

01

Explain how search works

Walk any colleague through the four stages a page passes before it can appear: crawling, indexing, ranking, and serving.

02

Read a results page

Tell organic listings, sponsored ads, map packs, featured snippets, and AI overviews apart at a glance — and know which one you can win.

03

Match pages to intent

Sort a keyword into informational, navigational, commercial, or transactional, then choose the page type that satisfies it.

04

Judge your own site

Check whether a URL is indexed, find the queries it already earns impressions for, and rank the fixes worth doing first.

The Whole Model in One Diagram

How a Search Actually Works

Every query passes through the same four stages (crawl, index, rank, and serve). Almost every SEO problem traces back to a failure at one specific stage, so naming the stage delivers most of the fix.

STAGE 01

Crawl

Bots discover the URL.

Googlebot follows links, reads XML sitemaps, and revisits pages it already knows to find what has changed.

STAGE 02

Index

The page is stored and understood.

The engine renders the crawled page, splits it into words and entities, deduplicates it against near-copies, and files it in the index.

STAGE 03

Rank

Candidates are scored per query.

At the moment of the search, hundreds of signals score every indexed page that could answer the query, and an order falls out.

STAGE 04

Serve

A results page is assembled.

Ten blue links, a map pack, a snippet, an AI overview — the layout is built around what the query appears to need.

Chapter 01 Foundation 4 min

What a Search Engine Actually Is

A search engine combines three systems: a crawler that discovers pages, an index that stores them, and a ranking system that orders the index against a query.

Google, Bing, and DuckDuckGo all run the same three-part structure. The size of the index, the frequency of the crawler's return, and the signals the ranking system trusts separate one engine from another. A reader who learns the structure once can read every engine, including the newer AI answer engines.

The engine searches nothing live. A typed query never sends the engine out to read the web; the engine reads its own copy of the web, assembled in advance. Every technique in search engine optimization exists either to get a page into that copy or to give the ranking system reasons to place the page near the top.

The stored copy also explains the lag beginners notice. A page published this morning might rank this afternoon or six weeks from now, and the timing depends entirely on when a crawler reaches the page and how readily the index accepts it.

Key concepts

  • The crawler

    A program such as Googlebot requests pages the way a browser does, then follows the links it finds to reach the next URL.

  • The index

    A distributed database of processed pages, broken into words, entities, and signals so a query can be matched in milliseconds.

  • The ranking system

    Scoring happens at query time. The ranking system recalculates the order for every search, device, and location — no page holds a permanent position.

Takeaway

Search engines answer from a stored copy of the web, not the live web. Crawlability and indexability come before any attempt to rank.

Chapter 02 Foundation 5 min

Crawling: How Search Engines Find Your Pages

Crawling covers discovery. A bot such as Googlebot fetches a URL, reads the HTML, extracts every link on the page, and queues those links to fetch next.

Links form the road network of a site. A page with no internal link pointing at it and no entry in a sitemap leaves the crawler without a route, however good the content is. Orphan pages quietly earn nothing for months for exactly that reason.

Two files shape the journey. An XML sitemap hands the crawler a list of URLs worth fetching, which matters most on large sites and brand-new domains. A robots.txt file does the opposite and tells compliant bots which paths to leave alone. A directory blocked in robots.txt ranks among the most common and expensive accidents: the pages inside stop being crawled, so their content never reaches the index.

Crawlers also budget their time. Google allocates each site a rough crawl capacity based on server response speed and the demand the site's pages generate. Slow responses, endless faceted URLs, and long redirect chains burn that crawl budget on pages nobody searches for.

Key concepts

  • Internal links

    The most reliable discovery path. Anything reachable in a few clicks from the homepage gets found early and re-crawled often.

  • XML sitemap

    A machine-readable list of the URLs a site wants crawled, submitted through Google Search Console and Bing Webmaster Tools.

  • robots.txt

    A plain text file at the domain root that permits or blocks crawler access per path. The file controls crawling, never indexing.

  • Crawl budget

    The practical ceiling on how many URLs a bot fetches per visit. Fast servers and tidy URL structures stretch it further.

Takeaway

A URL the bot cannot reach earns nothing downstream. Link internally, publish an XML sitemap, and check robots.txt before blaming the content.

Chapter 03 Foundation 5 min

Indexing: How Pages Get Stored and Understood

Indexing covers comprehension. The engine renders a crawled page, parses it into words and entities, compares it against near-duplicates, and files it — or discards it.

Crawled does not mean indexed. A filter sits between the two stages and asks whether the page is worth keeping. The index drops thin pages, near-identical variants, and content that merely restates what is already indexed, and Google Search Console reports them under coverage states such as Crawled — currently not indexed.

Rendering matters at this stage. Many sites build their content with JavaScript in the browser, so the raw HTML a crawler receives arrives close to empty. Google renders JavaScript, but on a second pass and after a delay. Content that exists in the initial HTML reaches the index faster and more reliably.

Canonicalization settles duplicates. When several URLs carry the same content — a printer version, a tracking parameter, an HTTP and HTTPS pair — the engine picks one URL to represent the set. A rel=canonical tag states a preference; the engine treats the tag as a strong hint rather than an order and can overrule it.

Key concepts

  • Rendering

    The engine executes the page the way a browser would run it, so scripts, templates, and lazily-loaded text resolve into final content.

  • Canonicalization

    The engine chooses one URL to represent a set of duplicates. Signals earned by the others consolidate onto the chosen one.

  • noindex

    A meta robots directive that permits crawling but forbids inclusion. Useful for thank-you pages, filters, and staging content.

  • Structured data

    Schema markup labels what a page contains — a product, a review, an FAQ — making the content easier to classify and eligible for rich results.

Takeaway

The URL Inspection tool in Google Search Console answers the only question that matters at this stage: is this exact URL in the index, and if not, why not?

Chapter 04 Core 5 min

Ranking: How Results Get Ordered

Ranking happens at query time. The ranking system scores every indexed page that could plausibly answer the search against hundreds of signals, and the order falls out of that scoring.

Google describes its systems as weighing many signals rather than following one master rule, and the weighting shifts by query. A medical question leans hard on expertise and source reputation; a query for pizza leans on proximity and opening hours. No single leaderboard exists — position 1 for a searcher in Tampa on a phone differs from position 1 for a desktop searcher in Dublin.

Four groups make the signals manageable. Relevance asks whether the page addresses the query. Quality asks whether the source deserves trust, largely through links and reputation. Usability asks whether the page works once opened. Context adjusts the result for who is searching, from where, and on which device.

Two habits protect beginners at this stage. Treat every published ranking-factor list as a rough map rather than a specification, and measure movement over weeks, because results fluctuate daily as tests run and competitors publish.

Key concepts

  • Relevance

    Does the page address the query and the intent behind it? Language, topic coverage, and related entities all feed this judgment.

  • Quality and trust

    Links and citations from reputable sites act as endorsements. Google's raters use the E-E-A-T framework — experience, expertise, authoritativeness, trustworthiness — to describe the same idea in human terms.

  • Usability

    Mobile rendering, HTTPS, layout stability, and load speed decide whether an otherwise relevant page is pleasant to actually read.

  • Context

    Location, device, language, and recent activity reshape results. The engine rewrites local intent queries around proximity almost entirely.

Takeaway

Ranking runs as a per-query calculation, not a fixed table. Compete on relevance and trust for the specific searches your customers run.

Chapter 05 Core 4 min

The Results Page: What Searchers Actually See

A modern search engine results page (SERP) assembles a layout rather than a list. The blocks a query triggers tell you what winning that query would even look like.

Sponsored results sit at the top, and advertisers buy them through Google Ads. No budget buys the organic results below them — that separation counts as the single most useful fact on this page. Features drawn from the index sit between and around the two: map packs, featured snippets, People Also Ask panels, image and video carousels, and increasingly an AI-generated overview.

SERP features change the prize. When a query returns a map pack, the businesses inside it collect most of the clicks, and a Google Business Profile, not the website, earns that place. When a query returns a featured snippet, the answer often satisfies the searcher outright — a zero-click result — so the value shifts toward being the cited source.

AI overviews extend that logic rather than replacing it. An AI overview summarises and cites pages the engine has already crawled, indexed, and judged trustworthy, so the groundwork stays unchanged: be in the index, be clear, be citable.

Blocks you'll see on a results page

  • Organic results

    Earned listings ordered by the ranking systems.

  • Sponsored ads

    Paid placements, labelled, sold by auction.

  • Map pack

    Three local businesses plus a map, driven by proximity and profile quality.

  • Featured snippet

    An extracted answer promoted above the first organic result.

  • People Also Ask

    Expandable related questions, each citing a source.

  • AI overview

    A generated summary that cites indexed pages beneath it.

Takeaway

Run your target query before writing anything. The layout that comes back tells you which asset you actually need to build.

Chapter 06 Core 5 min

Keywords and Search Intent

A keyword names the phrase typed into the search box, and search intent names the reason behind it. A page ranks well when it answers the intent, not merely when it contains the keyword.

Keywords fall on a spectrum. Head terms run short, enormous in volume, and brutally competitive. Long-tail keywords run longer, quieter, and far more specific — and because specificity implies a decision nearly made, long-tail keywords convert better. Most new sites earn their first traffic on the tail.

Intent decides page type. A searcher typing how to fix a leaking tap wants instructions, and a service page ranking there gets skipped. A searcher typing emergency plumber near me wants a phone number now. An instruction manual served to that second searcher triggers an immediate bounce, no matter how well optimized the page is.

Practical keyword mapping follows a simple discipline: one primary intent per URL, one page per intent, and no two pages chasing the same phrase. When two of your own pages target one query, the engine picks between them — usually the weaker one — and both underperform. SEO calls that self-inflicted problem keyword cannibalization.

Intent What the searcher wants Example query Page that wins
Informational To learn or understand something how do search engines work Guide, explainer, or hub page
Navigational To reach a specific site or brand keever seo pricing The brand's own page
Commercial To compare before deciding best seo agency for dentists Comparison, review, or case study
Transactional To act, buy, book, or call seo audit near me Service or product page with a clear next step

Takeaway

Pick the keyword, then read the top ten results already ranking for it. Their shared format shows the format the engine has decided satisfies that intent.

Chapter 07 Applied 5 min

The Three Pillars of SEO

Every optimization task belongs to one of three pillars: on-page SEO covers what sits on the page, off-page SEO covers what the rest of the web says about it, and technical SEO covers whether the machine can process it at all.

On-page SEO covers everything a visitor and a parser both read — the title tag, the headings, the body copy, image alt text, internal links, and the structured data describing the page. The on-page pillar carries the fastest feedback loop, because changes ship the moment the page is re-crawled.

Off-page SEO covers signals earned elsewhere: backlinks from relevant sites, brand mentions, reviews, and citations of your name, address, and phone number across directories. Off-page signals resist faking and, unsurprisingly, carry the most weight for competitive queries.

Technical SEO provides the substrate: crawlability, indexation, site speed, mobile rendering, HTTPS, clean redirects, and a URL structure that mirrors the site's logic. None of it wins a ranking on its own, and all of it can prevent one.

Key concepts

  • On-page SEO

    Titles, headings, body copy, internal linking, and schema markup — the parts of relevance under direct editorial control.

  • Off-page SEO

    Backlinks, unlinked brand mentions, reviews, and local citations that establish how the wider web regards the site.

  • Technical SEO

    Crawl paths, indexation rules, Core Web Vitals, mobile rendering, and redirect hygiene that let the other two pillars count.

Takeaway

Fix technical blockers first, then sharpen on-page relevance, then earn off-page authority. Reversing that order wastes the most money.

Chapter 08 Applied 5 min

Measuring SEO: The Metrics That Matter

Two free tools answer nearly every beginner question. Google Search Console reports how search sees the site; Google Analytics reports what visitors do once they arrive.

Start with impressions rather than rankings. An impression records one appearance of a page in someone's results, and impressions climb before positions do, which makes impressions the earliest honest signal that the work is landing. A single keyword position, watched instead, misleads you for weeks.

Then read clicks and click-through rate (CTR) together. High impressions with a low click-through rate usually mean searchers see the listing and skip it, so the title tag and meta description need the fix rather than more content. Low impressions mean the relevance or authority problem still sits upstream.

Finish with conversions. Traffic that never enquires counts as a vanity number, and a page ranking third for a term that converts beats a page ranking first for one that never does.

Metric What it tells you Where to find it
Impressions How often your pages surfaced in results Search Console — Performance
Clicks and CTR Whether the listing persuades people to open it Search Console — Performance
Average position Roughly where a query places, averaged over time Search Console — Performance
Indexed pages How much of the site is eligible to rank at all Search Console — Pages
Core Web Vitals Whether real visitors experience the page as fast and stable Search Console — Core Web Vitals
Conversions Enquiries, calls, and sales attributable to organic search Google Analytics 4

Takeaway

Verify the site in Google Search Console before anything else. Without Search Console, every judgment about search performance rests on guesswork.

Unlearn These First

Six Beliefs That Cost Beginners Money

Bad SEO advice survives because it sounds plausible. Each claim below circulated widely, and each one misdirects a budget in a specific way — the correction beside it names what actually holds.

Myth 01

Meta keywords still influence rankings.

What's true

Google confirmed in 2009 that it ignores the meta keywords tag. The tags that still earn their place are the title tag, which is a ranking and click factor, and the meta description, which shapes clicks alone.

Myth 02

You can pay Google for a higher organic position.

What's true

Ad spend buys placement in the sponsored block, clearly labelled and separate. The ranking systems award organic positions, and no one sells them — not Google, and not anyone offering guarantees.

Myth 03

SEO is a one-time project you finish.

What's true

Competitors publish, algorithms update, and content ages. Positions decay without maintenance, which is why search performance behaves like a garden rather than a build.

Myth 04

More pages automatically means more traffic.

What's true

Indexes reward pages that resolve a question. Thin, overlapping pages compete with each other, dilute internal linking, and often leave a site with more URLs and fewer visits.

Myth 05

Ranking first is the goal.

What's true

Qualified enquiries are the goal. First place for a term nobody buys from is worth less than third place for the query your best customers actually type.

Myth 06

AI answers have made SEO obsolete.

What's true

AI overviews and chat assistants summarise pages that were crawled, indexed, and judged trustworthy first. The groundwork did not change — the reward simply moved toward being the cited source.

Reference

SEO Glossary

The glossary defines the 28 terms that recur in every SEO conversation, one sentence per term (crawler, index, canonical tag, featured snippet, and more). Keep the filter open while you read the chapters.

301 redirect
A permanent forward from one URL to another that passes the original page's ranking signals to the new address.
AI Overview
A generated summary placed above search results that answers a query and cites the indexed pages it drew from.
Algorithm
The set of systems a search engine uses to score and order indexed pages for a given query.
Alt text
A written description of an image, read by screen readers and used by search engines to understand what the image shows.
Anchor text
The visible, clickable words of a link. It signals to search engines what the destination page is about.
Backlink
A link from another website to yours, treated as an endorsement when it comes from a relevant, reputable source.
Canonical tag
An HTML tag naming the preferred URL among duplicates, so ranking signals consolidate on one version.
Core Web Vitals
Google's measures of loading speed, interaction responsiveness, and visual stability, sampled from real visits.
Crawl budget
The practical number of URLs a search engine will fetch from a site in a given period.
Crawler
An automated program, such as Googlebot, that requests pages and follows their links to discover more URLs.
CTR
Click-through rate — the share of impressions that turn into clicks, expressed as a percentage.
E-E-A-T
Experience, expertise, authoritativeness, and trustworthiness: the framework Google's quality raters use to describe credible sources.
Featured snippet
An answer extracted from a page and promoted into a box above the first organic result.
Impression
One appearance of your page in someone's search results, whether or not it was clicked.
Index
The searchable database of processed pages a search engine consults when a query is entered.
Internal link
A link from one page of a site to another. Internal links spread discoverability and authority through a site.
Keyword
A word or phrase entered into a search box, and the term a page is optimized to answer.
Keyword cannibalization
Two or more pages on one site targeting the same query, forcing the engine to choose and weakening both.
Long-tail keyword
A longer, more specific phrase with lower volume, less competition, and stronger commercial intent.
Meta description
The summary shown beneath a listing. It influences click-through rate rather than ranking position.
Nofollow
A link attribute telling search engines not to pass endorsement to the destination, common on ads and user-generated content.
Organic result
An earned listing placed by the ranking systems, distinct from the sponsored ads above it.
robots.txt
A file at the domain root that tells compliant crawlers which paths they may fetch. It governs crawling, not indexing.
Schema markup
Structured data that labels page content — a product, review, recipe, or FAQ — making rich results possible.
Search intent
The underlying goal behind a query: to learn, to navigate, to compare, or to act.
SERP
Search engine results page — the assembled screen of organic listings, ads, and features returned for a query.
Sitemap
An XML file listing the URLs a site wants crawled, submitted through Search Console to speed up discovery.
Title tag
The HTML title of a page, used as the headline of its search listing and weighed as a relevance signal.

Check Yourself

SEO Basics Knowledge Check

The knowledge check draws five questions from the chapters above and shows them one at a time. Pick an answer and the reasoning appears straight away, then the next question replaces it.

Your score

0 / 5

Nothing answered yet.

Question 1 of 5

  1. When you press enter on a search, what does the engine actually search?

  2. A page is blocked in robots.txt. What does that stop?

  3. Someone searches how to unclog a drain. Which page should win?

  4. Your page shows plenty of impressions but very few clicks. What do you fix first?

  5. Which of these can be bought directly from Google?

After the Basics

SEO Guides to Explore Next

The fundamentals carry a site a long way. The guides below pick up where the chapters stop — one per pillar (technical, on-page, and off-page SEO), plus the strategy, the mistakes to avoid, and the paid-versus-organic comparison that decide where to start.

FAQ

Search Engine Questions, Answered

Beginners ask these questions most often once the chapters land. The groups below follow the stages (search basics, ranking, and getting started), so you can jump to the one you are stuck on.

What is a search engine in simple terms?

A search engine is software that discovers pages on the web with automated crawlers, stores what it finds in a searchable index, and orders that index against whatever someone types. Google, Bing, and DuckDuckGo all follow that three-part pattern.

How does Google find a brand-new website?

Google reaches new sites through links from pages it already crawls and through XML sitemaps submitted in Google Search Console. A site with no inbound links and no sitemap can wait weeks for discovery, so submitting the sitemap on launch day is the fastest reliable route.

Why is my page not showing up on Google?

The usual causes, in order of frequency: the page has not been crawled yet, it is blocked by robots.txt or a noindex tag, it duplicates another URL and was consolidated away, or it is indexed but ranking too low to notice. The URL Inspection tool in Search Console identifies which of the four applies.

What is the difference between SEO and paid search?

Paid search buys labelled placements in the sponsored block and stops the moment the budget stops. SEO earns organic positions that cannot be bought, take longer to build, and keep returning traffic after the work is finished.

Still have questions? Get in touch — we're happy to help.