Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Mastering the Essentials of Information Retriev...

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest. →

Mastering the Essentials of Information Retrieval, Machine Learning and SEO Without the MSc Degrees

Bridging the Gap Between Information Retrieval Science and Modern SEO

Moving beyond surface-level blog advice and speculative 'GuesSEO', this presentation delivers a blueprint for applying academic Information Retrieval (IR), Machine Learning, and Computer Science principles to modern SEO strategy, without the £25k+ price tag of formal degree programmes.

Designed to transition marketers from a traditional T-shaped skill set to a technically grounded Y-shaped framework, this talk breaks down how search engines work at a foundational level. From inverted indexes, BM25 scoring, and vector embeddings to modern agentic RAG, generative IR, and Google DeepMind research, attendees will gain a deep, research-backed understanding of search algorithms.

What You’ll Discover:

The Core Science of Search: A breakdown of classic and modern IR foundations, including hybrid retrieval models, phrase-based indexing, and multi-stage search pipelines.
Research over Rumours: How studying academic papers and attending gold-standard IR conferences allows you to anticipate algorithm updates like BERT and dense retrieval years before they become SEO industry buzzwords.
The 'Zero-Tuition' Curriculum: A step-by-step, 4-phase self-study roadmap leveraging free, open-source academic tools, research databases, and public computer science courses to elevate your critical thinking and critical analysis.
Practical SEO Transferability: Actionable strategies for transferring academic concepts—such as knowledge graphs, entity relationships, and vector link architecture—into everyday execution.

Whether you want to eliminate guesswork, critically analyse industry claims, or future-proof your strategies against generative search evolution, this talk provides the roadmap to becoming a research-backed search practitioner.

Avatar for Dawn Anderson

Dawn Anderson

October 05, 2026

More Decks by Dawn Anderson

Other Decks in Marketing & SEO

Transcript

  1. Mastering the Essentials of Information Retrieval , Machine Learning and

    SEO Without the MSc Degrees  The journey of a sort-of pracademic in SEO
  2. 3

  3. 4 But… mostly an SEO consultant. A ‘pracademic’ SEO Consultant

    Combining information retrieval and research knowledge with practical application and SEO strategies.   Academic Rigour Practical Strategy Grounded in information retrieval, computer science, and peer-reviewed research. Hands-on industry application, technical execution, and commercial results.
  4. In 2011, I was an SEO focused on the T-shaped

    marketer profile Broad digital marketing knowledge across disciplines, anchored by deep technical SEO expertise.  Broad Knowledge Across Marketing Disciplines (The Crossbar)  Content Strategy  Paid Acquisition  UX & Design  Analytics & Data  Social & PR Copywriting, storytelling, content taxonomy, and editorial planning. PPC campaigns, social ads, bidding strategy, and ad creative. User journeys, CRO, wireframing, and accessibility best practices. GA4, event tracking, funnel analysis, and attribution modeling. Brand positioning, community management, and digital outreach.  Deep Core Expertise (The Vertical Stem)  Strategic & Technical SEO Mastery over search engine architecture and organic growth drivers. Technical Architecture Crawl budget, indexing pipeline, JavaScript rendering, site architecture, structured data, and Core Web Vitals optimisation. Intent Understanding Searcher intent mapping, semantic keyword clustering, entity relationship modelling, and vector search alignment. Authority & Content Strategy Topic authority, link graph analysis, programmatic SEO scaling, and organic revenue performance attribution.
  5. 6 I was an SEO obsessed with crawling and learning

    about broader digital marketing .  Obsessed with Crawling Deep interest in search engine mechanics, indexing, and technical architecture.  Broader Marketing Continuously expanding skills into full-stack digital strategy and acquisition.
  6. My ‘Pracademic’ Education Timeline 2011 Google Squared Digital marketing and

    leadership qualification. 2012−2014 2015−2017 2024–2027 PgDip Digital Marketing MSc Digital Marketing Postgraduate diploma in digital marketing strategy. Master of Science focused on advanced digital strategy. MSc Computer Science with Data Science Advanced degree combining computer science and data analytics foundations.
  7. 8 Plus lots of other courses and qualifications along the

    way.  Prince 2 Project Management  Level 3 AET Award in Education and Training   Agile Project Management Change Management   Google AdWords Google Analytics
  8. 9 Whilst studying for my MSc Digital Marketing Strategy dissertation

    in 2016, I discovered academic papers on web crawling and patents from Google and other researchers . SPEAKER PRESENTATION INDUSTRY PUBLICATION CRAWLING MECHANICS
  9. 10 Core Technical Foundations The Fundamentals of Search Engine Crawling

    in Great Detail Direct insights from Google engineers and leading web researchers Breadth-First Crawling Graph traversal mechanics & scalable link discovery architecture Crawl Scheduling Policies Importance scoring algorithms, recrawl optimization & politeness constraints Production Architecture Real-world web crawler design by Christopher Olston & Marc Najork
  10. 11 FOUNDATIONAL DISCOVERY IR SYSTEM ARCHITECTURE I discovered the fascinating

    world of information retrieval (IR) The Science Behind Search Moving beyond surface heuristics to understand the mechanical engineering of search engines. • Crawling & Inverted Indexing: Structuring unstructured web data at scale. • Scoring & Relevance Algorithms: Mathematically evaluating query-document fit. • Vector Space & Neural Models: Powering semantic comprehension in modern AI search.
  11. I learned about the core foundations of Classic IR Inverted

    Index TF-IDF & BM25 The core data structure maps every word/token to the exact documents where it appears. Statistical scoring functions that weigh term frequency against document rarity across the corpus. Vector Space Model Link Topology Analysis Representing documents and queries as geometric vectors to measure spatial closeness. Algorithms like PageRank and HITS that evaluate authority via graph structural analysis.
  12. And the 3 Pillars of Modern Information Retrieval Vector Embeddings

    Hybrid Retrieval Transformer Models Mapping words, sentences, and pages into Combining traditional sparse lexical search Attention mechanisms decode context, dense multi-dimensional mathematical space. with high-speed dense vector similarity entities, and deep semantic intent across search. passages.
  13. I Learned About the 3−Stage Search Pipeline 1. Crawl &

    Index 2. First-Stage Retrieval 3. Multi-Pass Re-Ranking Tokenisation, stemming, stop-word Rapid candidate generation retrieving Heavy deep learning models (like removal, and building the inverted index top 1,000 matching documents using BERT/RankTransformer), user context, to store document tokens efficiently. lightweight BM25 or ANN vector search. location, and freshness applied to top results.
  14. And the Growth of Dense Retrieval Adoption Transition rate of

    commercial search systems moving from keyword matching to neural vector search.
  15. And the Evolution of Machine Learning in SEO Search systems

    shifted from simple keyword counts to complex neural networks measuring natural language semantics.
  16. I discovered this was a world where engineers and researchers

    shared learnings and developments about search beyond what I could learn in SEO. FIRST IR CONFERENCE NETWORKING & EVENTS RESEARCH INSIGHTS 17
  17. 18 With Gold Standard IR Conferences I could attend or

    follow Long-established academic conferences shaping search algorithms & information retrieval CHIIR ECIR Human Information Interaction Focuses on human-centered aspects of information retrieval systems and interfaces. European Conference on IR The major European forum for presenting scientific research in information retrieval. SIGIR Research & Development in IR The premier international conference on information retrieval theory and practice. ESSIR Summer School on IR Advanced teaching and foundational learning for researchers and specialists. Also including premier venues: WSDM Web Search and Data Mining) and NeurIPS / NIPS Neural Information Processing Systems)
  18. Google engineers and researchers regularly present at IR conferences. This

    is the world of science behind search engines. 19
  19. 20 I attended an IR conference. I was an outlier.

    An interloper. “SEOs do not attend our conferences." “Who are you, and why are you here?” (Implied)
  20. 24 8 years later… I delivered a keynote talk on

    the relationship between the worlds of SEO and IR.
  21. 25 Example learnings from studying information retrieval and following researchers

    and their papers     Duplicate & near-duplicate content Web crawling Query intent detection Query intent shift     Conversational search Natural language processing & computational linguistics Recommender systems Contextual search     Result diversification Human in the loop Similarity & relatedness (co-occurrence) Loads more
  22. By 2024, it was time to think more like a

    Y-shaped marketer Bridging Commercial Marketing & Technical Computer Science with Deep SEO Expertise  Marketing & Business Strategy (Left Arm)  Computer Science & Data (Right Arm)  Content Strategy  Paid & Social  Development  Data Science Copywriting, storytelling, content taxonomy, and editorial planning. PPC, paid social campaigns, brand positioning, and attribution strategy. HTML, CSS, JavaScript, rendering pipelines, and Web Architecture. NLP, vector embeddings, ML models, and search algorithm mechanics.  Deep Core Expertise (The Vertical Stem)  Strategic & Technical SEO Mastery over search engine architecture and organic growth drivers. Technical Architecture & Indexing Crawl budget, indexing pipelines, JS rendering, site architecture, structured data, and Core Web Vitals. Information Retrieval & Intent Searcher intent mapping, semantic keyword clustering, entity relationship modeling, and vector search alignment. Authority & Commercial Growth Topic authority, programmatic SEO scaling, link graph analysis, and organic revenue performance attribution.
  23. MSc Computer Science & Data Science Curriculum Overview & Technical

    Module Focus     Data Science Data Mining Informatics Data Visualisation Advanced data analysis, predictive modeling, and computational frameworks. Pattern extraction, knowledge discovery, and large-scale dataset analysis. Information processing systems, human-computer interaction, and structure. Visual communication of complex data, dashboards, and graph design.     Statistics for AI & Data Science Project Management AI Technologies Database & Security Agile methodologies, software lifecycles, delivery, and team execution. Machine learning models, neural networks, and algorithmic intelligence. Relational/NoSQL databases, data protection, and security protocols. Probability distributions, statistical inference, and mathematical foundations.
  24. 29 And limiting ‘GuesSEO’ GOLDEN RULE Support all your decisions

    with data or respected and credible sources. • Eliminate Guesswork: Replace gut feeling and unverified tactics with empirical evidence. • Anchor in Research: Base technical and content strategies on official documentation and search science. • Drive Predictable Results: Ensure algorithm changes become manageable evolutions rather than surprises.
  25. And Rely Less on Blog Advice THE PROBLEM Correlation vs.

    Causation Most SEO guides rely on observational correlation rather than mechanical understanding. When search engines update, trial-and-error tactics break down. THE ADVANTAGE Mechanical IR Understanding Understanding Information Retrieval IR) allows you to step into the engineer's shoes and predict how search systems evaluate content mechanics under the hood.
  26. CORE PRINCIPLE When you understand the science of search ,

    algorithm updates become less stressful surprises – and more predictable evolutions .
  27. "The IR world was talking about BERT before BERT was

    a ‘thing’ to the SEO world." IR research is always ahead of the SEO curve. 01 / ACADEMIC IR PAPER 02 / CULTURAL CONTEXT 03 / SEO ADOPTION 32
  28. 33 Google Research • SIGIR 2023 Keynote 2023 — Generative

    IR from Google Engineers Pioneering research presented mostly BEFORE the SEO world was aware of Generative Information Retrieval. DeepMind July 2023 Marc Najork Early Industry Signal Distinguished Research Scientist at Google DeepMind delivering the SIGIR 2023 keynote address. Outlined the structural transition to Generative Information Retrieval ahead of mainstream adoption.
  29. 34 Recommender Systems Research Google’s Supersession-Decay Filter for Google Discover

    A dual-filter framework resolving content staleness at industrial scale 1. Dual-Mechanism Filtering Combines relational staleness (detects item supersession) with predicted traffic ratio models forecasting relevance decay. 2. Deployed in Production at Scale Applied upstream of ranking in Google Discover, serving hundreds of millions of daily active users. 3. 54.9% Reduction in Stale Content Reports User-filed staleness feedback dropped by 54.9% over a two-year deployment while significantly reducing serving costs.
  30. 35 DEPLOYMENT VS. PUBLICATION TIMELINE Paper published Aug 2026 —

    It’s been in production for 2 years already. No SEOs noticed. 2 Years Aug 2026 In Production Paper Published Fully deployed and active in Google Discover long before public research disclosure. Formal publication detailing the Supersession-Decay Filter SDF architecture.
  31. 36 RESEARCH & ARCHITECTURE SURVEY The Full Evolution of RAG

    to Agentic RAG CORE OVERVIEW Defining the Current Era of Retrieval-Augmented Generation An extensive survey detailing the full evolution of RAG to agentic RAG, which is the current era of retrieval-augmented generation.
  32. 37 AI FORENSICS & CONTENT INTEGRITY Google paper on dealing

    with AI slop from 2026 and scaled content abuse . CORE TAKEAWAY Research is ALWAYS ahead of production. Forensic mechanisms and automated detection frameworks appear in academic literature long before full production rollout across consumer platforms.
  33. But… There’s a Gap in Modern SEO Education Why surface-level

    blog advice fails in the era of modern AI search algorithms and why formal degrees aren't the only answer. Pillar 01 Pillar 02 Surface-Level Blog Advice Formal Academic Degrees Relies primarily on observational correlation and outdated heuristics rather than mechanical IR system understanding. Carries prohibitive costs and long academic timelines that struggle to mirror industry velocity. • High cost barrier for specialised search knowledge • Breaks when AI algorithms change retrieval mechanics • • Promotes trial-and-error tactics over fundamental principles Curricula often lag behind modern AI search innovations • Not the only route to mastering engineering-level SEO • Lacks deep technical and algorithmic rigour
  34. Deconstructing the Dual MSc Curriculum  MSc Digital Marketing Strategy

    Core Focus Consumer intent, brand authority metrics, content taxonomy, user experience optimisation, and multi-channel marketing attribution.  The SEO Takeaway Aligns search strategy with real business objectives, revenue impact, and user satisfaction metrics.  MSc Computer Science & Data Science Core Focus Information retrieval theory, machine learning, natural language processing NLP, vector embeddings, clustering, and algorithmic complexity.  The SEO Takeaway Decodes how search crawlers parse, index, score, and re-rank web documents at scale.
  35. The Value Proposition DIRECT SAVINGS Democratising Advanced Search Science TRADITIONAL

    ACADEMIA £25k+ Tuition Pounds Saved Equivalent value of a formal master's degree delivered through targeted self-study. Formal university master's programmes offer incredible depth in computer science and digital strategy, but 80% of the academic coursework is non-essential for practical SEO strategy. THE SMART PATH By focusing specifically on relevant research papers, patent analysis, and core algorithms, you can gain an informed competitive edge at zero tuition cost.
  36. 42 Before We Begin MINDSET SHIFT We have to think

    a bit differently. CRITICAL THINKING ‘Question Everything’ Rigorously evaluate underlying assumptions and deconstruct concepts rather than accepting statements at face value. Moving beyond simple memorisation toward active, high-order strategic synthesis. Core Objective Not just remembering facts — but understanding how and where to apply them. STRATEGIC THINKING Employ Strategic Application Connect distinct domains and determine precise execution strategies for optimal real-world impact.
  37. Master Level Thinking - Bloom’s Taxonomy COGNITIVE DEPTH Higher-Order Cognitive

    Synthesis Goes beyond remembering information to critically synthesising and understanding where application of the information would be most appropriate and optimal. Critical Synthesis Evaluating core principles, deconstructing complex ideas, and connecting distinct domains. Optimal Execution Knowing precisely when, where, and how to apply insights for maximum strategic impact. 43
  38. Consider Too: The Hierarchy of Bull**** CRITICAL EVALUATION Ensure Your

    Source is Credible & Reliable Avoiding GuesSEO whenever possible by grounding analysis in verified, high-rigour data sources. Core Principles • Prioritise peer-reviewed academic research • Rely on reputable organizations & primary data • Filter out unverified speculation and SEO noise 44
  39. 4−Phase Self-Study Roadmap Phase 1 Phase 2 Phase 3 Phase

    4 IR Basics Inverted Index, BM25, Tokenisation Patents & Research Google IR papers, arXiv preprints ML & NLP Embeddings, BERT, Vector Databases Generative IR LLM RAG & AI Overviews
  40. Get the Computer Science Technical Foundations for Free HTML →

    CSS → JavaScript → PHP Building basic code literacy is invaluable, even in a low-code/no-code world.  Academic Foundations (Harvard CS50)  Practical Web Dev (W3Schools) Comprehensive introduction to computer science & programming principles Interactive tutorials & hands-on code exercises Key Features & Workflow: • Live interactive code sandbox for instant feedback • Step-by-step guides for HTML, CSS, JS, and PHP • Track progress with streaks, quizzes, and certificates 47
  41. Gamify Ongoing CS & Data Science Learning Sustain momentum with

    bite-sized challenges, interactive problem solving, streaks & XP  Hyperskill (JetBrains)  Sololearn Key Highlights & Gamification: Key Highlights & Gamification: Project-based deep learning with hands-on practice • • • Build real-world portfolio applications step-by-step Knowledge map tracks topic mastery and progress Daily streaks, gems, and lives system to build habit Bite-sized mobile exercises & community leaderboards • • • Short 5-minute bite-sized lessons accessible anywhere XP rankings, league tables, and global leaderboards Peer code bits, social profile badges, and daily targets 48
  42. Phase 1: Must-Read IR Literature "Introduction to Information Retrieval" Manning,

    Raghavan, Schütze) The definitive free textbook from Stanford University covering foundational indexing and scoring algorithms. The Original PageRank Paper Page & Brin, 1998 Understanding how graph theory and link probability matrices revolutionised search authority. "The Probabilistic Relevance Framework: BM25 and Beyond" Learn why BM25 remains the primary first-stage ranker across search engines today.
  43. Phase 2: Decoding Search Patents Phrase-Based Indexing Patent US7536408B2 Learn

    how search engines identify related ngram phrases and predict document quality without keyword density. User Node Graph & EEAT Patents Discover how author authority and entity node relationships are calculated across document networks. Query Revision & Expansion Patents Understand how user queries are rewritten using synonym matrices before hitting the index.
  44. Phase 3: Applied ML & Vectors Word2Vec & Sentence Transformers

    Learn how continuous bag-of-words models convert text into numerical vectors. Cosine Similarity & Distance Metrics Understand how machines measure semantic similarity mathematically $cos(\theta) = \frac{A \cdot B\|A\| \|B\|$. Vector DB Foundations FAISS, Qdrant) Explore how nearest neighbour algorithms perform ultra-fast retrieval at web scale.
  45. Phase 4: Generative IR & LLMs Retrieval-Augmented Generation RAG How

    search engines combine document retrieval with LLMs to generate direct answers. Groundedness & Citation Models How AI search engines verify facts against indexed web sources to minimise hallucination. Information Density & Passages Why structured, passage-level answer blocks win placement in AI overviews.
  46. Free & Open-Source Toolkit Google Scholar & arXiv Set alerts

    for keywords like "information retrieval", "neural ranking", and "passage retrieval". Build your own library on Scholar. Python NLP Libraries Use SpaCy, NLTK, and HuggingFace Transformers to run vector similarity scripts on your content. Google Patents Search Track patent filings from top search scientists (e.g., Pandu Nayak, Anna Patterson).
  47. Open Access Education - Totally Free  Research Papers Full

    open access to research articles and publications  Conference Proceedings on IR Complete access to SIGIR proceedings and IR workshops 54
  48. Mostly Free Certification on Generative AI & LLMs Via Google

    Skills Boost & Microsoft Learn ★ Featured Platforms Google Skills Free micro-learning courses covering Generative AI fundamentals, LLM architecture, Prompt Design, and Responsible AI. Microsoft Learn Comprehensive learning paths for Azure AI, Copilot integration, and foundational machine learning credentials. 55
  49. 56 Search Result Diversification - Low Cost Books and e-books

    Explores why search result diversification is important and how search engines deal with serving the right content to the right person at the right time. Includes late-stage re-ranking after initial shortlisting of search results. An easy-to-consume short book. ★ Recommended.
  50. Build a Google Scholar Strategy  Custom Library Build your

    own Google Scholar library.  Researcher Tracking Set up follow alerts on Google researchers 57
  51. 58 Regular Periodicals on Information Retrieval and Trends Journal 

    Foundations Key periodical publications & issue archives Directions & Trends  Research Subject areas, affiliations, and top cited authors Subject Area Analytics Track high-impact topics including ranking models, neural networks, and query processing. Author & Institutional Leaderboard Monitor contributions from leading industry labs and academic research centers.
  52. Dig Into the Current Key IR Research Themes Identify core

    topics from thousands of accepted paper titles across ECIR, SIGIR, CHIIR, and WSDM 59
  53. IR to SEO Knowledge Transfer: Entity & Knowledge Graphs and

    Vector Link Architecture ACADEMIC CONCEPTS SEO STRATEGY ACTION Named Entity Recognition (NER) Entity Disambiguation Extracting discrete real-world objects and building subject-predicate-object triples (e.g., Brand Offers Service]). Use clean Schema.org markup, consistent Wikipedia/Wikidata alignments, and internal link co-occurrences to establish definitive brand node authority. Semantic Vector Proximity Contextual Inbound Hubs Pages that share mathematical vector space pass stronger contextual relevance than topically distant documents. Structure internal links using semantically aligned anchor variations to reinforce document topic clusters without keyword stuffing.
  54. 61 Pro Tips   Focus on Concepts Smart Reading

    Strategy Donʼt worry about the maths side of things — the concepts are what matters most. Read the abstract, intro and summary to papers only most of the time.
  55. Get Started - Your 30−Day Action Plan WEEK 01 WEEK

    02 WEEK 03 WEEK 04 Foundational Reading Patent Analysis Script Development Semantic Audit Read Chapters 16 of Stanford's free "Introduction to IR" book. Analyse 3 foundational search patents on Google Patents Search. Build a basic Python script using SentenceTransformers to measure content overlap. Audit your site's core pages for entity clarity and semantic depth. nlp.stanford.edu/IR-book/
  56. 64 MICRO-LEARNING Answer… A bit at a time. Make learning

    fun, a habit, and gamify where possible. • Streak Chasing: Get chasing streaks, for example.
  57. The Mindset of a Research-Backed SEO "Stop asking, 'What works

    in SEO?' and start asking, 'How would an information retrieval engineer solve this computational challenge at scale?'"
  58. Thank you… Please don’t eat an elephant • X @dawnieando

    • LinkedIn: msdawnanderson • Threads: @dawnieando • BlueSky: @dawnieando • Website: bertey.com