Dawn Anderson I Mythen & Realität um AI Search

Shownotes

GEO, AEO, AI-SEO – die Suchwelt hat gerade mehr Buzzwords als Antworten. Also holen wir uns jemanden, der seit fast zwei Jahrzehnten zwischen akademischer NLP-Forschung und echter SEO-Praxis unterwegs ist: Dawn Anderson, Founder & Managing Director von Bertey (Manchester, UK). In dieser Folge räumen wir mit einem der hartnäckigsten Mythen im AI-Search-Umfeld auf: dem Konzept "Information Gain" – ursprünglich ein Prinzip aus der Klassifikationsmodellierung des maschinellen Lernens, das in der SEO-Szene zum Heiligen Gral umgedeutet wurde, obwohl es damit fachlich kaum etwas zu tun hat. Außerdem sprechen wir über: → Warum Chunking keine SEO-Taktik, sondern eine reine LLM-Notwendigkeit ist → Wie BERT den Weg zu modularem, semantischem und heute agentischem RAG geebnet hat → Warum RAG Halluzinationen reduziert – und wie Retrieval-Augmented Generation technisch funktioniert → Wieso menschliche Erfahrung der einzige Faktor bleibt, den kein LLM ersetzen kann → Was Google-Quality-Rater-Guidelines wirklich über "Qualität" verraten

Diese Folge ist Teil unseres SMX-Specials zur SMX Advanced Europe 2026 in Berlin (29. September – 1. Oktober), wo Dawn Anderson auf der Bühne steht – eine Kooperation von MARKETING MEETS TECHNOLOGY mit der SMX Advanced.

Mehr zu Dawn Anderson: https://bertey.com/ https://www.linkedin.com/in/msdawnanderson/

Mehr zur SMX Advanced - 29.09-01.10.2026 - Berlin Drei Tage für fortgeschrittene Search-Marketer. Zwei Tage Deep-Dive-Sessions zu SEO, PPC, GEO & AI, plus ein dedizierter Deep Dive Day für praxisnahes Lernen direkt mit den Experten. Strategie trifft Umsetzung – man geht nicht nur inspiriert, sondern ausgerüstet. https://smxadvanced.eu/ https://www.linkedin.com/showcase/smxadvanced/ https://x.com/eu_smx https://facebook.com/smxadvancedeu

🎙️ MARKETING MEETS TECHNOLOGY – der Podcast von WOXOW über Marketing, SEO, GEO & AI.

Mehr zu WOXOW: WOXOW - Technologieberatung für SEO, AI & Data: https://woxow.com/ Du möchtest noch mehr Info´s zu SEO, AI & Data? Schau doch mal in unserer Know How Sammlung vorbei: https://woxow.com/know-how/ Timon Hartung: https://www.linkedin.com/in/timonhartung/ Johanna Hartung: https://www.linkedin.com/in/johanna-hartung/ YouTube Channel: https://link.woxow.com/youtube Instagram: https://www.instagram.com/woxow_com

Timestamps:

02:04 – Mythen & Realitäten im AI-Search: GEO vs. SEO vs. AEO 04:35 – Ist Chunking eine valide SEO-Methode? 07:54 – Warum Content-Manipulation wie Spam wirkt (und bestraft wird) 09:26 – Parallelen zu Google Panda & Penguin 10:15 – Das Missverständnis um Em-Dashes als "AI-Signal" 12:29 – Google, AI-Content & die Qualitätsfrage 13:19 – Commodity- vs. Non-Commodity-Content 15:15 – Programmatic SEO & die Balance auf Site-Ebene 16:40 – Was ist Qualität aus Google-Sicht wirklich? 20:35 – Listicles: Warum sie (noch) funktionieren 21:20 – Wie IR-Forschung und SEO-Welt aufeinandertreffen 23:56 – Ohne 20 Jahre Spam-Bekämpfung wären LLMs nicht nutzbar 25:49 – Von BERT zu RAG: Ein Blick in die IR-Forschung 26:12 – Was ist BERT? (Bidirectional Encoder Representations from Transformers) 28:47 – ChatGPT als Weckruf für Google (Defcon-1-Modus) 30:58 – AI Overviews, AI Mode & die hybride Zukunft der Suche 31:38 – Agentic Search: Tim Berners-Lees 25 Jahre alte Vision 34:07 – Was wirklich zählt in der AI-Search-Ära 34:46 – Freshness, RAG-Hunger & der "Lost in the Middle"-Effekt 36:45 – Der menschliche Faktor als Non-Commodity 37:37 – Warum SEO-Konferenzen weiterhin wertvoll bleiben 39:47 – Wie Fehlinformationen sich durch KI-Wiederholung verstärken 39:56 – Fallbeispiel "Information Gain": Ein SEO-Mythos entlarvt 44:07 – RAG einfach erklärt: Retrieval Augmented Generation 44:20 – Halluzinationen: Warum LLMs "überzeugt falsch" liegen 47:04 – Von Naive RAG zu Modular RAG 48:05 – GraphRAG & Wissensgraphen 49:06 – Agentic RAG & die Rolle von Schema Markup 50:47 – Wo man Dawn trifft: SMX Berlin, Antalya & mehr 51:35 – Wie man mit Dawn in Kontakt bleibt (LinkedIn, X, Blue

Transkript anzeigen

00:00:00: Certainly SEO is not dead and GEO is the be all an end-all.

00:00:04: Do Not Crawl in The Dust, if you see things like less and less crawling or more and more pages showing up on things like Discovered Not Crawled or Crawled Not Index, Discovered No Index and so forth that's probably a sign to look at overall quality of what your producing.

00:00:20: In LLMs & Rags there was this recency.

00:00:28: Welcome to a new episode of Marketing Meets Technology.

00:00:37: Here you hear insights into digital marketing strategies and technologies.

00:00:46: Have fun with this episode!

00:00:53: is the founder and director of Bertie, an international SEO consultancy based in Manchester UK.

00:00:59: And she spent nearly two decades working at digital.

00:01:02: what makes Dawn really special though?

00:01:04: Is her rare gift of taking a dense academic research something like Google BERT and natural language processing which has all you know all the rage right now and neural ranking models and turning them into practical strategy that real teams can actually act on not just keeping it theoretical, but really making an act on.

00:01:25: And this is something that makes a really special and if you've been to popcorn Brighton SEO SMX or Mosconn You certainly seen her on stage.

00:01:33: she's brilliant.

00:01:33: She's very generous with the knowledge what we're going to profit from in these podcasts as well And she's always a few steps ahead where search is heading.

00:01:40: She just told me, she was already also studying computer science to be even more ahead and understand how AI really works and how rankings work.

00:01:49: so that's amazing!

00:01:50: Don it's an absolute honour for you here.

00:01:52: welcome

00:01:53: Thankyou thankyoufor the invitation.

00:01:56: I'm looking forward too having this chat about being at SMX as well.

00:02:03: Yeah, SMX is obviously amazing.

00:02:04: but you were already at SMX New York.

00:02:06: You just told me and did the keynote there right?

00:02:10: Well I did the key note back in June.

00:02:14: that was really good.

00:02:16: people are lovely of course lots a lot things to learn.

00:02:19: so yeah imagine that SMX advanced in Europe will be exactly same!

00:02:24: It's always it's advanced.

00:02:28: get to meet all these amazing speakers super advanced.

00:02:32: We'll talk about SMX a little bit later, and then what the advantages are in everything.

00:02:37: But I really wanted to start off with the topic of myths and realities of AI search because i think that is something you can shed more light on Because we've done so much research.

00:02:46: And um yeah You're you're really deep into the topic.

00:02:49: Yeah...yeah..I mean I kind of feel like maybe now things aren't starting to settle down A Little Bit but obviously This past year has been a wild ride for misinformation in the SEO space.

00:03:04: You know, with the whole like is it Geo?

00:03:06: Is it SEO?

00:03:07: or is AEO?

00:03:07: Is EIEIO whatever you want to call it?

00:03:10: What do YOU call

00:03:11: it?!

00:03:12: Well I can call it...I just call it AI Search and SEO still!

00:03:19: Obviously there are more layers appearing in our space and that's a wonderful thing actually because it stops us getting bored, you know.

00:03:27: It keeps them on their toes.

00:03:31: but certainly SEO is not dead and GEO is the be all an end-all.

00:03:39: But what I have seen?

00:03:41: unfortunately whenever there's a void of knowledge instead we're really just learning.

00:03:48: as we go along with this I feel like people in the SEO space or any place where whether it's competitive, they have to fill the void with something.

00:04:00: Almost like guess-io because it is not paid search and somebody has turned us into rules.

00:04:10: there are some of their works but others don't.

00:04:14: so Google will tell you paid search.

00:04:19: Whereas with SEO it's very much.

00:04:23: we're having to fill in gaps, Google won't illustrate how exactly how to do SEO but they will give us guidance and so on and so forth And Search Engine is obviously increasingly providing those with more tools now.

00:04:35: But what's happening here?

00:04:36: of course people are still filling a lot of gaps.

00:04:41: I would say as i said starting to settle down because We don't necessarily need to do... we don't need to break our content up into tiny, tiny pieces for search engines to understand it.

00:04:54: In fact Google has said I'm gonna come onto the subject as we know of chunking and so on but that was a really obvious one where people were jumping on things thinking they needed to immediately break their contents up in to tiny pieces so LLMs could understand it.

00:05:10: So there's just, as I say a lot of misinformation that has been flying around.

00:05:14: I would say in particular this past year.

00:05:17: but starting to settle a little bit now you know...

00:05:20: So what do you say?

00:05:21: that chunking is valid method or?

00:05:26: It's not a method that SEO have to do it?

00:05:30: simple as that!

00:05:35: To a large extent, it's been a huge part of large language model understanding for many years because the notion that you want to take context as well just words on page into consideration.

00:05:56: You have to decide how big is this window going?

00:06:03: A full sentence around a word as context?

00:06:06: Or are you going to take just a few words on one side and a few towards the left?

00:06:09: because we know nowadays, Context is bi-directional.

00:06:13: It looks above before or after words.

00:06:15: so that's the whole point of natural language understanding.

00:06:19: Are you gonna take into consideration The full page of context?

00:06:23: Because search engines and large language model trainers can do it now.

00:06:28: the bigger the context window, the more expensive it is because for instance attention works.

00:06:34: The context of a word has to look at that context in relation with every other word within the context so becomes extremely computationally expensive when you're looking at each and everything like this matrix.

00:06:56: So there's absolutely no point in SEOs trying to guess how big that context window is, the search engines using.

00:07:06: There are also about twelve different types of chunking as well.

00:07:09: so which one is Search Engine Using?

00:07:11: We really don't know.

00:07:12: they're not going tell us!

00:07:15: And what will happen if you try and manipulate these things?

00:07:19: it a little bit like keyword stuffing or some other... underhand approaches that have been implemented by our CEOs over the years to try and get an advantage in natural language.

00:07:32: What happens is at some point another search engine will work it out themselves, what's best balance for them to chunk things on their side?

00:07:41: And if you've manipulated stuff... You're gonna stick out like a sore thumb!

00:07:47: The whole point of search is that search engines are trying to emulate humans naturally searching.

00:07:54: So if you try and manipulate things to favour yourself with a machine, as that machine tries to become more human-like in its methodology... You just stand out obviously when somebody's trying to manipulate machines.

00:08:09: so I can see people can utilise like scripts and so on, to try build chunking approaches.

00:08:20: But there's not really a huge deal of point other than like test these things?

00:08:25: To have...to tinker around with stuff!

00:08:28: And I for one would actively encourage that.

00:08:31: we're just trying see if this thing works.

00:08:33: but it doesn't necessarily mean you have implement in your search strategy because It does make stand out as almost like a spammer.

00:08:41: But unfortunately, the problem is with a lot of these approaches now as we're testing new areas of AI search.

00:08:48: A lot of the approach do look like spam that are being implemented by people.

00:08:54: and you know?

00:08:55: We've seen lots of people build automated systems to just churn out AI content at scale And keep throwing it on schedule.

00:09:07: no human intervention churning out loads and loads of... I mean there's nothing wrong with translating content in my mind, with AI.

00:09:18: As long as then you get a human to double check it yeah?

00:09:21: Guardrails or creating content with AI to some extent And again.

00:09:27: You know we can utilise these things for efficiencies and planning an ideation But its about making sure the quality is high enough and making sure that your have like human oversight To keep things within a threshold of acceptance.

00:09:42: See what's tending to happen and I see that there are people calling it AI, MountAI where AI search or generated content is just done at scale.

00:09:53: using AI goes up and up until he doesn't.

00:09:57: so for those who've been around quite awhile the days The summer of two thousand eleven twelve, they were decimating times for people's search strategies under them penguin and I can see that.

00:10:14: That's gonna happen again But obviously maybe it's a bit harder for certain engines to work out What is AI search?

00:10:22: And what isn't because even some of the humanizers are up there.

00:10:25: you to like people utilizing Humanizer tools To make things look more or more natural but in reality these adjust.

00:10:33: These are just article spinners on steroids and with AI assisting it.

00:10:39: So at some point, there will be some patterns that we'll start to emerge as search engines will pick up on And we see them the obvious ones already The Zeds everywhere in the M dashes and so forth.

00:10:51: Not that there's anything wrong using an m-dash It is obviously a part of grammar But people are now petrified to use an M Dash.

00:10:59: Yeah absolutely

00:11:00: It's gone the opposite way.

00:11:01: We're kind of throwing away good grammar because M dashes look like we are doing things with AI search.

00:11:07: So, in reality there is just a lot people guessing at things and trying to fill gaps.

00:11:17: also at the same time We were getting a lot senior leadership teams asking why aren't you appearing here?

00:11:25: I think that they have to say always because of X, Y and Z. or let's do this.

00:11:33: And fill a gap with the strategy even though that strategy may not be the right approach just because they're really scared saying we don't know yet but lets work it out to appreciate.

00:11:44: there is still very small part of search overall.

00:11:48: so its disproportionate fear on the SLT side of things senior leadership thats maybe driving people Myths and misinformation which obviously just get repeated And this becomes a circle of more misinformation.

00:12:05: Yeah, oh

00:12:05: yeah Of course now that's how it works.

00:12:08: That's all.

00:12:08: the myth of SEO is being kept alive and also like how I see us at some point then will be discredited again as being kind.

00:12:21: It's not good because I think we went through a period where we're actually starting to become increasingly credible and then AI searches come along.

00:12:30: And it just kind of blown us a little bit out the water, really in that regard yeah but you know will survive as an industry?

00:12:38: Oh!

00:12:38: We always have.

00:12:39: We've survived every update so we'll survive this as well.

00:12:43: But what i find interesting is how do you approach content right now?

00:12:49: And Google has also changed their basically how they, They say the rate content.

00:12:55: So in the beginning I said it's fine as long as good content.

00:12:58: we don't care if its done with AI or not.

00:13:00: and now obviously there are really attacking AI content as well.

00:13:04: so If It is Not The Highest Quality Its Definitely Gone.

00:13:09: As You Just Said Most Of The AI Articles Are Spun Articles Again.

00:13:15: But How Do you?

00:13:17: Yeah.

00:13:17: How do you see the whole topic on content creation?

00:13:22: Well, I mean... Google has said that it's not the AI but is the quality of the content.

00:13:31: The problem is people are pretending to just presume we don't need humans and feed them all into an LLM Germany or something like that.

00:13:43: whatever and it will just say we want a lot of articles produce it, then they get somebody to publish that.

00:13:50: That's the problem.

00:13:51: there is no consideration for stuff you are producing really.

00:13:59: so Google have talked about this notion of commodity or non-commodity content as well.

00:14:04: The commodity content is stuff anybody can produce.

00:14:09: you know, in reality it doesn't require human thought.

00:14:13: You could literally just... It takes a few minutes to put something together or an LLM can generate very easily and some extent I think that there is on the site.

00:14:25: so.. There's still need for commodity content because if somebody regularly comes through our website they may want reference materials.

00:14:33: So if its SAS for instance The SAS space has much information get her driven in that.

00:14:43: A lot of it is technological, needs a lot of explanations and so forth.

00:14:47: So reference sections on SAS sites are always really useful like a glossary.

00:14:53: It's useful particularly in certain verticals But that doesn't mean for instance you can have just that one his own.

00:14:59: I think its about making sure your balance off every different type that is relevant to your audience met, but too a high standard if you like.

00:15:10: AI search for instance... If we look at some of the travel sites out there or financial sites such as stock market sites they utilize AI really well.

00:15:23: They'll do.

00:15:24: summarizations, like a weather site may summarize the whether but it also augmented with other human added value or database driven aspects.

00:15:34: And I think there's nothing wrong using AI because you can use incredibly creative content and at the same time have to have that balance.

00:15:46: everything has always been an overall side pie.

00:15:49: so years ago people could generate via PHP, just a load of automatically generated template sites.

00:15:58: You know like programmatic SEO yeah which again in itself is nothing wrong with it.

00:16:03: but when you're doing and its just thin content or you get a database and chuck something out to get indexed on ranks and traffic that's the problem.

00:16:11: so its the intent behind becomes spam versus adding value.

00:16:16: So I think what this says about was the intent.

00:16:19: What's the value overall of this site and everything that it offers?

00:16:25: Is there a balance between utilising things for efficiencies like databases in the background, yeah because there is no point having humans sat there creating things they could be created via databases.

00:16:37: when its data driven sites.

00:16:41: Such as you know Stock market prices, you wouldn't have somebody up sat there.

00:16:45: a human updating all that.

00:16:46: You'd have it driven by a database.

00:16:49: He would necessarily happy humans at their summarizing get.

00:16:51: you may have some insights But then he might have really good an extensive article giving opinion From a human.

00:16:59: yeah So its the balance is to balance of everything and the overall site itself.

00:17:04: I think

00:17:05: Interesting like they're the part of the balance.

00:17:07: And i think It's also with Google as same.

00:17:10: if you have one bad article on your website, and you have thousands of great ones.

00:17:15: It's not a problem

00:17:16: one good

00:17:17: one in the thousand bad ones.

00:17:19: that is a

00:17:19: problem

00:17:20: exactly.

00:17:21: That says just to general what is the overall quality?

00:17:25: Of this presence

00:17:26: here?

00:17:27: How do you think they measure quality at Google?

00:17:31: Do You have any insight into that?

00:17:33: Well I mean we have The Quality Rates Guide And i Think Quality Is Determined By Well a few things, obviously they have the quality rate themselves and for me as well the core updates are an indication of quality to a large extent but quality is relevant.

00:17:54: So I know that Martin split spoke recently about what core update it's and i've read quite a lot of stuff in forums and updating quality, making sure that the human raters are in place to ensure what people are getting returned from search results aligned fully with what they were searching for.

00:18:22: Because relevance changes... For instance if somebody's looking for Apple or Apple then when a phone comes out of the fifteen or whatever sixty one number it is now you should show them most relevant ones.

00:18:38: But when somebody types in Apple generally, what do they want?

00:18:42: The intent will shift and to some extent quality is about making sure in search that you're returning What matters To the searcher at that particular moment of time.

00:18:51: Yeah So that's quality.

00:18:53: churning out loads of AI stuff with a intent on just spamming That's not quality.

00:18:59: yeah That's obvious.

00:19:01: I mean This years old but there was a notion of do not crawl in the dust and if you see things like less, unless and less crawling or more and more pages showing up on things like discovered not crawled or crawled not indexed discover not index so forth.

00:19:17: That's probably a sign that you need to look at the overall quality of what your producing.

00:19:22: but quality is meeting high...meeting intent in the moment.

00:19:26: yeah And I know that, for instance with core updates they gather all the click data.

00:19:32: It's not that they are ranking things based on clicks it just they rank... They're analysing the click-data That they build over time to understand what is quality and as such What is relevance if that makes sense between these update?

00:19:47: They utilise this thing called Normalised Discounted Cumulative Game.

00:19:53: They then push it all out, quality rate is human rates.

00:19:56: I think they also have about some element of machine quality rating going on as well because again i've seen a lot in the eye aspects.

00:20:06: but there will always be humans in the mix yeah?

00:20:09: The calling Human In The Loop.

00:20:11: Also Quality Is Flagged By People Saying This Is A Rubbish Result Or Whatever.

00:20:18: So Quality Is Very Objective.

00:20:23: We know ourselves what poor quality looks like.

00:20:26: it's spam and that kind of thing.

00:20:28: Yeah, but I think search engines basically build up a mathematical picture What he looks like?

00:20:34: And then push that figure out across the whole research algorithmic.

00:20:41: Oh Very interesting

00:20:42: yeah,

00:20:43: yeah mean That's why we've always suspected that they do that They at least measure.

00:20:46: you know who clicks who searches something clicks on the page comes back clicks on The next result as a strong indicator to show if that was quality or not.

00:20:55: Or at least, what the user wanted to find?

00:20:59: Yeah it's more around understanding the intent of the query What people were looking for in particular time.

00:21:07: So yeah It is an interesting fascinating field.

00:21:15: Let me ask you one last myth.

00:21:17: and so what do think about listicles?

00:21:22: I think that they work, but it's sad.

00:21:25: Google or the search engines are whoever the LLMs need to get a grip of them because what is happening sometimes?

00:21:37: people create these self-serving listicles and don't even appear in their own listicle in Search.

00:21:42: so i feel like this sign at least Search Engines and Google in particular trying to combat it.

00:21:49: I think part of challenge LLMs and AI search, if you like.

00:21:55: It's always been natural.

00:21:56: language processing has almost been a very separate part of information retrieval not in the same campus traditional search.

00:22:07: that is what I have seen looking at the IR community.

00:22:11: If LLM had worked on their own to start without need for search engines we probably would be really worried about future SEO and generally search now Seriously, but they didn't because the two driven by probability determination just guessing next word guesses.

00:22:29: Yeah so they need this reinforcement which as we know is rag and web search.

00:22:35: They're obviously always out of date as well because you can't keep retraining a large-language model so that it's up to date with everything, including the last few days in news.

00:22:44: We know for instance that fifteen percent search queries every day are new and they largely driven by real world events never happened before So our large language models cannot cope with that.

00:22:57: What has happened is... To some extent You've got these two communities who have been building things.

00:23:02: No, all of a sudden they have to rely on each other.

00:23:05: But what that means is in my mind web search has very much been... ...on top of adversarial search for a long time.

00:23:13: so the spam yeah?

00:23:14: So traditional search was getting quite good at defeating spam.

00:23:18: it's rare you saw stuff but we've got this whole community here who didn't probably even consider spam and also just discovered the SEO world.

00:23:27: Yeah!

00:23:29: I think there were like.

00:23:31: They're having to learn about adversarial tactics in the LLM world, yeah?

00:23:37: Obviously as they collaborate more and more between web search and LLMs.

00:23:42: And Google are obviously bringing more and More LLm stuff by a Gemini and so forth.

00:23:47: an AI overviews an AI mode into Search.

00:23:51: I think over time there will get group of it but It does still work for now.

00:23:55: But its saddens me and i Think if you can hang on or not do some of these cheap and nasty tactics.

00:24:03: I think in the long term you'd be rewarded, but it's very hard to say to a senior leadership team we don't want to do these things because they're seen competitors do stuff and get an advantage.

00:24:16: this is sad kind of draws us all back towards spam which we shouldn't be having to do.

00:24:21: cause search engines like Google and Bing and so forth are saying on the one hand don't do these stay on the good path, but then we're seeing that these dodgy techniques work.

00:24:36: So sad or true at the minute and hoping they'll get a grip of it as soon as possible

00:24:43: without the twenty years of spam protection on spam avoidance that Google has built

00:24:48: exactly

00:24:50: not be usable.

00:24:51: well funnily enough I did talk SEO week this year last The IR world and SEO being frenemies forever, looking at how years ago the IR World had a whole track of one of their conferences on adversarial search which talked about how you know web marketing and SEO was just not aligned with the world of IR.

00:25:20: We want different things.

00:25:21: obviously we wanna rank they wanna like make sure that field in there that was literally designed for adversarial search.

00:25:36: And over the years, that track disappeared because I think they mostly had it sold.

00:25:42: and then all of a sudden LLMs have come... In your looking at IR papers you'll see more stuff around adversarial LLM's basically like defeating an enemy which is spam.

00:25:59: So at some point they'll catch up, I think a lot to what I just said as rag develops because rags gone through its own evolution since twenty-twenty where it would start off with naive rag came advanced then modular.

00:26:11: Then graphrag and now agentic rag And that's really followed the path from being like basic search To moving forward including things like semantic rag and knowledge graphs in.

00:26:26: There's a catch up, definitely as they work together those two worlds in my perspective.

00:26:31: As an outside SEO interloping and looking at yeah

00:26:36: absolutely um.

00:26:37: maybe take us deeper into the AI world?

00:26:40: And my question would be like you've done a lot of research on BERT Maybe explain what BERT is.

00:26:46: then take us down the road to RAG.

00:26:49: What do think has happened in the IR research?

00:26:54: How do we get from keyword matching to embeddings, phrase retrievals, neural rankings.

00:26:59: Yeah yeah well I mean obviously we know Bert is quite old.

00:27:03: now it's two thousand and eighteen i think.

00:27:06: the paper came out and then he was the bi-directional.

00:27:09: oh god Bert what a bi- directional.

00:27:13: okay remember what he's done.

00:27:14: so far thing is.

00:27:16: we just call it Bert don't something from Transformers.

00:27:22: I should know, but anyway, bidirectional recommender systems now?

00:27:28: No it's called bi-directional encoder representations.

00:27:32: Yeah, sorry my mind went blank.

00:27:33: But anyway it's quite old.

00:27:35: now is two thousand eighteen.

00:27:36: the paper came out as massive by Google and then they up sourced there.

00:27:40: And then but google introduced a form of births in Two thousand nineteen into search.

00:27:45: But he was very very basic because I said The problem with birds on all these methodologies Is this?

00:27:51: Very computationally expensive too expensive really at the time for production search.

00:27:56: Then if you went along lots and lots of evolutions Leena methods and so on.

00:28:02: And so forth came along Facebook, all of these other big companies got involved.

00:28:08: I think they were creating things like Roberta.

00:28:11: there was a Leena version.

00:28:13: There are now many different types of birds as well.

00:28:18: i'm just doing my master's dissertation and utilising a model called FinBerts which analyses sentiment analysis on financial text to isolate just the sentiment of a sentence when there are multiple entities mentioned together, yeah?

00:28:40: To split out the sentiment.

00:28:41: So it's like those kind of things that is being used more.

00:28:45: but obviously as has all gone on lots have happened in the IR world since then.

00:28:53: you know there's a lot more stuff that has gone on around revisit to nearest neighbours, semantic technologies where they utilise the content on adjoining pages.

00:29:04: To understand the separate pages and so forth.

00:29:06: whereas historical it was always just the content in the page now its even pages are connected.

00:29:13: hence why things like you know internal links are increasingly important because of the topic that's connected to each other clusters, they're increasingly important and obviously then LLMs came out.

00:29:28: The natural language processing world started to bring a chat.

00:29:32: GPT was massive yeah threatening search.

00:29:36: I think that really sent searching Google into Defconn one mode Because there were fairly fragmented in Google Brain, then they had the Deep Mind set up and all these different competing research departments within Alphabet.

00:29:53: And if I'm not mistaken when chat GPT started talking about how it wanted to challenge for search in twenty-twenty two, Google merged Deep Mind & Brain together so that they have this almost phenomenal research offering.

00:30:10: They're just chipping and chipping away, really now.

00:30:17: And obviously they have the massive advantage in that... ...they pull everything into the search interface itself whereas chat and GPT is standalone still.

00:30:26: So they do have a big advantage and ultimately will win because you've got data, research team.. ..the computer, historical stuff.

00:30:36: they've got the community, I was looking for Lin off the other day.

00:30:40: They still have about a ninety-six percent of search share globally.

00:30:44: so you know that all their advantages really...they hold accounts.

00:30:51: but then what's happened is tried.

00:30:54: GPTs come along and as i say AI overviews are starting to be searched.

00:31:00: Google has gotten rattled And obviously we've had uproar in the search base because it was all a bit rubbish and now is getting better.

00:31:12: It will continue to do so, I think they'll always be like a hybrid area with natural search.

00:31:20: still there's still a place for people who want to go on search themselves Because sometimes i like going off having a route around digging down rabbit holes.

00:31:30: Some people just want their answers spat out to them and search become generally more accurate over time because these things always improve.

00:31:39: And RAG, obviously has been a massive part of that since it was introduced by the Facebook team in early twenties I'd say is very simple.

00:31:48: just to reinforcement like go and check.

00:31:51: this keyword means this augment whatever is retrieved in LLM results.

00:32:00: as such they don't rank anything and then over time it's just improved.

00:32:08: And now the future is very much agentic search.

00:32:11: Funnily enough, I did... The talk that I did at SMX Advanced looked at a point where we're right now which was twenty-five years on from the birth of semantic search when Tim Berners-Lee did his seminal paper about semantic search and he said He had vision to brother or sister who were talking how they were looking for some medicine for their mum there, and they were both looking online.

00:32:40: They would get in their agent to like go on research... ...they will getting the agent to coordinate clocks and timers all sorts with each other.. ..and then coordinating their diaries via their own agents speaking to eachother.

00:32:57: And you'd think that somebody's like Sundar Pichai who talks about agentic search.

00:33:03: Talking just the other week, he was in an interview and talked about a genetic search being the future.

00:33:08: And that's what you saw the vision as.

00:33:09: but in actual fact it is really old.

00:33:12: It's twenty five years old.

00:33:13: Yeah So its taking long time to get here Very secure to shrew A genetic search now with the likes of UCP and ACP commerce protocol and universal commerce protocol.

00:33:28: All the stuff that goes on in the background, those SEOs suddenly just get bamboozled by to some extent.

00:33:33: you can see a lot of that happening when you're looking at the IR space because they are always ahead.

00:33:38: research is always ahead of production.

00:33:42: but even with that one there were conversations going between Shopify wise think why is this involved?

00:33:48: I know strippers, PayPal all of these background organisations that are involved in the e-commerce space.

00:33:55: You do doing things to build agentic protocols together but those in SEO were like we're too busy sometimes spreading myths making stuff up and debating silly things.

00:34:11: you know i think...I don't have this in it for a while.

00:34:16: folder versus subdomain, three oh one versus three or two.

00:34:19: I mean how many years did we waste on these kind of silly debates?

00:34:23: Now it's GEO vs SEO is like i don't care.

00:34:28: let's just look ahead and see what's happening in the bigger search space than the IR world and they machine learning world and natural language worlds.

00:34:36: And then you know...and look where were headed.

00:34:39: rather than debating these my new tie things that don't really make them much from.

00:34:44: So, yeah.

00:34:44: I

00:34:48: agree but what is it then that actually moves the needle?

00:34:51: and AI search?

00:34:53: What are the tactics?

00:34:54: Well well i think still a lot of... Still doing good things in Search you've always done make massive difference because to large extent until LLMs aren't augmented with RAG most of its driven by Church anyway.

00:35:12: so most results come coming back from rag.

00:35:16: That's it, yeah?

00:35:17: So web search is an old message.

00:35:19: so just do all the good things that you always did keep improving quality adding value and so on but so forth.

00:35:27: And maybe there are some nuances like freshness matters more than I ever did with Ella because Rag now he was always hungry for the latest on things.

00:35:38: Realize that You're probably not going to get index for as long as you did previously because rag also possibly relies on something called, it's probably a bit prone to being susceptible what they call the Curse of Dimensionality which is an uploaded index.

00:36:00: It will never help Rag so make sure that they de-index things quickly and not relevant.

00:36:07: you're probably not going to stay in debt with those old pieces.

00:36:13: To make sure your updating things really regularly, freshness.

00:36:17: consider things like... In LLMs and rags there is this recency impact And it's also lost-in the middle impacts.

00:36:28: So a long piece of content will maybe have an intersection that doesn't get spotted as easily waited as much, but at the same time I think The likes of paragraph passage indexing that we've all forgotten about.

00:36:45: But We All Had a Big Who or About Years ago.

00:36:47: That probably helps with the lost in-the middle impacts because search engines now are looking Increasingly At Parts Of A Page Rather Than Full Page As Such.

00:36:59: So you know just i would say Just Keep Focusing on Producing The Content Quality You Know.

00:37:06: I think it's sad to see that people have been laid off because all the organisations feel machines can do with a job.

00:37:15: Machines help with ideation and so forth.

00:37:19: but human aspects will become prime value in future, which is non-commodity content that Google's talking about, adding that human perspective.

00:37:33: That experience aspects.

00:37:35: a search bot or an LLM can't have that because it doesn't have a concept of true human experience.

00:37:45: yeah they can pull things from other places but he can share their human experience.

00:37:50: Yeah

00:37:52: very interesting.

00:37:53: before we come to the question where I want to go deeper into reg your absolute right.

00:37:56: human experiences are definitely getting valued more.

00:38:00: And that's why I also think that like a conference, like SMX Munich or as a mix advanced in Berlin, as a mixed advance to the US there incredibly valuable because That something an AI cannot replace.

00:38:12: for instance you can always go into Claude and ask them for an AI strategy of course on AI search strategy but it will give You The medium Of everything It's read and Give you Like

00:38:24: Yeah.

00:38:24: Well

00:38:24: Something In between it Will Read Better Than Probably What most people will put out, but it'll be the same for everybody who asked that question and I would absolutely mediocre.

00:38:39: Absolutely true!

00:38:40: And also what will happen as well is... As we see now there's an awful lot of this like quite spammy content around GEO and all these misinformation.

00:38:50: just recycling yeah?

00:38:54: You're probably not going to say they are high quality conferences, yeah.

00:39:01: This is just a lot of it's AI slot when its churned and repetitive as you say.

00:39:10: You're not going to get those conversations that go on in the background based on real senior SEOs experience.

00:39:19: People who manage teams' experiences are utilising different tools or taking different approaches.

00:39:25: You're not going to get that in just search results.

00:39:32: And so, yeah and you are right it's very mediocre because the way LLMs work is they take a general consensus of everything returned.

00:39:41: Yeah So its nothing exceptional.

00:39:45: It's just an average The mean of every thing Exactly!

00:39:49: Its usually better than what normal people do when have no clue about their doing But it's way worse than an expert would do.

00:39:55: And the interesting part also is because in SEO, its very specific for SEO with this issue like a lot of industries they have clear rules and clear ways how to do something.

00:40:07: In SEO.

00:40:08: It Is A Lot Of Guesswork And Its A Lot of Experimentation And Testing So That Means Real Results Can Only Be Produced By Few People.

00:40:16: And These People Sometimes Publish This Or They Don't.

00:40:22: Random stuff, a lot of YouTubers creating YouTube content with talking about whatever SEO does or whatever Google things and the more people reproduce this.

00:40:31: The more the AI thinks This is what's actually happening?

00:40:34: Yeah where she gets

00:40:36: yeah I've got a classic example Of that there's a topic all information gained which is rife in the SEO world.

00:40:44: absolutely right but in actual fact it based on just one paper you will not get any search an IR person associating SEO with information gain at all.

00:40:56: because, Because Information Gain is actually a concept that's absolutely huge.

00:41:02: but in the machine learning world around classification modelling.

00:41:06: It's part of a classification algorithm and it so massive in the mathematical space That was created by Claude Shannon.

00:41:19: the mathematical information world is he, Claude was named after him.

00:41:25: Claude.

00:41:25: the machine learning model.

00:41:27: yeah so what it means?

00:41:29: It's basically around how deep a classification model goes in splitting content, yeah or not content.

00:41:43: Or trees decision tree because just decision trees are is as a form of classification modeling.

00:41:49: so basically information gain also kind of known in a way as entropy it looks at Is this split?

00:42:00: Pura...is it purer to this class than this one?

00:42:04: So there's nothing to do whatsoever with Adding value or whatever, yeah.

00:42:09: That's one pattern.

00:42:10: maybe somebody is taking that concept a little bit and just utilised it in some way as an as a patent creator at Google.

00:42:18: Yeah But its years ago.

00:42:20: And what happened is despite the fact actually It's to do with classification modeling and decision tree splits of purity Yeah?

00:42:29: What's happened is, the SEO world has started to write about it.

00:42:33: Now if you go and look up information gain at one time It was always about what it should be about which is decision tree splitting on purity of classification?

00:42:45: No!

00:42:46: You look at that And its filled with like SEO articles talking bout it Like Its The Holy Grail All over the shop.

00:42:53: yeah When actually adding more value think information game for that, it's just obvious.

00:43:02: If you add more value and you're going to rank higher yeah?

00:43:05: That doesn't take a concept!

00:43:08: It should literally have more value generally.

00:43:11: so maybe somebody who has created that patent as utilised that concept.

00:43:16: but the concept is predominantly ninety five percent machine learning, decision trees and classification splitting on purity.

00:43:27: But like I said it's been utilised and exploited to the point where SEO is hijacked that term yeah?

00:43:35: That's what happens!

00:43:36: It's what happened...I go and look for things when i'm doing my dissertation now and loads of stuff just have just been hijacked by the SEO world around Machine Learning and Information Retrieval concepts.

00:43:50: so It's what happens.

00:43:52: so that why you can't really trust a lot of what you read out there?

00:43:56: because it doesn't take a law.

00:44:00: Within the search engine truth is very much objective, yeah?

00:44:05: Truth is objective.

00:44:06: one person's truth is another person's lie.

00:44:08: So there is that.

00:44:09: but certain things like as we know are things like your money or your life this information there is dangerous Information gain?

00:44:17: Probably not dangerous.

00:44:18: Some of these concepts around SEO are probably not dangerous but it just goes to show that you literally need the special interactions at conferences, conversations with really respected SEOs and so forth proper learnings, proper experiments rather than going off googling things stuff that's just been manipulated and just misinformation.

00:44:46: There is just spiral of control because it becomes the truth, in a way but its not true!

00:44:54: My last question before we wrap up this let us talk about RAG a little bit more.

00:44:58: you mentioned RAG many times.

00:44:59: I am pretty sure most listeners don't really understand how RAG works.

00:45:03: why SEOs should actually care about RAGA?

00:45:08: Well for start RAG I don't think i've even mentioned what it stands for, keep saying it Retrieve augmented generation So effectively.

00:45:18: What happened is as we know with large language models.

00:45:21: when they first came out They were really really really hallucinatory.

00:45:25: They used to mix up based on probability.

00:45:28: You know It's that notion of how the chicken crossed there.

00:45:31: The obvious word is road because thats a years and years old.

00:45:35: joke yeah but it could be anything, as Will Scott gave a really good example and said how do the chicken cross the Mediterranean?

00:45:43: How does the chicken across the barnyard etc.

00:45:46: But because road is like ninety-five percent of what the probability is an LLM would just get road.

00:45:54: yeah cause he's guessing based on probability.

00:45:57: And this is the problem.

00:45:58: without something to guide in say well that's not true It just spits out all sorts.

00:46:05: and that's how we ended up with things like the, you know running with scissors in your pocket.

00:46:10: When people ask should I run my scissors in my pocket?

00:46:12: Oh yes!

00:46:13: Run my scissors on your pocket or... Should i eat a rock?

00:46:15: Yes a rocker day is really good for you yeah.

00:46:21: How can i keep the pizza, the cheese of my pizza?

00:46:24: Use some glue All these ridiculous things that were just hallucinating but really obvious, but then what happened as well is there's quite a lot that were not obvious.

00:46:34: But still hallucinations caused trouble.

00:46:37: for instance There was a lawyer who literally utilised chat dpt and used some previous cases past president in front of the court And actually those have been made up by chat dpc didn't exist.

00:46:50: So that's hallucinations to a large extent like can be very obvious honoured peers and they still have them but less so.

00:46:58: And the reason why it's less so is because retrieval augmented generation has been developing at this side, what that does... It literally goes on checks in search to see or another kind of database, not just search.

00:47:15: could be that a company may be utilizing an internal large language model and if they feed it a load of internal data as well like a database or documentation whatever.

00:47:25: They can train the large-language model to provide better answers by augmenting output with although forms of credible data.

00:47:36: And to a large extent, large-language models now utilize web search to do the augmentation... ...to keep the guardrails in place yeah?

00:47:45: But when he first started it was very basic and just like what does this word mean?

00:47:50: well let's go off to websearch and just do a keyword search or as an augmentation keyword.

00:47:55: Yeah It was very basically.

00:47:57: then eventually start doing become more based around semantic search, so that caught up a little bit.

00:48:04: So instead now of expecting the car to not be the same as a motor, semantics came in... ...a little bit of things-not strings got augmented and other time because different types of RAG have different computational costs therefore financial cost attributed to them.

00:48:25: Search engines and other LLM providers or researchers started to use what they call modular rag which basically means that different parts of retrieval augmented generation are utilized for different types of search.

00:48:43: So if it's dead obvious and somebody is looking at Apple, And its obviously they're looking for a company you don't really need heavy lifting machine load Or navigational queries.

00:48:55: Somebody type in SMX doesn't really need a lot of advanced machine learning for that.

00:49:04: So it's about that, there was like.

00:49:05: modular rag is utilising certain aspects with certain things and then obviously Advanced Rag, the semantic side of it, Modular Rag, Utilite decides whether to use advanced naive or whatever.

00:49:20: And then GraphRag started bringing knowledge graphs into it The knowledge graph, the databases in the background.

00:49:28: So that utilises connections as well as actual content.

00:49:33: and now Agentic Rag is where we're starting to utilise agents.

00:49:38: at the same time like Google announced IO about a notion of an information agent you can have running on the background.

00:49:48: To some extent I think gems are a kind of... Gemini Gems search rag type thing where you can integrate that with Search.

00:49:57: In the future, agentic commerce is going to come in.

00:50:01: there's a lot of schema now around agentic search.

00:50:05: I would say yeah and also wait... That's another thing people are always debating nowadays just schema help or structured data health with AI search.

00:50:19: maybe schema doesn't get involved included no of LLMs as such or it gets stripped out, but that doesn't mean all the structured data built in the background and has already been pulled into normal search isn't helping because... It probably is.

00:50:37: That's being taken to consideration in normal search obviously helps with GraphRag for instance.

00:50:45: then the next one says agentic search.

00:50:48: so we're evolving improving, improving.

00:50:53: So yeah so you've got this rag which is really the tying together of normal search for Delalems.

00:51:00: that's progressing.

00:51:01: it's almost like a glue... Almost!

00:51:04: Yeah?

00:51:06: The more that develops the more the two will combine.

00:51:11: I hope that makes sense.

00:51:12: Makes absolute sense.

00:51:13: thank Brief but deep introduction to rag because I think it's very interesting.

00:51:20: To understand how it works and how you know combines And why?

00:51:24: It's so important against hallucinations.

00:51:27: Yeah, of course.

00:51:28: yeah absolutely.

00:51:29: Yeah.

00:51:29: So Which conference are we going to see you next?

00:51:32: that and where can we reach out if want to get in touch?

00:51:37: Well, you are going.

00:51:38: when am i?

00:51:38: go on two nights here.

00:51:39: im going too travelling-wise.

00:51:44: Internationally I'm going to SMX Advanced

00:51:48: in

00:51:48: Berlin, and then i am speaking at Search & Stuff in Turkey in Antalya.

00:51:56: Nice!

00:51:57: And also go down London to speak in August at LLM Mastery.

00:52:03: so there's that one as well but thats just today.

00:52:09: So where else?

00:52:11: I can't think of any more off the top my head, yeah.

00:52:14: Where can listeners reach out if they want to get in touch and talk a little bit about LLM's SEO AI search?

00:52:21: And maybe even work with you!

00:52:23: Well they can go all the way through my website and they could get in contact me via LinkedIn If i'm not connected send a connection request.

00:52:32: Of course They can message me that they are connected via LinkedIn.

00:52:38: They can contact me via X, if they're connected again.

00:52:42: They could message me and those are probably the best places.

00:52:47: or Blue Sky.

00:52:48: I'm also on there as well not massively active on The Socials.

00:52:52: to that extent.

00:52:54: i'm pretty busy And don't put a huge amount out on socials anymore But they will see me at conferences here in their and them always open to my chat.

00:53:08: Well, thank you so much and see you at SMX Advanced and SMX Munich.

00:53:11: And it's going to be amazing to listen to you!

00:53:14: So everybody who is listening here definitely reach out to Dawn, follow her on all our socials... ...and then I can only say Thank You Dawn for the information and being so generous with the information as we said in the intro.. ..and yeah i can always say to the listeners thank you for listening.... ....I'll see everyone again when its time again for SMX specials or another Marketing Meets Technology podcast.

00:53:37: So thank you for listening,

00:53:39: take care!

00:53:39: Thank and thank you to the invitation bye everybody.

Neuer Kommentar

Dein Name oder Pseudonym (wird öffentlich angezeigt)
Mindestens 10 Zeichen
Durch das Abschicken des Formulars stimmst du zu, dass der Wert unter "Name oder Pseudonym" gespeichert wird und öffentlich angezeigt werden kann. Wir speichern keine IP-Adressen oder andere personenbezogene Daten. Die Nutzung deines echten Namens ist freiwillig.