Is AI a platform shift or a paradigm shift? | Benedict Evans
The MAD Podcast with Matt Turck · with Benedict Evans, Technology analyst and former partner at Andreessen Horowitz
Benedict Evans is the Technology analyst and former partner at Andreessen Horowitz. We cover why generative AI needs dedicated tools rather than blank chat interfaces, why predictions about AGI lack a theory comparable to physics, and how AI bias often comes from hidden patterns in training data rather than explicit demographic fields.
Chapters
- 1:06 — The AI platform shift in 2024
- 5:54 — Gen AI in 2024 vs. PC-boom in the 80-s
- 13:24 — Until AGI happens, there will be vertical-specific apps
- 15:12 — Should companies have an AI strategy?
- 21:04 — Platform shift OR paradigm shift?
- 23:55 — How should we think about AGI in 2024?
- 34:08 — Is gen AI grossly overhyped?
- 36:27 — AI bias and the hidden problems in data
- 44:56 — Apple Vision Pro and the future of AR/VR
Transcript
The AI platform shift in 2024
Matt Turck [0:53] Benedict, welcome.
Benedict Evans [0:54] Thank you.
Matt Turck [1:16] So talking about generative AI, you have said that almost everyone in tech agrees that it is the thing, but it's much less clear what it means. So in 2024, how do you think about generative AI as a platform shift, in particular in contrast with prior forms of machine learning?
Benedict Evans [1:40] Yeah, I actually think it's interesting to draw a direct comparison with the last wave of machine learning, which starts in, say, sort of 2013, 2014, so kind of a decade ago, in that it starts with demos of image recognition, or it starts with ImageNet. And, oh my God, this works. And then it kind of generalizes to a bunch of other computer science problems, quote-unquote AI problems that hadn't worked. So language and translation and so on, which people were trying to solve with different techniques.
Benedict Evans [2:09] And so it generalizes to that as well. But then you'd go to a big company and you would say, I would say, okay, you can do image recognition now. And they would say, yeah, but we don't have any images. So that's very clever, but what are we supposed to do with this? And it took a while to work out that the right level of abstraction was to think that this is pattern recognition. And like every software company for the last five years at least, maybe longer, has basically been saying, we've realized that this is a problem, and we've realized that we can solve it, we can turn it into pattern recognition, and we can solve it with AI.
Benedict Evans [2:49] And then there's a classic company-building, platform-building model. You unbundle the incumbent, you change the model, you peel something out of Salesforce or Oracle or Excel. You work out your go-to-market. Is this a feature or a company? All those kinds of standard venture startup questions. But the core of it is to work out what is this? How do you conceptualize what this is? And in fact, I remember having a meeting with a big media company, a big TV company, in 2015 or '16, and they were saying, look, we've got all this videotape, all this video of our people interviewing athletes.
Benedict Evans [3:26] They had a sports channel, and theoretically we could use machine learning to index all of those. Obviously we could transcribe it, but we could do more. We could index all the video as well. What would we do with that? What questions would we ask? What questions might we answer? And as we kind of sat around the room, this kind of long pause. And to be honest, I think the answer might be nothing. Here we are 10 years later.
Benedict Evans [3:48] And the reason I kind of make this point is you get this new thing, and to begin with, there's a couple of really obvious problems you have right now where you can see how this would solve that problem. But then your understanding of what the technology is kind of evolves, and your understanding of what problems would map against it kind of evolve. And so you could go back and talk about SQL and say, imagine explaining SQL to a supermarket in the late '70s.
Benedict Evans [4:19] It would not be obvious why this was useful or what you're going to be able to do with this stuff. And we're kind of at that point with generative machine learning now, in that we had the amazing demo, like, oh, I can go and get it to write me a song. I can make cat pictures. Like, last wave of machine learning, you can recognize cat pictures. Now I can make cat pictures. Okay, great. Why do we care? But we're still—and we had, I think, in 2023, that wave of the initial, oh my God, you can use it for that right now things, which is basically coding and brainstorming.
Benedict Evans [4:57] So on the one hand, all the code auto-suggests things. On the other hand, it's kind of an interesting contrast. On the one hand, coding, and on the other hand, advertising, which are like the two diametrically opposed industries. So, brainstorm me 200 ideas for a headline, brainstorm me 200 survey questions, brainstorm me 200 slogans or taglines, brainstorm me 50 different ideas for what this image might look like, and then we'll pick those up. And those are like the things where it works out of the box, but we haven't got the rest, and we're still sort of sitting and trying to work out, okay, then what?
Benedict Evans [5:27] And there are people who would say, you don't get it, it's everything. This is just a completely fundamental change. This is like a completely generalized compute platform. This is a shift of the magnitude of the GUI, say. And then there's a slightly less excited view that says, no, this is kind of more like smartphones or the web or PCs or cloud. Yeah, everything will get built around it, and then in 10 years there'll be something else, which is kind of what the last wave of machine learning was.
Gen AI in 2024 vs. PC-boom in the 80-s
Benedict Evans [5:54] It kind of got subsumed into everything, and now everything is AI and machine learning, but you don't really notice it. But I think that's the kind of pivot point now of: there's going to be 10,000 products that use this stuff, or is there just going to be one product that uses this stuff?
Matt Turck [6:19] Yes, precisely. That question of use cases seems to be the 2024 question, where 2022 was, okay, this exists. 2023 was, okay, let's experiment. 2024 seems to be the beginning of use cases. You've talked about bundling versus unbundling. To build on what you just said about, is it going to be one product or multiple products, how do you think about it?
Benedict Evans [6:45] Well, so the analogy that I was thinking about the other day, and my last answer was basically analogies, but the analogy that kind of occurred to me recently was to think about spreadsheets. So PCs emerge in the late '70s and really take off in the early '80s. And if you read Dan Bricklin talking about creating VisiCalc, he shows—and the young people of today won't know this, before either of us were born—spreadsheets were paper. You can still buy them on Amazon.
Benedict Evans [7:12] You can buy a pad of spreadsheet paper, preprinted paper, like A4-size paper. And so you've got all these accountants whose job is basically making spreadsheets by hand, maybe with an electronic calculator, but you're filling in each cell in the spreadsheet by hand, over and over again. And he shows them VisiCalc, and they're like—it blows their mind, because suddenly something that would have taken a week takes like 10 minutes. You iterate, recalculate, recalculate, recalculate, because it's not building the model, it's changing the assumptions.
Benedict Evans [7:46] That's so transformative. And so there are all these stories of accountants who are saying, "I did a week's work in an afternoon." And at the time, remember, PCs at the time are like $5,000, $10,000, which is an interesting comparison, incidentally, with the Vision Pro, which is—the original Mac was $7,500 in today's money. But the point is, if you looked at VisiCalc and you're an accountant or banker, that blows your mind. If you look at it and you're a lawyer, you say, "Well, I can see how that would be useful for them, but I don't do that."
Benedict Evans [8:16] So what do I do with this? And it's clearly not—maybe in the future you might imagine—and obviously word processors existed, so you could think, "Well, I need a word processor for this." If you're an architect, maybe that's a better example. What am I supposed to do with this thing? It's very clever, but—or you're a graphic designer, or you're a newspaper editor or something—it's not at all clear. You can sort of see, okay, it's very cool, but what do I do with it right now?
Benedict Evans [8:46] And that takes time. And I think this is a sort of paradox to looking at ChatGPT or Gemini or all these chatbots, is theoretically it's this completely open interface where you can do anything, except you kind of can't, actually. And there's a whole set of stuff where you need buttons and tooling. And I don't want to shut my eyes and think for 30 seconds to work out what the three options might be. That's somebody else's job. I want the person whose job it is, who knows about this stuff, to have worked out what the four options are.
Benedict Evans [9:06] You want GUI, you want tooling, you want somebody else to have thought about how this should work and how you should go through this task. You need auditing and records and tracking and import and export and all this stuff. Yeah.
Matt Turck [9:14] Which is what happened with enterprise software, right? It was not just about having a database, it was about having the tooling on top of it. So you think this is the same pattern?
Benedict Evans [9:32] Yeah. Well, so there's an old joke that every Unix function became a company. And I think you could say the same about if you hit File, New in Excel, you get all these templates. And those are partly suggestions, like, what could you do with this? But they're also tooling. We made you one. But of course, all of those become companies too. And theoretically, you could run FirstMark on Excel with a bunch of VLOOKUPs and IF statements, but you probably don't.
Benedict Evans [9:40] I mean, that would probably not be a great idea.
Matt Turck [9:41] Or maybe you would.
Benedict Evans [9:56] Yeah, well, I had this slide in the presentation I did about this saying I met this consultant years ago. I think I spoke to them on Twitter, but of course now I can't find it. He said that half of their jobs were telling people who use Excel to use a database, and the other half were the other way around. And so the point is, yes, theoretically you could run the entire world on Gmail, Excel, and Oracle, but in practice we don't.
Benedict Evans [10:25] And why is that? Why is it that the typical enterprise today has 400 to 500 SaaS apps? They're all doing stuff that you could do with Excel. I mean, clearly not at scale when you've got 200,000 people, but they're all doing stuff you could do with Oracle. They're all doing stuff you could do. It's basically every enterprise software company, for the sake of argument, is unbundling Oracle, Gmail, or Excel. And you could do Gmail with Oracle. In fact, my first job was, I worked at a bank that used Lotus Notes, which basically was trying to do email with a database.
Benedict Evans [10:49] It's like trying to build Gmail with Access. Terrible idea. And so all of those things kind of get peeled out and unbundled and turned into individual use cases. And a lot of that is just about telling people what the job task is and how to do it, rather than just giving somebody a blank sheet of paper and saying, "Go off, there you are." This is also the problem. I mean, the same point about Excel you could make about no-code.
Benedict Evans [11:04] No-code apps have kind of a natural ceiling, whereas I know I don't want to give 5,000 back-office invoice processing people a copy of Notion and say, "Here you are, you can build your own tool." I want to make the tool so that they can't screw it up.
Matt Turck [11:20] As you said, I think in one of your talks, maybe the difference with generative AI and the OpenAIs of the world is that they go very high in the stack. So there is a temptation to think that they could do all the things.
Benedict Evans [11:33] I think we kind of split two things apart. So let's kind of get to the core statement of the thesis. I mean, this is a quote from Bill Gates that he's seen two transformative demos: the GUI and ChatGPT. And the point about the GUI is that before GUIs, you had to have learned command lines and learned commands, and you'd get these kind of card overlays you put on your keyboard telling you what all the commands were, because if you wanted to save, you had to know what the keyboard command was to do that thing.
Benedict Evans [12:15] And with a GUI, suddenly anybody—so this is fundamental—with a GUI, you don't need to learn the command. So you have this huge change in how many tasks can be turned into software and how many people can use it, because suddenly the software can be much richer and many more people can use it. So far more tasks can be pulled into software. However, someone still needs to have written the GUI. So if you live in Australia and you want to use software to do your taxes, yes, you don't need to learn command lines, but someone needs to have written software that knows the Australian tax code and made a GUI for it so that you can click on stuff or tap on stuff.
Benedict Evans [12:52] In theory, you could go to ChatGPT and say, "Hey, talk me through how I do my Australian tax return." And it would go and find the Australian Tax Office's website and work out how the tax code works and come back and talk you through it in 15 steps and say, "Well, give me an image of your payslip for the end of the year." Obviously, it's going to be multimodal, and upload—give me a login to your bank account and I'll work it out.
Benedict Evans [13:20] Okay, here's your tax return, and I'll generate an image of the tax return form. None of this is science fiction exactly. You could sort of imagine that being possible, but you kind of know at the same time that, in practice, what I've just described would kind of require AGI. In practice, there's like 10 ways that that's going to break. And never mind the hallucination problem, which is a whole separate conversation. That's just kind of not a realistic description of the systems that we have today.
Until AGI happens, there will be vertical-specific apps
Matt Turck [13:42] So, is that fair to play it back? So, yes, there is a world where OpenAI, one of those top models, could do all the things, but that would basically be AGI. Until that happens, then it's going to be a series of vertical-specific AI-powered SaaS solutions.
Benedict Evans [14:03] Yeah. So, this—how far up the stack—there's a very binary difference between it's the top of the stack and it's kind of two or three levels underneath. And obviously, if you're an investor or a startup, how far up the stack matters a lot. You don't look at Snap or Pinterest or name 10 companies, you don't look at that and say, well, that's just an AWS wrapper. So this is kind of the challenge in this concept, this phrase: that's just a thin ChatGPT wrapper.
Benedict Evans [14:32] Well, yes, that's clearly a meaningful statement, but how thin? Because you don't look at Pinterest and say it's just a thin AWS wrapper, or whatever cloud they run on. Maybe they run their own cloud. Anyway, you don't look at it and say it's just a thin Intel wrapper. So how far up the stack does this stuff go? And this is clearly what happened with the last wave of machine learning. There was a brief moment where people said it needs all this data.
Benedict Evans [14:59] Only Google's got all the data. There's going to be like three people who've got enough data to do AI. And that turned out to be completely wrong. And if you ask today, how many machine learning models are there? Or if you'd asked in 2021, how many machine learning models are there? That would be like asking, how many databases are there? It would be just like a meaningless question, like millions, and who cares? The fact that LLMs are so big and heavy and capital-intensive shifts that model another level.
Should companies have an AI strategy?
Benedict Evans [15:13] But it's not clear how far. It's not clear how many models there'll be, and it's not clear how far they'll go up the stack.
Matt Turck [15:26] When you talk to large corporations, what do you tell them? What kind of conversations do you have with them in terms of their AI strategy and what it should be?
Benedict Evans [15:51] So I had this conversation with the head of a big industrial company last summer, who said, "Remember when everybody needed a 5G strategy?" And I remember I wrote something about it. At the time, it was kind of hilarious because it was like every big company had read about 5G on the plane in The Economist or Businessweek or something, and they land and they send everyone an email saying, "What's our strategy for this?" And the answer for most people was, "You don't need one."
Benedict Evans [16:19] I mean, the 5G hype, in hindsight, was very weird because it was just a faster pipe. It wasn't like a transformative new technology the way people talked about it. Clearly AI—or rather, let's get specific, generative AI, generative machine learning—clearly is a fundamentally different thing in a way that 5G wasn't. And a lot of people need to think about it or think about what it might mean for their company. After that, you go off and you think, okay, so do we pay Bain, BCG, McKinsey?
Benedict Evans [16:47] Do we pay WPP or Publicis? Do we pay Cognizant or Infosys or Accenture? Accenture would be really upset that I put them in that category, but you know what I mean. What kind of a question is this? Do we pay Adobe? And again, this comes back to this kind of platform shift conversation. So imagine you are—pick a name out of thin air. Imagine you're Caterpillar. What's your generative AI strategy? Well, some parts of that will be in software from your vendors.
Benedict Evans [17:15] So there will be generative AI capabilities inside stuff from Google and Microsoft and Oracle. Some of that you may not know is there. So, like, your network security intrusion detection system will use generative AI. You will not need to know about that. And you will use the same network security software as Citibank and Exxon—again, companies in completely different industries. There will be startups coming and trying to sell you this super-cool, new, expensive software that's more efficient and uses generative AI.
Benedict Evans [17:53] There will be people coming to you and selling you new CAD software, and your CAD vendor will be giving you new generative AI features. So there'll be stuff that is in general-purpose software that everybody uses and you don't know about it. There'll be people selling you task-specific things that aren't specific to your industry. There'll be people selling you task-specific stuff that is specific to your industry. The incumbents in your industry will be adding this as features. I mean, the classic pattern of the platform shift is the incumbents always try and make it a feature.
Benedict Evans [18:24] So we've seen Google and Microsoft spraying, and Adobe spraying, generative AI all over their products in the last year, 18 months. And then sometimes startups come along and unbundle something, or they find some new fundamental problem that you hadn't seen, and they solve it with generative AI. This is why large enterprises have 500 pieces of software. SaaS is the route to market that makes it much easier to deploy. But it's also like there's another 500 of those or 1,000 of those as people go and find all those individual problems.
Benedict Evans [18:55] So that's one axis. The other axis, kind of at a right angle, is: does this actually change the nature of your product? Does this change how you do your business, how you run your business? Does this change the nature of your product? Does this change the nature of the market and the competitive set that you might face? And clearly cloud did that to Oracle. And on the other hand, cloud didn't do that to Caterpillar.
Benedict Evans [19:24] Cloud changed the way Caterpillar does their business. And there's probably—I don't know much about their business—but there's probably, if you look at John Deere, John Deere is doing a lot of stuff in remote management and how they control the devices and how they control the equipment and so on. Is this a vector for a new competitor to enter your market? Is this a vector for fundamental change in the nature of your market? And if you are selling very large pieces of tightly wound copper to large railway companies, the answer is no.
Benedict Evans [19:53] If you're an airline, the answer is no. But for the sake of argument, everything is kind of disruptive to someone. So, like, online flight booking is very disruptive to travel agents, not disruptive at all to airlines. It kind of changes how airlines think about how they run their business and how they sell tickets and maybe pricing and maybe loading or something. But the fundamental business of the airline is owning or leasing the airplanes, owning the slots, buying fuel, and then filling the planes.
Benedict Evans [20:25] It's not really a software business. Wind back a minute: yes, there's a lot of software in how you do that, but software is not a vector to change how you compete there. So the challenge, of course, is it's not always obvious what that's going to look like. So Airbnb is not a hotel company. It changes what it means to be a hotel. Uber is not a car company or a taxi company. They don't sell software to taxis.
Benedict Evans [20:44] They change what it is to be a taxi. So you've got that sort of axis there of how do you operationalize this? How do you work out what you would do with this in your company? Some of which is to do with how you work, some of which is specific to your industry, some of which isn't. But then you've got this kind of other axis, which kind of intersects in the middle, maybe: how is this actually changing the nature of what you are and how you might work?
Benedict Evans [21:03] Which kind of goes back to my, like, is this a Bain kind of a question or an Accenture kind of a question?
Platform shift OR paradigm shift?
Matt Turck [21:15] So we've talked about platform shift, but I heard you ask the question whether that might be more than that, a paradigm shift.
Benedict Evans [21:39] Well, this is this point about—I've mentioned this number a couple of times. This is a number from Productiv, which helps people work out how much software they have as opposed to how much they think they have. Big companies have 400 or 500 apps. Okta has the same kind of data. A lot of that is about route to market, that SaaS is a new route to market. It's radically easier to deploy software because you don't have to go and install a Windows app on every Windows PC in the company and put it in the company's data center.
Benedict Evans [22:14] And so this unlocks this huge opportunity. And SaaS companies have spent the last 10, 15, maybe 20 years using that route to market, finding many more smaller use cases where, because there's so much less friction, you can go and build a business there. And machine learning kind of did the same, then became a thing that every SaaS app used. And now generative AI will be kind of that as well. But there's another argument here that says if you don't actually have to write all the software one at a time, you can pull many more tasks into dedicated pieces of software that maybe wouldn't support a dedicated SaaS app, just as SaaS lets you pull many more things into software that wouldn't have supported a dedicated on-prem installation.
Benedict Evans [23:02] So you remove this layer of friction in turning something into software. So you have way more stuff that gets done in software. And I don't know, I struggle with this because, as I said, in my actual job, I can't work out something that I would actually use ChatGPT for. I don't do things where the things that ChatGPT is useful for and reliable for today are useful. I don't write code, I don't brainstorm. I don't, in that way, need text written for me.
Benedict Evans [23:30] That's not what I do. What I would like to do is say, look, I've got 45 more or less identical PDFs with slightly varying formatting, each of which give an amount of money and a date. And I would like you to go through and compile all of those for me. Can I do that with ChatGPT now? No. I should be able to. Absolutely. Is that a ChatGPT use case or is that a third-party use case? Probably that's a ChatGPT use case.
How should we think about AGI in 2024?
Benedict Evans [23:55] There's always that kind of, is it a feature or a company kind of a thing? But then there's, again, this is the classic platform shift. When a platform shift comes, some of it becomes features. Some of it becomes separate companies. And this is a conversation we had earlier. I don't feel like everything will just get subsumed into the one thing. I feel like, no, you need buttons.
Matt Turck [24:06] You have a very nuanced and interesting view on AGI. Can you talk to that? How, in 2024, should we think about that?
Benedict Evans [24:24] Well, all my opinions are nuanced and interesting. Come on. It's funny. So my grandfather was a science fiction writer in the sort of '20s, '30s, '40s, '50s, and he wrote a story called A Logic Named Joe in, I think, 1946. And the premise is everybody has a home computer, which is called a logic computer. At the time, it was a job title. So everyone has a logic, and they're all connected to a global network, and they're connected to these global databases in the cloud called tanks, I think.
Benedict Evans [24:55] And so you can sit at this thing and do your banking or book your flights or do online dating or look up any piece of information and answer any argument. So it's basically describing the internet. And one of these things has some kind of manufacturing defect, which means it starts just being helpful and answering any question that anybody asks. And for some reason, because of the way the network works, it sees all the questions. I don't think my grandfather quite thought through the network architecture.
Benedict Evans [25:22] So this thing sees any question that's asked anywhere on the network, which doesn't seem very realistic, but anyway, it just starts answering them. Like, any question. Like, how do I murder my wife? Someone types this in as a joke, and it pauses and says, "What color is her hair?" And then it suggests an undetectable poison that only kills blondes and says, "This is not currently known by science. I've just invented it for you."
Matt Turck [25:24] It's amazing. That was in 1946.
Benedict Evans [25:42] Yeah. And they're, like, screaming in panic, like, "The censorship circuits are broken. Wait for the trust and safety thing." And, "How do I rob a bank?" And, "Give me a foolproof way of making money," and so on and so on. In the end, they work out which one it is and unplug it, which is probably not what the doomers have in mind as an easy solution. I think the challenge here is—so all that's just kind of a long digression.
Benedict Evans [26:08] I think the fundamental intellectual challenge is that if you took the specs for the Apollo program and gave them to Isaac Newton, he could have done the maths and told you whether it would get to the moon. Like, maybe not literally, maybe literally, but certainly, like, theoretically, you could have given the specs for this thing to somebody in 1750 and said, "Will it get to the moon?" And they could have done the maths and told you, like, this much weight, this much thrust, this much fuel, this is how far away the moon is, this is the rocket.
Benedict Evans [26:42] The point is, we had a theory of gravity, we had a theory of orbital mechanics, we had a theory of physics. And you knew how the rocket worked, and you knew what would happen if you put more fuel in, and you knew when it would explode, and you could calculate the tolerances of the rocket engine and the pipes and everything else, and you could work it out. We don't have any equivalent set of theories for intelligence or artificial intelligence. We have a lot of theories of how some bits of it might work, but we do not have a theory of what we have and in what senses what we have is different from and the same as a dog or an octopus or a horse or a mouse.
Benedict Evans [27:12] And we don't have a theory, actually, of how LLMs work. I mean, which is a kind of funny thing to say, but it's like, we do. But we also, at a very mechanistic level, we know what they're doing, but we also don't really know why it works. We don't have a theory of whether or not they would stop scaling. So this kind of goes back to the rocket point. Like, you remember Jules Verne wrote A Voyage to the Moon, and they use a cannon.
Matt Turck [27:17] Yep.
Benedict Evans [27:42] And people in, whenever it was, 1880, could have sat there and calculated, okay, number one, that much explosive in a cannon made of wrought iron or bronze, the cannon will burst. People have done that maths quite a lot. And plus, the G-force will be like 150 G and everyone will die. And everyone in 1880 could have done those maths and told you, no, it won't work. The cannon will explode and the people will die anyway, even if it doesn't. You could do the same with the Apollo program.
Benedict Evans [28:03] Like, will the rocket explode on the pad or not? If you double the size of the engines, what will happen? We can't do that with LLMs either. We don't know what will happen if you put double the data in. Or why. Or we don't know why it works with this much data or not. So the point of all of that is you can't make a chart. You can't make a chart. You can't kind of do a scatter plot and say, well, people are here and dogs are here and a horse is there and an octopus is here and ChatGPT is here and ChatGPT-3 was there and 4 is here.
Benedict Evans [28:17] And on the 17th of December, 2027, at current growth rates, it will hit dogs.
Matt Turck [28:19] Yeah, just add more data and we'll get there.
Benedict Evans [28:38] We don't have any of those kinds of theoretical models. And so that means you kind of can't do, like, a prediction. There's no Moore's Law here where you can say, well, it'll get to that power of compute level at this point. Set aside the fact that we actually don't have enough data to give it 10 or 100x more data, unless synthetic data. Anyway, so the point is, so that means that all of these conversations about this stuff become like a hunt for analogies.
Benedict Evans [29:04] And of course, talking about the Apollo program is an analogy. So people say, well, it's like nuclear weapons, or it's like meteorites, or it's like this, or it's like this. And they say, well, imagine if it was that, then you would know what to do. And the problem with those statements, of course, always is it isn't that. It's this thing. It's rather like when we had that kind of great panic about Facebook and people were saying, well, a restaurant wouldn't do this, a newspaper wouldn't do this.
Benedict Evans [29:36] Well, that may be true, but Facebook isn't a restaurant. It's a global social communications platform with 3 or 4 billion users. It's not a restaurant. It's also not a newspaper. It's not a phone company. It's Facebook. And you have to analyze it as that. Then it's the same thing here. An LLM is not a nuclear weapon, it's not a meteorite, it's not a car, it's an LLM. We don't actually know how they work. And so then everything becomes a sort of a hunt for metaphors, but it also becomes kind of a question: well, how is it that you think about a fundamentally unknown and unknowable risk?
Benedict Evans [30:08] There's an urban legend from, I think, the Cuban Missile Crisis, that there was a rumor that the missiles had launched, and everyone on the stock exchange starts selling, and one guy goes out and starts buying, and he says, "Look, it's binary. Either the rumor is true and we're all dead anyway, or it's not true and the stocks are cheap." And this is kind of the situation now. You can either look at this and you can say, well, there is a nonzero possibility that this thing is going to scale and kill us all.
Benedict Evans [30:41] And therefore we should freak out. Or you can say we have absolutely no way of knowing whether that's true or not. So this is no different fundamentally from saying we should all prepare for—here comes another analogy—we don't know that the meteorite isn't going to hit New York tomorrow, yet we all live our lives and we do not shut down the economy and build meteorite scanning systems and put nukes into orbit to do something about it. How do you think about unknown risks?
Benedict Evans [31:05] Now, this gets kind of hilarious because you have all these conversations where people are saying, well, what P(doom) do you assign to this? Which, to me, is a fundamentally invalid exercise because you're attempting to ascribe a numerical value to something that's fundamentally unknown. It's like saying, what's your probability of the existence of God? Well, you can have an opinion about it, but the only way to find out is to kill yourself and see what happens.
Benedict Evans [31:38] And that's only, like, that would only be kind of a negative proof, which gets you to the kind of Pascal's wager. Like, you'll find out because you're in hell; otherwise, you won't know either way. So again, Pascal's wager, I think, is kind of a funny one, because then people kind of start dredging up all their half-forgotten undergraduate philosophy. So you get, like, Plato's cave and Pascal's wager. And I always kind of like Anselm's ontological proof. Do you know this one?
Matt Turck [31:39] No.
Benedict Evans [31:59] Okay, so I love this. So this would be, like, one thing that people learn, if they learn this, the forgotten undergraduate philosophy. So Anselm says, okay, premise one is—maybe it's axiom, I forget what the terminology is—but okay. First proposition is that God, by definition, is the greatest thing that there can possibly be, because if there was anything greater than that, that would be God. So there cannot be any—God must be the greatest possible thing in any possible axis that you could define.
Benedict Evans [32:17] That's what God is by definition. Secondly, a God that doesn't exist is less great than one that does exist. A God that is actually real would be more of everything on any possible axis than one that didn't exist. Yeah. Therefore, God exists.
Matt Turck [32:17] Interesting.
Benedict Evans [32:40] Yeah. And about 30 seconds later, all the other theologians in medieval Europe said, "But this is obviously bullshit." And Anselm says, "Yes, but try proving it." And I think Bertrand Russell said it's actually much more interesting to talk about why it's hard to prove that it's wrong than the fact that it clearly is wrong. And this is kind of the way I look at the AGI argument, which is you can kind of define an AGI as something that's all-powerful and would kill us all, and then say, therefore, it's all-powerful and would kill us all.
Benedict Evans [33:18] And it's like, well, how can you know any of this stuff? Yeah, I feel—I don't know. I feel like all these conversations are best had after a bunch of kind of weird psychological, psychotropic chemicals in a group house in the Berkeley Hills where you live, even though you're in your 30s, you live with a group of other people and talk about AGI all day.
Matt Turck [33:24] Yes. And we're actually just saying there is a whole scene, quote, end quote, that does just that.
Benedict Evans [33:47] Yeah, I mean, the sort of sociology of Silicon Valley, there is an AGI scene of a certain kind of person that has a certain kind of lifestyle and lives kind of on the periphery of the tech industry and thinks that this is all really important and interesting and talks about it a lot. There are other scenes, like there was a VR scene. To some extent there still is, although it's out of reach of hobbyists now. But there was a VR scene. Palmer Luckey made the original Oculus himself out of components he bought on Amazon.
Benedict Evans [33:58] There was a nootropic scene. There was a crypto scene. There were all these sorts of scenes. The Homebrew Computer Club was a scene.
Matt Turck [34:03] Yes. Do you see the same people going from scene to scene? I don't know.
Benedict Evans [34:07] One of the reasons I left Silicon Valley is I couldn't deal with this kind of thing.
Is gen AI grossly overhyped?
Matt Turck [34:20] So, a little bit to the it's-hard-to-predict-and-the-data question: is there an argument to be made that actually generative AI may be grossly overhyped?
Benedict Evans [34:39] Well, so this is back to, like, how excited about this should we be? Every platform shift is a kind of—it's a classic Gartner hype cycle thing. And then there are other people who say, no, it's just going to keep growing. We've got this exponential growth, and this one isn't going to do that. It's going to go straight through to AGI and kill us all. Clearly, if we're not in a bubble now, we're going to have a bubble, because that's just, like, the nature of the cycle of life.
Benedict Evans [35:11] There will be a bubble around each new technology. There was a bubble around iPhone apps, a bubble around cloud, a bubble around every new thing. There is a bubble of some kind, which is kind of interesting because you would know more about this than me, but clearly venture fund investment has kind of gone down radically since two years ago. Investments have gone down radically. So you've kind of got a crash and a bubble kind of happening at the same time.
Matt Turck [35:12] Yes. Weird times.
Benedict Evans [35:33] How can I put this? I think what I would caution people against is doing the thing of saying, well, that doesn't work perfectly yet, therefore this is all completely useless. And you certainly see a bunch of people who sort of think this is like NFTs or something, that these are all just kind of con artists and it's just word prediction and it doesn't really work and this is all a bunch of nonsense. And you don't have to believe this is going to go to AGI, and you don't have to believe that we won't have apps, we'll just have one piece of software, in order to think this is a really important, profound change in how everything works and a kind of step change in what you can do with computers.
Benedict Evans [36:09] I, by default, go back to the quote from Larry Tesler that AI is anything that doesn't work yet. AI is whatever hasn't been done yet. Today, you don't look at image recognition and say, "That's AI." You can go onto your phone and you can go to the photo app and you can type in a word and it will find a picture of a book on a bookshelf behind somebody that you took 15 years ago. And you don't look at that. Five years ago, 10 years ago, that would have been witchcraft.
AI bias and the hidden problems in data
Benedict Evans [36:27] Totally impossible today. Yeah, of course, it's just image recognition. And that's what happened with the last wave of machine learning. By default, I think that's what will happen with this one. But that doesn't mean that image recognition was overhyped.
Matt Turck [36:41] If generative AI is going to be around us everywhere, obviously a key question is bias. And you've been thinking about bias in AI for a long time. What's your 2024 view of it?
Benedict Evans [37:11] So it's interesting. I wrote about this in my newsletter this week or last week. This week, I think, because Bloomberg did a story where they looked at people using ChatGPT to screen resumes, and it finds bias. And there was a thing in 2018 where Amazon had an internal project to use machine learning to filter resumes, CVs, because obviously they're hiring at huge scale. And what they found is that historically they mostly hired men. And so the pattern of a successful candidate is a man.
Benedict Evans [37:44] And meanwhile, it's not that it was looking at gender equals male in a database. It was looking at what sports people played and even more subtle things like what language people would be using to describe their accomplishments. And it doesn't have a model of male and female. It just has a model of all of the people we hired played football and none of them played lacrosse. Sorry, I'm stereotyping. In Britain, lacrosse is a girls' sport. In America, maybe it's a boys' sport.
Benedict Evans [37:47] I don't know.
Matt Turck [37:50] Yeah, it is more so.
Benedict Evans [38:15] Yeah, but you get the point. We tended to hire people who used more direct affirmative language in describing their accomplishments. It doesn't have a concept of man or woman. It's just doing pattern matching. The analogy here I really like is, there's a dumb, naive reaction to this, which is to say, maths is maths. It can't be biased. Yes, but the training data can be biased. What data have you put into it? The equally dumb, naive reaction is to say, this is because your hiring teams and your engineers are all white men who live in the Valley, and they haven't looked at this stuff.
Benedict Evans [38:49] And this is true, but only in a very limited sense. And the reason it's true in a limited sense is there was another really interesting thing that came out, a machine learning bias issue, again, like five, six years ago, which was someone was building a skin cancer-recognizing system. And so the obvious way you can screw this up is clearly there's no difference. It's not seeing man or woman, but you could have different skin tones. So if you don't have the right distributions of skin tones, you might get false positives and false negatives for people that are kind of a smaller group.
Benedict Evans [39:24] However, the problem that actually came up was that dermatologists tend to put a photo of a ruler in the picture of the skin cancer. And so if your pictures of skin cancer have a ruler and your pictures for your sample set of healthy skin don't, then guess what? What's the most statistically obvious difference between the two sample sets? It's not the shape of the little blemish on the skin. It's the giant great ruler. Push this a little further. Imagine if your pictures of healthy skin are taken under incandescent light and your pictures of unhealthy skin are taken under fluorescent light.
Benedict Evans [39:53] Imagine if mostly you use Samsung cameras for this and mostly Sony cameras for that. You're not even going to be able to see that. And so hiring more Black people or more women, yes, you should. But that's not what the issue is. That's not where the bias is coming from. The bias is coming from the stuff in the data that you didn't know was the stuff in the data. And so there's an interesting turn on this, however, which is that DeepMind did a project with Moorfields, which is an eye hospital in the UK, and they were looking at retinas, and their system discovered a difference between male and female retinas.
Benedict Evans [40:34] And apparently medical science didn't actually know there was a difference between male and female retinas. And so part of what you're seeing is, if you kind of systematize this, your machine learning system will use the patterns in the data. Those patterns might or might not be things that shouldn't be there. They might be things that are there but don't reflect society, like the ruler. They might be things that are there and do reflect society, but you don't want to use them. Like you only hire men.
Benedict Evans [40:58] They might be things that are there and do reflect something that you didn't know and would like to know. Like there's a male-female correlation in this data that you weren't aware was there. One of the ways I used to talk about machine learning is that it gives you infinite interns. Like, you would like someone to listen to every call coming into the call center and tell me if the customer is angry. Like, you've got a million calls a day, you don't have enough interns.
Benedict Evans [41:15] But it's also, what if you had one intern who's infinitely fast? This one intern could listen to every call coming into the call center and say, "When I heard the third billionth phone call, I noticed this interesting pattern." And that's obviously what the medical research point is, that it's finding patterns in the data that you didn't know were there, as opposed to look at the X-ray and save time from a human radiologist.
Benedict Evans [41:44] So the point of all of this is, what machine learning systems are doing is seeing patterns in the data. Some of those are patterns that reflect what you want, some of them are not. What does that mean? Where is that useful? Not useful. Helpful. Not helpful. I'm tending to monologue here, but there's kind of another interesting, important point here, which, in fact, this is probably the one thing anyone listening to this should take away.
Benedict Evans [42:00] So don't edit it out. So do you know about the Post Office scandal in the UK?
Matt Turck [42:01] A lot.
Benedict Evans [42:24] Okay, so in the UK, most of the Post Office is a sort of franchise system. So the individual branches are owned by independent small business people. Like, it's the bodega, it's the local pharmacy, has a post office counter. And so the Post Office rolls out a new point-of-sale computing system built for them by Fujitsu. The system has bugs. It starts showing shortfalls in cash, like big shortfalls. Like, this person paid us five, ten, twenty grand less than our system says they should have paid us.
Benedict Evans [42:54] And they think, okay, these people have been robbing us for years. Now we've got them. There are prosecutions. Maybe 1,000 people are prosecuted. Maybe 100 people go to prison. There are suicides, bankruptcies. People lose their homes. Meanwhile, Fujitsu and the Post Office know there are bugs in the system. And so they are prosecuting people and going to court on oath saying there's no errors in the system. The court's accepting this. And at a certain point, and this now becomes a legal question, at what point does this go from a sort of institutional blindness that they know there are bugs and kind of don't accept it, to it's actually like criminal conspiracy territory, and that they're kind of hushing it up.
Benedict Evans [43:13] But the point of this is, this is not—
Matt Turck [43:14] And what happened?
Benedict Evans [43:34] There's now a public inquiry going on. So I don't know what the American equivalent would be, but there's a debate about how much compensation, how many convictions get overturned, and so on. So this became a huge scandal. And the point is, this is 1970s technology. This is databases. This is Fujitsu. It's not Google or DeepMind or someone or OpenAI. And you don't look at this and say, well, obviously the solution is that we need to have a database regulator that makes sure databases don't have bugs.
Benedict Evans [44:07] Rather, you look at it and say, well, this is an institutional failure, A, in Fujitsu and the Post Office, and B, in the legal system not properly testing the evidence. And the same thing with AI bias. You have these people talking about this stuff as though we need to have a code of ethics so that people will make sure that there's no bias in their database system. Can I swear on this podcast? The fuck are you talking about? You're going to write—you're going to get people to commit that there's no mistakes in the code.
Benedict Evans [44:36] Do you have any idea? Do you not know anything about the software industry? The correct answer is to say, yes, you have to train people to be aware that this can happen, just as they need to be aware of how bugs happen, like buffer overflows and SQL injection, all the other ways that you can screw up software. This creates, like, a whole other class of ways that you can screw up your software, and people will do it. What you need then is to have the broader awareness from everybody that computers can be wrong.
Benedict Evans [44:55] This is another way that computers can be wrong. And do you have the processes to deal with it? Because you will do this, just as you will have more of these Post Office scandals. So it was kind of a very long answer to explain AI bias in five minutes.
Apple Vision Pro and the future of AR/VR
Matt Turck [45:18] No, fascinating. Maybe as a last part of this conversation, you also cover other areas as part of your thinking, many other areas, but one of them is AR/VR. What do you make of the Vision Pro, where that fits into that whole picture that might be emerging or maybe not emerging?
Benedict Evans [45:41] Well, so I think there's a sort of symmetry here in that 10 years ago this month, Facebook bought Oculus, and Oculus has this impractical, bulky device that's clearly not ready for consumers. It wasn't very expensive, but you needed, like, a god-level PC to run it. So you need an expensive PC to run it. But you have the demo. It's amazing. It's clearly part of the future. Like, oh my God, this works. Now, 10 years later, Apple launches this thing.
Benedict Evans [46:09] It's expensive, it's impractical, it's bulky. It's clearly not ready for consumers. You have the demo. Oh my God, it's amazing. This is clearly part of the future. Meanwhile, 10 years of work and $15 billion a year of R&D budget, Meta has the Quest 3, which is a perfectly good, credible consumer device that does not have traction. It's probably got, I don't know, maybe 10 million active users, huge abandonment rate in the past. It's not really good for anything other than games.
Benedict Evans [46:24] I said this on Twitter, and Zuck replied to me and said, no, like, the top five apps are all social apps. I don't know, is that self-selection in the user base or what? Anyway, this thing is clearly—
Matt Turck [46:24] This is—
Benedict Evans [46:48] We're bumping along the bottom. We haven't got the hockey stick yet. And so the question is, well, the real question here is: what Apple have done is they've made a device that lets you use apps. You can have an iPad app in the room with you that looks like it's really there, which you can't do on the Quest yet. Now, set aside the fact that it's heavy, it's expensive, blah, blah, blah, blah, blah. Also, in fact, is that useful?
Benedict Evans [47:15] That's really the kind of core question. Is it useful to have an iPad app floating in the room with you, or three iPad apps floating? Is that better than holding an iPad in your hand? Is the future of computing bigger screens? Is the future of computing seeing—it's a caricature—more rows and columns in your spreadsheet at once? Is there a way that you can turn those apps into 3D? Because that's really the transformative thing. And I puzzle about this because I think, like, taking desktop stuff and putting it smaller and vertical on an iPhone was useful even before you then get the native mobile stuff like Uber and TikTok and everything else.
Benedict Evans [47:52] But just having the email from your desktop on your phone was already useful. We get that with BlackBerry and so on. The point is, moving desktop from desktop to mobile makes it better. Moving from 2D to 3D, much less easy to say why that's better. What is 3D email? What is a native 3D experience? And the point is, a native 3D experience is a harder thing to see than a native mobile experience. But the point is, even taking the desktop stuff to mobile was better.
Benedict Evans [48:20] Taking the 2D phone stuff to Vision Pro is not really much better, except for a very small number of people who want millions of screens. And that might be like looking at a colour screen in 1985 and saying, well, I don't need colour in my spreadsheets, I don't need colour in my documents. Text is black and white. I'm not sure about that. And so there's a real puzzle here of, like, how big does this become? And you can get all these kind of deterministic frameworks.
Benedict Evans [48:46] They're like, people look at the iPhone, people always look at the new thing and say it's not useful, it's a toy. You have to ask yourself, okay, yes, but presume it gets very cheap and very lightweight, which is my point. Ignore the price of the Vision Pro for the sake of argument. Ignore the fact that it's heavy. It can get light, it can get cheap. Will it be useful then? And for the iPhone, yes, it wasn't obvious at the time, but it was easier to see that.
Benedict Evans [49:14] I think there's, like, some sort of deterministic tools here. Maybe one of them is: you're not going to wear a headset, no matter how light and cheap it is, all day outdoors. I would not have worn it walking here, even if it weighed 100 grams and had completely perfect passthrough. Therefore, it can't replace your phone. Therefore, we are looking at a sort of iPad-y kind of market opportunity. It's a phone accessory, which is not a bad market.
Benedict Evans [49:42] It's a couple of hundred. How many people have got an iPhone? 400 million people, maybe. It's a big thing, but it's not like the universal compute platform that replaces the phone. And for that, you would need glasses. And we don't know how far away glasses or the optics for glasses are. The fact that Apple has shipped this suggests that Apple doesn't think it's got them yet. Meta is spending $15 billion a year trying to make them. Apple maybe too.
Benedict Evans [50:10] But this is, again, it's almost like the AGI question. How far away into the future is something that looks like the glasses that we're both wearing that could put an iPad app here that we could both see? That feels like that could be a universal thing more than the headset. There's another backstop here, though, which is something could be amazing and part of the future, but not necessarily a universal part of the future. So games consoles are amazing. If you'd seen a PlayStation 5 in 1980, or not even 1990, oh my God, this is amazing.
Benedict Evans [50:37] Turns out the install base of games consoles is like 250 million units, and AAA PC games are like another 100 million, maybe. So it's like 300 or 400 million people. Pick a number. You can argue about this with Matthew Ball. Maybe it's 500. I don't know. It's not 5 billion people. Most people look at the AAA games and are saying, very pretty, well done, I don't care. So something can be amazing, but only a relatively small part of everything.
Benedict Evans [51:06] The same thing. I mean, the extreme case here would be like drones or 3D printing. A couple of years ago, we all bought a drone. Five days after Christmas, you say, okay, I've seen the roof of my house now. You buy a 3D printer, you make a little Eiffel Tower. There's no consumer use case. And so it's very easy to say, yes, of course, this stuff in some form: doctors, architects, engineers, CAD. Yeah, yeah, yeah, yeah, yeah. Is this a universal device?
Benedict Evans [51:25] No, not yet. A, it's not clear if the use case is universal, which is the games console point. B, given that it cannot replace your phone, it can't replace your phone. It's almost like a circular point. Like, it can't be the universal platform that replaces your phone if it can't replace your phone.
Matt Turck [51:30] Benedict, it has been a fascinating conversation. Thank you so much for doing this.
Benedict Evans [51:31] Thanks a lot.