You.com: An AI Chatbot to Search The Web with CEO Richard Socher
The MAD Podcast with Matt Turck · with Richard Socher, CEO, You.com
Richard Socher is the CEO at You.com. We cover why he argues search deteriorates under monopoly conditions, how retrieval-augmented generation lets language models reason over current web facts, and why organizations should not train their own LLMs from scratch without hundreds of millions of responses.
Chapters
Transcript
Full episode
Matt Turck [1:19] Richard, welcome. It's a real pleasure to have you. I've been looking forward to the conversation. So, to set it up, you are a man of many talents. You are an AI researcher, and your research has received over 150,000 citations and many awards. You're also an entrepreneur, having started an AI company called MetaMind that was acquired by Salesforce. Actually, you and I connected initially in the context of MetaMind. You kindly spoke at our Data Driven NYC event back in the day.
Matt Turck [1:43] I think that was 2014. You are also the founder of You.com, an AI search engine, which we are going to talk about extensively. And then, last but not least, you are also an investor as a founder and managing partner at AIX Ventures. So thanks for making the time in your busy schedule. We really appreciate it.
Richard Socher [1:47] My pleasure.
Matt Turck [1:54] I mentioned MetaMind a minute ago, and I thought that could be a good place to start. Can you talk about what led from MetaMind to You.com?
Richard Socher [2:32] Yeah. So I started MetaMind after my PhD, and I actually thought I was going to become a full-time professor and accepted a faculty job. And then I thought, okay, I'll do this company for a year. Maybe I'll teach a little bit on the side already at Stanford also, because no one at the time, and this was 2014, '15, no one was teaching neural nets for NLP. And I thought, well, but they're the right way of doing all of NLP. So someone's got to teach a class about it.
Richard Socher [3:01] And so I started teaching at Stanford. But my main job was CEO, CTO, and founder of MetaMind. And the idea there was: it's the right technology, so we should make it easier for people to use. This was before there was PyTorch, before there was TensorFlow. There were a few smaller frameworks, and none of them were general-purpose, worked really well, and were easy to use. And so we thought we could make it as easy as: you just drag and drop some images into the web browser, and then you get an image classifier.
Richard Socher [3:42] And so at MetaMind, we built a bunch of neural nets, both for natural language processing and computer vision, including 3D computer vision. I think we ran the first FDA-approved trial of a deep learning classifier in a production radiology system. It was triaging brain bleeds, intracranial hemorrhage, and a bunch of other really exciting work. And we eventually got acquired by Salesforce, where I became chief scientist and, after two years, executive vice president. At Salesforce, I had a really phenomenal time because we could do some pure research again, but also have a massively scaled impact on thousands and thousands of different large companies that exactly needed this platform and needed it to be easier to build their own classifiers, things like prediction builders, but also running search within enterprise and a bunch of other areas.
Richard Socher [4:40] And so I had a phenomenal time. I really enjoyed the work with many incredible people there, and especially Marc Benioff, my boss at Salesforce and CEO there. And so I had a great time, but also felt like it's kind of crazy how much progress we've made in natural language processing, yet the biggest application of NLP, which is search, hasn't really improved much. And as part of the research team, we actually invented prompt engineering in 2018. And we had some of the best language models in 2016 that used pointer mechanisms, which is essentially a form of attention mechanism, to do better language modeling.
Richard Socher [5:19] So there are a bunch of different ideas we've had, and we thought, man, at some point, hopefully Google picks this up and makes actual search better than, here's a list of blue links, and then here's a list of mostly advertisements followed by a list of micro-SEO microsites. And it sort of felt like search was getting worse for the first time a few years ago, rather than better. And I think it's what almost inevitably happens when you have a monopoly and there's not enough competition in a space.
Richard Socher [5:56] And so ultimately, users were suffering from that situation and couldn't get the answers as quickly. So we started You.com to improve web search for general-purpose questions. Then we were the first last year to bring chat into the search context also by having the first LLM shipped to the world, as far as we know, that connected and had a connection to the internet. And so it could be up to date, factual, and give you citations. And that idea got copied many, many times by basically a bunch of the biggest companies, Bing and others, and copied by a bunch of other small startups also throughout this year.
Richard Socher [6:09] You.com.
Matt Turck [7:09] You.com as a consumer search engine, but as one looks at it, it's actually a lot more than that. It's a lot more than what people would typically think of as a search engine. And in fact, it's basically a suite of AI tools. So there's YouChat, there's YouCode, YouWrite, YouImagine, and also a new service that you just released called YouPro. Can you give us a tour and double-click on what those different products do?
Richard Socher [7:37] Basically, when you think about search and chat, there are lots of different intents that people have. Sometimes people have the intent to just find a piece of information quickly, or they want to just have a quick navigational search where they just really want to go to another website. In fact, when you observe normal users outside of Silicon Valley, probably not part of your podcast listeners either, they open a browser and they put Google.com in that first URL bar, and they're mixing up the browser and the search engine.
Richard Socher [8:06] And there's all kinds of confusion in normal users. At You.com, we're chat-first, and you can basically get more things done. We want to help you be more efficient, right? If you're a student, if you're a developer, if you're creative, we want to help you be more productive. And so what does that mean? Well, sometimes you ask, what's the stock price? Most LLMs will just make up a bunch of numbers. We figured that problem of hallucinations out before anyone else.
Richard Socher [8:42] And so we launched the first multimodal chat where sometimes you have chat and sometimes you have chat and a stock ticker app. So there's a bunch of different apps that are incorporated into the chat experience to be an actual replacement for your search engine. And I think we're still the only feature-complete chat-first search engine out there. People separate the two a lot, and we're trying to show people you can actually default to a chat-first experience. But when you want to be the default, the bar is very, very high, right?
Richard Socher [9:13] You gotta have weather and apps, and you gotta have stock tickers, and there's a lot of little features that you have. And then we also don't want to just replicate all the existing features. We want to show people we can do even more novel things. So last year, for instance, we were the first search engine that had a TikTok app and a Reddit app. And some people said, "Oh, I love the Reddit app. I hate the TikTok app." And so we allowed people to actually have agency over it, and they can vote on saying, "I want to see this app less or more often in my search results," whenever the AI generally feels like it's relevant.
Richard Socher [9:46] And so why do we have things like YouImagine? Well, sometimes you say, "I'm looking for an image of a dog with sunglasses in a club." And then if that image doesn't exist, what's the best action? Well, it's to make the image into reality and generate the image. And so that's why we created YouImagine.com. And so there are a couple of features, and we're still, I feel like, iterating on streamlining them more and more.
Richard Socher [10:16] You can be like, how can I generate an image with AI? And then it'll tell you about it in text, and it'll show you that YouImagine app too. And then you can have, similar to how Google eventually started having Gmail and Maps and a bunch of other things, we're kind of building out a suite of those kinds of apps too. But the next generation of useful apps that help you be more productive.
Matt Turck [10:48] So you mentioned YouChat and YouImagine. What about YouWrite? What does that do?
Richard Socher [11:12] YouWrite is essentially also an essay writer that helps you write essays and blog posts, and it's a little bit easier to use and will write very good content. And we're now basically offering all of those AI services under the YouPro umbrella. So YouPro basically allows you to quickly have access to all the latest and most useful generative AI tools under there. And you can default to YouChat, the chat experience, in your browser through our Chrome extension, for instance, and then you have easy access to all of these latest AI tools.
Richard Socher [11:35] And this year we'll have some new modalities come out that will be very exciting to people, I think. Great.
Matt Turck [12:04] I'd love to double-click on some of what you just mentioned. How does that work to be a chat-first search engine, in particular in connection with the hallucination problem? When do you know when the LLM hallucinates, and therefore you should be grabbing information from a live app or a live website? How does that work practically?
Richard Socher [12:19] Yeah, it's a great question. It's actually pretty non-trivial to know when is a person looking for something purely factual, and you want to have as many citations as possible, or are they just trying to jam on something novel? You could have some people who want to talk about an alien invasion of Berlin in 2030 because they're trying to write a short story about it, and other people write about an alien invasion of Berlin in 2020 because they think the politicians are all lizard people, and they think it's some kind of conspiracy thing going on.
Richard Socher [13:07] And so I think there is a very subtle difference. In one case, you really want to be factual. In the other one, it's fine to jam. So what we often do is first build an intent classifier that understands sort of what is the intent that the user has. Those can get pretty fine-grained. And then as you go into it, you can essentially prime the large language model with a retrieval backend. And this is, I think, where the world is going—every major LLM company is asking to work with us, and we're actually going to start working with them now and supporting them in what's called RAG, retrieval-augmented generation.
Richard Socher [13:45] And the idea here is that you don't throw away everything from search and say, all right, the LLM now has to memorize everything and know all the facts in the world and be perfect at retrieval and reasoning, but instead you think of the large language model as the reasoning engine. And it's kind of like your uncle. He might not remember all the different facts. So your uncle will be more helpful if he has access to the internet or access to his phone and looks at photos and so on and is like, oh, when did we go on this hike?
Richard Socher [14:14] Well, exactly on this date, because you look it up. And so instead of thinking about, wow, maybe it was in this timeframe, 15 years or something like that, and maybe embellish the story. And so, long story short, in RAG, retrieval-augmented generation, you have a search backend. It will surface everything from the internet, and then you infuse those facts into the prompt, and you tell your LLM that sometimes you should use those facts if they're relevant to give the right answer, and you let the LLM kind of be the reasoning engine.
Richard Socher [14:32] With all the facts in mind.
Matt Turck [14:58] And how do you almost mechanically grab those facts and put them into the LLM? In other words, how do you connect to all of those? Do you crawl those and put them in some kind of vector database, effectively, or do you query them at search time?
Richard Socher [15:26] Yeah, it's a great question. There's a bit of a hybrid system. We actually have a larger and larger index ourselves, and maybe at some point we'll make that available to others as well. But yeah, right now there's a mix of some of our own index and retrieval mechanisms and ranking mechanisms and so on and crawling, as well as some external API partners. Okay.
Matt Turck [15:44] So you have built your own large language model called YouBot. Can you go into that, like how you built it, what it's based on, any kind of techniques you've used? Obviously, I'm sure there's a lot of secret sauce there, but whatever you can talk about.
Richard Socher [16:16] Yeah, at a high level, that also is a hybrid system. That basically enables users also to choose. So one of the other many advantages of YouPro is that you can actually use GPT-4, so the same model that powers ChatGPT, for your answers, but it only costs $10 rather than $20 that OpenAI charges. And of course, it's connected to the internet, and you have YouImagine and YouWrite and all these other capabilities. And you can also choose to not see advertisements, and there's lots of goodness.
Richard Socher [16:50] But at a high level, the LLM is also a mix of our own LLM for some of the questions that's just cheaper for us to serve. And we use one of the foundation models that was open-sourced, and then modify it and improve it for our use cases, as well as some OpenAI models, and we even have other backup models from other LLM providers.
Matt Turck [17:03] And do you imagine yourself in the future sort of integrating a combination of commercial, open source, and having kind of like a meta model that sits on top? Is that part of how you see things evolving?
Richard Socher [17:28] Yeah, I think more and more organizations will want to have their own LLMs and fine-tune them, but it's also not always clear necessary, especially when you understand the power of the retrieval backends in giving you the right facts and priming the answers through the prompt with that retrieval backend.
Matt Turck [18:05] And for the LLM that you grabbed from open source and worked on, I assume you fine-tune it with some data sources? Can you talk about that model of further training it for your purposes? And, for example, YouImagine is a copyright-free image generator. So I assume that the underlying training data for it was specifically curated to enable the copyright-free aspect of this?
Richard Socher [18:38] Yeah, I mean, when you generate, like, a brand-new image that's very different to anything out there, I don't think they generally have copyrights associated with them. And yeah, on the LLM side, we have millions of votes of people saying this was a good answer or this was not a good answer. They can tell us why. And we see if people click on a citation, for instance, and verify that fact. And then we have a billion-plus queries with web links that were clicked on, of course anonymized.
Richard Socher [19:05] Even DuckDuckGo keeps queries around and knows what clicks happen, just not associated with the user. And so that can all feed into systems to then improve the overall retrieval, search, and LLM answers.
Matt Turck [19:41] Maybe as a more general question for people that are thinking about deploying generative AI and doing work at the core model level, like with open source, any lessons that you can share that would not be super secret or anything about what worked, what didn't, maybe roads you went down on and turned out to be unproductive? Not that, by the way, you're an old company. You've been around for two years, so I don't know how many roads you had time to travel, but curious if anything comes to mind.
Richard Socher [20:15] I guess in terms of LLM training, I think it's a tough one to do for small organizations, and to do it really well and to really do it so much better than anything that's available through an API from OpenAI or Cohere or Anthropic. And I think you need a lot of training data for it. You need a lot of feedback from people to update those models. And, yeah, I don't think you should attempt it unless you have hundreds of millions of responses to really train your own model versus just fine-tune it a little bit.
Richard Socher [21:01] Certainly not from scratch. I think it's also interesting in that there's so many new open-source models that come out every week that it feels like if you now kind of focus on a model and fine-tune a foundation model from two months ago, there are probably better foundation models two months later. And so we were very careful in choosing what kind of foundation model we might want to rely on. And you want to be careful about the size of it.
Richard Socher [21:28] It's just beyond a certain size, it's harder to run it on a single GPU. But that will also change in the future as NVIDIA builds larger and larger-memory GPUs so that you can then run larger and larger language models on a single GPU. So that will change the equation, I think, again, significantly.
Matt Turck [21:57] Switching tacks a little bit and talking about business models and go-to-market: YouPro is a $9.99-a-month service. And as far as I know, you don't do ads on You.com. Is that correct? And if so, just maybe walk us through the thinking. Obviously, the most famous search engine in the world has generated many, many billions of revenue doing ads. So why did you choose to not do ads?
Richard Socher [22:37] Yeah, so it's a great question. Google is an ad company when it comes to revenue, and then everything else tries to fuel that ad engine. And so we tried very hard for quite a while to try other things, like YouPro. Truth is, people don't like paying for things, and they like it even less than advertisements. And so we feel like it's important to give users choice. And so we explored a few ads here and there, and we didn't find any ad partner so far that was trustworthy enough and created a good ad experience for our users.
Richard Socher [23:17] So we shut—we turned off some of the ad partners we had tried before. And now we're actually partnering with similar partners that DuckDuckGo partners with for private advertisements with reasonably high quality. And so we will, in the free version, support continued free usage through advertisements. And then if you use YouPro, we can turn the ads off.
Matt Turck [23:32] You.com. And I remember that at the time it was around being the best search engine for developers and tech people in general, but coders in particular. And you seem to have evolved towards a broader strategy. You.com?
Richard Socher [23:57] Developers who use You.com are way more efficient than developers who use Google. So if you're a developer and you're still using Google, you're shooting yourself in the foot and your own productivity.
Matt Turck [24:13] And just quickly as an aside on that, why is that? Is that because you have more connectors and third parties to code-specific kinds of resources? Why is that?
Richard Socher [24:24] It's best explained in a side-by-side comparison. Do you think I can share my screen really quick for your viewers?
Matt Turck [24:31] Yeah, we can do that for the YouTube version, and maybe walk us through it verbally for the podcast-only version.
Richard Socher [24:31] Sounds good.
Matt Turck [24:32] Yeah.
Richard Socher [25:03] I mean, here's why you're so much more efficient. If you look for some programming thing, like you want to use Python to implement a Fibonacci direct computation function, you can go to Google, get a bunch of lists of links, get People Also Ask, but basically don't get an answer. You have to open a bunch of tabs, and you go through those tabs, try to find the right answer. On You.com, it will just describe to you the formula. And then you have the code snippet, you have a copy-and-paste button right there, and you just saved yourself five, 10 minutes.
Richard Socher [25:38] And if you keep doing that many times a day, then you're just much more efficient if your search engine helps you and answers your code questions for you, helps with error messages more, and so on. And so that's kind of the answer. It's very intuitive when you see the result difference. You're like, oh wow, this is the code snippet I'm looking for. And so we're still working with developers. I think ultimately a good chunk of developer queries will move to be directly in the IDE, in integrated development environments that coders are using, but not all of them.
Richard Socher [26:20] You.com will be the best search engine for developers. We also saw a lot of students starting to use us this year, and it's a very useful resource for students. Again, citations, being more factual, being up to date. We also have various apps, from encyclopedia entries and Wikipedia and so on to other sources that are useful for students. And so we are doubling down on that. Part of the ambassador program that we have is with students. And then we're also seeing a lot of creatives come in.
Richard Socher [26:56] We have 10,000-plus images getting generated every day. If you imagine the latest Stability AI Stable Diffusion model, the YouImagine got even better. I mean, you can have four images now at the same time, download high-res versions of it and whatnot. And so we're seeing a lot of creatives too. So those are the three personas that we really work with: students, developers, creatives. And then across those, of course, there are also different countries. We're seeing a lot of growth in Latin America, for instance, which is very interesting to see.
Richard Socher [27:11] I'm excited. And so we're doubling down on that in some ways as well. And we'll show you more about that in the next week.
Matt Turck [27:36] Yeah, great. Looking forward to it. And doubling down means that you sort of go back and forth. You see a group of people that express interest, so you understand what else they may need and you develop those features, and then you create that flywheel, sort of community by community. That's what you mean?
Richard Socher [27:37] That's right.
Matt Turck [27:43] Okay, very good. And maybe talk about the ambassador program. You.com ambassador, what does that mean?
Richard Socher [28:13] Yeah, if you want to know about the latest and greatest features and then help us spread the word and help build the best search chat engine in the world, we'd love to work with you. And we have different folks. Some help more on the product side, some help more with community outreach. And, yeah, we're going to announce our first batch of ambassadors in, I think, also a week or two. A lot of things that were sort of—
Matt Turck [28:14] Busy.
Richard Socher [28:23] When things are a little bit slower, and especially students are out and about and getting ready for the new year. Great.
Matt Turck [28:41] Take your You.com hat off and put on your investor hat. So maybe tell us a little bit about AIX as an outfit, the story and the focus for anybody that may be raising money and looking for great investors in the space.
Richard Socher [29:04] Yeah, happy to. So AIX is, I think, the most knowledgeable venture fund when it comes to AI already. We are growing and just had the first close of Fund 2.
Matt Turck [29:06] Congratulations.
Richard Socher [29:41] And started to transition into Fund 2, and it's been very exciting. We have several incredible companies in our portfolio, and we're always looking to invest in AI companies started by founders who have deep AI expertise, but also industry insights and understand where the world is headed in a particular way. And so, just to give a few examples, I also transitioned my angel portfolio into the fund. And so one of my angel portfolio companies is Hugging Face. I think the founders also, they took my class back in the day remotely, so they weren't official students at Stanford, but took the class remotely and had a study group for the first-ever neural net for NLP class.
Richard Socher [30:19] And they were just so smart, had some really good ideas, and mildly pivoted also from those very early ideas into even better ideas, and have been growing from those first-round validations quite incredibly. And another company is Weights & Biases, help you train your own models now.
Matt Turck [30:24] Actually, just had Lukas on the pod last week.
Richard Socher [30:52] Yeah, he's super great. We have companies like Athelas. They measure blood samples and can count your red and white blood cells at home, but now moved into full-on oncology care at home, especially at a time like COVID. It's really suboptimal if you're immunocompromised and you have to go and take your blood pressure or other vitals inside the hospital. It's much, much better if you can do it at home. It's much faster. It became a companion diagnostic for different drugs and medications.
Richard Socher [31:20] So, incredible company. Those are all unicorns now. We also have companies in a broad range of applications. Again, AIX is kind of short for AI plus X, and X is a parameter. And there are all kinds of things. We think AI will change every industry out there. And yeah, I'll give you just a few more examples. One is called Machina Labs. They basically allow you to form sheet metal very, very easily and are growing very, very well and make it much more general-purpose.
Richard Socher [31:58] At a high level, if you're a car manufacturer and you want to have that exact piece of metal in this particular shape, like 10 million times, you build a special factory, massive machines that just give you tons of that exact form. But if you have the need to have maybe 500, it's incredibly hard to do because each one is almost manually built. And so you want to have an AI that is in between, can be much more flexible. And it's a little bit—it's not exactly like 3D printing.
Richard Socher [32:24] You can think of it as 3D printing, where you are able to build very solid metal pieces, but much more general-purpose machines that can still very quickly work for sheet metal and things like that. Another company is called Metha AI. They reduce the methane production of cows. Cows fart and burp a lot of methane. It takes a lot of energy of the cows, so they also produce less milk by producing all this methane, and methane is 80 times more potent as a greenhouse gas than CO2.
Richard Socher [32:49] And hence, it's much better for the planet. It's better for the cows, it's better for the farmers, and so it's a really great company, and they're doing very well also. And I love it.
Matt Turck [32:55] This is awesome. What is the AI part of reducing cows' emissions?
Richard Socher [33:25] Great question. Yeah, you actually have—it's pretty complicated. There's some other cow food supplement companies that just add algae, for instance, into the cow diet, and that reduces methane. But then what's crazy is, after a while, the cow digestive systems and stomachs update, and they start going back to their original methane. So it's actually constantly evolving, and you have to change depending on the cow. You have a fairly complex model and simulations and data collections and so on, but then build a model to change the diet adaptively every month for that set of cows.
Richard Socher [34:06] And so that's where the AI comes in, to actually understand and predict what the best supplement is to change the diet. So that's another company. And maybe one more example that I'm really excited about is Profluent. They use large language models to generate new protein sequences. And I think that will change all of medicine in the next decade in a massive way. Ali and a few others and I had started in 2018. We trained the largest language model for proteins five years ago, and now that work has really influenced a lot of folks.
Richard Socher [34:32] There are several companies in that space that have started. And yeah, there's some very exciting things in the pipeline there that I think have the chance to win a Nobel Prize at some point and certainly improve a lot of people's lives.
Matt Turck [34:44] And you guys are seed investors. Is that your favorite stage to come in?
Richard Socher [34:46] That's right. Yeah, we're pre-seed, seed. Yeah.
Matt Turck [35:11] Okay, great. You mentioned AI founders, or desirable kind of AI founders, as being people with AI expertise and industry insights. Is that how you think about it? Do you want to have, what, two founders, one with AI, the other one with insights? And how do you, I guess, measure AI expertise?
Richard Socher [35:37] Yeah, it's a great question. And there's no hard and fast rules. One thing that's beautiful about AI is that it's getting easier and easier for everyone to learn. The AI community has been very open for many years and publishes papers, lots of tutorials, a lot of classes online. My Stanford lectures from many years ago, and people learning NLP, were online already back then. And that whole field's gotten more and more professional too.
Richard Socher [35:57] And so it's very easy to learn. And the bar is getting lower and lower for people to get through and build impactful systems. Sometimes it's just a single founder who knows a lot about AI but has also spent a considerable amount of time building insights into a space. So we definitely have solo founders who have had some experience, built some great teams in the past, and then are able to apply their AI expertise to a new application area.
Richard Socher [36:34] But yeah, and in the case of Ali, he learned a lot of the biology over the years back when we were at Salesforce. He has some background that has also built large AI systems and knows them. So we do have some pretty incredible founders. There's just sort of interesting statistics. A lot of companies die because the founders don't get along, but the best companies are usually founded by two founders. And so there are a few exceptions like Amazon and Facebook, but the vast majority of really strong companies have had multiple founders.
Richard Socher [37:04] And so, yeah, we do look at the full founding team and sometimes also at their dynamics. I had once this dinner with some really incredible security researchers, and each of them was a rockstar. But after dinner, I kind of sensed some disagreements, and they kind of interrupted each other a little bit here and there. And so I put just a small check into the company, and indeed, a few months after, they just gave me the money back and said, "We couldn't quite agree on the vision."
Richard Socher [37:42] And so, yeah, it's the founders and the founder risk, and there's sort of the mental power and also just the sort of inability to give up, or an almost pathological optimism that you need as a founder to keep on going when things are tough. I sometimes euphemistically call the startup founding journey an emotional roller coaster, which is a euphemism for sometimes it's really like you're really far down and it's very depressing. And you just gotta keep on going.
Richard Socher [38:09] I actually love a lot of your tweets. They're very funny and spot-on. But yeah, so we're looking for that, and then, of course, the AI expertise. But there's also sometimes you have people with a ton of AI expertise, but they lack that user empathy and sort of common sense in the sense that they wanna, for instance, in the past cycle, speech recognition started to work, and all of a sudden everything was a speech problem.
Richard Socher [38:38] And I was like, oh, so what is your assistant going to do? And they said, oh, it's gonna tell you about restaurants and anything you wanna look for. And I'm like, well, let's double-click into this. Like, if I'm in a city and I'm like, tell me a good restaurant around me, is speech really the right system to answer that question? Like, the answer would then be, there are 200 restaurants in your city. One mile away, it has four stars on Yelp. Its favorite dishes are pad Thai and pad see ew. Or, and then like 10 hours later, you would have the list.
Richard Socher [39:01] Was that really the better experience than just looking at a map and seeing how far they are and intuitively knowing whether you wanna walk all the way over there if it's like a Michelin star or something really nice? And the truth is, even though I love AI and I do think it'll change pretty much every industry out there, you do need to still be clever and have a lot of user empathy and think about the market, think about your go-to-market, think about your distribution.
Richard Socher [39:39] And all the potential challenges in that. And we hope that not every founding team has to know everything in the beginning. We've had some pretty incredible pivots already also in the portfolio, but they need to have strong convictions, loosely held, go after something, try it hard, see if it works or not, and then adapt if they have to.
Matt Turck [40:15] Great. This mental model that a lot of investors thinking about the generative AI space in particular seem to have these days, which is effectively a three-layer cake where you have the LLM foundation models at the bottom, in the middle, I guess middleware, for lack of better terms, like anything from LangChain to prompt engineering kind of systems, and then at the top, application-level kind of opportunities. As an investor, what do you think is interesting, less interesting from your perspective?
Richard Socher [40:59] Yeah, it's a good one. I do think there will be probably only a few dozen horizontal companies that are going to be really major breakout successes, but there will be thousands of vertical applications, what you call it, sort of vertical companies that are successfully using AI for their industry. Across the full stack, like Hugging Face, Weights & Biases. We've also been in the first seed-stage round of Chroma Vector Database. They already had an up round as well. And so we do have a few strong players in that space.
Richard Socher [41:14] But we are looking, and most of our investments, I think, are, or yeah, I would say are in that vertical space.
Matt Turck [41:24] Yeah.
Richard Socher [41:31] Wildlife, e-commerce, cow methane production, and everything in between. Yeah.
Matt Turck [41:55] And maybe to wrap up and zoom out with a big theme, and I'm sure we cannot do it justice in just a few minutes, but let's give it a go. Where do you think we're heading? Are we having a moment now and things are gonna just taper off a little bit, or are we sort of early in that exponential moment? The related question is, is that a question of just feeding ever more data to more powerful models?
Matt Turck [42:19] And does that get us to this continued acceleration? Or at some point, do we get to a stage where, well, there's only so much data we can feed, or the models can only do so many things with whatever amount of data we feed them?
Richard Socher [42:44] Yeah. I actually think there's sort of this high-level exponential that we're on in terms of all of the AI capabilities. And there are sort of sub-functions of the overall progress. And when you start at the beginning of an exponential, you never know how long it's going to last, right? And people have, I think, in some cases, inflated expectations. And I think it's really hard to navigate. And it's part of why I'm trying to write a book on the side also about all of this.
Matt Turck [42:52] Yeah, because you have some free time on top of all of that.
Richard Socher [42:56] That's right. Yeah, my Saturday evenings, not enough things to do.
Matt Turck [42:58] From 1 a.m. to 3 a.m.
Richard Socher [43:24] Yeah. And so basically it's really hard to navigate because AI is so exciting. There's so much change and so many improvements and breakthroughs that are happening. So many industries can be influenced, similar to electricity. You can apply it to many different industries out there. At the same time, I think people are a little bit overly optimistic about, oh, it'll be self-conscious, or it is already self-conscious. Oh, it could replace all the jobs or destroy us all, and all of these things, be even more intelligent than we are.
Richard Socher [43:57] And it is still very much a tool. It's an incredible tool, but it is a tool, and it will stay that way for a while. And so I think we can already, for instance, see a flattening out of that exponential. And most exponentials eventually flatten out and become sort of S-like curves. And one example is in image generation. If you look at Midjourney, where do you go from there? Like, it looks perfectly photorealistic. You can't be 10x superhuman, like, photorealistic, right?
Richard Socher [44:15] There's only one reality out there, and humans have only so many eyes to look at it and so many cones and whatever. Like, this is what it is. And so it's flattened out. And now, in the grand scheme of things, you can go to videos, right? And have longer features. And there's some interesting research that is required in making diffusion models have some discrete, time-insensitive components where you have, like, no variance, so that the same person that started in the video five minutes later is still that person and looks the same.
Richard Socher [45:00] And no one has figured out some of those research breakthroughs yet for longer video generations. But I'm sure they will. Similarly, human language, it's hard to be 10x different or better when it comes to human language. I mean, sure, you can translate better than the best translator in 50 different languages, because no human can translate 50 different languages really well. But ultimately, language is consumed by humans, and we need attention for it, and humans can only read so much.
Richard Socher [45:27] And so fast. And so even if you can write faster than a human, humans can't consume it that much faster than before. And so there's certain limits that people think, and they say, oh, this AI could be so smart, it'll influence all of us to potentially be this existential threat to us. And I'm like, I think humans are smarter. If you really think that, like, humans would be influenced by the smartest person out there all the time, I don't think that's what history told us.
Richard Socher [45:58] is actually the case. And so anyway, there's a lot of fun stuff you can talk about there. It is a hard space to navigate, a lot of nuance, a lot of interesting facets. And it's all in the backdrop of, indeed, some industries will get massively changed, which will mean that some industries will be more and more expensive. Like, probably at some point, your plumber might get more than your marketer because the marketer can use tools and be way, way more efficient.
Richard Socher [46:21] But the plumber AI is gonna take a very long time because none of the plumbing companies are collecting data, having robotics. Physical robots are still very, very hard. Yeah. So it'll be an interesting next couple of decades.
Matt Turck [46:28] That's for sure. Fantastic. Well, that feels like a wonderful place to leave it. Richard, thank you so much. Appreciate it.
Richard Socher [47:04] Thanks so much for having me. Have a good one. Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.