Poolside: AI for Software Development with CTO Eiso Kant
The MAD Podcast with Matt Turck · with Eiso Kant, CTO and Co-Founder, Poolside
Eiso Kant is the CTO and Co-Founder at Poolside. We cover why executable code lets models learn from scalable feedback rather than human evaluation, why Poolside expects most developers to adopt AI pair-programming assistants within 24 months, and why natural-language planning helps models find solutions that single-shot code generation misses.
Chapters
Transcript
Full episode
Matt Turck [1:02] Welcome to The MAD Podcast. So you are the co-founder and CEO of a very buzzy, very ambitious generative AI startup company called Poolside, and you are building, and I quote, "the world's most capable AI for software development and the applications to unlock the potential of developers." The company is just a few months old, but you've already raised $126 million in seed money. And this is actually the first time, I believe, that you all are talking publicly about what you've been building at Poolside.
Matt Turck [1:17] So we're particularly grateful that you're joining us today and very much look forward to the conversation.
Eiso Kant [1:26] No, no, thank you so much for having me, Matt. And we've known each other for a while now, so it felt like a good first place to have a long-form conversation about what we're doing at Poolside. All right.
Matt Turck [1:29] Wonderful. So first things first, tell us about you and your journey.
Eiso Kant [1:53] Yeah. So I'm essentially a computer geek. I started programming when I was about 10 years old through the graces of someone in my school who was crazy enough to write an operating system, but smart enough not to want to write the drivers. So he taught me how to program, and that unlocked a passion that's kind of held true throughout my entire life till now, of building software. That's gone through different journeys. I found myself founding companies for a while. There's probably one in particular that's very relevant for Poolside in today's conversation, and that was Sourced.
Eiso Kant [2:18] So back in 2016, I was building Sourced. To the best of my knowledge, we were the first company in the world that was dedicated to applying deep learning to source code and software. And so this was in a time pre-transformers. We ended up starting with RNNs and then graduated to LSTMs, and we went down a whole host of different solutions. And it's also a little bit how I met my co-founder. I'm actually not sure if we've talked about this on any recording ever, but Jason—for those of you who don't know—is my co-founder.
Eiso Kant [2:49] In 2017, Jason had just become CTO of GitHub, and GitHub made an acquisition offer for Sourced. At the time, Jason and some of the other members on the team already had a forward-looking view that AI was going to become highly relevant in the process of building software. And we had a lot of things running, things that today look like Copilot, where you could get code suggestions based on natural language instructions. And while our board couldn't get to an agreement on the number with GitHub, which I like to joke with them they later regretted when GitHub sold like a year later for an incredible multiple.
Eiso Kant [3:16] Nonetheless, Jason and I became friends, and that led us to building a friendship over the last six, seven years, and also our own podcast. And yeah, that gets us to earlier this year when we started Poolside. But before I dive into that, happy to take guidance from your questions here before I narrate like a founder.
Matt Turck [3:20] No, I think that's great. Maybe just another word on Jason's background.
Eiso Kant [3:40] So he was CTO of GitHub, and before that, he was the engineering leader, VP of Engineering at Heroku. And before that, he was at Canonical, running a big part of the Ubuntu Linux team. Both of us have had a long-term passion for software development, how we do it, and for developers themselves.
Matt Turck [3:52] Yeah. And after that he was a VC at Redpoint. So it's an example of somebody who was a VC and went back to operating, presumably getting very familiar with the joys of being a VC on a daily basis.
Eiso Kant [4:07] Yeah, I got to see that journey up close as a friend. And I think Jason, at heart, is a builder. While I know he really enjoyed his time at Redpoint and it gives you this broad view of what's happening, I know at the end of the day, Jason wants to get in and build. I think that sounds very similar. And yeah, that culminated in us having conversations at the beginning of this year, which led, by the end of April, to us starting a company together.
Eiso Kant [4:15] All right.
Matt Turck [4:23] So let's jump right into the overall vision for the product. What is it that you guys are building, and why did you choose that specific problem to go after?
Eiso Kant [4:39] So I think it's worth taking a step back and kind of looking at where Jason and I found ourselves at the beginning of this year. We like to jokingly say amongst ourselves that Poolside exists because of an OpenAI-induced existential crisis. And we are incredibly deeply grateful to them for it. I felt a very long-held belief, and you can find old YouTube talks of me talking about this, that it's very likely that in our lifetime, neural networks will become capable of learning anything and everything that we are capable of as humans.
Eiso Kant [5:16] And earlier this year, Jason and I found ourselves in long conversations, seeing that we were on a trajectory with everything that had happened with ChatGPT and subsequently GPT-4, which I think was even more impressive, that was putting us on a trajectory where this is very likely to occur in our lifetime, and probably rather sooner than later. I'm notoriously careful with putting timeline predictions on such big events. While others might say five or 10 years, I like to reserve a little bit of room for seeing what's going to happen.
Eiso Kant [5:49] But there is something that a lot of people will hear this and say, okay, yeah, I believe that that's going to happen in the next five or 10 years, but then they're not changing their lives. And the truth is, if you really believe that we will create systems in silicon that are above and beyond our own capabilities and skills and intelligence, that will become more economically valuable and useful than we are ourselves, it really warrants a genuine existential crisis. And I think there's a gap between people who are saying so and then going on with their lives as if nothing is happening, and maybe rightfully putting on blinders, versus probably Jason and I.
Eiso Kant [6:20] We really got to a point where we said, okay, if you believe that, you find yourself kind of with two paths in front of you. Do you decide to say, "I've gathered enough nuts as a squirrel. I'm gonna do just the things that I'm passionate about in life, right? I'm gonna sail or paint or things like that." Or do I want to try to contribute to that future? Do I want to try to push forward what we are capable of, try to accelerate that timeline?
Eiso Kant [6:46] And that's where we ended up. So it really started from the point not of the problem for a product or developers. It started really from this point of, how do we help the world get to sufficiently advanced intelligence in artificial systems quicker? Now, the next set of discussions, which lasted several months, were: do we have our own unique point of view on this? It's one thing to contribute to the world and do what everyone else is doing.
Eiso Kant [7:13] It's not particularly useful to try to follow the same path OpenAI or Anthropic are following. And the world's already doing that. So the more Jason and I discussed, the more we realized—and this is obviously biased by our backgrounds, but hopefully in the right direction—that we actually believe there's a lot of value in focusing on a single capability instead of going general-purpose. Today, with models like GPT-4 and Gemini, their whole notion is to be general-purpose, useful for everyone.
Eiso Kant [7:43] We kind of took the step back and said, what if we only focused on software development as a capability? Now, on one hand, that comes from software development, in our opinion, being a pretty good proxy task for a big part of the spectrum of intelligence. You need an understanding, a model of the world, because we build software for the world. You need to be particularly strong at reasoning, and you need to be absolutely excellent at planning, right? These three major pillars that I think are a good way to divide artificial intelligence, or intelligence in general, up into.
Eiso Kant [7:52] Well, it doesn't cover everything. I'm maybe already jumping ahead to people's comments, like, it's not a robot, it's not embodied, but our view is that if you're able to solve for a person coming with a higher-level objective about something that they want to see in software, an AI-led system that's able to ask the right questions to then bring you to a full piece of running software that you've created together with the AI, you've had to solve for reasoning, you've had to solve for planning, you've had to solve for having a good model of the world.
Eiso Kant [8:40] But all of this is in the nice theoretical realm. There's a very specific technical detail that made us say, okay, let's dedicate our lives to this. Let's go into this capabilities race that seems almost impossible to enter into when you have juggernauts like Google and OpenAI with massive amounts of funding. And the particular technical detail that, for us, is what the company's really built around is this notion that source code is one of the very few things that we generate with neural networks.
Eiso Kant [9:10] That we can actually execute and introspect. It doesn't require human feedback to evaluate it. It doesn't require humans to step-by-step reason through it. We have built compilers and interpreters, right? We actually have the ability to run code. And the ability to run code, for us, is something that opens up this massive possibility to improve these models. Well, let me take a pause here because this is really what got us started, but there's a lot that went between there and where we are today.
Matt Turck [9:32] Yeah, very exciting and very ambitious, but indeed broad. So the obvious next question is, how do you narrow down the aperture so that you can turn this big vision into an actual product?
Eiso Kant [9:54] Yeah, well, I think it's first noting that if there's one thing I've learned over the years, it's sequencing, right? It's not enough to be able to have a big vision. The sequencing that you follow as a company ends up really determining if you're around to try to have the opportunity or even the chance to go for that big vision. So for us, the sequencing, we looked at it from a couple of angles. One is, what do we think is happening in the world in the coming years?
Eiso Kant [10:23] Well, the first thing that I think is at this point almost uncontested, but frankly, when I was building Sourcegraph, most people didn't think would happen, is the fact that the majority, if not all developers in the next couple of years in the world will adopt some form of an AI pair programming assistant. As these systems and the models behind them and the applications become more capable, more useful, it's going to feel incredibly subpar. It'll be like trying to work without tools, almost, to be able to create software.
Eiso Kant [10:47] And so, knowing that every developer very likely in the next 24 months is going to adopt an AI pair programming assistant, this was really the part that we said, okay, this is a good place to start, right? This is assisting people. We're not at this autonomous future yet, but it forces us to focus on solving for the big things we need to solve. Because if developers are comfortable and trust that they can give instructions to a system and the system reliably does it, fixes the bugs for them, writes the code for them, writes and runs tests, et cetera, sets up infrastructure, then you are getting in the direction of having to solve for reasoning, planning that can get us to a more autonomous future.
Eiso Kant [11:33] So we see this as kind of a two- to three-year window where it's really around assisting developers, but the rate at which these underlying models are increasing in capabilities, we don't think we're more than three years out to that point where a non-engineer, someone with no coding, software development background, is able to come and give one of those high-level vague things of what they want built, right? Because that's what we kind of do as humans all the time. Like, I want an app that does X, Y, or Z.
Eiso Kant [11:50] And then a system that can then actually ask you the questions to refine what it is that you want to then actually build software for you. And that first couple of years of an AI assistant is powerful, but the world isn't standing still, right? We've seen lots of startups and people inside companies building on top of these models, trying to build more autonomy in them, trying to build agents, trying to not just make it work as a pair programming assistant, but trying to drop this everywhere in the software development lifecycle, which is why also it's important to us to see ourselves as kind of a company that does two things.
Eiso Kant [12:32] We deliver this AI pair programming assistant, full application stack with the models, I think, fully vertically integrated with our own foundation models trained from scratch, but then also exposing those models with an API to allow other people to build all the other experiences on top, because we think there's just too much to be built on what a single company can do.
Matt Turck [12:43] Okay, great. So are you starting with something specific, be it a language or frontend or backend, or is the idea to go very broad from the beginning?
Eiso Kant [13:07] If we work our way a little bit up the stack, the first thing for us that was really important is that the capabilities of these products are determined today by the capabilities of the models underneath at about 80, 90%. There's some things we can do with the application layer to improve the experience. But fundamentally, who's still talking to Siri versus who's talking to ChatGPT, right? And so we treat the world as kind of two races that are happening. There's a race for users, right?
Eiso Kant [13:31] A go-to-market race. How do you get in front of every developer in the general domain, but also in enterprises? And then there's a capabilities race. And you do not want to be Siri while OpenAI is bringing out GPT-5 or Google's coming out with a more capable model. Because your whole business falls apart because, at the end of the day, right, why we build products is to deliver value to the end user. And the nice thing about developers is that they're the first people to leave you if something else is more capable.
Eiso Kant [13:46] And AI is only going to get so good at writing a sales email that the salesperson is not going to get a differentiation anymore between one model and the other. But we have a huge journey ahead of the capabilities of these models to assist in software development, that that difference in capabilities makes up for a big part of the experience, which, for us, is why the starting point was: let's build a general-purpose model fully oriented only on software development as a capability and completely from scratch.
Eiso Kant [14:32] So, from determining our own training data to even the architecture of the model, to how we then improve it via reinforcement learning from code execution feedback, something we'll end up talking about more, and a lot of other areas. But in terms of what we're delivering to the end user, we're starting with a direct competitor to GitHub Copilot X. And so, a natural language, multi-turn dialogue system where developers give instructions of what they want done while they're in their editor. That's our starting point and what we're going to be bringing out next year to users.
Matt Turck [15:05] Okay, so thank you. Let's dive into how it all works. So you mentioned you're building and training your own LLM. And as I was prepping for this, I read that part of the idea is to allow, and I'm quoting, "allow your LLM to improve by completing millions of tasks in tens of thousands of real-world software projects." And you call this approach reinforcement learning from code execution feedback. Do you want to get into this and explain what that means?
Eiso Kant [15:27] Yeah. So we can kind of look at the training of our model as two parts. One is the actual training of the base model. So everything that we do there, and we can dive into that. And then after we've trained our base models, what do we do to really improve them specifically for being more accurate at going from instructions to code that is useful to the end user? We throw this under the umbrella term of RLCEF, reinforcement learning from code execution feedback, but it really comes down to an idea that even started back in source{d} in 2016 or 2017, which is that it's not enough to show these models, essentially have them learn how to code from showing them massive amounts of examples.
Eiso Kant [16:13] And I want to emphasize one thing: our models are not only trained on code. Our data split is about 50-50 between language and code. We process web-scale data, but our point of view is that, well, if you treat pre-training, the training of a base model, as kind of the model reading the textbooks, when does it do the exercises at the end of the textbook, right? And today we don't. We've heavily invested in the world in reinforcement learning from human feedback, right? Human preference, which allows us to say what answers we prefer, which to some extent is a little bit like judging the answers to questions at the end of the textbook, but it's not actually asking these models to complete the exercises.
Eiso Kant [16:52] So what we have seen from our work is that if you take a large language model, a sufficiently good base model, and you give it an instruction to write some code or fix a bug that it's not able to successfully complete today, very often with the right guidance and environment, you can get it to successfully complete it. So how does that manifest in practice? We have an environment that we're building. Actually, we should probably update the website soon because we will go past tens of thousands of real-world software projects, almost already are.
Eiso Kant [17:30] We're already at the scale that is 20K-plus, and we'll be scaling up even further. What we've done is we've taken real-world, high-quality software projects. So think libraries across a multitude of languages, across different domains, projects that we have seen have been created by developers who have been very active and collaborated with other people who are very active in the community. And we've made sure that we understand the relationships between the code and the tests that cover them. Right? So as developers, we write massive amounts of tests, starting with unit tests and then integration tests.
Eiso Kant [17:58] And these unit tests have coverage, right? They cover certain pieces of code. And unit tests are incredible because unit tests are fundamentally rules that we have written as humans to see if the code successfully does what we intended to do. So you can probably see where I'm heading here already, right? We have these millions and tens of millions of tasks, and next year scaling up into the hundreds of millions, that developers have previously done for which they've written rules to test them.
Eiso Kant [18:28] Now, this is where we start. We take advantage of that. We synthetically generate instructions, instructions being tasks, things like write code that does X, fix this bug, et cetera. Synthetically create these extremely large instruction datasets. And we ask the model to try to complete this instruction. Now, this is where you could already make a lot of headway. If a model is able to complete it or not, the next step is giving it the feedback from the environment, which is the interpreter or compiler, right?
Eiso Kant [19:00] I run my tests, the code doesn't compile, throws an error, or only two out of five tests pass. But this is still pretty sparse feedback and sparse rewards. And it's also a singular back-and-forth, right? It's the model and it's the interpreter, and it's going back and forth. And this is useful. This helps models. And you can even try this yourself with any model that's out there. Right? Usually, by pasting the error from your interpreter or compiler, it might get a step further in its reasoning.
Eiso Kant [19:27] We've taken this several steps further. On one hand, we've made massive investments in planning algorithms that are very unique and specific to Poolside that allow us to actually efficiently explore different solutions to a task. Think here a little bit of what we see with AlphaGo and these kinds of systems where there's an exploration step. The way that we tackled this is particularly interesting because what we've really learned is that if we think about how we as humans solve a programming instruction or task that we're given, we don't explore seven variations of code straight away.
Eiso Kant [20:03] We actually build these kind of natural language plans in our head. We start writing code, we start thinking a little bit, okay, I'm going to do this next. We build essentially a set of steps. We use data to inspire what we call natural language planning. So for every single problem or instruction, we'll explore different possible plans on how to solve it. And this is great because this solves for an entropy collapse problem that you get if you're trying to get large language models to probabilistically explore code.
Eiso Kant [20:38] At some point, it's just going to put out garbage. And then for each of these plans, we explore different code solutions. There's some secret sauce that I won't go too much into that goes beyond this, but fundamentally the idea is that we built this extremely large, constantly growing real-world dataset of tasks that developers have already completed, many of which we exclude from our training data—not all, but quite a few—so that we have a good ability to understand how good these models are getting.
Eiso Kant [21:07] And then we use the model to explore both a natural language solution space and a code solution space. And what we have found consistently is that the models are able to start finding solutions to instructions that are not in the training data, that are more diverse. Often they'll find multiple solutions, different ways with code, and even often sometimes radically different ways to find a solution to the task that was given. But more so, they find more and more solutions by having this exploration and planning algorithms.
Eiso Kant [21:40] What before, with even a single-shot or a multi-shot back-and-forth, would be incapable of finding that answer, now it comes up with the answers. Now, this is, of course, highly reliant on how good your base model is, right? Your base model has gained a set of knowledge and some limited skills like reasoning and such. And if there's no good reasoning capabilities in your base model, which also comes from language and not just code, you'll never find the solution. We use all of this data with a whole set of rewards and reward models and a couple of other things that we do that we're a little bit more, say, not as transparent about, to then feed back to the model when it succeeds, when it fails.
Eiso Kant [22:19] And it's that RL loop that is very interesting because, since it's programmatic, since we have an oracle of truth, we can scale this up far larger, right? Magnitudes larger than what you can do with human feedback today. And it changes a little bit the mindset, right? We often think about reinforcement learning from human feedback as not really teaching the model something new, more kind of correcting it or aligning it to our preference. But our view is, and we plan to show this as we scale up even further, that reinforcement learning can teach these models new things.
Eiso Kant [22:44] It can essentially take away the fact that it's had so much confusing data during its pre-training that we can actually get to a point where these models are becoming better and better at going from instructions to code that works. And that's kind of the starting point to build on top of.
Matt Turck [22:54] Are you finding that the model solves problems in a way that a human coder would not already, or do you expect to find that in the future?
Eiso Kant [23:19] One of the things that is, I think, surprising for people is that the model is often more eloquent or stronger at explaining what the instruction or problem is. This is one thing that really surprised me. Synthetically generated instructions based on ground truth of code, by almost all accounts, feel better than what most people, engineers, would write themselves. This is also, by the way, a problem to the end user, right? Because the end user will not write perfect things. They'll be flawed in what they ask from these models to do.
Eiso Kant [23:46] And so I think there's a lot of progress that needs to be made for these models not to just come with instant solutions to what you ask, but also to become more and more capable of asking for additional information. Because often we are the flawed part of the process. It's actually quite hard to voice an instruction of what you want without going all the way to either writing pseudocode, and then you might as well just write the code. There are solutions that you see that these models come up with that just range across a diversity of coding styles.
Eiso Kant [24:12] Also, some things that you would just, like, they're functionally correct, but you're like, oh, I don't like this. But also the opposite side. We're right now still at the realm where they're solving the kind of problems that humans are very good at solving. So I don't think we're at a space yet where we can start seeing truly incredible feats of programming. And this has a lot to do with we're still really living in the realm of kind of function-file-level, what these systems are able to create.
Eiso Kant [24:30] We're not yet at the place where we can truly have generation that spans to something far more complex. I don't think we're that far away from it, but there's some work to be done there.
Matt Turck [24:58] Yeah. So you alluded to some of this, but maybe to double-click, let's talk about the data aspect of this, in particular synthetic data generation, which you mentioned on your website that, while seemingly counterintuitive, it works and works particularly well for code. So your dataset—you mentioned it's not just code, there's some language as well. So is there a part of just crawling existing repositories? And then it sounds like there's a part that's generating the data. How does it all work?
Eiso Kant [25:24] Yeah, so I think first we spend way too little time in our space talking about data quality. To me, the one thing that's very interesting, and I think if you look at what gets leaked at the major companies, at OpenAI, at Anthropic, what gets leaked: the architecture, the attention mechanisms, et cetera. But one thing you can barely ever see a leak about is: what's the data? What are we doing to the data? How are we gathering it, et cetera, right? And it makes sense that these are kept as very close-guarded secrets because fundamentally, the one rule of machine learning that's held true 20 years ago, and it still holds true today, is garbage in, garbage out; quality in, quality out.
Eiso Kant [26:01] But it's easy for us to forget that with large language models, because even when we throw a lot of general data at it, including a lot of garbage, these models are particularly good at learning and they'll spit out answers that are still highly relevant. And then we can steer them a bit with human or AI feedback so that we can oversee the fact that, well, I've given training data to this model that's highly contradictory, that's messy, that includes everything from well-formed arguments to 4chan comments, to make maybe an analogy.
Eiso Kant [26:29] So our point of view is that not every token is equal. Not every sample of your training data is equally valuable. This sounds like a lot of common sense, right? It's the same thing how we learn as humans. If I give you a high-quality textbook in a domain, it's going to be a lot better for you to learn from than if I just throw 500 random papers at you that are roughly related to it. So how do we get to that good-quality data?
Eiso Kant [26:56] Well, our belief is that it needs to start with trying to gather all the data, trying to gather everything that's ever been published on the web, everything that's published in open-access research, trying to build as large possible datasets as you possibly can, which also makes it far more comfortable for you to filter down, right? Because there is an amount of data, there is a clear relationship with the duration of training. And we undertrain our models. So training it on more data in the right way almost always has your loss go down, right?
Eiso Kant [27:22] It improves the capabilities of your models to come closer to the original distribution, which doesn't, by the way, always mean closer to the actual quality. Because if your distribution of your data is garbage data, well, then it doesn't really matter how good your loss function is. But fundamentally, yeah, it starts with trying to build these very large datasets. And we do that, right? Essentially, think here: the crawled web, special data partnerships that give us more access to data that isn't publicly available.
Eiso Kant [27:49] And then in the case of code, it's actually very straightforward, right? We have massive amounts of Git repositories. There's over 330 million in the world, 151 million if you remove forks. And from that, if you filter down, a subset of that is quality. And I guess the next logical question is, what's quality? We think that humans are very good at judging the quality of training data. You can look at a piece of text and you yourself are going to be pretty good at judging if this piece of text is likely to add something to what the model can learn from.
Eiso Kant [28:21] Is there reasoning in here? Is there knowledge in here? Is there something that will help a model learn? Or is it pure garbage? With code, interestingly enough, it's for an engineer capable to do the same thing. We can look at certain files and pretty quickly reason to the fact that this is unlikely to teach the model as much as another file, right, comparably. So our starting point has always been humans are good at judging this. Now, the other thing is, if humans are good at judging quality and high quality, then should systems be good at judging the inverse of that, right?
Eiso Kant [28:52] Like, if once you know what high quality is and what high quality is during training, is there data that lives in your training that allows you to understand the inverse of what's quality? Part of that's what humans can label, but part of that also starts becoming something you can actually automatically infer. We don't go too much in depth here into some of our mechanisms because this is some of our quote-unquote secret sauce. But I think for anyone in this space, you'll very quickly understand that the more that you invest in building datasets that determine what's quality and what's not, you can end up, of course, using that data to then train models that become very capable at making that decision for you.
Eiso Kant [29:28] And I think I'm mentioning all this to your question of synthetic data generation because I think it starts with: what's the data that exists in the world? How do you get as much of it as possible so that you're comfortable filtering it down? Because everyone's biased by the fact that they need a certain amount of tokens at the end to train. Then the next step becomes, how do I judge this data? And there's rule-based systems that are very useful, and we all build them.
Eiso Kant [29:47] But then at the end of the day, we actually want to use the capabilities of these large language models to do some of this judgment for us. Because, right, we operate in the space of tens of billions, even 100 billion-plus documents. Right? So it's an extremely large filtering process. And then the next step comes, which is that you'll very quickly realize that a lot of the data you have looks something like this: bad, good, bad, good, good, good, bad in the same document.
Eiso Kant [30:16] This is very often not because of something profound. It's because a bunch of text was scraped from the website, ended up in the wrong place, or the bibliography is not really relevant for the model to learn from when it sees a research paper. Right? It will just end up outputting random references. So there's a lot of benefits. If you would draw the midwit meme on essentially large language models and data, on each end of the spectrum of the midwit meme, you would have: just look at lots of samples of data, right?
Eiso Kant [30:44] In the middle, you would get some crazy academic theory. We spend a lot of time looking at the samples of the data and then determining how to get models to make that judgment. But now you get to the next step. If you know that you have something that has partial good and partial bad data, or it's good text, but it could even be done better, right, for a model to learn from, that's where you start getting into this slight area of synthetic, right?
Eiso Kant [31:12] It's where you start getting into the realm of: can I use models to modify this data? You're not fully asking for new generations yet, which is what I would put under the headline of synthetic, but you're asking it to start modifying and improving it. This is to a large extent also—there's been some amazing work on distillation in the series of Phi papers by Sebastian Bubeck and his team that show this on a smaller scale. Like, what happens if you start asking for textbook-quality data?
Eiso Kant [31:39] All of a sudden, you see that even smaller models start becoming capable of learning things that previously they weren't. So the next step is modifying it. And now this is where a lot of compute goes into. I think that when you're dealing with tens of billions of documents, and you try to start modifying data and you start using models, your inference cost gets very expensive. So you have to kind of make this trade-off of what we're able to do. But at some point, you very much see the path that newly generated documents by models can make for great training data.
Eiso Kant [32:08] Now, there's a lot of people here who will then say, ah, but it's impossible. How can a model learn from something that was improved from its training data? I like to argue that we do this all the time as humans. We'll have professors, and they'll read a whole bunch of stuff, and then they'll write a textbook. And that textbook is then what we use to teach people. But with code, you can't argue with it. And what I mean by that is, to our earlier point of reinforcement learning from code execution feedback, when I see our model combined with planning algorithms exploring a space where it's finding novel solutions or new code that wasn't in our training data to solve for an instruction, that's pure synthetic data, right?
Eiso Kant [32:46] Like, no one can argue that if a piece of code that solves for a problem in a codebase that you can search and not find back in the training data by any similarity measure, that it is synthetic. And I think this is easier for people to reason about when they see this with code. And then you very often get, yeah, but that will never happen in language. The process is the same thing. I think with images, we see this as well.
Eiso Kant [33:03] There's a point right now where you might look at something that's generated by DALL-E or Midjourney, and you can no longer argue that this is a valuable picture, right? It's unique. It didn't exist before. So why shouldn't it be training data? And so this is kind of our view. Now, if you extend this a couple of years out, and this is where some people will probably have objection with us, we think that this will go so far that we will kind of recycle all of the world's real data into higher-quality new synthetic data to a point where you can't reconstruct source data anymore.
Eiso Kant [33:43] And us running out of data in the world isn't actually so much of a problem. We will just end up generating from essentially recycling it, cleaning it up, kind of like how a professor will read documents, articles, and then turn it into something of a higher quality. We think this will happen at scale with pre-training. The reason it doesn't happen today is because it's too compute-expensive, but the world solves for that with time. Great.
Matt Turck [34:05] So it's becoming reasonably well known in the world of generative AI that, in addition to the model and to the data, a lot of the core work that needs to take place to make things happen has to do with engineering rather than machine learning. Do you want to talk about what you all have done or plan on doing around the engineering behind the company?
Eiso Kant [34:29] Yeah, so it starts with culture, right? Who do you hire? Who do you bring on? And first of all, I want to say here, I think there's different paths that can work. Just like I think there's different paths towards AGI, one focused on general purpose, others focused on software development like us. I think there's different paths of building teams and cultures that can get to building these incredible artifacts and products and companies following. The path that we believe in is: in the early foundation of the company, every single person who's here needs to be a strong engineer if they are a researcher.
Eiso Kant [34:57] Right? It's not enough to just be a researcher in the, "I live in a Python notebook and I can kind of make things work," but I can't do good engineering at scale. And the reason that is, is because it allows people to go from idea to an implementation. And most of the valuable things that you learn with large language models and reinforcement learning at scale, right, trying to make them stable, no matter if that's on your pre-training or on your RL and others, come from the actual productionization.
Eiso Kant [35:29] And that requires good engineering. Now, the other thing that we've seen that I don't think is spoken enough about in our space is pairing up really good researchers who are engineers with very, very strong distributed systems engineers is a magic formula. Because then, all of a sudden, when you're building your data pipelines or you're building your pre-training or things like this, you are not just looking at it from the more researcher point of view. You're also looking at this from, hey, this is a large distributed system that needs to scale.
Eiso Kant [35:57] It needs to go through different sequences of maturity over time, right? We can't build the perfect end system straight away because we do need to get things out. And this is probably, for us, one of the best decisions we made early on, was by bringing in some absolutely incredible distributed systems engineers: a person who was a responsible staff engineer for Uber Payments, a person who built his own cloud company over 11 years, and just several people like this who are extremely strong engineers.
Eiso Kant [36:22] And what you end up seeing is that you can then start getting to iterations of your systems, your data pipeline, your code execution environment, your distributed RL, your pre-training, that are from the ground built up to scale. And it's not scale for the sake of scaling because we're engineers and we like to build scale. No, but to say that everything we do today is already at a very large scale. The other thing that surprises a lot of people about us technically, and again, multiple paths to the same goal, is that we didn't take Megatron or DeepSpeed or the big open-source frameworks for training large language models as our base.
Eiso Kant [36:55] We built our own from scratch. We built our own distributed training for our pre-training stage. And this really was for a very specific reason. If you want to be on the frontier of models, right, we are, we want to beat OpenAI, Anthropic, Google. We very publicly will say this: we want to be better than anyone else at software development, right? This is, to your quote earlier, trying to build the world's most capable AI for software development.
Eiso Kant [37:24] You need to, from the ground up, understand your architecture. You need to be able to modify it. You need to be able to try things. And there's so many layers that we have built into some of these frameworks because we've generalized them to lots of different model types. We've had lots of people contribute to them that you don't really know why your loss is exploding anymore if you're using Megatron. Is it because it's a wrong implementation? Is it because of the architecture?
Eiso Kant [37:51] Is it because of the containerization that people put underneath? And you can just keep—like, it's likely not the containerization, but you can kind of see that there's a lot of complexity. And so we built this from scratch. This also led us to do something that is another big surprise, is that we've decided to scale up RevNets, reversible layers. So, to the best of our knowledge, I think we're training the world's largest models using reversible layers, that we've seen at least, and nothing else that we're aware of publicly.
Eiso Kant [38:17] And this is something that a lot of people were skeptical of, but on our side has proven to work incredibly well and has opened up a path for us of scaling large language models pre-training far more compute-efficiently than what we've seen some of the more traditional architectures like LLaMA and others have been able to do. So, choosing to make some of these big bets early on on the engineering, having people who truly understand it full stack, having researchers that build things end to end from their idea to working, all of those things add up to giving you a lot of velocity.
Eiso Kant [38:32] And, as hopefully we'll show in time, add up to building things that are of really high quality.
Matt Turck [38:53] And speaking of engineers, and switching tacks a little bit, one of the things that caught people's attention in the days of Poolside, in addition to all the above, is that you guys decided to build the company largely out of Europe and France, although I'm sure you have engineers elsewhere. Can you talk about what is the reasoning there?
Eiso Kant [39:14] Yeah, so in the very early days of the company, talking, we started April 25th officially. And so we were just a couple of people, and we went out and started recruiting, and we were primarily recruiting in the Bay Area in the first weeks of the company. And one of the things that we had done, and our very first hire, was a head of people, someone I've known for a very long time. She's an absolute machine at what she does.
Eiso Kant [39:41] And together with her and one of the early founding engineers, we built this very large list of talent we found relevant, right? This was just doing the work, like looking at GitHub, research papers, LinkedIn, all ourselves. And we went through about 3,200 people that we ended up putting on the list. And we had this location column, like, think genuinely Google spreadsheet, like, where's someone based? And at some point we looked at all these locations and were like, we should try to group this between Europe and North America.
Eiso Kant [40:10] And now let's filter it on all the people we considered good and all the people we considered incredible. That was kind of, it was like reject, good, or incredible was our criteria. And then we found a 50-50 split between Europe and the U.S., or Europe and North America. And that was surprising to us. And Jason at the time, I think, was the one who said, he has this phrase I really like. He always says, "We should zig while others are zagging."
Eiso Kant [40:37] If you can't—and not for the sake of it—but you can't compete with an OpenAI if you're doing exactly the same things. And so throughout these weeks in the Bay Area, we started also interviewing some people in Europe, and we kept hearing the same thing over and over again. The best people who were deciding to leave one of the major FAANGs or who were in a place that was less known, none of them wanted to join OpenAI or Anthropic. And we asked why.
Eiso Kant [40:57] They said, "Well, I don't really want to be in the secondary semi-commercial office in London, far away from the center of where things are happening. If I'm going to choose to leave my incredibly paid job at Google or go for something new or join a riskier early-stage company, I want to be at the heart of things, right? I want to have an impact." And that really for us was like a switch that was like, "Oh, wait a second, we have the numbers, we have the data, we're hearing qualitatively that these are people we're not competing with others in the space for."
Eiso Kant [41:30] And at the same time, we were in the Bay Area, realizing we were going to be competing, obviously, for every single person with OpenAI, with Anthropic, with some of the others out there. And so we, through a whole bunch of other series of events, ended up on a plane to Paris and had an incredibly warm welcome from other founders and people there. And it became very clear to us that there would be a huge advantage if we built from Europe.
Eiso Kant [41:57] We are at our heart a global company, right? I'm European, Jason's American. We have members of the team in the US today already. We don't see ourselves as a European player or US player. We see, hey, we want to build a global company. But it was clear to us that from an engineering and research perspective, there would be a huge initial advantage in building up that team in Europe. And that's proven true so far. Now we've built a team that's slightly shy of 25 people.
Eiso Kant [42:12] And we've interviewed, I think, about 700-plus people to build that team, not an exaggeration. And yeah, we've just been incredibly grateful and inspired by how much talent there is across Europe.
Matt Turck [42:21] Wonderful. All right. So to wrap this up, maybe walk us through what happens next in the next few months or maybe year at the company.
Eiso Kant [42:42] What should we expect? Yeah. So the most important thing that you need to expect from us and hold us to is a product, right? We want to get things in the hands of developers in your editor that are backed by our own model that is going to make you more productive, more essentially offload a lot of the work that you do in combination with an AI system. So that's what you can expect from us next year: a productization, a model coming out both behind an API and as a product experience.
Eiso Kant [43:17] You're going to see us continuously push the boundaries of the capabilities of these models. So we don't see ourselves as the one-in-a-six-month update. I think we're heavily inspired by Midjourney-style efforts of constantly improving and bringing out new models that back these application experiences. So that's what you can expect. At the same time, what you can expect from us is that we believe it's important to build a real business with real revenue. It's massively costly to be in the capabilities race because of the underlying compute.
Eiso Kant [43:48] And while it's incredible that there's an amazing funding environment that allows for extremely large amounts of funding upfront, similarly to how people back SpaceX or biotech companies, right? Very different than SaaS software. We also know that the best way to back ourselves is through business that makes revenue. And so you're going to see a strong emphasis from us on also bringing Poolside into enterprises, into complex environments where there might be 10,000 or 100,000 developers, where the model needs to continuously learn and be fine-tuned on their data, where it's not enough to just bring something that's more general-purpose software development in.
Eiso Kant [44:24] And so you'll see a strong emphasis from us on not just bringing things out in the regular world via our cloud product, but also very, very strong effort on doing things in the enterprise. A lot of stuff that we're already starting today. And that's how we think we will blaze a path to building a business that becomes valuable and worth its own valuation.
Matt Turck [44:27] Very exciting. Eiso, thank you so much.
Eiso Kant [44:49] Thank you to everyone who has joined us for season 1 of The MAD Podcast. We'll be taking a short break for the winter holidays, and we'll be back with an exciting new lineup of great speakers for season 2 on Wednesdays in January. If you like the show, you can find the video recording of this episode, along with many more, on the Data Driven NYC channel on YouTube. Important links are in the show notes.