Dataiku's Secret to Scaling AI in Global Enterprises | Florian Douetteau, CEO, Dataiku
The MAD Podcast with Matt Turck · with Florian Douetteau, Co-founder and CEO, Dataiku
Florian Douetteau is the Co-founder and CEO at Dataiku. We cover why enterprises need one platform for analytics, predictive models, and agents, why scaling hundreds of agents makes centralized guardrails and cost controls essential, and why generative AI will not replace predictive models built to optimize specific business decisions.
Chapters
- 2:08 — Florian's life before Dataiku
- 6:58 — Creation of Dataiku
- 12:08 — Secret behind the Dataiku's name
- 12:47 — How does Dataiku stay insightful about the future?
- 14:46 — Building a platform, not just a tool
- 17:26 — How to sell to the enterprise from the beginning
- 20:09 — Dataiku platform today
- 26:55 — Data is always the problem
- 28:50 — LLM Mesh
- 36:02 — Will Gen AI replace ML?
- 39:41 — Managing Gen AI and traditional AI on one platform
- 40:37 — Gen AI deployment in the enterprise
- 48:33 — Dataiku's roadmap
- 50:28 — What has changed with the company's growth?
Transcript
Florian's life before Dataiku
Matt Turck [1:23] Hey, Florian, welcome.
Florian Douetteau [1:24] Hi, Matt. How are you?
Matt Turck [1:34] I'm good. So we're recording this today in New York, but you're on the tail end of a long trip, right? You were in Asia, you were in SF. What was your trip?
Florian Douetteau [1:37] I did Tokyo, Sydney, Melbourne, SF, Vegas.
Matt Turck [1:41] And Dataiku has offices in where?
Florian Douetteau [1:46] We have customers and offices in most of those locations, but not in Vegas.
Matt Turck [1:47] Not in Vegas.
Florian Douetteau [1:48] No casinos as customers, I think.
Matt Turck [2:16] Yes. Vegas was for the Goldman Sachs private conference, right? I've had the good fortune of working with you for a number of years now, and we also had various conversations, typically in the format of Data Driven NYC, which is the sort of meetup I've been running for a while. But we never quite did this in long-form format, kind of podcast format. So I'm excited for it. We can get into some parts we never covered. We've never talked about the founding story of Dataiku or even your story.
Matt Turck [2:32] And I thought maybe that would be a great opportunity to go into that. So tell me the story, like, I don't know, you as a teenager and then school, and what was the journey into Dataiku?
Florian Douetteau [2:49] As a student, I was mostly into, well, a little bit of everything, like philosophy and reading and writing and math and physics and so forth. But ultimately also a bit of a geek, as in I got my first computer when I was 5.
Matt Turck [2:51] You were 5?
Florian Douetteau [3:02] Yeah, and it was '85. And back then, the best way to play with a computer was actually to do some programming for fun with it, or to try to hack some games in order to be able to copy floppy disks.
Matt Turck [3:03] What computer was it?
Florian Douetteau [3:28] Amstrad CPC 6128. 128 standing for the amount of random-access memory in it, 128K. Three-inch floppy disks, very specific to Amstrad. And most of the programming was BASIC. I'm not sure if there was Pascal there. It was a very first step to programming. And got my PC at 10, started to do more C and C++, which was great or awful. Then, as a kid, or as a teenager, got into functional programming because it was a thing in France.
Florian Douetteau [4:02] In France, we have this programming language called OCaml, a great functional programming language, which is French and taught in France instead of Pascal or Java, believe it or not. I got into being a functional programming geek when I was like 17, 18, 19. And when I was 20, I was at this school with lots of math geeks, as in, like, in France, where—
Matt Turck [4:28] Yeah, most of the podcast and YouTube listeners are American. So ENS is, my words, not necessarily yours, but like this super-elite, what's known as a grande école in France, where you have a tiny, tiny number of people that make it through. And that's known both for philosophy and very advanced math. Is that a fair description?
Florian Douetteau [4:30] Yeah. And physics too.
Matt Turck [4:31] And physics.
Florian Douetteau [4:56] Yeah. And very smart people. I think you're in a group of 50 where the top 10 people there are probably expecting to get a Fields Medal in their lifetime and have a decent chance of it. And so you're like, am I really good enough in math to research? Maybe not. And then at that time, 2000, I switched to startups very quickly. I was 20.
Matt Turck [5:01] But you considered at some point doing pure math research as a career?
Florian Douetteau [5:28] Yeah. And my focus was mostly in logic and programming languages and functional languages and type systems and compilers and compiler optimization and a lot of good stuff when I was 20. But then I switched to a startup in search engines. And so I switched to linguistics. Of course, I was interested, as lots of people in programming back then, in this mix of programming, logic, linguistics, and all of that good stuff. And so switched to more like linguistics and language models and applied to search and then indexing technologies and all of that good stuff.
Florian Douetteau [5:40] And I did that for about 10 years.
Matt Turck [5:42] And the startup in question was Exalead.
Florian Douetteau [5:43] Yeah.
Matt Turck [5:59] I guess most people don't know the company, including in France today, but Exalead was sort of the OG enterprise software success in France. Is that fair? I mean, I guess with Business Objects, there's a whole mafia that came out of Exalead.
Florian Douetteau [6:28] Yeah, 20, 30, 40 good startups actually came up out of it. And indeed, the company ran from like '99 to 2010, when it got acquired. But it was, yeah, those days, meaning those are the years where everything was changing. You had Google pushing new technologies. You had those days of MapReduce, those days of deep learning starting to work, the days of—of course, Moore's Law was still working back then. And so the amount of things you could do and the memory and the storage you had back in 2000 and in 2010 was very different.
Creation of Dataiku
Florian Douetteau [7:06] The starting point was like classic hard drives and solving for their failures and so forth. The end state was very fast and with very different types of setups. So yeah, hardware was changing quite a bit. Everything was changing around. So those ten years were actually fantastic for everything, which was like storage and compute, and also the raw state of what you could do in terms of language models and linguistics.
Matt Turck [7:10] That led to the creation of Dataiku in 2013.
Florian Douetteau [7:11] Yeah.
Matt Turck [7:17] So what was the journey? So you were at Exalead. That's where you met Clément, in particular, your co-founder. Is that right?
Florian Douetteau [7:18] Correct.
Matt Turck [7:36] Especially back then. I mean, now France has progressed leaps and bounds in terms of how vibrant the tech ecosystem is over there. But at that time, it was still somewhat of a strange idea to start a company. So you guys met, and then what led to the creation of the company?
Florian Douetteau [8:05] I took, like, two gap years between 2010 and 2012. Gap years in the sense that I was working in smaller companies or as a freelancer, working on data platforms, data strategy, data science setup for tech companies in France, working as a CTO of a social gaming platform on Facebook, which was a thing back then. So lots of actually very interesting things to understand in practice: what was the realm of the possible in terms of data? And I think that just hands-on experience is always great in order to switch your mindset and think about what are the real pains of any type of ecosystem.
Florian Douetteau [8:42] The starting point was that Dataiku—we were actually four co-founders—with, let's say, similar backgrounds in terms of our experience of the pains of working with data, scaling data, and so forth, with a mix of experience in business intelligence, in data science, at scale, in natural language processing, and all of those good things. Actually, back then, data science for everyone. You had a need for the enterprise to do more in terms of data science and ML, but a skill gap because there were not enough data scientists on planet Earth in order to fill that gap.
Florian Douetteau [9:15] And so the starting point was like, we need to democratize data science and, through software, find a way where this can be simple enough. And all of this was in this early ecosystem of data, where companies were setting up their first Hadoop clusters on top of Hortonworks and Cloudera and MapR back in those days.
Matt Turck [9:16] Back in the big data days.
Florian Douetteau [9:29] Yeah, the big data and the three Vs of big data. Yes. Yeah, I really sound like an OG now. That was really our starting point: data science for everyone. How can you make it work? And how do we solve for this problem of making this accessible?
Matt Turck [9:36] And that was very early to the trend, I guess. The concept of data scientist was just starting to be a thing.
Florian Douetteau [9:42] Yeah, 2011. I think it was coined in 2011 as a term, probably in either LinkedIn or Yahoo. I don't remember now.
Matt Turck [9:56] I think that was DJ Patil and a couple of others that came up with the term, but that was very nascent. And so the intuition was that there needed to be a platform for those people to work and collaborate.
Florian Douetteau [10:14] The main intuition was, in data science for everyone, that you needed actually to expand the boundaries so that people outside of data science could actually do data science. That when doing this type of modeling, you need to do the translation from a business problem to a data problem. A lot of the knowledge is within the business. And some of my starting points were just, like, meetings I saw in real-life companies where you had, on one side, people from the business; on the other side, people from tech, newly named data scientists.
Florian Douetteau [10:40] And there was just a gap in terms of communication. They were just not talking the same language, and they had very different expectations. Lots of passive-aggressive type questions, such as, "Do you really understand what kind of data we've got here?"
Matt Turck [10:41] Yes.
Florian Douetteau [11:07] Why aren't we doing like Facebook, what Facebook does, and, like, working with graphs? And this gap in terms of understanding the technology, what's the data, what can you expect? Our thinking was it should not be solved by slides or consulting or whatsoever. It should not be solved by hiring more data scientists, because anyway, there are not enough data scientists, and if they don't understand the business, there is no point. It should be solved by finding a way for people from the business that are very much into data and people in data science that need to understand more of the business to just work together.
Florian Douetteau [11:27] And work together is like, just use the same tool instead of working in silos and sending emails and meeting once a week to fight one another.
Matt Turck [11:50] So it's funny, this conversation has a lot of echoes to the conversation we had here with Olivier from Datadog. A big part of the story of Datadog was the developer people and the operations people sort of being good friends, but sort of hating each other in a work context. That led to that need for a collaboration platform precisely to make people work together.
Secret behind the Dataiku's name
Florian Douetteau [12:15] Yeah, it's very similar dynamics. Meaning people from DevOps trying to call themselves dev, trying to, from the perspective of devs, tell them, like, you're not really developers. And on the other end, people with actual experience of production being like, developers who actually have no clue of how to scale or what are the constraints. And this gap in terms of translation can be solved by having some way, a place to meet.
Matt Turck [12:18] What's the story behind the name of the company?
Florian Douetteau [12:18] Why?
Matt Turck [12:20] Or the story of the thinking?
Florian Douetteau [12:28] I'm a bit into poetry, and maybe was more before. At least after 10 years of startup, you read less poems.
Matt Turck [12:31] You're not really into anything.
How does Dataiku stay insightful about the future?
Florian Douetteau [12:54] You're into what you do. I found the combination of data, which was especially back then all about size, and haiku, which was—it's just like a haiku is a Japanese poem in three verses, 5-7-5 in terms of number of syllables, which means to say a lot by saying not too much. So, data haiku, Dataiku. Yeah, it worked for me.
Matt Turck [13:29] What's remarkable is what you described, which is this collaborative platform that enables technical people and business people to collaborate around data, data science, machine learning, AI problems. That was then, that was 2013, and that's still very much what the company does today. So you're actually one of the rare companies that had a very precise insight way ahead of a forming wave. And fast forward to the latest stage of the journey as a pre-IPO company, you still very much do that.
Florian Douetteau [13:46] We have to be cognizant of the fact that it's always easy to rationalize things after the fact. But in our case, we have this ambition in terms of what is the problem to be solved, and we are probably fairly open-minded in terms of the way to solve it. Instead of focusing on a very specific or narrow product concept, think about this as a realm of possible things to do, where you solve for the democratization of data science, of ML, of whatever you need in order to create value out of data, and see that as a big pain point to solve and then build software around it.
Florian Douetteau [14:35] Whenever you start a company, the question is, are you kind of sure that the problem you're solving will still be a problem whatever happens in two years, in five years, in 10 years? Because if you're back to the fundamentals, most companies need 10, 15, 20 years to be big. I think that we were lucky to have this ability to have a pretty wide ambition, but still be very pragmatic in terms of how to do it. And I think it was because my co-founders were very much grounded in the reality of the business, the reality of data science, the actual pains.
Building a platform, not just a tool
Matt Turck [15:12] It's interesting, the parallels with the Olivier discussion, because we had that discussion as well. You're supposed to be a tool before you're a platform, because you need a wedge into a market, and it's hard to sell a whole aircraft carrier to people. But you guys were a platform from the very beginning. Was there any thinking to this, or was that just a reality of, okay, well, there is a business problem, and this is a way of solving the business problem?
Florian Douetteau [15:28] This is maybe not always a relevant way to think about it, because it's mostly about what is the pain, and is there someone in the organization with the budget and with the ability to decide for which there is clarity about this as a pain? And so in our case, it was like, you need some layer, some orchestration layer on top of whatever you do with data, especially because data was moving from existing databases to, in our case back then, Hadoop, and starting to move to the cloud.
Florian Douetteau [16:12] So how do you actually build a layer on top of that? And indeed, early adopters of big data technologies were clear that they could not build everything on top of SQL queries in Hadoop, because for them it was clear that, indeed, you needed an abstraction layer between various teams and various business units and so forth in order to build things at scale. And so we were fitting that need. It was indeed controversial because explaining that to the market or to the early VC market required lots of nuance.
Florian Douetteau [16:48] The market back then had lots of tools that were focusing on narrow value propositions in terms of AutoML, or just AutoML, or even AutoML for specific models. Or MLOps—that's not really a thing yet, arguably, back then—but data prep was, and visual data prep was a thing. There were multiple startups just doing visual data prep, but without the idea of workflows or combining data prep with something else. The pushback against Dataiku from an investor perspective was indeed, if you're not the best at this or that, how can you win in the market?
Florian Douetteau [17:15] Because you wouldn't fill existing criteria. And our point was like, those markets are too tiny anyway. So why would we actually want to be the best at this or that if we believe the end market will be something else, kind of like a bigger category? I think that it required then some will or independence of mind, which was easier for us because I'm not sure if I had had the will or the mind power to keep this line of thinking if I had been in the Bay Area.
How to sell to the enterprise from the beginning
Florian Douetteau [17:33] Because around me, everyone would have had a different opinion.
Matt Turck [17:44] So the relative isolation of being in France helped do something that was non-traditional in terms of approach to building a company.
Florian Douetteau [17:45] Yeah, I think so.
Matt Turck [18:06] A common playbook for the last few years has been that, in addition to being a tool before your platform, you sell to other tech companies before you start selling to the enterprise. But you guys started selling to the enterprise pretty much right away. And by enterprise, I mean Global 2000, what's known in France as CAC 40.
Florian Douetteau [18:31] My thinking is, and I think it applied back then but still applies today, that to some extent, tech companies and non-tech companies have a different way to build and maintain over time things related to data or ML or AI. When you're a tech company, it kind of makes sense to think of your company as a big engineering group where, whatever happens, half of your company is a form of engineer, whether that be to build internal capabilities, to build a customer-facing product, and so forth.
Florian Douetteau [18:55] But yeah, you've got half of your company, at least, that is able to do some coding, ballpark. When you're a non-tech company, it's not the case. Maybe you've got 5%, 10% of your company who is able to code. And so it leads to very rational, different dynamics. Tech companies, I think, in the field of data, from our perspective back then, were not that eager to buy platforms because, from their perspective, it was pretty much about buying infrastructure and then building on top, relying on open source and packages, and leveraging open source as much as possible, which was making sense in the context where 50% of your company is engineering per se, or has engineering skills.
Florian Douetteau [19:47] But our perspective was, how do we solve for the problem of every other company, which will not follow the same route, which cannot follow the same route? And take the perspective of, like, we will never try to sell aggressively to tech companies per se. So how do we solve for everyone else? And the reality is that if you're not in the Bay Area anyway, 99% of the companies are not tech companies. So, not really a thing.
Matt Turck [19:58] And then you just had to be able to sustain the early, longer sales cycles that are necessarily involved when selling to those companies and just be patient.
Dataiku platform today
Florian Douetteau [20:12] Arguably, it was also because I had done 10 years of enterprise software and enterprise selling before. I was not afraid of building an enterprise sales team and an enterprise sales company. So I was like, yeah, let's do it.
Matt Turck [20:19] And then fast forward to today. How do you describe Dataiku?
Florian Douetteau [20:47] Yeah, the platform evolved quite a bit because we started as probably two or three key capabilities, AutoML, data prep, to more like 15, 20 today, all things being compared. So it's almost hard to describe the platform as a set of capabilities because it's a long list. Dataiku now is kind of like an orchestration layer enabling companies to build and maintain and govern any type of AI applications. It's operating with mostly a no-code type of paradigm, but actually extended no-code, meaning you can do no-code, low-code, full-code within the platform, which is important for us in terms of collaboration.
Florian Douetteau [21:18] And within this platform, you can do this mix of analytics, meaning building metrics. You can do predictive ML to actually build models, and you can do generative AI to build agents or any type of agentic or LLM-based workflows.
Matt Turck [21:47] That very much reflects what you were saying in terms of just generally extensibility and being open-minded to whatever new technology comes in. As the next thing comes, which I guess for the last couple of years has been generative AI, then that becomes something that you add to the platform. So you follow whatever the market is, because fundamentally the point of the platform is to enable, to orchestrate the whole thing, govern the whole thing, and enable people to collaborate around it.
Florian Douetteau [22:12] Rather than talking about Dataiku in particular, I think it's something about the ecosystem of AI that I find super interesting, which is that I've got the perspective, but maybe because I know less of other domains, that in the space of analytics and AI, you have a new technology popping in every year or so, like with something actually different, a new way to think, new types of models and so forth. It's not a stale type of ecosystem. I think it's less stale compared to others.
Florian Douetteau [22:41] It's moving quickly. So the underlying technology is moving. When you have the perspective of a builder, I think the people building the technology, you're very excited by it. You're like, oh, new technologies, we can build new stuff. Let's go, let's do stuff and so forth. Great, great, great. When you take the perspective of the enterprise, sometimes the reception of those new technologies is more nuanced. It's more like, oh my God, yet another technology. Oh my God, yet another thing.
Florian Douetteau [22:58] Do we do it? It's like decision-making. Do we need to integrate it or not? Is it a thing? Are we laggards if we don't do it? Do we have some risk if we do? Does that mean that we need to replace something else? Yeah, it's a question. It's a stress factor, actually. Innovation is a stress factor, but the positive side of innovation, as builders view it, is okay. But the side of, there is new technology popping in, is actually kind of like a risk.
Florian Douetteau [23:33] Our perspective was like, because we are building this platform, which is this layer in between the business focus of enterprise and any type of data technologies, like the orchestration layer where you can connect things together, we are almost acting as a dampener. Like, there is a new technology popping in, open source or whatever, new type of database, new type of models, new type of way to think about evaluation of models or what else. They should come up in Dataiku after a few months.
Florian Douetteau [23:59] The things pop up in Hacker News or what else, and it's becoming a thing. It's a new open-source cool thing. You should be able to do it in Dataiku. Maybe not one day after the news on Hacker News, maybe like a few months later, but in a way which is more enterprise-ready, with an actual UI on top of it, so that people in the enterprise can start using it and incorporate it into whatever they do. And so this constant integration of technologies has been part of the DNA of Dataiku.
Florian Douetteau [24:26] And it's uncomfortable when you do enterprise software because it means that every quarter or so, twice a year at least, you've got something else going on and you're not controlling those three- or five-year roadmaps. Because the way we think about it, in three years or five years anyway, there will be new technologies popping up. The fact of being able to integrate them is what is driving us.
Matt Turck [24:37] And part of the orchestration layer is that you follow the entire lifecycle of data. So you basically do all the things that one needs to do to build, maintain, and deploy machine learning.
Florian Douetteau [25:04] Yeah, including the ops. And that's also something that was a bit controversial back then. But the core thinking here is twofold. When I started Dataiku, I was mesmerized by a couple of ideas. One was self-service, which is pretty much about democratization. How can you enable people in the business? Essentially, the idea of assets and Lego. When you're building in data science, people tend to redo from scratch, or even analytics and so forth. How can you build something where people can build on top, like Lego?
Florian Douetteau [25:26] Probably because I love Legos. Let's make it composable and so forth from a real perspective, not just on a slide. How can you actually stitch things together, put them in a box, and then you can reuse the box, and have this Lego type of mentality, which is very fun. The notion of end-to-end assets. If we all think that AI is going to change the world and change the enterprise, the actual things being built in terms of data and ML are actual assets, meaning that they will derive lots of long-term value for the enterprise.
Florian Douetteau [26:03] If an enterprise is becoming data-driven or AI-driven, it means that every piece of analytics is actually bearing lots of value long term because you're rebuilding lots of intelligence in your company based on this. Even if it's a simple metric that you're building, building it in the right way is very important because it will be used and reused for maybe 10 years in multiple locations in your company. So this building-assets-end-to-end type of mentality, which leads to the idea of, like, in Dataiku, you should be able to do everything end-to-end.
Florian Douetteau [26:39] Like manage those assets as real assets. A metric is an asset, a model is an asset, an agent is an asset. So you should be able to manage the full lifecycle from ideation even, and design and testing it and pushing it in production and monitoring of it in a consistent manner, and manage this full lifecycle instead of having multiple tools. And it's back to the question of dev versus DevOps that we talked about when mentioning observability and Datadog. You can have, in some organizations, just change the perspective that there are some people building and some others operating, hence they should have different tools and platforms.
Data is always the problem
Florian Douetteau [27:03] But here we were like, no, if they do that, they'll just lose lots of efficiency by trying to redo and translate, and it would be very hard to actually understand overall. So can we actually do things the other way around?
Matt Turck [27:30] A part of this at the beginning of the lifecycle of data in the enterprise is data prep, and you have very strong capabilities. I mean, the joke is that whether that's data science back then, machine learning in 2016, '17, or generative AI today, ultimately, a core part of the whole thing for the enterprise is data and having it ready to go.
Florian Douetteau [27:50] If you've got a great AutoML tool or even a great RAG-type tool and so forth, ultimately, lots of those technologies can become, in isolation, pretty much commoditized. Meaning, it's not hard to build RAG or an agent builder. Meaning, what's hard is actually to have an LLM-first part. But the RAG builder or the UI is not hard. What is then hard for the enterprise is that you need to continuously have the right data in it and test it and manage all of the edge cases, all of the boring stuff.
Florian Douetteau [28:30] What's hard in many data science projects is not the model tuning and so forth. In most instances, in fact, it's getting the customer data right and managing all of the edge cases or situations where you've got empty columns or not, and whether they matter in terms of measuring the performance of your model and whatever else. So lots of initial tools in those domains get stuck because, when deployed in real-life organizations, they don't fix the pipelining before and after. And we were like, we need to build great data prep capabilities and all of the things associated with them in order to solve for that.
LLM Mesh
Florian Douetteau [28:57] Because if we want our software to be used by real people and not send consultants with them so that they fill the gap and prepare data for them so that they can just click on the button and do the AutoML, if we really want to pursue the self-service approach, we need those capabilities.
Matt Turck [29:10] Let's talk about the LLM Mesh, which is what you all have called your specifically generative AI sort of effort over the last couple of years. So what is it? What does that do?
Florian Douetteau [29:37] You provide for the enterprise the best type of underlying environment in order to build agents that matter, and what are the actual equivalents of the pipelining and data prep problems that you have for ML related to LLMs themselves. So when you want to build agents, you've got some form of data problems, like you need to get the right data in and so forth. Great. But you also have the overall management of the models and so forth. What we built is a way for the enterprise to get, first, easy connectivity and abstraction over all of the models out there and their versions and all of that good stuff.
Florian Douetteau [30:16] This first part is, I think, very important because what matters, I think, for the enterprise is not just to have the variety of models you can have within one provider. As in, within AWS Bedrock, you can have lots of different models, and same on the Azure AI Studio thing. But what you want in many instances is to understand that you can switch from OpenAI and Anthropic and Llama and this or that provider, switch from one cloud to the other, discover if you could be more efficient by switching to new cloud providers with cheaper GPUs, and all of that good stuff.
Florian Douetteau [30:47] If you just build one application, you can just maintain and fix the code of your application and so forth. But our vision is that enterprise at large will not have one agent, but more like 500 or 1,000 of them. And so you just need a central place to understand which agents are talking to which LLM, where you can actually manage this fleet at large. And then comes the interesting part. Once you start to understand and abstract the relationship between an agent and an LLM, you can add capabilities that are required for real enterprise readiness, which are all of the guardrail-type services you want on top of the agents themselves.
Florian Douetteau [31:27] You want to control and understand their cost profile. You want to understand and manage the content they're producing, and you want to make sure that they don't get hacked. And so, how can you build a security layer around LLMs? How can you actually check for the content and leaks of confidential information or personal data, managed properly? And how can you actually manage and control the cost? One of the core ideas we have there is that, when the ecosystem is building and companies are starting to build agents because it's an early technology, as point solutions.
Florian Douetteau [32:05] And then every business unit will go and build their own agents, either by themselves or with a consulting firm popping in to build the agent and so forth. So you've got all of this happening. Some things are working, most are not, but it's an early point. But as the ecosystem will evolve and companies will start to have lots of agents in production, the actual problem from an IT perspective will be how to manage all of this complexity. Which is when you will actually need a platform, because you don't want to have as many implementations of guardrails as you've got applications in your company.
Florian Douetteau [32:40] You want to be able to update them centrally. You want to be able to have a central vision of the costs of all of them and how you actually spend money on GPUs and models at large. You want to be able to have a single point of security where you understand all of the profile of all types of injection attacks on your LLMs and all of that good stuff. All of those capabilities could be startups by themselves. But the same way we thought about the platform required for the enterprise back in the day, 10 years ago, we think that for the world of LLMs, all of those capabilities will just become features of platforms like ourselves.
Matt Turck [33:07] Because the fundamental problem, regardless of the technology evolution, just to play it back, is also humans and process and governance and control and visibility and transparency.
Florian Douetteau [33:31] Whenever you build software, it's kind of interesting to think of what will be the actual problem two years or five years from now. And I, for instance, don't think that the problem two years or five years from now of agents will be relevance per se, as in the quality of the models. The problem will be very boring stuff, such as: Is it the right data, or have we actually updated the data for the RAG every hour or so? Or is it stale data that is like one month old, hence it provides bad results?
Florian Douetteau [33:51] Do we have any type of security problem or security leakage because we haven't identified what is given to whom in terms of answers? Do we actually understand how we are spending the money? And do we have some users that generate 10% of the users generating 90% of the cost? I think these types of issues that are the very boring issues you get when scaling things in the enterprise will be happening in agents, and that, I think, will be the remaining of the problems.
Florian Douetteau [34:13] Because the quality of the models themselves will just keep improving enough so that the question of whether it's the right summary or the right answer will not be the key one.
Matt Turck [34:31] Your thought is that right now, in the period where people are experimenting on the one hand, and then you have a bunch of model builders that create increasingly impressive generative AI, but then the next wave is like, okay, what do we do with it and how do we deploy it?
Florian Douetteau [34:58] The next phase would be, what do we do with it? And the next phase will be, okay, how do we actually deploy and maintain it and keep building stuff without spending 90% of our time maintaining or fixing the existing stuff? And in terms of predictive type of AI and traditional, meaning machine learning models, we have customers having a few hundred of them in production. And for some of them, saying 10 years ago that they would have 500 models in production was crazy.
Florian Douetteau [35:06] Data science?
Matt Turck [35:06] What?
Florian Douetteau [35:29] No. And so I think that the same kind of thing, with which timeline, in fact, is it like two years, five years, 10 years, will happen in the enterprise too. You will have lots of processes involving generative AI and agents that will be in production here or there. And then the actual practical question would be the level of control in the world of generative AI. And it's maybe a bit of a poetic way to think about it.
Florian Douetteau [35:56] It's no longer the creative act itself that is the limiting factor, because you don't have to code things, you write prompts, and it's kind of easier in the grand scheme of things, or it will be easier, at least for sure. And so the limiting factor is more like all of the control you need to put around the generation itself, around the prompt or the playbook or explaining to the agent what it should be doing. But how do you make sure that it's the right data in, the right data out, the testing, the framework, the connectivity, and so forth?
Will Gen AI replace ML?
Florian Douetteau [36:09] So the control of AI will be the limiting factor for creating more AI.
Matt Turck [36:42] And you mentioned companies, your customers, having gone on that journey from data science not really being a thing to now having hundreds of machine learning models in production. In your view, in the future, is the end result a combination of those machine learning models and generative AI sort of all working together? There is a little bit of that narrative that you hear from time to time that generative AI is completely replacing all of traditional machine learning. What's your experience, and what do you see on the ground working with customers?
Florian Douetteau [37:13] I think that there is indeed some misconception on the fact that generative AI, as it is built today, is not meant to provide any consistency in terms of how statistically you reach a business goal through the output or through your predictions. I mean, when you build a model, a predictive model, what you're optimizing for is making a decision with the best trade-off so that you optimize a business goal. Like, I don't know, you're doing anti-money laundering, you flag transactions, but you don't flag them all because you want to optimize for the cost of the business.
Florian Douetteau [37:52] So you make a trade-off associated to probabilities and so forth in order to flag the one with the risk that you need to manage. Same for any type of business goals or pricing. You're providing discounts to your customers when you're building models to optimize your sales or revenue. So all of those models take into account not just one decision, but thousands of decisions, and are built on optimizing for the probability of best outcome for the business. And that's what data science is about, which is why you need the history and the data and the backtesting and so forth.
Florian Douetteau [38:02] Those concepts do not exist in generative AI.
Matt Turck [38:31] So there's data science and machine learning. I guess we're all struggling in today's industry to put labels and terms on those different waves. But that's the world of, let's call it, machine learning, which is a terrible use of the term from a scientific standpoint, but just maybe for clarity. So that first wave of machine learning models, that's what they do. They do optimization, typically on structured data. Not always, but typically. Is that fair?
Florian Douetteau [38:53] Yeah, typically on structured data or a mix of structured data and unstructured data. But ultimately, you optimize for a given business goal very narrowly, with understanding the risk and the risk profile you're building your model against. You will still always need, as an organization, to be able to build those models because there are actual business decisions in them, like actual risk profiles with conscious decisions in them.
Matt Turck [38:58] And just to double-click on it, generative AI will not do that.
Florian Douetteau [39:25] Will not do that. But what we will build are agents that are combining all those models together in order to do more autonomous tasks end to end. Again, it's back to the topic of building blocks. It's like, how can you build building blocks where you're building new metrics in order to have the right data? You're building models that can take some form of long-term decisions for your business in a repetitive way, in a good way, by predicting something. And how can you build agents using large language models in order to actually build more and solve for more autonomous tasks?
Managing Gen AI and traditional AI on one platform
Florian Douetteau [39:48] But you need to stitch those things together so that the agent can actually have the right data, like the one that was validated and tested, and have the right models that have the good long-term characteristics in order to help make or automate the right decisions.
Matt Turck [40:15] And then through the Dataiku platform, people are able to seamlessly manage both types of AI, whether that's, again, struggling to figure out the right term, predictive AI or traditional AI or machine learning on the one hand, and then generative AI on the other hand. There's enough consistency in terms of how those need to be managed that you can manage both of those types of AI in one platform?
Gen AI deployment in the enterprise
Florian Douetteau [40:42] There is a need, actually, to have platforms where, in fact, you need the three types of things. Really, it is about having analytics so that you've got a clear, detailed view about the past. It's about having predictive capabilities so that you can build forward-looking models. And it's about having generative-type capabilities so that you can actually build agents, automate workflows, and actually integrate and interact with the rest of the world.
Matt Turck [41:01] Metrics, models, and agents. Talk a little bit about what you see on the ground in terms of the reality of generative AI deployments within Dataiku customers. And again, you tell me, I'm never quite sure which customers are public or not, but we're talking about some of the largest companies in the world.
Florian Douetteau [41:33] We indeed have the luck to work with north of 700 enterprise customers, among which Morgan Stanley, Michelin, Novartis, Perdue Farms. So, across all sectors and industries, really, those customers are all going through the journey of modernizing their data, meaning being better at doing analytics and so forth, building more and more models, and getting into usage of generative AI. Those use cases are all very exciting, meaning there's lots of excitement about, oh, you can build more with AI. There is also a fear or question of understanding the risks.
Florian Douetteau [42:08] So, on one side, the nature of the market is that companies want to understand how they will build and scale generative AI-type projects, how they're going to scale agents, understand the development model for it, how you have the right guardrails, the right lifecycle, and so forth, which is what we help them with. And on the other end, there is lots of excitement about key projects and things where they were able to automate repetitive tasks or very difficult tasks using generative AI.
Matt Turck [42:29] And without naming names, can you talk about any kind of examples of what people have actually done? It seems that so far we are in the kind of low-hanging-fruit, easy-win part of the market cycle, where people do things like search and chatbots and that kind of stuff.
Florian Douetteau [42:43] Yeah. And I think that's not really where the excitement is coming from. There are always three buckets in this type of innovation. You've got the false low-hanging fruits, you've got the very hard stuff, and you've got things in the middle. And what's interesting to me are the things in the middle. The low-hanging fruits, like chatbots, have the issue that even if you deploy to the whole enterprise a way to be better at summarizing emails or doing some task and you get some level of productivity out of it, it's very hard to keep long-term excitement.
Florian Douetteau [43:07] And I think it's not necessarily the role of the data teams or the AI teams to make that happen.
Matt Turck [43:10] And why is that? Why is it hard to keep long-term excitement about it?
Florian Douetteau [43:43] Because even if you've got 5% or 10% productivity for everyone because you were better at delivering productivity or searching for how to submit expenses, meaning this kind of use case, ultimately those capabilities will become built-in capabilities of our day-to-day productivity tools. And so I don't think it's long-term excitement. That's not how you build or transform your company. And it's also not necessarily how you think about transforming your organization, changing the P&L of your organization or your business unit or whatsoever. What we've seen, on the other end, are situations where people are focusing on something in the middle, as in, like, not for the full company, not completely changing a given job, but focusing on repetitive hard tasks that can be automated.
Florian Douetteau [44:27] And I think that those types of projects need to be pursued by the business, as in people having a good understanding of what is the end-to-end process and helping with that. We've seen, for instance, some legal teams using our platform in life sciences in order to better understand the patentability landscape, or to make some arguments, or to analyze the overall state of the art in order to generate all of the documents they needed to generate. And suddenly, you can have the work of 20 days of someone that can be very much facilitated because the machine can actually do all of the information gathering, the summarization, if you define the proper business rules and understand where to look for.
Florian Douetteau [45:08] Those topics are not easy because all of those domains were things that are being done by very talented professionals that needed to look into some structured database of molecules or demographic data, needed to look into prior art, internal databases, and so forth, needed to challenge some of the output and so forth. And all of this can be steps into an agent or into a workflow that can be built within our platform. But you need this translation from the process and then what the person was doing in those 20 days into the system, as in the various steps that you implement.
Florian Douetteau [45:46] We've seen scenarios of operators of big plants that had to do every day one hour or so of summarizing the operation of the day and build a daily report of the day, where they were able, using our platform, to automate the generation of this report and turn that one hour into 10 minutes of checking the data. And again, it's harder than summarizing your email because when you build those reports, those reports would be used for safety purposes down the line. As in, when there is a safety issue, someone will look at past reports and past daily reports to look for similar cases.
Florian Douetteau [46:21] And so you need to make sure that you've got the right data, meaning it's the right metric, the right pressure. You don't want hallucination there. So you need to build it in a specific manner and with specific tests to guarantee that. Because it's an actual enterprise business process. It's not like I'm generating an image to post on Twitter. There are more strings attached. You can get tons of productivity. And I think that enterprises will transform themselves by multiplying all of those agents that will, step by step, automate the most repetitive tasks requiring gathering information and fact-checking and so forth in the loop.
Florian Douetteau [46:42] So that jobs can be augmented and humans can focus on the most creative or important or impactful part of their job.
Matt Turck [47:13] So what does it look like? Sort of fast-forward here to wrap it all up. So this combination of different types of models, like lots of generative AI, lots of agents that operate on both predictive machine learning kind of models and generative AI models. And then there'll be some vendors that will provide AI in the box, sort of like the way SaaS companies have been doing it, and then a lot of homegrown custom automation. Is that how you think about it?
Florian Douetteau [47:36] I think that first you will have lots and lots of capabilities coming from existing vendors and platforms, meaning your CRM will have agents in it, your HR system will have agents in it, your ERP will have agents in it, and your support platform will have agents in it. Yeah, great. And this will happen and will provide productivity out of the box. Then enterprises, in order to differentiate and transform themselves, will also need to build their own agents on everything, which is about stitching together the data from one vendor to the other, or the overall process, or the thing they actually built, meaning what makes them different.
Florian Douetteau [48:20] And especially you have all of those agents that are about helping the enterprise and helping people make the right decisions or come up with the right information to support the decision. All of the agents related to, essentially, decision-making. Today, lots of those tasks are very cumbersome and require lots of thinking. You've got lots of people helping with that. And the agents that could help an enterprise to be smarter, I think, are the ones that many enterprises will want to actually keep building themselves.
Dataiku's roadmap
Florian Douetteau [48:41] Because if you delegate your own intelligence to someone else, what do you do? And indeed, we think that as Dataiku, we can help all of our customers to actually get there.
Matt Turck [48:50] So what's next? Maybe to wrap things up, what does the next year at Dataiku look like in terms of whatever you can talk about? What are you launching? What are you thinking about?
Florian Douetteau [49:17] Lots of interesting developments where we'll keep integrating with a fast-moving agent ecosystem and new models and new model paradigms. You have a new technology popping up, meaning every big conference, whether that be the one from Microsoft or Amazon or Databricks or Snowflake, provides new technologies that we can integrate in order to further push the bar of what's simple to do in the enterprise. And that's our focus. A few hundred enterprise customers were focusing last year quite a bit on building their first use case or success with generative AI.
Florian Douetteau [49:53] And I think this year many of them were like, how do you replicate? How do you scale? How do you maintain? And I think we have the opportunity to help them with that. And that's very exciting because it's not necessarily about building an LLM itself and fighting for benchmarks, which lots of companies do these days. But it's more about, like, can you actually make it real for many people for whom it is changing their day-to-day life?
Matt Turck [50:09] What a journey, right? As it turns out, it takes 10, 12, maybe 15 years to build a great company. And you've been at it for a number of years now, but in many ways it sort of feels like you're just getting started, right?
What has changed with the company's growth?
Florian Douetteau [50:35] I think it's at least 20. I think it's almost probably very comparable to the number of years in order to get a kid in or even out of college, in and out of college, and maybe being financially independent. That's kind of like the same time. It's not 10 to just get a phone and become teenagers. You need probably 20 so that the company is really an adult.
Matt Turck [51:07] Yeah. How did you navigate that whole thing? So, going back to the beginning of the conversation, it's one thing to start a company on a vision and all the things we talked about 10 years ago. Fast-forward to today, Dataiku is a thriving company. You've got, what, numbers you've most recently disclosed, but you're in the hundreds of millions in ARR. You've got, as you said, 700 enterprise customers. It's a very different job for you today than it was then, presumably.
Florian Douetteau [51:27] At the end of the day, it's about having great people around you. And so indeed, initially, when you start a company, the great people are the one, two, three other people with whom you start the company, which is great. But then, at the end, I think what I find exciting is to know that we can keep hiring and getting new people in, meaning people who will keep changing the game, keep helping us to change our vision of the market, the way we operate, and so forth.
Florian Douetteau [52:04] So it's all about adding, meaning keep getting new people in. The thing is, when you grow, the great thing is that you can indeed attract more and more great talent. Meaning it would have been impossible for me, when we were four of us, to just hire the type of people I've got in today. That's for real.
Matt Turck [52:04] Yeah.
Florian Douetteau [52:32] They would be like, what? They wouldn't even have answered the call and even opened the email, or even worse than that, I don't know. And I think that's actually the good part of it. When you scale, you can actually have great people around you. And that's, I think, what makes the job interesting, because on top of working in technology or thinking in abstracto about the market, you can actually have the engineering talent to make it real. And you can have the people around to actually scale the business.
Matt Turck [52:45] That feels like a wonderful place to leave it. Congrats on an amazing journey so far. Excited for what's next, and thanks for doing this.
Florian Douetteau [52:45] Thank you.
Matt Turck [53:05] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.