2024 will be the year of Generative AI in production in the enterprise | Florian Douetteau from Dataiku

The MAD Podcast with Matt Turck · with Florian Douetteau, Co-founder and CEO, Dataiku

Florian Douetteau is the Co-founder and CEO at Dataiku. We cover why enterprises need to plan to switch LLM providers as applications mature, how generative AI supplies the last mile of personalization on top of traditional recommendation systems, and why cost monitoring can prevent an internal application from reaching a $100,000 run rate.

Watch on YouTube

Chapters

  1. 1:09 — What is Dataiku?
  2. 2:03 — Is the market ready for AI?
  3. 4:33 — Traditional AI vs Generative AI
  4. 8:33 — What a company should know before diving into Generative AI?
  5. 10:18 — Cost of Generative AI adoption
  6. 12:10 — What blocks the AI adoption?
  7. 14:31 — Dataiku product tour
  8. 16:34 — How to build one product for different audiences
  9. 17:45 — LLM Mesh: what is it?
  10. 21:10 — Evolution of platform building with Gen AI
  11. 22:17 — Enterprise AI motion in 2024
  12. 23:28 — Dataiku's partnerships
  13. 24:24 — Being platform-first as a startup

Transcript

What is Dataiku?

Florian Douetteau [1:42] So our company is trying to democratize AI for the enterprise, building a platform to make it simple to prep your data, to build your models, and then move them into production. So it's a platform that we provide, focusing quite a bit on the large enterprise. To share with you some numbers, we are still private, so we share numbers only when we need or want to. And last time we shared numbers was last September, and back then we crossed 600 customers. We were at 230 million of ARR back then, and we have more than 1,000 employees.

Is the market ready for AI?

Florian Douetteau [2:03] So we focus quite a bit on medium to large enterprises, helping them essentially make AI models in production, prep the data, like getting all of these boring things required in order to get value out of data.

Matt Turck [2:23] A lot of people seem to think that 2024 is going to be the year when AI becomes real in the enterprise. From your perspective, from the ground, what is the state of readiness of, let's call it, Global 2000 companies for AI in general and generative AI in particular?

Florian Douetteau [2:43] I would say that it probably needs to be a differentiated answer compared to AI in general and generative AI and non-generative AI. So first, I think we should all have a common name for non-generative AI. Call it conventional AI, non-AI, traditional AI. I don't know.

Matt Turck [2:45] Tabular.

Florian Douetteau [3:05] Tabular AI. But it's not only tabular. If you get to a specific use case, like getting the images of wafers for a microchip company and predicting quality control, it is traditional AI. It's non-generative AI, but it's deep learning. So is it generative AI or not? We don't know. Well, that's the issue. But I would say that from my perspective, at least, the maturity of AI, as in having predictive models in production in the enterprise, has grown significantly in the last 10 years.

Florian Douetteau [3:44] And I see multiple customers having hundreds of models in production and being happy with it. I would say it was not the case 10 years ago, and I think there was some maturity in terms of moving analytics to the cloud, in terms of understanding how to manage the cost of machine learning, or to distribute and give access to the data more broadly. And so not every enterprise is there, but there has been some progress. It's a continuum in terms of progress that I've seen in the last 10 years.

Florian Douetteau [4:17] On generative AI, I would say that for the enterprise, and here I'm talking very specifically of enterprises leveraging generative AI for their own internal kind of use cases, building their own generative AI solutions, I would say that it's a mixed situation where most enterprises are testing stuff, are doing some proof of concepts and so forth. But we've got about 100 customers that use our platform for some LLM-associated use cases, and we've seen some use cases in production, but it's not the majority yet.

Traditional AI vs Generative AI

Florian Douetteau [4:33] And I suspect that indeed 2024 will be the year where we see way more use cases in production and see some durable, real use cases associated with generative AI for enterprise use cases.

Matt Turck [5:02] I think it's really interesting, and it's sort of poorly understood, which is those two different families of AI. So traditional tabular, columnar, structured-data AI versus generative AI. So what is that first category? What does it do? How is it different? What kind of data does it operate on? And most importantly, what are the specific use cases? Versus generative AI: what does generative AI do, and what are its use cases?

Florian Douetteau [5:38] It's possibly obvious or not to the audience here, but indeed today in the enterprise, most large enterprises would have quite a bit of machine learning in production, leveraging traditional AI in order to manage use cases ranging from predicting the churn of your customers, segmenting them, targeting them, understanding affinity of your product with customers to recommend some products or do dynamic pricing, optimize the quality control of your assembly lines, manage supply and demand, and all of those use cases. And it's becoming fairly common. Why is it fairly common?

Florian Douetteau [6:02] Because it's doable now. And that's the best way to actually optimize your business, especially when you reach a certain size as an organization. So this is mature in the sense that, in each industry, you would have the major player of such industry having done and implemented traditional AI in production. And that's how the economy is working these days. So AI is already making this work, meaning probably the electricity in this building is provided with—not AI per se—but there are indeed some AI models, machine learning models, involved in predicting stuff associated with the production itself.

Florian Douetteau [6:39] Almost sure that's how it works. And on the other side of the spectrum, generative AI, which is how AI gets to be known by the general public, which is fun for us if you're a practitioner. Like, you've got AI as you know it, and now someone else is coming and giving another definition of AI, and your grandmother is talking to you about AI. I think for us that generative AI might be this additional layer on top of traditional AI, reaping and adding more benefit to it, more automation, and lots of promises there.

Florian Douetteau [6:53] But yeah, it's the early innings of it.

Matt Turck [6:58] And different use cases, or even the use cases are not completely clear in the enterprise.

Florian Douetteau [7:19] Most important use cases are not fully scoped yet, but there is a mix of use cases that are brand new, use cases that are not associated at all with tabular data, which are about finding information, very much about documents and so forth. You've got also many use cases that we see that are just, let's say, natural extensions to existing traditional use cases. As in, if you already have in production some product recommender system, one issue with the last mile is that if you start having some sophistication there with multiple products to push to someone with limited channel, maybe your last mile of personalization is not so obvious.

Florian Douetteau [7:59] If it's advertising, it's like, how do you craft the content as a display? If it's email, what is the content? If it's text, how do you actually craft a text message pushing for two products that would be personalized enough? And so this kind of last mile can be done now with generative AI. And we could not do that before because human time to actually personalize the content was not realistic. Meaning, if you have two unrelated products and you want to craft a very small text message which is very personalized per the profile and demographic of the individual, like, you don't have time to do all of those combinations and make them very personalized.

What a company should know before diving into Generative AI?

Florian Douetteau [8:40] And now you actually can. But in order to do so, you still need traditional AI because if you don't have a product recommendation system, it wouldn't work. So in that case, generative AI is providing this last mile, this additional benefit in terms of traditional AI: the last mile of, let's say, personalization, for instance; the last mile of accessibility in other contexts and in other situations. It's like brand-new use cases, indeed.

Matt Turck [8:50] What do you tell Global 2000 large enterprises when they come to you and ask, "Generative AI sounds amazing. How do I get started?"

Florian Douetteau [9:14] I think that, if I'm intellectually honest, they have so many people telling them so much information that I'm trying to focus on one aspect that I truly believe in. Because scoping use cases is the job of already too many people on the planet, from my perspective. And so I'm focusing on one thing, which is back to basics. Well, you must be aware that you need some ROI to your use cases. Good. You're not stupid, so I'm not telling you that.

Florian Douetteau [9:36] But you must be aware that, probably because of the evolution of technologies and LLMs out there, it's very likely that you will have to switch from one provider to the other over the course of your application, if your application is successful. Well, if your application is not successful, we don't care about it, of course. But if you plan to run it for one or two years, it's very probable that you'll switch from OpenAI today to Anthropic tomorrow and maybe Mistral—they're French—on Friday.

Florian Douetteau [10:08] And maybe you will have to, for this or that country, if you really deploy it globally, you might have to self-host an open-source model because of whatever regulation. What do you think about this complexity moving forward? And what do you think about the overall state of your company when you don't have one application in production, but like five, 10, 20, 50, 100? Because it's actually fairly realistic for an enterprise to have hundreds of generative AI applications in production at some point.

Cost of Generative AI adoption

Florian Douetteau [10:18] More talking about this aspect of things.

Matt Turck [10:32] One of the moving pieces in the generative AI landscape is a question of cost. What do you tell companies, especially when they tell you that they're concerned about wasting money on generative AI investments?

Florian Douetteau [10:57] Well, first I'll tell them that it's true: they might waste money. That's just a fact. And so, from my perspective, the way to approach it is pretty basic. Today, you might have multiple providers, multiple applications. Some of them are in dev, in prod. And the issue with generative AI is also that you may have unforeseen costs when moving to production because you don't anticipate the size of the context you're pushing at. You don't anticipate some user behaviors and so forth.

Florian Douetteau [11:27] So you need some actual monitoring of cost and attribution towards users and apps. So my first step is telling them you actually need to have a gateway between your application and your LLMs in order to track all of the costs and understand, before moving things to production, how much you can anticipate in terms of cost. A real-life scenario is that even us internally, at some point, we built some internal use cases associated to sales calls or what else. And we realized that the running cost would be like $100K.

Florian Douetteau [11:49] I was like, "$100K? No way. No way we are doing that. We have to optimize this and that." But indeed, without having some control, we could have clicked on a button and spent maybe not $100K, but like $20K or $30K before realizing it was completely stupid. And I think that it does happen quite a bit on the market right now. If you fast-forward to a world where you move those LLM applications into production, managing the cost is indeed a bit boring, but it's just one requirement, especially when we move from a state where we focus on the easy, obvious applications that are reaping lots of benefits to multiple applications for which the ROI might be more uncertain day one.

What blocks the AI adoption?

Matt Turck [12:26] What are the other blockers you see to enterprise adoption? The cost could be one. Defining the use case could be one. Are you seeing things around skills? We don't have the people internally, or governance, or data transparency.

Florian Douetteau [12:56] First layer of concern associated to security, to putting guardrails to your application, understanding what are the risks associated to that leakage of information. A lot of those things. I think there are solutions to that, but it requires some discipline. Potentially one of the roadblocks is indeed around skills, ease of use, and ultimately making things work. Because if we're again a bit honest, it's not that easy to make generative AI applications work in practice. We are playing a little bit dark science there where you get your LLM, you fine-tune, you RAG your stuff around, and you decide, like, okay, good, this size of context, chunking yes or no, this way or not.

Florian Douetteau [13:24] And it works, and you're like, oh yeah, it's like cooking. It ended up working, and you stop because once it's good, it's like a sauce. You stop cooking. You're just like, yeah, I'm good. So it looks like that really in real life these days. So from my experience of this type of field, it's very realistic that six months or one year forward, we will look at us now and we'll be like, okay, they were playing a bit with new technologies, and maybe we will have more sophisticated ways to do things that will work better.

Florian Douetteau [14:06] But today, there is a part of experimentation there. It's part of the roadblock. I think it's important for many enterprises to do so because that's how you innovate. That's how you will learn faster than others. That's how you will actually maybe get one year or two years ahead in terms of adoption of AI. But indeed, it's a bit of a risky business because not everything is working. And so us at Dataiku, we are working on making things simpler for the enterprise, but we just have to be realistic on the fact that not everything is working today in terms of making LLMs the next generation of things doing everything in an enterprise.

Florian Douetteau [14:24] It's not magic yet.

Matt Turck [14:30] So my main takeaway is that building generative AI is a little bit like the movie Ratatouille. That's the general idea.

Dataiku product tour

Florian Douetteau [14:31] Yeah, Ratatouille.

Matt Turck [14:44] So let's talk about Dataiku and the platform itself. The platform is very broad, does lots of different things. Maybe give us a product tour of what it does, from data prep to AI.

Florian Douetteau [15:02] We've added LLM and GenAI capabilities recently, so of course they are the most exciting ones. It's like connecting to the various LLMs out there, doing prompt engineering, and an easy no-code UI to build your app. That's all fine, and I think I talked about it already quite a bit. So a lot of the interesting aspect of our platform is we decided to build one platform where you could do data prep and all of the data engineering very well, but in a way that can be very accessible by the business, so that business can be independent in terms of doing data prep, and to build a platform where you can also do AutoML at scale, AutoML and MLOps.

Florian Douetteau [15:53] The vision is to have, in one platform, the full end-to-end lifecycle of data as it should be. As in, you connect, you prep, you build your models, you test them, you have all of the backtests there, you move it to production, you can visualize things at every step, and you can lifecycle all of that. Just because I do believe, as a practitioner, that it's easier to do it this way than having multiple different tools or different teams or stages responsible for that. And the other thing of the platform is, I do believe that there is a need to make this more accessible towards non-technologists, as in non-nerds, meaning people that don't love doing Python.

How to build one product for different audiences

Florian Douetteau [16:34] And I say that loving doing Python, but also getting old. So I do know that I'm a very bad coder today. And so I sympathize with everyone that actually needs and has some domain expertise, but won't get into code anytime soon, even with Copilot or whatsoever, and that still need to leverage their data in order to get things done. I think there is this gap that we are fulfilling in the market in terms of providing this platform for the enterprise.

Matt Turck [16:54] And by the way, any lessons learned building this over the last many years about what it takes to build a product which serves different kinds of audiences? Because you have an audience that's data analysts, an audience that's much more technical, and then you're doing the full lifecycle. Yeah, lessons learned building product.

Florian Douetteau [17:15] One of the lessons was the lesson of cloud ETL, because it's not an obvious bet. Because in theory, you should focus on one aspect of data and just do the best product at it. But my own take, which was to some extent very personal, is that data is a plumbing problem where most of the time you waste is connecting things from one step to the other. And so, consequently, there is more value by helping people do each step, but consistently, so that you have one platform where you can get everything done.

LLM Mesh: what is it?

Florian Douetteau [17:45] And this bet, for instance, was, from a product design perspective, a bit controversial indeed. But nonetheless, we made it, and it led us to make some compromises because, indeed, you can't be the best platform at everything while doing everything at the same time. I would say it's a controversial product bet to do so.

Matt Turck [17:57] So the big launch of 2023 was the LLM Mesh, which is your generative AI initiative. Do you want to talk about what it is? What does that do, and where does that fit in the overall picture?

Florian Douetteau [18:26] Some of this audience would be familiar with data mesh. Okay, so why not LLM Mesh? Why not? Because really, if you think about the enterprise, fast-forward a few months, if not a few years, I think that most enterprises will leverage multiple LLMs coming from multiple vendors. You will have large ones, small ones. So technically, you should not call them LLMs, but whatever, easier for everyone. You would have some that would be fine-tuned. You would have some virtual LLMs where you have some RAG, and so they become an LLM augmented with some knowledge or information.

Florian Douetteau [19:04] And so you have to manage all of those LLMs together. And ideally, you would want any kind of application from your company to be able to tap into any of those LLMs with some flexibility. So, having proper routing there, a virtualization layer for that, the same as a data mesh is a contract or virtualization layer for data. And so, talking about contract, it means that you need to add some security or conditions there in order to help them move to production. And then the idea is, can you add a universal layer to ensure privacy filtering, to ensure content filtering, to ensure cost control, given that privacy, content control, and cost are the things that are today the most variable from one LLM to the other.

Florian Douetteau [19:33] And so if you want to virtualize them, you need to be able to have basic guarantees in terms of making sure they don't go crazy, either in terms of cost or in terms of content they push to the outside.

Matt Turck [19:38] And you have multiple partners in the LLM Mesh. Do you want to mention a few?

Florian Douetteau [20:17] Yeah. So we built this in order to essentially integrate with all of the vendors on the market, integrating with Hugging Face, AWS Bedrock, OpenAI, and Google Gemini, of course, Anthropic, and AI21 Labs. Also with vector databases that can provide RAG for virtual LLMs, such as Pinecone, for instance, as a local vendor, but also, of course, LlamaIndex. And so all of this being an ecosystem of partners and vendors that our customers are using in order to build GenAI applications.

Matt Turck [20:34] So the way it would work is I use Pinecone for my vector database to bring in my data. I can bring any of the large language model vendors, and then Dataiku would provide the governance and orchestration layer.

Florian Douetteau [20:54] Exactly. And if you want to switch vector databases or LLMs between your open-source ones, your local ones, or from one vendor to the other, it's more like a dropdown when doing prompt engineering, easily test from one to the other. These kinds of things are important for the enterprise when moving things to production. We see quite a bit of usage patterns where people would start with GPT-4 for design and then decide to move to essentially something cheaper, like Mistral self-hosted or Mixtral self-hosted, when moving to production.

Evolution of platform building with Gen AI

Florian Douetteau [21:11] And so I think that you have to help these kinds of usage patterns.

Matt Turck [21:34] How does architecture evolve, to the earlier part of the discussion, between traditional AI and generative AI? Did you end up with a hybrid kind of architecture where a generative model would call a predictive analytics model in some kind of chain, or do they remain somewhat separate for the foreseeable future?

Florian Douetteau [21:59] When you build the models, especially the RAG models, there are lots of integrations because of the, let's say, data analytics part and the generative AI part. Because indeed, a big problem when doing an actual RAG is that you need actual data. And you need to clean your data and so forth in order to have the relevant data to use. And so at the end of the day, you've still got big data pipelines which can involve some traditional AI in them. And then you've got applications such as personalization or recommender system-type applications where you kind of mix traditional AI, a product recommender system providing the basic inputs that then feed into the generative AI layer to generate the output.

Enterprise AI motion in 2024

Florian Douetteau [22:17] And so, in that sense, you chain them.

Matt Turck [22:32] What does an enterprise AI motion look like for Global 2000 companies? Is there a part of education and services that's involved, or is the market sort of mature in buying? How does it work for Dataiku?

Florian Douetteau [22:57] Well, from my perspective, on the traditional side, it's a fairly mature market where a lot of the movement is associated with the modernization of analytics, which is, from my perspective, a big movement of analytics requiring moving to larger datasets, moving to the cloud, more sophistication involving more ML than before, or more sophisticated ML techniques. And so, it's all about moving from the desktop to the cloud with more stuff. And so, we are part of this movement where typically our customers would also set up cloud data warehouses or data lakes such as Snowflake and Databricks.

Dataiku's partnerships

Florian Douetteau [23:28] And we would be a layer on top, helping the overall orchestration and ease of use by the business. That's a big movement we see in the market. And from my perspective, this movement is, let's say, fairly mature. On the GenAI side, it's, I would say, today more consultative-oriented because many companies are still figuring out the scoping, the cost, and don't have the resources or the bandwidth to do so.

Matt Turck [23:40] And speaking of Databricks and Snowflake, talk about your partnership strategy. Last year was a big year for Dataiku on that front as well. You won three key awards in the space.

Florian Douetteau [24:03] Yeah, we were honored to be last year both AI and ML Partner of the Year for Databricks, Snowflake, and AWS, which is great, meaning we can work with everyone—Snowflake, Databricks, AWS, and so forth. Our focus is actually indeed enabling a broader user base in the enterprise to get faster there, like get faster to, let's stop, just do my own things on my laptop with my old tools, and move to the cloud, get more sophisticated faster, which is complementary to, well, and actually generating consumption and a faster drive for the cloud platforms below us.

Being platform-first as a startup

Matt Turck [24:53] I really like this conversation we had around building a platform, and we talked about it from a product perspective. But from an entrepreneurial perspective, the common wisdom that VCs like me say is that you need to be a tool before you become a platform, meaning you need a wedge to get to product-market fit. And then, over time, you add functionality, and then you end up being a platform. Whereas Dataiku, back in the day, in 2013, at its creation, started as a platform. Do you think that was a moment in time because you were early to the industry, or is it still possible to be a platform-first startup?

Florian Douetteau [25:07] I think it's really a matter of vocabulary. Do you think that VC can be a tool or should be a platform day one?

Matt Turck [25:12] As a VC, we certainly have a reputation for being tools.

Florian Douetteau [25:41] Okay, I didn't say that. And so, again, on our side, it was really about focusing on what is the problem at hand. The way I see it, enterprises need to see analytic pipelines, ML pipelines, and their apps as assets themselves, like IP assets they have to manage over the course of one, two, five, 10 years. And so previously, those assets were not considered as assets, more like programs done by some people you don't talk to on their laptop. But because they are becoming more important, they really need to be managed.

Florian Douetteau [26:02] And so if you want to be the platform managing them, you need to have the ambition fairly early on to be kind of end-to-end there, hence being a platform. I think it's specific to a market where there was this kind of disruption of the lifecycle.

Matt Turck [26:04] Florian, thank you so much. Really appreciate it.