Synthesia: Creating Video with Generative AI with CEO Victor Riparbelli
The MAD Podcast with Matt Turck · with Victor Riparbelli, CEO, Synthesia
Victor Riparbelli is the CEO at Synthesia. We cover why Synthesia’s videos replace text rather than video production, how enterprises use them to turn 40-page training manuals into multilingual video, and why production AI requires workflows for fact-checking rather than a simple API call.
Chapters
Transcript
Full episode
Matt Turck [1:22] Hey Victor, welcome to this podcast. I've been looking forward to the conversation. We're going to talk about generative AI a lot, but maybe to set it up upfront, Synthesia is the world's number-one-rated AI video creation platform, meaning that you use AI to create professional videos without microphones, cameras, or actors. The platform enables one to turn text into high-quality videos with AI avatars, which is one of the things that the company is mostly known for, and voiceover in over 120 languages.
Matt Turck [2:10] And I should add that Synthesia just became a unicorn, which is much rarer these days, even in AI, reaching a billion-dollar valuation in a Series C round of financing that literally just closed and that was led by our friends at Accel. And I should also add that I have been a very proud investor and board member at the company for a couple of years now, but hopefully we can make this conversation objective and interesting to anyone that may listen to it. So I'd love to start with the origin story of the company because obviously generative AI is all the rage right now.
Matt Turck [2:46] But even when you and I first met, the term did not exist. And almost by definition, you had started the company probably a couple of years before that. So what was the vision at the time? Because in retrospect, it seems like a completely crazy idea to start a generative AI company at a time when generative AI just did not exist.
Victor Riparbelli [3:08] Yeah, back then we called it synthetic media. We were hoping that that was going to be the term that would latch on. Unfortunately, it wasn't. We still have some decks where we say generative AI, like, '19, I think. Should have stuck with that. That would have been great for SEO. So the origin story, my path to starting Synthesia: I grew up in Denmark, in Copenhagen. That's where the accent's from. And I figured out in my late teens that I love building products and that I one day wanted to start my own company after doing something like four or five years in the Danish startup ecosystem.
Victor Riparbelli [3:39] But what I also figured out during that time was that I just wasn't super passionate about building accounting software or business process tools, which were mainly the things I'd been involved with at that point in time. I'm a huge nerd in my spare time. I love science fiction. I love the weird, wonderful edges of technology. And I've just always been drawn to that in my life. And I wanted to see if I could combine that with also building a great business and making awesome digital products.
Victor Riparbelli [4:06] So, to make a very long story very short, I decided to move to London back in 2016. I knew I wanted to start a company. I knew I wanted to do something with deep tech. And I basically just spent nine to 12 months working with a bunch of really awesome people, some of whom became my co-founders. It was a lot about VR at the time, which had its—I don't know what hype cycle we're in now. I guess we're almost about to enter the next one.
Victor Riparbelli [4:28] But this is back when Oculus came out, which was a huge milestone for VR. It used a lot of deep learning—I guess more machine learning almost back then—to understand where in the room you are and make the headset work way better than anything we'd seen before. But I wasn't convinced that VR was going to be an interesting enough market to build a big company in at that time. But I got very interested in this underlying breakthrough in AI that had happened.
Victor Riparbelli [4:52] So this is, for those of you who are technical, back when CNNs were the big thing, right? And we'd finally taught a computer to recognize an image of a cat with 70% accuracy. Today we can generate pictures of cats in one million different colors and shapes, and we don't bat an eye at it. But that was one of the big things that happened back then. And while I was working on VR, I met a bunch of really cool people.
Victor Riparbelli [5:17] And one of them was a guy called Matthias Niessner, who was a professor at Stanford at the time. And he'd done, I guess you could say, the seminal work on using deep learning to generate video. Deepfakes is obviously a bit of a dirty word, but that was how most people thought of the technology back then. And when I saw his research paper called Face2Face, I just felt like I saw magic for the first time. I think probably a bit the feeling that a lot of people have had chatting with ChatGPT or hopefully making a Synthesia video, creating images with Stable Diffusion.
Victor Riparbelli [5:49] You feel like there's something so powerful here, and it's just obvious that it's going to change the world. So I could just not let that idea go. And I started to build a thesis around how this would impact video production. We knew video back then was really, really important and was a massively growing market, right? But we also knew that production of video was really difficult. So we got a bunch of people together. Matthias Niessner, the professor from Stanford, became my co-founder.
Victor Riparbelli [6:12] So did Lourdes Agapito from UCL London, also a professor in computer vision and AI. And Steffen, my Danish sidekick, who's a bit more on the sales and finance side of things. And the idea we all gathered around is the same then as it is today. We want to make it easy for people to make video content, but we don't think of that as making smaller, more affordable cameras. We don't think of that as a better app that works on your phone for editing video.
Victor Riparbelli [6:39] Those have been the two main vectors for making video easier in the last decade or two decades or so, right? We are building technology to eventually replace the entire physical production process. Now, six years ago, almost seven years ago, that was very crazy. And the whole journey to get to where we are today had a lot of bumps on the road, to say the least. But I think we've made great strides towards that now. We run the world's biggest AI video platform, have more than 50,000 customers, and work with more than 35% of the Fortune 100.
Victor Riparbelli [7:05] And we do exactly what we set out to do, right? We help people make video in a really easy manner. You just open up your browser, select an avatar, type in the text, and you can build your video around it. And then we'll give you a video in just a matter of minutes. That said, there's still a very long way to go. I see ourselves as being 5 to 10% into the overall roadmap. At one point, we want to make people able to produce much richer and more complex video content than what we're known for today, which is this more green-screen-style presenter, dare I say, maybe slightly robotic type of content that we're doing today.
Victor Riparbelli [7:23] But the end goal is to be able to produce much richer and more complex content.
Matt Turck [7:47] Great, great. Yeah, I think it's worth double-clicking on what the product does today. I would say both the avatar part and the non-avatar part of the stuff around the avatar, because I think that's something fascinating about the product and something that people may not have realized, as well as how much product goes into the actual video creation process in addition to the avatar.
Victor Riparbelli [8:01] Yeah, absolutely. So I think what we're most known for, and we are at our core, we're an AI company, right? We have massive R&D. We're working on bringing these avatars even more to life than what they are today. And that part of the company is essentially about these avatars. Like, you go in, you select one of them, you can create yourself with three to four minutes of footage, and then you type up the script and they'll kind of read that out to the camera.
Victor Riparbelli [8:13] Or perform that to the camera. That's maybe a better way of thinking about it.
Matt Turck [8:16] You have stock avatars and you can create your own.
Victor Riparbelli [8:37] Stock avatars, you can create your own. Exactly. And some interesting insights, I think, when we initially built the technology, text-to-video is what we're calling it, right? And we built it and we were sort of like, this is definitely very cool. There is no doubt about that, right? And this is jaw-droppingly cool. The big question was, who actually has use for this? We had a slight inkling towards that, and we ended up being right on that, which was that in the first phases of Synthesia, we were working a lot with advertising agencies, the film industry.
Victor Riparbelli [9:09] And we had this technology that was cool. It wasn't very scalable. It took a PhD three days to make even just, like, a 10-second clip. It only worked if you looked directly at the camera. And we tried to roll this out into the market by saying, hey, this marketing video that Coca-Cola just did, let's make it in 15 different languages, overdub it, change the facial expressions to match that as well. And that's a way of scaling your video content. And in that process, we actually managed to build a fairly interesting business.
Victor Riparbelli [9:35] I think we did, like, $700,000 of revenue on that in a year. It wasn't terrible, but it just wasn't scalable. It was clearly much more of a service business. But what we figured out during that time was that actually what's much more interesting is that there's billions of people in the world who are desperate to make video content, right? Their house is on fire. Coca-Cola's advertising agency's house is not on fire at all, right? They're already making lots of great video content.
Victor Riparbelli [10:01] And what we found out was that for those people, if we could give them a thousand times more affordable and a thousand times more scalable way of making this content, they would be okay with kind of lowering the quality threshold a little bit, especially for particular types of content. So we built this product, we put it out to the market, knowing that the quality, of course, isn't one-to-one with a real camera, especially not back then. It still isn't 100% there. What we saw was just that these types of videos essentially became not a replacement for video production, but a replacement for text.
Victor Riparbelli [10:27] And that's a very important part of why we're so successful today, especially in the enterprise. Because what we're seeing is that if you're, like, one of the world's biggest fast-food companies, for example, then you have a lot of internal comms, you have a lot of training. That could both be training frontline workers, but it could also be training a sales team and many other things. And what they used to do is, if we take the training example, is you're onboarding millions of frontline workers in your restaurants every single year.
Victor Riparbelli [10:54] How do you do that? You send them a 40-page manual that they have to sit at home, they have to read it, they have to comprehend it, and they have to remember it when they go to work. I think everybody knows that that is not a great way of teaching people. Most people read through it and they'll just forget about it. What they can do now is they can make video content instead. It can be in the native tongue of whoever is consuming it.
Victor Riparbelli [11:21] It can be measured if they're actually finishing it. If they're having difficulties with it, they quiz people on it afterwards. That itself is amazing because they can make video content. Video content is a much better way of communicating with people in 2023. But I think the real unlock that we figured out then was that because it's so easy to use Synthesia, the same people who wrote that 40-page handbook can now make those videos instead. So you're kind of bypassing the video production department.
Victor Riparbelli [11:47] And I think that led us to sort of, I guess, the holy grail of every entrepreneur, is that you become one of the first to build an entirely new market by unlocking new capabilities for a new set of people. And I think that's very much how you should think about the product. Instead of thinking of this as, like, Adobe Premiere, visual effects, cool Hollywood things, I think that's the best way to understand why the product is powerful, right? Because we can go into these very large companies and we can essentially enable anyone at that company to make pretty good video content that is definitely better than text.
Victor Riparbelli [12:17] Might not yet be better than a real recorded video, but we'll eventually get there. So this idea of replacing text with video is the key driver of the product. And that's what we're building around. So we have the AI part, which is around the voices and the avatars. But most videos are not just an avatar talking to the screen, right? That's quite a boring video, and that's definitely not better than text. So we've built this entire creation tool around it.
Victor Riparbelli [12:30] You can think of this as PowerPoint, maybe something like Canva: import your own fonts, put on text, animate things, put in your screen recordings, all those things that make a video a video, right? And I think as much as the AI part is interesting and jaw-dropping, and that's what most people are drawn to, I think the whole product that sits around it in terms of the creation of the video, also the collaboration that we've rolled out now with Teams, workspaces, commenting, the less flashy features, all those things really contribute to making a product that delivers real utility and business value, not just novelty, right?
Victor Riparbelli [13:14] And that's something that we're obsessed with at Synthesia. It's amazing to build technology that's cool and jaw-dropping, but we want to sell utility. We don't want to sell novelty to our customers. And I think if you want to sell that utility piece, you really have to think about the entire workflow of the user and build around that. And I would humbly say that I think that's something we've been fairly good at doing.
Matt Turck [13:42] Okay, wonderful. So many great thoughts here and different directions to unpack this, but maybe you mentioned some use cases, so maybe let's talk about this. So there are two ways people can buy the product. They can self-serve online, or there is an enterprise kind of version. How do people use the product in both cases?
Victor Riparbelli [14:08] So I'd say on the self-service plans, we enable a lot of smaller businesses to make video content. So this is a group of customers who generally haven't had any kind of budget or capability to do video before. This can be anything from, I think recently I was looking through some of our latest signups, like a barber in Amsterdam and one in Brazil who are just making videos for their Facebook page, for example, all the way to maybe just, like, a smaller team in a big company who's just trying out the product at first.
Victor Riparbelli [14:38] But what I generally say is that, for where the technology is at today in terms of the realism and capabilities, what we're seeing is that instructional video content is what it really excels at, right? It's great for information dissemination. So you're making a video for someone who has to understand something, learn something, and you want to do that in a better way than just sending them a long Word document. So that could be customer support, it could be customer success, it could be sales enablement, it could be training of frontline workers, things like that.
Victor Riparbelli [15:14] For those types of use cases, the product really excels. As the technology gets better and better, I think we'll see much more storytelling and emotional types of video content start to take off. But I think the honest answer is that these technologies just aren't really there yet. And this is the first market that we've really seen a very strong product-market fit in. On the enterprise side of things, it's generally one team that'll start working with us. A lot of it, again, revolves around instructional video content.
Victor Riparbelli [15:30] That could be the example I described before with something like a fast food company, for example, but it can also be one of our customers who are using Synthesia to train and enable their 4,000-person sales force. This is, again, where every time you communicate with someone on your sales team about a new feature, a change in the product landscape or market landscape, if you can increase the information retention by four or five times versus sending them an email or pointing them to a Word document where they can read about it, that's massive.
Victor Riparbelli [16:16] That's huge value. And that's, again, about more information dissemination than it is about emotional storytelling and entertaining you, and creating that kind of content. I would say that that's the red thread between the customers. It's a lot about instructional video content more than it is about very salesy, emotional storytelling content. But that'll definitely be there, I think, before the end of the year as we roll out the next generation of our avatar technology. We'll start to see even more marketing, sales, and those types of use cases emerge as well.
Matt Turck [16:38] Great. I think an interesting question for the company, but for AI companies and generative AI companies in general, is: it seems that there is something new that comes up every day, or sometimes several times a day. As you build a truly AI-native company where you've developed all your core technology, how do you think about balancing sort of homegrown versus making sure that you're flexible enough to bring in anything new and interesting that may pop up in the world?
Victor Riparbelli [17:18] It's a great question. I don't think we have the answer. I don't think anyone really has. The pace of AI development the last 12 months has just been crazy. I think I saw a tweet today that today is, I think, 12 months since Stable Diffusion got released, which was a pivotal moment for generative AI. And it just feels like so much has happened since then. I think one thing we think a lot about is: what do we want to be the best in the world at?
Victor Riparbelli [17:46] That's definitely not going to be training generalized large language models for writing the script of your video, for example. What we want to be best in the world at is digital humans: creating insanely photorealistic humans that cannot be distinguished from real video, that work every time, not one out of 30 times when you give it the right prompt, and that are kind of world-leading in terms of capabilities and what they can do, right? So that's very much how we think about our roadmap on the AI side of things.
Victor Riparbelli [18:15] We use lots of LLM providers now. For example, we have a script-writing functionality where you can just type in the topic of your video, and we'll help you write the script. That's obviously not technology we've developed from the ground up. I think it's really important to see all these other things that are happening in generative AI as force multipliers on what we can do really well, and make sure that what we're doing really well, we're world-leading at. And I think that puts you into an interesting product space where, because everything is moving so fast and you want to build these AI-native products where you build around AI capabilities, you don't just bolt them on.
Victor Riparbelli [18:42] It is really hard, right? I mean, I think we've definitely screwed up a few times where we thought a piece of technology was further ahead than it actually was, and also the opposite way around. But I think overall it's just exciting for builders right now because these technologies unlock so much, right? And it means that you can, for us, for example, not just innovate on removing the camera and enabling you to make these videos of humans you can use in your video editing tool, but if you kind of combine that with the capabilities of LLMs and diffusion models for generating images and other kinds of things, right?
Victor Riparbelli [19:11] You can really build an entirely new product that is very, very different from what video editing has been before. And I think that that's the real opportunity for everyone right now. In terms of the AI side of things, that is hard. And I think it's interesting to see that there's this new class of companies now where the R&D teams are not a satellite team that's sitting in the corner somewhere and people are just waiting for them to come up with something amazing to put into the product.
Victor Riparbelli [19:49] For someone like us, that is the core of the product. And having two teams, you have one AI research team for us that is doing deep fundamental research, advancing the core capabilities of avatars and voices. And then you have another team which is more like your traditional SaaS product engineering team that builds the platform, builds the video editor, and all those types of things. And they kind of weigh equally to some extent. That's a very interesting kind of structure of company and how those two teams work together.
Victor Riparbelli [20:10] I think we'll see a new paradigm of how you build these types of companies well because it hasn't really been done before. And every other CEO I talk to who has a company structured like this is trying to figure out what's the right interaction model, what's the right org structure, and all that stuff.
Matt Turck [20:35] Yeah, that's good. You anticipated my next question here. How do you balance moonshot research with deliverables for an AI team? And do people get together and agree on the roadmap, on what might be feasible? And how do you know when to push or cut your losses on an AI research project?
Victor Riparbelli [20:59] It's really, really difficult. I just want to preface that by saying that I don't think we have the right answer. But I think the mental model that we're sort of moving towards now is basically looking at teams in terms of how far removed they are from product. On the one hand, you have blue-sky teams that are really doing things that might not hit the product for two years because this is betting on the next really big step change. In the middle, you might have something which is trying to advance the capabilities of what you have right now.
Victor Riparbelli [21:27] And then you might have a team that's almost more AI engineering, which is taking more stuff that already exists and quick wins where, in three months, you might be able to ship a feature. And then looking at those three things, figuring out how much resource do you place where, how do you manage the interaction with the product engineering teams, what do you do when a piece of research is done, how does that get put into the product. You want to try and parallelize those two things to ship faster.
Victor Riparbelli [21:51] But it's also really difficult because when you're doing AI research, nobody really knows if it's going to take three months or six months or nine months or 18 months, or if it's even possible. So it's very hard to plan for. And that's one of the things I think is difficult. In terms of figuring out the next technical direction, it's again this sort of weird balance of not chasing shiny new objects. If you do that, then every week a new paper comes out, you might want to—holy f—, we should just drop everything we have and just go for this thing.
Victor Riparbelli [22:11] That's a bad idea. It's also a bad idea to stay in your, "Oh no, we're doing it this way. This will definitely work." You have to find somewhere in the middle. The way we've done it is a little bit of spread betting. So, kind of like a very core roadmap that we're working towards, but then also making sure that we're exploring other interesting ways of solving these problems, at a bare minimum, just to understand how they work and what their capabilities are and what the limitations are.
Victor Riparbelli [22:38] And then making sure that you're flexible on the roadmap. But it is really difficult, especially when things move as fast as they are right now. But I would say one thing we've definitely seen is chasing shiny new objects is a very real trap. I think a lot of companies fall into this, and it's important to make sure you get the balance right of being open-minded to new ways of solving problems, but also not getting pulled in a new direction every second month.
Victor Riparbelli [22:54] That's not going to yield the results that anyone is after either.
Matt Turck [22:55] Yeah.
Victor Riparbelli [23:16] I think it's interesting. Maybe something I would add to this as just an interesting observation is because things have moved so fast the last year, and I think a lot of people this week and last week have argued that the AI hype is waning, which I definitely think is true. It's also, we have so much new technology that came out, and there's so many interesting things you can do with it. It's so accessible, right? Like LLMs, everyone could go in and try and use it and be impressed at how good they are.
Victor Riparbelli [23:36] Same thing with image generation, Synthesia videos, and things like that. But I think what we're also finding now is that actually it's very different doing a cool Twitter demo and actually putting something into production. If you take the large language models, for example, it's magic technology. It's amazing. I'm by no means downplaying it, but I can say, as someone who's definitely using them in production, they're not production-ready for 95% of the tasks that people think that they want to use them for.
Victor Riparbelli [24:09] They'll probably get there, but there's still a lot of problems with these technologies, and they will just take longer, I think, than people anticipate. I think that's also when things are compressed so much that people start all these projects and the future looks so bright, you could build everything with it. And then you get into this sort of realm where we're at right now, where people are freaking out. How much do you actually want to trust the output of these models? Do you really want to put this on high-stakes things in your company?
Victor Riparbelli [24:33] And if you don't want to do that, then all of a sudden you have to start building a whole UX around making sure that you are fact-checking these models or have the right kind of product that sits around it in place to make them work in production. And that's a very different thing than just an API call to OpenAI, for example, which I think a lot of people thought they could just build entire workflows around.
Matt Turck [25:00] Yeah, it seems that the winning model that seems to be emerging is this concept of full-stack AI company. And I mean, that's very much what Synthesia is, where there's work at the foundational level, at the AI level, and then there's a software layer around workflow and collaboration, and then there's another layer around verticalization and user-touching kind of features.
Victor Riparbelli [25:02] Do you—
Matt Turck [25:18] I mean, would you agree with that? Is that one of the most likely-to-succeed ways of building a company? Or are there other potential models for success in building an AI company?
Victor Riparbelli [25:40] Yeah, I think I agree. I think it depends a bit what your desired outcome is, right? I think if you're bootstrapping and just want to build a great business for yourself, you want to have 15 people employed, probably you could build a great wrapper company. But if you're building just a wrapper company, you probably shouldn't be 300 people and raise a lot of VC money. That might not end out the way that you envisioned. I think in terms of building a full-stack AI company, I agree.
Victor Riparbelli [26:05] I do think that an added dimension to this is if you really want to build something big, I think it's quite important that you're not just building a company which is kind of an existing product with some AI bolted onto it. Because I think incumbents also, as easy as it is to build something on the OpenAI API for someone who's just starting out, as easy as it is, in theory at least, for a big company to do the same thing. I think it's really important to think about the product that you're building, building that AI-first, and thinking deeply about what are you doing differently than incumbents and how does AI change not just one feature in the product, but the value proposition of the product.
Victor Riparbelli [26:43] If you take something like customer support, for example, which is an obvious use of LLMs, it's fairly low-stakes. You can fine-tune the model on your knowledge base and you can probably get something out that works pretty well. There's probably people out there right now starting companies around doing that, not realizing that if you're one of the big chat support companies out there today, they can also just take the OpenAI API and they already have built the collaboration. They are enterprise-ready.
Victor Riparbelli [27:07] They have all these things. So if that's the only thing you're changing, it's just that in the chat box that all companies already have on their website because they're using Ada, Intercom, something like that, that's not a very defensible company, even if you build a full stack. You want to rethink something fundamentally. And I would humbly say that with Synthesia, I think we've kind of hit on that a little bit because we're not actually replacing video production. We're not replacing anything in the enterprise.
Victor Riparbelli [27:31] And I feel like it's a good kind of acid test for what you're doing. Are you replacing something in an enterprise, or are you selling something that's net new to them? I don't go to my customers and say, "Hey, you should stop using Adobe Premiere or you should stop using whatever other video editing application they're using." This is something that's net new for you, and this is a new technology you should adopt. It doesn't have to be completely like that, of course, but I think if you imagine there's a scale of selling something that's completely new, no one has seen it before in the enterprise, and you're trying to rip out a piece of existing software because you have some smart AI features, I think you want to have some skew more towards being something that's kind of net new, right?
Victor Riparbelli [27:53] I think that's one interesting lens to view it through.
Matt Turck [28:24] No, that's definitely a very interesting lens. And that tends to lead to the conclusion that actually the universe of AI-native or generative AI startups that one can build is more limited than people thought six months ago, where everything was possible. But in reality, the opportunity for entrepreneurs and the investors who love them is actually more narrow, because you need to create something that's new and different, as opposed to enhancing something with a new technology.
Victor Riparbelli [28:45] I agree. And, yeah, I think it's kind of interesting, right? Because history repeats itself with this. And I think what we've seen the last one year is that 95% of people immediately go to first-order thinking, right? Like, oh, this could make this chatbot better, this could make my scriptwriting thing better, or whatever. And that's just rarely how technology pans out. I think what we have yet to see—we're slowly starting to see it—is how all these technologies will fundamentally change, like in my world, video, for example.
Victor Riparbelli [29:17] Video today is generally a broadcast medium. You have one video and you show that to millions of people. That's very much a constraint of having to work with a camera, because you're not going to film a million different videos for a million different people. That would be completely unfeasible to do with a camera. But now that we're starting to generate video and image and audio content via code, that opens up a bunch of new possibilities. We don't really yet know how it's going to pan out, but I'm pretty sure that in five years' time, AI video is going to look very different from a normal video, just as I think that AI voice is going to look very different from what a voiceover actor does today, right?
Victor Riparbelli [29:51] It opens up a lot of new markets and possibilities. And to give a very different analogy to that, because music production is my hobby, I think it's very interesting to see how technology has been adopted in that world. If you think back to when drum machines were invented 20, 30, 40 years ago, they were invented to actually replace drummers, right? The idea was that you can just use a drum machine and you don't have to record real drums anymore.
Victor Riparbelli [30:21] Same thing with synthesizers and other kinds of technologies. And it turned out that nobody really wanted to replace their drummers or their real pianos with these types of software, but they opened up a bunch of new things. It became a new genre: electronica, house music, techno music. All those things derive from those technologies. That didn't happen the first day the drum machine was released. It took some time for people to figure out what to do with this stuff. But I think there is also something to take for people in the business world of not just falling into the first-order thinking, the obvious answer to what these technologies are going to get used for.
Victor Riparbelli [30:54] I think if you try and take a step back and think from first principles, how is this going to fundamentally change this technology or this medium, and build a thesis around that, I think you'll have a higher chance of success. Because if there's one thing we've seen, it's that technology is hard to predict, but it rarely pans out the way that the McKinsey consultants predict in the first report. That's rarely how the world unfolds.
Matt Turck [31:29] Let's talk about the data part a little bit as well, because obviously that's one major topic in AI and for startups. So there is obviously that big race on the LLM front to just crawl the internet and also have access to some private databases, and all sorts of ethical and legal and copyright questions that come up with that. How have you thought about the data aspect of building Synthesia?
Victor Riparbelli [31:48] So I think we're at a really interesting point in time right now. Five years ago, or one year ago maybe even, most companies doing AI were building much smaller models than what we're seeing today. You'd get a smaller quantity of high-quality data, and then you'd train a synthetic voice or predictive system for whatever thing that you wanted to do. And in that world, it was generally at least a lot easier to be data-compliant and just go out and procure the data that you actually needed to train your systems.
Victor Riparbelli [32:16] To give you one example, let's say you're training a text-to-speech system. Then, with the old-school paradigm, how you would do that, maybe you would need a couple of hundred hours of data to train your initial model. Then you'd need 30 minutes of your data to train your voice. That is within the realm of what's feasible to get, and pay your way to get. You pay the voice actors to build that initial dataset, you could build your technology, and then you have a completely clean dataset.
Victor Riparbelli [32:40] That's basically how we've always run the company. Now what we're seeing is that the bigger the models get, the more data they ingest, the better they become. Some of the state-of-the-art text-to-speech systems today, they're not trained on 300 hours of data, they're trained on half a million or a million hours of data. Once you get to that scale, it becomes very difficult to attain that data without basically scraping the internet, which is of course what a lot of these companies have done.
Victor Riparbelli [33:16] And if you go to large-language-model-type scale, basically these technologies are not going to work unless you have internet-scale data. And at this point, it's never going to be possible to get consent from everyone on the entire internet to train your system like OpenAI has done. So I think that's interesting because we now have a class of technologies that are basically impossible to build unless you scrape the internet. And it's to be determined whether that's legal or not. But there's a lot of lawsuits going on right now.
Victor Riparbelli [33:34] And I think it presents a lot of interesting questions for people who build AI companies. What I'm really just hoping for is clarity as soon as possible. We've kind of decided to take a route where we only train on clean, compliant data. There's other companies who are not doing that. And right now, you're flipping a coin that either maybe this scraping business is going to be okay—we're going to determine that as long as you're not reproducing any of the original content, it's okay.
Victor Riparbelli [34:08] Your computer can kind of browse the internet and understand what an image looks like or what a human voice sounds like. Or it's going to be a different world where if you're training on data you don't have access to and it's not compliant, you're going to get in big trouble. And that's going to be interesting to see how that's going to pan out. As I said, from the beginning, we've sort of always procured our datasets. We are only training on clean data, and I want to keep it that way.
Victor Riparbelli [34:36] Even though it's a lot harder, I feel like it's morally the right thing to do first and foremost. But I also think that as we progress further and further into generative AI, the really big enterprise customers will care a lot about this stuff. We've already seen that today. A lot of big enterprises are restricting people from using ChatGPT because they're afraid that it'll spit out copyrighted content. They're not comfortable with using models where you haven't trained on clean data.
Victor Riparbelli [35:03] And I think we're going to see that a lot in the next couple of years, that this is going to be a requirement for big companies to work with generative AI. Adobe is another good example of this. They trained Firefly, which is their image generator, entirely on compliant data. I forgot which stock provider they work with, but they also have their own stock universe. So they trained it on all that data. And that basically means that if you go into the Adobe Firefly functionality in your Photoshop and you type in, "Make me an image of Spider-Man," it's not going to understand what you mean because there are no images of Spider-Man in the original dataset, which means that you as a customer are completely—there's absolutely no chance you'll generate something that's potentially copyrighted or might get you in trouble, right?
Victor Riparbelli [35:31] And this is what I think we're going to see a lot more of in the coming years.
Matt Turck [35:58] While we are on the topic of legal matters and ethical matters, you mentioned the term deepfake at the beginning of this conversation. How do you think about this right now? I mean, obviously, that could be a gut reaction that people could have when they see avatars and imagine that you can make Tom Cruise speak in your voice and all the things. What safeguards have you put in place, and how do you think about the topic right now?
Victor Riparbelli [36:29] Yeah, so I think AI safety is, of course, a huge topic right now and something that we've been thinking about for almost seven years, since we founded the company, on our ethical framework. I think there's lots of things to unpack here. I think for us, or for me, there's sort of two angles I look at this through. One is, how do we secure Synthesia from being misused by malicious actors? I think that one is very much around consent. We'd never create an avatar of someone without their full, actual, explicit consent.
Victor Riparbelli [36:59] You can't just go in and upload video images of someone that you don't know. And the second one is content moderation. So that's something we're investing heavily in: trust and safety. Basically, we take a quite hard stance on what kind of content you're allowed to create on Synthesia and what kind of content you're not allowed to create on Synthesia. As with anything content moderation, we're not perfect; we're always improving. And it is really difficult to ascertain if someone is making great videos about how blockchain works or if they're trying to push you into a get-rich-quick scam, right?
Victor Riparbelli [37:33] That's hard to do automatically. But those are some of the things where we go in and actually moderate the content at the point of creation, right? So you can't even create the video if you're creating content about the topic that we don't allow you to. Then I think the second part of it is: what can we as society do to ensure that these technologies are not misused? They definitely will be misused. There's no doubt about that. This week there's been a big story about WormGPT, which is an open-source GPT clone that's basically been fine-tuned for cybercriminals to use.
Victor Riparbelli [38:04] I think everybody knew that was going to happen, but it's really sad to see that out there in the world right now. And I think what we'll see is that most big companies will put guardrails on the technology that will ensure that harmful use is minimized. OpenAI does the same thing. You can't get OpenAI ChatGPT to say a lot of things. If you ask it how to make a bomb, it will not respond. But the open-source world, obviously, we can't really gatekeep these things.
Victor Riparbelli [38:34] And I think we'll see that a lot of harmful use of these technologies will emerge from the open-source world. That's not because I'm against open source. I'm very pro-open source, but that's just a reality we have to figure out what we want to do about, right? So that could be a very long-winded sort of topic, but I think we essentially have to rethink how the entire internet ecosystem works if we really want to build a trusted internet. We probably want to have something like an SSL-type standard, but for content, so that when you watch a video on YouTube or Facebook or TikTok, whatever, you have some kind of an idea of the provenance of that video: who created it, how was it created, has it been altered since then, and essentially give it a green checkmark if it's trusted.
Victor Riparbelli [39:08] And if it's not trusted, we don't know where it's from, we don't know how it was created, you'll get a kind of a red checkmark or a warning or something like that. But that's a pretty humongous project to undertake. We're part of the C2PA.
Matt Turck [39:17] Which, by the way, is a fascinating idea because it means flipping the logic, like assuming that all content is fake except if it's proven not fake.
Victor Riparbelli [39:30] Exactly. And I think it's a massive task, right? And it's going to take a long while, I think, for this to materialize. But YouTube has done it, right? If you upload a video to YouTube today, they'll listen to the soundtrack. And if there's a copyrighted piece of music in there, which they check against a database of all the world's copyrighted music, they'll slap an ad in there and they'll pay something back to the rights holders.
Victor Riparbelli [40:02] So it does work for music today. Of course, the catalog of all the world's copyrighted music is still small compared to all the content in the world. But we have seen a system like this actually work in production before. And I always talk about this like Shazam, but for every piece of content. You want to be able to have a piece of content and just get the information, like how and when was it created. But we're part of the Content Authenticity Initiative, which is led by Adobe.
Victor Riparbelli [40:16] And that's all about essentially fingerprinting content. And I do think that this is a very, very interesting solution to making the internet and content on the internet more trusted as these technologies evolve.
Matt Turck [40:44] Great. And maybe to wrap up, zooming out, leaving your Synthesia hat off, what else do you find interesting in AI these days or generative AI, whether that's, I don't know, a company, a product, a project, any sort of tip and recommendation for people listening to this who may want to explore other things?
Victor Riparbelli [40:59] I think there's so much interesting stuff that's going on right now. I'll say one thing I'm very interested in right now are companies who aren't selling the tools but actually creating the content using generative AI. We've seen some pretty wild successes with this. A company creating an entire manga fiction series, for example, using generative AI, where essentially the value prop here is that you can build amazing content at 1,000 times the speed and 1,000 times less cost than you would otherwise have to do.
Victor Riparbelli [41:35] And I think it's a great example of it's still about telling a great story. The people who are behind this are still great storytellers. They've just automated the production process, which allows them to scale up significantly the production and create much more kind of side content as opposed to kind of just the main storyline. And this is something that's commercially already very successful. AI has probably surprised a lot of people where they're not selling the tools to build chat interfaces.
Victor Riparbelli [42:04] They're actually selling you a chat experience where you get to talk to a computer. And a lot of people really, really like this. And I think it goes a bit back to the point I made earlier, right? The most successful companies in this will do things that seem weird and that are kind of off-center. And it's essentially a new media format. And I think that's what's being demonstrated by something like Character.AI, right? That's actually a product where you're talking to a computer, and people are paying for it, and people love it.
Victor Riparbelli [42:32] That's not the first thing people would imagine when you think of these technologies. And I think there's lots of other things like this. But I really think that one of the biggest opportunities right now is actually not necessarily selling the tools; it's actually creating the content and monetizing that content, because the tools are becoming so good now that this becomes very, very viable.
Matt Turck [43:01] Yeah, could not agree more on Character.AI. And again, I'm saying this as somebody that was chatting over the weekend with French Emperor Napoleon and telling him that the movie was done after him, and he asked me who was playing him, and I said Joaquin Phoenix, and he was very pleased. He thought that was a good choice. That's how I spent my Saturday evening, which is not at all a weird experience. But anyway, that feels like a really good place to leave it.
Matt Turck [43:27] Thank you so much, Victor. Really enjoyed the conversation, as always. And as somebody, again, who has a biased opinion on this, you're building an absolutely incredible company, and it's been a remarkable journey. So I look forward to your continued success and really appreciate your spending time with us today.
Victor Riparbelli [43:29] Thanks, Matt. Thanks for having me. Fun.
Matt Turck [43:34] Thanks. Bye.
Victor Riparbelli [43:59] Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.