Why generative AI will unlock new art forms | Cris Valenzuela from Runway
The MAD Podcast with Matt Turck · with Cris Valenzuela, Co-founder & CEO, Runway
Cris Valenzuela is the Co-founder & CEO at Runway. We cover why generative AI should develop into a new art form rather than copy existing media, why creative software must move beyond manipulating tracks and pixels toward natural-language interaction, and how Runway gives creators control over video generation through Motion Brush and camera controls.
Chapters
- 0:55 — What is Runway?
- 3:09 — Runway started before the GenAI boom. How?
- 4:41 — What do people get wrong about GenAI?
- 7:18 — How AI is going to change creative software?
- 8:44 — What is Gen-2?
- 12:02 — Runway's role in creating Stable Diffusion
- 14:25 — Gen-1: a model or a product?
- 15:11 — Runway's evolution from image generation to video
- 18:18 — Runway partnered with Getty. Why?
- 19:52 — How has the AI video generation ecosystem evolved?
- 21:58 — Adoption cyсle for AI video generation. Where are we now?
- 24:45 — Challenges of building a research-focused company
- 26:25 — How to build and maintain a soul in a startup?
- 28:27 — "It's like an invention of new art form" -
Transcript
What is Runway?
Cris Valenzuela [1:20] We've been working on Runway for almost five and a half years, but the idea of Runway actually dates to like seven, eight years ago. And the reason really why we started was it was perhaps the time that AlexNet was published and a few breakthroughs in this wave of AI was taking shape and form. And me and my co-founders were just a few blocks from here at NYU, tinkering with these convolutional neural networks and models, and mostly at the time simple versions of what we have right now.
Cris Valenzuela [1:59] And I think we realized that a lot of the potential applications of those models were much more focused on things that we weren't really interested in. We were actually really interested in art and creativity and filmmaking and storytelling. And so we thought to ourselves, why not just do the research to try to get these models to do these kinds of things? And that's what took us like five and a half years. And so the origins of the company are pretty much rooted in that idea of building a new suite of creative tools using these algorithms as the basis for it.
Cris Valenzuela [2:40] It's been definitely a journey. Trying to even hire for generative AI seven years ago was very hard. Yeah, I get that reaction a lot. People were laughing that it's like the images were just bad, and even the term wasn't really that common. And now I think we've just been very persistent on it, very stubborn. And now we have a great team of researchers that's made really meaningful progress in the field. We have a team—we're like 82 people here total, 85% just research engineering.
Cris Valenzuela [3:08] We're based in New York. We just opened an office in San Francisco. We have millions of users. We work with creative teams across the board, from Madonna's creative teams to filmmakers that are making blockbuster movies to creators and TikTokers and YouTubers. Podcasters, maybe. And yeah, it's been quite a journey so far.
Runway started before the GenAI boom. How?
Matt Turck [3:25] And by the way, it truly is amazing that you started the company this many years ago, before the whole generative AI thing. What gave you some level of comfort that it could be done then? Was it Transformers? What was the problem?
Cris Valenzuela [3:51] We didn't know it could be done. I guess I just wanted to try to see if we could. Transformers were not really that popular at the time. I think we had LSTMs and RNNs at the time to generate sequences. Image was working really well—not really well, sorry. GANs were the state-of-the-art thing at the time. But for me, it wasn't really one algorithm or one technique. It was more when I started delving into AI, it felt very early. PyTorch was four months old and TensorFlow was a year old, and it all felt really early.
Cris Valenzuela [4:28] And then we went from Hot Dog, Not Hot Dog—I don't know if you'll remember that algorithm, right?—from that to the first text-to-image model that I saw in 2016. There was a GAN-based model that generated images in really low resolution. But I felt like we went from an idea to this in very short amounts of time. This is going to continue to trend interestingly, and so we should just invest time in it. And I think that was more of the conviction, less of a model, but more of the trajectory of how things were progressing.
What do people get wrong about GenAI?
Matt Turck [4:48] Fast forward to today. I want to ask you about something really interesting you wrote recently. What do people get wrong about generative AI media?
Cris Valenzuela [5:12] I think there is a common misconception that the first way to think about images and videos and content and audio and all sorts of multimedia things that you can generate with these models is that you're going to make the previous thing that we make, but AI. So you hear a lot about AI films or AI novels or AI images, right? But that, for me, is only the first step towards something much more interesting. And I keep comparing this to perhaps a very similar moment that we had as a society, specifically when it comes to storytelling, when the invention of the camera happened, like, 120 years ago.
Cris Valenzuela [5:50] The camera is not a better paintbrush. It doesn't work in the same way that you use a paintbrush. You can get inspiration from how painting works, and that's actually how filmmakers used to think about the camera and how photographers used to think about the camera. But then eventually, the camera evolved as its own medium. It evolved in its own ways with its own mechanisms and metaphors and primitives and artists. Eventually, it actually created an entire new art form that we call cinema.
Cris Valenzuela [6:13] Cinema is a new art form that was created entirely by a technological breakthrough. And so, for me, we're in that early journey of understanding AI as some sort of a new camera. And the first thing we think is, like, oh, this is going to be like painting, so AI paintings, right? But really, it's not. And I think what's going to be different now is that if you compound how fast things are moving, we're going to get to a world really soon where we're going to be able to generate rich multimedia formats and audio and videos in almost real time, which means that you can start visualizing things that perhaps used to take you months or even years in a very distributed manner.
Cris Valenzuela [7:02] And I think that resembles much more of a video game experience than films or images. I don't think we have the name for it yet. I was calling it world forging because you're creating worlds, and anyone can create worlds, and those worlds can get shared and distributed. But you're also going to have artists that are going to do the world-building kind of simulation design. And you're also going to have filmmaking. We still have paintings, but they're just different mediums.
How AI is going to change creative software?
Cris Valenzuela [7:19] And so we're going to branch, I think, at some point. And I think for me, the most interesting thing is to try to build the next frontier of things, not to try to incrementally make better the films that we have right now.
Matt Turck [7:19] 0?
Cris Valenzuela [7:49] So creative software really hasn't changed. If you think about everything we use these days, from nonlinear editors to vector design to image manipulation software, they were all primitives that were invented 20 years ago based on this idea that you can manipulate photography and you can manipulate videos, and you can edit them and you can combine them and you can composite them and so on and so forth. But if you think about this new category of things that you can do with computers, you can interact with them with natural language. You can think with them and reason with them, perhaps in the same way that we reason. Then the type of interfaces and software approaches are not going to fit really well with this idea of having to move tracks and pixels around.
Cris Valenzuela [8:36] Which is, for me, the old way. That's, again, it's like painting. You're not going to do photography in the same way that you're painting. And so when it comes to software, I think it's one of those biggest-moment opportunities where it's like a greenfield and you can do a lot of things, but you need to break the mold of thinking about this as a new version of a database or a new version of the previous thing. And for us, that's really exciting because there's so many things to think and build and experiment and do that we haven't even thought of before.
What is Gen-2?
Matt Turck [8:54] Gen-2 is your latest video generation model. What does it do, and how is that different from prior versions?
Cris Valenzuela [9:20] So around 10 months ago, we released the first-ever commercial video generation model called Gen-1. I remember sitting with a friend, also a founder who's deeply technical, and he was asking me, "What about video?" Because images were very popular at the time, and it was like a year ago, and I knew Gen-1 was coming. And so he told me, like, "You're gonna be able to generate video in a few more days, or very soon?" And he was like, "No, it's decades away."
Cris Valenzuela [9:49] And then here we are. And now we're complaining that, like, "Well, you can actually generate more than 45 seconds." So you will, you will be able to generate all those things. And so Gen-1 was the first model that we released like 10 months ago. Gen-2 was the second model we released, the second version of that model, too, kind of. We haven't actually formalized it that way, but there's some big updates. And then, of course, Gen-3 will come and Gen-4 will come.
Cris Valenzuela [10:08] And every step of the way, those big updates represent a good improvement in quality and resolution, but also in control. So with Gen-2 today, you can do video generation, but also you can control a lot of that generation and how it happens.
Matt Turck [10:16] And you can create video using Runway with eight different kinds of inputs?
Cris Valenzuela [10:41] Yeah, that's critical for us. We really think about who's using this and for what. And if you're creative, if you wanna tell a story, if there's something you wanna say, you want to have full control. If you don't have control over the way you're using something, then it probably won't matter because it will be just like a system creating things for you without you actually being the one with the agency. And so we have nine different ways of controlling those models, and not all of them will last.
Cris Valenzuela [11:12] Some will change, we'll come up with new ones. And so we're experimenting with this new medium, but there's, for example, Motion Brush. So Motion Brush is this idea that we train a model and a way of manipulating video just by—this is inspiration that comes a lot from how filmmakers and art directors actually give reference on films. When you have a photogram, you sometimes just take a pen and draw on top of it and define how you want movement to happen. And so we took that inspiration.
Cris Valenzuela [11:38] You have an image, you can basically draw on top, define how things you want to move, and then the model will actually move them. And so that's Motion Brush, and you have like eight different pens you can have for that. Another one is camera controls. You have an image, let's say you wanna generate a video around it, but you want the camera to start moving to the left. You can actually just do that. You could just move the camera to the left, and you have options and buttons for that.
Cris Valenzuela [11:52] So really, a lot of the work on those models is not only the technical foundation layer, but also the way you interact with them is equally important.
Runway's role in creating Stable Diffusion
Matt Turck [12:14] How does it all work behind the scenes? The models and what enabled the jump from Gen-1 to Gen-2. You guys were the original co-creators of that very famous model called Stable Diffusion. Walk us through the history and what happens behind the scenes from a core technology perspective.
Cris Valenzuela [12:44] So we do a lot of research collaborations. We've done with the University of Washington, with Carnegie Mellon, with LMU Munich, which is the one we co-authored a lot of the research for Latent Diffusion with, which is one of the first versions of, I guess, what's known as Stable Diffusion now. That model was published like two years ago, and it was pretty much under the radar at the time. It was pretty well cited and recognized among some small research communities. But then we just scaled the training of that same model, and that became Stable Diffusion, just a larger version of Latent Diffusion.
Cris Valenzuela [13:12] We open-sourced that model, and that model just took a life of its own. I think the reason for that was, well, a few things. It was the right thing at the right moment. And so people were really excited about contributing to open source and taking a model that was easy to use and easy to fine-tune as well. But the second thing is it introduces this idea that you can generate images in the latent domain instead of the pixel domain.
Cris Valenzuela [13:49] And that just speeds up a lot of the processes, which makes it much easier for someone to customize this without having to have a huge infrastructure. So there's one engineer, for example. He fine-tuned a version of Stable Diffusion with his own 4090. Incredible work. And he figured out things that no researchers had figured out at the time. Of course, we hired him. But that proved the point that sometimes you don't have to have large compute resources. You just have to have a lot of creativity and ingenuity to figure things out.
Cris Valenzuela [14:19] But that model in itself now has become one of the very popular open-source models that a lot of people have built on. But it was kind of an unpredictable path at the beginning. It was more us open-sourcing research, us trying to go back to the vision of what we were saying before: hey, we want to build these tools and these systems. There's nothing yet. Well, let's start from the basis and start doing it. That's it.
Gen-1: a model or a product?
Matt Turck [14:27] And so, do you use Stable Diffusion still at the core? Is Gen-1 a product, or is that a model?
Cris Valenzuela [14:50] No, it's both. The model is—we published a paper for Gen-1. We haven't yet published a paper for Gen-2, but Stable Diffusion is now like a backbone, I would say, of a lot of image generation tools. Maybe not the entirety of it, but some ideas. And that's actually how research tends to evolve. It's not that you take the model and just implement it. Actually, you can take the ideas of the backbone, or the ideas of how the data is fine-tuned, or some ideas on how you start the training and the noise schedules.
Cris Valenzuela [15:10] And so, Stable Diffusion, I think, and Latent Diffusion, which was the first model, was a good set of ideas on how to get to good-resolution images in fast and reliable ways. Yeah.
Runway's evolution from image generation to video
Matt Turck [15:22] And the evolution from image to video, was that just—is that a cousin, or is that, like, a descendant? Is that the same kind of nature of problem, or is that a whole different problem?
Cris Valenzuela [15:47] It's the same. Again, if you think about our problem and the reason as to why Runway exists, Runway is not an image company. We're not a video company. We're not a 3D company. We're a company—our mission statement, and it's on our website and in the foundation of the company itself, is that we want to advance creativity. We want to help tell stories. We want to help people tell the best stories they've ever told. And to get there, well, you need to do a lot of work, and we're going to start with images, and we're going to start with videos, and we're going to start with making sure that all of those things work in a combined way, because the vision of Runway is to make sure that if you have something you want to say, you can use these tools to get that out.
Cris Valenzuela [16:34] And the best way to get that out is by using a combination of different media formats. And so we've never thought of Runway as a text-to-image company. I think there's a lot of companies that are focusing on specific functions or modalities. I think, in the speed of how things are moving, that's very transitory, and it's like a local maximum. It won't matter how you generate things. It will matter that you can generate things. It won't matter. People used to reflect and emphasize that you make things with Photoshop.
Cris Valenzuela [16:56] Now you just edit pictures. You don't think about it. And so, in a way, you want to get to a point where you stop actually emphasizing how things are made and you just focus on things, which is, for me, the core aspect of—that's for me when we're going to be closer to achieving our mission.
Matt Turck [17:10] Do you have just access to some large amounts of compute? Do you have a partnership that—I don't know if that's disclosed or not. How does that work? Are you throwing a lot of compute at the problem?
Cris Valenzuela [17:14] There's parts I can go deeper into.
Matt Turck [17:17] Yeah, yeah, don't say, obviously, anything that you—
Cris Valenzuela [17:41] No, so, yes, of course there's a lot of data, there's a lot of compute. Those two things are the bitter lesson for everyone doing research. You just need to scale things up. So data and compute really matter. I think more than data and compute—sure, those are baseline elements—is you need to have a team and a set of people that have tried a lot of things, have experimented a lot. There's accumulated knowledge in the way that you train these models, in the ways that you optimize them, and the ways that you deploy them, that it's just like we've tried 100 things and 90% of them haven't worked, but we've learned so much about not working.
Runway partnered with Getty. Why?
Cris Valenzuela [18:18] And so now I see a lot of companies who are trying the first things we tried five years ago. And we know where we're going to go, and we know they're going to fail because we were there before. And so data, compute, good ideas, but also just expertise and somehow experience have also been very important for us.
Matt Turck [18:36] And on the data front, you have a partnership with Getty. I mean, obviously the whole debate around what feeds generative AI, and is that image data that's crawled from the internet, like all the things. But you guys clearly chose to go with licensed datasets. Can you talk about the partnership?
Cris Valenzuela [19:03] Yeah, so the Getty Images partnership, if you're not aware, it's the partnership we've struck with Getty to allow companies, enterprises who might have some data lying around—so think about a film studio, a production team, advertising companies—that might not have hundreds of millions of videos or images, but maybe a few thousand. They can fine-tune the model. So you can take the existing Getty data and add your own data. And that basically gives you an edge because now you have a model that's able to consistently maintain the style or the guides or the artistic directions that you have, which for us has been very interesting to see.
Cris Valenzuela [19:33] We work mostly with media companies and studios in Hollywood, and these companies have a lot of data, like a lot, but they've never thought about it. Because why would you ever think about it? It's just some storage that some of them actually haven't even visualized. But if you do and you train a model, then you can open a bunch of new applications, not only on speeding up production cycles, but also offering perhaps new ways of interacting with some of the old IP or characters or styles.
How has the AI video generation ecosystem evolved?
Cris Valenzuela [19:52] And so the Getty partnership really taps into that potential, I would say.
Matt Turck [20:23] How do you see the whole ecosystem evolve? I loved how, when OpenAI released Sora, whenever that was, mid-February, you tweeted, "Game on," which I thought was very badass. But there seems to be a number of companies. Just today, there was, I think, one that was announced that just got some funding by two former DeepMind people. So it's a very vibrant and exciting space. Where do you fit, and how do you think different players are different from one another?
Cris Valenzuela [20:51] First of all, I think it's great that more companies are thinking about this. OpenAI, I mean, they're a great research team, a brilliant set of people, and it's great to have more people thinking about what's needed and the challenges and also the limitations and the opportunities. There's so many things that you have to do. And so getting a brilliant team like the OpenAI team to think about it together with us, that's great. I think part of it was there's an excitement that you're not running alone anymore because you have someone else trying to also do similar things.
Cris Valenzuela [21:21] I think it's also a realization that it's very early. Most of these things haven't yet been released. They're still in research phases. People are asking questions around scale and inference times and quality and control. And I want more people to think about those things. We can't be the only ones doing it. And so actually, for me, LLMs happened or had a similar momentum like two and a half years ago, where we started also building and seeing infrastructure being built around training and deploying language models.
Adoption cyсle for AI video generation. Where are we now?
Cris Valenzuela [21:58] And I really expect the same thing to happen with videos. I would love to just grab a data tool that can help us train models. If I don't have to build it myself, maybe there's a startup or a company that can help us do that. But to get to that, you need more companies trying to solve the problem. So I welcome more companies trying to do what we're doing.
Matt Turck [22:19] Where are we, from your perspective, in the adoption cycle for a technology like generative images, generative video? Who are your early users? You mentioned Madonna. I read A$AP Rocky and Slash, and celebrities like those. But are they hobbyists? Are they celebrities? Or are you already in the enterprise?
Cris Valenzuela [22:47] I measure adoption cycles in reactions of executives at Hollywood companies. So I pitched or spoke with some folks in a large production team five years ago. And this idea that you can generate images and videos was so far-fetched that it felt like, it's like, come on. It's like the metaverse. I don't know. It's some obscure thing. I talked with that same guy like a year and a half ago, and he was like, "Oh yeah, keep me posted. It's coming into—I saw one."
Cris Valenzuela [23:14] Yeah, it's kind of cool. And then I had, I'm not kidding, a chat with him last week, and he was like, "What should I do? It's freaking me out. I'm late." He actually left his current production team, where he's worked for 20 years, and he started a new company. He's like, "Should I not start that studio? Should I invest in—what should I do?" And I think that I'm seeing more of that. I'm seeing a lot of people and studios and companies realizing that it's just the tipping point of a major transformation for them.
Cris Valenzuela [23:50] Some of them are freaking out a little bit because their business is at risk if you don't catch up. So no one wants to be Blockbuster. You're seeing streaming for the first time, and you're like, better catch up. And so adoption for us is tied to that, I would say. You're starting to see more companies get a better understanding of how these tools work. Also realizing that a lot of their teams are already using it in some way. That's how we work with some of these companies.
Cris Valenzuela [24:21] We just see who else is using it. Sometimes they have dozens or hundreds of people, and then just reach out to them and sell them a bunch of licenses. And then on the creative side, most of the adoption is done by creatives themselves. These are people who are tinkering. So yeah, you're thinking Madonna and Pennywise made a video with Runway. Filmmakers have done it for movies, and that really helps drive awareness and adoption. But I still believe, even though all of those things, it's still pretty early.
Cris Valenzuela [24:27] Very early, yeah.
Matt Turck [24:39] And in VC speak, you have a bottoms-up kind of model, right? So anybody can show up, go to the website, and become a user of the product?
Challenges of building a research-focused company
Cris Valenzuela [24:45] Yes, it's free. You can just log in. You have a website and an iPhone app, so you can also use it on your phone as well.
Matt Turck [24:51] What are the unique challenges of building a research-focused kind of company?
Cris Valenzuela [25:22] Where should I start? There's so many. Research, I mean, it depends on what you want to do. I think for us, we always go back to the mission. I think research is sometimes too obsessed with novelty, with finding novel ideas that will get you published somewhere and introduce some complicated concept about some good idea. But sometimes you don't need those complicated sets of principles; you just need to be more pragmatic. And I think a lot of things we've learned are sometimes the best research actually comes from engineers who want to step in and try this thing called research because they come with a much more pragmatic view of how the world works.
Cris Valenzuela [25:59] And actually, I think I was speaking a lot with researchers these days, and a lot of them are going through an existential crisis of sorts because maybe a lot of things you were trying to solve and work on are just going to be solved by training larger models. And so it's a hard realization of, like, well, what should I work on next? And my answer is, well, there's a lot of things you can work on next, but maybe the way you phrase the problem is very different.
How to build and maintain a soul in a startup?
Cris Valenzuela [26:25] And so one of the challenges, I guess, is trying to keep that mindset in a heavy research culture, be very engineering-focused, and actually try to bring engineers to do research as well. I think you're getting a lot of value from actually trying that.
Matt Turck [26:41] Another thing you wrote, which I really liked, that I'd love you to expand upon: you said a lot of companies have a body but no soul. How do you build and maintain soul in a startup?
Cris Valenzuela [27:11] You've been reading my tweets. I tweeted that like six hours ago. Yeah, so I was coming from that from—I think that I've noticed a very interesting way of building products these days that feel very momentum-driven. There's a new technique, and there's a new model or a new modality or something, and there's a company that does that for X. And I think that's interesting as a concept, but how many of those are we going to have long-term? And I think that that's where someone was—an investor was asking me what companies I'm more excited about, early-stage companies.
Cris Valenzuela [27:44] And I'm mostly excited about companies that have a vision, companies that are not driven by some new framework or model or training or whatever. People who are like, "I want to solve that thing, and I want to do that thing really well." And how? Well, I don't know. I have all these possible approaches and what else? And I think we're coming from that. I think we're sometimes, as technologists and software engineers and researchers, too obsessed with technology, but then we get to a point of, "And now what?"
Cris Valenzuela [28:19] We built this thing, and I don't know for what, maybe for creatives. And I think that's like—I would think about it the other way around. And I think Runway kind of started as a company the other way around. It's like we knew what we wanted to do. We just had to start from the other side. So if you're starting a company today, look for a soul, I guess.
"It's like an invention of new art form" -
Matt Turck [28:28] Generative AI is going to change many things, and you had a call to action to entrepreneurs that I thought was particularly inspiring.
Cris Valenzuela [28:54] I mean, it riffs, I guess, based on what I mentioned just now, which feels like, for me, a completely new paradigm. It's like an invention of literally a new art form. We're interacting with computers and software in ways that we never dreamt of before. And sometimes I feel like the only thing we can think about is chatbots, which feels like, damn, we're missing so much. And I think that's convergence. Sometimes breaking ideas requires just breaking a little bit of comfort.
Cris Valenzuela [29:19] And actually, the best ideas I've seen were from seven years ago, where people were less risk-averse in trying new things and weird things. Because it wasn't—you were not trying to raise money for a company, or you weren't trying to build a product. You were just trying to define some interesting space. And so I guess my call to action was, like, if you're building something early and you're trying a model or experimenting with an idea, it's the opportunity for you to create a whole new market, create an entire thing.
Cris Valenzuela [29:48] And I think sometimes we limit ourselves too much by thinking about it incrementally. And so the invitation was, like, yeah, be weird, be strange, build stuff that no one has ever seen before. That, I think, is much more interesting.
Matt Turck [29:51] Cris, thank you so much. It was wonderful.
Cris Valenzuela [29:52] Thank you.