Assembly AI: Generative AI for Speech Recognition with CEO Dylan Fox
The MAD Podcast with Matt Turck · with Dylan Fox, Founder & CEO, AssemblyAI
Dylan Fox is the Founder & CEO at AssemblyAI. We cover why transcription creates more value as input to downstream features than as an endpoint, why AssemblyAI reserves tailored models for use cases with product-market fit, and why scaling labeled audio data still leaves speech recognition with roughly 15% error rates on many datasets.
Chapters
Transcript
Full episode
Matt Turck [0:55] All right, welcome, Dylan. So you are the CEO of AssemblyAI, which is an API platform for state-of-the-art AI models with at least an initial focus on speech recognition. I guess we're going to be talking about speech a lot. Yeah, the company is fully remote. You're actually here in New York?
Dylan Fox [1:00] Yeah, personally, but the company's all over.
Matt Turck [1:11] Yeah, which is wonderful. At least on Twitter, there's always a debate of San Francisco versus New York when it comes to AI and all those things. So I'm just using the opportunity to make the point that there's a lot of AI in New York.
Dylan Fox [1:13] I used to live in San Francisco.
Matt Turck [1:14] Okay, but you're here.
Dylan Fox [1:15] Yeah.
Matt Turck [1:42] And you've raised about $63 million in venture capital, most recently a $30 million Series B that was announced last year. So congrats on all of this. So you've had an exciting last 12 months, including announcing the Series B that I just mentioned, also releasing key enterprise features like Auto Chapters, and just also launching your own model. So I'd love to talk about all of this. And maybe starting with the core of what you do, what is speech-to-text, I guess, and speech recognition?
Dylan Fox [2:27] Yeah, yeah. So at the core of what we're doing right now is focusing on making these AI models that can transcribe and understand spoken audio data at scale. So we've processed almost 2 billion audio files through our system to date. There's something like 100 million-plus a month flowing through our system. And these are a mix of virtual meeting recordings, user-generated audio and video content, podcasts, contact center phone call recordings—it spans the gamut. And who we work with are product teams that are trying to build features and products on top of the audio data that's either being generated or flowing through their system.
Dylan Fox [3:01] And so we're training and creating these AI models that can do that really well and reliably at scale, with all the bells and whistles and features that product teams need to ship really quickly. I think I can stop there. I can go more than that.
Matt Turck [3:12] Yeah, yeah, no, no, that's great. So maybe to double-click on some of the things in terms of use cases. So there's audio content, there's video content, there's virtual meeting content.
Dylan Fox [3:13] Yeah.
Matt Turck [3:21] There is conversation intelligence. So just give us more color on what exactly people do.
Dylan Fox [3:52] Yeah, so there's a couple of examples I can give. So we've got over 1,000 customers, tens of thousands a month of developers that are building with the API. And this is everything from contact centers that are trying to provide insight into the massive amount of phone calls that are coming through their contact centers, that support agents and customers are having, so they can be analyzed at scale. This is video editing platforms adding subtitles to videos, basic use cases like that.
Dylan Fox [4:39] We have hiring intelligence platforms that are recording interviews. Maybe you guys have experienced that, and providing automatic notes, follow-ups, action items, all based on the transcription. Really what companies are doing, there are some use cases that product teams are leveraging our models for that are just like, you're automatically converting audio into text and you're displaying that as subtitles, or you're displaying the transcript for readability or accessibility. But where a lot of value is created is you're taking the transcription and then you're using it as an input to do something else.
Dylan Fox [5:13] Like, I think of a contact center platform that's building automatic text message follow-ups that, when you go into your UI as a customer, you can just click, like, send, send, and they're all automatically generated and customized based on the transcription of a voicemail or of a phone call recording that happened. So it's transforming applications that are being built on top of the text, on top of the—like, you're turning the audio and video into a more pliable format that you can build with.
Matt Turck [5:40] Yeah, and as I was prepping for this, I noted a whole list of really interesting features: summarization, sentiment analysis, entity detection, topic detection, content moderation.
Dylan Fox [5:41] Yeah.
Matt Turck [5:45] Maybe pick one or two of those that you think are particularly helpful.
Dylan Fox [6:18] Yeah, so I think more broadly what we're seeing is the company's been around for a while, but there's just this insane amount of demand right now, and product teams are really trying to figure out how to leverage all this AI tech to build new features, to tell a story to their customers and to the market that they're forward-thinking and they're leveraging AI in their products. And there's a lot of exploration that's happening. And I think similar to the comment that was just made, product-market fit is a question. Like, are summaries of a virtual meeting helpful?
Dylan Fox [6:52] Is a transcript of a virtual meeting helpful? Is an automatic text message follow-up of a phone call recording helpful? There's a lot of exploration that has to happen and there's a lot of iteration that has to happen. And so our value proposition—this comes back to your question about the models that we offer—I think it's more about what we're trying to do is help product teams iterate very quickly to figure out what's going to have product-market fit and what's going to have commercial success for them, create a lot of value for their end users, for their customers, and help them win.
Dylan Fox [7:43] And so what we think about is we're trying to help product teams execute 10 times faster and ship 10 times faster, so they can figure out where they need to focus. And so, to your question, where our customers have found product-market fit, we then lean in and support those use cases with really tailored, expert models. And so, for example, we have summarization models that are really good for two-party conversations in contact centers. So if you have a two-sided phone call, support agent and a customer, we have summarization models that can do those really well.
Dylan Fox [8:21] We have a lot of video editing and video hosting platforms and podcast hosting platforms that use our models. And we can create summaries that are really good for automatic chaptering of spoken content, so really catchy titles, really good descriptions of what happened within this time-coded segment. Where there's product-market fit, we focus on creating tailored models for those use cases. Then we also have an LLM that is more tailored for conversational data that is not an expert, but our customers can use and product teams can use to explore and see what even works well.
Dylan Fox [8:58] And so we try to focus across the board. The reason why we spend so much time on our automatic speech recognition models is because there's a lot of product-market fit there. And so it makes a lot of sense to be incredible at that and really help our customers excel there and continue to deliver value there. Some of these other things that are more exploratory, we're not going to put a ton of effort into those until we see that our customers, the developers that use our API, are actually able to find value and create value with those.
Dylan Fox [9:45] So, back to your original question. Today, we have a number of different models around PII detection and redaction, entity detection, sentiment analysis, automatic chaptering of content, summarization of content. We can detect sensitive content that's spoken, so hate speech, sensitive content that trust and safety teams are using to automate content moderation at scale when there's spoken audio in a platform. So it really runs the gamut, but it's really all around transcribing and understanding audio with AI models and making that available to product teams and developers through our API that's really easy to work with.
Matt Turck [10:23] Yeah. And how do you think of the balance between building your own models and leveraging open source? The fundamental value proposition is that, hey, you want to solve this problem, and AssemblyAI will bring the best state-of-the-art model to the task. You don't have to worry about it. Or is the value proposition like, we will build the best model?
Dylan Fox [10:43] So this is where there's so many—you asked a question in the last chat—there's so many market maps of AI right now. And I think it's really—the picture is fuzzy. And I think people are trying to make sense of it, which is why they're making all these maps. But the picture is still fuzzy. And I was prepping for this. I think you had some tweet where you were like, this is what an AI company is, and there's, like, shades of gray in between.
Dylan Fox [10:52] Not to put you on the spot.
Matt Turck [10:56] I will require going forward, like, every speaker needs to read my tweets before coming here.
Dylan Fox [11:24] Yeah, exactly. So what we'd like to go out and tell customers that we talk to, and our stance is like, we're not a research lab. We're not trying to just work on AI research and develop secret molecules. We're really focused on helping create commercial success stories, leveraging AI models. So that's why when I talk about our customers, our customers are product teams. And yeah, there's a ton of developers that use our API to play around, to tinker, and ultimately build startups, build projects that turn into actual businesses.
Dylan Fox [12:12] And then, like, the people that we end up working with there are the product managers and the product teams because we're helping them ship products and ship features. And so, to your question, sometimes we will train models from scratch when we need to make something better than what's available. Other times, we'll take something open source and we'll fine-tune it with a dataset that we accumulate, or we'll make modifications to it to make it faster, more performant, more scalable. It really depends. And so, we try to be very transparent about that.
Dylan Fox [12:23] So we shipped this Conformer-1 model a couple of weeks ago. It's this large-scale—go ahead, sorry.
Matt Turck [12:28] Yeah, yeah, no, I'd love to spend a good amount of time on Conformer-1.
Dylan Fox [12:29] Yeah. Yes.
Matt Turck [12:34] So, which is your own LLM? So a lot of the—
Dylan Fox [12:36] It's our own automatic speech recognition model.
Matt Turck [12:48] Yes. Yeah. A lot of the models that you talked about so far, and correct me if I'm wrong, were pre-LLM wave, all the way back to, like, early 2022.
Dylan Fox [12:49] Yeah, yeah, yeah.
Matt Turck [12:53] So, but now you're, in addition to this, you're adding your own LLM.
Dylan Fox [13:27] Correct. Correct. So we think about our product, we have kind of these three categories. So we have our—if you go to our product page or pricing page—there's the core transcription. So those are our automatic speech recognition models: audio in, text out, asynchronously or real-time streaming over WebSockets. And then we've got what we call these audio intelligence models. And so this is what I was referring to earlier: these task-specific, really streamlined models. So PII detection and redaction, sentiment analysis, tailored summarization models for specific types of summaries.
Dylan Fox [14:08] These are lightweight, so they're cheap to use, they're cheap to host and run. They're not these giant models because they're just focused on specific tasks where our customers are finding product-market fit, leveraging them within their applications and their products that they're building. And then where there's more exploration happening, we leverage large language models that can be applied at many different tasks with varying degrees of quality, but they're flexible. But back to the Conformer-1 model, with that model, for example, we really wanted to push forward accuracy with speech recognition models.
Dylan Fox [14:31] Particularly to make them more robust, have lower variance, just be overall better. And so we took this neural network architecture called the Conformer, published by Google Brain, I think in, like, 2021.
Matt Turck [14:36] Which was a speech-to-text version of Transformers, right? Is that correct?
Dylan Fox [14:41] Yeah, it's a transformer-based neural net for automatic speech recognition.
Matt Turck [14:47] And they published Transformer in 2017, then Conformer in 20—'18, you just said?
Dylan Fox [14:50] Yeah, like, '19, early 2021, I think.
Matt Turck [14:51] Late '21.
Dylan Fox [15:18] Yeah, yeah, late 2020. I'm not exactly—it's around there. And we made some modifications to that neural net. Like, that was just a paper. Like, the weights—there's no model release, there's just a paper. But we did a literature review and chose that neural net for some specific reasons. We made some modifications to it. And then we just tried to scale it up. So we trained it on, like, 60 terabytes of audio data, like labeled audio data.
Dylan Fox [15:55] So it was, I think, something like 650,000 hours of audio data. And our models prior, and most commercial speech recognition models, trained on, like, 50,000 hours. So this is, like, an order of magnitude larger. But whatever this successor will be is training right now, and that's something around 4 million hours of labeled audio data. So we're just scaling it up pretty aggressively to increase the robustness and accuracy of the models, because that's our core, primary value proposition right now, is our automatic speech recognition models.
Dylan Fox [16:30] And it's not just the model. So we provide a lot of features around it. So if you go to our API, you can hit an API endpoint to get, like, all the sentences split out. You can get the text broken out into paragraphs. You can quickly redact the text. You can fan out audio files to process a ton in parallel. You can real-time stream. You can get speakers annotated and labeled. Get really precise word timings and confidence scores.
Dylan Fox [16:46] There's, like, dozens of features that we provide around the model to make it really easy to build with and work with. And so that's really our flagship model that we offer today. And then the other models that we offer around summarization, those are really for customers that are trying to work with a single partner to help them just ship really quickly so they can explore and figure out how they're gonna leverage this AI tech to build new features, expand their customer base, grow their revenue, like have a success story that has product-market fit.
Matt Turck [17:29] Yep. And out of curiosity, for Conformer-1, how did you get the training data? Is that internet data or is that customer data, presumably with all the privacy and safety?
Dylan Fox [18:03] Yeah, it's a mix. So it's a mix of data that is from the internet. It's a mix of data that's been shared with us from our customers. So some customers don't care and prefer that we train on their data. So for the ones that do, like I said, we've processed—yeah, it's like over 100 million audio files a month that are flowing through the API, and that's growing pretty quickly. And then there's a lot of open-source datasets. We kind of just group it all together and use a combination of that.
Matt Turck [18:23] Yeah, it's really interesting. I'm a big fan of the concept of data network effect and always interested in examples of companies that work collaboratively with customers to pool data to help the AI get better.
Dylan Fox [18:57] And it's, yeah, it's like what we see with Conformer-1, for example, Conformer-2 that will launch, and there might be, like, a series of launches before we switch to a different neural network architecture. You can throw more data at the model and you get better robustness, and it's, like, generally more accurate. But proper nouns, for example, those are really important for applications that are built on top of automatic speech recognition models. And the percentage of proper nouns in datasets is actually low distribution. And so you might see a big reduction in just overall error rates, but you might not have really moved the needle much in email addresses or phone numbers or proper nouns.
Dylan Fox [19:38] So we also look to focus on making—because we're focused on—because we're not a research lab, because we're really focused on shipping models that product teams can just quickly go with and build with and not have to deal with a lot of the headache. And there are definitely some companies that want to build everything themselves and figure it out, and that's fine. But our opinion is that the tech is turning over so fast. So if you're a company and you're in the contact center space and you want to try to ship summarization in your contact center so that when your customer support managers are logging in and reviewing calls, they don't need to listen.
Dylan Fox [20:18] To every single call. They can just, boom, quickly go through some two-, three-sentence summaries or automatically have some calls flagged that are potentially problematic. If you take six to 12 months to ship that and then there's no product-market fit, that's a huge waste of time and resources. And so our opinion is: partner with a company like us. You can get that out in two months, three months—it depends how slow they are.
Dylan Fox [20:50] You can get that out quickly and see if there's product-market fit with that and if that's something you should even consider building in-house and bringing in-house. But right now, companies have to iterate quickly if they're going to stay competitive. And if they try to do everything themselves and build everything themselves, they're not going to be able to go quickly enough. There's also just a big delta in being able to train a model on a single GPU and actually fan out the experiments you need to run across clusters of GPUs and try all these different types of hyperparameters and groupings of data.
Dylan Fox [21:26] That's a whole different level. And so that's really what you have to commit to if you're going to start to take this stuff in-house, because none of the tech is done yet, right? State-of-the-art automatic speech recognition still has a 15% error rate on a lot of datasets. And so if you try to build all this in-house, what are you going to do as the tech continues to improve around you and you don't have the capability to keep up?
Dylan Fox [22:04] So that's what we try to offer as a value proposition. We'll be the expert, we'll deliver all that to you so that you can just keep shipping. And so that's why we focus on not just—how do we make these—we don't just focus on vanity metrics, but we focus on things like, okay, we want our models to be really good at email addresses, right? They're right now not, right? So in the future versions, they'll be better at email addresses and domain names and the long tail of things that are spoken.
Matt Turck [22:48] Yeah, I really like the positioning that you've explained very well, which is to deliver to product people as opposed to developers. For founders and investors like me, that's really that question of how you build companies in the space that have a sustainable competitive advantage. And so presumably there's some competitive pressure from the cloud vendors, who all have some kind of a speech-to-text product, but selling to developers. Would it be fair to say you're sort of the last mile that's focused on the application layer and the no-code aspect of this while—
Dylan Fox [23:23] So we do—I mean, it is an API, right? So the developer has to do the integration, but I think, like Datadog, right? They're selling to a VP of Engineering, right? You need to implement application performance monitoring. Likewise, you're probably selling to a VP of Engineering. The people that sign our contracts are VPs of Product. It's a different customer profile, but the developers are involved, right?
Dylan Fox [23:58] And especially in the long tail and startups, the developer is the product person. They're the founder. Back when I started AssemblyAI, I was the developer and the CEO, and you're just doing everything. And so we focus on really good—we try to focus on really good developer experience. It's easy to just get up and running. I actually don't—yeah, there's a ton of room for improvement that we still have to make, but we're seeing tens of thousands of developers register to the API every month.
Dylan Fox [24:35] But to your point, because we're really focused on—we're really customer-focused, we're building the features into the model, around the model, into the API that just make it easier to ship with and build with. And I think that customers feel that, developers feel that, and that's ultimately why they choose us. In a way that I just don't think—I'm surprised the big cloud companies, and apologies if anyone here works there, I'm surprised they can't ship better developer products.
Dylan Fox [24:52] But I think they're maybe just too big at this point. Yeah.
Matt Turck [25:11] Great. So maybe as a last question from me, until I open up to folks, I'd love to take a big step back and just talk about what you find exciting in the space, whether that's products or projects or companies or what—
Dylan Fox [25:15] Did you see the AI Drake song that was just released?
Matt Turck [25:16] Yeah, Drake and The Weeknd.
Dylan Fox [25:40] Yeah, yeah, yeah, that was exciting. I like—so my wife was using ChatGPT over the last week to build a Chrome extension. She's not a developer. Built a Chrome extension. It's published on the Chrome Web Store. Is now building a whole Python app with Flask. It's crazy. So I think it's just a really exciting space to be in, and it's a really exciting time.
Dylan Fox [25:53] Regardless of where it goes.
Matt Turck [25:53] Yeah.
Dylan Fox [26:16] To your question of how soon are we going to have AGI, right? I think there's a lot to still figure out. Regardless, I think there's just a lot of really cool things happening. I mean, the fact that you can create images like that, the fact that you—if you go to our website and you click, there's a link on the top, it says Playground. You can just throw in a YouTube link.
Dylan Fox [26:34] You can drop that in and then pretty quickly get a very accurate transcription and summarization of that. There's just a lot of really exciting developments happening. And I think for me, that's just—it's a cool space to be in and to just be in that. Yeah.
Matt Turck [26:42] And by the way, the website is particularly good. There's lots of great content there. You guys seem to be doing content very, very well. Wonderful. Thank you. Thanks.
Dylan Fox [26:52] Yeah, appreciate it. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. firstmark.com/events/data-driven.