Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
The MAD Podcast with Matt Turck · with Lin Qiao, CEO, Fireworks AI
Lin Qiao is the CEO at Fireworks AI. We cover why GenAI lets companies start building without an ML team, how Fireworks searches more than 80,000 configurations across quality, speed, and cost, and why AI infrastructure costs should fall by an order of magnitude to prevent startups from scaling to bankruptcy.
Chapters
- 1:20 — What is Fireworks AI?
- 2:47 — What is PyTorch?
- 12:50 — Traditional ML vs GenAI
- 14:54 — AI’s enterprise transformation
- 16:16 — From Meta to Fireworks
- 19:39 — Simplifying AI infrastructure
- 20:41 — How Fireworks clients use GenAI
- 22:02 — How many models are powered by Fireworks
- 30:09 — LLM partitioning
- 34:43 — Real-time vs pre-set search
- 36:56 — Reinforcement learning
- 38:56 — Function calling
- 44:23 — Low-level architecture overview
- 45:47 — Cloud GPUs & hardware support
- 47:16 — VPC vs on-prem vs local deployment
- 49:50 — Decreasing inference costs and its business implications
- 52:46 — Fireworks roadmap
- 55:03 — AI future predictions
Transcript
What is Fireworks AI?
Matt Turck [1:19] Lin, welcome.
Lin Qiao [1:20] Hi, thanks for having me here.
Matt Turck [1:28] We are going to talk about all things Fireworks AI today. But as a level set, what is the elevator pitch for Fireworks AI?
Lin Qiao [2:06] Fireworks AI provides a developer platform for application developers to build on top of GenAI technology. They need to hypothesize what kind of product is interesting, and they need to use certain models to integrate and build these great, innovative user experiences. Oftentimes, they have to solve multiple problems. One is quality: is the model delivering the quality they want? And also interactive, real-time speed, because many of those products are consumer- and customer-facing. And then, when they hit product-market fit and scale the business, they have to have a sustainable way to scale.
Lin Qiao [2:39] So cost efficiency is very important. Oftentimes, we see application developers heavily focused on three-dimensional optimization: quality, speed, and cost. And we want to take away all this complexity from them so they focus on thinking about what is the best user experience, what is the best product idea.
Matt Turck [2:45] So, infrastructure as a service: you abstract away the whole complexity of what needs to happen behind the scenes to deploy those models.
What is PyTorch?
Lin Qiao [2:47] Exactly. Okay, great.
Matt Turck [3:08] Maybe a few words on your story, your background. In particular, you were one of the key people on PyTorch at Meta, or Facebook at the time. Yeah, Facebook at the time, I guess. What is PyTorch, for anybody that doesn't follow these things in great detail? And then how did it all come about? What was your role there?
Lin Qiao [3:41] Yeah, that was an amazing journey. When PyTorch happened, what was the kind of industry context here? So I joined Meta in 2015. At that time, Meta was called Facebook. Facebook was going through the transition of finishing mobile-first to starting AI-first. The fundamental reason why it sequenced that way: after Meta transitioned its application from desktop to mobile-first, and it's widely accessible everywhere using your phone, you can connect with people. And that drives a lot of user engagement and, therefore, a lot of data being created.
Lin Qiao [4:15] And then we all know data fuels AI. So quickly after the mobile-first transition, we started AI-first. At that time, there was no software, no hardware, there were no GPUs, no people building all this infrastructure. So we bootstrapped the whole thing from the ground up. And it was an interesting time. It feels like right now there's so many companies building their own AI framework, a gazillion number of AI frameworks. Even within Meta, there were three different flavors: one for mobile, one for research, one for production.
Lin Qiao [4:54] So it's very confusing. And we decided to unify all of that and have one framework aiming to be the best tool for researchers to create innovative model architectures, but also one framework to deploy and deliver those innovative models into production to power all the product needs. So these two actually have huge tension across them because, for researchers, you want flexibility, you want ease of use, you want them to just think about what's possible, right? And for production, it's a constraint problem-solving, as in you have latency budget, you have cost budget, you want to scale, you want to be reliable.
Lin Qiao [5:34] So it's kind of, hey, you have all these constraints, let's kind of optimize. It's an optimization problem. So how to create one framework to solve both problems with big tension in between is a very, very hard challenge. I would say we had a lot of fun thinking about that. It was deemed as mission impossible at the beginning, and many people felt this was so experimental, probably about to fail, but it's worth a try. And I would say our initial attempt was not successful.
Matt Turck [5:41] When was that, 2016 or...?
Lin Qiao [6:05] From 2015 to 2016. And because the idea is great. The idea was simple because PyTorch has a great frontend, very simple, Pythonic. And we had another framework internally. It had high-performance kernels backend. And you would think we can zip them together so we get the benefit of both.
Matt Turck [6:06] How hard can it be?
Lin Qiao [6:41] How hard can that be, right? So let's just zip them together. Internally, it was actually called Zipper. And it turns out it's really hard to bring two frameworks completely designed with different goals and different interfaces and internal APIs to kind of work seamlessly together. It's like building a bridge with both ends moving. It's going to be an extremely fragile bridge. So then we decided, hey, making both sides happy is not the goal. We need to build an extremely strong product.
Lin Qiao [7:14] And we decided we were going to rebuild the whole entire backend from this beautiful frontend and bring it to production. It took us five years. It took us five years to get to the stage supporting almost all internal needs using deep learning at massive, massive scale. So it also has been deployed to not just data centers—Meta owns its own data center—but also mobile phones. If you use any of Meta's apps, that means our PyTorch infrastructure is running on your mobile phone and also AR/VR devices.
Lin Qiao [7:57] It's kind of ubiquitous everywhere. At that time, PyTorch externally, as an open-source project, had grown from a baby GitHub project into more and more massive adoption. So through that, we got exposed to many other companies using PyTorch. And it's fascinating to see the growth of all different use case adoption, from image recognition to ranking recommendation to robotics, self-driving cars. And GenAI is nascent, but it's literally everywhere.
Matt Turck [8:01] And it's used by all the research labs, right? Like all the—
Lin Qiao [8:08] It started from research labs. And the fun fact is OpenAI switched to use PyTorch fully.
Matt Turck [8:09] From what?
Lin Qiao [8:14] From TensorFlow. Right. So TensorFlow was dominating at that time.
Matt Turck [8:16] TensorFlow being the Google framework.
Lin Qiao [8:40] Google framework, yes. Google also put a lot of effort behind open-sourcing this effort, growing the community. But simplicity wins. PyTorch is so simple, and the user interface, the debuggability, and people can easily change the model architecture as they want. It's dynamic and flexible.
Matt Turck [9:06] To drive home the simplicity point: so it does, because you had the sort of low-level and high-level stuff, so it does anything from whatever the research scientists need to do in terms of, I don't know, like a backprop, whatever, on one hand, but on the other hand, it's going to do GPU usage optimization. Is that fair? So it spans all the things that you need to do to develop faster and simpler.
Lin Qiao [9:38] Right. So I think the API kind of adopts Python as the programming language. That's why it's called PyTorch. The idea is from Torch, LuaTorch time. And Python is a very simple, easy-to-use programming language for many researchers. And they think about deep learning neural networks in code, right? You need to think about that as graphs, as nodes and edges and how to tie them together. When this graph becomes like hundreds of thousands of nodes, then it's how do you even kind of manage that thinking process?
Lin Qiao [10:21] But representing that in code, in Python code, is very easy to manage. And then you become a software engineer to think about a neural network. And it's very easy to debug because all Python tooling, the toolchains, can all work. So that actually unblocks a lot of progress made in modeling innovation. And it also, of course, can run on GPU, it can run on different hardware SKUs. We have many different hardware vendors plug in at a lower level to support PyTorch, and that includes TPUs.
Lin Qiao [10:42] So, and because Google has a compiler that can directly connect with PyTorch and lower the program into kind of the GPU runtime.
Matt Turck [10:43] Yeah.
Lin Qiao [11:17] So that enabled significant advancement from the model research side. And the interesting part of model research is a lot of research has been published as open-source projects. And Hugging Face is one example of adopting and becoming the centralized repo for new research ideas. And it's very easy for people to take this model architecture. So there's a pre-GenAI and post-GenAI era. Pre-GenAI, the model code is open source, but there's no weights, right? Everyone has to curate their own data and train from scratch.
Lin Qiao [11:48] That just means any company that wants to invest in deep learning, they have to first hire people who understand how to curate data first, and then who understand how to train, how to manage GPU fleet, and so on. So that takes a lot of time to hire those people, which come from a very, very small pool, and takes a lot of capital investment to get this going. So that's why pre-GenAI, not all companies can access this technology, even though it's great.
Lin Qiao [12:27] And usually it concentrates in hyperscalers or whoever has big resources to be able to do it. And then post-GenAI, GenAI is basically built on top of foundation models which can learn from world knowledge, internet knowledge. And these models are good by themselves. So applications can develop directly, run on top of those models. So you don't need a machine learning team to begin with to start imagining what are the new user experiences. Or you can have a small machine learning team who just do fine-tuning or reinforcement learning, which requires much smaller samples to curate, and the post-training process is much simpler.
Traditional ML vs GenAI
Lin Qiao [12:50] So I would say that shift is fundamental. It's monumental to unblock a lot of new innovations in the application space.
Matt Turck [13:12] Yeah, just to play it back, it's something that's super obvious to people like you at the forefront of the field, but that a lot of people don't quite realize, this fundamental difference. And just to play back exactly what you said, between traditional machine learning and generative AI, it's not just a different technology. It's exactly what you said: you had to build your models from scratch before, more or less, and now you have this concept of foundation models that you can build on top of.
Matt Turck [13:30] Therefore, the whole complexity has moved to how you deploy the models, which was going to lead us into Fireworks. But just to play it back, because it's such an important concept that people may or may not have internalized, how that changes everything.
Lin Qiao [14:10] That's exactly right. So that's why we see an explosion of adoption of this technology from the inception of our company. And there are AI-native startups. There's no box around what's possible, and they create completely new user experiences that we would never imagine. But also the enterprises, whether it's digital-native or traditional enterprise, they are all joining this transformation and rethinking their existing product experience, or rethinking how to power their internal productivity or customer service, many different surface areas. So I would say, across the board, what's most surprising to me from the beginning, when I started this company, I had this sequential go-to-market motion in my head.
AI’s enterprise transformation
Lin Qiao [14:55] is to go to AI-native startups first because they are the most technically advanced, and then digital natives because they're more open-minded about embracing new technology, and then traditional enterprises because they want to see enough proof points before entering. Now it's happening across the board, everywhere. That's partly, I would attribute to, GenAI models being so easy to use and so accessible.
Matt Turck [15:23] And just to close on PyTorch, and then we'll go into Fireworks in detail, as a passing thought. And that sort of echoes a conversation I was just having most recently with Douwe from Contextual AI on this podcast, who was probably at FAIR around the same time you were. And we were just talking about how incredible an impact FAIR, Facebook, Meta have had on the AI ecosystem. So PyTorch being another fantastic example of just open-sourcing all of this fundamental framework that everybody uses and now is a fully independent project.
Matt Turck [15:48] Right. I think the governance was transitioned completely out of Meta at the end of 2022 or something. It's just incredible, the impact that all of you had on this ecosystem.
Lin Qiao [16:10] Yeah, it transitioned to the PyTorch Foundation, expanded to industry governance. But I believe Meta still has a very representative presence there to continue fostering PyTorch growth in the future.
Matt Turck [16:13] So, transitioned, but not completely transitioned.
From Meta to Fireworks
Lin Qiao [16:17] Yeah, I think it's working. Yeah, it is a progressive transition.
Matt Turck [16:25] How did your experience at Meta and then PyTorch influence the idea for Fireworks in the first place?
Lin Qiao [16:55] I think one of the biggest successes we saw from the PyTorch experience is simplicity scales. And this simplicity is user experience simplicity. Because we have a very strong engineering team, we can solve any infrastructure challenges and complexities. But when it comes to adoption, people don't want to spend time figuring out how to make things work. They just want it to work. And that's kind of where PyTorch shines. And that's why, in the heavy competition or very noisy market where every company was building their own AI framework, I think there were like tens of those AI frameworks at that time, none of those were focusing on the simplicity part.
Lin Qiao [17:43] Everyone was focused on production and all the nitty-gritty details to get into production. But to the point, it is very hard for researchers to adopt. And a kind of interesting funnel we saw from the PyTorch journey is researcher adoption is the top of the funnel because they create new model architectures which get open-sourced, and they get adopted by industry practitioners for experimentation. Maybe one out of 10 of those experiments will be successful, and then it goes down to the next phase of the funnel of full production.
Lin Qiao [18:27] And once it goes to production, people start to transition from PyTorch into other, more production-focused optimized frameworks. But that transition is too hectic. They're going to rewrite the whole model in a completely different message, and then you lose precision. It's too much hassle, and it's much easier to just use PyTorch into production and let them optimize PyTorch. So we saw this journey. Now, retrospectively, we saw how this evolved. And at the time, it was not clear who was going to win and how it was going to win.
Lin Qiao [19:06] But now it's very clear: leverage simplicity to get to adoption, because that's what the end user cares about. And then we focus on optimization behind the scenes. And that's the way to drive impact. So we carry the same mindset when we think about Fireworks. We want to keep the Fireworks user interaction really, really, really simple, and we take the heavy lifting of production optimization. And production optimization is so complicated. And I would like to elaborate on that.
Lin Qiao [19:16] And so that's kind of the founding principle of designing Fireworks' product: simplicity first.
Matt Turck [19:22] Because presumably at Meta, to manage PyTorch, you had, what, hundreds of engineers?
Simplifying AI infrastructure
Lin Qiao [19:39] We have hundreds of engineers building PyTorch and infrastructure around PyTorch. But at the same time, I believe PyTorch within Meta probably has thousands of users. Outside of Meta, it will be hundreds of thousands of developers.
Matt Turck [20:00] But in terms of the infrastructure required to just manage something like this, if you were one of your customers, like with Fireworks, you would need to hire dozens of folks or hundreds of folks maybe to just maintain and manage models, right? Is that fair in terms of what you abstract away?
Lin Qiao [20:31] First of all, hiring GenAI infrastructure engineers is a very small pool, and the optimization we do is very deep. You probably wouldn't be able to get anybody with our depth. But also, we manage huge fleets. The scalability is very fast. We saw, for example, AI-native startups, they can ramp up extremely quickly without even being able to predict what's going to come next.
Matt Turck [20:31] Yeah.
How Fireworks clients use GenAI
Lin Qiao [20:41] We have seen digital-native companies transition their online traffic to run on top of GenAI models extremely quickly.
Matt Turck [20:48] And famously, one of your customers is Cursor, right? Which is like the poster child example of super-fast growth.
Lin Qiao [21:13] Exactly. And from what I know, everyone I talk to uses Cursor. It's very interesting. So today I talked to one of our customers. They even use Cursor to schedule meetings. Very, very creative. Another example is that we work very closely with DoorDash and Uber and Samsung. So they are incorporating GenAI into their main products because they are marketplaces, and search relevance is extremely important. In order for them to accurately target and give recommendations to their end users, the consumers, they have to have high-quality data and high-quality supplier information, product information.
How many models are powered by Fireworks
Lin Qiao [22:03] All of those were manually corrected. And GenAI actually can have a model which can see, extract information, and enrich through LLMs. So we are renovating their whole entire internal setup to have a high-quality end-user experience.
Matt Turck [22:20] Let's get into the weeds of the product itself. So, starting with models, you sort of offer, power hundreds of models. What is the latest number of models that you offer?
Lin Qiao [22:54] Yeah, so our product has multiple layers. At the lowest layer, we have been working on a distributed inference engine from the beginning, and that's basically a customized PyTorch inference infrastructure we set up for GenAI specifically. And that distributed inference infrastructure was designed very composably and building-block-based. The fundamental reason is we want to launch any state-of-the-art model very quickly. As we just discussed, we have the track record of a new open model getting released, and the same day we launch it, right?
Lin Qiao [23:31] So that's because we kind of took the PyTorch idea and built our inference infrastructure in a composable way. And then we can pick and choose what are the best options and setup for that particular model. But the search space is very big. So we think about, first of all, our design principle is we don't believe in one size fits all. We don't believe in one size fits all for quality, as in we don't believe one model will solve all the problems in GenAI in the best way.
Lin Qiao [24:14] That setup doesn't exist. We also don't believe there's one inference configuration that's best for all different use cases. Instead, we believe in one-size-fits-one. We believe in a platform which can automatically do customization across three dimensions. As I mentioned, application developers, they usually are concerned about these three dimensions. One is quality, one is speed, one is cost. And it's a three-dimensional space, a three-dimensional curve, and different products, different use cases, they want to land in different spots and they want to make different trade-offs.
Lin Qiao [24:46] It's all use-case-specific. So what we want to provide to our end user is to give them the option, find the best option based on their requirements, specific target, and their workload, and give them the optimal point.
Matt Turck [24:52] And do you do that through consulting? Do you do that automatically? How does that work?
Lin Qiao [25:27] So we have a product called FireOptimizer. It's a platform, it's a SaaS platform, and it takes the inference workload for a particular use case and the objective function across quality, speed, and cost—where you want to land. It will do its own magic and spit out a new model plus a new inference deployment configuration and go into the rest of the CI/CD infrastructure.
Matt Turck [25:34] A new model, or it would pick one of the hundreds of models that you have on the platform?
Lin Qiao [25:51] So right now it's given a model, and there are different ways to improve quality or make quality-speed-cost trade-offs. For example, we have five different ways to quantize. We can quantize different parts of our model.
Matt Turck [25:53] Quantizing means making models smaller.
Lin Qiao [26:15] Making a model run in lower precision, where it's much more computationally efficient or memory-bandwidth-efficient or communication-efficient. But you can run, instead of 16-bit floating-point arithmetic, you can do 8-bit or even 4-bit. But you can apply these to different parts of the model. You can apply to weights, you can apply to activation, you can apply to KV cache, you can apply to the communication part, you can apply to different parts, and it will have different impact to the end sensitivity of the quality change.
Lin Qiao [27:02] At the end, sensitivity of quality change—who is the judge? The product is the judge, right? So the end user is the judge. So we would like to kind of fit the best into what the end user wants. So that's one example, and we can do runtime quantization, as in you still train, pre-train, or post-train to the full precision, and then we can quantize on the fly during runtime. Or we give you tools to do quantization-aware training where you decide the precision during the pre-training or post-training time.
Lin Qiao [27:43] And then you generate much higher-quality scalability that way. So there are different levers we give our customers to tune. But also, for example, we have three to four different ways to do speculative execution. The idea here is to conceptually think about speculation as, you have a way to predict not one token at a time; you have a way to predict three or four tokens ahead of time in one shot. And that's more efficient than predicting one at a time, right?
Lin Qiao [28:19] So because you can predict four at a time very accurately, then it's four times more efficient. And the physical mechanism of doing that, there are so many different ways. We support many different approaches. One approach is to pair a small model with a large model, right? And ask the small model the question first and have it do the speculation. And if it's correct, then it's great. Then you get the small model speed. If it's wrong, you always ask the large model to judge.
Lin Qiao [28:58] If it's wrong, then you ask the large model to give you the answer, right? So basically, the end result is you get the large model quality and the small model speed in combination. So that's a very good example just to talk about FireOptimizer, because in order for this to work, the small model and large model need to align towards your use case and workload pattern. If it doesn't align, then the small model predicts something, it's always off. Let's say, extreme case, it's 100% off, then you get the worst.
Lin Qiao [29:37] Situation where you always fall back to a large model. You have the overhead of running both models, but you have the speed of the large model, quality of a large model. There's no benefit. So this alignment process is basically to make the small model always predict higher and higher with higher and higher probability. We have seen cases improving the prediction hit rate from 30% to 90%. And that's huge speed saving. So yeah, those are kind of another examples of options.
Matt Turck [29:48] And do people on the customer side need to have the level of sophistication to understand those various options?
LLM partitioning
Lin Qiao [30:09] The idea here, so the number of options is big, right? So I just talked about two different areas, but we also have five different ways to partition the model and six, seven different ways to allocate different parts of the model to different hosts in a distributed way.
Matt Turck [30:16] Let's cover some of this quickly. This is fascinating. So what does that mean, different ways of partitioning the model, you said?
Lin Qiao [30:56] Right. Because we look at the model in—let's take a large language model, for example. Some parts of the model are doing prompt processing, and some part of the model is predicting what is the possibility of the next token. And during runtime, they are actually bottlenecked by different things. I'm talking about that in an extremely simplified way: prompt processing is bottlenecked by computation, and generating the next, predicting the next token is bottlenecked by memory bandwidth, right? So if you think about the model in one chunk and scale them all together, then you always get stuck in one bottleneck versus the other.
Lin Qiao [31:45] So we've seen similar kinds of infrastructure problems in other domains, like databases, and the way to solve that problem is to pull those parts that are bottlenecked by different system resources to scale them completely differently. So if it's compute-bound, then we should just add more GPUs to it, right? If it's memory-bottleneck-bound, then there's a different way to scale it. If it's computation-bound, then we make computation more efficient. So by scaling those different parts of model execution independently, then we remove all possible bottlenecks.
Lin Qiao [32:28] And that's the most efficient way to run those models. So that part, I think we are on the pretty cutting edge. I think recently we started to see open source, like people paying attention to those kinds of setups. But at the end, it's a complex system problem. There are libraries that get open-sourced to cover parts of the execution. But putting these all together and solving a lot of real system problems is the most difficult. I'll give you one example.
Lin Qiao [33:06] Before I give you an example, actually, I want to close the complexity on the optimization space. And all these different options add up together; it can lead to more than 80,000 possible ways to optimize. So then how does anyone even come up with the optimal point for a particular workload? And that's a search problem. Basically, we want to find the one out of these 80,000 possible options that is the best for a particular use case. So it's a search problem.
Lin Qiao [33:40] And this is not new. This problem exists in areas like compilers, right? A compiler is solving a search problem. This exists in areas of database query optimization. And speaking of databases, very interesting: in the early days, databases were designed to be very hard to optimize. And it created a new career called DBA. It's a very well-paid career. It's a very well-paid job. People get a certificate to qualify to be a DBA and enter the industry.
Lin Qiao [34:25] But over time, database vendors started to build automation. They built query optimizers where it takes the query workload and spits out, rewrites the SQL, and spits out the most efficient new way of running the SQL. But it's one-dimensional optimization. It only optimizes for speed. And here I'm talking about three-dimensional optimization across quality, speed, and cost. So it's a much more complicated problem, a much more complicated search problem we're solving through FireOptimizer. And that's kind of the value-add we want to—we want to free up app developers to not need to understand any of this complexity and get the best quality, speed, latency, and cost.
Real-time vs pre-set search
Lin Qiao [34:44] Cost for their new user experience.
Matt Turck [34:58] That search problem, do you solve it on a per-task basis, sort of a priori for a specific use case, or are you going toward real-time optimization? Or maybe you do that today?
Lin Qiao [35:07] Right. So we have two modes, similar to agentic development. I think agents are a very hot topic these days.
Matt Turck [35:08] Yeah, I'm hearing.
Lin Qiao [35:45] Yes, everybody's talking about agents. And there are basically two extremes of thinking about agent development, right? So one extreme is autonomous, right? The goal here is to replace humans and generate the same quality or even better results than what humans can do. And they are tapping into human resource budgets of different professions and so on. So that's one extreme. The other extreme is human assist, to make the professional much more productive. Right. So similarly, when we think about this optimization, we also have two modes.
Lin Qiao [36:26] One is human in the loop. The second is more automated. So, human in the loop: take DeepSeek, for example. DeepSeek is a very, very complicated model architecture because it has more than 250 experts. And DeepSeek, actually, that company itself was running, and is still running, this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many more replicas. It's a very big distributed systems problem to solve.
Reinforcement learning
Lin Qiao [36:56] So that's why that model itself is very hard to tune. So we enabled supervised fine-tuning for DeepSeek, where people need to be in the loop to label, give labeled data, and feed it into supervised fine-tuning. We have seen excellent results tuning DeepSeek models to customize to a specific domain. So that's human in the loop.
Matt Turck [37:03] And not to go into too many different rabbit holes, but that's the reinforcement learning with rewards.
Lin Qiao [37:42] So reinforcement learning with verifiable reward is a very simple version of reinforcement learning. And there are many domains where getting a verifiable reward is simple, is easy. For example, in the coding space, it's very easy to do grammar checks. And then it's very easy to run the code and see if it's runnable. Does it produce the right results or not? We have also seen other domains that are more vibe-based.
Lin Qiao [38:28] It's not like math, where it's absolutely correct or not. For vibe-based domains, people start to develop a good practice of providing some kind of API where there are, for example, designers sitting behind the scenes to rate certain kinds of design, whether it's good or not good. And that can feed into the reward to automate. So we start to see very active development of providing those reward functions, whether it's programmable or, in a sense, human in the loop. But I believe that's the direction that the industry will be moving towards.
Function calling
Lin Qiao [38:56] And we want to build out this stack in a way—we don't want to be opinionated about how different companies want to plug in their reward function, but we want to make the integration extremely easy, and the rest of the feedback loop should be automated.
Matt Turck [39:06] It seems that another important part of the platform is function calling. Do you want to talk about what it is, how you do it, and why it's important?
Lin Qiao [39:38] Right. Talking about that, we need to kind of talk about agentic development. So, on GenAI, I think the first wave of application development has been focusing on building user experiences for humans to consume, because the nature of the content, the results being generated, are human-consumable content, right? And the model was trained to generate, optimize for human-consumable content. So building agents is more complicated than being directly human-facing because those models need to speak to each other.
Lin Qiao [40:21] Or connect with external APIs to enrich the results. So the fundamental need here is to make the model output feed into some kind of programmable framework. That's a fundamental change in this new agent development. So first of all, we need to teach models to generate programmable output. It's called constrained output generation, and to follow a certain schema. Usually it's JSON format for most API integrations. But we have seen many other customers—they want to generate XML, they want to generate HTML.
Lin Qiao [41:05] Those are all schematized formats. So we support a new mode called grammar mode. Any BNF format we can support. It's very generic. So that empowers developers to extend beyond JSON. That's the base foundation. And on top of that foundation, there are two different ways to satisfy agentic development needs. One is people build an orchestration layer to pull multiple models. Here, multiple models come across multiple different modalities. I can talk about examples there.
Lin Qiao [41:43] Or integrate models with databases, with search engines, with knowledge graph databases, and so on and so forth. So they take full control of the orchestration scheduling. The other approach is to expand LLMs, let large language models dynamically orchestrate, create a plan, create execution, and tie the result back and continue to serve the rest of the answers. But LLM is the main brain behind this. So these are two different modes. For the first mode, we provide a composable framework, especially across multiple modalities.
Lin Qiao [42:29] We have seen use cases, for example, medical claims. Medical claims processing historically has a lot of legal involvement because lawyers need to review. People submit pictures of their medical bills, and there is a lot of patient information. It's very complicated document processing, and they have to review a lot of information, and then they will need to generate a report that follows some kind of format, legal format, to decide go, no-go, and what are the implications, and so on. It's a very lengthy process, and we know lawyers' time is very precious.
Matt Turck [42:36] And expensive.
Lin Qiao [43:10] But a significant part of this process can be automated, combining a vision model which can read and see with a large language model that is tuned toward generating legal reports and composing them. And you can pick and choose which vision model you want to use because, as I mentioned, no one size fits all. Our prediction is the future will be a lot of expert models focused on solving certain problems really well. There are vision models really good at reading handwritten documents. There are vision models really good at processing PDFs, processing tables, weird tables spanning across pages, and so on.
Lin Qiao [43:48] So there are a lot of complexities in the document format. And there are language models really good at generating summarization in general, or specializing in legal terms, or specializing in finance terms, and so on and so forth. So this composability allows you to pick and choose the best and solve your use case the best. So we have this multimodal composable framework. So that's one for you to orchestrate. And then when it comes to large language models doing dynamic orchestration, that's where tool use or function calling comes into the picture.
Lin Qiao [44:15] We have been working on function calling for a long time, and we are the first one to enable function calling for DeepSeek models. And DeepSeek models come in as raw models, and we add function calling to the V3 model, which created a lot of interest.
Matt Turck [44:17] That was recent, right? That was announced just a couple of weeks ago.
Low-level architecture overview
Lin Qiao [44:23] That was announced recently, and many people have been using it and gave us very good feedback.
Matt Turck [44:36] Let's talk a little bit about the low-level architecture part of the platform, in particular, sort of the GPU layer and the work you've done.
Lin Qiao [45:07] So we have been investing in overall speed and cost optimization. They're two sides of the same coin, shall we say, for a long time. And from the PyTorch days, we actually have a lot of customized kernels. We did a lot in-house. And we actually wrote PTX. So a lot of kernels are built on top of CUDA, but we also build kernels at the lower level. You think about PTX as writing assembly code, and those are extremely high-performance kernels to maximize software performance toward the hardware limit.
Cloud GPUs & hardware support
Lin Qiao [45:48] So we have been building that for large language models. Of course, large language models now have different classes. Mostly, the biggest buckets are chat models and logical reasoning models. They have different characteristics. And then we have been optimizing for communication, any possible bottleneck that we have optimized for.
Matt Turck [46:00] And the GPUs themselves, they sit on some cloud, or how does that work? Do you run on all sorts of NVIDIAs and AMDs, or are you agnostic that way? Or how does that work?
Lin Qiao [46:36] Yeah, similar to the PyTorch idea, we want to build on top of the best hardware possible. And different hardware SKUs have different strengths. So we build on top of NVIDIA GPUs from many generations to the most recent one, B200. We also build on top of AMD's hardware, and we optimize on top of ROCm as their CUDA equivalent layer. So we have been deployed to more than 15 regions globally. Currently, we're on top of five different clouds.
VPC vs on-prem vs local deployment
Lin Qiao [47:16] We plan to expand to possibly 10 different clouds and a lot more regions globally. And in terms of deployment, we have our hosted, fully managed single-tenant API. We can also deploy into your VPC and into your cloud for privacy concerns. Or we can connect VPC to VPC through PrivateLink or VPC peering. So all these possible ways of deployment we have enabled for various different enterprises and startups.
Matt Turck [47:48] Are you seeing robust demand for the latter, meaning enterprises want to deploy AI in VPC, on-prem, perhaps even locally? What's your sense of the market demand? There's this whole thread of conversation around the fact that people just don't want to send their sensitive data to OpenAI or any—not to pick on them, but any kind of API provider. What's the reality of that from your perspective, based on what you see in the field?
Lin Qiao [48:16] I feel there are market demands or constraints that are not movable. And there are constraints that are movable. So I think the default mode when we work with enterprises, the default mode is no data can leave our premises, for good reasons. But at the same time, the market is very interesting. So first of all, model depreciation has never been faster.
Matt Turck [48:23] Every week. Every week.
Lin Qiao [48:25] You never know tomorrow what's going to happen.
Matt Turck [48:29] Quite literally. It's not an expression.
Lin Qiao [49:04] So catching up with these fast-moving trends of model innovation happening across all different modalities, all different spaces, is really hard and interesting. And the second is hardware differentiation is happening also very fast. In the past, every three years there's new hardware. That's kind of the cadence we're thinking about. New SKUs coming up. And this year there are three, like each vendor probably has three SKUs planning to get into production. And with the moving target of new model innovation and new models typically optimized for newer hardware and new hardware advancement, enterprises ask us, hey, it probably is better to use our hosted API because we always move towards the top of the state-of-the-art models, state-of-the-art hardware, and they want to get onto the state-of-the-art technology.
Decreasing inference costs and its business implications
Lin Qiao [49:50] So we have seen the trend of enterprises moving to our single-tenant, heavily secured hosted API more and more. So that's very interesting, while we still deploy into their cloud, their premises, to run high-production traffic.
Matt Turck [50:16] From a business standpoint, there is this wide expectation across industry that the cost of inference needs to keep going down and will keep going down. What does that mean for you in terms of business as an inference infrastructure provider? Is there an expectation that you're going to be cheaper and cheaper and cheaper? And if so, how do you build your business over time?
Lin Qiao [50:50] So I believe overall AI infrastructure would go down by an order of magnitude in terms of cost. And they should. Where it is today is—I don't think it's a stable state to be where it is today. And this whole trend of infrastructure should become a utility is really, really good for the industry. Because today we see a lot of cases that I think today the infrastructure costs basically start to separate out two different concepts. One is product-market fit, as in, is this product useful?
Lin Qiao [51:22] Does this product add value? It is a different concern to: can you build a sustainable business? The fundamental reason is, even if people can have product-market fit, they get really good feedback from their users, they cannot scale it to the maximum degree because this is so expensive. And we literally have seen cases that a startup can scale to bankruptcy. And for big companies, for very big companies, like, the GPU bill surfaced to their CEO, and they're, like, asking questions: what the heck is going on here?
Lin Qiao [52:08] Why are we burning so much money? And we should really think about ROI here. So because of that high cost of infrastructure, infrastructure for GenAI, it basically creates a high bar of what kind of application can be a viable business. If this bar can be lowered by 10 times, you can imagine there are so many more—it will be 100 times more applications entering this arena to create a brand-new experience for end consumers and prosumers. And by that, we'll see a much bigger consumption across the industry.
Fireworks roadmap
Lin Qiao [52:46] So I don't think I worry about my business. I actually want this cost to go down. It has been our primary focus to help application developers build a sustainable business through Fireworks because we hyperfocus on cost efficiency, beyond just speed and quality. So from a total cost of ownership side, they can have the longevity of a durable business. And that's what we care about.
Matt Turck [53:10] So maybe to close, some forward-looking, future-guessing kind of thoughts. So one, for Fireworks specifically, in the next 12 to 18 months, what's on the roadmap? What are you excited to build that you can talk about today? And then for the broader industry, any thoughts for where things may be going from your perspective?
Lin Qiao [53:39] Yeah, I'm super excited. We are doubling down on the Flywheel Optimizer. It's our customization engine across quality, speed, and cost. We have made a tremendous amount of progress in the past year, and we are getting that to mass production with many customers. They now have a viable path to both product-market fit and durability. And we are going to focus a lot more on the quality side for Flywheel Optimizer because we're already doing really well on speed and cost, and we want to maximize the quality gains through customization. The quality should be much better than if they build on top of any generic model, any generic model API, because it should, right?
Lin Qiao [54:41] So if you fuse use-case-specific data into it, it should just shine. And for application developers, I think fundamentally, what is that moat, right? Their moat is probably not the user experience, because it's very easy to copy. Anyone can study the product and copy. The moat is data. It's the data they curate, or the insight into usage, like usage patterns, and kind of create this flywheel. And this flywheel should feed back into the GenAI model they use.
AI future predictions
Lin Qiao [55:04] And we want to kind of really double down to make customization heavily push on the next-level quality. So I'm super excited about that level of investment. And we will launch a series of product features along those lines.
Matt Turck [55:18] And then, any thoughts on the future of the industry at large, like the next 12 to 18 months? Is that the year of agents that everybody's talking about? Is that the year of adoption? What's your prediction? What are your predictions?
Lin Qiao [55:53] So there's no doubt 2025 is the year of agents. Literally, no matter where you go, people are building all kinds of agents. So I think there will be thousands of agents being built focused on solving specific problems. At the same time, I also think 2025 is the year of open models. And it's clear that DeepSeek has created a big dent there. But the meaning is not about DeepSeek itself. It's the precedent it has set in the world about setting a huge, very solid baseline for open models, for open-model providers, that it just needs to be better before anyone can open-source new models.
Lin Qiao [56:17] It also set a precedent for all model providers to be better.
Matt Turck [56:18] Better.
Lin Qiao [56:58] I'm very bullish on open models because I have seen the power of open source that PyTorch has been able to leverage. DeepSeek, for example, just within one month of releasing their new models, there are, despite DeepSeek models being extremely hard to tune and optimize, extremely hard, more than 500 variants published on Hugging Face, optimizing for local devices, optimizing for cloud infrastructure. People tune DeepSeek models, people distill all DeepSeek models into various different kinds of models. Companies like Perplexity also tune DeepSeek models for—
Matt Turck [57:03] 1776.
Lin Qiao [57:36] Yes. And we launched that model in partnership with them to offer much more unbiased, accurate information for deep research. Linq is another customer of ours. They focus on—they launched a DeepSeek model, a tuned DeepSeek model, for the financial assistant they're building. So we see a lot of vibrant energy from the open-source community. And that's why we believe the future of modeling sits on the open-model side. And I believe that side is going to be much more active in creating those hundreds or maybe thousands of expert models that are specialized in delivering much better quality in certain domains.
Lin Qiao [58:32] And the interesting challenge then is between those active developments of agents and active developments of open models. Across these two, there will be heavy spaghetti in the middle, and that complexity of navigating and doing the last-mile quality alignment between these two and delivering real-time experience for the agentic user to consume is going to be the next-level challenge. And we want to solve, take over, solve a big chunk of those complexities and simplify it for the agentic developers in the future. And that's also going to be our this year's focus for Fireworks.
Matt Turck [58:41] So do you tell your friends and parents that you are in the spaghetti business?
Lin Qiao [58:47] We are. We love spaghetti. We love to simplify spaghetti.
Matt Turck [58:52] That feels like a wonderful place to live in. Thank you so much, Lin. This was terrific.
Lin Qiao [58:53] Yeah, thank you.
Matt Turck [59:14] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.