Inside the Paper That Changed AI Forever - Cohere CEO Aidan Gomez on 2025 Agents

The MAD Podcast with Matt Turck · with Aidan Gomez, CEO, Cohere

Aidan Gomez is the CEO at Cohere. We cover why transformers remain dominant because infrastructure and chips are optimized for them, why reasoning models deliver a large intelligence uplift at far less cost than pre-training, and how agents can cut financial research and hedge proposals from weeks to four or eight hours.

Watch on YouTube

Chapters

  1. 2:00 — The Story Behind the Transformers Paper
  2. 3:09 — How a Cold Email Landed Aidan at Google Brain
  3. 10:39 — The Initial Reception to the Transformers Breakthrough
  4. 11:13 — Google’s Response to the Transformer Architecture
  5. 12:16 — The Staying Power of Transformers in AI
  6. 13:55 — Emerging Alternatives to Transformer Architectures
  7. 15:45 — The Significance of Reasoning in Modern AI
  8. 18:09 — The Untapped Potential of Reasoning Models
  9. 24:04 — Aidan’s Path After the Transformers Paper and the Founding of Cohere
  10. 25:16 — Choosing Enterprise AI Over AGI Labs
  11. 26:55 — Aidan’s Perspective on AGI and Superintelligence
  12. 28:37 — The Trajectory Toward Human-Level AI
  13. 30:58 — Transitioning from Researcher to CEO
  14. 33:27 — Cohere’s Product and Platform Architecture
  15. 37:16 — The Role of Synthetic Data in AI
  16. 39:32 — Custom vs. General AI Models at Cohere
  17. 42:23 — The AYA Models and Cohere Labs Explained
  18. 44:11 — Enterprise Demand for Multimodal AI
  19. 49:20 — On-Prem vs. Cloud
  20. 50:31 — Cohere’s North Platform
  21. 54:25 — How Enterprises Identify and Implement AI Use Cases
  22. 57:49 — The Competitive Edge of Early AI Adoption
  23. 1:00:08 — Aidan’s Concerns About AI and Society
  24. 1:01:30 — Cohere’s Vision for Success in the Next 3–5 Years

Transcript

The Story Behind the Transformers Paper

Matt Turck [2:01] Aidan, welcome.

Aidan Gomez [2:03] Thank you. Thank you. Thanks for having me.

Matt Turck [2:28] So to get started, I'd love for you to tell the story of the Transformers paper. You're famously one of the eight co-authors of it. And of course, Transformers and "Attention Is All You Need" is the seminal paper that led to everything that we're experiencing today in generative AI. What was your journey to it? How did you become part of the eight?

Aidan Gomez [2:54] I was a student at the University of Toronto, and I had been working on deep learning because a lot of the early work in the field was obviously done by Geoff Hinton and others at the school. So I'd been getting quite close to it, and I was reading up on papers, and I just kept seeing Google Brain repeatedly across these papers, again and again and again. Researchers from Brain. Eventually, I ended up reaching out to them, reaching out to some of the researchers there, and saying, "Hey, listen, I read your paper."

How a Cold Email Landed Aidan at Google Brain

Aidan Gomez [3:10] I have this idea of how to extend it. One thing led to another, and I got an offer to join.

Matt Turck [3:14] You just cold emailed them? You found their email address and just cold emailed them? That's awesome.

Aidan Gomez [3:43] Yeah. Well, on papers, you have the email address under the author name. I could just reach out to them to say, "Hey, super cool experiments. Did you try this? What if we did this next?" And they replied to me and said, "Hey, why don't you come down to Mountain View and do some work with us?" So I got hired as an intern on Łukasz Kaiser's team, and I was sat next to Noam Shazeer. And that was sort of the start of it.

Aidan Gomez [4:17] The end of it, if I skip all the way to the end, I was leaving Google after the internship had finished up, and they were throwing a little goodbye Aidan party, whatever, like some sweets and stuff. And my manager Łukasz is like, "Okay, everyone, Aidan's going back to his PhD. Aidan, how many more years have you got left?" And I had to be like, "Oh no, I need to finish third-year undergrad." And Łukasz was like, "What? We don't hire undergrad students."

Aidan Gomez [4:52] I think I got in through an administrative mistake because my manager thought I was a PhD student. So that's sort of how I got there. The process of the Transformer, I mean, it was incredible. The velocity that came together was—I haven't seen anything like it since. So I showed up, the project that I was supposed to work on, it's a paper that is out and it was released at the same time, called One Model to Learn Them All. And it was like an omni-model.

Aidan Gomez [5:28] So you can feed in text, audio, images, everything, and it could output the same. So super multimodal. Again, this was eight years ago, and so primitive compared to what we have today, but certainly prescient to where things were going. But as we were working on that, Łukasz and I built this framework called Tensor2Tensor. And it was used for doing big training jobs, distributing over a bunch of different GPUs, making things super efficient. Because we were sitting next to Noam, we convinced Noam to use it.

Aidan Gomez [6:09] So Noam joined Tensor2Tensor. And then Noam was in conversation with folks over at Google Translate, which was like the second group of people who were working on these projects. And we realized folks were kind of working on the same thing. We were all looking at text-based autoregressive models that heavily leveraged attention or were much more pure attention, as opposed to these previous RNN models, LSTM models, which were quite complicated and kind of, in some ways, ugly. So we wanted to strip all that back and just create the most simple, efficient, attention-based language model.

Aidan Gomez [6:25] And so then we just decided to team up and join forces. And that happened about a month into my internship.

Matt Turck [6:44] So that was purely organic, like Noam happened to be around, and then you had some conversations with other folks. Was that how Google Brain operated? Meaning Google Brain allowed people to just organically form groups and teams?

Aidan Gomez [7:07] Yeah, totally. That was it. It was a group of people who were researchers with full academic freedom to do whatever interested them. And you would sort of congeal around projects or ideas. And so that's what happened. We were just chatting to folks, saw a good idea, and then teamed up to take it on.

Matt Turck [7:38] It seems that it's a defining characteristic of successful research organizations. At some point recently, we were chatting with Daryl Aquila, now of Contextual, and he was talking about FAIR at the time, and it sounded like it was a little bit like that as well. Do you think that was a moment in time when all those labs were authorized to, or allowed to operate that way and given free rein to explore anything? Is that still true today? Has it changed?

Aidan Gomez [8:11] I don't know. I've been out of Google for long enough that I'm not sure how the culture has shifted. I would say that the economic relevance of this work is very different to when I was interning eight years ago. So I imagine things would have to shift out of necessity, especially because of the product implications, the amount of resources that are being thrown at these projects and these models. It's much more consolidated, I would imagine and expect. Certainly the way we run things at Cohere, it's much more like a product organization.

Aidan Gomez [8:47] Like, you have very clear work streams. There's scope to experiment and try to find new alpha, but towards the ends of the product, right? So it's focused and more narrow. But back then, it was very greenfield. And so you could work on whatever you were excited about. And yeah, I do think that is like a crucial component of successful research organizations. And it worked. It produced incredible technology.

Matt Turck [8:54] So going back to the story, the eight of you got together, and then how long was that process of writing the paper?

Aidan Gomez [9:25] Super, super fast. So probably about a month in, we decided to consolidate and all work on the Transformer together. And it just became a mad dash towards the NeurIPS conference deadline. So NeurIPS is like the biggest AI conference for academics where you submit your papers to. And so we were just all-out sprinting. And it was a lot of, honestly, it was a lot of throwing shit at the wall and seeing what sticks. So many different things were tried, so many little bugs kind of hackily patched.

Aidan Gomez [10:03] One example is pure attention architecture. The model can't tell the difference between positions of the elements in the same way that an LSTM could, because it could consume each one one by one. And so then Noam just came up with this idea. I remember the day I was sitting next to him, and he was talking to me about it, of just throwing these sinusoids into the embeddings and having that represent the position. And it stuck. And I think we've moved on a little bit, but shockingly, we're still quite close to that strategy today.

The Initial Reception to the Transformers Breakthrough

Aidan Gomez [10:39] So that's how stuff worked. It was just fixing bugs one by one as quickly as possible. And then whatever we were left with at the last moment, that's what we submitted to NeurIPS. And so one of the big shocks is how, over the past eight years, how little things have changed. It's really surprising to me that the transformers we train today look so similar to what was back then.

Matt Turck [10:51] And then when you guys submitted the paper and got accepted into NeurIPS, what was the general reception to it? Was it clear to folks that this was going to be a big deal or not?

Aidan Gomez [11:11] Yeah, I think folks noticed that and were pretty excited about it, but it was still like the folks who were in NLP or translation, a subset of a subset of the AI community. And so it was fairly limited in terms of reaction, but the folks who knew were quite excited about it.

Google’s Response to the Transformer Architecture

Matt Turck [11:35] And then the much-debated question of why did Google not immediately jump on this? And eventually, famously, OpenAI is the company that leveraged the Transformer architecture faster. What's your insider perspective on that?

Aidan Gomez [12:09] No, they jumped all over it. So it went to production inside of Search, inside of Translate, the existing product suite. And so to say that they didn't adopt Transformer architecture would not be correct. To say they didn't lean hard enough into language modeling, like just pure sequence modeling of text on the internet, that's, I think, the accurate statement. That's what OpenAI did early and uniquely well. But certainly it was everywhere across Google quite quickly. BERT and the Search folks, they figured out how to make use of the Transformer super fast.

The Staying Power of Transformers in AI

Matt Turck [12:45] And to the point that you were making a second ago, why do you think Transformers have had so much staying power? Is that because it's the gift that keeps on giving, and the more data and compute you feed it, the better it performs, and we haven't reached that moment when things get less exciting? Or is there something in research where people are, I don't know, for whatever reason, not as productive in terms of new ideas?

Aidan Gomez [13:15] I don't think people are not productive in terms of new ideas, but I do think it's sort of a reinforcing loop, or like a self-fulfilling prophecy, where the community got super excited about the Transformer. They built so much infrastructure specialized to the Transformer. And so it's like we dug ourselves into this well. We now have chips that are being optimized explicitly to that architecture. And so, to move architecture, it requires so much effort, energy, lift to rewrite everything and start from scratch.

Aidan Gomez [13:35] That new architecture needs to present something extraordinarily compelling, like a very good reason to move. And we just haven't found that architecture yet.

Emerging Alternatives to Transformer Architectures

Matt Turck [14:10] Yeah. So the bar is super high. And I don't know if that's an unfair question because you're now a CEO, so you're presumably, I'm sure, focused on building a fantastic company pretty much all day, every day. So I don't know how long you spend looking at research papers, but is there anything that you find exciting? So that's, Yann LeCun has proposed alternative architectures to LLMs. There was discussion about state-space models, all that stuff. Is there something that's emerging on your radar as a post-transformer architecture, possibly even if it doesn't work today, but it sounds promising?

Aidan Gomez [14:49] Yeah. So I have for a long time believed and kind of hoped for a replacement to the transformer. And I think most of the transformer paper team, like, we're researchers, we want to see our stuff surpassed. It would be terrible if the best we could do is this paper from eight years ago. That's just bad for humanity, right? We want to see progress. So I've been hoping for that. So much so that when we opened the New York office for Cohere, I named one of our meeting rooms SSM because I was like, this is it, it's going to get replaced.

Aidan Gomez [15:31] And it hasn't. It hasn't. It turns out the transformer is a great artist. It copies all the good ideas that it sees out there. And so it just hasn't been replaced. When SSMs came out, the good ideas from that got ported over to the transformer. And we kept going with transformers. Now there are these discrete diffusion models, which do diffusion, which has been super popular for image understanding, image generation. It's doing that same process for language models, but I still don't see that replacing the transformer.

Aidan Gomez [15:44] So I'm waiting like everyone else is, but I am hopeful that we'll get something.

The Significance of Reasoning in Modern AI

Matt Turck [16:17] And the whole reasoning/test-time compute that, at least for the non-AI researchers among us, seems to have come out of nowhere in the last sort of like five, six months in particular. From your perspective, is that a somewhat obvious idea that was sort of—it was a matter of time until it was going to be implemented, and it's good, but not completely groundbreaking? Or is that a major development?

Aidan Gomez [16:46] Well, it's been worked on for a very long time. So we've known it was coming for years, for the past three years. And it is kind of obvious. It is kind of obvious. Because if you think of the pre-reasoning world, the input space to a language model is everything. It's all of language. And so you can ask it very simple questions like one plus one, or extremely complex ones like, "Go cure cancer." Those are two strings or requests that you can ask it, and you really don't expect it to spend the same amount of energy and time on those two different problems.

Aidan Gomez [17:27] One, it should respond immediately. The other one, it should probably—it might take years of thinking and trying to accomplish. But we didn't have that reality before reasoning. We had an input and then an immediate response. And so both of those got the same energy and effort put into them. So it had to come at some point, this notion of different amounts of energy or time being spent on problems: test-time compute. I think the effectiveness of it was surprising.

Aidan Gomez [18:03] It was really quite incredible to see how much gets unlocked. These models actually can, with very little supervision, very little data from humans saying, "This is how you think through problems," sort of figure it out for themselves. So that's been incredible to watch. I think the other thing that's been incredible is it's actually really easy to do. It's easy to create a reasoning model. It's dramatically cheaper than pre-training, and so it's accessible. And so there's this huge intelligence uplift that comes for really quite little effort.

The Untapped Potential of Reasoning Models

Matt Turck [18:16] Is there more juice to squeeze there? Should we expect more progress specifically under that paradigm?

Aidan Gomez [18:50] Totally. We've just scratched the surface. At the moment, it's mostly focused on math problems and this sort of thing. There is a whole world of applications that we need to make it work in: medicine, everything from the pure sciences, physics, chemistry. There's so many—all the interesting problems require reasoning. And so there's a lot of white space for us to go after.

Matt Turck [19:11] And maybe to unpack this for people, the reason for that is because each time you work on test-time compute, there's an element of bringing a domain into the effort. So people have done that for math, but they haven't done that for other areas. Is that the right way to think about it?

Aidan Gomez [19:39] Yeah. So they've spent a lot of time teaching the model to reason in the domain of mathematics, in the domain of computer science and coding. But that effort hasn't been applied to biology yet, or some of these pure sciences. And it hasn't been applied to automation in the enterprise world. So using the tools that enterprises use to create value in our world, to get things done.

Matt Turck [19:43] So after the Transformers paper, what was your journey to starting Cohere?

Aidan Gomez [20:14] So I bounced around Google for a while. I went back to Toronto. I started working at Google for Geoff Hinton there. I met my co-founders, Nick and Ivan. And then I started my PhD in England, and I was still working at Google. I was flying back and forth between London and Berlin because in Berlin, like London, it was sort of exclusively DeepMind, which was another arm of Google that did AI research. And so Brain didn't really exist there.

Aidan Gomez [20:51] But one of the Transformer co-authors, Jakob, had opened a Brain office in Berlin. And so I would bounce back and forth to go see him. And we were doing work with Geoff Dean on scaling up infrastructure, so training on networks of TPU pods. So instead of having one supercomputer, you chain together a bunch of supercomputers and you can train something dramatically larger. And that's where we started to see the first instances of scaling up large language models.

Matt Turck [20:53] And that was what year?

Aidan Gomez [21:18] That would have been 2019 or 2020. No, 2019. It would have been just before GPT-2. So the first time computers started writing in a way that was compelling, almost like a human, you read it and you're like—that was the moment I felt shock reading what I was reading.

Matt Turck [21:47] That's amazing. And just to unpack that, that was a surprise to you. I mean, that's one of the things I found the most fascinating about this whole thing, is that even people who are super deep in the field would be shocked, because there's a whole train of thought that says, well, the rest of us are impressed because effectively the computer does what the computer always did, but now we can emotionally relate to it because it uses language. So I find it fascinating when somebody of your caliber says, well, no, it was a shock to me as well.

Aidan Gomez [22:23] Yeah, no, I think it's like, before you press go on one of these experiments, you have some belief, you think it'll get somewhere. But then when you actually get there, it takes you a while to get used to it. I had the same feeling with some of these voice models that are able to inject emotion—the experience of interacting with a machine where you hear it inhale before saying something. You hear its lips smacking as it talks to you. You hear it hem or haw.

Aidan Gomez [22:53] It tickles a part deep down in your brain. The experience is just so incredible. You can't prepare yourself for it when it's actually in front of you. You just can't wipe the smile off your face because it's so crazy. I had the same thing with reading the first outputs of these models, and just—it's so delightful. It's shocking. It's so creative. The first sample that I got was sent to me by Łukasz, my manager.

Aidan Gomez [23:27] And I've told this story a bunch before, but it was like an email, and the subject was, “Aidan, look at this.” And then in the body, it was a Wikipedia article titled “The Transformer.” It was about this Japanese punk rock band that had gotten together. I was just reading through the story of this. And then at the end, Łukasz was like, “I just wrote ‘The Transformer.’ The machine wrote the rest.” And I was just like, “What the fuck?”

Aidan Gomez [23:56] What do you mean? This is completely written by a human. So those moments, they're not that frequent, but they're pretty frequent. It's like once a year. Same with reasoning. When you're reading through what this model is thinking, you're like, “Holy shit, it's so solid.” The agentic part? Yeah, it has a monologue. It's talking to itself. It's thinking through stuff. “Oh, I messed up. Let me try this,” and figuring it out.

Aidan Gomez [24:02] It's so beautiful to see. It's very cool.

Aidan’s Path After the Transformers Paper and the Founding of Cohere

Matt Turck [24:14] It's very cool. So you were in England working with Geoff Dean and others, and then at some point you decided to leave and start the company. Where was the next hop?

Aidan Gomez [24:47] Yeah. So when computers started to get quite compelling in language using these language models, I called up Nick and Ivan and said, "Guys, we need to do something here. There's a very interesting direction of travel. Let's see if we can raise some money and build a company that builds models of the web." Because that's what this project really was. That's what these language models were. They were just models of the content on the internet. And so that's what we started doing.

Choosing Enterprise AI Over AGI Labs

Aidan Gomez [25:16] And then very soon after that, we wanted to explore enterprise applications of it, stuff like customer support and chatbots, this type of thing. So we quickly became an enterprise company, and the rest is history. We're now, I think, approaching 500 people, with offices in SF, Toronto, New York, London, Tokyo, Seoul. And yeah, it's been really incredible to scale the company.

Matt Turck [25:34] Amazing. Why did you guys decide early in the company to build an enterprise company? Certainly, given your credentials, you could have been one of those AGI labs. Was there a specific reason why you decided not to do that and focus on the enterprise instead?

Aidan Gomez [26:02] Yeah, I never liked the vibes of the whole AGI, effective altruist, this whole ecosystem. It never resonated with me. It felt like cosplaying. It felt like people were LARPing a new religion and all of this stuff, like, create God. I just didn't like that. I didn't like that ethos. And frankly, enterprise maybe has, like, a rap of being boring, but I think it's way more important. I think the idea of increasing human productivity, letting humans do more, increasing supply, driving costs of things down, letting humans do more and accomplish more, that's what inspires me much more than building God or saving the world from AI.

Aidan Gomez [26:51] I want to save the world with AI. I want to put it to work to actually make healthcare better. I want doctors to spend less time writing up notes and filling out paperwork. I want them to be with the patient, thinking through problems. I want to help them solve those problems, give them agents that can help them do research on something they've never seen before. So I want to put the technology to work in the global economy. And that's really what inspired Nick, Ivan, and I.

Aidan’s Perspective on AGI and Superintelligence

Matt Turck [27:13] Do you spend any time thinking about AGI and superintelligence? Bearing in mind what you just said, do you think that's something where we actually—do you think that's a tractable problem and we're getting closer, or nobody knows?

Aidan Gomez [27:39] The goalposts keep moving constantly, and we get to the place that we thought AGI was, and we say, "Oh, actually, I guess it's a little bit harder than that. No, it's out here now." Well, to your first question, can I avoid it? Do I spend time thinking about it? I can't avoid it. I do spend time thinking about it, of course. And many regulators and policymakers ask me questions about it, et cetera. So I have to think about it.

Matt Turck [27:45] That sounds like fun. Yeah, talking to regulators about AGI. I'm sure that's your favorite activity.

Aidan Gomez [28:15] It's a joy. Yeah, I love it. I'm glad people are thinking about it. I don't think it should be the center of the discourse in the way that it historically has been. I think there's been positive change to more of a practical focus. But certainly early on in the language modeling game, it was doomsday. It was, "We're going to save the world." It was, "This is so risky, we have to shut down everything." And so I think that era has passed us, which is good.

The Trajectory Toward Human-Level AI

Aidan Gomez [28:37] And I am glad that people think about that stuff. The long-tail risks are important, and they're worthy of academic inquiry. And so I'm glad, but I'm very relieved we're past the point where that's the only thing people seem to talk about.

Matt Turck [29:02] But doomerism aside, this concept of AI continuing to be on some kind of exponential curve towards, whether you call it human-level intelligence or superhuman, is that something that you see progress? Or effectively, we've made a lot of progress, but nobody really knows when that's going to stop or whether it's going to accelerate, or what's your sort of expert take on it?

Aidan Gomez [29:34] Listen, the models are going to continue to get better. They're going to continue to get better, and they're going to do some incredible things. And there will be specialist models that emerge to help in things like pharma, right? The creation of new drugs, in material science, for advanced materials, the models will be able to be incredibly helpful to us. The definition of ASI and AGI is so hazy and ill-defined. It's hard for me to give a concrete answer to what you're asking, aside from it must be a continuum and not a discrete bit flip where suddenly it's ASI.

Aidan Gomez [30:11] Or AGI. I think we already have AGI to a large extent. If you have the choice between—you right now, you have some symptoms, and the only option is Aidan prescribes me drugs or Cohere's model Command prescribes me drugs—the logical and actually correct answer is to have the model do it. I promise you, it knows more than I do. Now, it shouldn't be prescribing anyone drugs, but it's smarter than me at that thing and many other things, right?

Aidan Gomez [30:21] Like most things, actually.

Matt Turck [30:29] I personally thought we had AGI when Google Photos was able to recognize my family in thousands of photos. Sorry, my bar is low.

Aidan Gomez [30:53] Okay, yeah, so you're already there. And then the ASI thing, like, better than humans? Of course it will get better than humans at some things. Like I just said, it's better than me at this. Is it better than the world's best doctor at prescribing drugs? Probably not. Will it get there? Probably. So I think betting against progress is bad. I think it's, what does that mean? What does that progress mean? What are the tangible effects?

Transitioning from Researcher to CEO

Aidan Gomez [30:58] Anyone who's selling you doom and gloom, I think, is wrong.

Matt Turck [31:12] One thing that's really interesting about you guys as a founding team is that I believe the three of you are researchers, right? Like, you all met—that's what you were describing. All of you met at Google Brain. And it's something that I've been thinking about, the concept of AI researchers as founders and entrepreneurs, and it seems that there was a whole wave of people doing this, but equally there was a wave back of people just going back to the labs and sometimes going directly through employment, sometimes selling companies.

Matt Turck [31:53] So whether that's Adept or Inflection or Character.AI, what was your personal evolution to go from a world-class person writing one of the most important papers of all time to suddenly, oh, you have to worry about HR and fundraising and making customers happy?

Aidan Gomez [32:15] Yeah, no, HR is the worst. By far the worst. What was the transition? I would say it was gradual. It wasn't immediate. I was still doing loads of research for the first year and a half, two years of the company. But I've slowly been weaned or pushed off of research. I'm probably more annoying now to the modeling team than I am helpful.

Matt Turck [32:18] Guys, I can still do it.

Aidan Gomez [32:46] Yeah, guys, listen to this idea that I did. Yeah, no, it was gradual. I would say I love my job. I think being a CEO is such a privilege. You get to see so many different parts of the world. And not just, obviously, geographically, but, like, you get to just see so much of what happens out there. Different sectors, different types of people, private sector, public sector. And so it's been a huge privilege, but certainly a journey.

Aidan Gomez [33:08] And I've learned a lot. I've been lucky to have really good mentors. So I have an incredible board, Mike Volpi and Jordan Jacobs. They've been on my board since the very beginning and have really taught me everything I know.

Matt Turck [33:09] Great folks. Yes.

Aidan Gomez [33:26] And then I've met great founder CEOs, like Jensen and many others. And so it's been a lot of learning very fast and lots of failures, which I've had to adapt to and react to, but it's been a privilege.

Cohere’s Product and Platform Architecture

Matt Turck [33:58] So let's get into Cohere in more detail. You have built this enterprise AI platform, and I will let you correct me if those are not the right words, but interestingly, it's very vertically integrated. So you do the model part, which is Command, I believe, and then on top of that, you build other products, and I believe the most recent one is North, which is an agentic platform. So, in better words, how do you describe the company, what it does, and then that product architecture where all the pieces fit in?

Aidan Gomez [34:39] Yeah, I mean, for North, the fast way to describe it is it's like an AI agent platform where you can build agents, plug those agents into all the software and data that the humans inside your organization have access to, and then ask them to go do things. So they use tools, they use sales software, HR software, whatever, like your email, your docs, and it can just go out and accomplish tasks for you. So that's North. But I think it's more interesting to build up from the base, which is first and foremost, Cohere builds models.

Aidan Gomez [35:21] So we have our Command model, which is like our generative model, competes with GPT-4o, LLaMA, et cetera. We also have another set of models, which are our search models. These are our search models that can see, that can sift through data, surface information for you. And so those two form the backbone of North. They're what manages the model's ability to interact with the user, think through problems, use tools, and then find the data that it needs to accomplish a task.

Matt Turck [35:28] How are those trained, the Command and the re-ranker?

Aidan Gomez [35:50] The generative model, we actually released a very detailed paper on Command A, which describes exactly what we did. But similar to others, we do a large pre-training phase, which involves data from the web, synthetic data that we generate, and we do a big training run. And then we start doing SFT and RLHF. And so this is the part where you have a model that knows a bunch, and now we want to start making it capable of using that knowledge, using that intelligence towards some task, whether it's using tools, performing stuff like RAG, where you need to search databases, pull back information, respond with that information.

Aidan Gomez [36:38] So that's how we build that side. On the search side, it's similar, right? We have to search over really complex data. In particular, enterprise data isn't usually out on the web, so you can't find it out there. So we need to go create data that looks like it synthetically and train on that instead. And those models are fully multimodal. And I think I can say, and I think most people would agree, our re-ranker and embedding models are the best out there.

Matt Turck [37:03] And re-ranker, just to make this educational, that's part of the search architecture. And basically, that enables the customer to decide which results should be given priority versus others. Is that fair?

Aidan Gomez [37:15] Exactly, exactly. You give it a bunch of stuff, and there's some needle inside the haystack, and the re-ranker surfaces it. So it pulls it up to the front so that you can just grab that piece and pass it along.

The Role of Synthetic Data in AI

Matt Turck [37:48] What's the deal with synthetic data? A year or two ago, everybody was saying it doesn't work. It doesn't enable you to train as well. And I may be wrong, but that's what I would hear. And tell me if that's not correct. And this year, it seems to be—people seem to be viewing synthetic data as something that's completely ready for prime time and that they are increasingly using in just everything AI. So one, is that fair or not? And two, if that's fair, what happened?

Aidan Gomez [37:59] Yeah, there was a period where there were a lot of people saying, I forget the word, like the snake eating the snake, the ouroboros or whatever.

Matt Turck [38:00] Yes.

Aidan Gomez [38:31] Like a human centipede of data. And yeah, I think it just got decisively proven wrong. Synthetic data is incredibly effective. It's now the majority of the data that we train on for creating something like Command A. And in many instances, it's actually more useful to the model than human data. The most obvious example of that is stylistically. So humans, if you ask them to respond to a question, we're lazy. If I ask, "What's 9 plus 2?" a human is just going to say, "11."

Aidan Gomez [38:45] But what the human actually wants to see is, "Aidan, that's a great question. That's so interesting. So it's actually—"

Matt Turck [38:47] This is the first time I'm being asked this question. Yeah, yeah.

Aidan Gomez [39:27] Yeah, you're incredible. But stylistically, humans actually prefer models' answers that are incredibly empathetic, positive, patient to human answers. And so what we have to do is we have to get the models to rewrite the human answers because the humans are lazy and they're kind of like, "It's 11." And it makes you feel like shit. You're like, "Okay, fine, thanks." So that's the most concrete place where, yeah, synthetic data is just way better. But obviously, we're applying it all over the place.

Custom vs. General AI Models at Cohere

Matt Turck [39:59] Yeah. Because one of the many interesting things that you guys do is that you seem to have a development model where you work very closely with some key design partners per industry. So I believe you had RBC for banking. So I guess that's North for banking, the agentic platform. And then I believe you have a customer for telecom, and so on and so forth. Is part of the idea that you can train the model working with these partners, but ultimately, obviously, their data is their data?

Matt Turck [40:21] And if you want to generalize what you're doing to other financial customers, then you need to create synthetic data that looks like that data. Is that part of the idea?

Aidan Gomez [40:47] Yeah, that's exactly correct. So sometimes our customers either can't or don't want to train a model on their data. They actually don't want to train on their data. Either they don't have permission to, or they're just not comfortable with it. And so in that case, synthetic is the only option. But given a few examples of the ground truth data, synthetic data works extremely well. You can create huge quantities of fake data that is very, very faithful to the ground truth.

Aidan Gomez [41:13] And so that's being used all over the place for us. We're lucky in that most customers do trust us because of our deployment model. So we can deploy completely privately, like on-premises. We can air-gap it. So it's just way more secure than some of the other fine-tuning options that are out there. But if they still don't have comfort, we have this option that we can develop a bunch of synthetic data and show uplift in performance without having to actually train on real user data or real patient data.

Matt Turck [41:39] But ultimately, you have one big model, right? So what you learn in finance, what you learn in healthcare, what you learn in telecom, all that goes to the same model versus having specialized versions of Command?

Aidan Gomez [42:02] No, we do have specialized versions, so we can create custom versions for an enterprise. Now, of course, we want to be helpful to an industry, and if we know that a particular industry cares about a use case, we're going to make sure that our general model performs at that use case. And yeah, synthetic data will be a huge part of that. But for a particular customer, oftentimes they want a dedicated model for them, for their use case, for whatever it is they care about.

The AYA Models and Cohere Labs Explained

Matt Turck [42:30] And they want this to happen at the model level, not just the agents or interface, or they want— okay, interesting. So that's Command and Rerank. I read somewhere that the Aya models, is that on the side? Is that part of the research arm?

Aidan Gomez [42:33] So it's based off of our Command models.

Matt Turck [42:33] Okay.

Aidan Gomez [42:54] But it's extended and trained for many more languages, so like over 100 different languages. So it's our multilingual effort. That effort is being run by Sarah Hooker. She leads Cohere Labs, which is like our research nonprofit, and that used to be called Cohere For AI.

Matt Turck [42:55] Is that what it is?

Aidan Gomez [42:57] Cohere For AI, yeah.

Matt Turck [43:00] Okay. And before that, it had another name?

Aidan Gomez [43:01] Just For AI.

Matt Turck [43:03] Just For AI. Okay, okay.

Aidan Gomez [43:14] Yeah. That's what Ivan and I started, For AI, when we were in undergrad, just because we wanted to do more AI research. And we needed some money for GPUs.

Matt Turck [43:27] And so we built a little organization, and that's still around now in the form of that organization. Okay, great. So that's research. Is that the research part of Cohere?

Aidan Gomez [44:01] I mean, yes, it is, but so is Cohere proper. So everyone does research. Cohere collaborates with Cohere Labs very, very closely. A lot of the stuff that you see in Cohere Labs papers, you'll see pop up in the next version of the Command model or other models. And so, like, there's an Aya Vision model that we released today. So Command A isn't currently multimodal, but you can imagine very soon it will be.

Matt Turck [44:02] Mm-hmm.

Aidan Gomez [44:10] So Cohere Labs usually runs a little bit out in front of Cohere proper, but with deep collaboration to the org itself.

Enterprise Demand for Multimodal AI

Matt Turck [44:19] Interesting. Do you see enterprise demand for multimodal, or is that more anticipation of what may come?

Aidan Gomez [44:52] Totally. No, there's lots of demand. Some of the use cases, I mean, there's all the OCR-type stuff, but multimodal is essential for understanding enterprise data like PDF documents where there's graphs and this type of thing, or understanding slide decks. A lot of the modalities that enterprises work in are visual. So it's sort of table stakes. The other thing that vision is crucial for is computer use, right? It's a GUI.

Aidan Gomez [45:12] It's a visual experience. Trying to use a computer or even surf the web looking at HTML—you could try doing it if you want—is a horrible experience for language models too. And so the ability to see is essential to being able to navigate it.

Matt Turck [45:31] Multilingual seems to be something that you guys care about a lot as well. I read that you apply some of the same design principle or design kind of go-to-market with some key partners like Fujitsu for Japanese, I believe.

Aidan Gomez [45:33] LG for Korean, yes.

Matt Turck [45:40] So walk us through the thinking, how it came about, what's special and perhaps difficult about doing this.

Aidan Gomez [46:13] Yeah, I mean, the markets there are extremely underserved. They have populations that don't speak that much English. And the current technology that exists doesn't serve their needs, especially in the enterprise world. Maybe you can get them to speak good-enough Japanese at chitchat, but for actual enterprise documents, everything breaks down immediately. And so we focused on teaming up with regional champions. So Fujitsu is the largest SI in Japan, LG CNS is one of the largest SIs in Korea, and obviously one of the large chaebols.

Aidan Gomez [46:59] And so we team up with them to create models specifically for the market. So native in Japanese and Korean, focused on the enterprise use cases that economy cares about, whether it's manufacturing, finance, whatever is relevant in that space. So that's what we've been doing. And it's been extremely successful. It gives that market access to something that it would have had to just wait and wait or develop internally at massive cost overnight.

Matt Turck [47:18] All right, so that's the model layer. On top of that, do you have an API layer? I mean, it seems that part of the business is to provide the models to companies like Oracle or Notion. Is that the models themselves? Is it a separate product?

Aidan Gomez [47:46] No, we provide the whole platform. So serving of the models, optimization on different hardware like AMD, we're up and running on Cerebras, Groq, of course, NVIDIA. So all of that we provide, and we can deploy anywhere. And that's been quite unique. I think one thing that's different about agents and AI compared to other SaaS is that usually you're trying to do something that a human is doing in the organization. And to do that, you need the same context that human has.

Aidan Gomez [48:00] Et cetera. And that is a huge security risk, a very unique one compared to CRM software or HR software, that type of thing. And so our security posturing, the fact that we don't say, "Hey, send your data over to us, hit our API, trust us, we're SOC 2," the fact that we don't say that, and instead we say, "We're going to ship our models directly on your hardware, whether it's in your VPC on a cloud or, for regulated industries, in your data center."

Aidan Gomez [48:31] That's been a huge unlock. People get comfortable plugging in much more, and so they can actually do more with the product.

Matt Turck [49:13] Interesting. So that's one benefit of vertical integration. So it's not just that you control the whole experience, it's also that you can make the claim with a straight face, because obviously it's true that you're not sending data anywhere, right? As opposed to some direct or indirect competitors to Cohere who would say, "Hey, we can connect to all your sources, but ultimately we power it by Gemini, OpenAI, or Claude."

Aidan Gomez [49:18] And that's on Bedrock or Vertex or model as a service.

On-Prem vs. Cloud

Matt Turck [49:37] Whereas here, you control the whole thing. Okay. Are you seeing a lot of demand for on-prem and VPC deployments versus cloud? Is the overlap 100% with the regulated industries versus non-regulated industries? Or do you see some nuances there?

Aidan Gomez [49:51] I would say the most interesting work we do requires VPC or on-prem. It touches the most critical data to the business, and so we're having the most significant impact.

Matt Turck [50:12] And that's regardless of the regulated industries that kind of have to do it, but are you seeing non-regulated industries that still decide to use you guys on-prem or VPC because their data is so sensitive? I'm just curious, because this whole theme of cloud repatriation could be accelerated by AI. I'm just curious what you see in your daily reality.

Cohere’s North Platform

Aidan Gomez [50:31] Yeah, I mean, in startups that use us, the ones who care about on-prem and in VPC, they are serving regulated customers. That's what they're doing. And so it comes from a security standpoint. And I do think on-prem and in VPC resonates the most with the folks who have the most restrictions on what they can do with their data and where it can live.

Matt Turck [50:45] So, talking about North, what are some of the capabilities it has and doesn't have yet? Perhaps express in terms of use cases. What can it do? What can it not do?

Aidan Gomez [50:53] Yeah, so there's a few different modalities in North. We're still in early access, so we opened that up in January. Right, right.

Matt Turck [50:54] It's brand new.

Aidan Gomez [51:37] Yeah, hopefully we'll GA that product quite soon, and I'll be able to say more about it. But yeah, you can imagine it as, like, a fully private version of your favorite consumer chatbot, with reasoning natively supported, but much more customizable. So you can plug in literally anything. If Accenture or Deloitte built your supply chain team a custom piece of software, you can integrate that into North, and the model can start using that software to accomplish tasks. And it does a lot of interesting work on deep research and this type of spending time thinking, researching things, sifting through data, taking five, 10, 20, 30 minutes to accomplish a task.

Aidan Gomez [52:05] It does that extremely well because of all the data it has access to. It's not just web search; it's sifting through your entire organization.

Matt Turck [52:28] Is part of the idea that you're going to evolve towards multi-agent workflows, which seems to be the cool thing to talk about on Twitter in 2025? And if so, what would be some examples of what an agentic workflow looks like? We have to take a shot, by the way, each time we say the word agentic. That's the rule of podcasts.

Aidan Gomez [52:53] Yeah, I'll be asleep by the end of it. So, for use cases, like, right now, the coolest stuff is usually in time-sensitive areas. So there's one really cool use case, which I think is going to have a huge, huge, huge, huge impact on the economy, which is, like, doing research for folks in finance, for example, wealth managers, when an event breaks. So what happens today is, if I'm a wealth manager, I'm managing, like, 10 to 15 clients, a war breaks out somewhere, or there's some announcement about a new tariff, I get calls from all those clients saying, "Oh my God, what are we going to do?"

Aidan Gomez [53:44] And I have to spend maybe a week, maybe four weeks researching and coming up with a hedge proposal. So I come up with a portfolio hedge, and then I need to implement that across my client base. At that point, like, a week, four weeks, the world has changed, the market has dropped 20%, the conflict is resolved, whatever. And so the velocity and the importance of being able to act with speed in that space is crucial. And so that's what these agents can do.

Aidan Gomez [54:15] They can help do that job dramatically faster. That research, they can come up with a proposal way faster than a human ever could because it can read 100,000 times faster than a human. So we can go out, read analysis, read articles, and come back with a concrete proposal of what to do. And then the human can take over and make whatever edits they want to make, and then take that forward. So we can take something that used to be a month and bring it down to four hours, eight hours.

How Enterprises Identify and Implement AI Use Cases

Matt Turck [55:05] That's a fascinating example. The beauty and the curse of this whole agentic, sort of horizontal kind of possibility, the fact that even in the enterprise, agents can do all those things, feels like both a blessing and a curse, meaning that the risk is that customers could be just overwhelmed by the possibilities and not know where to start. How does that work in practice? Do you find yourself effectively doing consulting and sitting down with customers to help them come up with use cases? Or do some people know exactly what they want to do?

Matt Turck [55:12] What's the reality of that part?

Aidan Gomez [55:37] It used to be much more like that, where you'd come in and they'd say, "Hey, this AI stuff is super cool. My board is giving me a bunch of pressure. What am I going to do about GenAI? What should I do?" And we would have to help them think through the opportunity space. We'd have to learn about their business, try to help them identify the opportunity. But increasingly, it's actually changed dramatically, and the competency of the customer is much higher.

Aidan Gomez [56:01] They know exactly what they want to do. They know their business. They know what will count. And they just want to go implement it. And so they need a partner to help them accomplish that roadmap that they've decided on. So there's that phase of POC or figuring things out. It feels like it's passed us by now. Most organizations know the opportunity, they know what they want to do, and they really just need help to go execute on it.

Matt Turck [56:42] The $2 billion we collectively spent on Accenture did pay off. Now our customers are ready to truly move forward. Okay, that's fantastic news for the industry. Particularly interesting because, again, the agentic aspect of this actually, in some ways, makes it more complex because you have more possibilities. So it's fascinating to hear that people actually, despite the scope of this being widened, have a more precise idea of what it is that they want to do. Okay, great.

Aidan Gomez [57:10] Yeah, it's not like they come with only one idea. So the scope is still broad. We've had multiple large enterprises come to us and say, "We have 700 use cases that we've identified. Can you help us accomplish them?" It's like, "Okay, yeah, we can. But let's start on the most important five, and then expand from there." But with a company like Oracle, which has all of this workplace software in Fusion Apps and NetSuite, they've implemented hundreds of use cases themselves using Cohere's models.

The Competitive Edge of Early AI Adoption

Aidan Gomez [57:49] And so the progress for the early adopters, it is pretty staggering. Like, it's incredible how far this technology is reaching. And of course, it's in the hands of hundreds of millions, potentially over a billion now, of consumers. But for workers, for employees, it's very rapidly approaching similar scales.

Matt Turck [58:16] And you're starting to see a breakout between the companies that adopted this early, did the work, joke aside, brought in Accenture, and are starting to deploy this at scale, and the laggards following the typical adoption cycle, because that's something that we as an industry have been, like, collectively warning, quote-unquote, let's say, the Global 2000 that would happen. Are you seeing it happen?

Aidan Gomez [58:43] I definitely think so. I think there's advantage being conferred to those who have access, who have adopted early and given their employees this augmentation, and whose employees are, by the day, getting more and more competent at integrating models and agents and AI into their work. The organizations that have the workforce best capable of doing that, they're gonna win.

Matt Turck [58:51] What are agents not quite ready for? What would you advise customers to not do?

Aidan Gomez [59:19] Well, there's all the sensitive use cases, like medicine and these sorts of places, lots in finance. For those, you want a human in the loop. You don't want to just hand it over to an agent and let the model go crazy. So there's places where it's not ready just because we need good oversight, and it may never be ready. There's a large swath of things that we always want a human in the loop for. In terms of technological limitations where the model is not yet smart enough to accomplish something, I think one of the things that's said quite often now and is a really good example of how far the bar has been raised is people are saying these models haven't discovered new science.

Aidan Gomez [1:00:07] They haven't solved some millennium problem or something like that. That is somewhere that the models are not yet super helpful. If you're a postdoc, it might be useful to your productivity of reading papers and preparing talks and that type of thing. But actually helping you discover a new compound, I'm not sure how useful it is yet, but I'm very confident it will be useful very soon.

Aidan’s Concerns About AI and Society

Matt Turck [1:00:28] It's been an awesome conversation. Maybe to close, zooming out, what keeps you up at night? Progress in, or lack of progress in AI research, or the world moving in a specific direction, or micro problems that you deal with every day at Cohere as a CEO?

Aidan Gomez [1:01:03] I think for me, what keeps me up at night is politics. I think it's the fracturing that we're seeing around the world. And of course, those have implications on me and the business, but I'm mostly concerned about our societies, liberal democracies. I'm afraid for liberalism and the progress that's been made over the past century. I'm afraid for that. So that's mostly what keeps me up at night, which is not a technical answer. But I'm really optimistic about AI. I think it can actually be a big force for good in making sure the good guys win.

Cohere’s Vision for Success in the Next 3–5 Years

Aidan Gomez [1:01:30] And making sure that some of the economic issues the world has been facing over the past 15 years, in terms of slow productivity growth, stagnation, wealth not reaching the population evenly, I'm optimistic that AI can play a role in helping to resolve some of that.

Matt Turck [1:01:35] Success for you guys over the next three years, five years, what does that look like?

Aidan Gomez [1:01:59] I want to see GDP-impacting productivity gains. I want to see this technology integrated across the globe, and I want it to become a part of everybody's workday, not just their fun time or searching up stuff. I want it to help people accomplish more, and I want this technology to make stuff much cheaper, much more abundant.

Matt Turck [1:02:02] Aidan, terrific. Thank you so much for doing this.

Aidan Gomez [1:02:02] Thanks so much.

Matt Turck [1:02:23] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.