How Data Happened - A Conversation with Chris Wiggins, Author and Chief Data Officer at The New York Times

The MAD Podcast with Matt Turck · with Chris Wiggins, Chief Data Officer, The New York Times

Chris Wiggins is the Chief Data Officer at The New York Times. We cover how statistics originally meant statecraft rather than mathematics, why early digital computers were built to process messy codebreaking data at Bletchley Park, and how deep neural networks decisively beat other image-labeling methods in 2012.

Watch on YouTube

Chapters

  1. 0:00 — Full episode

Transcript

Full episode

Matt Turck [1:01] Chris, welcome, or I should say welcome back. You are a frequent guest at this event. And I think we first had you in 2013 and then 2015. I forget. You might actually be the most frequent guest I guess we ever had.

Chris Wiggins [1:01] Amazing.

Matt Turck [1:18] Yes, but always love the conversation. Welcome back. So you are a man of many talents. You are the Chief Data Scientist at The New York Times. You are an associate professor of computational math.

Chris Wiggins [1:19] Applied mathematics.

Matt Turck [1:22] Applied mathematics at Columbia.

Chris Wiggins [1:22] Yes.

Matt Turck [1:42] And you are also a prolific book author. And we are going to spend most of the conversation today talking about your new book right here, How Data Happened. Maybe towards the end, we'll talk about The New York Times and data science and machine learning at The New York Times, but we're going to spend most of the time on this. So to jump right in, what is the book about, and why is it an important book to write about and read in 2023?

Chris Wiggins [2:14] Yes, so this book attempts to situate data and how it came to pass that so much of our lives are shaped by data-empowered algorithms—our personal realities, our professional realities, our political realities as well. And it's really the result of my own journey. I grew up as a physicist and then started working in biology, and then biology became a data science back in a previous millennium. And now, as Chief Data Scientist at The New York Times, I've tried to understand: how did it get that way?

Chris Wiggins [2:34] And the it includes data science, machine learning, statistics, artificial intelligence, and the general dream for centuries of trying to put data to work to understand our world and understand our society.

Matt Turck [2:39] And the book was co-written with Matt Jones, who was one of your colleagues at Columbia.

Chris Wiggins [2:58] Yes, I am a mere fan of history. I think of history as one of the best ways of understanding root causes, but Matt is an actual historian. So I met Matt about a decade ago at Columbia when he was giving a talk on the history of machine learning. And we've been collaborating on a few projects over the years, including a class. The class was very difficult to sculpt, right, because you have to take some subject and carve it into 13 chunks.

Chris Wiggins [3:15] But once we had the history of data in 13 chunks, we realized that could also be a book, a general-interest book just for anyone who's curious about data, about data and how it got that way.

Matt Turck [3:35] So to start from the beginning, the book starts at the end of the 18th century or beginning of the 19th century, and you make the point that data, the earliest form of data, was used by states. Tell us more about that. What did states do, and how did it all start?

Chris Wiggins [3:56] Yes, but no. The book doesn't actually start in the 19th century, right? The book starts in the classroom. So the book starts with me trying to convince the students that understanding data and how it got that way would be relevant for understanding their present day, and I'm trying to explain some mathematical concepts, and one of the students raises his hand and says, "Can we talk about Facebook now?" And the reason he said, "Can we talk about Facebook now?" is because five years ago, Mark Zuckerberg was in front of Congress testifying.

Chris Wiggins [4:24] In fact, today, Sam Altman is testifying in front of Congress in exactly the same way. Society was concerned. Where did these algorithms come from? How did it get this way? And so five years ago, I was teaching this class, and students wanted to understand, how do we understand how it got that way, how it came to pass that we've got these algorithms shaping our personal and political realities? And the claim of the book is that it's helpful to understand the present day, to look at that arc and understand 200 years of people trying to make sense of the world and society through data.

Chris Wiggins [4:57] But to get back to Matt's question, yes, one of the things we try to do is to try to help people understand why these words are so confusing and malleable and used by different communities over different times. Why is it that your investor might ask you for your AI strategy today, but would have asked for your ML strategy six months ago, and maybe in a different universe would have been asking for your data science or big data or statistics strategy? And each one of those things is a chapter in the book, by the way, as well as a chapter about venture capital.

Chris Wiggins [5:05] I don't know if you got to that chapter yet.

Matt Turck [5:07] Of course, I read it all.

Chris Wiggins [5:31] Okay, brilliant. That's part of what the book is trying to help people make sense of. And statistics in particular, people forget that statistics entered the English language to mean statecraft. It had nothing to do with math, and it certainly had nothing to do with data. It was about trying to make sense of the state. And trying to use data to make sense of the state was considered vulgar statistics, or mere table statistics. And that fight is one of the examples of a fight from 200 years ago that I think is illustrative of fights of this present day.

Chris Wiggins [5:54] Because even today, you can go into some community where they understand a thing, like picking wine, or rating movies, or rating baseball players, and then somebody shows up and says, "I have data about that thing, and data should have a seat at the table for understanding that thing," and you get fights. You get fights between people who think that there's a way of quantifying knowledge, and other people who think that that is mere vulgar treatment to understand knowledge, and a higher statistics, for example, is to understand the greatness of the men who rule the different countries, which is the view from 1806, for example.

Matt Turck [6:19] And fast forward through that early phase, when did data become math?

Chris Wiggins [6:48] Yeah, so there's a chapter on data's mathematical baptism, where data takes on the sacredness of the academy, and in particular, this scientific way of knowing things by applying mathematics to it. That chapter opens up with the hottest IPO in the late 19th century, which was Guinness. So Guinness, the beer company, IPO'd in late 1886, and literally people were breaking the doors down to try to get in on that IPO. Guinness had all of the money, and so they could afford the hottest tech of the day, and they hired the hottest nerds of the day, who were the statisticians, except they called them brewers.

Chris Wiggins [7:04] Brewer was the great title, like Chief Data Scientist of the late 19th century.

Matt Turck [7:05] The sexiest job of the 19th century.

Chris Wiggins [7:29] Right, actually, I just got to meet Hal Varian yesterday and thanked him for that line. Yes, so Guinness actually was the industrial use of putting data to work. And then over the next 25 years, it became a concern of the academy when it was clear that there was really a lot of profit to be made in understanding data. It became a concern not only of companies, but also of states that were trying to maximize agricultural output. And not long thereafter, it became a branch of mathematics.

Chris Wiggins [7:53] And that itself was a big transformation, that the way you make sense of data was not just about getting your hands on a lot of data or looking back at what happened in the past, but actually using mathematics to try to understand data. And that fight about how to make sense of data and how to use data to make sense of what's true keeps going all the way through World War II. And then we turn to Bletchley Park and the birth of computation.

Matt Turck [8:13] And around that time, you start seeing some of the dark turns of data already, right, at the hand of Darwin's cousin, Galton, that starts using data and stats to tell the story he wants to tell.

Chris Wiggins [8:34] Yeah, I'm thinking, is that one of the first dark turns? There's so many dark turns in data. So yes, that chapter is one of the first sort of obviously dark turns of data because the creators of mathematical statistics in the late 19th century were trying to make their empire great again, and they were clearly thinking about the role of data in not only understanding society but shaping it. They really wanted to not only understand the world, but to drive it using data.

Chris Wiggins [9:04] And so there was a lot of mathematics around policy decisions, like trying to understand, well, what causes poverty in England, for example? And they really did it using a form of supervised learning. We would recognize it today. But in particular, one of the founders of mathematical statistics we look at is Sir Francis Galton, a distant cousin of Charles Darwin, who gives us the word regression, gives us the word correlation, and gives us the word eugenics. And we include that story not because it's like shooting fish in a barrel to point out that these people in Victorian England were advancing eugenics, but because they weren't writing about themselves like, "We're the baddies, and we really want to oppress the crap out of people."

Chris Wiggins [9:40] They wrote about themselves like, "We're going to do a solid for society, and we're going to make society better with data." That's the thing that I think is valuable about that story, is by what right do we, other than our own biases, look back on the concerns and misuse of data under the name of eugenics, for example, and think that we ourselves are free of any prejudice? So for we data scientists, we think we're doing something that's technically sweet, to use Robert Oppenheimer's phrase, but we have to think about what is actually the impact of things.

Matt Turck [9:55] What happens during World War II with codebreaking? How is that important?

Chris Wiggins [10:15] So computers were born of a data science problem, which is a story that's not often told, and in fact was intentionally opaque in history for about 75 years. Moreover, I grew up as a physicist thinking that physics really won World War II, but now that I'm a data scientist, I realize that it was actually data science that won World War II. But that story was classified for about 75 years, which is the story of how the first digital programmable computers were created at Bletchley Park for dealing with streams of messy data.

Chris Wiggins [10:46] Any of you who deal with streams of messy real-world data will know that that's a huge pain in the ass. They had that pain in the ass in Bletchley Park, which was a remote little place in England. It's sort of the Los Alamos of England, right? It's between Cambridge and Oxford, but you can't get there from here, and so you put something secret there. Anyways, so they had to invent special-purpose digital hardware and electronic hardware in particular for solving the problem of dealing with streams of messy data.

Chris Wiggins [11:19] It's a story that's completely occult. And it was done entirely by people who were absolutely not statisticians, right? It was this mix of puzzle programmers, mathematicians, and people who worked for the telecommunications industry in England. That story had its own mirror on the other side of the Atlantic in Bell Labs, and how Bell Labs played a crucial role in scaling up codebreaking as a computational problem. That pairs Bell Labs with the nascent intelligence community, which goes on to fund IBM 701, IBM 704.

Chris Wiggins [11:34] All the beginnings of the birth of computation in the United States were really born of this messy data science problem funded by the intelligence community.

Matt Turck [11:47] And carving out AI, because we're going to talk about it in a second, bring us home from the '50s, '60s, '70s to the era of big data.

Chris Wiggins [12:05] So we break the book into three parts. Part one is really about data in the service of what is true, and that includes data's mathematical baptism. Part two was really about data in the service of engineering and problem solving, and that includes Bletchley Park, as well as a number of problems of the present day. Part three is really about our present milieu, which includes an ad economy, including the role of venture capital in accelerating that ad economy, the battle for data ethics—there's a chapter about companies trying to define ethics and design for ethics in a way that responds to actually many of the things that were mentioned in congressional testimony this morning.

Chris Wiggins [12:49] And then what are the actual forces that constrain and guide data and power? And that chapter is really about the balance, the unstable three-player game, to rip off Bill Janeway, among corporate power, state power, and people power, and how that game has a bunch of forces that are as yet unresolved, and the resolution of those contests is going to shape the future. En route, we do introduce people to artificial intelligence, which is a great story. The reason why it's such a wicked term and such a movable term is because the term itself is kind of meaningless, and it's named after an aspiration rather than the method of how you're going to get artificial intelligence.

Chris Wiggins [13:20] And we try to trace it. We try to explain to people why is that term so difficult to work with, in part because the guy who invented the term is on record as saying, "I made up the term to get money." Second of all, really, for the first half of the life of artificial intelligence, people thought it had nothing to do with data whatsoever. They thought it was really about logic and just going to find an expert and then just programming it, stating so precisely what it is that we do when we think that you could put it on a computer.

Chris Wiggins [13:49] The problem is we don't really know how we think. We think we know how we think, but the story about how data triumphed, and how data was discarded for the first half of the life of artificial intelligence as the wrong way to get intelligence and then has triumphed so much in the last 25 years, is part of what that chapter is all about.

Matt Turck [14:20] Yeah, maybe double-click on this because it's so interesting. It would be an understatement to say that we're oversaturated with generative AI stories today. Not that many people actually know the history of artificial intelligence. And obviously, that would take a whole evening to talk about it, but maybe, like, the last sort of 40 years: some key moments, key milestones that people should know.

Chris Wiggins [14:46] Well, one of the things we talk about that situates the battle to even define what even is artificial intelligence is this great paper by Herb Simon from 1984. It was a talk delivered in '83, and he gives this provocative—so a young professor named Tom Mitchell invites him, Herb Simon, the only person ever to win a Turing Prize and a Nobel Prize, invites him to speak at his conference on machine learning. And Herb Simon gives a talk called "Why Should Machines Learn?" which is basically saying, if we want to get artificial intelligence, we shouldn't do it using machine learning.

Chris Wiggins [15:06] Obviously, the way to understand artificial intelligence is via schema. So in the '80s, you get a story, which some of you have probably heard, about how artificial intelligence for a long time had nothing to do with data whatsoever, let alone neural nets. And that's part of the exciting story, is the true believers, the very few true believers, who were banging on neural nets all the way through the '80s and '90s and even 2000s, up until 2012 or so, when people realized that large neural nets, really big neural nets, could get the job done.

Chris Wiggins [15:29] So many things had to come together there, including the creation of really big computers and also really, really massive datasets.

Matt Turck [15:31] And what happened in 2012?

Chris Wiggins [15:56] In 2012, there was a particular example of a common task framework, meaning a competition. So science often advances from these sort of engineering-like projects where somebody says, here is a standardizable competition, and we're going to get a bunch of people together and see who can quantitatively win at this competition. So the Netflix Prize is one good example from 2007 or so. But another one is from the computer vision community, where they had this competition to see who could tag images successfully.

Chris Wiggins [16:23] And 2012 was the year when the winner was a very deep neural network. So the idea of a neural network, the idea that you can take processing information and make little tiny information-processing units and chain them together, was something inspired by the human nervous system, first postulated in 1943, mechanically realized in 1959. But for most of its life, derided as a ridiculous idea. And to be fair, it's really, really hard, and you need really big computers, and you need a lot of data to do it.

Chris Wiggins [16:54] But by 2012, really deep neural networks, meaning many, many layers of neural networks, had decisively beat every other method. Now, so for people in the community, it's not a big surprise that deep neural networks are really, really good at what they do. They saw it happen with image processing in 2012 and image labeling. They saw it happen with machine translation over the next few years. And certainly, we've seen it in the case of transformers and large language models since 2017, and most visibly instantiated as a really cute product last November.

Matt Turck [17:31] Great. All right, so to keep it going, hopefully that gives a flavor for this really interesting book that I've enjoyed reading, How Data Happened. I'd love to use the few minutes to zoom out and talk about The New York Times. So again, you are the chief data scientist. What does data science mean at The New York Times?

Chris Wiggins [17:55] So I started as chief data scientist at The New York Times 10 years ago. Actually, this summer is my 10-year anniversary there. So data science still means sort of a more orthodox definition of data science from 10 years ago, which is developing and deploying machine learning. So the data science team is about a 22-person team that develops and deploys machine learning for newsroom and business problems. Most of the projects are things that are relevant to many different companies, certainly to many subscription companies, like machine learning that actually controls the paywall, that decides when you should be asked to become a paying subscriber, recommendation engines, which is not just personalization but also identifying what's trending and then serving it in a variety of different surfaces.

Chris Wiggins [18:36] Fancy ad products, so we can create advertising that's useful to marketers but is also privacy-forward. Marketing, so when we market on other advertising platforms, that is done not using guessing and pointing and clicking, but using Python and optimization. That's a variety of things. We have a couple of things that are editor-facing to help editors understand the relationship between stories and how they're promoted and how people engage with the stories, but a variety of problems that are all about developing and deploying machine learning.

Chris Wiggins [18:52] Very good.

Matt Turck [19:08] And as a heads up, I'll open up to questions in, like, a couple of minutes if anybody has a question in mind. LLMs, is that something that you guys at The New York Times think about or use or plan on using?

Chris Wiggins [19:32] I definitely have nothing publicly to disclose about any of that. I would say that the people in the data science team are generally aware—and actually, not just in the data science team. The truth is, people who are at The New York Times read the papers. So I would say everybody I've talked to, not just the data science team, is well aware of LLMs and is thinking about what it means both for The New York Times and beyond. There's obviously many different potential use cases, right?

Chris Wiggins [20:01] It's a technology. And we open up the book by quoting Kranzberg's First Law: technology is neither good nor bad, nor is it neutral. A technology is a capability, and that capability is made mobile, and different people can use that capability for different things. So The New York Times has a variety of potential use cases for this. It's also been very useful because it's sort of turned everybody's attention to artificial intelligence writ large, and how machine learning more generally can be useful.

Chris Wiggins [20:26] So language models have been around for a while, and natural language processing has been around for a lot longer. So it's also opened up to people the curiosity to think, oh, how could I use natural language processing in new ways, generative artificial intelligence as well as just natural language processing itself in my process? So it's opening up a lot of really interesting and creative conversations.

Matt Turck [20:41] To the extent you can, can you talk about the tech stack for the data science team and data team in general at The New York Times? In terms of, I don't know, data warehouse tools that one uses?

Chris Wiggins [21:09] Yeah, that's been a wonderful journey. So when I showed up at The New York Times in 2013, if you wanted to get your hands on data, you needed to write your own MapReduce jobs in Hive and hit buckets of unstructured JSON sitting in S3. Then it was decided that we should build our own Hadoop on-premises, which was the style of the time. Then all of that went away. And through a story that we don't have time to go into, we started kicking the tires on GCP, Google Cloud Platform.

Chris Wiggins [21:43] And at this point, we have fast, reliable SQL access via Google Cloud and via BigQuery, which has made life so much less painful. That said, there is also a lot of work being done in AWS and plenty of developer work happening on Amazon's cloud. So the data stack is—in my team, the data stack is SQL and scikit and occasionally Go. So scikit-learn is a particular module, and Python, where most of the machine learning you're going to want to do is already done.

Chris Wiggins [22:11] A lot of containerization. We still rely heavily on BigQuery, because a lot of times we want to score something and put it in a table so that the analysts have fast, reliable SQL access to the output of those models. I think that's about it. Occasionally, we code in Go when things really need to be performant.

Matt Turck [22:16] Chris, thank you so much. Really appreciate it. This is so wonderful. Thank you.

Chris Wiggins [22:26] Thank you, Matt. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. FirstMark.com/events/data-driven.