Todd Underwood is a Director at Google, where he leads Machine Learning SRE and is also Site Lead for Google’s Pittsburgh office. ML SRE teams build and scale internal and external machine learning services, and they are critical to almost every significant product at Google. In that role Todd works at the frontier of a relatively new discipline: applying the rigor of site reliability engineering to systems whose behavior depends on models, data, and training pipelines as much as on code.
He is a co-author of Reliable Machine Learning, published by O’Reilly with Kranti K. Parisa, Cathy Chen, and Niall Murphy. The book distills lessons from running ML systems in production, covering how teams can make machine learning systems as dependable as the rest of their infrastructure. It has become a key reference for engineering leaders building reliability cultures around ML.
Todd’s career in operations and infrastructure runs deep. Before Google, he held a variety of roles at Renesys, an Internet intelligence company now part of Oracle’s Cloud, where he was in charge of operations, security, and peering. Before that he was Chief Technology Officer of Oso Grande, an independent internet service provider in New Mexico. That foundation in running critical internet infrastructure informs his approach to keeping machine learning, one of today’s most critical workloads, reliable at scale.
Todd Underwood