LASR Colloquia - John Wilkes/Google, "Cluster management at Google"

Contact Name: 
Leah Anderson-Wimberly
GDC 6.514
Jun 25, 2013 10:00am - 11:00am

Talk Audience: UTCS Faculty, Graduate Students, Undergraduate Students, and Outside Interested Parties

Host:  Michael Walfish and Michael Dahlin

Talk Abstract: Cluster management is the term that Google uses to describe how we control the computing infrastructure in our datacenters that supports almost all of our external services. It includes allocating resources to different applications on our fleet of computers, looking after software installations and hardware, monitoring, and many other things. My goal is to present an overview of some of these systems, introduce Omega, the new cluster-manager tool we are building, and present some of the challenges that we're facing along the way. Many of these challenges represent research opportunities, so I'll spend the majority of the time discussing those. 

Speaker Bio: John Wilkes [] has been at Google since 2008, where he is working on cluster management and infrastructure services. He is interested in far too many aspects of distributed systems, but a recurring theme has been technologies that allow systems to manage themselves. In his spare time he continues, stubbornly, trying to learn how to blow glass.