Managed Service Engineering Lab

Department of Computer Science, The University of Texas at Austin

Managed Service Engineering Lab

We study how to make software programmable, observable, and governable when it is delivered as a service.

Each square is one running copy of an application, used by a different customer, team, or student, each needing to be isolated, operated, and audited on its own.

What is managed service engineering?

More and more software reaches its users as a service: the provider does not just ship it, but runs it for them. That means creating many isolated copies of the same software and keeping every one of them monitored, upgraded, secured, and accounted for.

Managed service engineering is the discipline of building the systems, abstractions, and tools that make this manageable for any software, not just for applications with a hand-built control plane. Our central idea is to wrap software in an automatically generated API: a programmable control surface that spans both what the software is and every instance of it that is currently running.

Many of our problems come from real deployments of cloud-native applications in web services, AI workloads, geospatial systems, and distributed enterprise platforms, often with industry partners. The lab develops and maintains open-source systems, including KubePlus and KubeProvenance, that serve both as research artifacts and as platforms for student projects. Students gain experience building real distributed systems, contributing to open-source communities, and working on problems grounded in industry practice.

Programmable services

Automatically generating control surfaces for applications, and using them to deliver multi-instance multi-tenancy: isolated instances per tenant across compute, storage, and network.

Work: KubePlus

Composable platforms

Designing Operators and platform components that work well alongside each other on shared clusters, and composing application-specific platforms from them declaratively.

Work: Operator Maturity Model, Platform-as-Code

Observable, intelligent operations

Using operational evidence and AI systems to troubleshoot, find root causes, and carry out day-2 tasks, while preserving reliability, auditability, and human oversight.

Work: AI-assisted platform operations

Governable services

Access control, provenance, policy enforcement, and service tiers, so that providers can answer who did what, to which instance, and when.

Work: KubeProvenance, provenance for PaaS and key-value systems

Projects

Active
2020 to present
750+ GitHub starsProgrammable services

KubePlus: multi-instance multi-tenancy on Kubernetes

KubePlus is an open-source Kubernetes Operator that turns an application's Helm chart into a new API for creating and managing isolated instances of that application. Each tenant gets its own instance, with isolation across compute, storage, and network resources, and every instance supports monitoring, license management, and upgrades through the same API.

KubePlus has been adopted by engineering teams at Atomic Maps and Verizon Communications, and has delivered per-student CI/CD instances in UT's cloud computing courses.

Active
2024 to presentObservable, intelligent operations

AI-assisted platform operations

We investigate how AI systems, particularly local LLMs, can support operating managed services: troubleshooting, root-cause analysis, and day-2 tasks. A guiding question is how much information an AI agent actually needs to diagnose a failing instance, and how to keep its actions reliable, auditable, and under human oversight.

This work is connected to the lab's participation in the CNCF AI-Conformance Working Group, which examines how AI capabilities can be standardized in cloud-native environments.

Active
2018 to presentGovernable services

KubeProvenance: operational provenance tracking

KubeProvenance is a declarative model for tracking and querying provenance information about actions performed on Kubernetes APIs. It supports auditability and governance in managed service delivery, and informs current work on access control for managed services.

It builds on the lab's earlier research on provenance for key-value systems (TaPP'13) and Platform-as-a-Service environments (HotCloud'15).

Active
2020 to presentComposable platforms

Operator Maturity Model and Platform-as-Code

Managed services on Kubernetes are built from Operators, and a real platform runs many of them side by side. The Operator Maturity Model defines objective criteria for Operators that behave well in a multi-Operator cluster. Platform-as-Code builds on it with declarative mechanisms for composing application-specific platforms and workflows from these components.

The broader goal is to move platform construction from ad hoc engineering toward a systematic, programmable discipline.

Past
2012 to 2015Governable services

Provenance for PaaS and key-value systems

Studied provenance in Platform-as-a-Service environments and key-value stores. Developed the Key-Value Provenance Model (KVPM), which supports both data-level and schema-level provenance, and proposed mechanisms for collecting provenance at the PaaS level. This work laid the foundation for the lab's current research on governable managed services.

Questions we'd like to work on with you

Managed service engineering is a systems problem, but many of its open questions sit at the edges of other fields. We welcome collaborators from across UT.

Computer Science and AI

Agentic managed services

What is the minimum information an AI agent needs to find the root cause when a software instance fails?

Statistics and Data Science

Signal in operational data

How much telemetry is enough for diagnosis? What can causal inference over logs and metrics, or analysis of metering data, tell us about running services?

Information

Provenance and governance

How do we capture and query who did what, to which instance, and when? How should service tiers be defined, enforced, and changed?

Faculty, postdocs, and students interested in any of these questions can reach the lab at devdatta@cs.utexas.edu.

Publications

Peer-reviewed industry conference presentations

2020
Being a Good Citizen of the Multi-Operator World
Kulkarni, D.
KubeCon North America, Cloud Native Computing Foundation
2019
Operators and Helm: It Takes Two to Tango
Kulkarni, D.
Helm Summit, Cloud Native Computing Foundation
2016
Application CI/CD on OpenStack: Building a Solution Using Jenkins and OpenStack Solum
Kulkarni, D., & Jain, A.
OpenStack Design Summit, Austin, TX

Workshops and posters

2015
Provenance Issues in the Platform-as-a-Service Model of Cloud Computing
Kulkarni, D.
USENIX HotCloud'15, Santa Clara, CA
2013
A Provenance Model for Key-Value Systems
Kulkarni, D.
USENIX TaPP'13, Lombard, IL
2013
A Fine-Grained Access Control Model for Key-Value Systems (poster)
Kulkarni, D.
ACM CODASPY, San Antonio, TX
2006
Exception Handling in CSCW Applications in Pervasive Computing Environment
Tripathi, A., Kulkarni, D., & Ahmed, T.
Advanced Topics in Exception Handling Techniques, Springer LNCS 4119
2004
Context-Based Secure Resource Access in Pervasive Computing Environments
Tripathi, A. R., Ahmed, T., Kulkarni, D., et al.
PerCom Workshops, 159–163

Peer-reviewed journals

2012
A Generative Programming Framework for Context-Aware CSCW Applications
Kulkarni, D., Ahmed, T., & Tripathi, A.
ACM Transactions on Software Engineering and Methodology (TOSEM), 21(2), 11:1–11:35
2010
A Framework for Programming Robust Context-Aware Applications
Kulkarni, D., & Tripathi, A.
IEEE Transactions on Software Engineering (TSE), 36(2), 184–197
2007
Autonomic Configuration and Recovery in a Mobile Agent-Based Distributed Event Monitoring System
Tripathi, A., Kulkarni, D., et al.
Software: Practice and Experience, 37(5), 493–522
2005
A Specification Model for Context-Based Collaborative Applications
Tripathi, A., Kulkarni, D., & Ahmed, T.
Pervasive and Mobile Computing (PMC), 1(1), 21–42

Peer-reviewed conferences

2009
Resource-Aware Migratory Services in Wide-Area Shared Computing Environments
Tripathi, A., Padhye, V., & Kulkarni, D.
IEEE Symposium on Reliable Distributed Systems (SRDS)
2008
Context-Aware Role-Based Access Control in Pervasive Computing Systems
Kulkarni, D., & Tripathi, A.
ACM Symposium on Access Control Models and Technologies (SACMAT), 113–122
2008
Application-Level Recovery Mechanisms for Context-Aware Pervasive Computing
Kulkarni, D., & Tripathi, A.
IEEE Symposium on Reliable Distributed Systems (SRDS), 13–22
2002
Dynamic Network Information Collection for Distributed Scientific Application Adaptation
Kulkarni, D., & Sosonkina, M.
International Conference on High Performance Computing (HiPC), 555–563
2002
A Framework for Integrating Network Information into Distributed Iterative Solution of Sparse Linear Systems
Kulkarni, D., & Sosonkina, M.
VECPAR, 436–450

People

Devdatta Kulkarni

Devdatta Kulkarni

Lab Director; Assistant Professor of Instruction, Computer Science

Researcher, educator, and technologist working on managed service engineering, multi-tenancy, and provenance for distributed systems. Founder of CloudARK and lead developer of KubePlus. Active contributor to the Cloud Native Computing Foundation.

Current students

Harshith SadhuUndergraduate research, Spring 2026 to presentAI-enabled managed service delivery on Kubernetes
Bryan ZhaoUndergraduate research, Spring 2026Support for GitOps in a multi-instance managed services framework

Alumni

Annie HuSpring 2026Access control in managed services for cloud-native applications
Pranav VenkateshMS, 2025Scaling Kubernetes with LLMs: AI-assisted DevOps techniques
Tony Nguyen2025Application-specific metrics for multi-tenant cloud-native applications
Om Goswami2024Application upgrades in multi-instance multi-tenancy; KubePlus CLI plugin
Yizhou Li2023CLI plugins for provider and consumer permissions in KubePlus
Emil Arslan2023Improved API registration mechanisms in KubePlus
Daniel Moore2018Provenance tracking for CRUD operations on custom resource types in Kubernetes
Mohamad N El-Zein2018Evaluation of Operator development frameworks (Operator SDK vs. Kubebuilder)

Partners

Industry

Founding collaboration

CloudARK

CloudARK builds a managed service delivery platform for cloud-native applications on Kubernetes and was founded by the lab director. The lab's research informs CloudARK's platform, and operational challenges seen at CloudARK motivate key research questions in multi-tenancy and managed service delivery.

Deployment partner

Verizon Communications

Engineering teams at Verizon Communications have adopted KubePlus in application-specific platform deployments, providing real-world validation and feedback for the lab's research on multi-tenant service delivery.

Deployment partner

Atomic Maps

Atomic Maps uses KubePlus for application-specific platform deployments, helping the lab understand operational challenges in geospatial and data-intensive cloud-native workloads.

Open source and community

Open-source community

Cloud Native Computing Foundation (CNCF)

Active participation in the Kubernetes Multi-tenancy Working Group, including discussions on multi-tenant platform design and the Kubernetes multi-tenancy manifesto, and in the CNCF AI-Conformance Working Group, which examines how AI capabilities can be standardized in cloud-native environments.

Open-source community

OpenStack

The lab director previously served as Project Team Lead (PTL) for OpenStack Solum, the open-source Platform-as-a-Service initiative in the OpenStack ecosystem. Contributions to CI/CD pipeline design and PaaS abstractions continue to inform the lab's research.