University of Wisconsin–Madison

Matt Sinclair receives Department of Energy Early-Career Award

By Karen Barrett-Wilt

Matt Sinclair headshot
Matt Sinclair

Matt Sinclair, assistant professor in the University of Wisconsin–Madison Department of Computer Sciences, has received a US Department of Energy Early Career Award in recognition of his pioneering research in computer architecture and high-performance computing (HPC) systems. The honor, which supports exceptional early-career faculty poised to make significant impacts in their fields, is among the most competitive federal research awards given to scholars in many fields, including computer scientists. Sinclair is the first person in the Computer Sciences department to be given this distinction. 

Sinclair’s research explores how to make large-scale computing systems — those that power everything from scientific simulations to machine learning breakthroughs — more efficient and reliable. As these systems increasingly rely on accelerators such as Graphics Processing Units (GPUs) to deliver performance, they face a growing challenge: performance variability, where identical components behave differently in ways that may waste energy and hurt performance. Sinclair’s work embraces that variability instead of fighting it, turning a growing problem into an opportunity for smarter, more efficient computing. 

The award will provide five years of support for Sinclair’s research and education efforts, helping him expand his work with graduate and undergraduate students at UW–Madison.

Computer Sciences Chair Paul Barford is “thrilled for Matt.” He continues, “Early Career awards are highly prestigious and very competitive. This award is a reflection of the great work that Matt is doing, which will very likely have long term impacts in his field and in society.”

From cooking to computing 

For non-computer scientists, Sinclair explains his work using a simple analogy. “Imagine you and a friend are following the exact same recipe, but you’re cooking in different kitchens,” he says. “You’re cooking in Denver at a high elevation, and your friend is cooking the same dish near sea level. Even though you’re following the same steps, there are many factors that change how the dish turns out. For example, did you measure the exact same amount of ingredients like salt?” Sinclair explains that environmental factors — like elevation or temperature — can change how the dish turns out. “Computers are the same way: even when we set everything up the same, their ‘environment’ — things like temperature, cooling, and how power is managed — affects how they perform,” he says. 

That variability, Sinclair notes, matters even more at scale. In modern HPC clusters, hundreds or thousands of accelerators work together on the same problem. If one of them slows down, the rest must wait, like a delayed airplane that causes further delays and issues down the line. This “straggler effect” leads to longer run times, lower system utilization, and higher power costs. 

“What my research group and I are doing,” Sinclair says, “is finding ways to understand, predict, and even leverage that variability. Instead of treating it like a problem to eliminate, I’m designing mechanisms that harness it to our advantage.” 

Embracing imperfection for efficiency

Instead of treating variability like a problem to eliminate, I'm designing mechanisms that harness it to our advantage.

The key insight of Sinclair’s research is that differences in performance are inevitable, so systems should be designed to account for them. His group is developing techniques that span applications, hardware, and software, from the chip level, to the scheduling software that decides where work should run, to the applications themselves.

By collecting fine-grained data about each accelerator’s performance and behavior, Sinclair’s team can design “variability-aware” mechanisms that group similar accelerators together. This way, workloads that need to synchronize frequently are less likely to be held up by slower and/or hotter devices. In the longer term, his work could influence how future systems are designed — integrating awareness of variability directly into chips and system architectures.

At the hardware level, Sinclair’s team also studies how to let accelerators temporarily exceed strict power limits in a safe, controlled way, akin to letting an oven run slightly hotter for a moment without burning dinner. This approach allows the system to safely use available headroom more intelligently, further improving performance and energy efficiency.

Real-world impact

While Sinclair’s work is deeply technical, its implications reach far beyond computer architecture. National laboratories, large technology companies, and scientific researchers all depend on large-scale computing systems to advance discovery in fields such as climate modeling, plasma physics, and artificial intelligence.

“The goal is to make these systems not just energy efficient, but smarter,” he says. “By improving efficiency, we can accelerate scientific progress and improve energy efficiency of computing at scale.”

To help ensure this work impacts the relevant stakeholders, Sinclair is collaborating with partners across the computing ecosystem, including hardware vendors, national laboratories, and application developers. “It’s important to have all three perspectives represented,” he says. “The people running the systems, the people using them, and the people building the hardware all stand to benefit from managing and leveraging variability.”

Sinclair’s colleague CS Associate Professor Shivaram Venkataraman notes that  “this award is a great indicator of the importance and impact of Matt’s work. His work on large accelerator rich clusters will directly impact large ML deployments for researchers like me, who are working on machine learning systems.”

For Sinclair, the recognition affirms the importance of tackling foundational problems with broad implications. “It’s an incredible honor,” he says. “This award gives us the opportunity to push the boundaries of what’s possible in computing — and to train the next generation of researchers who will carry that work forward.”

________________________________________

Matt Sinclair is an assistant professor in the Department of Computer Sciences at the University of Wisconsin–Madison. His research primarily focuses on how to design, program, and optimize future heterogeneous systems. He also designs the tools for future heterogeneous systems, including serving on the gem5 Project Management Committee and the MLCommons HPC, Power, and Science Working Groups. His research has been frequently recognized by prestigious organizations, including an ACM Doctoral Dissertation Award nomination, a Qualcomm Innovation Fellowship, the David J. Kuck Outstanding PhD Thesis Award, and an ACM SIGARCH–IEEE Computer Society TCCA Outstanding Dissertation Award Honorable Mention.