The SyFI Lab at the University of Washington builds efficient and resilient infrastructure for the future of AI. As applications grow more complex, we bridge the gap between next-gen models and heterogeneous hardware through cross-stack innovation, delivering scalable, open-source systems validated by industrial partners.

Our research targets three key areas:

  • Efficient AI: Optimizing algorithms and systems to maximize performance for training and inference.
  • Flexible AI: Architecting systems that seamlessly adapt to diverse tasks, strategies, and model structures.
  • Resilient AI: Ensuring AI system reliability at scale while leveraging AI to improve infrastructure robustness.

Publications

Ekka: Automated Diagnosis of Silent Errors in LLM Inference

Yile Gu, Zhen Zhang, Shaowei Zhu, Xinwei Fu, Jun Wu, Yida Wang, Baris Kasikci — International Conference on Machine Learning (ICML) (2026)

PDF
TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Kan Zhu, Mathew Jacob, Chenxi Ma, Yi Pan, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci — (2026)

Piper: A Programmable Distributed Training System

Megan Frisella, Shubham Tiwari, Andy Ruan, Yi Pan, Parker Gustafson, Mat Jacob, Gilbert Bernstein, Stephanie Wang — (2026)

Read more »

Blog Posts

VibeSys Builds a Qwen3.5-397B Engine on MI300A, 2.3× Tuned SGLang

October 08, 2026

VibeSys's agents built a serving engine from scratch for Qwen3.5-397B-A17B on four AMD MI300As, taking it from 12.8 to 2,242 tok/s, 2.33× the goodput of an SGLang that we first had to port and tune.
Introducing ServingStudio: An Integrated Workbench for Simulating, Analyzing, and Optimizing LLM Serving Systems

September 24, 2026

ServingStudio's Simulator predicts LLM serving performance and analyzes execution costs from measured GPU kernel timings. Its Agent uses these results to implement promising changes in real serving frameworks and validate them on hardware.
Breaking Down Bash Commands in TraceLab

July 24, 2026

TraceLab V2 looks inside coding agents' Bash calls, comparing which executables Claude and Codex invoke and where those commands spend their time.
Read more »

Talks

Enabling ASIC AI Chips for Cryptography Primitives

June 12, 2026 Jianming Tong — Georgia Tech

ML for Systems in the Real World: Challenges Beyond Training the Models

May 29, 2026 Jiayi (Jane) Chen — The University of Texas at Austin

Extreme Codesign for Efficient AI: Algorithms, Software, and Hardware

May 22, 2026 Mohamed Abdelfattah — Cornell University

Read more »