Schedule

Class calendar

To see the full list of papers and their details, please refer to the Reading List tab.

Explanations of the presenter, advocate, and critic group roles are available in the Resources tab.

This schedule is subject to change and currently serves as a placeholder while we finalize discussion-group assignments and a few remaining class activities.

Date Classroom Agenda Slide deck Presenter Group Advocate Group Critic Group
August 25
Instructor lecture - course overview
Slide deck
August 27
Instructor lecture - hardware for AI
Slide deck
September 1
Instructor lecture - language model basics
Slide deck
September 3
Paper discussion
Efficient Memory Management for Large Language Model Serving with PagedAttention
Slide deckzhong45, jinruih2, sehyeon2neilas3, yrao5, mfp7kaidifu2
September 8
Paper discussion
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Slide deckxz105, cmo8tylerba2neilas3, yrao5, mfp7
September 10
Paper discussion
Splitwise: Efficient Generative LLM Inference Using Disaggregated Compute
Slide decknsatch3, bikrant2, dsun19qpang2, fl22, jinjig2, zhang430zhong45, jinruih2, sehyeon2
September 15
Paper discussion
Fast Inference from Large Language Models via Speculative Decoding
apandey8, evanhz2, enguang2, yehyas2siddc2, coibe2, aryanns2, boddeti2haorany7, mf46, hongyic5
September 17
Paper discussion
Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations
neilas3, yrao5, mfp7javierd3, hieunn2, nikunjm2, kk103nsatch3, bikrant2, dsun19
September 22
No class
September 24
Paper discussion
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
qpang2, fl22, jinjig2, zhang430kaidifu2xz105, cmo8
September 29
Paper discussion
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
siddc2, coibe2, aryanns2, boddeti2apandey8, evanhz2, enguang2, yehyas2qpang2, fl22, jinjig2, zhang430
October 1
Google Guest Lecture
Suvinay Subramanian
Open details

Talk title: Codesigning Computing Systems for Artificial Intelligence

Abstract: The rapid advancement of artificial intelligence has ushered in an era of unprecedented computational demands, necessitating continuous innovation in computing systems. This talk highlights how codesign has enabled innovative solutions and state-of-the-art performance in Google's AI computing systems, namely Tensor Processing Units (TPUs).

It presents several codesign case studies across hardware, systems, software, algorithms, and the datacenter. The discussion examines how TPUs have made judicious, yet opinionated, design choices that have both kept pace with the rapid rate of change and enabled many breakthroughs in AI.

Speaker bio: Suvinay Subramanian is a Senior Staff Engineer at Google DeepMind, where he works on architecture and codesign for Google's ML supercomputers, Tensor Processing Units. His work has directly influenced innovative architecture and systems features in multiple TPU generations, supporting high-performance training and serving for Google's research and production AI workloads. He received a Ph.D. from MIT and a B.Tech. from the Indian Institute of Technology Madras. He is a recipient of the ACM SIGMICRO Early Career Award in Computer Architecture and co-hosts the Computer Architecture podcast, which spotlights developments in computer architecture and systems.

October 6
Intel Guest Lecture
Mrittika Ganguli
Open details

Talk title: The CPU Is the Agentic Critical Path: The Case for Heterogeneous Core Profiles

Abstract: Serving a language model is a forward pass. Serving an agent is a loop: reason, retrieve, act, admit, context, commit. Five of those six phases run on general-purpose cores. The CPU in this workload is not the host attached to an accelerator; it is the critical path. This talk argues that the six phases have CPU signatures different enough that no single core profile serves them all, and presents the measurements behind that claim.

Starting with what happens when a message is sent to an AI assistant, the talk traces where a single pass becomes a loop. It then examines how to count what an agent system actually delivers: trials per minute, turns per second, sandbox work seconds, and agents held inside an SLO. Aggregate throughput is the wrong axis to buy on. A capacity model decomposes delivered capacity into routing, balance, and offload efficiency, drawing on familiar operational laws including Amdahl's law, Little's law, the universal scalability law, and roofline.

The empirical core is phase-level measurement on server CPUs, including two counterintuitive findings: under agent density, the scheduler binds before memory bandwidth does; and moving tool execution to another node costs almost nothing. Both point toward racks whose nodes are matched to phases rather than uniform. The talk closes with implications for core profiles, on-die accelerators, and a memory hierarchy organized around sessions; a short critique session on three public capacity claims, including Intel's; and a set of open problems suitable for a semester project.

Speaker bio: Mrittika Ganguli is a Senior Principal Engineer and Platform Software Architect in Intel's Data Center Group. She leads Xeon platform software optimization and strategy, as well as software-enablement workstreams for Xeon 7 and 8. Her work spans accelerator enablement across QAT, DSA, and IAA, workload characterization, and cross-organization prioritization of software investments. She also works on agentic AI architecture, focusing on how CPU platforms, accelerators, and system software can deliver efficient, scalable infrastructure for emerging agentic AI workloads. She holds a master's degree in computer science and has published more than 80 patents.

October 8
Paper discussion
Parrot: Efficient Serving of LLM-Based Applications with Semantic Variable
hsutzut2, ambikas2, junyep2zhong45, jinruih2, sehyeon2dagraw2, ssashi2
October 13
Paper discussion
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
devansh9, vinaykp2nsatch3, bikrant2, dsun19apandey8, evanhz2, enguang2, yehyas2
October 15
NVIDIA Guest Lecture
Vikram Sharma Mailthody
October 20
Instructor lecture - agentic AI basics and power optimizations
October 22
No lecture
October 27
Gimlet Labs Guest Lecture
Ankita Nayak
October 29
Paper discussion
Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
javierd3, hieunn2, nikunjm2, kk103dagraw2, ssashi2siddc2, coibe2, aryanns2, boddeti2
November 3
Paper discussion
Timeloop: A Systematic Approach to DNN Accelerator Evaluation
lifan3, mjbove2, vpal2, cyz4hsutzut2, ambikas2, junyep2devansh9, vinaykp2
November 5
NVIDIA Guest Lecture
Benjamin Klenk
November 10
Paper discussion
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions
kaidifu2xz105, cmo8tylerba2
November 12
Paper discussion
Corsair: An In-Memory Computing Chiplet Architecture for Inference-Time Compute Acceleration
tylerba2devansh9, vinaykp2lifan3, mjbove2, vpal2, cyz4
November 17
Paper discussion
Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
haorany7, mf46, hongyic5lifan3, mjbove2, vpal2, cyz4hsutzut2, ambikas2, junyep2
November 19
Cerebras Guest Lecture
Ishita Chaturvedi
November 24
Fall break
November 26
Fall break
December 1
Paper discussion
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
dagraw2, ssashi2haorany7, mf46, hongyic5javierd3, hieunn2, nikunjm2, kk103
December 3
AMD Guest Lecture
Muhammad Awad
December 8
Final project demonstration