Schedule
Class calendar
To see the full list of papers and their details, please refer to the Reading List tab.
Explanations of the presenter, advocate, and critic group roles are available in the Resources tab.
This schedule is subject to change and currently serves as a placeholder while we finalize discussion-group assignments and a few remaining class activities.
| Date | Classroom Agenda | Slide deck | Presenter Group | Advocate Group | Critic Group |
|---|---|---|---|---|---|
| August 25 | Instructor lecture - course overview | Slide deck | |||
| August 27 | Instructor lecture - hardware for AI | Slide deck | |||
| September 1 | Instructor lecture - language model basics | Slide deck | |||
| September 3 | Paper discussion Efficient Memory Management for Large Language Model Serving with PagedAttention | Slide deck | zhong45, jinruih2, sehyeon2 | neilas3, yrao5, mfp7 | kaidifu2 |
| September 8 | Paper discussion Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve | Slide deck | xz105, cmo8 | tylerba2 | neilas3, yrao5, mfp7 |
| September 10 | Paper discussion Splitwise: Efficient Generative LLM Inference Using Disaggregated Compute | Slide deck | nsatch3, bikrant2, dsun19 | qpang2, fl22, jinjig2, zhang430 | zhong45, jinruih2, sehyeon2 |
| September 15 | Paper discussion Fast Inference from Large Language Models via Speculative Decoding | apandey8, evanhz2, enguang2, yehyas2 | siddc2, coibe2, aryanns2, boddeti2 | haorany7, mf46, hongyic5 | |
| September 17 | Paper discussion Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations | neilas3, yrao5, mfp7 | javierd3, hieunn2, nikunjm2, kk103 | nsatch3, bikrant2, dsun19 | |
| September 22 | No class | ||||
| September 24 | Paper discussion Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs | qpang2, fl22, jinjig2, zhang430 | kaidifu2 | xz105, cmo8 | |
| September 29 | Paper discussion FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness | siddc2, coibe2, aryanns2, boddeti2 | apandey8, evanhz2, enguang2, yehyas2 | qpang2, fl22, jinjig2, zhang430 | |
| October 1 | Google Guest Lecture Suvinay Subramanian Open detailsTalk title: Codesigning Computing Systems for Artificial Intelligence Abstract: The rapid advancement of artificial intelligence has ushered in an era of unprecedented computational demands, necessitating continuous innovation in computing systems. This talk highlights how codesign has enabled innovative solutions and state-of-the-art performance in Google's AI computing systems, namely Tensor Processing Units (TPUs). It presents several codesign case studies across hardware, systems, software, algorithms, and the datacenter. The discussion examines how TPUs have made judicious, yet opinionated, design choices that have both kept pace with the rapid rate of change and enabled many breakthroughs in AI. Speaker bio: Suvinay Subramanian is a Senior Staff Engineer at Google DeepMind, where he works on architecture and codesign for Google's ML supercomputers, Tensor Processing Units. His work has directly influenced innovative architecture and systems features in multiple TPU generations, supporting high-performance training and serving for Google's research and production AI workloads. He received a Ph.D. from MIT and a B.Tech. from the Indian Institute of Technology Madras. He is a recipient of the ACM SIGMICRO Early Career Award in Computer Architecture and co-hosts the Computer Architecture podcast, which spotlights developments in computer architecture and systems. | ||||
| October 6 | Intel Guest Lecture Mrittika Ganguli Open detailsTalk title: The CPU Is the Agentic Critical Path: The Case for Heterogeneous Core Profiles Abstract: Serving a language model is a forward pass. Serving an agent is a loop: reason, retrieve, act, admit, context, commit. Five of those six phases run on general-purpose cores. The CPU in this workload is not the host attached to an accelerator; it is the critical path. This talk argues that the six phases have CPU signatures different enough that no single core profile serves them all, and presents the measurements behind that claim. Starting with what happens when a message is sent to an AI assistant, the talk traces where a single pass becomes a loop. It then examines how to count what an agent system actually delivers: trials per minute, turns per second, sandbox work seconds, and agents held inside an SLO. Aggregate throughput is the wrong axis to buy on. A capacity model decomposes delivered capacity into routing, balance, and offload efficiency, drawing on familiar operational laws including Amdahl's law, Little's law, the universal scalability law, and roofline. The empirical core is phase-level measurement on server CPUs, including two counterintuitive findings: under agent density, the scheduler binds before memory bandwidth does; and moving tool execution to another node costs almost nothing. Both point toward racks whose nodes are matched to phases rather than uniform. The talk closes with implications for core profiles, on-die accelerators, and a memory hierarchy organized around sessions; a short critique session on three public capacity claims, including Intel's; and a set of open problems suitable for a semester project. Speaker bio: Mrittika Ganguli is a Senior Principal Engineer and Platform Software Architect in Intel's Data Center Group. She leads Xeon platform software optimization and strategy, as well as software-enablement workstreams for Xeon 7 and 8. Her work spans accelerator enablement across QAT, DSA, and IAA, workload characterization, and cross-organization prioritization of software investments. She also works on agentic AI architecture, focusing on how CPU platforms, accelerators, and system software can deliver efficient, scalable infrastructure for emerging agentic AI workloads. She holds a master's degree in computer science and has published more than 80 patents. | ||||
| October 8 | Paper discussion Parrot: Efficient Serving of LLM-Based Applications with Semantic Variable | hsutzut2, ambikas2, junyep2 | zhong45, jinruih2, sehyeon2 | dagraw2, ssashi2 | |
| October 13 | Paper discussion Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures | devansh9, vinaykp2 | nsatch3, bikrant2, dsun19 | apandey8, evanhz2, enguang2, yehyas2 | |
| October 15 | NVIDIA Guest Lecture Vikram Sharma Mailthody | ||||
| October 20 | Instructor lecture - agentic AI basics and power optimizations | ||||
| October 22 | No lecture | ||||
| October 27 | Gimlet Labs Guest Lecture Ankita Nayak | ||||
| October 29 | Paper discussion Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference | javierd3, hieunn2, nikunjm2, kk103 | dagraw2, ssashi2 | siddc2, coibe2, aryanns2, boddeti2 | |
| November 3 | Paper discussion Timeloop: A Systematic Approach to DNN Accelerator Evaluation | lifan3, mjbove2, vpal2, cyz4 | hsutzut2, ambikas2, junyep2 | devansh9, vinaykp2 | |
| November 5 | NVIDIA Guest Lecture Benjamin Klenk | ||||
| November 10 | Paper discussion SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions | kaidifu2 | xz105, cmo8 | tylerba2 | |
| November 12 | Paper discussion Corsair: An In-Memory Computing Chiplet Architecture for Inference-Time Compute Acceleration | tylerba2 | devansh9, vinaykp2 | lifan3, mjbove2, vpal2, cyz4 | |
| November 17 | Paper discussion Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC | haorany7, mf46, hongyic5 | lifan3, mjbove2, vpal2, cyz4 | hsutzut2, ambikas2, junyep2 | |
| November 19 | Cerebras Guest Lecture Ishita Chaturvedi | ||||
| November 24 | Fall break | ||||
| November 26 | Fall break | ||||
| December 1 | Paper discussion DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency | dagraw2, ssashi2 | haorany7, mf46, hongyic5 | javierd3, hieunn2, nikunjm2, kk103 | |
| December 3 | AMD Guest Lecture Muhammad Awad | ||||
| December 8 | Final project demonstration |