COMS 6998, Fall 2026
Serving a language model at scale is a networking problem. Below is one request moving through a prefill GPU and a decode GPU. Change the link speed and watch where the time goes.
Shown 12x slower than real time. Numbers assume an 8B model (128 KB of KV cache per token), 0.35 ms/token prefill, and 20 ms per decoded token.
ssh <your-uni>@mv.cs.columbia.edu
Read the server guide