networks
for AI

COMS 6998, Fall 2026

Serving a language model at scale is a networking problem. Below is one request moving through a prefill GPU and a decode GPU. Change the link speed and watch where the time goes.

Time to first token
-
KV cache over the link
-
Decoding
-

Shown 12x slower than real time. Numbers assume an 8B model (128 KB of KV cache per token), 0.35 ms/token prefill, and 20 ms per decoded token.

ssh <your-uni>@mv.cs.columbia.edu Read the server guide