Blog
Notes on LLM inference systems: quantization, kernels, serving, and profiling.
-
Let's optimize Nemotron 3.5 Lightning - a worklog
I came across NVIDIA’s Nemotron 3.5 Lightning model recently. I rented a RTX PRO 6000 box, thanks to Huggingface and Modal for their GPU credits! I had Claude Fable. So yeah, spent my weekend on squeezing as much as p...