Notes on LLM inference systems: quantization, kernels, serving, and profiling.
First post coming soon.