A progressive map
Moves from foundations through core, extended, and advanced topics for sequential or reference reading.
transformer.ngjoo.com · ONLINE
A four-part, 33-chapter path through the Transformer stack, from attention fundamentals to KV Cache, MoE, LoRA, quantization, and inference optimization.
Moves from foundations through core, extended, and advanced topics for sequential or reference reading.
Explains Q/K/V, masking, multi-head attention, position encoding, and sampling through data flow and implementation.
Connects key ideas to papers, historical implementations, or teaching code with clear source boundaries.