← All projects

transformer.ngjoo.com · ONLINE

Transformer Atlas

A four-part, 33-chapter path through the Transformer stack, from attention fundamentals to KV Cache, MoE, LoRA, quantization, and inference optimization.

Content structure

CORE

A progressive map

Moves from foundations through core, extended, and advanced topics for sequential or reference reading.

Mechanisms unpacked

Explains Q/K/V, masking, multi-head attention, position encoding, and sampling through data flow and implementation.

Traceable sources

Connects key ideas to papers, historical implementations, or teaching code with clear source boundaries.