0A Transformers and Mixture of Experts#
Chapter 0 showed what a model does: predict the next token, again and again. This page opens it up: the design almost every model shares, the transformer, and the trick that lets the largest ones run cheaply, mixture of experts. This page is pre-reading, with no session of its own.
▶ Open the chapter 0 notebook in Colab (section 3)Learning objectives
Describe the three steps a transformer takes to predict the next token, and why attention makes long inputs cost more.
Explain what a mixture-of-experts model saves, and what its “experts” are not.
0A.1 The transformer#
The neural network inside an LLM has a particular design, the transformer. Almost every LLM in use today is one; the design was published by Google researchers in 2017.
Step through one sentence the way a transformer reads it:
Because every token looks at every token before it, doubling the text more than doubles the work. That is one reason the context window (chapter 0, section 0.5) has a limit.
0A.2 Mixture of experts#
Many recent models, including DeepSeek, Mixtral and several GLM and Qwen models, split each layer into many small blocks, the experts, with a router that sends each token to only a few of them. Send tokens through and compare it with a dense model, where every token uses everything:
The firm works the same way: a question about Deere goes to the industrials analysts, not to all 14. One difference: the experts are not neat specialists such as “finance” or “grammar”. The router learns its own way of dividing the work, which is why the same token always lands in the same place but a topic can be spread across several experts.
Checkpoint.
DeepSeek-V3 has 671 billion weights but uses about 37 billion for each token. What does that buy?
Further reading#
Ashish Vaswani and others, Attention Is All You Need (2017) — the paper that introduced the transformer. Technical; the abstract and first figure are enough.
Hugging Face, Mixture of Experts Explained — how a router sends each token to a few experts, and what that saves.