Home / Chapter 7 · Production and serving
    Last edited · 5 min read

    Use with AI

    Mixture of Experts

    In a Mixture of Experts (MoE) model, the MLP layer is replaced by many “experts”, and a router picks a few of them for each token. The model has a huge number of parameters but computes with only some of them per token: compute like a small model, memory like a large one.

    In plain wordsA clinic. Reception sends you to two of its eight doctors. They all have to be in the building, but you pay for only two appointments. When a crowd arrives every doctor has patients, and reception has to make sure the whole queue doesn’t end up outside one door.

    Click the tokens and watch which experts the router picks. At the bottom: how many experts the whole sentence needs

    2 of 8experts compute for one token
    experts needed by the whole sentence, i.e. the batch
    13B of 47Bparameters active in Mixtral 8×7B

    Simplified: one layer, 8 experts and 2 per token, as in Mixtral 8×7B (2023), and the routing is illustrative, not taken from a real model. The bar under each expert shows how often the tokens in this sentence picked it. The router looks at the token’s state in context, not just its text, so the same token in a different position can go elsewhere.

    How routing works

    Consequences for serving

    MoE or a dense model

    Check yourself

    What is MoE, and what are its consequences for serving?

    In MoE the MLP layer is replaced by many experts, and a router picks a few for every token in every layer. Only the active parameters are computed, for example 37 of 671 billion in DeepSeek-V3, so the quality is closer to a large model at the compute of a small one. The price is memory, because every expert must be loaded, and serving: at larger batch sizes tokens hit almost every expert, so it takes much bigger batches, expert parallelism with all-to-all communication between GPUs, and care to keep the expert load balanced.

    Po polsku

    W MoE warstwę MLP zastępuje wielu ekspertów, a router dla każdego tokena w każdej warstwie wybiera kilku. Liczy się tylko aktywne parametry, na przykład 37 z 671 mld w DeepSeek-V3, więc jakość jest bliższa dużemu modelowi przy obliczeniach małego. Ceną jest pamięć, bo wszyscy eksperci muszą być załadowani, i serwowanie: przy większym batchu tokeny trafiają prawie do wszystkich ekspertów, więc potrzeba dużo większych batchy, expert parallelism z komunikacją all-to-all między kartami i pilnowania równego obciążenia ekspertów.

    Follow-up questions (3)
    An MoE model has 5 billion active parameters. Will it behave like a 5B model?
    It computes like 5B, but the whole model has to sit in memory. With one conversation, decode is as fast as a small model. At a medium batch size each step reads almost all the experts, so it costs as much as reading a large model while doing little computation. Only a very large batch brings the cost per token close to a 5B model.
    Why balance expert load during training?
    Without it the router collapses onto a few favourite experts, which get better and are picked more and more often, while the rest of the parameters go to waste. In serving, uneven load means the most loaded expert sets the step time.
    How would you spread a large MoE across GPUs?
    Attention usually with tensor or data parallelism, experts with expert parallelism: each GPU holds some of the experts, and tokens are exchanged all-to-all in every layer. What matters most is a fast interconnect between GPUs and duplicating the most frequently picked experts. DeepSeek-V3 does this on 320 GPUs for decode.

    Sources

    Report an error · Suggest a fix