Why Transformers Need Something Like MoE
In transformer architectures, the attention block is followed by a feed-forward (FF) layer: a fully-connected network with a hidden layer and a nonlinearity, typically ReLU. This layer carries most of the model's weights — the hidden dimension is often 4x the embedding depth. That's not accidental: attention mixes token embeddings to capture their relationships, while the FF layer performs the actual per-token computation. Repeating this block dozens of times makes the parameter count grow quickly, which motivates the sparsely-gated mixture of experts approach.
The MoE Layout
Instead of one large FF layer, MoE splits it into NEXP separate FF blocks called experts, each transforming vectors of size D to vectors of size D. A router — a fully-connected layer of shape (D, NEXP) — scores each expert for a given token. Only the top K experts process the token, and their outputs are combined via a weighted average (weights derived from softmax over the router's scores) to produce the final output of size D.
With NEXP=8 and TOPK=2, experts #1 and #5 win the routing competition for the token above; all other experts are skipped entirely. This is the essence of MoE: increase total model capacity without proportionally increasing compute per token. The skipped experts contribute nothing on either the forward or backward pass.
This explains naming conventions like Mixtral 8x7B: with eight experts, it would be misleading to multiply 7B by 8, since not every parameter participates for every token. Per the Mixtral paper, only about 13B parameters are active for each token.
MoE increases the model's capacity without proportionally increasing its computational cost.
A Minimal Numpy Implementation
The core parameters and logic are straightforward. The router is a learned linear layer followed by softmax; expert selection uses argpartition to grab the top K indices without a full sort:
# Parameters for a feed-forward layer with a fixed activation function.
@dataclass
class FFParams:
Wh: np.ndarray
Wo: np.ndarray
# Parameters for a Mixture of Experts (MoE) layer.
@dataclass
class MoEParams:
# Embedding dimension of each token (a.k.a. model dimension, Dmodel)
D: int
# Hidden dimension in FF layers
DH: int
# Total number of experts
NEXP: int
# K in the top-k selection of top experts per token
TOPK: int
# List of experts: each expert is a forward layer with FFParams.
ff_weights: List[FFParams]
# Router weights: a linear layer (D, NEXP) that maps input to expert scores.
router_weights: np.ndarray
def moe(x: np.ndarray, params: MoEParams):
"""Mixture of Experts (MoE) layer.
Args:
x: Input tensor (B, N, D).
params: MoEParams.
Returns:
Output tensor (B, N, D).
"""
# Run input through router to get expert scores for each token.
expert_scores = x @ params.router_weights # (B, N, NEXP)
# Select the top-k expert scores and their indices for each token.
top_scores, top_experts = topk_lastdim(expert_scores, params.TOPK) # (B, N, TOPK)
# Apply softmax to the top scores to get weights that sum to 1.
weights = softmax_lastdim(top_scores) # (B, N, TOPK)
out = np.zeros_like(x)
for b in range(x.shape[0]):
for n in range(x.shape[1]):
# Unvectorized implementation: for each token in the batch and
# sequence, select the top-k experts and apply them with the
# calculated weights.
for expert_idx, weight in zip(top_experts[b, n], weights[b, n]):
expert = params.ff_weights[expert_idx]
out[b, n] += weight * feed_forward_relu(x[b, n], expert.Wh, expert.Wo)
return out
Note that the expert computations themselves aren't vectorized across tokens — they run one token at a time. This reflects MoE's inherent sparsity: different tokens in a batch may route to different experts, making bulk vectorization difficult and hardware-specific. A notable GPU-oriented effort in this direction is the MegaBlocks paper (2022); the problem remains an active research area.
def topk_lastdim(x, k):
"""Get the top k elements and their indices.
x is an arbitrary array with at least two dimensions. The returned
array has the same shape as x, but its elements are the top k elements
across the last dimension. The indices of the top k elements are also
returned.
"""
idx = np.argpartition(x, -k, axis=-1)[..., -k:]
return np.take_along_axis(x, idx, axis=-1), idx
def softmax_lastdim(x):
"""Compute softmax across last dimension of x.
x is an arbitrary array with at least two dimensions. The returned array has
the same shape as x, but its elements sum up to 1 across the last dimension.
"""
# Subtract the max for numerical stability
ex = np.exp(x - np.max(x, axis=-1, keepdims=True))
# Divide by sums across last dimension
return ex / np.sum(ex, axis=-1, keepdims=True)
Keeping Experts Busy: Load Balancing
A recurring concern with MoE is load balancing. Without countermeasures, the model tends to funnel most tokens through a small subset of experts, wasting capacity. Common remedies include injecting noise into the top-k selection to add randomness, and adding a training loss term that encourages roughly equal sample counts across experts.
Code
The full implementation is available on GitHub.



