Implement layer normalisation over the last dimension without calling
nn.LayerNorm or F.layer_norm: normalise each vector to zero mean and unit
variance, then scale by gamma and shift by beta.
Use the population variance (unbiased=False). PyTorch's .var() defaults
to the sample variance, dividing by n-1, which is not what LayerNorm uses and
is the single easiest way to get this subtly wrong.
LayerNorm normalises across features within one example, which is why it does not care about batch size -- the property that makes it work for transformers.
Input
x =
tensor([[ 1., 2., 3.],
[10., 10., 16.]])
gamma = tensor([1., 1., 1.])
beta = tensor([0., 0., 0.])
eps = 1e-05
Output
tensor([[-1.2247, 0.0000, 1.2247],
[-0.7071, -0.7071, 1.4142]])