x is (B, D). Normalise each feature across the batch to zero mean and
unit variance, using the population variance and no affine parameters.
Do not use nn.BatchNorm1d. The contrast with LayerNorm is the axis: BatchNorm
reduces over the batch dimension, which is why its behaviour depends on batch
size and why it needs running statistics at inference.
Input
x =
tensor([[ 1., 10.],
[ 3., 20.],
[ 5., 30.]])
eps = 1e-05
Output
tensor([[-1.2247, -1.2247],
[ 0.0000, 0.0000],
[ 1.2247, 1.2247]])