Create a Linear(in_features, out_features) and re-initialise its weights
with Kaiming (He) normal initialisation configured for ReLU -- fan-in mode,
ReLU nonlinearity -- then zero the bias.
Seed with torch.manual_seed(seed) immediately before the Kaiming call. Return
(weight, bias).
Kaiming differs from Xavier by assuming a ReLU throws away half the signal, so it scales variance up to compensate.
Input
in_features = 4
out_features = 3
seed = 0
Output
(tensor([[ 1.0896, -0.2075, -1.5406, 0.4019],
[-0.7669, -0.9890, 0.2852, 0.5926],
[-0.5086, -0.2852, -0.4219, 0.1287]]), tensor([0., 0., 0.]))