Create a Linear(in_features, out_features), then re-initialise it: weights
with Xavier (Glorot) uniform initialisation and bias set to zeros.
Call torch.manual_seed(seed) immediately before the Xavier call, so the draw is
reproducible. Return the tuple (weight, bias).
Xavier scales the bound by both fan-in and fan-out, which is what keeps the forward and backward variance balanced for symmetric activations.
Input
in_features = 4
out_features = 3
seed = 0
Output
(tensor([[-0.0069, 0.4967, -0.7620, -0.6813],
[-0.3566, 0.2483, -0.0183, 0.7341],
[-0.0822, 0.2450, -0.2798, -0.1820]]), tensor([0., 0., 0.]))