x has shape (B, T, D). Build the feed-forward block a transformer uses
between attention layers -- Linear(D, 4D), then GELU, then
Linear(4D, D) -- and return its output, which keeps the shape (B, T, D).
Seed with torch.manual_seed(seed) immediately before creating the layers, in
that order. Run under torch.no_grad().
Input
x =
tensor([[[1., 1., 1., 1.],
[1., 1., 1., 1.],
[1., 1., 1., 1.]],
[[1., 1., 1., 1.],
[1., 1., 1., 1.],
[1., 1., 1., 1.]]])
seed = 0
Output
tensor([[[-0.1423, 0.1748, 0.1783, 0.0001],
[-0.1423, 0.1748, 0.1783, 0.0001],
[-0.1423, 0.1748, 0.1783, 0.0001]],
[[-0.1423, 0.1748, 0.1783, 0.0001],
[-0.1423, 0.1748, 0.1783, 0.0001],
[-0.1423, 0.1748, 0.1783, 0.0001]]])