Premium problem71. Transformer Feed-Forward Block

Medium Locked

x has shape (B, T, D). Build the feed-forward block a transformer uses between attention layers -- Linear(D, 4D), then GELU, then Linear(4D, D) -- and return its output, which keeps the shape (B, T, D).

Seed with torch.manual_seed(seed) immediately before creating the layers, in that order. Run under torch.no_grad().

Input

x =
tensor([[[1., 1., 1., 1.],
         [1., 1., 1., 1.],
         [1., 1., 1., 1.]],

        [[1., 1., 1., 1.],
         [1., 1., 1., 1.],
         [1., 1., 1., 1.]]])
seed = 0

Output

tensor([[[-0.1423,  0.1748,  0.1783,  0.0001],
         [-0.1423,  0.1748,  0.1783,  0.0001],
         [-0.1423,  0.1748,  0.1783,  0.0001]],

        [[-0.1423,  0.1748,  0.1783,  0.0001],
         [-0.1423,  0.1748,  0.1783,  0.0001],
         [-0.1423,  0.1748,  0.1783,  0.0001]]])

Premium problem

This one's part of Premium. Unlock the full PyTorch track plus every other premium problem on the site.

Implement solve(...)