Premium problem70. Multimodal Fusion Network

Hard Locked

Fuse two modalities. Given image_features of shape (B, image_dim) and text_features of shape (B, text_dim):

  • project the image features to hidden_dim with a Linear
  • project the text features to hidden_dim with a second Linear
  • concatenate the two projections along the feature dimension
  • map the result to num_classes with a third Linear

Seed with torch.manual_seed(seed) immediately before creating the layers, and create them in that order -- image projection, text projection, classifier -- since the order determines which random weights each one receives. Run under torch.no_grad() and return the (B, num_classes) logits.

Input

image_features =
tensor([[1., 1., 1., 1.],
        [1., 1., 1., 1.]])
text_features =
tensor([[0.5000, 0.5000, 0.5000],
        [0.5000, 0.5000, 0.5000]])
hidden_dim = 5
num_classes = 3
seed = 0

Output

tensor([[ 0.2981,  0.2158, -0.6512],
        [ 0.2981,  0.2158, -0.6512]])

Premium problem

This one's part of Premium. Unlock the full PyTorch track plus every other premium problem on the site.

Implement solve(...)