Fuse two modalities. Given image_features of shape (B, image_dim) and
text_features of shape (B, text_dim):
hidden_dim with a Linearhidden_dim with a second Linearnum_classes with a third LinearSeed with torch.manual_seed(seed) immediately before creating the layers, and
create them in that order -- image projection, text projection, classifier --
since the order determines which random weights each one receives. Run under
torch.no_grad() and return the (B, num_classes) logits.
Input
image_features =
tensor([[1., 1., 1., 1.],
[1., 1., 1., 1.]])
text_features =
tensor([[0.5000, 0.5000, 0.5000],
[0.5000, 0.5000, 0.5000]])
hidden_dim = 5
num_classes = 3
seed = 0
Output
tensor([[ 0.2981, 0.2158, -0.6512],
[ 0.2981, 0.2158, -0.6512]])