embeddings is (B, T, D) and mask is (B, T) with 1 for a real token and
0 for padding. Return the (B, D) mean over valid tokens only.
Averaging over T instead would let padding drag every sentence toward zero, and
longer padding would distort it more -- which is the bug this problem exists to
prevent.
Input
embeddings =
tensor([[[ 0., 1.],
[ 2., 3.],
[ 4., 5.]],
[[ 6., 7.],
[ 8., 9.],
[10., 11.]]])
mask =
tensor([[1, 1, 0],
[1, 0, 0]])
Output
tensor([[1., 2.],
[6., 7.]])