Return the (n, n) boolean mask where entry [i, j] is True exactly when
position i is allowed to attend to position j, which for causal attention
means j <= i.
For n = 3 that is a lower triangle including the diagonal. This is what stops a
language model reading its own answer.
Input
3
Output
tensor([[ True, False, False],
[ True, True, False],
[ True, True, True]])