Compute log-softmax over the last dimension of x without calling
torch.softmax, torch.log_softmax, F.softmax or F.log_softmax.
It must stay finite for very large logits. Exponentiating first overflows to
inf and then produces nan; subtracting the row maximum before exponentiating
leaves the result unchanged mathematically and keeps every intermediate in
range.
Input
tensor([[1., 2., 3.],
[0., 0., 0.]])
Output
tensor([[-2.4076, -1.4076, -0.4076],
[-1.0986, -1.0986, -1.0986]])