Compute cross-entropy with label smoothing. Instead of putting all the
probability mass on the true class, the target distribution gives it
1 - epsilon + epsilon/C and every other class epsilon/C, where C is the
number of classes.
logits is (B, C) and targets is (B,). Return the mean loss. At
epsilon = 0 it must equal ordinary cross-entropy exactly.
Smoothing stops the model driving the true logit to infinity, which is where overconfidence comes from.
Input
logits =
tensor([[2.0000, 1.0000, 0.1000],
[0.5000, 2.5000, 0.3000]])
targets = tensor([0, 1])
epsilon = 0.1
Output
tensor(0.4369)