Implement binary focal loss from raw logits:
loss = -alpha * (1 - p_t)**gamma * log(p_t)
where p = sigmoid(logits) and p_t is p for a positive target and 1 - p
for a negative one. Return the mean over all elements.
The (1 - p_t)**gamma factor shrinks the contribution of examples already
classified confidently, so training attends to the hard ones. Note that alpha
here multiplies every term, positive and negative alike.
Input
logits = tensor([ 2.0000, -1.0000, 0.5000])
targets = tensor([1., 0., 1.])
alpha = 0.25
gamma = 2.0
Output
tensor(0.0077)