Given true labels y (0 or 1) and predicted probabilities, return the mean
binary cross-entropy.
This is the negative log-likelihood of the Bernoulli model: minimising the loss and maximising the likelihood are the same operation.
Input
y = [1, 0, 1]
probabilities = [0.9, 0.1, 0.8]
Output
0.14462152754328741