A classifier's true accuracy is p. You measure it on n independent
examples. Return the probability the measured accuracy comes out at least
threshold.
The number correct is , so sum from upwards. On a small test set this is uncomfortably far from certain even when the model is genuinely good.
Input
n = 100
p = 0.9
threshold = 0.85
Output
0.9601094728889168