Premium problem137. Backprop Gradients for a Two-Layer MLP

Hard Locked

Return the four gradients (dW1, db1, dW2, db2) for this network:

  • z1 = X @ W1 + b1, a1 = relu(z1), z2 = a1 @ W2 + b2
  • the loss is mean softmax cross-entropy against labels y
  • no autodiff

Premium problem

This one's part of Premium. Unlock the full NumPy track plus every other premium problem on the site.

Implement solve(...)