grads is a list of gradient tensors, one per parameter, with None where a
parameter received no gradient.
Return the global L2 norm: the square root of the sum of the squared L2 norms of
every tensor that is not None. Entries that are None are skipped entirely
rather than counted as zero-length.
This is the quantity gradient clipping compares against a threshold.
Input
[tensor([3., 4.]), None, tensor([[1.],
[2.]])]
Output
tensor(5.4772)