Build a Linear(len(grad_values), 1) and set its weight gradient to
grad_values as a single row, leaving the bias gradient at zero. Clip the
model's global gradient norm to max_norm using PyTorch's own clipping utility,
then return the weight gradient afterwards as a 1-D tensor.
Clipping rescales all gradients by one shared factor when the global norm exceeds the threshold, and leaves them untouched when it does not -- it is not a per-element clamp.
Input
grad_values = [3.0, 4.0]
max_norm = 1.0
Output
tensor([0.6000, 0.8000])