Perform exactly one training step on a Linear(4, 1) built after
torch.manual_seed(seed), using SGD at learning rate lr and mean-squared-error
loss against targets y.
The order matters: clear the gradients, forward, compute the loss, backward, then
step. Return the tuple (loss_value, weight_after_step) where loss_value is
the scalar loss from before the update and weight_after_step is the layer's
weight once the optimizer has stepped.
One step only. Stepping before backward(), or forgetting to zero the
gradients, are the two classic ways to get this wrong, and both change the
answer.
Input
x =
tensor([[1., 1., 1., 1.],
[1., 1., 1., 1.]])
y =
tensor([[1.],
[2.]])
lr = 0.1
seed = 0
Output
(tensor(5.1235), tensor([[0.4378, 0.7097, 0.0300, 0.0735]]))