Write a module with two pieces of state:
scale, a learnable scalar parameter initialised to scale_initrunning_offset, a registered buffer holding offset_initwhose forward pass returns x * scale + running_offset.
Return the tuple (parameter_names, buffer_names, state_dict_keys, output),
each name list taken in the order PyTorch yields it and state_dict_keys sorted.
A buffer is state that moves with the model, is saved in state_dict and
follows .to(device), but is not a parameter and takes no gradient. Running
statistics are the usual example; registering them as parameters instead would
have the optimiser trying to train them.
Input
x = tensor([1., 2.])
scale_init = 3.0
offset_init = 0.5
Output
(['scale'],
['running_offset'],
['running_offset', 'scale'],
tensor([3.5000, 6.5000]))