image is (H, W) and kernel is (K, K). Convolve them with stride 1
and no padding, and return the 2-D result of shape (H-K+1, W-K+1).
F.conv2d works on 4-D batched input, so the single image and single kernel
have to gain batch and channel dimensions on the way in and lose them again on
the way out.
Note that F.conv2d computes a cross-correlation rather than a true
convolution -- it does not flip the kernel -- which is what every deep learning
framework means by "convolution".
Input
image =
tensor([[ 0., 1., 2., 3.],
[ 4., 5., 6., 7.],
[ 8., 9., 10., 11.],
[12., 13., 14., 15.]])
kernel =
tensor([[ 1., 0.],
[ 0., -1.]])
Output
tensor([[-5., -5., -5.],
[-5., -5., -5.],
[-5., -5., -5.]])