In the functional space, and with the right kernel to compare functions, neural net learning can actually be regarded as a nice gradient descent, which is even convex for some common loss functions, as discussed by Arthur Jacot, PhD candidate in mathematics at EPFL.
https://people.epfl.ch/arthur.jacot
Check Arthur's 2018 NeurIPS paper on the neural tangent kernel
• Neural Tangent Kernel: Convergence and Gen...
https://arxiv.org/abs/1806.07572