Table of Contents:
00:11 - Problem statement: matrix-vector multiplication
00:36 - Naive implementation of matrix-vector multiplication
01:20 - Why temporal locality is important for vector b?
01:40 - Problem with naive implementation
02:04 - Loop tiled code implementation for matrix-vector multiplication
02:34 - Code explanation
02:52 - Complications of new implementation
03:05 - Parallel reduction
03:20 - Strip-mining, parallel reduction, alignment and compiler hints
03:26 - Next optimization step
04:13 - New implementation - expanding parallel iterations space
04:18 - Double tiling
04:31 - iTile is imperical
04:45 - Orders of the tiled loops
04:56 - Choosing tiling optimization parameter
05:44 - Performance results for matrix-vector multiplication
06:32 - Cache-oblivious method results