Table of Contents:
00:12 - Expanding Iteration Space
00:35 - Application must have enough parallelism to keep all cores busy
00:52 - Example: Matrix elements summation
00:58 - Problem statement
01:13 - Naive implementation
01:24 - Code explained
01:31 - Thread parallelism
01:38 - Vectorized inner loop
02:03 - Short but wide matrix
02:07 - Loop tiling can lead to t
02:20 - VTune analysis of the naive implementation
02:51 - Problem with the current implementation
03:02 - Redistribution parallelism
03:10 - Optimized implementation
03:20 - Code explained
03:28 - Strip-mining technique
04:02 - #pragma omp for collapse(2)
04:34 - Summary of changes in the code
04:46 - Avoiding race conditions
04:57 - Redoing analysis in VTune for improved code
05:08 - Intent of the optimization
05:15 - Other strategies to try
05:46 - Performance results
06:41 - STREAM: https://www.cs.virginia.edu/stream/