Episode 6 of Inside the CPU: Memory Lab. The last episode followed a virtual address through the page tables and met the TLB, the cache of translations; this one is about the other cache, the one that holds your data. Why a read from RAM costs about 300 CPU cycles, what L1, L2 and L3 are with their typical sizes and latencies, why the cache moves whole 64 byte lines and never single bytes, what a hit and a miss cost, spatial and temporal locality, and how the prefetcher guesses the next read. In the lab we read the cache sizes of the machine from /sys, then sum the same 16 MiB array in order and through a shuffled index: same O(n), same sum, often several times slower. Every number in the video is a typical order of magnitude; the samples print the numbers of your own machine.
Covered:
Why caches exist: one cycle is about 0.3 ns at 3 GHz, one RAM read about 100 ns
L1 (per core, data and code, about 1 ns), L2 (per core, about 4 ns), L3 (shared, 10 to 40 ns), and how a miss falls to the next level
The 64 byte cache line: 16 ints per line, pay for one read and get the neighbours for free
Hit and miss: the line number is the address divided by 64, and how many lines a loop touches
Spatial and temporal locality: stride 1 against stride 16
The prefetcher: constant strides are fetched ahead, random walks and linked lists are not
Reading level, type, size and coherency_line_size from /sys/devices/system/cpu/cpu0/cache, and lscpu --caches
Sequential sum against a shuffled index with bench_ns, fixed seed, the ratio labelled this machine
The same random walk over 16 KiB, 256 KiB, 4 MiB and 64 MiB: the edges of L1, L2 and L3
Next episode: alignment, padding and locality, because the layout decides the hits
The slides, the C++ sample code and the full captions are on GitHub:
https://github.com/tomjnet/inside-the...