#computerarchitecture #riscv
Reference:
Most of lecture material is derived from freely available course i.e. UC Berkeley course CS 152
https://inst.eecs.berkeley.edu/~cs152...
Single Cycle Processor
(0:00) Intro
(4:18) A simple processor design
(22:47) RISC-V single cycle processor design
Pipelining and branch prediction
(1:00:00) Pipeline intro and critical path
(1:09:42) Pipelining single cycle RISC-V
(1:17:30) Hazards: Structural, Data & Control Hazards
(1:22:33) Interocks / Stalls
(1:44:06) Bypassing / Forwarding
(1:51:48) Iron Law of processor performance
(2:00:47) Control Hazards
(2:03:18) Speculate PC+4
(2:09:54) Branch delay slots
(2:15:59) Branch History Table (BHT), Branch Target Buffer (BTB)
(2:49:59) Subroutine return stack
(2:52:40) Interrupts
Memory - Caches and VM
(3:08:19) Memories overview, SRAM, DRAM
(3:21:41) Cache intro. Multilevel Memory
(3:38:25) Placement policy: Directly mapped, Set associative,Fully associative
(3:53:00) Replacement policy, Least recently used, Random, FIFO, NLRU
(3:56:50) Block size, Cache Performance
(4:05:40) Write policy, Pipelined cache write, Write buffer, RAM-TAG cache
(4:15:34) Multi-level cache, Inclusion policy
(4:20:38) Examples Itanium-2, IBM Power 7, z196
(4:25:03) Victim Cache, Way prediction, Early restart, Non-Blocking Cache
(4:31:27) Pre-fetching and Compiler Optimizations
Virtual Memory System
(4:46:50) Base and Bound Scheme
(4:57:48) Fragmentation and Paged Memory System
(5:09:27) Page table size and two level page table
(5:20:13) TLB, Demand paging, Aliasing
(5:45:44) Hashed page table
Complex Pipelines, Superscalar, Scoreboard
(5:52:20) Complex Pipelines motivation, FPUs, Memory subsytem
(6:03:00) Complex In order pipeline
(6:06:50) Data Hazards (RAW, WAR, WAW)
(6:14:45) Scoreboard and Out of order completion
(6:35:36) Tomasulo: Out of order issue and Register renaming
(6:56:09) Exception and Branch issues, In order Commit
(7:07:32) Unified Physical Register file
(7:28:26) Memory dependencies
VLIW, Loop unrolling, Software pipelining
(7:41:29) Superscalar control logic scaling
(7:44:23) VLIW - Very Long Instruction Word
(7:48:39) Loop Unrolling & Software pipelining
(8:02:02) Trace scheduling
(8:10:07) Predicated execution
(8:12:00) Speculative execution
Multithreading, coarse and fine grained
(8:14:52) Multi processing, Multi tasking, ILP and TLP
(8:22:06) Fine grained multithreading
(8:33:20) Coarse grained multithreading
(8:39:11) Simultaneous Multithreading
Vector Processors
(8:51:59) Supercomputers and CDC 6600 by Seymour Cray
(9:00:40) Vector Processors
(9:17:50) Stripmining, Mask registers, Reductions, Scatter/Gather
(9:42:40) Multimedia extensions, SIMD
Synchronization and Sequential consistency
(9:44:34) Symmetric Multiprocessors, Synchronization, Producer/Consumer, Mutex/Locks
(10:00:30) Sequential Consistency
(10:08:24) Memory Fences, Membar, Locks, Critical Section, Atomic Operations
(10:15:57) Dekker's algorithm, Lamport's Bakery algorithm
(10:27:45) Semaphores, Atomic Read-Modify-Writes, Test&Set, Fetch&Add, Swap
(10:38:16) Non blocking Synchronization, Compare&Swap, Load Reserve & Store Conditionals
Snoopy Cache, Cache Coherence
(10:48:43) Sequential consistency issues, Cache coherence and Snoopy Cache
(11:00:58) Snoopy cache for Multicore systems, MSI Protocol and MESI Protocol
(11:22:27) Directory based cache