Computer Architecture Tutorial
A complete CSE 203 path - Computer Architecture and Organization from abstraction and ISA through ALU, single-cycle and pipelined datapaths, hazards, caches, virtual memory, DMA, and multicore coherence.
Curriculum
Work through each section in order. Every lesson ends with practice and key points so the idea sticks.
Course Orientation
- 1Welcome to Computer Architecture
Meet CSE 203 - why hardware DNA matters for every software engineer, and how this InTelleX path is organized.
12 min - 2Course Map, Labs & Tools
Set up your mental map of CSE 203 activities: Assembly Duel, Bottleneck Audit, Hardware Visualization, and Clock-Cycle Stress Test.
11 min
Architecture of Abstraction
- 3Von Neumann vs Harvard Architecture
Compare the classic stored-program model with Harvard separation of instruction and data memory.
12 min - 4Levels of Program Abstraction
Walk the stack from problem statement through HLL, assembly, machine code, microarchitecture, and digital logic.
11 min - 5The Iron Law of Performance
Master CPU time = Instruction Count × CPI × Clock Period - and what each term really means.
13 min - 6CPI, Clock Cycles & Execution Time
Calculate total cycles, effective CPI for mixed instruction mixes, and wall-clock execution time.
12 min - 7Project: The Architectural Profiler
Design a lightweight profiler concept that tracks instruction mixes and graphs CPI for software blocks.
16 min
Instruction Set Architecture
- 8What Is an Instruction Set Architecture?
Define the ISA as the hardware/software interface: operations, registers, memory model, and encodings.
12 min - 9RISC vs CISC Philosophies
Contrast Reduced vs Complex Instruction Set designs and why modern machines borrow from both.
12 min - 10Registers, Operands & Calling Conventions
Learn register files, operand types, and the RISC-V / MIPS calling convention essentials.
13 min - 11RISC-V Instruction Formats
Decode R-Type, I-Type, S-Type (and friends): fields, opcodes, and why formats stay regular.
14 min - 12Memory Addressing Modes
Understand base+offset addressing, alignment, and how addressing modes affect hardware complexity.
11 min - 13Translating C into RISC-V Assembly
Compile nested loops and conditionals by hand into clean assembly with minimal redundant instructions.
15 min - 14Lab: The Assembly Duel
Manually trace register state and memory contents step-by-step - no emulator crutches.
14 min - 15Project: The Native Translator
Design a lightweight disassembler that maps binary machine words to human-readable assembly.
16 min
ALU & Binary Logic
- 16Fixed-Point & Signed Number Systems
Represent unsigned and signed integers, two's complement, and overflow conditions.
12 min - 17Binary Addition, Subtraction & Logic Ops
Build intuition for full adders, subtract-via-negate, AND/OR/XOR, and shifts as ALU building blocks.
12 min - 18Carry-Lookahead Adders
Design faster addition with generate/propagate signals - escape the ripple-carry bottleneck.
14 min - 19Multipliers & Wallace Trees
See how array multipliers and Wallace trees compress partial products for high-speed multiply.
13 min - 20IEEE 754 Floating-Point Arithmetic
Decode sign, exponent, mantissa; normalize; and handle corner cases that break naive intuition.
15 min - 21Project: The Custom ALU
Specify a simulated ALU supporting arithmetic, logic, and shifts with clear control encodings.
18 min
Single-Cycle Datapath
- 22Datapath Elements
Identify PC, instruction memory, register file, ALU, data memory, and the muxes that steer them.
12 min - 23The Instruction Execution Cycle
Trace Fetch → Decode → Execute → Memory → Writeback for ALU, load, store, and branch instructions.
14 min - 24Control Units: Hardwired vs Microprogrammed
Compare hardwired combinational control with microprogrammed control stores and when each shines.
12 min - 25Project: The Blueprint Simulator
Model a single-cycle datapath that routes data correctly from opcode-driven control signals.
18 min
Pipeline Paradigm
- 26Why Pipelining? Instruction-Level Parallelism
See how overlapping instruction stages multiplies throughput without magically shortening individual instruction latency.
12 min - 27The Classic 5-Stage Pipeline
Map IF, ID, EX, MEM, WB onto hardware resources and pipeline registers between stages.
13 min - 28Pipeline Performance & Throughput
Compute pipelined speedup, pipeline CPI with stalls, and the cost of unbalanced stages.
12 min - 29Lab: The Pipeline Factory
Manually plot multi-cycle execution charts and calculate speedup versus non-pipelined execution.
15 min - 30Project: The Stage Monitor
Specify a visualization that shows instructions propagating through a 5-stage pipeline in real time.
16 min
Mid-Term Checkpoint
- 31Mid-Term Review: Digital Logic & ALU
Consolidate adders, number systems, FP basics, and ALU control before the mid-term checkpoint.
14 min - 32Mid-Term Review: ISA & Assembly
Rehearse formats, addressing, and hand translation under timed conditions.
14 min - 33Mid-Term Review: Datapath & Control
Trace control signals and mux settings for core instructions on a single-cycle blueprint.
13 min
Pipeline Hazards
- 34Structural Hazards
Recognize resource conflicts when two stages need the same hardware in one cycle.
11 min - 35Data Hazards: RAW, WAR, WAW
Classify read-after-write, write-after-read, and write-after-write dependencies in pipelines.
13 min - 36Control Hazards
See why branches disrupt the pipeline and how delay slots or prediction respond.
12 min - 37Forwarding & Bypassing
Wire EX/MEM and MEM/WB results back to ALU inputs to slash many RAW stalls.
14 min - 38Pipeline Stalls & Bubbles
Insert NOPs / freeze pipeline registers safely when forwarding is not enough.
11 min - 39Branch Prediction Basics
Compare predict-not-taken, static predictors, and 1-bit/2-bit dynamic predictors.
13 min - 40Lab: The Bottleneck Audit
Analyze assembly for stalls and reschedule instructions to improve throughput.
15 min - 41Project: The Hazard Detector
Design a module that scans assembly, flags dependencies, and proposes forwarding paths.
18 min
Memory Hierarchy
- 42Principle of Locality
Explain temporal and spatial locality - the reason caches and hierarchies work.
11 min - 43Direct-Mapped Caches
Map addresses to a single set: tag, index, offset - and compute hits/misses.
14 min - 44Set-Associative & Fully Associative Caches
Compare associativity levels, replacement policies, and the hit-rate vs hardware cost trade-off.
13 min - 45Write-Through vs Write-Back
Choose write policies and write-allocate vs no-write-allocate strategies with clear trade-offs.
12 min - 46AMAT & Miss Rate Analysis
Compute Average Memory Access Time from hit time, miss rate, and miss penalty - including multilevel caches.
13 min - 47Lab: Hit or Miss
Walk address streams through a small cache, counting hits/misses and tracking line updates.
15 min - 48Project: The Cache Configurator
Build an algorithmic cache simulator to sweep line size and associativity against AMAT.
18 min
Virtual Memory & I/O
- 49Virtual Memory & Page Tables
Separate virtual from physical addresses and walk the role of page tables in translation.
14 min - 50TLBs & Fast Address Translation
Use Translation Lookaside Buffers to avoid walking page tables on every access.
13 min - 51Memory-Mapped I/O
Control devices by reading/writing reserved addresses - unify CPU access paths.
11 min - 52Interrupts vs Polling
Compare busy-wait device service with interrupt-driven I/O and when each is appropriate.
12 min - 53Direct Memory Access (DMA)
Let devices transfer bulk data to memory without per-word CPU babysitting.
13 min - 54Project: DMA Controller Simulator
Build an automated transfer module that models DMA setup, bus grant, and completion interrupt.
18 min
Multicore & Parallelism
- 55Amdahl's Law
Quantify the speedup ceiling when only part of a workload can be parallelized.
12 min - 56Flynn's Taxonomy
Classify SISD, SIMD, MISD, and MIMD - and map them to CPUs, GPUs, and clusters.
12 min - 57SIMD, MIMD & Multicore Processors
Differentiate data-level parallelism from thread-level parallelism in practical terms.
13 min - 58Cache Coherence: MSI & MESI
Track shared line states across cores so every reader sees a consistent memory story.
15 min - 59Project: The Coherence Arbiter
Simulate a multi-core environment tracking shared line state changes across processing cores.
18 min
Capstone & Exam Prep
- 60Hardware-Software Co-Design
Think jointly about ISA, compiler scheduling, and microarchitecture when chasing performance per watt.
12 min - 61Time-Space-Power Trade-offs
Practice SIU engineering discipline: every design choice balances latency, area, and energy.
12 min - 62Grand Challenges in Multicore Scaling
Tackle memory bus contention, synchronization overhead, and diminishing returns at scale.
13 min - 63Final Review: Assembly & Datapath
Comprehensive drill on ISA encodings, hand translation, and single-cycle control before the final.
15 min - 64Final Review: Pipelines, Hazards & Caches
Drill hazard charts, forwarding, AMAT, and associativity questions for the comprehensive final.
16 min - 65Capstone: Architecture Synthesis
Integrate ISA, datapath, pipeline, cache, and I/O into one coherent system narrative and portfolio packet.
20 min