Distributed Systems

Subpage of Operating Systems

Diving deep into the design and structure of Operating Systems.

Phase 5 — Distributed Systems (Weeks 17–22). Treat this as a mini-course of its own. The Raft paper is mandatory reading. Lamport's original clock paper is short and rewarding. Take your time on consensus.

As you go along, you need to finish the following topics:

  • Distributed Systems — Motivation and Failure Models
  • Remote Procedure Calls (RPC) — Mechanisms and Semantics
  • Distributed File Systems — GFS, HDFS, Ceph
  • Naming and Directory Services
  • Time, Clocks, and Event Ordering
  • Logical Clocks and Vector Clocks
  • Consistency Models — Strong, Sequential, Causal, Eventual
  • The CAP Theorem
  • Consensus — Paxos Explained
  • Consensus — Raft Explained
  • Replication Strategies — Primary-Backup, Chain, Quorum
  • Fault Tolerance and Failure Detection
  • Distributed Transactions and Two-Phase Commit
  • MapReduce and Distributed Computation
  • Gossip Protocols and Epidemic Algorithms

/ Continue

Follow the technical trail.

Use the dense notes as the source material, then move through the guided route, writing, or project proof when you want a cleaner entry point.