A framework-agnostic system design fieldbook built around production contracts, protocol mechanics, failure recovery, capacity models, and evidence from papers and operating systems.
Read the interactive fieldbook
·
Download PDF or EPUB
"Most system design resources are unorganized and overly simple. This repository aims to change that."
Contracts before components - Start with invariants and observable promises
Depth over breadth - Follow mechanisms through overload, failure, recovery, and migration
Single-owner chapters - Cross-reference neighboring concepts instead of repeating generic checklists
Framework-agnostic reasoning - Use products as evidence, not as the architecture
Quantitative honesty - Source reported figures and label illustrative assumptions
Verifiable decisions - Include failure traces, observability, rollout, and tests
Evidence labels used throughout the fieldbook:
Documented - A cited primary source states the behavior for its named version or publication scope.
Inference - The conclusion follows from stated mechanisms and assumptions; it is not attributed to an undisclosed deployment.
Reference design - A reusable architecture proposed by this book; thresholds and worked numbers remain illustrative until measured.
Part 2: Distributed Databases
Part 7: Real-Time Systems
Part 11: Observability and Operations
Part 12: Service Connectivity and APIs
Batch Execution: DAGs, Shuffle, and Safe Reprocessing
Stream Execution: Time, State, Recovery, and Backpressure
Change Data Capture: Snapshot, Tail, Apply, and Repair
Lakehouse Table Formats: Snapshots, Commits, and Maintenance
ML System Fundamentals : control planes, training/serving parity, maturity model
Feature Stores : point-in-time correctness, offline/online parity, freshness SLOs
Model Serving : latency budgets, batching, routing, degradation ladders
Model Monitoring : drift detection, label delay, slice guardrails, triage flow
Training Pipelines : TFX/Kubeflow/Airflow DSLs, reproducibility, promotion gates, distributed training
Model Deployment and Rollouts : canary sizing, shadow isolation, kill switches, threshold migration, rollback playbooks
Recommendation Systems : two-tower retrieval, Wide & Deep, ANN indexing, DPP diversity, exploration, cold start
Online Experiments : power analysis, SRM checks, CUPED, sequential testing, interleaving
ML Risk and Governance : risk tiers, model cards, proxy detection, audit logs, incident response, lifecycle
Label and Ground-Truth Systems : label delay, selective labels, human review, weak supervision, append-only label stores
Dataset Management and Versioning : immutable snapshots, manifests, split reproducibility, lineage, privacy deletion
Offline Evaluation and Metric Design : metric selection, thresholds, calibration, leakage, slices, uncertainty
Model Registry and ML Metadata : artifact identity, serving contracts, lifecycle states, promotion gates, rollback metadata
ML Capacity and Cost Planning : GPU sizing, batching, feature fanout, training cost, queueing, headroom
Distributed Training Internals : parallelism topologies, collectives, ZeRO/FSDP, MFU, failure math, checkpointing
Agent Fundamentals : durable loops, action identity, authority, evidence, and bounded autonomy
Orchestration Patterns : deterministic workflows versus adaptive branches, typed state graphs, latency and cost composition
Multi-Agent Systems : delegation, shared state, message semantics, scheduling, and correlated failure
RAG Patterns : corpus publication, authorization-aware retrieval, evidence packets, grounded generation, and citation provenance
LLM Infrastructure : gateway and model control planes, regional admission/routing/streaming, provider and self-hosted fleets
Prompt Engineering : versioned request compilation, structured outputs, tool schemas, prompt-injection test boundaries
Fine-Tuning Patterns : method choice, dataset/release identity, serving adapters, privacy and evaluation
Context Management : context revisions, compaction, memory, token budgets, and concurrency
Harness Engineering : action transactions, permission boundaries, sandbox/secret broker, recovery and rollout
LLM Evaluation and Observability : estimands, datasets, calibrated evaluators, uncertainty, traces and production feedback
GPU Inference Internals : phase rooflines, KV/device memory, kernels, parallelism, and accelerator isolation
Agent Inference : multi-turn session scheduling, KV residency, cache-aware routing, fan-out and task-level goodput
Part 18: Workflow & Job Systems
Workflow System Fundamentals : workload taxonomy, durable state model, ownership and engine selection
Background Jobs and Worker Pools : queue/worker contracts, acknowledgement, concurrency and overload
Distributed Scheduling and Timer Services : timer state, claiming, misfires, recurrence, sharding and clock semantics
Durable Execution and Workflow Engines : event histories, deterministic replay, commands, signals and versioning
DAG Orchestration : dependency graphs, artifact readiness, backfills and critical-path execution
Effect Commit Protocols for Workflows : ambiguous effects, inbox/outbox, idempotency, reconciliation and compensation
Priority Queues, Fairness, and Backpressure : admission, quotas, starvation bounds and tenant scheduling
Leases, Heartbeats, and Recovery : attempt ownership, fencing, timeout detection and reassignment
Workflow Observability and Replay : history/attempt correlation, stuck-state diagnosis, replay and forensic evidence
Part 19: Engineering Systems for Coding Agents
Symbol
Meaning
N
Total nodes/replicas
W
Write quorum size
R
Read quorum size
f
Failures tolerated
Book
Author
Topics
Designing Data-Intensive Applications
Martin Kleppmann
Replication, partitioning, transactions, distributed systems
System Design Interview Vol. 1
Alex Xu
Rate limiting, consistent hashing, key-value stores
System Design Interview Vol. 2
Alex Xu
Real-world systems, proximity services, stock exchange
Database Internals
Alex Petrov
B-trees, LSM trees, storage engines, distributed databases
Understanding Distributed Systems
Roberto Vitillo
Networking, coordination, scalability, resiliency
Building Microservices
Sam Newman
Service decomposition, integration, deployment
Distributed Systems Foundations
Company
Notable Posts
Netflix Tech Blog
Microservices, chaos engineering, streaming
Uber Engineering
Real-time systems, geospatial, scaling
Meta Engineering
TAO, distributed systems, ML infrastructure
Stripe Engineering
API design, idempotency, payments
Cloudflare Blog
Edge computing, DNS, DDoS mitigation
Discord Engineering
Real-time messaging, voice, scaling
Slack Engineering
Messaging architecture, search, reliability
Dropbox Tech Blog
Sync, storage, infrastructure
Pinterest Engineering
Recommendations, search, scaling
LinkedIn Engineering
Kafka, data infrastructure, ML
Twitter Engineering
Timeline, real-time, graph processing
Spotify Engineering
Streaming, personalization, microservices
GitHub Engineering
Git internals, availability, scaling
Shopify Engineering
E-commerce, flash sales, payments
Architecture & Design Patterns
Book
Author
Topics
Clean Architecture
Robert C. Martin
Dependency rule, boundaries, components, frameworks
Patterns of Enterprise Application Architecture
Martin Fowler
Domain logic, data source, web presentation patterns
Domain-Driven Design
Eric Evans
Bounded contexts, aggregates, ubiquitous language
Implementing Domain-Driven Design
Vaughn Vernon
Practical DDD patterns and techniques
Release It!
Michael Nygard
Stability patterns, capacity, networking
Software Architecture: The Hard Parts
Ford, Richards, Sadalage, Dehghani
Trade-off analysis, modularity, decomposition
Fundamentals of Software Architecture
Mark Richards, Neal Ford
Architecture styles, characteristics, decisions
A Philosophy of Software Design
John Ousterhout
Complexity, modules, abstractions, comments
Distributed Systems & Reliability
Book
Author
Topics
Site Reliability Engineering
Google
SLOs, error budgets, toil, monitoring, on-call
The Site Reliability Workbook
Google
Practical SRE implementation
Distributed Systems
Tanenbaum & Van Steen
Processes, communication, naming, coordination
Designing Distributed Systems
Brendan Burns
Patterns for scalable, reliable services
Database Reliability Engineering
Campbell & Majors
Database operations, infrastructure, recovery
Web Scalability for Startup Engineers
Artur Ejsmont
Practical scaling strategies
Book
Author
Topics
Streaming Systems
Akidau, Chernyak, Lax
Watermarks, windows, triggers, exactly-once
Kafka: The Definitive Guide
Shapira, Palino, et al.
Kafka internals, producers, consumers, operations
Making Sense of Stream Processing
Martin Kleppmann
Event sourcing, change capture, stream processing
Data Mesh
Zhamak Dehghani
Decentralized data architecture
The Data Warehouse Toolkit
Ralph Kimball
Dimensional modeling, ETL, BI
Performance & Optimization
Book
Author
Topics
Systems Performance
Brendan Gregg
Linux, observability, methodologies, tools
BPF Performance Tools
Brendan Gregg
Linux BPF observability and tracing
High Performance MySQL
Silvia Botros, Jeremy Tinley
Query optimization, replication, scaling
High Performance Browser Networking
Ilya Grigorik
TCP, UDP, TLS, HTTP/2, WebSocket, WebRTC
Consistency & Transactions
Distributed Data Structures
Scaling Distributed Machine Learning with the Parameter Server - Li et al., 2014
TensorFlow: A System for Large-Scale Machine Learning - Abadi et al., 2016
Hidden Technical Debt in Machine Learning Systems - Sculley et al., 2015
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform - Baylor et al., 2017
Data Validation for Machine Learning - Breck et al., 2019
TensorFlow Serving: Flexible, High-Performance ML Serving - Olston et al., 2017
Deep Neural Networks for YouTube Recommendations - Covington et al., 2016
Wide & Deep Learning for Recommender Systems - Cheng et al., 2016
Sampling-Bias-Corrected Neural Modeling for Large Corpus Item Recommendations - Yi et al., 2019
Diversity-Promoting Recommendation with Determinantal Point Processes - Chen et al., 2018
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models - Rajbhandari et al., 2019
Trustworthy Online Controlled Experiments - Kohavi et al., 2020
Improving the Sensitivity of Online Controlled Experiments by Utilizing Pre-Experiment Data (CUPED) - Deng et al., 2013
Overlapping Experiment Infrastructure: More, Better, Faster Experimentation - Tang et al., 2010
Model Cards for Model Reporting - Mitchell et al., 2019
Datasheets for Datasets - Gebru et al., 2018
A Unified Approach to Interpreting Model Predictions (SHAP) - Lundberg & Lee, 2017
NIST AI Risk Management Framework - NIST, 2023
Container & Orchestration
MIT License