Enterprise System Design
16 Deep-Dive Modules

System Design & Data Architecture

Master the architectural foundations required for Principal Data Engineers and Enterprise Data Architects. Covers storage engine internals (LSM-trees vs Columnar), distributed event logs (Kafka), Lakehouse table formats (Iceberg/Delta), dimensional modeling (Kimball), Change Data Capture (CDC), stream time semantics, and decentralized Data Mesh.

Storage & LSM
Kafka & Stream Time
Lakehouse & Iceberg
Observability & Mesh
Seminal Engineering Literature
Distilled Architectural Knowledge

Every module is built on the foundational principles of classic engineering texts: Martin Kleppmann's Designing Data-Intensive Applications, Ralph Kimball's The Data Warehouse Toolkit, Tyler Akidau's Streaming Systems, and Zhamak Dehghani's Data Mesh.

Architecture Reference Curriculum

Explore comprehensive modules categorized by engineering domain.

10 min read
Distributed Systems

Foundations of Scalable Systems

Reliability, Scalability, and Maintainability: The three pillars of modern software systems.

Ref: Designing Data-Intensive Applications (DDIA), Chapter 1
Read Deep-Dive
12 min read
Storage & Modeling

Data Models and Query Languages

Relational vs Document databases, Graph-like data models, and the impedance mismatch.

Ref: Designing Data-Intensive Applications (DDIA), Chapter 2
Read Deep-Dive
15 min read
Storage & Modeling

Storage Engines: Row vs Columnar vs LSM-Trees

How databases physically store bits on disk. B-Trees, LSM-Trees (SSTables & Memtables), and Columnar formats (Parquet/ORC).

Ref: Designing Data-Intensive Applications, Chapter 3 & Database Internals (Alex Petrov)
Read Deep-Dive
15 min read
Distributed Systems

Replication Strategies

Single-leader, Multi-leader, and Leaderless replication. Handling replication lag and eventual consistency.

Ref: Designing Data-Intensive Applications (DDIA), Chapter 5
Read Deep-Dive
11 min read
Distributed Systems

Partitioning and Sharding

How to split very large datasets across multiple machines. Key-range vs Hash partitioning.

Ref: Designing Data-Intensive Applications (DDIA), Chapter 6
Read Deep-Dive
16 min read
Distributed Systems

Distributed Transactions & Consensus

ACID vs BASE, Two-Phase Commit (2PC), Paxos/Raft consensus algorithms, and the Saga pattern for distributed microservices.

Ref: Designing Data-Intensive Applications, Chapters 7 & 9
Read Deep-Dive
8 min read
Distributed Systems

Caching Strategies

Cache-aside, Read-through, Write-through, Write-back. Eviction policies and caching layers.

Ref: System Design Interview – An Insider's Guide (Alex Xu)
Read Deep-Dive
15 min read
Streaming & Integration

Message Brokers & Event Streaming (Kafka Deep Dive)

Append-only commit logs vs message queues. Kafka partitions, consumer rebalancing, consumer groups, and exactly-once processing (EOS).

Ref: Kafka: The Definitive Guide (Gwen Shapira et al.) & DDIA, Chapter 11
Read Deep-Dive
15 min read
Streaming & Integration

Batch vs Stream Processing: Lambda, Kappa & Event Time

Lambda vs Kappa architectures, event time vs processing time, watermarking, and windowing strategies in Apache Flink and Spark Streaming.

Ref: Streaming Systems (Tyler Akidau et al.) & DDIA, Chapters 10 & 11
Read Deep-Dive
12 min read
Streaming & Integration

Change Data Capture (CDC) & The Outbox Pattern

Log-based CDC vs query polling. Debezium, Kafka Connect, and solving dual-write distributed data consistency with the Transactional Outbox Pattern.

Ref: Designing Data-Intensive Applications, Chapter 11 & Debezium Architectural Guide
Read Deep-Dive
16 min read
Storage & Modeling

Data Warehousing & Dimensional Modeling

Kimball vs Inmon methodologies. Star schema vs Snowflake schema. Fact types and Slowly Changing Dimensions (SCD Types 1, 2, 3, 6).

Ref: The Data Warehouse Toolkit (Ralph Kimball & Margy Ross)
Read Deep-Dive
15 min read
Storage & Modeling

Lakehouse Architecture & Modern Table Formats

Overcoming Hive-style Data Lake limitations. Apache Iceberg, Delta Lake, Apache Hudi, ACID transactions over object storage, and the Medallion Architecture.

Ref: Delta Lake & Apache Iceberg Architecture Specifications
Read Deep-Dive
14 min read
Streaming & Integration

Data Pipeline Orchestration & Idempotency

Directed Acyclic Graphs (DAGs), state management, idempotent pipeline execution, backfilling, and comparing Airflow, Dagster, and Prefect.

Ref: Fundamentals of Data Engineering (Joe Reis & Matt Housley) & Data Pipelines Pocket Reference
Read Deep-Dive
13 min read
Architecture & Governance

Data Observability, Quality & Contracts

The 5 pillars of data observability. Automated quality testing, end-to-end lineage, schema enforcement, and Data Contracts between software and data teams.

Ref: Fundamentals of Data Engineering & Data Observability: The Definitive Guide (Monte Carlo)
Read Deep-Dive
14 min read
Architecture & Governance

Data Mesh & Modern Architectural Paradigms

Decentralized data ownership vs centralized data teams. The four principles of Data Mesh: Domain Ownership, Data as a Product, Self-Serve Platform, and Federated Governance.

Ref: Data Mesh: Delivering Data-Driven Value at Scale (Zhamak Dehghani)
Read Deep-Dive
9 min read
Architecture & Governance

Microservices and API Gateways

Monoliths vs Microservices. The role of API Gateways, service discovery, and circuit breakers.

Ref: Building Microservices (Sam Newman)
Read Deep-Dive