System Design & Data Architecture
Master the architectural foundations required for Principal Data Engineers and Enterprise Data Architects. Covers storage engine internals (LSM-trees vs Columnar), distributed event logs (Kafka), Lakehouse table formats (Iceberg/Delta), dimensional modeling (Kimball), Change Data Capture (CDC), stream time semantics, and decentralized Data Mesh.
Every module is built on the foundational principles of classic engineering texts: Martin Kleppmann's Designing Data-Intensive Applications, Ralph Kimball's The Data Warehouse Toolkit, Tyler Akidau's Streaming Systems, and Zhamak Dehghani's Data Mesh.
Architecture Reference Curriculum
Explore comprehensive modules categorized by engineering domain.
Foundations of Scalable Systems
Reliability, Scalability, and Maintainability: The three pillars of modern software systems.
Data Models and Query Languages
Relational vs Document databases, Graph-like data models, and the impedance mismatch.
Storage Engines: Row vs Columnar vs LSM-Trees
How databases physically store bits on disk. B-Trees, LSM-Trees (SSTables & Memtables), and Columnar formats (Parquet/ORC).
Replication Strategies
Single-leader, Multi-leader, and Leaderless replication. Handling replication lag and eventual consistency.
Partitioning and Sharding
How to split very large datasets across multiple machines. Key-range vs Hash partitioning.
Distributed Transactions & Consensus
ACID vs BASE, Two-Phase Commit (2PC), Paxos/Raft consensus algorithms, and the Saga pattern for distributed microservices.
Caching Strategies
Cache-aside, Read-through, Write-through, Write-back. Eviction policies and caching layers.
Message Brokers & Event Streaming (Kafka Deep Dive)
Append-only commit logs vs message queues. Kafka partitions, consumer rebalancing, consumer groups, and exactly-once processing (EOS).
Batch vs Stream Processing: Lambda, Kappa & Event Time
Lambda vs Kappa architectures, event time vs processing time, watermarking, and windowing strategies in Apache Flink and Spark Streaming.
Change Data Capture (CDC) & The Outbox Pattern
Log-based CDC vs query polling. Debezium, Kafka Connect, and solving dual-write distributed data consistency with the Transactional Outbox Pattern.
Data Warehousing & Dimensional Modeling
Kimball vs Inmon methodologies. Star schema vs Snowflake schema. Fact types and Slowly Changing Dimensions (SCD Types 1, 2, 3, 6).
Lakehouse Architecture & Modern Table Formats
Overcoming Hive-style Data Lake limitations. Apache Iceberg, Delta Lake, Apache Hudi, ACID transactions over object storage, and the Medallion Architecture.
Data Pipeline Orchestration & Idempotency
Directed Acyclic Graphs (DAGs), state management, idempotent pipeline execution, backfilling, and comparing Airflow, Dagster, and Prefect.
Data Observability, Quality & Contracts
The 5 pillars of data observability. Automated quality testing, end-to-end lineage, schema enforcement, and Data Contracts between software and data teams.
Data Mesh & Modern Architectural Paradigms
Decentralized data ownership vs centralized data teams. The four principles of Data Mesh: Domain Ownership, Data as a Product, Self-Serve Platform, and Federated Governance.
Microservices and API Gateways
Monoliths vs Microservices. The role of API Gateways, service discovery, and circuit breakers.