Learn data engineering by building it
Seven hands-on courses covering the tools data teams actually run in production — from your first DataFrame to sizing a real cluster. Create a free account to start any course — it's also what saves your progress across devices.
The Spark Field Guide
Distributed data processing with PySpark, from architecture to Catalyst internals.
๐๏ธSQL & Data Modeling
Querying, window functions, and the star-schema modeling every warehouse is built on.
๐๏ธDatabase Systems
What happens under the hood: transactions, isolation, indexing internals, and replication.
๐Apache Airflow
Orchestrating pipelines as DAGs โ scheduling, sensors, retries, and production operations.
๐กApache Kafka
Event streaming: topics, partitions, producers/consumers, and exactly-once delivery.
๐งdbt & the Modern Data Stack
Transformation as SQL + version control โ models, tests, snapshots, and CI/CD for data.
๐๏ธCloud Data Warehouses
Snowflake and BigQuery architecture, cost control, and choosing between them.
๐ฆDocker & Kubernetes
Containerizing pipelines and running Spark/Airflow on Kubernetes in production.