Data Engineer
Master distributed computing with Apache Spark, Airflow orchestration, dbt, SQL, and Snowflake/BigQuery.
1. Beginner
Advanced SQL and Python data structures.
Advanced SQL & Window Functions
PARTITION BY, CTEs, ROW_NUMBER vs DENSE_RANK, query execution plans.
2. Foundation
Data modeling paradigms and file formats.
Data Modeling (Kimball Star & Snowflake Schema)
Fact vs dimension tables, slowly changing dimensions (SCD 1, 2, 3), Parquet vs ORC.
3. Core Skills
Distributed data processing engines.
Apache Spark & PySpark
RDDs, DataFrames, catalyst optimizer, shuffles, partition skew mitigation.
4. Framework
Pipeline orchestration and data transformation.
Apache Airflow & dbt
DAG design, idempotency in pipelines, data testing, incremental models.
5. Real Projects
End-to-end data pipeline from raw events to analytical warehouse.
Streaming Financial Ticker Pipeline
Kafka stream ingestion, Spark streaming aggregation, Delta Lake storage.
6. Advanced
Data mesh, lakehouses, and data governance.
Modern Data Lakehouse (Iceberg / Delta)
Time travel queries, schema evolution, catalog management, data lineage.
7. Interview Prep
Pipeline design, data debugging, and SQL screening.
Data Pipeline System Design
Designing metrics aggregation for 1B daily events, late arriving data.