Course – Data Engineering
PySpark Training Program
5 Days / 40 HrsBeginner to IntermediateClassroom / Live Virtual
Duration
5 Days / 40 hrs
Level
Beginner to Intermediate
Format
Classroom / Live Virtual
Domain
Data Engineering · Big Data
Training Methodology
Learning by Doing
Every module pairs Spark concepts with a hands-on lab against real datasets, culminating in a full ETL capstone pipeline.
01
Explore
Hands-on labs from Day 1.
02
Experiment
Work against real datasets.
03
Engage
Business reporting use cases.
04
Apply
Capstone ETL pipeline.
Who This Is For
What You’ll Be Able to Do
- Understand Big Data and distributed computing concepts, and configure Spark environments
- Develop PySpark applications and process large datasets using DataFrames
- Write Spark SQL queries and perform ETL operations
- Optimize Spark jobs for performance using caching, partitioning, and broadcast joins
- Work with structured and semi-structured data across CSV, JSON, Parquet, and ORC formats
- Integrate Spark with databases and cloud storage, and build end-to-end data engineering solutions
Prerequisites
- Basic programming logic and SQL familiarity helpful — a Python refresher module is built into Day 1
- No prior Apache Spark experience required
Curriculum
Day 1
Big Data & Apache Spark Foundations, Python Refresher
- Big Data fundamentals, 5 V’s, Hadoop vs Spark ecosystem
- Spark architecture — driver, executors, cluster manager, worker nodes; Spark modes
- Python essentials refresher — variables, data types, functions, lambdas, exceptions, file I/O
Labs: Install and configure Spark, create first Spark session, execute shell commands, build mini Python scripts.
Day 2
Getting Started with PySpark, DataFrame Transformations & Actions
- Creating DataFrames from lists, CSV, JSON, Parquet; DataFrame basics (show, printSchema, describe, count)
- Narrow & wide transformations (filter, select, withColumn, groupBy, join), actions, column functions
Labs: Load employee dataset, explore schema; rename/derive columns, filter and aggregate department data.
Day 3
Data Cleaning & Data Wrangling, Spark SQL
- Missing values, duplicates, invalid formats; string manipulation functions
- Temporary views, SQL queries, aggregate/window/date functions
Labs: Customer data cleansing with validation report; department-wise business reports and salary analytics.
Day 4
Files & Data Sources, Advanced Transformations & Joins
- File formats (JSON, Parquet, ORC), JDBC connectivity to SQL Server/MySQL/PostgreSQL
- Joins (inner, left, right, full, cross), window functions, rollup/cube aggregations
Labs: Load from CSV/SQL, export to Parquet/CSV; retail sales ranking, running totals, product performance.
Day 5
Performance Tuning, Structured Streaming & Capstone ETL
- Lazy evaluation, DAG execution, caching, partitioning, broadcast joins, Spark UI
- Structured streaming basics; Capstone: Retail Data Engineering Pipeline (ingestion, transformation, analytics, optimization)
Labs: Before/after optimization comparison; real-time log analysis; full capstone pipeline build.
Delivery Details
- Delivered as classroom or live virtual instructor-led — scheduled around your team
- 30% Lecture / 70% Hands-On Labs, built around employee, customer, and retail sales datasets