Trainings/PySpark Training Program
Course – Data Engineering

PySpark Training Program

5 Days / 40 HrsBeginner to IntermediateClassroom / Live Virtual
Duration
5 Days / 40 hrs
Level
Beginner to Intermediate
Format
Classroom / Live Virtual
Domain
Data Engineering · Big Data
Training Methodology

Learning by Doing

Every module pairs Spark concepts with a hands-on lab against real datasets, culminating in a full ETL capstone pipeline.

01

Explore

Hands-on labs from Day 1.

02

Experiment

Work against real datasets.

03

Engage

Business reporting use cases.

04

Apply

Capstone ETL pipeline.

Who This Is For

Data EngineersBig Data DevelopersPython Developers Moving into SparkData Analysts Working with Large Datasets

What You’ll Be Able to Do

  • Understand Big Data and distributed computing concepts, and configure Spark environments
  • Develop PySpark applications and process large datasets using DataFrames
  • Write Spark SQL queries and perform ETL operations
  • Optimize Spark jobs for performance using caching, partitioning, and broadcast joins
  • Work with structured and semi-structured data across CSV, JSON, Parquet, and ORC formats
  • Integrate Spark with databases and cloud storage, and build end-to-end data engineering solutions

Prerequisites

  • Basic programming logic and SQL familiarity helpful — a Python refresher module is built into Day 1
  • No prior Apache Spark experience required

Curriculum

Day 1
Big Data & Apache Spark Foundations, Python Refresher
  • Big Data fundamentals, 5 V’s, Hadoop vs Spark ecosystem
  • Spark architecture — driver, executors, cluster manager, worker nodes; Spark modes
  • Python essentials refresher — variables, data types, functions, lambdas, exceptions, file I/O
Labs: Install and configure Spark, create first Spark session, execute shell commands, build mini Python scripts.
Day 2
Getting Started with PySpark, DataFrame Transformations & Actions
  • Creating DataFrames from lists, CSV, JSON, Parquet; DataFrame basics (show, printSchema, describe, count)
  • Narrow & wide transformations (filter, select, withColumn, groupBy, join), actions, column functions
Labs: Load employee dataset, explore schema; rename/derive columns, filter and aggregate department data.
Day 3
Data Cleaning & Data Wrangling, Spark SQL
  • Missing values, duplicates, invalid formats; string manipulation functions
  • Temporary views, SQL queries, aggregate/window/date functions
Labs: Customer data cleansing with validation report; department-wise business reports and salary analytics.
Day 4
Files & Data Sources, Advanced Transformations & Joins
  • File formats (JSON, Parquet, ORC), JDBC connectivity to SQL Server/MySQL/PostgreSQL
  • Joins (inner, left, right, full, cross), window functions, rollup/cube aggregations
Labs: Load from CSV/SQL, export to Parquet/CSV; retail sales ranking, running totals, product performance.
Day 5
Performance Tuning, Structured Streaming & Capstone ETL
  • Lazy evaluation, DAG execution, caching, partitioning, broadcast joins, Spark UI
  • Structured streaming basics; Capstone: Retail Data Engineering Pipeline (ingestion, transformation, analytics, optimization)
Labs: Before/after optimization comparison; real-time log analysis; full capstone pipeline build.

Delivery Details

  • Delivered as classroom or live virtual instructor-led — scheduled around your team
  • 30% Lecture / 70% Hands-On Labs, built around employee, customer, and retail sales datasets

Request This Program

Email