AWS Data Engineering Expert (AI L2) — S3 | Glue | EMR | Kinesis | Redshift | Amazon Bedrock | Bedrock Knowledge Bases
Duration: 59 Hours
Target Level: L2 (Intermediate)
Delivery Mode: Classroom / Online (ILT)
AI Platform: Amazon Bedrock
Course Overview
This course equips Data Engineering & Analytics professionals with core AWS data engineering skills and the AI layer required for L2 proficiency. It is designed for practitioners who already work on AWS and want to master enterprise-grade data pipelines alongside modern Gen AI capabilities — specifically Amazon Bedrock, RAG systems, and AI-enriched pipeline patterns.
The course follows a learn-by-doing philosophy. Every module includes hands-on labs and flexible case study scenarios that trainers can adapt to their domain context.
L2 Outcome Statement
- Design and build end-to-end AWS data pipelines (ingestion → processing → governance)
- Integrate Amazon Bedrock into data engineering workflows for enrichment and automation
- Architect and deploy enterprise RAG systems using Bedrock Knowledge Bases and OpenSearch
- Embed Gen AI capabilities into both batch (Glue/EMR) and streaming (Kinesis/Lambda) pipelines
- Contribute to internal AI accelerators and POCs using AWS-native AI services
Prerequisites
- 1–2 years of hands-on experience working on the AWS platform in a data engineering or analytics role
- Working knowledge of Amazon S3 and at least one processing service (Glue, EMR, or Athena)
- Familiarity with Python or PySpark for data processing
- Basic SQL proficiency
- Understanding of data pipeline concepts (ETL/ELT, batch vs streaming)
Helpful but not mandatory: Exposure to Gen AI/LLM concepts (prompt engineering, GPT/Claude APIs) · Experience with AWS CDK, CloudFormation, or CI/CD pipelines
Course Structure at a Glance
| # | Module | Hours | Track |
|---|---|---|---|
| 1 | AWS Architecture & Storage Fundamentals | 8 | Data Engineering |
| 2 | Data Ingestion & Integration | 8 | Data Engineering |
| 3 | Data Transformation & Processing | 8 | Data Engineering |
| 4 | Pipeline Orchestration & Workflow Automation | 6 | Data Engineering |
| 5 | Governance, Security & Compliance | 6 | Data Engineering |
| 6 | AWS Gen AI Landscape & Amazon Bedrock Foundations | 7 | Gen AI |
| 7 | Bedrock with Redshift & Gen AI in Data Pipelines | 8 | Gen AI |
| 8 | Embeddings, Vector Stores & Bedrock Knowledge Bases | 8 | Gen AI |
| TOTAL DURATION | 59 Hours | ||
Data Engineering modules build the AWS platform foundation. Gen AI modules apply Amazon Bedrock and AI services in a data-engineering context. Both tracks run in sequence and are equally mandatory.
Detailed Course Modules
Part A — AWS Data Engineering Foundation
Module 1: AWS Architecture & Storage Fundamentals | 8 hrs
Topics Covered: AWS global infrastructure (regions, AZs, VPCs) · Core storage services (S3, EBS, EFS) · S3 data lake architecture (raw/curated/consumption zones) · AWS Glue Data Catalog (databases, tables, crawlers) · Storage security (IAM, bucket policies, KMS, VPC endpoints) · Performance tuning and file formats · Lake Formation permissions
Hands-On Labs: Provision an S3 data lake with 3-zone folder structure and bucket policies · Configure a Glue crawler to catalog a Parquet dataset · Set up Lake Formation column-level permissions
Case Study: Design a zone architecture with governance and access control for a retail org migrating an on-prem warehouse to an S3 data lake.
Module 2: Data Ingestion & Integration | 8 hrs
Topics Covered: AWS Glue ETL (jobs, triggers, workflows, bookmarks) · Amazon Kinesis (Data Streams, Firehose, Data Analytics) · Amazon MSK · AWS DMS (full load + CDC) · Amazon AppFlow · AWS DataSync · Event-driven ingestion with SNS, SQS, Lambda
Hands-On Labs: Build a Kinesis Firehose pipeline streaming data to S3 (Parquet) · Configure a DMS CDC job from RDS PostgreSQL to S3 · Set up an MSK topic with a Lambda consumer
Case Study: Design an ingestion layer handling mixed batch and streaming IoT sensor data for a manufacturing firm.
Module 3: Data Transformation & Processing | 8 hrs
Topics Covered: AWS Glue Studio (visual ETL, DynamicFrames) · Amazon EMR (clusters, Spark) · EMR Serverless vs Glue Spark · Spark optimization (partitioning, broadcast joins, AQE) · Delta Lake/Apache Iceberg on S3 · Amazon Athena · AWS Glue Data Quality (DQDL rules)
Hands-On Labs: Build a Glue Spark job transforming raw JSON to Parquet with schema enforcement · Run an EMR Serverless Spark job with Delta Lake · Write Athena queries with partitioning and cost optimization
Case Study: Build a near-real-time sales KPI pipeline for a logistics company combining batch and streaming layers on AWS.
Module 4: Pipeline Orchestration & Workflow Automation | 6 hrs
Topics Covered: AWS Step Functions (state machines, parallel execution) · Amazon MWAA (Managed Airflow) · AWS Glue Workflows · EventBridge · Lambda-based micro-orchestration · CloudWatch monitoring · IaC for pipelines (CDK/CloudFormation)
Hands-On Labs: Build a Step Functions workflow orchestrating Glue → Athena → SNS alert · Deploy an Airflow DAG on MWAA triggering an EMR Serverless job · Set up an EventBridge rule on S3 object arrival
Case Study: Design a fault-tolerant orchestrated pipeline for a daily financial reconciliation process with retry logic and alerting.
Module 5: Governance, Security & Compliance | 6 hrs
Topics Covered: AWS Lake Formation (fine-grained access, tag-based policies) · AWS IAM roles and permission boundaries · AWS Macie (PII detection) · Glue Data Catalog governance · KMS key management · GDPR/HIPAA patterns · AWS Config and CloudTrail
Hands-On Labs: Configure Lake Formation tag-based access control · Enable Macie and review PII findings · Set up CloudTrail + Athena to audit data access logs
Case Study: Implement a data governance framework for a BFSI regulatory audit, mapping lineage from raw ingestion to reporting.
Part B — AWS Gen AI for Data Engineers
Module 6: AWS Gen AI Landscape & Amazon Bedrock Foundations | 7 hrs
Topics Covered: AWS Gen AI landscape (Bedrock, Amazon Q, Comprehend) · Bedrock architecture and model providers (Anthropic Claude, Meta Llama, Mistral, Amazon Titan) · Bedrock API (InvokeModel, Converse API, streaming) · Foundation model selection · Prompt design for structured data tasks · Token management and cost optimization · Responsible AI and Guardrails · Bedrock Model Evaluation
Hands-On Labs: Deploy a Claude model via Bedrock and call via Boto3 · Write a prompt to extract structured schema from pipeline logs · Configure Bedrock Guardrails · Run a Bedrock model evaluation job comparing two foundation models
Case Study: Use Bedrock to auto-generate documentation from Glue job scripts, and run model evaluation to select the best-fit model.
Module 7: Bedrock with Redshift & Gen AI in Data Pipelines | 8 hrs
Topics Covered: Amazon Redshift architecture (clusters, Serverless, RA3) · Bedrock + Redshift integration (text-to-SQL) · Amazon Q in QuickSight and Redshift · Embedding Gen AI in batch pipelines (Glue + Bedrock) and streaming pipelines (Kinesis + Lambda + Bedrock) · Anomaly detection with Bedrock · Cost and latency considerations
Hands-On Labs: Build a Glue job calling Bedrock to classify and tag records in S3 · Implement a Lambda function invoking Bedrock for real-time sentiment tagging on a Kinesis stream · Use Amazon Q to generate SQL over Redshift Serverless
Case Study: Design an AI-enriched batch pipeline (S3 → Glue → Bedrock → Redshift) for auto-categorizing an e-commerce product catalogue.
Module 8: Embeddings, Vector Stores & Bedrock Knowledge Bases | 8 hrs
Topics Covered: Embeddings and semantic vs keyword search · Bedrock embedding models (Titan, Cohere Embed) · Vector store options (OpenSearch Serverless, pgvector on Aurora, MemoryDB) · Chunking strategies · Bedrock Knowledge Bases (managed RAG with S3 + OpenSearch) · Metadata filtering · Hybrid search (vector + BM25) · Similarity search mechanics
Hands-On Labs: Generate embeddings via Titan Embeddings · Index into OpenSearch Serverless and run similarity queries · Create a Bedrock Knowledge Base backed by S3 and OpenSearch · Attach metadata attributes and implement filter expressions
Case Study: Design an embedding pipeline and Knowledge Base with metadata filtering for a semantic search engine over 10,000+ internal support tickets.
Delivery Guidelines
- All labs run on an AWS sandbox account or the learner’s org AWS subscription
- Designed for Instructor-Led Training (ILT) — classroom or virtual/online
- Case studies are industry-agnostic, adaptable to BFSI, Retail, Healthcare, or Telecom
- Recommended batch size: 15–20 learners
- Online delivery uses breakout rooms for labs and screen-share for demos
Tools & Environment
- AWS Console + CLI/CDK
- Amazon S3 + Lake Formation
- AWS Glue + Glue Studio
- Amazon EMR Serverless
- Amazon Kinesis / MSK
- Amazon Redshift Serverless
- Amazon Bedrock
- Amazon OpenSearch Serverless
- AWS Step Functions / MWAA
- VS Code + Python (Boto3)