Get in Touch

Course Outline

Introduction, Learning Objectives, and Migration Strategy

  • Course goals, alignment with participant profiles, and success metrics
  • Overview of migration methodologies and associated risk assessments
  • Configuration of workspaces, repositories, and laboratory datasets

Day 1 — Migration Fundamentals and Architectural Concepts

  • Lakehouse principles, Delta Lake overview, and Databricks architecture
  • Contrasting SMP and MPP models and their implications for migration
  • Medallion (Bronze→Silver→Gold) architecture design and Unity Catalog introduction

Day 1 Lab — Converting a Stored Procedure

  • Practical migration of a sample stored procedure into a notebook
  • Mapping temporary tables and cursors to DataFrame transformations
  • Verification and comparison against original output results

Day 2 — Advanced Delta Lake Features & Incremental Loading

  • ACID transactions, commit logs, versioning mechanisms, and time travel
  • Auto Loader implementation, MERGE INTO patterns, upserts, and schema evolution
  • Optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning

Day 2 Lab — Incremental Ingestion & Performance Optimization

  • Building Auto Loader ingestion and MERGE workflows
  • Applying OPTIMIZE, Z-ORDER, and VACUUM commands; result validation
  • Evaluating improvements in read/write performance

Day 3 — SQL in Databricks, Performance Analysis & Debugging

  • Analytical SQL capabilities: window functions, higher-order functions, JSON/array manipulation
  • Interpreting Spark UI: DAGs, shuffles, stages, tasks, and identifying bottlenecks
  • Query optimization strategies: broadcast joins, hints, caching, and reducing spillage

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring complex SQL processes into optimized Spark SQL
  • Utilizing Spark UI traces to detect and resolve skew and shuffle issues
  • Pre- and post-implementation benchmarking and documentation of tuning actions

Day 4 — Applied PySpark: Replacing Procedural Logic

  • Spark execution model: drivers, executors, lazy evaluation, and partitioning methods
  • Transforming loops and cursors into vectorized DataFrame operations
  • Modularization strategies, UDFs/pandas UDFs, widgets, and creating reusable libraries

Day 4 Lab — Refactoring Procedural Scripts

  • Converting procedural ETL scripts into modular PySpark notebooks
  • Implementing parametrization, unit-style testing, and reusable functions
  • Code review sessions and application of best-practice checklists

Day 5 — Orchestration, End-to-End Pipelines & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error management
  • Architecting incremental Medallion pipelines with quality rules and schema validation
  • Integrating with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark logic

Day 5 Lab — Building a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
  • Implementing logging, auditing, retries, and automated validations
  • Executing the full pipeline, verifying outputs, and preparing deployment documentation

Operationalization, Governance, and Production Readiness

  • Unity Catalog governance, data lineage, and access control best practices
  • Cost management, cluster sizing, autoscaling, and job concurrency patterns
  • Deployment checklists, rollback strategies, and creation of operational runbooks

Final Review, Knowledge Transfer, and Future Steps

  • Participant presentations on migration work and key takeaways
  • Gap analysis, recommended follow-up actions, and handover of training materials
  • References, advanced learning paths, and support options

Requirements

  • Foundational knowledge of data engineering principles
  • Practical experience with SQL and stored procedures (Synapse / SQL Server)
  • Proficiency with ETL orchestration concepts (ADF or comparable tools)

Target Audience

  • Technology managers possessing data engineering backgrounds
  • Data engineers shifting from procedural OLAP logic to Lakehouse patterns
  • Platform engineers overseeing Databricks adoption
 35 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories