Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction, Learning Objectives, and Migration Strategy
- Course goals, alignment with participant profiles, and success metrics
- Overview of migration methodologies and associated risk assessments
- Configuration of workspaces, repositories, and laboratory datasets
Day 1 — Migration Fundamentals and Architectural Concepts
- Lakehouse principles, Delta Lake overview, and Databricks architecture
- Contrasting SMP and MPP models and their implications for migration
- Medallion (Bronze→Silver→Gold) architecture design and Unity Catalog introduction
Day 1 Lab — Converting a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Verification and comparison against original output results
Day 2 — Advanced Delta Lake Features & Incremental Loading
- ACID transactions, commit logs, versioning mechanisms, and time travel
- Auto Loader implementation, MERGE INTO patterns, upserts, and schema evolution
- Optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning
Day 2 Lab — Incremental Ingestion & Performance Optimization
- Building Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM commands; result validation
- Evaluating improvements in read/write performance
Day 3 — SQL in Databricks, Performance Analysis & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, JSON/array manipulation
- Interpreting Spark UI: DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Query optimization strategies: broadcast joins, hints, caching, and reducing spillage
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Utilizing Spark UI traces to detect and resolve skew and shuffle issues
- Pre- and post-implementation benchmarking and documentation of tuning actions
Day 4 — Applied PySpark: Replacing Procedural Logic
- Spark execution model: drivers, executors, lazy evaluation, and partitioning methods
- Transforming loops and cursors into vectorized DataFrame operations
- Modularization strategies, UDFs/pandas UDFs, widgets, and creating reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting procedural ETL scripts into modular PySpark notebooks
- Implementing parametrization, unit-style testing, and reusable functions
- Code review sessions and application of best-practice checklists
Day 5 — Orchestration, End-to-End Pipelines & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Architecting incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark logic
Day 5 Lab — Building a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, data lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and creation of operational runbooks
Final Review, Knowledge Transfer, and Future Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommended follow-up actions, and handover of training materials
- References, advanced learning paths, and support options
Requirements
- Foundational knowledge of data engineering principles
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Proficiency with ETL orchestration concepts (ADF or comparable tools)
Target Audience
- Technology managers possessing data engineering backgrounds
- Data engineers shifting from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption
35 Hours