Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Goals, and Migration Strategy
- Course objectives, alignment with participant profiles, and success metrics
- Overview of high-level migration approaches and associated risks
- Configuring workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, an introduction to Delta Lake, and Databricks architecture
- Differences between SMP and MPP models and their impact on migration
- Medallion (Bronze→Silver→Gold) architecture and an overview of Unity Catalog
Day 1 Lab — Converting a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Translating temp tables and cursors into DataFrame transformations
- Validating results and comparing them against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Techniques for OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization
Day 2 Lab — Incremental Ingestion & Optimization
- Building Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM; verifying the results
- Evaluating read/write performance gains
Day 3 — SQL in Databricks, Performance & Debugging
- Advanced SQL features: window functions, higher-order functions, and JSON/array manipulation
- Interpreting the Spark UI, DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Query optimization strategies: broadcast joins, hints, caching, and reducing spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Using Spark UI traces to locate and resolve skew and shuffle problems
- Conducting before/after benchmarks and documenting tuning actions
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning tactics
- Converting loops and cursors into vectorized DataFrame operations
- Modularization, UDFs/pandas UDFs, widgets, and creating reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Performing code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Building a Complete End-to-End Pipeline
- Constructing a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, lineage, and best practices for access controls
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations showcasing migration work and key takeaways
- Gap analysis, suggestions for follow-up activities, and distribution of training materials
- Providing references, further learning paths, and support options
Requirements
- A solid understanding of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Familiarity with ETL orchestration principles (ADF or equivalent tools)
Target Audience
- Technology managers with a data engineering background
- Data engineers looking to transition procedural OLAP logic to Lakehouse patterns
- Platform engineers tasked with overseeing Databricks adoption