Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction, Objectives, and Migration Strategy
- Course objectives, alignment with participant profiles, and definition of success metrics
- Overview of high-level migration approaches and associated risk factors
- Configuration of workspaces, repositories, and laboratory datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Distinctions between SMP and MPP models and their impact on migration
- Medallion (Bronze→Silver→Gold) design principles and an introduction to Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure to a notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Validation and comparison of outputs against the original implementation
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Storage optimization techniques: OPTIMIZE, VACUUM, Z-ORDER, and partitioning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM; verifying results
- Evaluating improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, and JSON/array processing
- Analyzing Spark UI, DAGs, shuffles, stages, and tasks to diagnose bottlenecks
- Query optimization strategies: broadcast joins, hints, caching, and reducing spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Utilizing Spark UI traces to detect and resolve skew and shuffle issues
- Conducting before/after benchmarks and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization techniques, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Implementing parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error handling
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment notes
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, data lineage, and access control best practices
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and operational runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations covering migration work and key takeaways
- Gap analysis, recommended follow-up actions, and handover of training materials
- References, further learning paths, and support options
Requirements
- Foundational knowledge of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Understanding of ETL orchestration principles (ADF or equivalent tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers shifting from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption
35 Hours