Get in Touch
 Duration 35 hours

Course Outline

Introduction, Objectives, and Migration Strategy

  • Course objectives, participant profile alignment, and success metrics
  • High-level migration methodologies and risk assessment
  • Configuration of workspaces, repositories, and lab datasets

Day 1 — Migration Fundamentals and Architecture

  • Lakehouse concepts, Delta Lake introduction, and Databricks architecture overview
  • Differences between SMP and MPP architectures and their migration implications
  • Medallion (Bronze→Silver→Gold) design principles and Unity Catalog overview

Day 1 Lab — Translating a Stored Procedure

  • Practical migration of a sample stored procedure into a notebook format
  • Mapping temporary tables and cursors to DataFrame transformations
  • Validation and comparison against original output results

Day 2 — Advanced Delta Lake & Incremental Loading

  • ACID transactions, commit logs, versioning mechanisms, and time travel features
  • Auto Loader, MERGE INTO patterns, upsert operations, and schema evolution
  • OPTIMIZE, VACUUM, Z-ORDER, partitioning strategies, and storage tuning

Day 2 Lab — Incremental Ingestion & Optimization

  • Implementation of Auto Loader ingestion and MERGE workflows
  • Application of OPTIMIZE, Z-ORDER, and VACUUM commands with result validation
  • Evaluation of read/write performance enhancements

Day 3 — SQL in Databricks, Performance & Debugging

  • Analytical SQL capabilities: window functions, higher-order functions, and JSON/array processing
  • Interpreting the Spark UI, DAGs, shuffles, stages, and tasks to diagnose bottlenecks
  • Query optimization techniques: broadcast joins, hints, caching, and reducing spills

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring a resource-intensive SQL process into optimized Spark SQL
  • Utilizing Spark UI traces to detect and resolve data skew and shuffle issues
  • Conducting before/after benchmarks and documenting tuning procedures

Day 4 — Tactical PySpark: Replacing Procedural Logic

  • Spark execution model: driver, executors, lazy evaluation, and partitioning methods
  • Converting loops and cursors into vectorized DataFrame operations
  • Modular code structure, UDFs/pandas UDFs, widgets, and reusable libraries

Day 4 Lab — Refactoring Procedural Scripts

  • Refactoring a procedural ETL script into modular PySpark notebooks
  • Incorporating parametrization, unit-style testing, and reusable functions
  • Conducting code reviews and applying best-practice checklists

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error management
  • Designing incremental Medallion pipelines with quality rules and schema validation
  • Integration with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark

Day 5 Lab — Build a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
  • Implementing logging, auditing, retry mechanisms, and automated validations
  • Executing the full pipeline, validating outputs, and preparing deployment documentation

Operationalization, Governance, and Production Readiness

  • Unity Catalog governance, data lineage, and access control best practices
  • Cost management, cluster sizing, autoscaling, and job concurrency patterns
  • Deployment checklists, rollback strategies, and runbook creation

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations on migration outcomes and key learnings
  • Gap analysis, recommended follow-up actions, and handover of training materials
  • References, further learning pathways, and support options

Requirements

  • A solid grasp of data engineering principles
  • Proficiency with SQL and stored procedures (Synapse / SQL Server)
  • Familiarity with ETL orchestration concepts (ADF or equivalent tools)

Target Audience

  • Technology managers with a data engineering foundation
  • Data engineers shifting from procedural OLAP logic to Lakehouse patterns
  • Platform engineers overseeing the adoption of Databricks

Number of participants


Price per participant

Upcoming Courses

Related Categories