Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Course objectives, participant profile alignment, and success metrics
- High-level migration methodologies and risk assessment
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Lakehouse concepts, Delta Lake introduction, and Databricks architecture overview
- Differences between SMP and MPP architectures and their migration implications
- Medallion (Bronze→Silver→Gold) design principles and Unity Catalog overview
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook format
- Mapping temporary tables and cursors to DataFrame transformations
- Validation and comparison against original output results
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning mechanisms, and time travel features
- Auto Loader, MERGE INTO patterns, upsert operations, and schema evolution
- OPTIMIZE, VACUUM, Z-ORDER, partitioning strategies, and storage tuning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementation of Auto Loader ingestion and MERGE workflows
- Application of OPTIMIZE, Z-ORDER, and VACUUM commands with result validation
- Evaluation of read/write performance enhancements
Day 3 — SQL in Databricks, Performance & Debugging
- Analytical SQL capabilities: window functions, higher-order functions, and JSON/array processing
- Interpreting the Spark UI, DAGs, shuffles, stages, and tasks to diagnose bottlenecks
- Query optimization techniques: broadcast joins, hints, caching, and reducing spills
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring a resource-intensive SQL process into optimized Spark SQL
- Utilizing Spark UI traces to detect and resolve data skew and shuffle issues
- Conducting before/after benchmarks and documenting tuning procedures
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning methods
- Converting loops and cursors into vectorized DataFrame operations
- Modular code structure, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Refactoring a procedural ETL script into modular PySpark notebooks
- Incorporating parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integration with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retry mechanisms, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Unity Catalog governance, data lineage, and access control best practices
- Cost management, cluster sizing, autoscaling, and job concurrency patterns
- Deployment checklists, rollback strategies, and runbook creation
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration outcomes and key learnings
- Gap analysis, recommended follow-up actions, and handover of training materials
- References, further learning pathways, and support options
Requirements
- A solid grasp of data engineering principles
- Proficiency with SQL and stored procedures (Synapse / SQL Server)
- Familiarity with ETL orchestration concepts (ADF or equivalent tools)
Target Audience
- Technology managers with a data engineering foundation
- Data engineers shifting from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing the adoption of Databricks