Course Outline
Databricks Platform and Lakehouse Essentials
- Architecture and components of the Databricks Lakehouse
- Managing workspaces and catalogs
Databricks Workspace and Notebook Development
- Navigating the workspace and developing with notebooks
- Organizing code into reusable notebook structures
Apache Spark Architecture and Execution Model
- Spark runtime architecture and execution mechanics
- Lazy evaluation and the directed acyclic graph (DAG)
PySpark DataFrames and the DataFrame API
- DataFrame abstractions and schema definitions
- Fundamental DataFrame operations and column expressions
Converting SQL to PySpark DataFrames
- Mapping core SQL clauses to DataFrame operations
- Utilizing window functions and aggregations in PySpark
Data Ingestion and Output in Databricks
- Reading data from standard file and database sources
- Writing and partitioning data within the Lakehouse
Delta Lake and Table Administration
- Delta tables and ACID transaction support
- Time travel capabilities and schema evolution
Data Cleansing and Transformation Strategies
- Data cleaning techniques and type casting
- Creating reusable transformation logic
User-Defined Functions and Modular Programming
- Python UDFs and pandas UDFs
- Encapsulating procedural logic into modular functions
Performance Tuning and Optimization Techniques
- Strategies for partitioning and caching
- Identifying bottlenecks using the Spark UI
Basics of Structured Streaming
- Distinguishing between batch and streaming processing models
- Working with streaming DataFrames and simple aggregations
Databricks Jobs and Workflow Orchestration
- Scheduling notebooks as automated jobs and tasks
- Constructing multi-step workflows with defined dependencies
Unity Catalog and Data Governance
- Unity Catalog architecture and namespace management
- Access control mechanisms and data lineage tracking
Testing, Debugging, and Production Standards
- Performing unit tests on PySpark logic
- Debugging processes and adhering to code quality standards
Comprehensive Financial Services Applications
- Developing a complete banking ETL pipeline
- Converting legacy SQL processes to PySpark
Migrating SQL Workloads to PySpark
- Migration strategies and planning approaches
- Step-by-step conversion of SQL workflows to PySpark
Requirements
- Proficiency in Python programming, covering functions and data types
- Knowledge of SQL concepts, such as joins, aggregations, and subqueries
- No prior experience with Databricks or PySpark is necessary
Target Audience
- Data engineers, data analysts, and data professionals
- Teams transitioning existing SQL-based workflows to Databricks and PySpark
Testimonials (1)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.