Get in Touch
 Duration 35 hours

Course Outline

Databricks Platform and Lakehouse Essentials

  • Architecture and components of the Databricks Lakehouse
  • Managing workspaces and catalogs

Databricks Workspace and Notebook Development

  • Navigating the workspace and developing with notebooks
  • Organizing code into reusable notebook structures

Apache Spark Architecture and Execution Model

  • Spark runtime architecture and execution mechanics
  • Lazy evaluation and the directed acyclic graph (DAG)

PySpark DataFrames and the DataFrame API

  • DataFrame abstractions and schema definitions
  • Fundamental DataFrame operations and column expressions

Converting SQL to PySpark DataFrames

  • Mapping core SQL clauses to DataFrame operations
  • Utilizing window functions and aggregations in PySpark

Data Ingestion and Output in Databricks

  • Reading data from standard file and database sources
  • Writing and partitioning data within the Lakehouse

Delta Lake and Table Administration

  • Delta tables and ACID transaction support
  • Time travel capabilities and schema evolution

Data Cleansing and Transformation Strategies

  • Data cleaning techniques and type casting
  • Creating reusable transformation logic

User-Defined Functions and Modular Programming

  • Python UDFs and pandas UDFs
  • Encapsulating procedural logic into modular functions

Performance Tuning and Optimization Techniques

  • Strategies for partitioning and caching
  • Identifying bottlenecks using the Spark UI

Basics of Structured Streaming

  • Distinguishing between batch and streaming processing models
  • Working with streaming DataFrames and simple aggregations

Databricks Jobs and Workflow Orchestration

  • Scheduling notebooks as automated jobs and tasks
  • Constructing multi-step workflows with defined dependencies

Unity Catalog and Data Governance

  • Unity Catalog architecture and namespace management
  • Access control mechanisms and data lineage tracking

Testing, Debugging, and Production Standards

  • Performing unit tests on PySpark logic
  • Debugging processes and adhering to code quality standards

Comprehensive Financial Services Applications

  • Developing a complete banking ETL pipeline
  • Converting legacy SQL processes to PySpark

Migrating SQL Workloads to PySpark

  • Migration strategies and planning approaches
  • Step-by-step conversion of SQL workflows to PySpark

Requirements

  • Proficiency in Python programming, covering functions and data types
  • Knowledge of SQL concepts, such as joins, aggregations, and subqueries
  • No prior experience with Databricks or PySpark is necessary

Target Audience

  • Data engineers, data analysts, and data professionals
  • Teams transitioning existing SQL-based workflows to Databricks and PySpark

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories