Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundational Concepts (Theory):

  • Architecture
  • RDDs
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring Fundamentals in the Databricks Environment (Hands-on Workshop):

  • Practical exercises utilizing the RDD API
  • Core action and transformation functions
  • PairRDD
  • Joins
  • Caching strategies
  • Practical exercises utilizing the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDF (User Defined Function)
  • Overview of the DataSet API
  • Streaming

Deployment Strategies in the AWS Environment (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Demonstrating sample jobs in both environments
  • Evaluating the advantages and limitations of each

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming skills (Python and Scala are preferred)

Fundamental SQL knowledge

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories