Get in Touch

Course Outline

Introduction

This section offers an overview of the contexts in which 'machine learning' is applicable, the key considerations involved, and the broader implications, including its advantages and limitations. It covers datatypes (structured/unstructured/static/streamed), data validity and volume, the distinction between data-driven and user-driven analytics, statistical models versus machine learning models, the challenges of unsupervised learning, the bias-variance trade-off, iteration and evaluation, cross-validation strategies, and the paradigms of supervised, unsupervised, and reinforcement learning.

MAJOR TOPICS

1. Understanding naive Bayes

  • Foundational concepts of Bayesian methods
  • Probability theory
  • Joint probability
  • Conditional probability and Bayes' theorem
  • The naive Bayes algorithm
  • Classification using naive Bayes
  • The Laplace estimator
  • Incorporating numeric features into naive Bayes

2. Understanding decision trees

  • The divide and conquer strategy
  • The C5.0 decision tree algorithm
  • Selecting optimal splits
  • Pruning decision trees

3. Understanding neural networks

  • Evolution from biological to artificial neurons
  • Activation functions
  • Network topology
  • Determining the number of layers
  • Information flow direction
  • Node count per layer
  • Training neural networks via backpropagation
  • Deep Learning

4. Understanding Support Vector Machines

  • Classification via hyperplanes
  • Maximizing the margin
  • Scenarios with linearly separable data
  • Scenarios with non-linearly separable data
  • Utilizing kernels for non-linear spaces

5. Understanding clustering

  • Clustering as a machine learning objective
  • The k-means clustering algorithm
  • Assigning and updating clusters using distance metrics
  • Selecting the optimal number of clusters

6. Measuring performance for classification

  • Handling classification prediction data
  • In-depth analysis of confusion matrices
  • Assessing performance through confusion matrices
  • Metrics beyond accuracy
  • The kappa statistic
  • Sensitivity and specificity
  • Precision and recall
  • The F-measure
  • Visualizing performance trade-offs
  • ROC curves
  • Predicting future performance
  • The holdout method
  • Cross-validation
  • Bootstrap sampling

7. Tuning stock models for better performance

  • Automated parameter tuning with caret
  • Constructing simple tuned models
  • Customizing the tuning workflow
  • Enhancing model performance through meta-learning
  • Concepts of ensembles
  • Bagging
  • Boosting
  • Random forests
  • Training random forests
  • Assessing random forest performance

MINOR TOPICS

8. Understanding classification using the nearest neighbors

  • The kNN algorithm
  • Distance calculation
  • Selecting an appropriate k value
  • Data preparation for kNN
  • The lazy nature of the kNN algorithm

9. Understanding classification rules

  • The separate and conquer approach
  • The One Rule algorithm
  • The RIPPER algorithm
  • Deriving rules from decision trees

10. Understanding regression

  • Simple linear regression
  • Ordinary least squares estimation
  • Correlations
  • Multiple linear regression

11. Understanding regression trees and model trees

  • Integrating regression into tree structures

12. Understanding association rules

  • The Apriori algorithm for association rule learning
  • Measuring rule relevance via support and confidence
  • Generating rule sets using the Apriori principle

Extras

  • Spark/PySpark/MLlib and Multi-armed bandits

Requirements

Proficiency in Python

 21 Hours

Number of participants


Price per participant

Testimonials (7)

Upcoming Courses

Related Categories