Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps Using Open Source Tools

  • Overview of AIOps principles and their advantages
  • The role of Prometheus and Grafana within the observability ecosystem
  • Positioning Machine Learning in AIOps: predictive versus reactive analysis

Configuring Prometheus and Grafana

  • Deployment and setup of Prometheus for time series data acquisition
  • Building dashboards in Grafana leveraging live metrics
  • Investigation of exporters, relabeling rules, and service discovery mechanisms

Data Preparation for Machine Learning

  • Retrieval and conversion of Prometheus metrics
  • Structuring datasets for anomaly detection and prediction tasks
  • Utilizing Grafana transformation features or Python-based data pipelines

Implementing Machine Learning for Anomaly Detection

  • Foundational ML algorithms for outlier identification (e.g., Isolation Forest, One-Class SVM)
  • Model training and validation using time series datasets
  • Displaying detected anomalies through Grafana visualizations

Metric Forecasting via Machine Learning

  • Construction of basic prediction models (ARIMA, Prophet, introduction to LSTM)
  • Anticipating system load and resource consumption
  • Leveraging forecasts for proactive alerting and scaling strategies

Connecting Machine Learning with Alerting and Automation

  • Establishing alert criteria driven by ML outputs or static thresholds
  • Configuration of Alertmanager and notification workflows
  • Initiating scripts or automated workflows upon anomaly identification

Scaling and Implementing AIOps Operations

  • Integration with external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace)
  • Deployment of ML models within observability data pipelines
  • Best practices for managing AIOps at scale

Conclusion and Future Actions

Requirements

  • A solid grasp of system monitoring and observability principles
  • Practical experience with Grafana or Prometheus
  • Knowledge of Python and fundamental machine learning concepts

Target Audience

  • Observability engineers
  • Infrastructure and DevOps teams
  • Monitoring platform architects and Site Reliability Engineers (SREs)

Number of participants


Price per participant

Upcoming Courses

Related Categories