Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open Source Tools
- Overview of AIOps principles and their advantages
- The role of Prometheus and Grafana within the observability ecosystem
- Positioning Machine Learning in AIOps: predictive versus reactive analysis
Configuring Prometheus and Grafana
- Deployment and setup of Prometheus for time series data acquisition
- Building dashboards in Grafana leveraging live metrics
- Investigation of exporters, relabeling rules, and service discovery mechanisms
Data Preparation for Machine Learning
- Retrieval and conversion of Prometheus metrics
- Structuring datasets for anomaly detection and prediction tasks
- Utilizing Grafana transformation features or Python-based data pipelines
Implementing Machine Learning for Anomaly Detection
- Foundational ML algorithms for outlier identification (e.g., Isolation Forest, One-Class SVM)
- Model training and validation using time series datasets
- Displaying detected anomalies through Grafana visualizations
Metric Forecasting via Machine Learning
- Construction of basic prediction models (ARIMA, Prophet, introduction to LSTM)
- Anticipating system load and resource consumption
- Leveraging forecasts for proactive alerting and scaling strategies
Connecting Machine Learning with Alerting and Automation
- Establishing alert criteria driven by ML outputs or static thresholds
- Configuration of Alertmanager and notification workflows
- Initiating scripts or automated workflows upon anomaly identification
Scaling and Implementing AIOps Operations
- Integration with external observability platforms (e.g., ELK stack, Moogsoft, Dynatrace)
- Deployment of ML models within observability data pipelines
- Best practices for managing AIOps at scale
Conclusion and Future Actions
Requirements
- A solid grasp of system monitoring and observability principles
- Practical experience with Grafana or Prometheus
- Knowledge of Python and fundamental machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and Site Reliability Engineers (SREs)