AI-Driven Observability: From Logs to LLM-Powered Insights Training Course
Conventional observability strategies often depend on static dashboards, threshold-based alerts, and the manual scrutiny of logs. AI-driven observability revolutionizes this approach by enabling natural language queries against telemetry data, leveraging LLMs for root cause analysis, utilizing foundation models for anomaly detection, and producing automated incident summaries that retain contextual understanding.
This live, instructor-led training session (available online or onsite) is specifically designed for observability and SRE professionals seeking to integrate LLMs and AI into their monitoring, alerting, and incident analysis processes.
Upon completing this training, participants will be equipped to:
- Create natural language interfaces for querying data within Prometheus, Elasticsearch, and SQL-based observability repositories.
- Develop pipelines for LLM-enhanced log analysis and anomaly detection.
- Produce automated incident summaries and postmortem drafts directly from raw telemetry data.
- Architect AI-assisted root cause analysis workflows that incorporate evidence chaining.
- Integrate foundation models to perform time-series anomaly detection and forecasting.
- Implement an AI-enhanced on-call experience featuring smart alert enrichment.
Training Structure
- Interactive lectures and facilitated discussions.
- Extensive practical exercises and hands-on practice.
- Real-time implementation within a live-lab environment.
Customization Options
- For tailored training requirements, please contact us to discuss arrangements.
Course Outline
The Landscape of AI in Observability
- From static dashboards to dynamic conversations: the evolution toward AI-augmented observability
- Relevant LLM capabilities: summarization, reasoning, and pattern matching
- Architectural patterns for embedding AI into existing observability stacks
Telemetry Querying via Natural Language
- Text-to-PromQL: converting natural language into monitoring queries
- Natural language querying for Elasticsearch, OpenSearch, and Loki log stores
- Generating SQL from natural language for structured telemetry data
- Developing query assistant agents with tool usage and context awareness
Log Analysis Powered by LLMs
- Automating log parsing and structuring using LLMs
- Detecting anomalies in log streams through embedding similarity
- Clustering logs and discovering patterns at scale
- Creating human-readable explanations from raw log sequences
Intelligent Alerting and Incident Enrichment
- Correlating and deduplicating alerts using semantic understanding
- Gathering automated incident context from runbooks, past incidents, and documentation
- Routing alerts intelligently based on content understanding and team expertise
- Mitigating alert fatigue through AI-driven noise reduction
AI-Assisted Root Cause Analysis
- Generating hypotheses through multi-source telemetry correlation
- Evidence chaining: linking symptoms across metrics, logs, and traces
- Guided troubleshooting via interactive AI diagnosis sessions
- Building root cause analysis agents capable of progressive investigation
Automated Incident Response and Communication
- Creating incident summaries and status updates derived from telemetry
- Automating postmortem drafting with timeline reconstruction
- Tailoring stakeholder communication for both technical and executive audiences
- Suggesting runbooks and providing automated remediation recommendations
Machine Learning for Observability
- Time-series forecasting for capacity planning and anomaly prediction
- Using foundation models for zero-shot anomaly detection on metrics
- Mapping service dependencies and discovering topology via embeddings
- Training and deploying lightweight ML models alongside observability pipelines
Production Deployment and Ethical Considerations
- Addressing latency and cost factors in real-time AI observability
- Data privacy: preventing LLMs from leaking sensitive telemetry data
- Human oversight: identifying when AI diagnosis requires operator validation
- Measuring impact: tracking MTTD, MTTR, and on-call experience metrics
Requirements
- Practical experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Understanding of log management and metrics principles.
- Foundational Python scripting skills for data processing.
Target Audience
- SRE and observability engineers incorporating AI-enhanced tooling.
- Platform engineers constructing next-generation monitoring pipelines.
- DevOps leads assessing the integration of LLMs into incident workflows.
Open Training Courses require 5+ participants.
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Booking
AI-Driven Observability: From Logs to LLM-Powered Insights Training Course - Enquiry
AI-Driven Observability: From Logs to LLM-Powered Insights - Consultancy Enquiry
Upcoming Courses
Related Courses
Agentic Development with Gemini 3 and Google Antigravity
21 HoursGoogle Antigravity serves as an agentic development environment, engineered to foster autonomous agents that leverage Gemini 3's multimodal capabilities for planning, reasoning, coding, and execution.
Offered as an instructor-led, live training session (available online or on-site), this course is tailored for advanced technical professionals seeking to design, build, and deploy autonomous agents within the Antigravity ecosystem using Gemini 3.
By the conclusion of this programme, participants will be equipped to:
- Create autonomous workflows that harness Gemini 3 for intricate reasoning, planning, and task execution.
- Develop agents within Antigravity capable of analysing tasks, generating code, and interacting with various tools.
- Integrate Gemini-powered agents seamlessly with enterprise systems and external APIs.
- Enhance agent behaviour, ensuring safety and reliability within complex operational environments.
Course Format
- Combination of expert-led demonstrations and interactive discussions.
- Hands-on experimentation focused on autonomous agent development.
- Practical implementation leveraging Antigravity, Gemini 3, and complementary cloud tools.
Customisation Options
- Should your team require domain-specific agent behaviours or bespoke integrations, please reach out to us to tailor the programme to your needs.
Advanced Antigravity: Feedback Loops, Learning & Long-Term Agent Memory
14 HoursGoogle Antigravity serves as a sophisticated framework designed for experimenting with long-lived agents and their emergent interactive capabilities.
This live, instructor-led training—available online or on-site—is tailored for advanced professionals looking to design, analyze, and optimize agents that can retain memory, refine performance through feedback, and evolve over extended operational periods.
By the end of this course, participants will be equipped to:
- Architect long-term memory structures to ensure agent persistence.
- Establish robust feedback loops that effectively guide agent behavior.
- Assess learning trajectories and address model drift.
- Incorporate memory mechanisms within complex multi-agent ecosystems.
Course Format
- Expert-led discussions complemented by technical demonstrations.
- Practical exploration through structured design challenges.
- Application of concepts within simulated agent environments.
Customization Options
- Should your organization require specific content or case-based examples, please reach out to tailor this training to your needs.
Advanced Mastra Integrations: APIs, Tools, Enterprise Data & External Systems
21 HoursMastra serves as a robust framework that enables seamless integration between AI agents, APIs, enterprise applications, and external data ecosystems.
This instructor-led, live training session, available online or onsite, is tailored for intermediate-level engineers looking to construct reliable, secure, and scalable connections between Mastra agents and the wider enterprise environment.
Upon completion of this program, participants will be equipped to:
- Establish API-driven connections between Mastra agents and third-party services.
- Link enterprise data repositories and tools with automated agent workflows.
- Adopt best practices for secure data exchange and authentication protocols.
- Architect integration layers that are scalable, maintainable, and ready for production deployment.
Course Format
- Interactive lectures and facilitated discussions.
- Practical exercises focused on integration engineering and API management.
- Live laboratory sessions utilizing real-world enterprise case studies.
Customization Options
- Tailored API scenarios, enterprise system mapping sessions, or data-integration workshops can be arranged upon request.
Interactive AI Agents: AgentCore Memory, Code Interpreter & Browser Tool in Action
14 HoursAgentCore equips AI agents with memory persistence, a secure code interpreter, and a browser tool, empowering them to deliver interactive, dynamic, and context-aware experiences.
This live, instructor-led training—available online or onsite—is tailored for intermediate to advanced technical professionals seeking to design and deploy AI agents that support long-term context retention, on-the-fly computation, and direct engagement with web UIs.
Upon completion, participants will be equipped to:
- Implement AgentCore memory to drive stateful, context-aware workflows.
- Utilize the secure code interpreter for dynamic calculations and data transformations.
- Integrate the browser tool to facilitate real-time data retrieval and UI interaction.
- Develop interactive agents suitable for analytics, customer support, and research applications.
Course Format
- Interactive lectures and group discussions.
- Practical lab exercises focused on AgentCore memory and tools.
- Case studies covering analytics, automation, and customer support scenarios.
Customization Options
- To arrange a tailored training session for this course, please reach out to our team.
Accelerating AI Agent Deployment with AgentCore Runtime & Gateway
14 HoursThe AgentCore Runtime & Gateway is an AWS service suite designed to package, deploy, and securely expose AI agents while facilitating seamless integrations with external systems.
This instructor-led, live training—available online or on-site—is tailored for intermediate-level engineering teams aiming to transition from agent prototypes to production environments. By mastering the AgentCore Runtime for deployment and the Gateway for secure connectivity and API integration, teams can achieve reliable, scalable outcomes.
Upon completion of this training, participants will be equipped to:
- Establish AgentCore Runtime environments and prepare agents for deployment.
- Publish agents via Gateway using authenticated, rate-limited endpoints.
- Incorporate external tools and APIs into agent workflows through stable contracts.
- Implement observability, logging, and usage monitoring for robust production operations.
Course Format
- Interactive lectures and group discussions.
- Hands-on labs focusing on Runtime deployments and Gateway integrations.
- Practical exercises emphasizing reliability, security, and deployment strategies.
Customization Options
- Please reach out to us to arrange a customized training program tailored to your specific needs.
Antigravity for Developers: Building Agent-First Applications
21 HoursAntigravity is a purpose-built development platform that facilitates the creation of AI-driven, agent-first applications.
This instructor-led live training, available either online or on-site, targets intermediate-level developers eager to construct real-world solutions utilizing autonomous AI agents within the Antigravity ecosystem.
Upon completion, participants will possess the skills necessary to:
- Architect applications powered by autonomous and collaborative AI agents.
- Leverage the Antigravity IDE, editor, terminal, and browser for seamless end-to-end development.
- Orchestrate multi-agent workflows effectively using the Agent Manager.
- Embed agent capabilities into robust, production-grade software architectures.
Training Delivery Style
- Combines detailed presentations with practical demonstrations.
- Features extensive hands-on sessions and guided practical exercises.
- Includes direct implementation tasks within the live Antigravity environment.
Customization Possibilities
- We can tailor the curriculum to align with your specific technology stack; please reach out to discuss a bespoke training arrangement.
Getting Started with Antigravity: An Introduction to Agent-First IDEs
14 HoursGoogle Antigravity represents a shift towards agent-first development environments, specifically engineered to optimize engineering processes via intelligent automation.
Targeted at beginners, this live, instructor-led training (available both online and onsite) provides a foundational understanding of Antigravity, demonstrating how agent-driven coding environments can significantly boost productivity.
By the end of this session, participants will have the capability to:
- Set up and configure Google Antigravity.
- Navigate the interface, distinguishing between the Editor View and Manager View.
- Leverage agents to streamline and automate routine development tasks.
- Utilize Antigravity to create, modify, and oversee project files.
Course Delivery Format
- Instructor-led insights paired with live, real-time demonstrations.
- Structured exercises emphasizing practical, hands-on agent interaction.
- Practical application of core Antigravity features within a controlled lab setting.
Customization Availability
- For those seeking a tailored training experience, please reach out to discuss customizing this program to your specific needs.
Antigravity for Web Automation & Browser-Based Tasks
21 HoursGoogle Antigravity serves as a comprehensive platform for developing agents that engage with web applications, browser environments, and complex multi-surface workflows.
This live, instructor-led training session, available either online or on-site, is designed for intermediate-level professionals looking to construct, automate, and validate browser-based workflows using Google Antigravity.
By the end of this program, participants will be equipped to:
- Develop agents that interact with web applications within a browser surface.
- Streamline end-to-end workflows across various browser contexts.
- Verify and resolve agent behavior issues in UI-driven environments.
- Deploy cross-surface automation strategies leveraging Antigravity.
Training Structure
- Facilitated learning reinforced by practical demonstrations.
- Hands-on practical tasks and scenario-driven exercises.
- Building agent workflows within an interactive laboratory setting.
Customization Possibilities
- Reach out to us for bespoke training solutions tailored to your specific business objectives.
Building Fully Managed AI Agents with AgentCore: From Concept to Production
14 HoursAgentCore streamlines the development, enhancement, and monitoring of fully managed AI agents by offering a cohesive suite of services designed for large-scale deployment.
Delivered as an instructor-led live session, either online or onsite, this training is tailored for practitioners ranging from beginner to intermediate levels who are eager to acquire practical experience in creating production-ready AI agents using AgentCore.
Upon completion of this program, participants will be equipped to:
- Grasp the fundamental capabilities of AgentCore for AI agent development.
- Architect and configure straightforward AI agents leveraging managed services.
- Integrate workflows to expand agent functionality.
- Roll out and monitor AI agents within production environments.
Course Format
- Interactive lectures and facilitated discussions.
- Practical labs utilizing AgentCore services.
- Structured exercises guiding you from agent concept to deployment.
Customization Options
- To arrange a tailored training experience for this course, please get in touch with us.
AI Agent Development with Mastra
14 HoursThis live, instructor-led training, available either online or onsite, targets intermediate software developers and engineering teams seeking to construct scalable, observable AI systems with Mastra.
Upon completing this course, participants will be equipped to:
- Grasp Mastra’s underlying architecture and its integration capabilities with LLMs and external APIs.
- Architect and build AI agents and workflows using TypeScript.
- Leverage Mastra’s observability and memory utilities to monitor and enhance agent performance.
- Launch production-ready AI applications by utilizing Mastra’s framework capabilities.
Mastra Debugging, Evaluation & Quality Assurance for AI Agents
21 HoursMastra serves as a comprehensive framework offering structured instruments to evaluate, debug, and guarantee the dependability of AI agents functioning within intricate workflows.
Delivered as an instructor-led, live training session (either online or on-site), this course is tailored for intermediate-level practitioners seeking to rigorously verify agent conduct, enhance system reliability, and establish quantifiable evaluation processes.
By the conclusion of this training, participants will be equipped to:
- Utilize debugging methods to pinpoint and resolve issues in agent behavior.
- Assess agents through structured metrics, benchmarks, and quality scoring.
- Deploy tools and workflows that monitor reliability, drift, and hallucinations.
- Formulate QA strategies that guarantee consistent and predictable agent performance.
Course Format
- Engaging lectures and facilitated discussions.
- Practical exercises focused on debugging and evaluation.
- Real-time lab analysis of agent behaviors using observability tools.
Customization Options
- Tailored reliability testing scenarios and industry-specific QA methodologies can be arranged upon request.
Mastra Ops & Production Engineering: Deploying and Scaling AI Agents
21 HoursMastra serves as an operational framework built to simplify the deployment, scaling, and lifecycle management of AI agents within production settings.
Delivered as an instructor-led, live session (available online or onsite), this course targets intermediate to advanced technical professionals who require the ability to operationalize AI agents with reliability and efficiency across production systems.
By the end of this training, participants will have the skills to:
- Deploy Mastra-based AI agents into stable, production-grade environments.
- Scale agents both horizontally and vertically by leveraging platform-native primitives.
- Establish observability pipelines to monitor agent behaviour and performance metrics.
- Tune runtime configurations to minimize latency, control costs, and mitigate operational risks.
Course Format
- Engaging lectures combined with interactive discussions.
- Practical exercises centered on realistic deployment scenarios.
- Live-lab implementation utilizing containerized and orchestrated environments.
Customization Options
- Topics, hands-on labs, and industry-specific scenarios can be tailored upon request.
Mastra Workflow Automation & Multi-Agent Orchestration
21 HoursMastra serves as a robust framework for advanced workflow automation and the coordination of multiple AI agents within distributed systems.
Designed for intermediate-level practitioners, this instructor-led training (available online or onsite) focuses on the design, orchestration, and operation of multi-agent workflows at scale.
Upon completion, participants will be equipped to:
- Construct complex workflows leveraging Mastra’s orchestration features.
- Manage multiple agents executing tasks that are either parallel or dependent.
- Deploy monitoring and debugging tools to oversee workflow execution.
- Fine-tune orchestration logic to enhance reliability, throughput, and automation efficiency.
Course Delivery Format
- Engaging lectures and group discussions.
- Practical exercises in workflow design and automation.
- Real-world implementation within a containerized live-lab environment.
Customization Options
- Tailored automation scenarios, enterprise integrations, or specific workflow patterns are available upon request.
Managing Agent Workflows in Google Antigravity: Orchestration, Planning and Artifacts
14 HoursGoogle Antigravity serves as an agent-centric development platform designed to oversee and coordinate AI-driven coding and automation workflows.
This live, instructor-led training—available both online and onsite—is tailored for intermediate professionals seeking to design, manage, and optimize multi-agent workflows within Google Antigravity.
By the end of this training, participants will be equipped with the capability to:
- Define agent responsibilities and orchestration pipelines using the Manager interface.
- Create and analyze Antigravity artifacts, such as task lists, plans, logs, and browser recordings.
- Apply verification strategies to maintain transparency and auditability of agent actions.
- Enhance multi-agent collaboration for intricate development and operational tasks.
Course Format
- Guided presentations and practical demonstrations.
- Scenario-based exercises addressing real-world workflow challenges.
- Hands-on experimentation in a live Antigravity workspace.
Course Customization Options
- For a tailored version of this course, please reach out to discuss customization possibilities.
Testing & Verifying Agent-Driven Code: Quality Assurance in Antigravity
14 HoursAntigravity is a framework designed to facilitate advanced agent-driven development workflows.
This live, instructor-led training, available online or on-site, targets intermediate to advanced professionals seeking to verify, validate, and secure the outputs generated by AI agents within Antigravity environments.
Upon completion, participants will be equipped to:
- Evaluate the accuracy and safety of code artifacts generated by agents.
- Utilize structured techniques to verify tasks executed by agents.
- Analyze browser recordings and trace agent activities with precision.
- Implement QA and security principles to guarantee the reliability of agent workflows.
Course Structure
- Instructor-led technical briefings and discussions.
- Practical exercises centered on verifying real-world agent workflows.
- Hands-on testing and validation performed in a controlled lab setting.
Customization Possibilities
- Scenarios, workflows, and testing examples can be tailored to specific needs upon request.