Trainings/ AIOps Engineering
Course – AIOps & Agentic Operations Engineering

AIOps Engineering — Building Intelligent Operations at Scale

24 Hrs (5 Days / 5 Hrs or 6 Days / 4 Hrs) Advanced Classroom / Live Virtual 75% Hands-On
Duration
24 Hrs / 5-6 Days
Level
Advanced
Format
Classroom / Live Virtual
Stack
Azure Foundry · Agent Framework · Azure Monitor
Training Methodology

Learning by Doing

Every session in this program is built around hands-on execution, not passive slides — you leave having built and deployed working AIOps agents, not just watched a demo.

01

Explore

Hands-on labs from Day 1 — real Azure Monitor and Foundry environments, not slide decks.

02

Experiment

Work against simulated IT environments that mirror production observability and incident scenarios.

03

Engage

Build detection, RCA, self-healing and FinOps agents across every module, not toy examples.

04

Apply

Capstone builds a working end-to-end AIOps prototype tied to your team’s actual stack.

Who This Is For

Practitioners with Prior Agentic AI Development Experience SRE & IT Operations Engineers Azure Cloud & Platform Engineers MLOps Practitioners

What You’ll Be Able to Do

  • Architect full-stack observability pipelines on Azure Monitor and design SRE-aligned SLIs, SLOs and error budgets
  • Build ML-driven signal processing pipelines that correlate, group and de-noise high-volume alert streams
  • Deploy topology-aware RCA agents on Azure Foundry that reason across service dependency graphs
  • Build closed-loop self-healing agents with Microsoft Agent Framework, including human-in-the-loop escalation and rollback guardrails
  • Secure agentic AIOps systems against prompt injection with least-privilege access, audit logging and red-teaming
  • Integrate FinOps cost observability and build agents that trigger automated scale-down on cost anomalies
  • Build RAG pipelines on Azure AI Search that ground RCA agents in operational runbook knowledge
  • Apply MLOps practices — model registry, drift detection and automated retraining — to operationalize AIOps models

Prerequisites

  • Completion of foundational Agentic AI training on the Microsoft technology stack, with working knowledge of agent design, development and deployment
  • Familiarity with IT Operations processes — incident, change and service management — and basic Azure cloud computing concepts
  • Hands-on experience with Azure services, Azure Foundry and Microsoft Agent Framework, plus basic Python proficiency and understanding of REST APIs and event-driven architectures

Curriculum

Module 1
AIOps Architecture & Observability Foundation
  • AIOps vs traditional IT Ops — architectural shift
  • Full-stack observability — logs, metrics and traces
  • Azure Monitor as the observability backbone
  • Telemetry data schema, normalization and instrumentation strategies
  • Observability maturity model
Lab: Configure Azure Monitor to collect logs, metrics and traces from a simulated IT environment.
Module 2
SRE-Driven Reliability Design
  • SRE principles and their role in AIOps
  • Defining SLIs, SLOs and SLAs
  • Error budget policies and burn rate alerts
  • Reliability as a design constraint in AIOps
  • Integrating SRE metrics into operational dashboards
Module 3
Intelligent Signal Processing
  • Alert fatigue and the noise problem in IT Ops
  • Event correlation techniques — rule-based and ML-driven
  • ML-based noise reduction approaches
  • Alert grouping and deduplication strategies
  • Threshold-based vs. anomaly-based alerting
  • Building signal processing pipelines on Azure
Lab: Build an event correlation pipeline that filters and groups high-volume Azure Monitor alerts into prioritized actionable signals.
Module 4
Anomaly Detection & Topology-Aware RCA
  • ML model types for anomaly detection in IT Ops
  • Operationalizing pre-built anomaly detection models on Azure
  • Topology mapping and service dependency graphs
  • Root cause analysis frameworks
  • Agentic AI for automated RCA reasoning
  • Integrating RCA agents with Azure Foundry
Lab: Deploy a pre-built anomaly detection model and configure an Azure Foundry RCA agent to reason across a service topology map.
Module 5
Predictive Analytics & Real-Time Telemetry Pipelines
  • Forecasting models for operational failure prediction
  • Real-time vs. batch telemetry processing
  • Azure Event Hubs and Stream Analytics for telemetry ingestion
  • Building scalable telemetry pipelines on Azure Foundry
  • Operationalizing forecasting models in production
  • Predictive alerting and capacity planning
Lab: Deploy a real-time telemetry ingestion pipeline with predictive alerting for capacity threshold breaches.
Module 6
Closed-Loop Self-Healing Systems
  • Closed-loop automation architecture
  • Detection → diagnosis → remediation design pattern
  • Building self-healing agents with Microsoft Agent Framework
  • Agent orchestration in multi-step remediation workflows
  • Human-in-the-loop escalation patterns
  • Safety guardrails and rollback mechanisms
Lab: Build a self-healing agent on Azure Foundry that detects a simulated incident, diagnoses root cause, and executes automated remediation.
Module 7
AI Agent Security & Prompt Injection Defense
  • Prompt injection attacks on LLM-based agents
  • Agent identity & access management — least-privilege design
  • Audit logging of autonomous agent actions
  • Secrets management for agents calling external APIs
  • Threat modeling for agentic AIOps systems
  • Security testing and red-teaming AI agents
Lab: Configure Azure Key Vault secrets management and audit logging for a self-healing agent, then simulate a prompt injection attack and validate defensive controls.
Module 8
Cost Observability & FinOps Integration
  • Cost as a first-class operational metric
  • Azure Cost Management API integration with Azure Monitor
  • Cost anomaly detection — budget alerts and spend spikes
  • Agent-triggered scale-down actions tied to cost policies
  • FinOps dashboards alongside reliability dashboards
  • FinOps governance and showback/chargeback models
Lab: Integrate Azure Cost Management API with Azure Monitor and build an agent that detects a cost spike and triggers an automated scale-down action.
Module 9
RAG for Runbook Knowledge
  • Vector indexing of operational runbooks in Azure AI Search
  • Grounding RCA agents with retrieved runbook context
  • Post-mortem ingestion pipelines for agent learning
  • Hybrid search — keyword and semantic for incident context
  • RAG pipeline architecture on Azure Foundry
  • Evaluating RAG quality — relevance, groundedness and faithfulness
Lab: Build a RAG pipeline that indexes operational runbooks in Azure AI Search and grounds an RCA agent with retrieved context during a simulated incident diagnosis.
Module 10
MLOps for AIOps Operationalization
  • MLOps principles in the AIOps context
  • Model registry, versioning and governance on Azure
  • Model drift detection and monitoring
  • CI/CD pipelines for model deployment
  • Agentic AI for automated model monitoring and response
  • Continuous model evaluation and retraining triggers
Lab: Configure a model monitoring pipeline where an Azure Foundry agent detects drift and triggers an automated retraining workflow.
Module 11
Capstone — Working AIOps Prototype
  • Capstone problem statement and scope definition
  • End-to-end system architecture design
  • Integration of observability, agents, RAG and MLOps components
  • Prototype build, testing and validation
  • Demonstration and peer review

Delivery Details

  • Delivered as classroom or live virtual instructor-led — scheduled around your team, 5 hrs/day over 5 days or 4 hrs/day over 6 days
  • 75% hands-on labs across Azure Foundry, Microsoft Agent Framework and Azure Monitor
  • Labs are illustrative and may vary by trainer approach, participant profile and Azure subscription access

Request This Program

Email