Primary Skill:
- Azure Cloud Operations Site Reliability Engineering SRE Observability Automation
- Experience in Years 7 to 12 Years
Must Have Skills:
- Strong experience in Azure Cloud Operations Platform Reliability and Production Support
- Handson expertise with Azure Monitor Log Analytics Application Insights Datadog and Observability Platforms
- Experience with CICD Pipeline Management using Azure DevOps GitHub Actions and deployment automation
- Strong scripting and automation skills using PowerShell Azure CLI Azure Automation and Infrastructure as Code TerraformBicepARM
Secondary Nice To Have Skills
- Experience with Kubernetes AKS container monitoring and application performance management
- Exposure to DevSecOps practices security baselines and platform hardening
- Experience in Azure Cost Optimization Capacity Planning and Performance Tuning
- Knowledge of AIOpsAIdriven Operations anomaly detection correlation and predictive incident management
Soft Skills
- Strong analytical and troubleshooting capabilities
- Excellent stakeholder communication and incident management skills
- Ability to work under pressure during critical production incidents
- Strong documentation collaboration and continuous improvement mindset
3 Qualifying Questions
- Describe your experience implementing and managing observability solutions such as Datadog Azure Monitor Log Analytics and Application Insights What outcomes did you achieve?
- How have you leveraged automation PowerShell Azure CLI Azure Automation Terraform etc to reduce operational effort and improve reliability?
- Can you share an example of a major production incident you handled including the RCA process corrective actions and reliability improvements implemented afterward?
Skills
Mandatory Skills : Azure DevOps, Azure Infra Services, Azure Log Analytics, Azure Monitor
Good to Have Skills : Azure AI, Azure App Service