Senior DevOps
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
We are looking for an experiencedSenior DevOps & Site Reliability Engineer (SRE)to design, build and operate highly available, secure, scalable and automated enterprise technology platforms.
This is a senior hands-on engineering role spanningDevOps, Site Reliability Engineering, Azure Cloud, Platform Engineering, Kubernetes, Infrastructure as Code, CI/CD, Observability and DevSecOps.
The successful candidate will work across engineering and delivery teams to improveplatform reliability, deployment velocity, resilience, automation, operational efficiency and production performance, while supporting mission-critical enterprise applications.
Key Responsibilities
DevOps & Platform Engineering
Design, build and maintaincloud-native infrastructure and platform services.
Develop and maintainInfrastructure as Code (IaC)solutions.
Automate infrastructure provisioning, configuration and operational processes.
Build reusable engineering tools, deployment templates and platform components.
Establish and standardise platform engineering practices across multiple delivery teams.
Identify opportunities to reduce manual intervention and increase engineering automation.
CI/CD & Release Automation
Design, implement and maintain enterprise-gradeCI/CD pipelinesfor application and infrastructure deployments.
Implement automated testing, security scanning, code-quality controls and release automation.
Enable automated deployments, rollback and recovery processes.
Improve deployment frequency while reducing change and deployment risk.
Continuously optimise software delivery and release-management processes.
Site Reliability Engineering
Implement and matureSite Reliability Engineering practicesacross production environments.
Define, monitor and manageService Level Indicators (SLIs), Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
Improve application and platformavailability, scalability, resilience and performance.
Lead production incident response, troubleshooting, problem management andRoot Cause Analysis (RCA).
Drive proactive reliability improvements and reduction of technical debt.
ImproveMean Time to Detect (MTTD)andMean Time to Recover (MTTR).
Azure Cloud Engineering
Design, implement and operate enterpriseMicrosoft Azureenvironments.
Work extensively with technologies such as:
Azure Kubernetes Service (AKS)
Azure App Services
Azure Networking
Azure Monitor
Azure Storage
Azure Identity Services
Design highly available and disaster-recovery-capable environments.
Optimise cloud environments forperformance, resilience, security and cost.
Support hybrid-cloud and multi-cloud environments where required.
Containers & Kubernetes
Build, deploy and support containerised applications usingDockerandKubernetes.
Manage Kubernetes environments, particularlyAzure Kubernetes Service (AKS).
Develop and maintain deployment configurations usingHelm.
Support container-platform reliability, scalability and operational performance.
OpenShift experience would be advantageous.
Infrastructure as Code & Automation
Hands-on experience with technologies such as:
Terraform
Bicep
ARM Templates
Ansible
Candidates should be comfortable using Infrastructure as Code to build repeatable, scalable and governed enterprise infrastructure.
Monitoring & Observability
Implement comprehensivelogging, monitoring, metrics, tracing and alerting.
Build operational dashboards and platform insights.
Establish enterprise observability standards.
Implement proactive and predictive monitoring.
Use observability information to improve application and infrastructure reliability.
Relevant technologies may include:
Dynatrace
Grafana
Prometheus
Elastic Stack / ELK
Splunk
Azure Monitor
OpenTelemetry
DevSecOps & Security
EmbedDevSecOps