DevOps & Cloud Infrastructure Manager
2 days ago
Cape Town, Western Cape, South Africa
Hepstar | Global Leader in Travel Ancillaries & Revenue Optimisation
Full-time
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
DEVOPS & CLOUD INFRASTRUCTURE MANAGER
Position: DevOps & Cloud Infrastructure Manager
Start date: As soon as possible
Duration: Permanent
Reporting into: Systems Architect
ABOUT THE ROLE
We are looking for an experienced, hands-on DevOps & Cloud Infrastructure Manager to take ownership of the day-to-day operation, reliability, security, cost efficiency, and continuous improvement of our cloud infrastructure and deployment processes. Our primary infrastructure is hosted on AWS, with additional AI and Big Data workloads running on Google Cloud Platform (GCP). Our development and deployment workflows are centred around GitHub, while our Linux-based workloads predominantly run on Ubuntu. This is a practical, hands-on role. In some areas, you will be recommending and implementing new practices based on your experience and industry best practice, while in others, you will take ownership of existing processes and refine and optimise them. The successful candidate will be comfortable moving between cloud infrastructure, Linux administration, CI/CD pipelines, monitoring, security, cost optimisation, and production deployments. Ideally, you will have experience in automation and be familiar with AI-augmented workflows. While the role carries ownership of our DevOps and cloud infrastructure function, we are looking for someone who enjoys being close to the technology rather than purely managing from a distance.
KEY RESPONSIBILITIES
AWS INFRASTRUCTURE & OPERATIONS Take day-to-day ownership of our AWS environment, ensuring that it remains reliable, secure, performant, and cost-effective. Responsibilities will include:
• Managing and maintaining AWS infrastructure and services.
• Proactively identifying and implementing AWS cost optimisation opportunities.
• Managing Reserved Instances / Savings Plans and ensuring commitments remain appropriate for current and forecast workloads.
• Improving infrastructure monitoring, dashboards, logging, and alerting.
• Continuously improving our AWS Security Hub posture and addressing identified security findings.
• Reviewing and improving AWS security configurations, IAM permissions, network security, and general cloud security practices.
• Monitoring infrastructure capacity, availability, and performance.
• Supporting incident investigation and root-cause analysis.
• Identifying opportunities to automate repetitive operational tasks.
• Maintaining appropriate backup, recovery, and business continuity processes.
• Keeping infrastructure components and operating systems appropriately patched and up to date.
• Simulating disaster recovery scenarios to ensure our systems are robust. GITHUB, CI/CD & DEPLOYMENTS GitHub is central to our software development and deployment processes. The role will:
• Take ownership of our existing CI/CD processes and continuously improve them.
• Maintain and enhance existing GitHub Actions workflows while creating new workflows where required.
• Work closely with developers to make software deployments reliable, repeatable, and efficient.
• Manage deployments across our various environments, including development, testing/staging, and production.
• Improve deployment controls, release processes, rollback capabilities, and visibility.
• Help troubleshoot build, deployment, and environment-related issues.
• Maintain appropriate management of secrets and credentials used within CI/CD pipelines.
• Promote automation and Infrastructure-as-Code wherever practical. LINUX / UBUNTU ADMINISTRATION Our server workloads predominantly run on Ubuntu Linux. The successful candidate must therefore be highly comfortable working from the Linux command line and will be responsible for:
• Maintaining and administering Ubuntu servers.
• OS patching, upgrades, package management, and general server maintenance.
• Monitoring server health, capacity, disk utilisation, memory, CPU, and system performance.
• Troubleshooting operating system, networking, application, and performance issues.
• Hardening servers and maintaining appropriate security configurations.
• Automating administration and operational activities. Strong Bash scripting skills are essential. The ability to quickly develop reliable scripts for automation, troubleshooting, deployments, maintenance, and operational tasks will be an important part of the role. GOOGLE CLOUD PLATFORM While AWS represents most of our infrastructure, our AI and Big Data platforms are hosted within GCP. The role will provide operational oversight and support for these environments, including:
• Maintaining GCP infrastructure and services.
• Monitoring availability, performance, security, and costs.
• Supporting AI and Big Data infrastructure requirements, including ML pipeline implementation.
• Working with engineering and data teams on infrastructure changes and deployments to ensure system performance is not compromised.
• Applying similar security, monitoring, governance, and cost-management principles across AWS and GCP. Deep GCP specialisation is not necessarily required, but candidates should have practical GCP experience and be comfortable operating in a multi-cloud environment. DOMAINS, DNS & CERTIFICATES Our domains and associated services are currently managed through GoDaddy. Responsibilities will include:
• Domain management and renewals.
• DNS configuration and maintenance.
• SSL/TLS certificate management and renewals.
• Troubleshooting DNS and certificate-related issues.
• Ensuring critical domains, DNS records, and certificates are appropriately monitored and documented. SECURITY & OPERATIONAL EXCELLENCE Security, reliability, and operational discipline will be important aspects of the position. The role will be expected to:
• Continuously improve our cloud security posture.
• Proactively identify vulnerabilities, configuration weaknesses, and operational risks.
• Ensure security findings are communicated, prioritised, and remediated.
• Improve monitoring and alerting so that issues are identified before they become significant incidents.
• Maintain appropriate operational documentation and runbooks.
• Improve disaster recovery and infrastructure resilience.
• Reduce manual processes through sensible automation.
• Establish clear standards and best practices for cloud infrastructure and deployments. REQUIRED EXPERIENCE & SKILLS We are looking for someone with strong practical experience across most of the following:
• Strong commercial experience managing AWS environments.
• Strong knowledge of AWS infrastructure, security, IAM, networking, monitoring, and cost management.
• Experience with AWS Security Hub and cloud security best practices.
• Experience managing AWS Reserved Instances and/or Savings Plans.
• Strong Linux administration, particularly Ubuntu.
• Strong Bash scripting skills.
• Strong experience with Git and GitHub.
• Experience building and maintaining automated CI/CD pipelines, preferably using GitHub Actions.
• Experience taking ownership of existing CI/CD workflows and developing new deployment workflows where required.
• Experience deploying applications across multiple environments, including production.
• Experience with monitoring, alerting, logging, and observability platforms.
• Practical experience with GCP.
• Understanding of DNS, domains, SSL/TLS certificates, and related internet infrastructure.
• Experience automating infrastructure and operational processes.
• Strong troubleshooting and root-cause analysis skills.
• Good understanding of cloud security principles.
• Ability to communicate clearly, drive improvements, make decisions, and solve problems. HIGHLY DESIRABLE Experience in any of the following would be advantageous:
• Infrastructure-as-Code using Terraform, CloudFormation, or similar.
• Docker and containerised workloads.
• Kubernetes / EKS / GKE.
• AWS Organizations and multi-account environments.
• Centralised logging and observability platforms.
• FinOps / cloud cost-management practices.
• Automated vulnerability and security management.
• Database infrastructure and administration.
• High-availability and disaster-recovery architectures.
• AI / ML or Big Data infrastructure experience.
• Python or another scripting/programming language in addition to Bash. WHAT SUCCESS LOOKS LIKE Within this role, success means that:
• Our infrastructure is stable, secure, and well maintained.
• AWS and GCP costs are understood, controlled, and continuously optimised.
• AWS Security Hub scores and security posture continuously improve.
• Monitoring and alerts identify genuine issues quickly without unnecessary noise.
• Production deployments are predictable, automated, and low risk.
• Ubuntu servers remain patched, secure, and healthy.
• Reserved capacity and cloud commitments are actively managed rather than simply renewed.
• Infrastructure and deployment processes become increasingly automated.
• Operational knowledge is documented rather than dependent on individual team members.
• Developers can deploy and operate software efficiently without infrastructure becoming a bottleneck.
• Infrastructure problems are identified proactively rather than only after customers or developers report them. THE PERSON We are looking for someone who combines technical depth with operational ownership. You should be comfortable taking responsibility for an environment, identifying what needs improvement, prioritising the work, and then implementing those improvements yourself. You will enjoy automation, dislike repetitive manual processes, pay attention to security and cost, and be comfortable troubleshooting complex problems across applications, operating systems, networks, CI/CD pipelines, and cloud infrastructure. This is an ideal role for an experienced DevOps / Cloud Infrastructure professional who wants meaningful ownership of a multi-cloud environment while remaining highly hands-on. INTERESTED? Please send your CV to Henry Hsu at henry@hepstar.com.
ABOUT THE ROLE
We are looking for an experienced, hands-on DevOps & Cloud Infrastructure Manager to take ownership of the day-to-day operation, reliability, security, cost efficiency, and continuous improvement of our cloud infrastructure and deployment processes. Our primary infrastructure is hosted on AWS, with additional AI and Big Data workloads running on Google Cloud Platform (GCP). Our development and deployment workflows are centred around GitHub, while our Linux-based workloads predominantly run on Ubuntu. This is a practical, hands-on role. In some areas, you will be recommending and implementing new practices based on your experience and industry best practice, while in others, you will take ownership of existing processes and refine and optimise them. The successful candidate will be comfortable moving between cloud infrastructure, Linux administration, CI/CD pipelines, monitoring, security, cost optimisation, and production deployments. Ideally, you will have experience in automation and be familiar with AI-augmented workflows. While the role carries ownership of our DevOps and cloud infrastructure function, we are looking for someone who enjoys being close to the technology rather than purely managing from a distance.
KEY RESPONSIBILITIES
AWS INFRASTRUCTURE & OPERATIONS Take day-to-day ownership of our AWS environment, ensuring that it remains reliable, secure, performant, and cost-effective. Responsibilities will include:
• Managing and maintaining AWS infrastructure and services.
• Proactively identifying and implementing AWS cost optimisation opportunities.
• Managing Reserved Instances / Savings Plans and ensuring commitments remain appropriate for current and forecast workloads.
• Improving infrastructure monitoring, dashboards, logging, and alerting.
• Continuously improving our AWS Security Hub posture and addressing identified security findings.
• Reviewing and improving AWS security configurations, IAM permissions, network security, and general cloud security practices.
• Monitoring infrastructure capacity, availability, and performance.
• Supporting incident investigation and root-cause analysis.
• Identifying opportunities to automate repetitive operational tasks.
• Maintaining appropriate backup, recovery, and business continuity processes.
• Keeping infrastructure components and operating systems appropriately patched and up to date.
• Simulating disaster recovery scenarios to ensure our systems are robust. GITHUB, CI/CD & DEPLOYMENTS GitHub is central to our software development and deployment processes. The role will:
• Take ownership of our existing CI/CD processes and continuously improve them.
• Maintain and enhance existing GitHub Actions workflows while creating new workflows where required.
• Work closely with developers to make software deployments reliable, repeatable, and efficient.
• Manage deployments across our various environments, including development, testing/staging, and production.
• Improve deployment controls, release processes, rollback capabilities, and visibility.
• Help troubleshoot build, deployment, and environment-related issues.
• Maintain appropriate management of secrets and credentials used within CI/CD pipelines.
• Promote automation and Infrastructure-as-Code wherever practical. LINUX / UBUNTU ADMINISTRATION Our server workloads predominantly run on Ubuntu Linux. The successful candidate must therefore be highly comfortable working from the Linux command line and will be responsible for:
• Maintaining and administering Ubuntu servers.
• OS patching, upgrades, package management, and general server maintenance.
• Monitoring server health, capacity, disk utilisation, memory, CPU, and system performance.
• Troubleshooting operating system, networking, application, and performance issues.
• Hardening servers and maintaining appropriate security configurations.
• Automating administration and operational activities. Strong Bash scripting skills are essential. The ability to quickly develop reliable scripts for automation, troubleshooting, deployments, maintenance, and operational tasks will be an important part of the role. GOOGLE CLOUD PLATFORM While AWS represents most of our infrastructure, our AI and Big Data platforms are hosted within GCP. The role will provide operational oversight and support for these environments, including:
• Maintaining GCP infrastructure and services.
• Monitoring availability, performance, security, and costs.
• Supporting AI and Big Data infrastructure requirements, including ML pipeline implementation.
• Working with engineering and data teams on infrastructure changes and deployments to ensure system performance is not compromised.
• Applying similar security, monitoring, governance, and cost-management principles across AWS and GCP. Deep GCP specialisation is not necessarily required, but candidates should have practical GCP experience and be comfortable operating in a multi-cloud environment. DOMAINS, DNS & CERTIFICATES Our domains and associated services are currently managed through GoDaddy. Responsibilities will include:
• Domain management and renewals.
• DNS configuration and maintenance.
• SSL/TLS certificate management and renewals.
• Troubleshooting DNS and certificate-related issues.
• Ensuring critical domains, DNS records, and certificates are appropriately monitored and documented. SECURITY & OPERATIONAL EXCELLENCE Security, reliability, and operational discipline will be important aspects of the position. The role will be expected to:
• Continuously improve our cloud security posture.
• Proactively identify vulnerabilities, configuration weaknesses, and operational risks.
• Ensure security findings are communicated, prioritised, and remediated.
• Improve monitoring and alerting so that issues are identified before they become significant incidents.
• Maintain appropriate operational documentation and runbooks.
• Improve disaster recovery and infrastructure resilience.
• Reduce manual processes through sensible automation.
• Establish clear standards and best practices for cloud infrastructure and deployments. REQUIRED EXPERIENCE & SKILLS We are looking for someone with strong practical experience across most of the following:
• Strong commercial experience managing AWS environments.
• Strong knowledge of AWS infrastructure, security, IAM, networking, monitoring, and cost management.
• Experience with AWS Security Hub and cloud security best practices.
• Experience managing AWS Reserved Instances and/or Savings Plans.
• Strong Linux administration, particularly Ubuntu.
• Strong Bash scripting skills.
• Strong experience with Git and GitHub.
• Experience building and maintaining automated CI/CD pipelines, preferably using GitHub Actions.
• Experience taking ownership of existing CI/CD workflows and developing new deployment workflows where required.
• Experience deploying applications across multiple environments, including production.
• Experience with monitoring, alerting, logging, and observability platforms.
• Practical experience with GCP.
• Understanding of DNS, domains, SSL/TLS certificates, and related internet infrastructure.
• Experience automating infrastructure and operational processes.
• Strong troubleshooting and root-cause analysis skills.
• Good understanding of cloud security principles.
• Ability to communicate clearly, drive improvements, make decisions, and solve problems. HIGHLY DESIRABLE Experience in any of the following would be advantageous:
• Infrastructure-as-Code using Terraform, CloudFormation, or similar.
• Docker and containerised workloads.
• Kubernetes / EKS / GKE.
• AWS Organizations and multi-account environments.
• Centralised logging and observability platforms.
• FinOps / cloud cost-management practices.
• Automated vulnerability and security management.
• Database infrastructure and administration.
• High-availability and disaster-recovery architectures.
• AI / ML or Big Data infrastructure experience.
• Python or another scripting/programming language in addition to Bash. WHAT SUCCESS LOOKS LIKE Within this role, success means that:
• Our infrastructure is stable, secure, and well maintained.
• AWS and GCP costs are understood, controlled, and continuously optimised.
• AWS Security Hub scores and security posture continuously improve.
• Monitoring and alerts identify genuine issues quickly without unnecessary noise.
• Production deployments are predictable, automated, and low risk.
• Ubuntu servers remain patched, secure, and healthy.
• Reserved capacity and cloud commitments are actively managed rather than simply renewed.
• Infrastructure and deployment processes become increasingly automated.
• Operational knowledge is documented rather than dependent on individual team members.
• Developers can deploy and operate software efficiently without infrastructure becoming a bottleneck.
• Infrastructure problems are identified proactively rather than only after customers or developers report them. THE PERSON We are looking for someone who combines technical depth with operational ownership. You should be comfortable taking responsibility for an environment, identifying what needs improvement, prioritising the work, and then implementing those improvements yourself. You will enjoy automation, dislike repetitive manual processes, pay attention to security and cost, and be comfortable troubleshooting complex problems across applications, operating systems, networks, CI/CD pipelines, and cloud infrastructure. This is an ideal role for an experienced DevOps / Cloud Infrastructure professional who wants meaningful ownership of a multi-cloud environment while remaining highly hands-on. INTERESTED? Please send your CV to Henry Hsu at henry@hepstar.com.