Now hiring

Senior Site Reliability Engineer (Bellville, Western Cape, ZA) @ Sanlam Life Insurance Limited

ZAOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

<p><span style="font-size:14.0pt"><strong>Who are we?</strong></span></p> <p>Sanlam Fintech is a newly established digital first business within the Sanlam Group on a mission to democratize financial advice and solutions for everyone across the African continent. We exist to pioneer inclusive financial confidence helping people build strong foundations to bridge the gap in generational wealth. Our culture us that of agility and constant deployment, we believe in learning fast, learning cheap and learning forward. Our aim is to provide a work environment where knowledge workers can accelerate the development of their ideas and bring innovation to market, at the same time provide compelling career and development proposition that will enable them to realize their dreams.</p><div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Position Overview</b></H2> </div><div><p>The Site Reliability Engineer (SRE) at Sanlam Fintech is responsible for ensuring the reliability, scalability, and performance of our cloud-native infrastructure and services. This role bridges software engineering and operations, applying engineering principles to solve complex infrastructure challenges. The SRE will focus on building and maintaining resilient systems on AWS, implementing comprehensive observability solutions, and driving automation across the infrastructure lifecycle.</p> <p>Operating in a DevOps environment, the SRE takes full ownership of the systems they build and operate, ensuring high availability and optimal customer experience. They work closely with Software Engineers, Platform Engineers, and DevSecOps teams to deliver infrastructure solutions that support Sanlam Fintech business objectives and uphold our commitment to operational excellence.</p></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>What will you do?</b></H2> </div><div><h2><span style="font-size:12.0pt">Reliability &amp; Resilience</span></h2> <ul> <li>Build highly available, fault-tolerant systems on AWS</li> <li>Define SLIs, SLOs and error budgets to track and improve reliability</li> <li>Plan and implement disaster recovery strategies (RTO/RPO)</li> <li>Lead incident response and root cause analysis</li> <li>Build self-healing systems with automated fixes for common failures</li> <li>Run chaos engineering tests to find and fix weaknesses</li> </ul> <h2><span style="font-size:12.0pt">Observability &amp; Monitoring</span></h2> <ul> <li>Set up metrics, logs and traces for full system visibility</li> <li>Build dashboards and alerts for fast incident detection</li> <li>Implement distributed tracing to spot performance issues</li> <li>Set monitoring standards and maintain operational runbooks</li> <li>Publish regular uptime and operational metrics reports</li> </ul> <h2><span style="font-size:12.0pt">Infrastructure Automation</span></h2> <ul> <li>Write and maintain Infrastructure as Code using Terraform and CloudFormation</li> <li>Automate provisioning, configuration and deployments with DevOps/Platform teams</li> <li>Build and manage CI/CD pipelines using GitHub Actions</li> <li>Implement GitOps practices and self-service automation to reduce manual work</li> </ul> <h2><span style="font-size:12.0pt">Cloud Infrastructure &amp; Architecture</span></h2> <ul> <li>Design and optimise serverless solutions (Lambda, API Gateway, Step Functions)</li> <li>Manage and optimise Kubernetes clusters</li> <li>Implement cloud-native patterns like event-driven and microservices architectures</li> <li>Optimise cloud costs and evaluate new AWS services</li> </ul> <h2><span style="font-size:12.0pt">Software Engineering &amp; Development</span></h2> <ul> <li>Build clean, well-structured automation tools and scripts</li> <li>Apply Clean Architecture and Domain-Driven Design to infrastructure code</li> <li>Improve internal tools to boost developer productivity</li> <li>Use AI tools (Claude, GPT) to automate routine tasks</li> </ul> <h2><span style="font-size:12.0pt">Collaboration &amp; Knowledge Sharing</span></h2> <ul> <li>Work with cross-functional teams using Jira, Confluence and JSM</li> <li>Participate in on-call rotations and incident handoffs</li> <li>Mentor junior engineers in SRE practices</li> <li>Document decisions, procedures and run blameless postmortems</li> </ul></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Qualification and Experience</b></H2> </div><div><p><strong>Requ</strong><strong>i</strong><strong>red</strong><strong> </strong><strong>Exper</strong><strong>i</strong><strong>ence</strong></p> <ul> <li>5+ years of experience in systems engineering, DevOps, or site reliability engineering roles </li> <li>3+ years of hands-on experience with AWS cloud services in production environments</li> <li>2+ years of experience with Infrastructure as Code (Terraform and/or CloudFormation) </li> <li>Demonstrated experience in incident management and on-call responsibilities</li> <li>Track record of implementing automation that reduced operational toil</li> </ul> <p> </p> <p><strong>Educat</strong><strong>i</strong><strong>onal</strong><strong> </strong><strong>Background</strong></p> <ul> <li>Bachelor&apos;s degree in Computer Science, Information Technology, Engineering or related field; or equivalent practical experience</li> <li>Relevant professional certifications are advantageous but not required</li> </ul></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>What will make you successful in this role?</b></H2> </div><div><div> <p><strong>Cloud Platforms &amp; Infrastructure</strong></p> <ul> <li>Strong expertise in AWS services including EC2, ECS, EKS, Lambda, API Gateway, Step Functions, S3, RDS, DynamoDB, CloudWatch and networking services such as VPC, Route53 and ALB/NLB</li> <li>Deep understanding of serverless architecture patterns and best practices</li> <li>Experience with Kubernetes cluster management, deployment strategies and service mesh concepts</li> <li>Knowledge of cloud security best practices including IAM, security groups and encryption</li> </ul> <p> </p> <p><strong>Infrastructure</strong><strong> </strong><strong>as</strong><strong> </strong><strong>Code</strong><strong> </strong><strong>&amp;</strong><strong> </strong><strong>Automat</strong><strong>i</strong><strong>on</strong></p> <ul> <li>Proficiency in Terraform for multi-environment infrastructure management</li> <li>Experience with AWS CloudFormation for native AWS resource provisioning </li> <li>Strong scripting skills in Python for automation and tooling development</li> <li>Experience with configuration management tools and practices</li> </ul> </div> <p> </p> <p><strong>Observab</strong><strong>i</strong><strong>l</strong><strong>i</strong><strong>ty</strong><strong> </strong><strong>&amp;</strong><strong> </strong><strong>Mon</strong><strong>i</strong><strong>tor</strong><strong>i</strong><strong>ng</strong></p> <ul> <li>Expertise in Datadog, Cloudwatch and OTEL for full-stack observability including APM, infrastructure monitoring, log management and synthetic testing and monitoring</li> <li>Experience designing and implementing SLI/SLO frameworks</li> <li>Proficiency in creating effective dashboards, alerts and runbooks</li> <li>Understanding of distributed tracing and correlation across services</li> </ul> <p> </p> <p><strong>Development</strong><strong> </strong><strong>&amp;</strong><strong> </strong><strong>Vers</strong><strong>i</strong><strong>on</strong><strong> </strong><strong>Control</strong></p> <ul> <li>Strong experience with GitHub for version control, code review and CI/CD workflows</li> <li>Understanding of Clean Architecture principles and their application to infrastructure code </li> <li>Familiarity with Domain-Driven Design concepts for complex system design</li> <li>Experience building and maintaining CI/CD pipelines using GitHub Actions</li> </ul> <p> </p> <p><strong>Tools</strong><strong> </strong><strong>&amp;</strong><strong> </strong><strong>Platforms</strong></p> <ul> <li>Proficiency with Atlassian suite (Jira, Confluence) for project management and documentation</li> <li>Experience leveraging AI tools (Claude, GPT) for code generation, documentation, and problem-solving</li> <li>Familiarity with containerisation technologies (Docker) and orchestration platforms </li> <li>Experience with Linux system administration and troubleshooting</li> </ul></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Nice To Have Skills</b></H2> </div><div><div> <p><strong>The following skills are desirable and will strengthen a candidate&apos;s application:</strong></p> <ul> <li>Experience with additional cloud providers (Azure and GCP) for multi-cloud strategies</li> <li>Knowledge of FinOps practices and cloud cost optimisation techniques</li> <li>Experience with chaos engineering tools (AWS Fault Injection Simulator, Gremlin and Chaos Monkey)</li> <li>Familiarity with service mesh technologies (Istio and AWS App Mesh)</li> <li>Experience with database reliability engineering and performance tuning</li> <li>Knowledge of compliance frameworks relevant to financial services (POPIA and PCI-DSS)</li> <li>Contributions to open-source projects or community involvement</li> <li>AWS certifications (Solutions Architect, DevOps Engineer or SysOps Administrator)</li> <li>Kubernetes certifications (CKA and CKAD)</li> <li>Experience with event-driven architectures using AWS EventBridge, SNS, SQS or Kafka</li> </ul> </div></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Knowledge and Skills</b></H2> </div><div><div>IT Data Analysis</div><div>IT product enhancements</div><div>Software design and deployments</div><div>Platform management and integration</div><div>Business Requirements</div></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Personal Attributes</b></H2> </div><div><div>Organisational savvy - Contributing through others</div><div>Manages complexity - Contributing through others</div><div>Plans and aligns - Contributing through others</div><div>Optimises work processes - Contributing through others</div></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Build a successful career with us</b></H2> </div><div><p>We’re all about building strong, lasting relationships with our employees. We know that you have hopes for your future – your career, your personal development and of achieving great things. We pride ourselves in helping our employees to realise their worth. Through its five business clusters – Sanlam Fintech, Sanlam Life and Savings, Sanlam Investment Group, Sanlam Allianz, Santam, as well as MiWay and the Group Office – the group provides many opportunities for growth and development.</p></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Core Competencies</b></H2> </div><div><div>Being resilient - Contributing through others</div><div>Collaborates - Contributing through others</div><div>Cultivates innovation - Contributing through others</div><div>Customer focus - Contributing through others</div><div>Drives results - Contributing through others</div></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Turnaround time</b></H2> </div><div><p>The shortlisting process will only start once the application due date has been reached. The time taken to complete this process will depend on how far you progress and the availability of managers. </p></div></div><div style="padding:10.0px 0.0px;border:1.0px solid transparent"><div style="font-size:16.0px;word-wrap:break-word"><H2 style="font-size:1.0em;margin:0.0px"><b>Our commitment to transformation</b></H2> </div><div></div></div></div><p>The Sanlam Group is committed to achieving transformation and embraces diversity. This commitment is what drives us to achieve a diverse, inclusive and equitable workplace as we believe that these are key components to ensuring a thriving and sustainable business in South Africa. The Group&apos;s Employment Equity plan and targets will be considered as part of the selection process.</p>

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores