About this role
Salary: £64,000 - 104,000 per year
Requirements: Strong production experience with AWS services, particularly EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, IAM, and VPC networkingProficiency authoring and maintaining Terraform modules for production infrastructureProficiency authoring and maintaining Puppet modules, or equivalent agent-based configuration management, for fleet managementSolid Python skills, including writing and maintaining production daemonsDeep Linux systems knowledge, particularly Ubuntu, with familiarity with Apache/Nginx, PHP-FPM, Varnish, systemd, filesystem mounts, and networking fundamentalsUnderstanding of distributed systems concepts, including consensus, leader election, distributed locking, eventual consistency, and related tradeoffsProduction experience building and maintaining observability pipelines using Prometheus, Grafana, Loki, or equivalent toolsComfort working in a GitLab-based CI/CD workflowClear communication skills, including documenting architectural decisions and explaining technical tradeoffs to technical and non-technical stakeholdersPreferred: Hands-on experience with distributed storage systems such as Ceph, GlusterFS, JuiceFS, CubeFS, or AWS EFS, particularly migration or evaluationPreferred: Familiarity with etcd or similar distributed key-value stores such as Consul or ZooKeeper, including watch APIs, TTL-based locking, and cluster operationsPreferred: Experience with Varnish and VCL, especially dynamic backend routing or multi-tenant configurationsPreferred: Working knowledge of PHP for understanding and maintaining infrastructure-to-application integration scriptsPreferred: Background in multi-tenant SaaS platform design, particularly database-per-tenant models on shared infrastructurePreferred: Familiarity with Moodle LMS or education technology platformsPreferred: Experience with secrets management solutions such as AWS Secrets Manager, HashiCorp Vault, or Parameter Store, and automated credential rotationPreferred: Experience designing zero-downtime deployment strategies for VM-based, non-containerized environments Responsibilities: Design, build, and maintain AWS infrastructure using services including EC2, RDS, S3, SQS, Lambda, ALB, ElastiCache, Route 53, and VPC networkingWrite and maintain Puppet modules to configure and manage EC2 instance fleets across multiple auto-scaling groupsMaintain and extend Python-based automation and tooling for platform operationsOperate and improve distributed service discovery and configuration management using etcdManage and tune a multi-tier caching strategy using Varnish, Redis/Valkey, and PHP OPcacheRun and scale the observability stack, including Prometheus, Grafana, Loki, Fluentd, and PagerDuty, and participate in on-call rotationsEvaluate and implement distributed storage solutions as the platform evolvesImprove deployment workflows and release processesCollaborate with internal teams on API contracts, integration patterns, and operational toolingParticipate in incident response, root cause analysis, and platform reliability improvements Technologies: APIAWSLambdaBackendCI/CDCloudCephEC2GitLabGrafanaIAMKubernetesLinuxNginxPHPPagerDutyPrometheusPuppetPythonRedisTerraformUbuntuZooKeeperDevOps More:
We are hiring a full-time Senior Cloud Infrastructure Engineer to help build, scale, and evolve our multi-tenant SaaS hosting platform on AWS. Our platform dynamically provisions, manages, and scales hundreds of Moodle LMS instances for education clients using custom orchestration tooling, distributed service discovery, and infrastructure as code. This is a hands-on infrastructure role spanning Terraform, Puppet, Python automation, and observability. The platform is not containerized and does not use Kubernetes; the role offers ownership and influence over our platform architecture and direction as we grow. We are an equal employment opportunity/affirmative action employer and consider qualified applicants without regard to protected characteristics.
last updated 40 week of 2026