About this role
Salary: £35,000 - 35,000 per year
Requirements: 5+ years operating distributed systems at scale in productionDeep expertise in Kafka, Ceph, or similar distributed infrastructureProven ability to design for scale and reliabilityExperience mentoring engineers and making technical decisionsWillingness to work 9 a.m.–6 p.m. U.S. Eastern TimeHelpful experience with multi-region replication and disaster recovery, infrastructure cost optimization, and hybrid on-premises/cloud operations Responsibilities: Design and optimize Kafka architecture, including topic governance, partition strategies, throughput, and latencyOperate Ceph, including pool design, placement optimization, and capacity planningAutomate operational work, speed up incident response, and build preventive systemsSupport SQL Server backup and recovery pipelines and provide basic cluster supportDevelop self-service tooling and observability for the Data teamOwn data infrastructure across architecture, deployment, capacity planning, and incident responseMentor engineers, make technical decisions, and help shape platform architecture and engineering standards Technologies: AnsibleArgoCDCloudCephGrafanaHadoopIcingaSupportKafkaPagerDutyPrometheusPuppetSQLTerraformKubernetes More:
We are hiring for PulsePoints Data Platform team, which supports systems processing billions of events daily using Kafka, Hadoop HDFS, and Ceph across bare-metal, cloud, and hybrid infrastructure. Our technology includes Terraform, Ansible, Puppet, ArgoCD, Prometheus, Grafana, Icinga, and PagerDuty. This is a fully remote opportunity to contribute to large-scale infrastructure serving multiple engineering organizations and help shape the platforms future. The role is intended as a technical contributor position rather than a ticket-driven operations job. Our hiring process includes introductory, technical, architecture, and leadership conversations.
last updated 40 week of 2026