Now hiring

Site Reliability Engineer @ CME- Group

Walker Drive 20, BelfastOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Salary: £24,500 - 36,500 per year

Requirements: Programming and scripting skills in high-level languages such as Python, Go, Java, or Bash to construct production-grade tooling.Proficiency with Linux-based systems, distributed systems, containerization (Kubernetes/GKE), and public cloud platforms (GCP/GCE).Understanding of modern CI/CD patterns and IaC tools such as Terraform, Ansible, or Kubernetes Config Connector (KCC).Knowledge of core systems and networking concepts (TCP/IP, UDP, HTTP, DNS, load balancing, and messaging protocols).A forward-thinking approach to automation, leveraging Generative AI and Agents (e.g., Gemini) to optimize platform operations.A data-driven mindset to troubleshoot complex, non-linear system behaviors in a fast-paced, high-pressure trading ecosystem.Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively.Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana is preferred.Comfort working within Agile frameworks and collaborative software development lifecycles is preferred.GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified Kubernetes Application Developer (CKAD) certifications are preferred.Experience in Financial Markets or other highly regulated, ultra-low latency, high-concurrency environments is highly beneficial though not essential. Responsibilities: Architect, operate, and support the migration of application platforms—including Messaging (Kafka, RedPanda, MQ, Pub/Sub), Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform.Manage cluster lifecycles, data replication, RBAC, and workload placement.Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana.Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection.Engage with urgency in live production incidents, take ownership of minor incidents, lead post-mortems, and ensure rapid system recovery.Actively identify operational toil and eliminate manual effort through code, automation, and systematic platform improvements.Contribute to disaster recovery (DR) strategies, continuous systems resiliency testing, and present reliability improvement suggestions to the Product backlog.Lead technical discussions for assigned scope, present solution options, collaborate across functional teams, and mentor junior SRE colleagues. Technologies: AIAnsibleArchitectBackboneBashCI/CDCloudGCPGrafanaHTTPSupportJavaKafkaKubernetesLinuxLoad BalancingOpenTelemetryPrometheusPythonRBACSplunkTCP/IPTerraformDevOpsFabricSecurity More:

We are CME Group, the worlds leading derivatives marketplace, and we are seeking a Site Reliability Engineer III to help engineer reliability for our Google Cloud infrastructure, Middleware Platform Engineering team, and core technology foundations supporting our Clearing, Risk, and derivatives applications. We offer a code-first engineering culture that values systematic, automated solutions, along with a robust compensation and benefits package, including bonus and equity programmes, employee stock purchase plan, private medical and dental coverage, mental health benefits, pension, income protection, life assurance, cycle to work, EV car benefit scheme, gym membership, family leave, education assistance, and ongoing employee development. This is a full-time hybrid role based in Belfast (Millennium House), where we invest in our people and provide the opportunity to grow within an organization transforming its production engineering approach.

last updated 36 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores