About this role
Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.
The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.
RESPONSIBILITIES
Reliability Engineering
• Define and manage SLIs, SLOs, and Error Budgets for critical services • Drive production readiness reviews and reliability requirements into architecture and design • Perform capacity planning, failure mode analysis, and dependency risk assessments • Identify systemic reliability risks and drive remediation before they cause customer impact Observability
• Design monitoring, alerting, logging, and tracing solutions using modern observability tooling • Improve signal-to-noise ratio and reduce alert fatigue • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics Incident Management
• Lead technical response for high-severity incidents • Drive blameless postmortems and root cause analysis focused on systemic fixes • Continuously improve detection, response, and recovery processes • Participate in an on-call rotation Automation & Toil Reduction
• Identify and eliminate manual, repetitive operational work through automation • Build self-healing systems, tooling, and scripts to reduce human intervention • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green) • Support Infrastructure as Code (Terraform, Bicep, or similar) Performance & Scalability
• Conduct load testing, performance benchmarking, and bottleneck analysis • Partner with engineering to design systems for horizontal scalability and fault tolerance Collaboration & Culture
• Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting) • Mentor engineers on SRE best practices • Promote a culture of engineering-driven reliability over reactive operations
EDUCATION AND EXPERIENCE QUALIFICATIONS
Required Qualifications
• 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering • 2+ years experience with Kubernetes and containerized workloads • 4-year degree in Computer Science or related field
Preferred Qualifications
• Experience with chaos engineering or resiliency testing • Experience with high-volume, high-availability transactional systems • Experience with AI-assisted observability or operational automation • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES
• Strong programming/scripting skills (Python, Go, Java, or Node.js) • Demonstrated experience defining and operating against SLOs/Error Budgets • Strong skills in leading incident response and root cause analysis for production systems • Solid understanding of distributed systems and microservices architecture • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP) • Expertise with observability platforms and monitoring strategy
This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.
Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.