About this role
Job Summary We are seeking a highly motivated Software Engineer II to transform and modernize the hardware break/fix operations that support one of the world's largest cloud storage fleets. This role is not a traditional operations position. The successful candidate will initially immerse themselves in the end-to-end hardware break/fix lifecycle, partnering with engineering, datacenter operations, vendors, and support organizations to understand current workflows, operational pain points, and reliability risks. They will then drive automation and AI-driven solutions to eliminate manual effort, accelerate hardware recovery, reduce operational risk, and improve service reliability. This is a unique opportunity to work at the intersection of cloud infrastructure, software engineering, automation, AI, and operational excellence. Job Requirements Own and Improve Hardware Break/Fix Workflows Develop deep expertise in the end-to-end hardware break/fix process for a large-scale cloud storage fleet. Coordinate and support hardware replacement activities involving storage controllers, motherboards, disks, networking components, and related infrastructure. Partner with engineering, datacenter operations, support teams, and hardware vendors to identify operational bottlenecks and opportunities for improvement. Analyze incidents and recurring failure patterns to improve repair workflows and fleet reliability. Drive Automation and Modernization Design and build automation solutions that eliminate manual operational tasks. Develop tools and services that automate hardware diagnostics, repair orchestration, case management, approvals, escalations, and reporting. Create self-service platforms and workflow automation that improve operational efficiency and reduce mean time to repair (MTTR). Integrate data from multiple operational systems to provide a unified view of hardware health and repair status. Leverage AI to Transform Operations Apply AI and machine learning technologies to improve operational decision making. Build intelligent systems that assist with: Incident triage Failure prediction Repair prioritization Root cause analysis Knowledge management Automated workflow execution Explore and implement generative AI solutions to improve troubleshooting and operational productivity. Engineering and Software Development Design, build, test, and maintain software services and automation frameworks. Develop scalable backend services, APIs, dashboards, and automation pipelines. Apply modern software engineering practices including CI/CD, testing, observability, and reliability engineering principles. Write high-quality, maintainable, and production-ready code. Reliability and Operational Excellence Identify systemic risks across the storage fleet and drive engineering solutions to address them. Develop metrics and dashboards to measure operational health, repair performance, automation effectiveness, and reliability outcomes. Partner with reliability engineering teams to drive continuous improvements in service availability and customer experience. Contribute to incident reviews and convert operational learnings into engineering investments. Preferred Qualifications Experience with cloud platforms such as Azure, AWS, or GCP. Experience with infrastructure automation tools and scripting. Familiarity with distributed systems, storage systems, networking, or datacenter infrastructure. Experience with AI, machine learning, LLMs, or intelligent automation platforms. Experience with reliability engineering, DevOps, SRE, or production engineering. Experience building dashboards, analytics solutions, or workflow orchestration systems. Knowledge of incident management and operational excellence practices. Education Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or related technical field. 2+ years of software development experience. Strong programming skills in one or more languages such as Python, C#, Java, Go, or similar. Experience developing automation solutions, backend services, or operational tooling. Strong problem-solving and debugging skills. Ability to learn and understand complex operational systems and workflows. Excellent communication and collaboration skills.