About this role
Annapurna ML Neuron is the team that delivers the software that powers the existing and upcoming Trainum families of Machine Learning Accelerated EC2 instances. We build a Compiler, Collective Communication, Drivers, Runtime and a fully integrated suite of Pytorch and JAX stacks, providing the highest-scale, distributed Training and Inference servers available. We work from the lowest level collaborating with chip and server design up to and including massive scale inference and training for frontier models. We're seeking a Principal Program Manager for the Annapurna ML, Neuron core team. In this role you will be owning the software technology bring up lifecycle from early hardware design ideation, to landing hardware in labs and data centers, developing, debugging, and qualifying full servers for customer and Neuron developers. You will work stakeholders to secure and allocate prototype and production Trainium instances. Establish and maintain dashboards, utilization metrics, and recurring review mechanisms to proactively identify new usecases, gaps, and drive resolution. You will be responsible for scoping and delivering large projects end-to-end. Responsibilities include collection of business and systems requirements from internal and external customers, writing specifications, driving project schedules from design to release, and managing the production launch. You will lead and coordinate design/implementation efforts between Neuron, Trainium hardware teams, as well as internal teams and external customers. You will be expected to make appropriate tradeoffs to optimize time-to-production, clearly communicate goals, roles, responsibilities, and desired outcomes to internal cross-functional and remote project teams. The right candidate will possess a strong technical and program management background, demonstrate experience in new hardware design and bringup, will have demonstrated experience leading medium to large projects, and will have a well-rounded technical background in at most of the following: current machine learning technologies, products spanning multiple technical domains from higher performance networking/compute and drivers to compilers and software development programs. You must be able to thrive and succeed in an entrepreneurial environment and not be hindered by ambiguity or competing priorities. This means you are not only able to develop and drive high-level strategic initiatives, but can also roll up your sleeves, dig in and get the job done. As a Principal TPM, you will anticipate bottlenecks, provide escalation management, anticipate and make trade-offs, and balance the business needs versus technical constraints. An ability to take large, complex projects and break them down into manageable pieces, develop functional specifications, then deliver them in a timely manner. In this role you will: · Drive execution of projects · Provide technical direction with limited assistance · Lead cross functional project meetings with VP level audience and stakeholders · Lead milestone reviews · Present project status to the executive team Key job responsibilities Work with Engineering leadership, Product Management, and Business Development to help define requirements and roadmaps, work with Engineering leadership to drive efficiencies in engineering development and ensure the right things are delivered to Customers at the right time. We work like a startup - moving fast and building new things. About the team Annapurna ML Neuron is the team that delivers the software that powers the existing and upcoming Trainum families of Machine Learning Accelerated EC2 instances. We build a Compiler, Collective Communication, Drivers, Runtime and a fully integrated suite of Pytorch and JAX stacks, providing the highest-scale, distributed Training and Inference servers available. We work from the lowest level collaborating with chip and server design up to and including massive scale inference and training for frontier models.