About this role
Salary: £100,000 - 100,000 per year
Requirements: Formal training, certification, or equivalent practical experience in software engineering conceptsHands-on experience with system design, application development, testing, and operational stability in production environmentsAdvanced proficiency in Python for building production-grade services and toolingProficiency with automation and continuous delivery methodsHands-on experience with AWS and Terraform for infrastructure delivery and lifecycle managementStrong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patternsPractical knowledge of observability and instrumentation across metrics, logs, and tracesComfort with on-call operations and production troubleshootingHands-on production experience operating LLM inference servers such as vLLM and llm-d, or directly equivalent serving stacksHands-on experience hosting and serving LLMs on Amazon EKS and/or Amazon SageMaker, and on local GPU infrastructureKnowledge of LLM reliability and risk considerations, including latency/throughput trade-offs, model and weight versioning, prompt/response logging, and safe rollout patternsExperience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patternsExperience building AI agents using frameworks such as LangChain, CrewAI, LangGraph, or similar orchestration platformsExperience operating or integrating serving platforms such as KServe, Ray Serve, NVIDIA Triton Inference Server, Text Generation Inference (TGI), alongside vLLM/llm-dFamiliarity with Amazon SageMaker JumpStart, SageMaker Endpoints, and Amazon Bedrock for managed model hostingExperience with online LLM quality monitoring, such as hallucination, toxicity, and drift detection, and tracing via OpenTelemetry conventionsContributions to open-source LLM serving or inference projects, such as vLLM, llm-d, Ray, KServe, or Triton Responsibilities: Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructureBuild backend services and APIs that enable reliable operation of AI infrastructure in productionOperate and scale LLM serving infrastructure, including model hosting, request routing, continuous batching, and KV-cache optimizationDeploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon EKS, Amazon SageMaker, on-prem, and local GPU clusters using reproducible infrastructure as code and continuous delivery pipelinesImplement observability with logs, metrics, traces, dashboards, and actionable alerting for LLM and GPU workloadsTune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization techniquesLead reliability engineering for LLM endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model-quality regressionsParticipate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-upsIdentify recurring operational issues and automate remediation to improve platform stability and developer experienceBuild and maintain multi-agent systems with strong orchestration where appropriateContribute to an inclusive team culture and help drive adoption of leading-edge technologies through communities of practice Technologies: AIAI AgentsAWSBackendCloudIncident ManagementSupportKubernetesLLMMachine LearningMarketingOpenTelemetryPythonTerraformvLLMGrafanaModel ServingPrometheus More:
We are partnering directly with JPMorgan Chase for this role on our AI and Machine Learning Platform team. We are building and scaling AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI, with a focus on reliable, cost-efficient LLM inference at scale. We work with cloud and Kubernetes-based deployments, deep observability, and production-grade operational rigor across AWS, Amazon EKS, Amazon SageMaker, on-prem, and local GPU environments. J.P. Morgan is a global leader in financial services, and our Corporate Functions teams support our businesses, clients, customers, and employees across finance, risk, human resources, marketing, and more. We value diversity, inclusion, and equal opportunity, and we make reasonable accommodations for applicants and employees religious practices and beliefs as well as mental health or physical disability needs.
last updated 35 week of 2026