About this role
Salary: £100,000 - 100,000 per year
Requirements: Experience with compiler techniques to improve AI model inference performance on CPU or CPU/XPU hybrid systemsExperience with LLVM/MLIR development is preferredExperience profiling and optimizing AI frameworks such as TensorFlow, PyTorch, ONNX, and llama.cpp is preferredStrong technical writing skills with prior publications or reports are preferred Responsibilities: We implement compiler-based performance optimizations to improve inference latency and throughput on CPU and CPU/XPU hybrid systemsWe optimize JIT-level compute graphs through operator fusion, memory allocation, and related techniquesWe profile end-to-end inference workflows to identify hotspots and bottlenecks in AI frameworks such as TensorFlow, PyTorch, ONNX, and llama.cppWe propose and implement optimization strategies such as kernel tuning and graph-level optimizationsWe track and analyze the latest advancements in AI and compiler research, including academic papers and open-source projectsWe produce actionable insight reports summarizing trends, benchmarks, and potential optimizations Technologies: AILLVMPyTorchTensorFlowBackend More:
We are seeking a skilled AI Compiler Optimization Engineer to help us optimize AI model inference performance through advanced compiler technologies. In this role, we focus on performance tuning for CPU and hybrid CPU/XPU heterogeneous architectures, profiling AI frameworks to uncover new optimization opportunities, and turning industry research into practical insights.
last updated 37 week of 2026