About this role
Inference System & Performance Engineer - Member of Technical Staff at Callosum. Location: London, England, United Kingdom. Role: optimizing inference, profiling kernels, building tooling Requirements: Deep knowledge of LLM inference internals and distributed GPU systems; proficiency with C++, CUDA, Python or Rust; experience optimizing multi-GPU/multi-node workloads and low-level debugging across GPU, networking, and Linux. Category: Engineering Seniority: Senior Level Tools: C++, CUDA, Python, Rust, Triton, NCCL, NVLink, RDMA, InfiniBand, RoCE, Linux Commitment: Full Time Workplace: Onsite Languages: English