Senior ML Systems Engineer

London, England
Exciting London based role with competitive salary!
Job ID: 470103

We’re partnering with a well-funded, research-driven organisation at the frontier of large-scale ML infrastructure. This is a hands-on technical role for someone who enjoys going deep on performance modelling, distributed systems, and real hardware behaviour — with direct influence over architecture decisions at scale.

What you’ll do

  • Build simulation models for compute, memory, interconnect, and communication behaviour across large-scale ML systems
  • Develop tools to simulate training and inference workloads across distributed accelerator clusters
  • Model distributed execution patterns including collectives, synchronisation, and communication bottlenecks
  • Run experiments and benchmarks on real ML systems to calibrate and validate simulation models
  • Analyse end-to-end performance: throughput, latency, scaling efficiency, and cost/performance tradeoffs
  • Collaborate with hardware, software, networking, and ML teams to communicate findings through design recommendations

What we’re looking for

  • Master’s or PhD in CS, Electrical or Computer Engineering, or related field
  • Strong background in ML systems, distributed systems, performance engineering, or simulation
  • Experience analysing compute, communication, and memory behaviour in large-scale ML systems
  • Hands-on benchmarking, profiling, and measurement of ML systems
  • Familiarity with distributed training concepts: data/tensor/pipeline parallelism, collectives, synchronisation
  • Proficiency in Python, C++, or Rust

To find out more please reach out to Charles Duran.