We’re partnering with a well-funded, research-driven organisation at the frontier of large-scale ML infrastructure. This is a hands-on technical role for someone who enjoys going deep on performance modelling, distributed systems, and real hardware behaviour — with direct influence over architecture decisions at scale.
What you’ll do
- Build simulation models for compute, memory, interconnect, and communication behaviour across large-scale ML systems
- Develop tools to simulate training and inference workloads across distributed accelerator clusters
- Model distributed execution patterns including collectives, synchronisation, and communication bottlenecks
- Run experiments and benchmarks on real ML systems to calibrate and validate simulation models
- Analyse end-to-end performance: throughput, latency, scaling efficiency, and cost/performance tradeoffs
- Collaborate with hardware, software, networking, and ML teams to communicate findings through design recommendations
What we’re looking for
- Master’s or PhD in CS, Electrical or Computer Engineering, or related field
- Strong background in ML systems, distributed systems, performance engineering, or simulation
- Experience analysing compute, communication, and memory behaviour in large-scale ML systems
- Hands-on benchmarking, profiling, and measurement of ML systems
- Familiarity with distributed training concepts: data/tensor/pipeline parallelism, collectives, synchronisation
- Proficiency in Python, C++, or Rust
To find out more please reach out to Charles Duran.