A well-funded AI engineering organisation is hiring a Senior Software Engineer to deploy, extend, and optimise a leading open-source LLM inference engine, across two core programmes: a scaling inference initiative and a high-assurance/regulated-environment initiative.
What you’ll do
- Deploy, instrument, and monitor open-weight models served via the inference engine
- Build new engine features to support novel hardware architectures
- Extend the engine for accuracy, explainability, and accountability in regulated settings
- Contribute upstream to the relevant open-source community
- Optimise inference performance (latency, throughput, hardware utilisation)
What they’re looking for
- Senior-level Python + low-level programming (C/C++, Rust, CUDA)
- Direct contributions to a major inference/serving engine or similar ML infra project (e.g. PyTorch, Ray)
- Strong grasp of LLM inference internals (KV caching, batching, quantisation)
- Experience upstreaming to active OSS communities
- Hardware accelerator optimisation experience (GPU/TPU/etc.)
- Genuine enthusiasm for AI-assisted development (LLMs/agents as core tools)
Nice to have: regulated-environment experience, AI safety involvement, MLOps/CI/CD depth.
If you think this role would be a good fit for you, please reach out to Charles Duran.