Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Full-Time Working from home possible
C
  • We're looking for a senior engineer to help build, maintain and evolve the training framework that powers our frontier-scale language models. This role sits at the intersection of large-scale training, distributed systems, and HPC infrastructure
  • You will design and maintain the core components that enable fast, reliable, and scalable model training - and build the tooling that connects research ideas to thousands of GPUs
  • If you enjoy working across the full stack of ML systems, this role gives you the opportunity and autonomy to have massive impact
  • Build and own the training framework responsible for large-scale LLM training
  • Design distributed training abstractions (data/tensor/pipeline parallelism, FSDP/ZeRO strategies, memory management, checkpointing)
  • Improve training throughput and stability on multi-node clusters (e.g., GB200/300, AMD, H200/100)
  • Develop and maintain tooling for monitoring, logging, debugging, and developer ergonomics
  • Collaborate closely with infra teams to ensure our cluster, container environments, and hardware configurations support high-performance training
  • Investigate and resolve performance bottlenecks across the ML systems stack
  • Build robust systems that ensure reproducible, debuggable, large-scale runs
  • You'll work on some of the most challenging and consequential ML systems problems today
  • You'll collaborate with a world-class team working fast and at scale
  • You'll have end-to-end ownership over critical components of the training stack
  • You'll shape the next generation of infrastructure for frontier-scale models
  • You'll build tools and systems that directly accelerate research and model quality
  • Sample Projects:
  • Build a high-performance data loading and caching pipeline
  • Implement performance profiling across the ML systems stack
  • Develop internal metrics and monitoring for training runs
  • Build reproducibility and regression testing infrastructure
  • Develop a performant fault-tolerant distributed checkpointing system

Benefits

  • Six weeks' paid vacation
  • Equity / stock options
  • RRSP, 401(k), and Pension Scheme contributions
  • Coverage for 100% of your insurance premiums across health, dental, vision, and travel
  • Additional coverage for accessing mental health providers/services
  • Six months of fully paid parental leave, including adoption and surrogacy
  • Financial support for egg freezing and IVF in Canada and the UK
  • A monthly fitness and wellness allowance
  • Globally dispersed company that supports a remote work culture
  • A $2,000 annual education benefit for professional development
  • A weekly stipend for meals when working remotely and catered lunch when working from one of our global offices
  • A monthly arts and culture allowance
  • A monthly quality time allowance
  • A track record of building tools that increase developer velocity for ML teamsExperience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)Experience with training LLMs or other large transformer architecturesDeep familiarity with JAX internals, distributed training libraries, or custom kernels/fused opsStrong engineering experience in large-scale distributed training or HPC systemsComfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelinesExcellent judgment around trade-offs: performance vs complexity, research velocity vs maintainabilityExperience with data pipeline optimization, sharded datasets, or caching strategiesContributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.)Experience working with containerized environments (Docker, Singularity/Apptainer)Background in performance engineering, profiling, or low-level systemsIf some of the above doesn't line up perfectly with your experience, we still encourage you to apply!Strong collaboration skills - you'll work closely with infra, research, and deployment teamsFamiliarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches)

#J-18808-Ljbffr

Senior Machine Learning Systems Engineer (Frameworks & Tooling) employer: Cohere

Cohere is an exceptional employer that fosters a dynamic and inclusive work culture, where innovation thrives and employees are empowered to make a real impact in the AI landscape. With generous benefits such as a weekly lunch stipend, comprehensive health coverage, and a robust education stipend, we prioritise employee well-being and growth. Our London office offers a collaborative environment with opportunities for meaningful engagement with enterprise clients, making it an ideal place for those looking to advance their careers in cutting-edge technology.

C

Contact Details:

Cohere Recruitment Team