β’ Turn machine learning engineer/scientist models into production services meeting latency, cost, and reliability targets
β’ Build and maintain SageMaker training, processing, and inference workloads
β’ Build pipelines that orchestrate SageMaker workloads
β’ Make training runs reproducible and configuration-driven so results can be rebuilt from code
β’ Build monitoring for model performance, data drift, and system health
β’ Ensure appropriate alerts are sent when model or system behaviour changes
β’ Ship infrastructure through code review using Terraform and pull-request workflows
β’ Document systems and raise engineering standards
β’ Own the platform and infrastructure that deploy and operate machine learning solutions across live hospitals
Requirements
- Expert Python
- Deep hands-on PyTorch experience; TensorFlow welcome
- Experience with Hugging Face Transformers, timm, scikit-learn, NumPy, pandas, and OpenCV
- Production experience on AWS, specifically Amazon SageMaker, including training and processing jobs, pipelines, model registry, and endpoints
- Practical MLOps experience with MLflow, hyperparameter optimisation, and reproducible, configuration-driven training runs
- Experience with Docker and CUDA-based GPU images
- Terraform experience sufficient to ship models through code review
- Strong software engineering fundamentals, including Git, pull-request workflow, automated testing, and CI/CD
- Ability to explain technical trade-offs clearly to non-engineers
- Desirable: event-driven and streaming architectures on AWS, including Step Functions, Lambda, Kinesis, ECS, and DynamoDB
- Desirable: model monitoring in production, inference logging, and drift detection
- Desirable: exposure to healthcare, medical devices, or another regulated industry
Core Competencies
Demonstrates expertise in building and maintaining machine learning production services, with a strong focus on AWS SageMaker, Python, and MLOps practices. Capable of ensuring model performance and system health through effective monitoring and alerting mechanisms.
Highest-signal resume keywords
- Expert Python
- AWS SageMaker
- MLOps Experience
- PyTorch
- Terraform
ATS Optimization Keywords
Hard Skills
- Python
- PyTorch
- TensorFlow
- MLflow
- Docker
- CUDA
- Git
- Automated Testing
- CI/CD
- Event-Driven Architectures
Soft Skills
- Clear Communication
Industry Keywords
- Healthcare
- Medical Devices
- Regulated Industry
Tools & Technologies
- SageMaker
- Hugging Face Transformers
- Scikit-learn
- NumPy
- Pandas
- OpenCV
- Step Functions
- Lambda
- Kinesis
- DynamoDB
#J-18808-Ljbffr
Senior MLOps Engineer in London employer: Jobtailor
As a Client Services Coordinator at our dynamic company, you will thrive in a supportive work culture that prioritises employee growth and development. We offer comprehensive training, opportunities for advancement, and a collaborative environment where your contributions are valued. Located in a vibrant area, our team enjoys a healthy work-life balance and the chance to engage with diverse clients, making every day rewarding and meaningful.