At a Glance
- Tasks: Own and optimise ML infrastructure, ensuring reliable deployment and monitoring of predictive models.
- Company: Join Circadia Health, a pioneering company transforming patient monitoring with AI.
- Benefits: Competitive salary, flexible work environment, and the chance to make a real impact.
- Other info: Dynamic startup culture with opportunities for growth and collaboration across teams.
- Why this job: Be at the forefront of healthcare innovation, improving patient outcomes through technology.
- Qualifications: 4+ years in MLOps or related fields, strong Python skills, and experience with AWS.
The predicted salary is between 36000 - 60000 £ per year.
As an ML Ops Engineer at Circadia Health, you will own the infrastructure and operational lifecycle of the machine learning systems that power our clinical monitoring platform. You will build and maintain the production ML pipelines, deployment infrastructure, and monitoring systems that enable Circadia's predictive models to identify early signs of clinical deterioration. Reporting to the Principal ML Engineer, you will work across ML, backend, data, and clinical teams to ensure models are reliably trained, versioned, deployed, and monitored in both cloud and edge environments. You will be a key driver in elevating Circadia's ML practice – from reproducibility and experiment tracking to CI/CD for models and operational observability. This is a high-ownership role at a lean company where production reliability, rapid iteration, and pragmatic engineering are essential. Your work will directly impact patient outcomes by ensuring our predictive models are always running, always accurate, and always improving.
Key Responsibilities
- Own and extend Circadia’s ML pipeline orchestration using Apache Airflow, including training, evaluation, and deployment workflows.
- Build and maintain automated pipelines for model retraining, validation, and promotion across development, staging, and production environments.
- Implement pipeline monitoring, alerting, and failure recovery to eliminate silent failures and ensure operational reliability.
- Design pipeline architectures that support rapid experimentation while enforcing production-grade reproducibility.
- Deploy and manage ML models on AWS infrastructure (e.g. AWS Batch for batch inference workloads).
- Support deployment of models to edge devices, including Circadia’s clinical monitoring hardware, working with firmware and embedded engineering teams as needed.
- Manage model versioning, promotion, and rollback workflows through the MLflow model registry.
- Evaluate and implement strategies for safe model rollouts (e.g. shadow deployments, canary releases) as the platform matures.
- Maintain and improve the MLflow-based experiment tracking and model registry infrastructure.
- Establish conventions for experiment logging, artifact storage, model metadata, and lineage tracking.
- Enable ML engineers to move seamlessly from experimentation to production deployment with minimal friction.
- Implement and maintain training data versioning and dataset management practices to ensure reproducibility of model training runs.
- Track dataset lineage, labeling provenance, and feature dependencies alongside model versions.
- Collaborate with ML engineers and data engineers to formalise dataset release and validation workflows.
- Build monitoring systems for model performance in production, including data drift detection, prediction quality tracking, and alerting on degradation.
- Implement operational dashboards for pipeline health, compute utilisation, and deployment status.
- Collaborate with data engineering to ensure upstream data quality and pipeline reliability for ML feature inputs.
- Develop incident response procedures and runbooks for ML system failures.
- Manage and optimise AWS compute resources (Batch, EC2, or similar) used for model training and inference.
- Design infrastructure-as-code solutions for reproducible ML environments.
- Drive cost optimisation across ML compute, storage, and data transfer.
- Support Snowflake integrations for feature generation and training data pipelines.
- Introduce and champion ML engineering best practices including CI/CD for models, automated testing for ML pipelines, and reproducible training workflows.
- Build internal tooling and templates that accelerate the ML development-to-production cycle.
- Document operational processes, architecture decisions, and onboarding materials for the ML platform.
- Participate in architecture discussions and technical planning to ensure ML systems scale with Circadia’s growth.
- Ensure all ML pipelines and infrastructure meet healthcare security and privacy requirements, including HIPAA and SOC 2.
- Apply best practices for handling Protected Health Information (PHI) in training data, model artifacts, and inference outputs.
- Maintain audit trails for model decisions, data access, and deployment history.
Required Qualifications
- 4+ years of experience in MLOps, ML Engineering, DevOps, or a closely related infrastructure role.
- Strong proficiency in Python for ML pipeline development, tooling, and automation.
- Hands-on experience with ML pipeline orchestration tools, particularly Apache Airflow.
- Experience with model registries and experiment tracking platforms (MLflow preferred).
- Experience deploying and operating ML workloads on AWS (Batch, EC2, S3, IAM, CloudWatch).
- Solid understanding of the ML lifecycle: training, evaluation, deployment, monitoring, and retraining.
- Experience with containerisation (Docker) and infrastructure-as-code.
- Proficiency with Git and version control workflows.
- Familiarity with SQL and data warehousing platforms (Snowflake preferred).
- Experience implementing monitoring, logging, and alerting for production systems.
- Strong debugging and incident response skills for complex distributed systems.
Preferred Qualifications
- Experience deploying models to edge or embedded devices.
- Background in healthcare, medical devices, or clinical data systems.
- Familiarity with model serving frameworks (e.g., TorchServe, TF Serving, Triton, or custom solutions).
- Experience with CI/CD systems for ML (e.g., GitHub Actions, Jenkins, or similar).
- Experience with data versioning tools (e.g., DVC, LakeFS, or similar).
- Experience supporting data science or ML research teams in a production context.
- Exposure to HIPAA compliance and healthcare security best practices.
- Experience with distributed compute frameworks (e.g. Apache Spark, Dask) for large-scale data processing.
- Experience with streaming or real-time inference architectures.
What You Bring
- You take ownership of ML infrastructure end-to-end — from training pipelines to production monitoring.
- You care deeply about reliability, reproducibility, and operational excellence in ML systems.
- You have strong opinions (loosely held) on how to build a great ML platform, and you’re eager to put them into practice.
- You are comfortable working in a startup environment where you’ll wear multiple hats and move fast.
- You communicate clearly across engineering, data science, and clinical teams.
- You’re motivated by building technology that directly improves patient care.
Why Circadia Health
Circadia Health is redefining patient monitoring through contactless sensing and AI-driven clinical insights. As we scale from tens of thousands to hundreds of thousands of monitored patients, our data infrastructure is central to everything we do. You’ll have the opportunity to:
- Work on real-world healthcare problems with measurable patient impact.
- Build data systems that power clinical-grade AI and ML.
- Take ownership in a fast-growing, mission-driven company.
- Collaborate with a highly skilled, multidisciplinary team.
ML Ops Engineer employer: MedTech Innovator
Circadia Health is an exceptional employer, offering a dynamic work environment where innovation meets healthcare. With a strong focus on employee growth, we provide opportunities for professional development and collaboration with leading experts in the field. Our commitment to meaningful work in health AI ensures that you will be part of a team making a real impact on patient care, all while enjoying the benefits of a supportive and forward-thinking culture.
StudySmarter Expert Advice🤫
We think this is how you could land ML Ops Engineer
✨Tip Number 1
Network like a pro! Reach out to folks in the industry, attend meetups, and connect with people on LinkedIn. You never know who might have the inside scoop on job openings or can refer you directly.
✨Tip Number 2
Show off your skills! Create a portfolio showcasing your ML projects, especially those involving pipelines and AWS. This gives potential employers a taste of what you can do and sets you apart from the crowd.
✨Tip Number 3
Prepare for interviews by brushing up on your technical knowledge and soft skills. Practice common ML Ops scenarios and be ready to discuss how you've tackled challenges in past roles. Confidence is key!
✨Tip Number 4
Don't forget to apply through our website! It’s the best way to ensure your application gets seen. Plus, we love seeing candidates who are proactive about their job search.
We think you need these skills to ace ML Ops Engineer
Some tips for your application 🫡
Tailor Your CV:Make sure your CV is tailored to the ML Ops Engineer role. Highlight your experience with ML pipelines, AWS, and any relevant tools like Apache Airflow. We want to see how your skills align with what we're looking for!
Craft a Compelling Cover Letter:Your cover letter is your chance to shine! Share your passion for ML and how you can contribute to our mission at Circadia Health. Be sure to mention specific projects or experiences that showcase your expertise.
Showcase Your Projects:If you've worked on any cool ML projects, don’t hold back! Include links to your GitHub or any relevant portfolios. We love seeing practical examples of your work and how you tackle real-world problems.
Apply Through Our Website:We encourage you to apply directly through our website. It’s the best way to ensure your application gets into the right hands. Plus, it shows us you're genuinely interested in joining our team!
How to prepare for a job interview at MedTech Innovator
✨Know Your ML Ops Inside Out
Make sure you brush up on your knowledge of ML Ops principles, especially around pipeline orchestration with tools like Apache Airflow. Be ready to discuss how you've built and maintained production ML pipelines in the past, as this will show your hands-on experience.
✨Showcase Your AWS Skills
Since you'll be deploying models on AWS, it's crucial to demonstrate your proficiency with AWS services like Batch, EC2, and S3. Prepare examples of how you've managed compute resources and optimised costs in previous roles to highlight your practical experience.
✨Emphasise Collaboration
This role requires working closely with various teams, so be prepared to share examples of how you've successfully collaborated with ML engineers, data engineers, and clinical teams. Highlight your communication skills and how they’ve helped drive projects forward.
✨Prepare for Problem-Solving Questions
Expect to face scenario-based questions that test your debugging and incident response skills. Think of specific challenges you've encountered in ML systems and how you resolved them, as this will showcase your critical thinking and operational excellence.