At a Glance
- Tasks: Lead the charge in making roads safer with innovative AI solutions.
- Company: Join CMT, a leader in telematics and AI for safer mobility.
- Benefits: Enjoy competitive salary, unlimited PTO, and flexible work options.
- Other info: Be part of a diverse team committed to innovation and safety.
- Why this job: Make a real impact by preventing crashes and saving lives worldwide.
- Qualifications: 7+ years in Site Reliability Engineering and strong AWS skills required.
The predicted salary is between 60000 - 80000 £ per year.
CMT is looking for a Principal Site Reliability Engineer I, Machine Learning to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — understanding and reducing risk, detecting crashes, and getting people life-saving help. The problems are hard. The impact is real. No matter your role, your work will matter at CMT. CMT is looking for a collaborative, customer-committed, and creative SRE team member who wants to join us in making roads safer by making drivers better!
Responsibilities:
- Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions.
- Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs.
- Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery.
- Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies.
- Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle.
- Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades.
- Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil.
- Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation.
- Complete any additional tasks as they arise.
Qualifications:
- Bachelor’s degree or equivalent years of experience and/or certification in a related field.
- 7+ years working in Site Reliability Engineering or Information Technology.
- Design and document systems, including writing and reviewing code, to automate away problems within your team’s domain.
- Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM.
- Intermediate to expert experience monitoring services and applications using tools such as CloudWatch Metrics, CloudWatch Logs, and Datadog, including defining and configuring alerts and SLO reports.
- Intermediate to expert experience maintaining the uptime and scalability of AWS compute services used for Machine Learning and Data Science workloads, specifically EC2 and EKS.
- Intermediate to expert coding skills in at least one programming language; we work primarily in Python.
- Intermediate to expert experience using Infrastructure as Code platforms and CI/CD pipelines, specifically Terraform, to manage AWS infrastructure and services.
- In-depth knowledge & experience with Linux operating systems (Amazon Linux, Ubuntu) on EC2 and Docker / Kubernetes.
- Experience with leading projects in system design, architecture changes, and technology selection.
Compensation and Benefits:
- Fair and competitive salary based on skills and experience, and annual performance bonus.
- Equity may be awarded in the form of Restricted Stock Units (RSUs).
- Medical, Dental, Vision and Life Insurance, matching 401k, short-term & long-term disability and parental leave.
- Unlimited Paid Time Off including vacation, sick days & public holidays.
- Flexible scheduling and work from home policy depending on role and responsibilities.
Base Salary Range:
The base salary range for this position is: $142,000 to $177,600. This range is specifically for Cambridge, MA.
Additional Perks:
- Work on a mission with real impact: crashes prevented, injuries avoided, lives protected around the world.
- Join an industry leader — 65 million drivers protected, powering 140+ programs across 25 countries.
- Recognized innovator in mobility AI, earning top honours including the TIME Industry Leader in AI, a Gold Edison Award, and the Artificial Intelligence Excellence Award for AI for Social Good. CMT is also Great Place to Work Certified.
- Be part of the team inventing the future of mobility and road safety.
- Move fast, own outcomes, do work that matters.
- High ownership, small teams, and direct access to leadership — no layers between your work and its impact.
- Unlimited PTO, flexible scheduling, competitive salary, annual performance bonus, RSUs, and full benefits including medical, dental, vision, and 401k match.
- Summer Fridays provide team members with half days to recharge.
- Join one of our employee resource groups: Black, AAPI, LGBTQIA+, Women, Book Club, and Health & Wellness.
- Comprehensive wellness, education, and employee assistance programs.
Commitment to Diversity and Inclusion:
At CMT, we believe the best ideas come from a mix of backgrounds and perspectives. We are an equal-opportunity employer committed to creating a workplace and culture where everyone feels valued, respected, and empowered to bring their unique talents and perspectives. Diversity is essential to our success, and we actively seek candidates from all backgrounds to join our growing team. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status or disability state.
CMT is headquartered in Cambridge, MA. To learn more, visit www.cmtelematics.com and follow us on Instagram @cmt.ai.
About Cambridge Mobile Telematics:
Cambridge Mobile Telematics (CMT) is the world’s largest telematics and AI company for safer mobility. Its mission is to make the world’s roads and drivers safer. The company’s AI-driven platform, DriveWell Fusion®, proactively identifies and reduces driving risk, leading to fewer crashes and injuries. To date, CMT’s technology has helped prevent over 126,000 crashes worldwide. CMT enables partners to measure risk, detect crashes, provide life-saving assistance, and streamline claims. Headquartered in Cambridge, MA, CMT operates globally with offices in Budapest, Hungary; Chennai, India; Seattle, Washington; Tokyo, Japan; and Zagreb, Croatia. Learn more at www.cmt.ai.
Principal Site Reliability Engineer, Machine Learning in Cambridge employer: Cambridge Mobile Telematics
CMT is an exceptional employer, offering a dynamic work culture where innovation meets impact. As a Principal Software Engineer I, Full Stack, you will not only lead critical projects but also mentor teams in a collaborative environment that values diversity and inclusion. With competitive compensation, unlimited PTO, and a commitment to employee growth, CMT empowers you to make a real difference in mobility safety while enjoying a flexible work-life balance in the vibrant city of Cambridge, MA.
Contact Details:
Cambridge Mobile Telematics Recruitment Team
StudySmarter Expert Advice🤫
We think this is how you could land Principal Site Reliability Engineer, Machine Learning in Cambridge
✨Join Local Tech Meetups
Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Cambridge Mobile Telematics or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!
✨Contribute to Open Source Projects
Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Cambridge Mobile Telematics.
✨Tap into Online Developer Communities
Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Cambridge Mobile Telematics.
✨Explore Job Boards Specifically for Tech Roles
Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Cambridge Mobile Telematics that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!
We think you need these skills to ace Principal Site Reliability Engineer, Machine Learning in Cambridge
Some tips for your application 🫡
Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.
Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Cambridge Mobile Telematics.
Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Cambridge Mobile Telematics and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!
Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!
How to prepare for a job interview at Cambridge Mobile Telematics
✨Brush Up on Your Coding Skills
For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.
✨Know Your Tools and Frameworks
Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Cambridge Mobile Telematics uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.
✨Showcase Your Projects
Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.
✨Prepare for Behavioural Questions
While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.