Principal Software Reliability Engineer, Product Engineering in London
Principal Software Reliability Engineer, Product Engineering

Principal Software Reliability Engineer, Product Engineering in London

London Full-Time 60000 - 84000 ÂŁ / year (est.) No home office possible
Entrust

At a Glance

  • Tasks: Drive reliability efforts and define the roadmap for product engineering.
  • Company: Entrust, a leader in identity-centric security solutions.
  • Benefits: Flexible work options, career growth opportunities, and a collaborative culture.
  • Why this job: Shape the future of security solutions and make a real impact.
  • Qualifications: 8+ years in software engineering with a focus on reliability.
  • Other info: Join a diverse team committed to innovation and inclusion.

The predicted salary is between 60000 - 84000 ÂŁ per year.

Join us at Entrust. At Entrust, we’re shaping the future of identity centric security solutions. From our comprehensive portfolio of solutions to our flexible, global workplace, we empower careers, foster collaboration, and build solutions that help keep the world moving safely.

Headquartered in Minnesota, Entrust is an industry leader in identity-centric security solutions, serving over 150 countries with cutting‑edge, scalable technologies. But our secret weapon? Our people. It’s the curiosity, dedication, and innovation that drive our success and help us anticipate the future.

This is a Product Reliability position, not an infrastructure SRE role. Our DevOps team manages the infrastructure platform; this role focuses on application and service‑level reliability, working directly with product engineers. This is the first role of its kind in product engineering. Reporting to the VP of Product Engineering for Consumer Identity, you’ll drive reliability efforts across the team: defining the roadmap, prioritising initiatives, and partnering with engineering directors and senior ICs to deliver them.

Why Join Us

  • Greenfield opportunity: You’ll define Product Reliability as a discipline here. Build the playbook, not inherit one.
  • High‑impact domain: Consumer Identity powers identity verification and biometric authentication for some of the world’s largest financial institutions. Our reliability directly impacts fraud prevention and customer onboarding at scale.
  • Real authority: Direct line to VP Engineering, budget for tooling, seat at architecture council and service reviews.
  • Strong foundation: We’re not firefighting. 99.98% uptime means you’re optimising, not triaging chaos.
  • Technical depth: Work across ML pipelines, computer vision systems, and mobile SDKs (not just YAML and dashboards).
  • Ownership culture: Engineers own their services end‑to‑end; you’ll amplify that, not replace it.

Experience Level

  • Staff SRE: 8+ years in software engineering, 4+ years in reliability/SRE. Drives reliability initiatives across multiple teams; hands‑on with complex systems.
  • Principal SRE: 15+ years in software engineering, 6+ years in reliability/SRE. Sets technical direction org‑wide; influences business‑unit‑level reliability strategy.

We’re open to either level. Scope and compensation will match your experience. Principal candidates should demonstrate cross‑org impact and a track record of building reliability programs from scratch.

Current State Incident Analysis (2020–2025)

  • Postmortem volume peaked in 2023, down 48% since then despite increased release cadence.
  • P0‑to‑P1+ ratio remains stable despite lower overall incident volume.
  • 65% change‑induced incidents (deployments, migrations, config changes); 35% organic (third‑party outages, expirations, attacks).
  • Change‑induced ratio improved modestly: 69% → 62%.
  • Detection time: 35 min → 18 min.
  • Customer‑first detection: 40% → 22%.

Availability Targets

  • 2024 & 2025 average uptime: 99.98% (as available in our public status page).
  • Goal: Consistent 99.99% (four nines) average uptime, SLO breach reductions.

System Simplification

We’re reducing system complexity to narrow the reliability target area:

  • Microservices (K8s deployments/rollouts) reduced 29% from peak, with further cuts planned for 2026.
  • Goal: Smaller footprint, higher reliability, lower cost for new regions.

Role Objectives

Primary goal: Improve release safety, reduce releases that cause downtime or SLO degradation. We already have foundational systems in place:

  • Automated test coverage and crowd testing.
  • A/B testing and dark canaries.
  • Progressive rollouts (infrastructure and application level).
  • Back‑testing against historical data.

To consistently exceed four nines, we need to mature these systems and build new capabilities.

Ideal Candidate Profile

Mindset

  • Passionate about reliability as a discipline, not just a checkbox.
  • Focused on reliability, not product features, but willing to learn the product to understand impact.
  • Hands‑on: eager to build tooling and systems.
  • Pragmatic about balancing reliability with development velocity.

Required Skills

Software Engineering

  • Strong software engineering in at least one of our backend languages (Python, Ruby, Node.js); able to navigate most of our codebase.
  • Experience building reliability tooling: progressive delivery, automated rollbacks, monitoring/alerting.

Reliability Patterns

  • Deep knowledge of resilience patterns: circuit breakers, bulkheads, back‑pressure, retries with backoff, rate limiting, load shedding, graceful degradation.
  • Solid incident management and blameless postmortem practices.

Observability

  • Proficiency with observability: distributed tracing, structured logging, metrics instrumentation.
  • Uses data to drive decisions: experienced with SLIs, SLOs, and error budgets.

Communication

  • Skilled at influencing without authority.
  • Able to hold deep technical reliability discussions with senior ICs.

Nice‑to‑Have

  • Experience with chaos engineering (fault injection, game days, controlled failure experiments).
  • ML system reliability experience (mixed I/O and CPU‑bound workloads, non‑deterministic behavior, model serving).
  • Familiarity with our specific stack (Datadog, Kubernetes, AWS, GitLab CI/CD).
  • Experience leveraging LLMs for code analysis, design doc review, or automated runbook generation.
  • On‑call experience in a high‑availability environment.

Our Stack

  • Backend: Python, Ruby on Rails, Node.js.
  • Frontend: React, TypeScript.
  • Mobile: Swift (iOS), Kotlin (Android), React Native.
  • Infrastructure: AWS, Kubernetes, Terraform, SNS, SQS.
  • Databases: PostgreSQL (Aurora), Redis, OpenSearch.
  • Observability: Datadog, Splunk, Sentry.
  • ML: PyTorch, TensorFlow.
  • CI/CD: GitLab (on‑prem).

At Entrust, we don’t just offer jobs – we offer career journeys. Here is what you can expect when you join our team:

  • Career Growth: Whether you’re a budding developer or a seasoned expert, we’re invested in your professional journey. With learning‑forward initiatives and exciting challenges, your growth is our priority.
  • Flexibility: Life is all about balance. Whether you’re remote, hybrid, or on‑site, we offer flexible options that fit your lifestyle.
  • Collaboration: Here, your voice matters. Our teams thrive on sharing ideas, brainstorming solutions, and working together to build a better tomorrow.

We believe in securing identities—but it doesn’t stop there. At Entrust, we’re passionate about valuing all identities. Our culture is built on diversity, inclusion, and respect. From unconscious bias training for our leaders to global affinity groups that connect colleagues across the globe, we’re creating a community where everyone is encouraged to be themselves.

If you’re excited by the prospect of innovating, growing your career, and collaborating in a dynamic environment, Entrust is the place for you. Join us in making a difference. Let’s build a more secure world—together.

Principal Software Reliability Engineer, Product Engineering in London employer: Entrust

Entrust is an exceptional employer that prioritises career growth and flexibility, offering a dynamic work environment where innovation thrives. With a strong commitment to diversity and inclusion, employees are encouraged to share their ideas and collaborate on impactful projects, particularly in the high-stakes domain of identity verification and security solutions. Located in Minnesota, Entrust provides a unique opportunity to shape the future of reliability in product engineering while enjoying a supportive culture that values every individual's contributions.
Entrust

Contact Detail:

Entrust Recruiting Team

StudySmarter Expert Advice 🤫

We think this is how you could land Principal Software Reliability Engineer, Product Engineering in London

✨Tip Number 1

Network like a pro! Reach out to current employees at Entrust on LinkedIn. Ask them about their experiences and any tips they might have for landing the Principal Software Reliability Engineer role. Personal connections can give you insights that job descriptions just can't.

✨Tip Number 2

Prepare for the interview by diving deep into reliability practices. Brush up on your knowledge of resilience patterns and incident management. Being able to discuss these topics confidently will show that you're not just a fit for the role, but passionate about reliability as a discipline.

✨Tip Number 3

Showcase your hands-on experience! Be ready to share specific examples of how you've built reliability tooling or improved uptime in previous roles. This is your chance to demonstrate your technical depth and how you can contribute to Entrust's mission.

✨Tip Number 4

Don't forget to apply through our website! It’s the best way to ensure your application gets seen by the right people. Plus, it shows you're genuinely interested in joining the Entrust team. Let's make this happen!

We think you need these skills to ace Principal Software Reliability Engineer, Product Engineering in London

Software Engineering
Python
Ruby
Node.js
Reliability Tooling
Observability
Distributed Tracing
Incident Management
Blameless Postmortem Practices
Resilience Patterns
Communication Skills
Chaos Engineering
AWS
Kubernetes
GitLab CI/CD

Some tips for your application 🫡

Tailor Your Application: Make sure to customise your CV and cover letter for the Principal Software Reliability Engineer role. Highlight your experience in reliability and software engineering, and show us how your skills align with our needs at Entrust.

Showcase Your Passion: We want to see your enthusiasm for reliability as a discipline! Share examples of how you've driven reliability initiatives in past roles and how you plan to bring that passion to Entrust.

Be Clear and Concise: When writing your application, keep it straightforward. Use clear language and avoid jargon where possible. We appreciate directness and clarity, so make it easy for us to see your qualifications.

Apply Through Our Website: Don’t forget to submit your application through our website! It’s the best way for us to receive your details and ensures you’re considered for the role. We can’t wait to hear from you!

How to prepare for a job interview at Entrust

✨Understand the Role

Before your interview, make sure you fully grasp what a Principal Software Reliability Engineer does at Entrust. Familiarise yourself with their focus on application and service-level reliability, and how it differs from traditional SRE roles. This will help you tailor your responses to show that you’re the right fit for this unique position.

✨Showcase Your Technical Skills

Be prepared to discuss your experience with backend languages like Python, Ruby, or Node.js. Highlight any specific projects where you've built reliability tooling or implemented resilience patterns. This is your chance to demonstrate your hands-on experience and technical depth, so don’t hold back!

✨Prepare for Scenario Questions

Expect questions that assess your problem-solving skills in real-world scenarios. Think about past incidents you've managed, how you approached them, and what you learned. Be ready to discuss your incident management practices and how you ensure blameless postmortems.

✨Emphasise Collaboration and Communication

Since this role involves working closely with product engineers and influencing without authority, be ready to share examples of how you've successfully collaborated with cross-functional teams. Highlight your communication skills and how you can hold deep technical discussions with senior engineers.

Principal Software Reliability Engineer, Product Engineering in London
Entrust
Location: London

Land your dream job quicker with Premium

You’re marked as a top applicant with our partner companies
Individual CV and cover letter feedback including tailoring to specific job roles
Be among the first applications for new jobs with our AI application
1:1 support and career advice from our career coaches
Go Premium

Money-back if you don't land a job in 6-months

>