Principal Site Reliability Engineer

Principal Site Reliability Engineer

Full-Time 80000 - 100000 £ / year (est.) No working from home possible
Jobgether

At a Glance

  • Tasks: Shape reliability engineering practices and drive scalable automation in a tech-driven environment.
  • Company: Join a fast-growing partner company at the intersection of finance and digital assets.
  • Benefits: Enjoy fully remote work, 35 days off, private health insurance, and career growth opportunities.
  • Other info: Inclusive culture with employee-led communities and exposure to international teams.
  • Why this job: Make a real impact on innovative projects while mentoring others in a collaborative setting.
  • Qualifications: Deep expertise in distributed systems and strong problem-solving skills required.

The predicted salary is between 80000 - 100000 £ per year.

This position is listed on behalf of a partner company, who manages all applications and next steps.

Our partner is looking for a Principal Site Reliability Engineer based in United Kingdom.

This is an exceptional opportunity to shape reliability engineering practices within a fast-growing, technology-driven environment operating at the intersection of finance and digital assets.

You will play a strategic role in defining how reliability, observability, and operational excellence are embedded across engineering teams.

The position combines hands‑on technical leadership with broad organizational influence, enabling you to drive scalable automation and resilient infrastructure practices.

Working closely with engineering, product, and platform teams, you will help build highly available systems that support mission‑critical services.

This role is ideal for an experienced engineer who enjoys solving complex distributed systems challenges while mentoring others and influencing engineering culture in a collaborative, globally distributed setting.

Accountabilities Define and promote Site Reliability Engineering principles, establishing frameworks for reliability, observability, service level indicators (SLIs), service level objectives (SLOs), and error budgets.

Drive the adoption of operational excellence practices and ensure reliability metrics are measurable and continuously improved.

Design and implement automation solutions that enhance system scalability, reliability, and deployment efficiency.

Conduct production readiness assessments and provide architectural guidance to engineering teams to ensure services are built for scale and resilience.

Lead initiatives to improve the lifecycle of distributed systems and microservices, from development and deployment to monitoring and optimization.

Identify performance bottlenecks, capacity challenges, and operational risks while implementing sustainable solutions.

Partner with engineering and product leadership to embed reliability considerations into product development processes.

Lead incident management improvements through blameless postmortems, root cause analysis, and systemic remediation initiatives.

Mentor engineers and advocate for best practices in reliability engineering, fostering ownership and operational maturity across teams.

Contribute to the future vision and strategic direction of the Site Reliability Engineering function.

Requirements The ideal candidate brings deep expertise in distributed systems, reliability engineering, and organizational leadership, combined with a passion for building scalable and resilient platforms.

Proven experience designing, operating, and troubleshooting distributed systems and microservices architectures.

Strong expertise in observability, monitoring strategies, incident management, and operational excellence frameworks.

Demonstrated ability to drive organizational change and influence engineering practices across multiple teams.

Extensive experience implementing reliability frameworks, including SLIs, SLOs, error budgets, and production readiness processes.

Strong problem‑solving skills with a structured and analytical approach to complex technical challenges.

Excellent communication and stakeholder management abilities, with experience collaborating across engineering and leadership teams.

Experience working with cloud environments, particularly AWS, is highly desirable.

Previous exposure to financial services, regulated industries, or mission‑critical platforms is considered an advantage.

Interest in blockchain technologies, digital assets, or decentralized finance ecosystems is a plus.

Master's degree in Computer Science, Engineering, or a related field is advantageous.

Benefits Fully remote opportunity within a globally distributed and collaborative environment. 35 days of paid time off annually, including public holidays.

Additional annual leave entitlement based on years of service.

Private health insurance coverage.

Opportunity to work on innovative technologies and large‑scale, high‑impact infrastructure projects.

Strong emphasis on learning, career progression, and professional growth.

Inclusive and diverse workplace culture with employee‑led communities and wellbeing initiatives.

Exposure to international teams and cross‑functional collaboration across multiple regions. #J-18808-Ljbffr

Principal Site Reliability Engineer employer: Jobgether

At Jobgether, we pride ourselves on being an exceptional employer that fosters a collaborative and innovative work culture. Our remote working environment allows for flexibility while providing ample opportunities for professional growth and development in the tech industry. Join us to make a meaningful impact in enhancing open-source technology adoption, all while enjoying the benefits of a supportive team and a commitment to your career advancement.

Jobgether

Contact Details:

Jobgether Recruitment Team

We think you need these skills to ace Principal Site Reliability Engineer

Site Reliability Engineering principles
Reliability frameworks (SLIs, SLOs, error budgets)
Distributed systems architecture
Microservices architecture
Observability and monitoring strategies
Incident management
Operational excellence practices