Senior Site Reliability Engineer, Production Engineering
Senior Site Reliability Engineer, Production Engineering

Senior Site Reliability Engineer, Production Engineering

London Full-Time 48000 - 84000 £ / year (est.) No home office possible
T

At a Glance

  • Tasks: Design and manage large-scale cloud systems, optimising for reliability and performance.
  • Company: Join Cisco ThousandEyes, a leader in Digital Experience Assurance, enhancing seamless digital experiences.
  • Benefits: Enjoy a hybrid work model, with flexibility to work from home and great corporate perks.
  • Why this job: Be part of a dynamic team improving digital experiences while leveraging cutting-edge technology.
  • Qualifications: Expertise in Kubernetes, cloud services, and software development required; 5+ years preferred.
  • Other info: Diverse backgrounds are encouraged; apply even if you don't meet every qualification.

The predicted salary is between 48000 - 84000 £ per year.

Please note that we have a hybrid approach to work and would like to find someone who can come into our offices in London at least one day a week.

Who We Are

Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences. ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.

About The Role

We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.

What You’ll Do

  • Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.

Qualifications

  • Expert-level knowledge of Kubernetes and its ecosystem.
  • Proficiency in software development with languages such as Python or Go.
  • In-depth knowledge of cloud providers, preferably AWS.
  • Proven ability to build and implement scalable and well-tested solutions.
  • Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols.
  • Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs.

Preferred Qualifications

  • Familiarity with best practices for operating a large-scale, highly available enterprise platform.
  • 5+ years of experience in a related role.
  • Excellent communication and documentation skills.
  • Strong sense of ownership, drive, and attention to detail.

Cisco values the perspectives and skills that emerge from employees with diverse backgrounds. That’s why Cisco is expanding the boundaries of discovering top talent by not only focusing on candidates with educational degrees and experience but also placing more emphasis on unlocking potential. We believe that everyone has something to offer and that diverse teams are better equipped to solve problems, innovate, and create a positive impact. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification. Research shows that people from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy. We urge you not to prematurely exclude yourself and to apply if you’re interested in this work.

Senior Site Reliability Engineer, Production Engineering employer: ThousandEyes (part of Cisco)

At Cisco ThousandEyes, we pride ourselves on being an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration. Our London office provides a vibrant environment where employees can thrive, with opportunities for professional growth and development in cutting-edge technologies. With a strong emphasis on diversity and inclusion, we empower our team members to contribute their unique perspectives, ensuring a rewarding and meaningful career journey.
T

Contact Detail:

ThousandEyes (part of Cisco) Recruiting Team

StudySmarter Expert Advice 🤫

We think this is how you could land Senior Site Reliability Engineer, Production Engineering

✨Tip Number 1

Familiarise yourself with the specific tools and technologies mentioned in the job description, such as Kubernetes, AWS, and Prometheus. Having hands-on experience or projects that showcase your skills with these tools can set you apart from other candidates.

✨Tip Number 2

Engage with the community around Site Reliability Engineering. Join forums, attend meetups, or participate in online discussions to stay updated on best practices and trends. This not only enhances your knowledge but also helps you network with professionals in the field.

✨Tip Number 3

Prepare to discuss real-world scenarios where you've implemented scalable solutions or improved system reliability. Be ready to share specific examples during interviews that demonstrate your problem-solving skills and technical expertise.

✨Tip Number 4

Showcase your soft skills, particularly communication and teamwork. As a Senior SRE, you'll be collaborating with various teams, so highlighting your ability to work well with others and communicate complex ideas clearly can make a significant difference.

We think you need these skills to ace Senior Site Reliability Engineer, Production Engineering

Kubernetes Expertise
Cloud Computing (AWS)
Python or Go Programming
Distributed Systems Design
Site Reliability Principles
Incident Response Management
Change Management
Deployment Strategies
Service Level Objectives (SLOs)
Unix/Linux Systems Knowledge
Automation and Scripting
Monitoring Tools (Prometheus, OpenTelemetry)
Scalability Best Practices
Strong Communication Skills
Attention to Detail

Some tips for your application 🫡

Tailor Your CV: Make sure your CV highlights your experience with Kubernetes, AWS, and any relevant programming languages like Python or Go. Emphasise your expertise in Site Reliability principles and any previous roles that align with the responsibilities outlined in the job description.

Craft a Compelling Cover Letter: In your cover letter, express your passion for digital experience assurance and how your skills can contribute to the ThousandEyes platform. Mention specific projects or achievements that demonstrate your ability to design and manage large-scale distributed systems.

Showcase Relevant Experience: When filling out your application, ensure you detail your experience in SaaS operations and any past roles where you collaborated with software engineers. Highlight your contributions to incident response and automation solutions, as these are key aspects of the role.

Be Authentic: Cisco values diverse backgrounds and perspectives. Don’t hesitate to share your unique experiences and how they have shaped your approach to problem-solving and innovation. This will help you stand out as a candidate who brings something special to the team.

How to prepare for a job interview at ThousandEyes (part of Cisco)

✨Showcase Your Technical Expertise

Be prepared to discuss your experience with Kubernetes, AWS, and other cloud-native tools. Highlight specific projects where you've optimised architecture for reliability and performance, as this role heavily relies on these skills.

✨Demonstrate Problem-Solving Skills

Expect to face scenario-based questions that assess your ability to handle incidents and improve system reliability. Share examples of how you've identified and resolved operational obstacles in previous roles.

✨Emphasise Collaboration

This position requires working closely with software engineers. Be ready to discuss how you've collaborated with development teams in the past to enhance platform reliability and performance.

✨Stay Updated on Industry Trends

Familiarise yourself with the latest best practices in Site Reliability Engineering and cloud operations. Showing that you are proactive about learning can set you apart from other candidates.

Senior Site Reliability Engineer, Production Engineering
ThousandEyes (part of Cisco)
T
  • Senior Site Reliability Engineer, Production Engineering

    London
    Full-Time
    48000 - 84000 £ / year (est.)

    Application deadline: 2027-05-25

  • T

    ThousandEyes (part of Cisco)

Similar positions in other companies
UK’s top job board for Gen Z
discover-jobs-cta
Discover now
>