Site Reliability Engineer - SRE Fleet

Site Reliability Engineer - SRE Fleet

Full-Time 60000 - 80000 Β£ / year (est.) No working from home possible
Cisco Systems Inc

At a Glance

  • Tasks: Develop automation solutions to enhance infrastructure reliability and efficiency across global cloud environments.
  • Company: Join Cisco, a leader in innovative technology and cloud solutions.
  • Benefits: Competitive salary, flexible work options, and endless growth opportunities.
  • Other info: Collaborative team culture with a focus on innovation and operational excellence.
  • Why this job: Make a real impact on global cloud infrastructure while working with cutting-edge technology.
  • Qualifications: 2+ years in SRE or related fields, experience with automation tools and cloud environments.

The predicted salary is between 60000 - 80000 Β£ per year.

Meet the Team

The SRE Fleet team is responsible for maintaining the stability, scalability, and efficiency of the infrastructure that powers our global cloud platform.

As a team of six engineers distributed across the US, Canada, and the UK, we combine deep infrastructure expertise with a strong focus on automation, reliability, and operational excellence.

We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide.

The team operates with a high degree of autonomy, giving engineers the opportunity to drive both critical initiatives and grassroots improvements that solve real operational challenges.

One of our most exciting areas of focus is expanding our ability to build highly automated regional, sovereign, and isolated cloud environments that support new markets and evolving regulatory requirements.

Everyone on the team has a voice, and engineers are encouraged to identify problems, propose solutions, and take ownership of improvements that make the platform more reliable and easier to operate.

Your Impact

Develop and maintain automation solutions that improve the reliability, scalability, and operational efficiency of infrastructure spanning more than 2,000 machines across global cloud environments.

Design and enhance deployment pipelines, testing frameworks, and operational tooling to support the continued growth of a platform serving millions of managed devices worldwide.

Troubleshoot complex infrastructure and distributed systems issues to ensure high availability while helping teams identify and address performance and scalability challenges.

Contribute to critical projects such as cluster build out by building automation that enables the rapid and repeatable deployment of new sovereign, regional, and purpose-built cloud environments.

Partner with other engineering teams, product management, and business partners across multiple teams and time zones to understand platform dependencies, seek opportunities for improvement, and deliver solutions that enhance reliability and reduce operational overhead.

  • Minimum Qualifications
  • 2+ years of experience in Site Reliability Engineering, Dev Ops, Infrastructure Engineering, or a related role supporting cloud-based production environments.
  • Experience developing and maintaining infrastructure automation using Ansible.
  • Experience programming in Ruby and developing automated tests using RSpec or comparable testing frameworks.
  • Experience administering and troubleshooting Linux-based systems and distributed infrastructure environments.
  • Experience designing, implementing, and maintaining CI/CD pipelines, including Git Lab CI.
  • Experience supporting large-scale infrastructure environments consisting of hundreds or thousands of systems.
  • Preferred Qualifications
  • Familiarity with AWS or other public cloud platforms and hybrid infrastructure environments.
  • Knowledge of monitoring, observability, and reliability engineering practices and tooling.
  • Familiarity with Kubernetes concepts and containerized application platforms.
  • Experience leveraging AI-assisted development tools to improve software development, automation, operational analysis, and engineering productivity.

Why Cisco?

At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era - and beyond.

We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds.

These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions.

Add to that our worldwide network of doers and experts, and you'll see that the opportunities to grow and build are limitless.

We work as a team, collaborating with empathy to make really big things happen on a global scale.

Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

#J-18808-Ljbffr

Site Reliability Engineer - SRE Fleet employer: Cisco Systems Inc

Cisco Systems, Inc. is an exceptional employer that champions innovation and inclusivity within the Higher Education sector. With a strong focus on personal growth and development, employees are encouraged to collaborate across teams and drive impactful strategies that align with market needs. The vibrant work culture and commitment to operational excellence make Cisco a rewarding place to build a meaningful career.

Cisco Systems Inc

Contact Details:

Cisco Systems Inc Recruitment Team

We think you need these skills to ace Site Reliability Engineer - SRE Fleet

Site Reliability Engineering
DevOps
Infrastructure Engineering
Cloud-based Production Environments
Infrastructure Automation using Ansible
Programming in Ruby
Automated Testing using RSpec