AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM in London

AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM in London

London Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
Scope AT Limited

At a Glance

  • Tasks: Develop and implement SRE methodologies for cloud environments, focusing on automation and operational excellence.
  • Company: Join a leading financial services firm with a focus on innovation and collaboration.
  • Benefits: Enjoy hybrid working, competitive salary, and opportunities for professional growth.
  • Other info: Be part of a team that values continuous improvement and offers excellent career advancement.
  • Why this job: Make a real impact by optimising infrastructure and enhancing system observability in a dynamic environment.
  • Qualifications: Experience in SRE methodologies, scripting (Python/Ansible), and cloud platforms like AWS.

The predicted salary is between 63000 - 77000 Β£ per year.

The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team.

Responsibilities include:

  • Driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Driving continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs.
  • Defining and enhancing frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues.
  • Developing secure high-quality production code, and reviewing and debugging code written by others.
  • Building out and enhancing GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform.
  • Providing on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
  • Ensuring risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.

Minimum Job-Related Experience Required:

  • Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause.
  • Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
  • Strong knowledge of at least 1 scripting language, preferably either Python or Ansible. PowerShell would also be a positive.
  • Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies.
  • Experience of Observability/APM tools (e.g. Grafana/Datadog/Dynatrace).

Permanent Role based in Canary Wharf - Hybrid Working.

AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM in London employer: Scope AT Limited

Scope AT Limited is an exceptional employer, offering a dynamic work environment in the heart of Glasgow's financial district. With a strong focus on employee development and a collaborative culture, team members are encouraged to grow their skills while working on innovative projects in equity swaps. The company also provides competitive benefits and a supportive atmosphere that values work-life balance, making it an ideal place for those seeking meaningful and rewarding careers in the financial services sector.

Scope AT Limited

Contact Details:

Scope AT Limited Recruitment Team

We think you need these skills to ace AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM in London

SRE Methodologies
Cloud Infrastructure Management
Python
PowerShell
AWS
GitOps
Terraform