At a Glance
- Tasks: Develop and implement SRE methodologies for cloud environments, focusing on automation and operational excellence.
- Company: Join a leading financial services firm with a focus on innovation and collaboration.
- Benefits: Enjoy hybrid working, competitive salary, and opportunities for professional growth.
- Other info: Be part of a team that values continuous improvement and offers excellent career advancement.
- Why this job: Make a real impact by optimising infrastructure and enhancing system observability in a dynamic environment.
- Qualifications: Experience in SRE methodologies, scripting (Python/Ansible), and cloud platforms like AWS.
The predicted salary is between 63000 - 77000 Β£ per year.
The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team.
Responsibilities include:
- Driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
- Driving continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs.
- Defining and enhancing frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues.
- Developing secure high-quality production code, and reviewing and debugging code written by others.
- Building out and enhancing GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform.
- Providing on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
- Ensuring risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.
Minimum Job-Related Experience Required:
- Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause.
- Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
- Strong knowledge of at least 1 scripting language, preferably either Python or Ansible. PowerShell would also be a positive.
- Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies.
- Experience of Observability/APM tools (e.g. Grafana/Datadog/Dynatrace).
Permanent Role based in Canary Wharf - Hybrid Working.
AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM in London employer: Scope AT Limited
Scope AT Limited is an exceptional employer, offering a dynamic work environment in the heart of Glasgow's financial district. With a strong focus on employee development and a collaborative culture, team members are encouraged to grow their skills while working on innovative projects in equity swaps. The company also provides competitive benefits and a supportive atmosphere that values work-life balance, making it an ideal place for those seeking meaningful and rewarding careers in the financial services sector.