Observability SRE in London

Observability SRE in London

London Full-Time On-site
H

Job/Group Overview:

  • SRE within the Group Platform Services & Engineering division which provides the common services to Development, Infrastructure and Production Services.

  • This is an SRE/support position responsible for administering and supporting the Production environment as well as engineering reliability into the products / services we support

  • i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction.

  • A fantastic opportunity for somebody with 3+ years IT experience to work with state-of-the-art technologies to deliver industry leading solutions in the Telemetry,Observability and Monitoring space.

  • The successful candidate would join a team of enthusiastic, creative and forward-thinking SRE in the UK who are working in tandem with the engineers to radically transform how the Group manages the operation of its estate.

  • The position is within a global team consisting of 20 team members, across Engineering and SRE, bringing change across the organisation.

  • The candidate will work closely with their peers in other regions as well as other teams to facilitate the strategic objectives of the team.

  • The challenges we strive to solve include availability, scalability and performance related to delivering a platform used by the entire Group.

The observability platform consists of a combination of platforms and frameworks from in-house, vendors, and open source. These include:

  • Grafana LGTM stack (open source)

  • RightITNow, EverBridge, Sentinel (3rd Party tools)

  • AMBER, Bing, MCM, CMS (homegrown)

Responsibilities:

  • Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the group of companies.

  • Gain understanding of the various tools and frameworks that together provide observability and notification service to the organization and assist development and production support teams with queries / issues related to their usage of our platform.

  • Act as custodian of production environment and engage within the team and outside, if need be, towards building and maintaining robust, scalable, highly

  • available production systems in accordance with our service level objectives

  • Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime as well as

  • possibility of recurrence.

  • Pushing out changes and releases to production environment reliably via effective change and release management

  • Quick and effective response to alerts before they become incidents, with an approach to prevent them from occurring ever again

  • Effectively triaging alerts, requests, emails such that things that needs attention get addressed first and in a timely manner in the order of their priority, the drivers for which should be production stability and user satisfaction.

  • Continuous and effective engagement with users, with the required empathy,providing the right guidance so as to provide a good customer experience

  • Collaborate in a global agile team environment using established support practices, participating in sprint planning, reviews, and continuous improvement initiatives

  • Build and maintain scalable, reliable monitoring solutions that support global infrastructure

  • Engage with engineers, architect towards contributing to architectural decisions that influence the future direction of observability platform

  • Champion observability best practices across the organization, helping teams leverage data-driven insights to improve system reliability and performance

  • Partner with engineers as needed to optimize operational efficiency and enhance system resilience

  • Effectively leveraging AI tools such as Claude, CoPilot etc. with adequate guardrails to bring efficiencies into operational processes in a consistent, repeatable and risk averse manner.

  • Mentor and guide other SREs, sharing your knowledge and expertise across other team members for the benefit of the team.

Requirements (indicate mandatory and/or preferred):
Mandatory:

  • Minimum 2 yearsโ€™ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.

  • At least 2 yearsโ€™ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands

  • Exposure to one or more of following โ€“ Python / Ansible

  • Production support experience โ€“ Request handling, incident management,problem management, change management, release management, on-call

  • handling, user engagement, responding to alerts etc.

  • Good communication and interpersonal skills

  • Strong analytical and trouble-shooting skills, with the ability to exercise mature judgement

  • Basic understanding of cloud platforms

  • Basic understanding of CI/CD tools such as GitLab, Jenkins, Ansible, Nexus etc.

  • Good Team player

Preferred:

  • Understanding of Open Telemetry standards

  • Understanding of containerization technologies such as Kubernetes, EKS, Docker etc.

  • Supporting a medium / large scale production environment

  • Exposure to AI tools and their usage for increasing work efficiency

  • Knowledge of ITIL

  • Decent understanding of DB Platforms โ€“ Sybase / MySQL / MSSQL โ€“ general RDBMS concepts, SQL

  • Collaboration Tools โ€“ Confluence / JIRA

  • Basic knowledge of / familiarity with other infrastructure technologies such as Middleware (ActiveMQ / Solace / EMS / Tibco etc.), Web servers, Load

  • balancers, Directory Services etc.

  • Experience working with a globally dispersed team

Observability SRE in London employer: HCL Europe

As a leading technology firm, we pride ourselves on fostering a collaborative and innovative work culture that empowers our employees to excel in their roles. Located in a vibrant tech hub, we offer competitive benefits, continuous learning opportunities, and a commitment to professional growth, making us an ideal employer for those seeking to make a meaningful impact in the field of cloud-native application architecture.

H

Contact Details:

HCL Europe Recruitment Team

ยฉ StudySmarter