Job/Group Overview:
-
SRE within the Group Platform Services & Engineering division which provides the common services to Development, Infrastructure and Production Services.
-
This is an SRE/support position responsible for administering and supporting the Production environment as well as engineering reliability into the products / services we support
-
i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction.
-
A fantastic opportunity for somebody with 3+ years IT experience to work with state-of-the-art technologies to deliver industry leading solutions in the Telemetry,Observability and Monitoring space.
-
The successful candidate would join a team of enthusiastic, creative and forward-thinking SRE in the UK who are working in tandem with the engineers to radically transform how the Group manages the operation of its estate.
-
The position is within a global team consisting of 20 team members, across Engineering and SRE, bringing change across the organisation.
-
The candidate will work closely with their peers in other regions as well as other teams to facilitate the strategic objectives of the team.
-
The challenges we strive to solve include availability, scalability and performance related to delivering a platform used by the entire Group.
The observability platform consists of a combination of platforms and frameworks from in-house, vendors, and open source. These include:
-
Grafana LGTM stack (open source)
-
RightITNow, EverBridge, Sentinel (3rd Party tools)
-
AMBER, Bing, MCM, CMS (homegrown)
Responsibilities:
-
Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the group of companies.
-
Gain understanding of the various tools and frameworks that together provide observability and notification service to the organization and assist development and production support teams with queries / issues related to their usage of our platform.
-
Act as custodian of production environment and engage within the team and outside, if need be, towards building and maintaining robust, scalable, highly
-
available production systems in accordance with our service level objectives
-
Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime as well as
-
possibility of recurrence.
-
Pushing out changes and releases to production environment reliably via effective change and release management
-
Quick and effective response to alerts before they become incidents, with an approach to prevent them from occurring ever again
-
Effectively triaging alerts, requests, emails such that things that needs attention get addressed first and in a timely manner in the order of their priority, the drivers for which should be production stability and user satisfaction.
-
Continuous and effective engagement with users, with the required empathy,providing the right guidance so as to provide a good customer experience
-
Collaborate in a global agile team environment using established support practices, participating in sprint planning, reviews, and continuous improvement initiatives
-
Build and maintain scalable, reliable monitoring solutions that support global infrastructure
-
Engage with engineers, architect towards contributing to architectural decisions that influence the future direction of observability platform
-
Champion observability best practices across the organization, helping teams leverage data-driven insights to improve system reliability and performance
-
Partner with engineers as needed to optimize operational efficiency and enhance system resilience
-
Effectively leveraging AI tools such as Claude, CoPilot etc. with adequate guardrails to bring efficiencies into operational processes in a consistent, repeatable and risk averse manner.
-
Mentor and guide other SREs, sharing your knowledge and expertise across other team members for the benefit of the team.
Requirements (indicate mandatory and/or preferred):
Mandatory:
-
Minimum 2 yearsโ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.
-
At least 2 yearsโ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands
-
Exposure to one or more of following โ Python / Ansible
-
Production support experience โ Request handling, incident management,problem management, change management, release management, on-call
-
handling, user engagement, responding to alerts etc.
-
Good communication and interpersonal skills
-
Strong analytical and trouble-shooting skills, with the ability to exercise mature judgement
-
Basic understanding of cloud platforms
-
Basic understanding of CI/CD tools such as GitLab, Jenkins, Ansible, Nexus etc.
-
Good Team player
Preferred:
-
Understanding of Open Telemetry standards
-
Understanding of containerization technologies such as Kubernetes, EKS, Docker etc.
-
Supporting a medium / large scale production environment
-
Exposure to AI tools and their usage for increasing work efficiency
-
Knowledge of ITIL
-
Decent understanding of DB Platforms โ Sybase / MySQL / MSSQL โ general RDBMS concepts, SQL
-
Collaboration Tools โ Confluence / JIRA
-
Basic knowledge of / familiarity with other infrastructure technologies such as Middleware (ActiveMQ / Solace / EMS / Tibco etc.), Web servers, Load
-
balancers, Directory Services etc.
-
Experience working with a globally dispersed team
Observability SRE in London employer: HCL Europe
As a leading technology firm, we pride ourselves on fostering a collaborative and innovative work culture that empowers our employees to excel in their roles. Located in a vibrant tech hub, we offer competitive benefits, continuous learning opportunities, and a commitment to professional growth, making us an ideal employer for those seeking to make a meaningful impact in the field of cloud-native application architecture.