As a preferred supplier to one of our biggest Clients, I am seeking for a SRE Engineer (Datadog) for a position in York (UK).
All candidates should make sure to read the following job description and information carefully before applying.
Key Responsibilities:
Experience: 10+ years of hands-on experience in Site Reliability Engineering (SRE), DevOps, or Systems Architecture, with at least 3+ years specializing deeply in Datadog administration and configuration.
Cloud & Container Expertise: Deep professional experience working with AWS, Azure, or GCP, paired with heavy production experience managing Kubernetes clusters.
Instrumentation & Coding: Proficiency in systems or scripting languages (e.g., Python, Go, Bash, or JavaScript) and experience instrumenting applications for APM.
Datadog Mastery: Deep understanding of Datadog's core pillars-Infrastructure, APM, Logs, Metrics, Synthetics, and Security Monitoring. Datadog Certifications are a strong plus.
Problem-Solving Mindset: Demonstrated ability to debug complex, distributed microservices architectures under high-pressure incident response scenarios.
Communication Skills: Excellent interpersonal and stakeholder management skills, with the ability to translate technical telemetry data into actionable business and engineering insights.
Essential skills/knowledge/experience:
Architecture & Implementation
Platform Ownership: Design, deploy, and manage Datadog agents, integrations, and custom metrics across multi-cloud (AWS/Azure/GCP) and containerized (Kubernetes, Docker) environments.
Observability Pipelines: Architect and scale high-throughput log processing, routing, and transformation systems using Datadog.
APM & Infrastructure Monitoring: Configure and optimize Application Performance Monitoring (APM), Distributed Tracing, Real User Monitoring (RUM), and Infrastructure metrics.
Governance & Best Practices
Standardization: Establish company-wide standards for dashboards, monitors, SLOs/SLIs, and alert routing (integrating with PagerDuty, Jira, Opsgenie, etc.).
Cost & Performance Optimization: Audit and optimize Datadog usage, index management, log retention policies, and custom metric volume to maximize ROI and control licensing costs.
Security & Compliance: Leverage Datadog Security products (CSPM, CWPP, Cloud SIEM, Container Security) to maintain compliance postures and mitigate runtime threats.
Collaboration & Enablement
Cross-Functional Mentorship: Act as the go-to escalation point and technical mentor for DevOps, SRE, and Software Engineering teams regarding troubleshooting and instrumentation. xrnqpay
Training & Documentation: Create internal documentation, runbooks, and training modules to elevate organizational proficiency in observability.
Vendor Management: Act as the primary technical point of contact for Datadog account teams,
Contract: 6 months+
Rates: Excellent
Location: York
Interview: 2 stages, 1 technical + 1 assessment
SRE Engineer in London employer: Confidential
As a leading employer in the data centre operations sector, we offer an exceptional work environment that prioritises employee growth and development. Our collaborative culture fosters innovation and accountability, while our commitment to safety and operational excellence ensures that you will be part of a team that values your contributions. With opportunities for international travel and the chance to lead critical operations across the EMEA region, this role provides a unique platform for impactful leadership and career advancement.