As part of the Cloudflare Infrastructure Engineering organization, our platform SREs are primarily responsible for production reliability
SREs are based in locations in Asia, Europe and the US enabling follow the sun coverage during daytime hours
SREs are supported by all engineering teams at Cloudflare who participate in on call schedules for their services
The SRE teams facilitate remediation and follow up of production issues and mature the tooling to enable all engineering teams to self-service on production
Incident follow up work across all engineering teams is prioritized above product innovation and the impact of production incidents influences the priority
SREs support two main environments: Edge SRE are focused on edge distribution where most client traffic is served
Core SRE are focused on the core services like control plane, data pipeline and other supporting supporting services
What you'll do:
We are looking for an Engineering Manager to lead the Edge SRE team in London
You will lead and develop a team of SREs that are responsible for Cloudflare edge production and building the tools for all teams to understand and interact with it
You will play a lead role in driving our Platform initiatives for edge services and will be tasked with leading engineers who build tools and best practices for engineering teams to debug in production, measure availability and performance indicators, track and report on thresholds
Lead a team of engineers who are working to keep the Cloudflare edge reliable and scalable
Mentor, grow, and empower your team by giving them the skills, confidence and motivation to make decisions
Help the individuals on your team to build and execute personal development plans that align with Cloudflare's goals and objectives
Take an active role in prioritizing the roadmap for the SRE Org
Drive cross-team and cross-org alignment in engineering, infrastructure and product teams
Partner with other Engineering Managers across Cloudflare to achieve reliability outcomes for their services
Participate in deep technical design discussions within your team, and across partner teams, and ensure that we're building the right systems and keeping the quality high
Benefits
Competitive pay: We offer a competitive total rewards package, where every employee is an owner of our stock.
Take-what-you-need vacation: We encourage employees to find a comfortable work-life balance by taking as many days off as they need while still being able to perform their jobs satisfactorily. (We really mean it!)
Paid maternity & paternity leave: Our global parental leave policy allows up to sixteen paid weeks of bonding leave time for all qualifying new parents.
Employee benefits: We offer a comprehensive benefits package including healthcare, life insurance, short- and long-term disability, pension plans in accordance with the market practice in our locations.
Commuter program: We are flexible about where we work and value connecting at the office. Our commuter benefits program is in place to support team members' commute to work via public transportation without the extra cost burden.
Wellbeing: Global offerings that provide mental health, childcare and family forming support to our employees across our offices.
you have 2+ years experience managing a team of 5 or more engineers on projects in the areas of: distributed systems, tooling, Linux, Internetworking, infrastructure security or infrastructure managementYou have 5+ years of software engineering, reliability, or operations experience in a customer-focused environmentYou are comfortable collaborating and co-ordinating on cross-team projects and workflowsYou can provide a strong technical vision for systems and infrastructure teamsYou are capable of leading a discussion with upper management, and are able to tailor the level of technical detail to suit your audienceYou have experience building services and systems, have successfully taken projects from inception to production, and are comfortable diving in to provide leadership for major projects when neededExperience running and maturing distributed systemsExperience using observability tools such as Jaeger, OpenTracing, ELK, Prometheus, Thanos, Grafana, ClickhouseExcel at planning and overseeing execution to meet commitments and deliver with predictabilityExperience leading and hiring a team that builds and runs tools and platformsExperience developing tools and APIsIncident root cause analysis and follow-upsHands-on experience with software or reliability engineeringIncident managementFamiliarity working with Proxies, DNS, Databases, Internet and SecurityComfortable managing teams/projections with deadlines and short release cycles
#J-18808-Ljbffr
Engineering Manager (Edge SRE) in London employer: Cloudflare
Cloudflare is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration among talented engineers. With a strong emphasis on employee growth, you will have access to cutting-edge technologies and the opportunity to contribute to a globally impactful distributed database system, all while enjoying the benefits of a supportive and inclusive environment in a fast-paced industry.