Engineer II, Site Reliability (Hybrid, London)

Engineer II, Site Reliability (Hybrid, London)

Full-Time 63000 - 77000 £ / year (est.) Home office (partial)
C

At a Glance

  • Tasks: Join our TechOps SRE team to ensure our platform runs flawlessly 24/7.
  • Company: CrowdStrike, a global leader in cybersecurity with a mission-driven culture.
  • Benefits: Competitive salary, wellness programs, professional development, and vibrant office culture.
  • Other info: Diverse team environment with excellent career growth opportunities.
  • Why this job: Make a real impact in cybersecurity while working with cutting-edge technology.
  • Qualifications: Experience in software engineering and large-scale production environments required.

The predicted salary is between 63000 - 77000 £ per year.

As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward.

CrowdStrike is looking to hire an Engineer II to the TechOps SRE team that will have a focus on our Commercial Cloud. We’re looking for a deeply-technical, hands-on engineer, who loves to develop automation and tooling through software to ensure delivery of mission critical solutions and services for large-scale distributed systems.

What You'll Do:

  • Have expertise with Linux engineering and administration for thousands of bare metal servers and virtual machines
  • Be responsible for all operational aspects of our platform - Availability, Latency, Throughput, Monitoring, Issue Response (analysis, remediation, deployment) and Capacity Planning with respect to Latency and Throughput
  • Work in a team of highly motivated engineers distributed across the globe
  • On-call rotation with other team members
  • Troubleshoot server hardware issues
  • Use your passion for technology to ensure our platform operates flawlessly 24x7
  • Obsess about learning, and champion the newest technologies & tricks with others, raising the technical IQ of the team
  • Have broad exposure to our entire architecture and become one of our experts in our overall process flow
  • Have an intrinsic drive to make things better
  • Bias towards small development projects and the occasional larger projects
  • Have experience with modern monitoring and telemetry stacks (ELK, Prometheus, Grafana, Zabbix)
  • Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding
  • Ability to lead incident analysis for incidents, champion incident response practices and assist in correlating incidents to systemic problems, and drive towards resolution

What You'll Need:

  • Bachelor's degree and/or equivalent experience in Computer Science
  • A minimum of five years of experience working in a large scale production environment
  • A minimum of two years of experience in software engineering
  • A minimum of two years of experience in one or more of: C++, Java, Python, Go
  • Experience with storage technologies (Examples: SAN, NAS, NFS, Object Storage, FreeNAS, iSCSI)
  • Experience with Infrastructure technologies (Examples: Linux, Windows, VMware, Docker, Kubernetes, etc.)
  • Experience writing technical documentation
  • Configuration management experience with one or more tools such as Puppet, Chef, Ansible
  • Solid understanding of application design, including operational trade-offs of various designs
  • Analytical skills coupled with a strong sense of urgency, ownership, and drive
  • Ability to work well in a diverse, team-focused environment with other SREs and Engineers
  • Ability to broadly communicate and present recommended conventions defined by the reliability team
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes

CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed.

Engineer II, Site Reliability (Hybrid, London) employer: CrowdStrike

CrowdStrike is an exceptional employer that champions a culture of innovation and flexibility, allowing employees to thrive in a mission-driven environment focused on cybersecurity. With comprehensive wellness programs, competitive benefits, and ample professional development opportunities, employees are empowered to grow their careers while contributing to a meaningful cause. The remote work model, complemented by occasional in-person support at the Reading office, ensures a balanced work-life dynamic, making CrowdStrike a top choice for those seeking impactful employment in the tech industry.

C

Contact Details:

CrowdStrike Recruitment Team

We think you need these skills to ace Engineer II, Site Reliability (Hybrid, London)

Linux Engineering
Virtual Machine Administration
Operational Management
Monitoring and Telemetry Stacks (ELK, Prometheus, Grafana, Zabbix)
Incident Response
Software Engineering (C++, Java, Python, Go)
Storage Technologies (SAN, NAS, NFS, Object Storage, FreeNAS, iSCSI)