Site Reliability Engineer

Site Reliability Engineer

Full-Time 63000 - 77000 £ / year (est.) No working from home possible
A

At a Glance

  • Tasks: Ensure reliability and performance of critical trading systems while collaborating with diverse teams.
  • Company: High-performing trading technology firm focused on operational excellence.
  • Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
  • Other info: Join a dynamic team dedicated to continuous improvement and cutting-edge technology.
  • Why this job: Make a real impact on trading outcomes in a fast-paced, innovative environment.
  • Qualifications: Experience with Linux, distributed systems, and proficiency in systems-level programming.

The predicted salary is between 63000 - 77000 £ per year.

A high-performing trading technology firm is seeking an SRE Engineer to drive reliability, scalability, and operational excellence across critical production systems. This role is suited to a hands-on engineer who operates across infrastructure, software, and platform environments, with a strong focus on resilience, automation, and engineering quality. You will work closely with software engineering, platform, security, and infrastructure teams to improve service availability, strengthen observability, and enhance operational practices in a high-performance, low-latency environment.

Responsibilities

  • Own the reliability, performance, and availability of business-critical production systems and infrastructure
  • Support incident response, service restoration, root cause analysis, and post-incident reviews
  • Implement SRE best practices including monitoring, alerting, and service health standards
  • Build and improve automation to reduce operational toil and enhance deployment consistency and recovery
  • Partner with engineering teams to improve system design, fault tolerance, and capacity planning
  • Contribute to production readiness for new services and infrastructure changes
  • Maintain and improve observability across systems (metrics, logging, tracing)
  • Continuously improve operational processes and platform reliability

Requirements

  • Strong hands-on experience operating and troubleshooting Linux-based production environments
  • Solid understanding of distributed systems and streaming architectures
  • Experience with messaging platforms
  • Proficiency in at least one systems-level language (C, C++, Rust, or similar)
  • Familiarity with large-scale data systems
  • Experience with CI/CD pipelines, build systems, and modern development workflows
  • Working knowledge of containerised environments
  • Proven ability to troubleshoot and resolve complex production issues in high-availability systems

Preferred Profile

  • Background in a high-scale or performance-sensitive engineering environment (e.g. trading, FAANG, or similar)
  • Experience supporting low-latency or highly distributed infrastructure
  • Track record of improving automation, observability, and system reliability
  • Strong production mindset with a focus on stability, performance, and continuous improvement

This is an opportunity to work on critical, real-time systems in a high-performance environment, with direct impact on production reliability and trading outcomes.

Site Reliability Engineer employer: Autonomai Recruitment

Join a leading HFT-style firm in London that prioritises innovation and excellence in network engineering. With a vibrant work culture that fosters collaboration and continuous learning, employees are encouraged to grow their skills in a fast-paced environment. Enjoy competitive benefits and the unique opportunity to work on cutting-edge technologies that drive the financial sector.

A

Contact Details:

Autonomai Recruitment Recruitment Team

We think you need these skills to ace Site Reliability Engineer

Reliability Engineering
Scalability
Operational Excellence
Incident Response
Root Cause Analysis
Monitoring and Alerting
Automation