Site Reliability Engineer in Brighton

Site Reliability Engineer in Brighton

Brighton Full-Time 70000 - 90000 £ / year (est.) No working from home possible
Falcon Smart IT (FalconSmartIT)

At a Glance

  • Tasks: Drive IT operations modernization and implement observability practices to enhance system reliability.
  • Company: Join a forward-thinking tech company in Hove with a hybrid work culture.
  • Benefits: Enjoy competitive salary, health benefits, and opportunities for professional growth.
  • Other info: Collaborative environment with a focus on continuous improvement and mentorship.
  • Why this job: Be at the forefront of innovation, automating processes and improving system efficiency.
  • Qualifications: 12+ years in IT operations with strong SRE and automation skills.

The predicted salary is between 70000 - 90000 £ per year.

  • Job Location : Hove, UK (Hybrid 3 days office)
  • Job Type
  • : FTE

Job Description

SRE will play a pivotal role in driving the modernization of IT operations by implementing observability practices and automating toil.

This position requires a deep understanding of Site Reliability Engineering (SRE) principles, modern observability tools, and automation techniques to ensure scalability, reliability, and efficiency in IT systems.

This role requires a strategic thinker with hands‑on expertise who can lead modernization efforts while fostering a culture of reliability and innovation.

Primary Responsibilities

  • Work closely with Product Engineering team and implement strategies for modernizing IT operations enhancing observability and toil reduction.
  • Architect and deploy observability platforms to monitor system health, performance, and reliability effectively.
  • Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR.
  • Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets.
  • Establish & create AIOPS roadmap for improving operational efficiency.
  • Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions.
  • Drive toil automation initiatives for automated incident responses & self‑healing automation for achieving autonomous operations.
  • Collaborate with cross‑functional teams to ensure systems are scalable, resilient, and maintainable.
  • Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations.
  • Partner with engineering, architecture, and product teams to enable shift‑left engineering practices ensuring reliability.
  • Mentor and guide teams on adopting SRE principles and tools.
  • Advocate for a culture of reliability, automation, and continuous improvement across the organization.

Key Skills

  • Strong expertise in implementing Site Reliability Engineering (SRE) principles.
  • Advanced knowledge of establishing observability using tools –

Dynatrace &Datadog (primary skills).

  • Proficiency in automation & scripting using
  • Python
  • Ansible

(primary skills).

  • Strong experience with cloud platforms –
  • AWS

Azure (primary skills).

  • Solid understanding of containerization and orchestration tools like
  • Docker and

Kubernetes .

  • Proficiency in cloud native distributed systems & microservices architecture.
  • Exposure to AI/ML techniques for predictive analytics and automated problem resolution.
  • Familiarity with CI/CD pipelines & enabling automated release & deployment engineering solutions.
  • Good to have experience with chaos engineering tools like
  • Gremlin or

Chaos Monkey and implementing automation frameworks for resilience tracking.

  • Ability to manage and prioritize multiple projects in a fast‑paced environment.
  • Strong interpersonal and communication skills to work effectively across teams.
  • Excellent problem solving, analytical thinking, and adaptability.
  • Strategic mindset balancing engineering excellence with business priorities.

Preferred Qualifications

  • 12+ years of experience in IT operations, SRE, or Dev Ops roles.
  • Proven track record of SRE experience in implementing observability and automation solutions in large‑scale environments.
  • Certifications in cloud platforms, observability tools & other SRE related areas.
  • #J-18808-Ljbffr

Site Reliability Engineer in Brighton employer: Falcon Smart IT (FalconSmartIT)

Falcon Smart IT is an exceptional employer that fosters a collaborative and innovative work culture in the heart of London. With a strong emphasis on employee growth, we provide opportunities for professional development through cutting-edge projects and the latest technologies, ensuring our team members thrive in their careers. Enjoy the unique advantage of a flexible work schedule with four days in-office, allowing for a balanced work-life experience while contributing to impactful solutions in the tech industry.

Falcon Smart IT (FalconSmartIT)

Contact Details:

Falcon Smart IT (FalconSmartIT) Recruitment Team

We think you need these skills to ace Site Reliability Engineer in Brighton

Site Reliability Engineering (SRE) principles
Observability tools (Dynatrace, Datadog)
Automation & scripting (Python, Ansible)
Cloud platforms (AWS, Azure)
Containerization and orchestration (Docker, Kubernetes)
Cloud native distributed systems & microservices architecture
AI/ML techniques for predictive analytics