SRE (Terminal)

SRE (Terminal)

Full-Time 60000 - 80000 £ / year (est.) No working from home possible
M

At a Glance

  • Tasks: Design and maintain high-availability cloud infrastructure for a decentralised crypto social network.
  • Company: Join a high-growth software development organisation at the forefront of crypto technology.
  • Benefits: Competitive salary, equity options, and a commitment to equality and accessibility.
  • Other info: Opportunity for professional growth and to work on innovative projects in Web3.
  • Why this job: Make a real impact in a fast-paced environment while working with cutting-edge technologies.
  • Qualifications: Expertise in infrastructure-as-code, cloud providers, and high-availability systems required.

The predicted salary is between 60000 - 80000 £ per year.

Compensation: Competitive Compensation

Location: New York, United States; London, United Kingdom (Office)

Role Overview

Our client is a high‑growth software development organization focused on a large decentralized crypto social network. To support expansion and ensure continuous uptime of its high‑throughput environment, the company seeks an experienced Site Reliability Engineering (SRE) Expert.

Key Responsibilities

  • Own Foundation & Architecture: Design, scale, and maintain highly available, multi‑region, or active‑active cloud infrastructure patterns
  • Incident Response & Reliability: Lead critical incident response efforts, participate in real on‑call rotations, and drive comprehensive, blameless post‑mortems to continually harden the system
  • Automation & Tooling: Write clean, production‑grade automation code (Python, Go, or similar) for infrastructure tooling, operators, and seamless systems integration
  • Risk & Security Management: Exercise judgment regarding system risks, balancing rapid deployment velocity with robust infrastructure safety and stability
  • Operational Excellence: Raise the engineering and operational bar through rigorous standards, modern tooling, and technical mentorship

Requirements

  • Deep expertise in infrastructure‑as‑code (Terraform/OpenTofu), network topology, high‑availability architecture, and system internals
  • Proven track record building foundational infrastructure (0→1) and running high‑availability environments where reliability is treated with financial‑system level seriousness
  • Advanced proficiency with modern cloud providers (AWS, GCP) and container orchestration platforms (Kubernetes)
  • Strong capacity to operate independently in high‑stakes environments, deciding when to gather consensus versus when to execute autonomously

Preferred Qualifications

  • Experience with infrastructure security hardening, IAM architecture, or compliance mapping (e.g., SOC2, ISO)
  • Hands‑on experience managing and scaling high‑throughput, low‑latency data backbones and event streaming systems (Kafka, Redpanda, PostgreSQL)
  • Understanding of Web3/crypto infrastructure patterns and comfort operating within them

Benefits

  • Competitive Base Salary
  • Equity and Token Allocation

Commitment to Equality and Accessibility: We are committed to offering equal opportunities to all candidates. We ensure no discrimination and provide accessible formats. If you need a reasonable adjustment or an accessible job advert, please let us know. Contact human‑resources@mlabs.city.

SRE (Terminal) employer: MLabs

At MLabs, we pride ourselves on being an exceptional employer, offering a collaborative and innovative work culture that empowers Plutus Developers to thrive. With generous paid time off, flexible contract options, and a commitment to employee growth, our remote team in Europe is dedicated to advancing the Cardano ecosystem while ensuring a supportive environment for all. Join us to make a meaningful impact in the world of smart contracts and blockchain technology.

M

Contact Details:

MLabs Recruitment Team

We think you need these skills to ace SRE (Terminal)

Site Reliability Engineering (SRE)
Infrastructure-as-Code (Terraform/OpenTofu)
Network Topology
High-Availability Architecture
Incident Response
Automation (Python, Go)
Risk Management