At a Glance
- Tasks: Join us to enhance the reliability of our cutting-edge DBaaS product, Neo4j Aura.
- Company: Neo4j, a leader in graph intelligence and data innovation.
- Benefits: Competitive salary, remote work options, and a vibrant team culture.
- Other info: Dynamic environment with opportunities for growth and collaboration across teams.
- Why this job: Be part of a transformative journey in data and AI with real-world impact.
- Qualifications: Experience in software development, especially with Go, and a passion for reliability engineering.
The predicted salary is between 60000 - 80000 £ per year.
About Neo4j: Neo4j is the graph intelligence platform that transforms data into knowledge to power the next generation of intelligent applications and AI systems. It includes enterprise-ready knowledge graphs for accurate, explainable, and governed AI; the most comprehensive, trusted, and easy-to-deploy graph capabilities across any environment and data source; and an unmatched ecosystem trusted by 84 of the Fortune 100 and supported by the world’s largest graph community.
Our Vision: At Neo4j, we have always strived to help the world make sense of data. As business, society and knowledge become increasingly connected, our technology promotes innovation by helping organizations to find and understand data relationships. We created, drive and lead the graph database category, and we’re disrupting how organizations leverage their data to innovate and stay competitive.
The Team: The Site Reliability Engineering team’s mission is to improve the reliability of Neo4j’s DBaaS product: Neo4j Aura. Operating at a global scale across all three major cloud providers, Aura runs hundreds of Kubernetes clusters and hosts thousands of Neo4j instances in production at any given time. We’re reshaping what SRE means at Neo4j Aura—and we want you to be part of that journey.
The Role:
- Automate for insight and scale: Build systems that make troubleshooting fast, safe, and scalable across thousands of Neo4j instances.
- Treat operations as a software problem: Replace tribal knowledge and ad-hoc scripts with tools and systems that codify best practices—making operations predictable, scalable, and repeatable.
- Design for resilience, learn from failure: Own and evolve the tooling and processes behind incident response.
- Champion reliability as a product feature: Help teams define and act on SLIs and SLOs, turning reliability into a shared, data-driven priority across engineering.
- Create signals, not noise: Shape an observability stack that tells us what matters, when it matters—so we can detect issues early and resolve them quickly.
We're interested in hearing from Engineers with deep experience in some of the following areas:
- Writing backend tools and automation in Go—our primary language—with an emphasis on sound architecture, testing, and maintainability.
- Strong software skills in other languages, like Python, are also welcome.
- Applying SRE practices in real-world environments: defining SLIs and SLOs, reducing toil through automation, and driving reliability through engineering.
- Collaborating with other teams to promote SRE thinking—educating on principles like observability, ownership, and service level objectives.
- Troubleshooting large-scale, cloud-based systems with confidence and curiosity.
- Monitoring distributed systems and understanding their performance characteristics.
- Designing systems with reliability, safety, and debugability as first-class concerns.
- Working with observability tools like OTel Collector, Prometheus, Grafana, and Google Cloud’s operations suite.
- Deploying and managing applications on Kubernetes; cluster-level administration is a plus.
- Managing infrastructure with Kustomize and Terraform—keeping it clear, modular, and easy to evolve.
- Building and maintaining CI/CD workflows—ours run on GitHub Actions.
- Participating in on-call rotations and incident response with a focus on improvement, not blame.
- Writing and contributing to postmortems that lead to meaningful, lasting changes.
Why Join Neo4j? Neo4j is, without question, the most popular graph intelligence platform in the world. We have customers in every industry globally, and our products are a proven product/market fit. Joining our team is an opportunity to shape the future of data and analytics.
Neo4j is committed to building awareness and helping to improve these issues. One of our central objectives is to provide an inclusive, diverse, and equitable workplace for everyone to develop their potential and have a positive, career-defining experience. We look forward to receiving your application.
Neo4j Values: Neo4j is a Silicon Valley company with a Swedish soul. We foster collaboration and each of us is empowered to contribute and put our innovative stamp on projects.
Neo4j is committed to protecting and respecting your privacy. Please read the privacy notice regarding Neo4j's recruitment process to understand how we will handle the personal data that you provide.
Software Engineer - Site Reliability Engineering in London employer: Neo4j
At Neo4j, we pride ourselves on being a leader in the graph database space, offering an inclusive and innovative work culture that empowers employees to shape the future of data and analytics. As a Remote Technical Curriculum Developer based in London, you'll enjoy flexible working arrangements, opportunities for professional growth, and the chance to engage with a vibrant community of developers and data scientists. Join us to make a meaningful impact while collaborating with talented teams dedicated to user success and continuous learning.