At a Glance
- Tasks: Join us in enhancing AI reliability and ensuring Claude's performance for users worldwide.
- Company: Anthropic, a mission-driven tech company focused on safe and beneficial AI.
- Benefits: Competitive salary, flexible hours, generous leave, and equity donation matching.
- Other info: Collaborative environment with opportunities for growth and diverse perspectives.
- Why this job: Make a real impact on AI systems that shape the future of technology.
- Qualifications: Experience in distributed systems and a passion for reliability in software engineering.
The predicted salary is between 63000 - 77000 £ per year.
About Anthropic: Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the Role: AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team's boundaries, so someone has to zoom out and look at the whole picture.
Responsibilities:
- Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
- Design and implement monitoring and observability systems across the token path.
- Assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
- Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
- Support the reliability of safeguard model serving -- critical for both site reliability and Anthropic's safety commitments.
You may be a good fit if you:
- Have strong distributed systems, infrastructure, or reliability backgrounds.
- Are curious and brave -- comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don't have deep expertise yet.
- Think holistically about how systems compose and where the seams are.
- Can build lasting relationships across teams.
- Care about users and feel ownership over outcomes, even for systems you don't own.
- Have excellent communication and collaboration skills.
- Bring diverse experience.
Strong candidates may also:
- Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems.
- Have experience operating large-scale model serving or training infrastructure (>1000 GPUs).
- Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium).
- Understand ML-specific networking optimizations like RDMA and InfiniBand.
- Have expertise in AI-specific observability tools and frameworks.
- Have experience with chaos engineering and systematic resilience testing.
- Have contributed to open-source infrastructure or ML tooling.
The annual compensation range for this role is listed below:
Annual Salary: 325,000—390,000 GBP
Logistics:
- Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience.
- Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.
- Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position.
- Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time.
- Visa sponsorship: We do sponsor visas!
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. We think AI systems like the ones we're building have enormous social and ethical implications. We strive to include a range of diverse perspectives on our team.
Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses.
How we're different: We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. We value impact — advancing our long-term goals of steerable, trustworthy AI.
Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues.
Staff Software Engineer, AI Reliability Engineering employer: Humanloop
At Anthropic, we pride ourselves on being an exceptional employer, fostering a collaborative and innovative work culture that empowers our employees to thrive. As an Applied AI Security Architect, you'll engage with top-tier clients in a dynamic environment, benefiting from competitive compensation, generous leave policies, and opportunities for professional growth in the rapidly evolving field of AI. Our commitment to safety and ethical AI ensures that your work will have a meaningful impact on society while you enjoy the flexibility of a hybrid work model in a vibrant location.