AI Reliability Engineer – Scalable, Incident-Driven Systems

AI Reliability Engineer – Scalable, Incident-Driven Systems

Full-Time 63000 - 77000 Β£ / year (est.) No working from home possible
Anthropic

At a Glance

  • Tasks: Design and implement monitoring for AI systems, ensuring reliability across various platforms.
  • Company: Join Anthropic, a leader in AI technology focused on user trust and safety.
  • Benefits: Enjoy competitive pay, flexible work options, and opportunities for professional growth.
  • Other info: Be part of a dynamic team with a focus on innovation and excellence.
  • Why this job: Make a real impact on AI reliability and user safety in a cutting-edge environment.
  • Qualifications: Experience in system design, monitoring, and strong collaboration skills required.

The predicted salary is between 63000 - 77000 Β£ per year.

Anthropic is seeking an AI Reliability Engineer to help keep Claude reliable across serving pathsβ€”from SDK through network, APIs, and infrastructure.

You will design and implement monitoring, contribute to high-availability infrastructure across regions and clouds, and lead incident response for critical AI services.

The role emphasizes holistically understanding systems, strong collaboration, and ownership of outcomes, with a direct impact on user trust and safety commitments.

#J-18808-Ljbffr

AI Reliability Engineer – Scalable, Incident-Driven Systems employer: Anthropic

Anthropic is an exceptional employer for those passionate about advancing reinforcement learning in a collaborative and innovative environment. With competitive compensation, generous vacation and parental leave, and flexible working hours, employees enjoy a supportive work culture that prioritises both personal and professional growth. Located in a vibrant office space, team members have the unique opportunity to engage directly with cutting-edge research while making meaningful contributions to the responsible scaling of AI.

Anthropic

Contact Details:

Anthropic Recruitment Team

We think you need these skills to ace AI Reliability Engineer – Scalable, Incident-Driven Systems

Monitoring Design and Implementation
High-Availability Infrastructure
Incident Response
Collaboration Skills
Systems Understanding
Ownership of Outcomes
User Trust and Safety Awareness