- Build resilient, scalable, fault-tolerant infrastructure for WRITER's high-traffic enterprise generative AI platform
- Move between SRE, DevOps, Infrastructure, and Platform initiatives as priorities shift
- Automate operational tasks and infrastructure management with Python or Go
- Design and operate infrastructure across AWS, GCP, and Azure
- Work with Kubernetes, Helm, Terraform, and cloud and AI tooling
- Use AI agents to investigate incidents, draft Terraform and Helm changes, write runbooks, scaffold tooling, and review pull requests
- Encode recurring infrastructure tasks as reusable internal skills for human and agent teammates
- Lead incident response, post-mortems, and root-cause analyses
- Own reliability, performance, and efficiency of core services end-to-end
- Define and uphold SLOs and error budgets and carry the on-call pager
- Balance immediate reliability work with long-term platform, observability, cost, and reliability investments
- Collaborate with product, security, and engineering peers on system design from conception through launch
- Report to the director of engineering
Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
- Experience running containerisation in production
- Experience with Helm and Terraform or Pulumi on at least one major cloud (AWS preferred)
- Good proficiency in Python or Go for automation and tooling
- Daily workflow already includes agentic tooling such as Claude Code, Droid, Codex, or internal skills; this is a hard requirement
- Demonstrated ability to challenge the status quo, identify systemic weaknesses, and propose innovative solutions to complex reliability problems
- Ability to reason from constraints and failure modes and articulate tradeoffs in business terms
- Ability to make reversible decisions, write rollback plans, work with monitoring and logging stacks, and stress systems safely
- Excellent communication, collaboration, and problem-solving skills
- Strong ownership and accountability for mission-critical systems
- At least one end-to-end 0-to-1 infrastructure build with an attached outcome metric
- Willingness and ability to work in person in the office 3 days per week
- Legally authorized to work in the country where the job is located
Core Competencies
Demonstrates expertise in building and operating resilient, scalable infrastructure for high‑traffic enterprise platforms, with a strong focus on automation using Python or Go. Proven ability to lead incident response and uphold reliability standards while collaborating effectively with cross‑functional teams.
Highest‑signal resume keywords
- Infrastructure Engineering
- DevOps Practices
- AWS, GCP, Azure
- Python or Go Automation
- Helm and Terraform
Hard Skills
- Infrastructure Engineering
- DevOps
- Automation
- Containerization
- Incident Response
- Root‑Cause Analysis
- SLO Definition
- Monitoring and Logging
- System Design
- High‑Availability Systems
Soft Skills
- Excellent Communication
- Collaboration
- Problem-Solving
- Ownership
- Accountability
Industry Keywords
- Generative AI
- High‑Traffic Platforms
- Production Systems
- Reliability Engineering
- High‑Growth Product Company
Tools & Technologies
- Kubernetes
- Helm
- Terraform
- Pulumi
- Cloud Tooling
- Agentic Tooling
#J-18808-Ljbffr
Infrastructure Engineer employer: Jobtailor
As a Client Services Coordinator at our dynamic company, you will thrive in a supportive work culture that prioritises employee growth and development. We offer comprehensive training, opportunities for advancement, and a collaborative environment where your contributions are valued. Located in a vibrant area, our team enjoys a healthy work-life balance and the chance to engage with diverse clients, making every day rewarding and meaningful.