Senior Site Reliability Engineer (GIPHY)

Senior Site Reliability Engineer (GIPHY)

Full-Time No working from home possible
Shutterstock
  • GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms
  • You will also work closely with our development teams to improve reliability, scalability, and operational efficiency of our systems, while ensuring code can be deployed safely and efficiently across the organization
  • You will also play a key role in responding to production incidents, troubleshooting complex issues across the infrastructure stack, and driving improvements based on what we learn from them
  • GIPHY serves internet-scale traffic every day, so the ideal candidate must be comfortable with operating large and distributed systems and solving problems at scale. Your work will directly impact billions of daily users, alongside some of the world’s largest social media platforms and apps
  • Operate and evolve GIPHY’s cloud and CDN presence, ensuring resources are reliable, performant, scalable and cost-efficient
  • Continuously improve and optimize our CI/CD platform to enhance developer experience and deployment reliability
  • Drive the implementation of multi-region Kubernetes clusters while operating and improving our existing Kubernetes environments
  • Partner with other Engineering teams to troubleshoot complex production and infrastructure issues, taking a holistic view across our platforms to ensure the reliability and availability of GIPHY’s systems
  • Stay current with emerging technologies, particularly agentic development and AI-assisted engineering, and identify opportunities to incorporate them into our platforms and engineering workflows

Benefits

  • Flexibility: We offer a hybrid model that’s centered around flexibility and enabling collaboration, innovation, and accountability—no matter where we are.
  • Connection: Whether online or on site, we make sure there are plenty of ways to connect and have fun.
  • Recognition: We never miss an opportunityto celebrate our achievements and commend great work through our global recognition program.
  • Empowerment: We offer generous and competitive compensation packages, PTO, wellness initiatives, on-and-off working hours, tuition reimbursement, and an employee referral bonus.
  • Growth: Our SkillUP Program lets you take ownership of your personal development through virtual learning offerings.
  • Belonging: Our goal is to build a workforce that’s representative of the diverse global community we serve, and where all employees can come to work as their authentic selves.
  • Expert-level knowledge and significant professional experience with Infrastructure as a Service (IaaS) operations and Infrastructure as Code (IaC), particularly Terraform and Atlantis
  • Demonstrate a high degree of autonomy and initiative, proactively identifying areas for improvement and independently driving work around optimization, cost reduction, patching, upgrades, and technical debt
  • Experience with agentic engineering and effectively applying AI-assisted development practices
  • Ability to design and implement cost-effective solutions and proactively identify cost-optimization opportunities
  • Proficiency in one or more programming languages, such as Python, Java, Go or Rust
  • Experience with CI/CD tools such as Jenkins or Github Actions and deployment tooling such as Helm, Spinnaker and ArgoCD
  • Strong systems engineering fundamentals, with deep knowledge of Linux and networking and the ability to diagnose complex reliability and performance issues across the OS, network, container, and cloud infrastructure layers
  • Strong problem-solving and communication skills, and the ability to work effectively in a team environment
  • Experience with observability and monitoring platforms such as Datadog and AWS CloudWatch, including diagnosing production reliability and performance issues
  • Extensive experience with containerization and orchestration technologies, mainly Docker and Kubernetes
  • Strong experience with operating, designing and troubleshooting AWS infrastructure, especially EKS, S3, EC2 and VPC

#J-18808-Ljbffr

Senior Site Reliability Engineer (GIPHY) employer: Shutterstock

Shutterstock is an exceptional employer that fosters a dynamic and innovative work culture, perfect for those looking to make a significant impact in the creative industry. With a strong focus on employee growth and development, team members are encouraged to explore new ideas and solutions while collaborating with leading brands across the UK, Ireland, and Nordics. Located in London, employees benefit from a vibrant city atmosphere, access to diverse networking opportunities, and the chance to be part of a cutting-edge technology company that values creativity and teamwork.

Shutterstock

Contact Details:

Shutterstock Recruitment Team