Junior Site Reliability Engineer in Cardiff

Junior Site Reliability Engineer in Cardiff

Cardiff Full-Time 63000 - 77000 £ / year (est.) Home office (partial)
Sadler Recruitment

At a Glance

  • Tasks: Support and optimise cloud platforms while working with cutting-edge AI technologies.
  • Company: Dynamic Cardiff-based tech company leading in AI operations.
  • Benefits: Competitive salary, training opportunities, remote work, and industry event attendance.
  • Other info: Flexible environment with real ownership and career growth potential.
  • Why this job: Join a fast-growing team and shape the future of AI-era cloud infrastructure.
  • Qualifications: Familiarity with AWS or Azure and scripting skills in Python or Bash.

The predicted salary is between 63000 - 77000 £ per year.

Office requirements: Remote (central Cardiff office available 2 days a week)

This company provides managed AI operations for technology businesses. The service is not a traditional managed service model of cloud support and resets. The company operates, secures and governs the cloud, observability and AI runtime layer behind mission-critical software. The client's team owns the product, application, and model behaviour; we own the operating layer that keeps it reliable, secure, cost-controlled and evidence-ready.

Clients range from start-ups and scale-ups building fast without formal engineering controls, through to regulated, mission-critical platforms in fintech, healthtech and insurance, where the operating model and evidence trail matter as much as uptime. You'll work across traditional cloud platforms like Azure and AWS, as well as newer, AI-native application stacks such as Lovable, and the modern cloud platforms that often sit behind them, like Vercel and Supabase.

The company is a Datadog Advanced Partner (UK) and holds the accolade of being the world's first accredited MSP powered by Datadog. Datadog is a Nasdaq-listed observability platform. The company's technical focus is on observability and LLM observability specifically.

This is a chance to grow with a Cardiff business on the front line of AI-era cloud operations, with bleeding-edge tech, AI used responsibly, a team worth learning from, and the flexibility to do your best work. Other awards include the Sir Michael Moritz Tech Start-up Award.

Key Responsibilities
  • Runtime Assurance Delivery: Support customers across a mix of traditional and modern cloud platforms, Azure and AWS, alongside newer AI-native application stacks and the cloud platforms behind them, such as Lovable, Vercel and Supabase. Work through an ongoing ticket backlog per customer, sized to allow both reactive support and proactive improvement work. When a customer's backlog is empty, proactively review their environment for improvements rather than waiting for tickets to be raised.
  • Observability & Datadog: Receive training on client environments and on the Datadog platform. Identify improvements independently, either by raising them to the team or by implementing them directly. Tickets may also be generated directly by Datadog alerts, covering both backlog items and live incidents.
  • Build & Improvement Projects: Deliver end-to-end build-and-run projects for managed service customers, from scoping through to live operation. Implement cloud cost optimisation strategies, working through the technical dependencies that can sometimes block them. Manage monitoring and observability infrastructure as estates grow, keeping platforms performant and well-maintained. Build and deploy new infrastructure in Azure, including infrastructure-as-code, Datadog implementation, and APM setup, covering the full lifecycle from onboarding through go-live stabilisation.
  • AWS Platform Work: Work hands-on with containers, EC2 and RDS instances across a range of customer environments, from day-to-day configuration and troubleshooting through to cost and performance optimisation. Also extends into AWS Bedrock for customers building AI-native products.
  • Internal Tooling & Process: Contribute to a growing suite of internal AI tools built to improve engineering efficiency, reduce context-switching across platforms, and modernise how we run cloud operations day to day.
A Day in the Life

Review ticket queues across the managed service accounts. Project time is spent on a mix of Azure, AWS and Datadog implementation work, alongside newer platforms like Lovable, Vercel and Supabase. Work includes cost optimisation, infrastructure builds, and Datadog implementation and APM setup on live applications. On AWS, this includes container work, EC2 and RDS configuration, and Bedrock. Python is used throughout. Depending on the customer, this work may also involve GitHub Actions, Azure DevOps pipelines, or PowerShell scripting. Time is also allocated to internal work covering internal AI tools and process improvements. The role involves working across managed service delivery, project builds, observability, and internal tooling within the same week, rather than working on a single area on an ongoing basis.

Experience Required
  • Working knowledge of at least one cloud platform (AWS or Azure)
  • Comfort with scripting: Bash, Python, or similar
  • Understanding of core observability concepts: metrics, logs, traces
  • Clear written and verbal communication: you'll work with customers
Nice to have
  • Hands-on Datadog experience (any tier)
  • Terraform or other IaC tooling
  • Kubernetes or containerised workload exposure
  • Experience in a managed service provider or multi-customer environment
  • Familiarity with ISO 27001 or similar compliance frameworks
  • Any cloud certification (AWS, Azure, or Datadog)
What's on Offer

Join a growing Cardiff business entering its scaling phase, with real ownership, visibility, and a chance to shape how a fast-moving modern MSP operates AI-era cloud infrastructure. On-call is documented, structured and remunerated in addition to salary. Training and tooling costs are paid for by the company. Attendance at industry events and courses is encouraged and funded as part of ongoing development. Employees are expected to identify and pursue improvements independently, without pre-defined limits on scope. Due to expected high demand, we'll be reviewing CVs as they come in and may close this role early once we've received sufficient applications.

Junior Site Reliability Engineer in Cardiff employer: Sadler Recruitment

Join a dynamic global luxury asset insurer in Cardiff Bay, where you'll enjoy a flexible working pattern of 3 days in the office and 2 days from home. With a strong focus on employee growth and technical autonomy, this role offers you the chance to lead DevOps initiatives in an innovative environment, ensuring that your contributions directly impact the architecture and scalability of our platform. Experience a collaborative work culture that values your expertise and encourages you to explore new solutions while insuring some of the world's most valuable assets.

Sadler Recruitment

Contact Details:

Sadler Recruitment Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Junior Site Reliability Engineer in Cardiff

Join Local Tech Meetups

Get out there and mingle with fellow developers by joining local tech meetups. It’s a fantastic way to meet people who might be working at Sadler Recruitment or know someone who does. Plus, you can pick up some trendy tech skills and trends while you're at it!

Contribute to Open Source Projects

Show off your coding chops by jumping into open-source projects. Not only does this give you practical experience, but it also gets you noticed in the dev community. You'll create a killer portfolio that speaks volumes about your skills to Sadler Recruitment.

Tap into Online Developer Communities

Don’t underestimate the power of online developer communities like GitHub, Stack Overflow, and even Reddit. Participate in discussions, share your projects, and build your visibility. We can often find opportunities through these channels that can lead to a full-time gig at companies like Sadler Recruitment.

Explore Job Boards Specifically for Tech Roles

Keep your eyes peeled on job boards that focus on tech roles. Sites like TechCareers or Stack Overflow Jobs can often have listings for companies like Sadler Recruitment that might not show up on broader job sites. Make it a habit to check these regularly, and don’t hesitate to apply directly through our website!

We think you need these skills to ace Junior Site Reliability Engineer in Cardiff

Cloud Platform Knowledge (AWS or Azure)
Scripting (Bash, Python, or similar)
Observability Concepts (metrics, logs, traces)
Datadog Implementation
Infrastructure as Code (Terraform or similar)
Containerisation (Kubernetes or similar)
Cost Optimisation Strategies

Some tips for your application 🫡

Show off your coding skills:When applying for a software engineering role, it's super important to showcase your coding skills. Make sure your CV includes your tech stack, any relevant programming languages you’re comfortable with, and examples of projects you've worked on. If you have a GitHub profile, link it up! We love to see code in action.

Tailor your portfolio:For a full-time role, we’d expect to see some solid examples of your work in your portfolio. Make sure to include at least two or three projects that highlight your problem-solving skills and your ability to work with different technologies. Focus on the projects that are most relevant to the position at Sadler Recruitment.

Craft a killer cover letter:Your cover letter is your chance to stand out—make it personal! Explain why you want to work at Sadler Recruitment and how your skills align with the role. Show us your passion for software development. We dig enthusiastic candidates who understand the value of collaboration and continuous learning!

Be clear and concise:When it comes to writing your CV and cover letter, clarity is key. Avoid jargon that could confuse us and stick to simple, direct language. Highlight your achievements with quantifiable results where possible, and keep everything easy to read. A well-organised application goes a long way!

How to prepare for a job interview at Sadler Recruitment

Brush Up on Your Coding Skills

For a full-time software engineering role, it's crucial that we stay sharp with our coding abilities. Expect technical questions that might involve solving problems on the spot or discussing algorithms. Practise on platforms like LeetCode or HackerRank to get comfortable with the types of questions that often come up.

Know Your Tools and Frameworks

Make sure we’re well-acquainted with the tools and technologies listed in the job description. Familiarise ourselves with any specific frameworks or programming languages mentioned. If Sadler Recruitment uses React or Node.js, for instance, be ready to discuss how we’ve used them in previous projects or coursework.

Showcase Your Projects

Bring along a portfolio that highlights our best work. This could be code samples, GitHub repositories, or any side projects we’ve built. Make sure we can talk through our thought process for each project, especially the challenges we faced and how we solved them—this shows our problem-solving skills in action.

Prepare for Behavioural Questions

While technical skills are key, full-time positions also require cultural fit. Be ready to discuss our previous experiences and how we handle teamwork, conflict, and deadlines. Brush up on the STAR method—Situation, Task, Action, Result—to clearly articulate our past experiences when discussing how we've contributed to a team.