At a Glance
- Tasks: Join Fitch Group to enhance service reliability and drive strategic platform initiatives.
- Company: Fitch Group, a global leader in financial information services with an inclusive culture.
- Benefits: Hybrid work model, learning opportunities, and a supportive environment for growth.
- Other info: Collaborate in a dynamic team and contribute to innovative AI/ML projects.
- Why this job: Make a real impact in financial markets while working with cutting-edge technology.
- Qualifications: Experience in SRE, DevOps, or Platform Engineering with AWS and Azure.
The predicted salary is between 63000 - 77000 £ per year.
This job is with Fitch Group, an inclusive employer. Fitch Group is currently seeking a Service Reliability Engineer to embed with Fitch Ratings development squads. The role is based out of our Manchester office. As a leading global financial information services provider, Fitch Group delivers vital credit and risk insights, robust data, and dynamic tools to champion more efficient, transparent financial markets.
Fitch's Technology establishes adoption guardrails. Define and enforce cloud guardrails and security controls (SCPs/IAM boundaries, OPA policies, tagging, centralized logging with AWS Config/CloudTrail/Security Hub) in partnership with Security and Risk. Influence cross‑functional roadmaps, lead complex release planning, and drive strategic platform initiatives across CI. Serve as an escalation point and participate in the L3 on‑call rotation.
You May be a Good Fit if:
- You have deep, hands-on experience in SRE, DevOps, or Platform Engineering across both AWS and Azure, with a strong track record operating Docker and Kubernetes in production environments.
- You’re highly proficient in administering both Linux and Windows, and have practical, enterprise-level experience supporting IIS/.NET applications as well as Java Spring Boot services.
- You have built and maintained CI/CD pipelines (primarily GitHub Actions; Bamboo experience a plus) with DevSecOps principles baked in—integrating security scans, policy-as-code, and compliance gates—and script confidently in Python, PowerShell, or Bash.
- You have experience with cloud security best practices (IAM, secrets management, container/image scanning) and understand core infrastructure fundamentals (networking, storage, DNS) and APM/telemetry tooling.
What Would Make You Stand Out:
- Practical experience with agentic AI for operations—incident triage, runbooks, and change management—with clear guardrails, auditability, and human-in-the-loop controls.
- Supporting AI/ML workloads at scale: SageMaker endpoints, GPU node groups, autoscaling, and Kubernetes-based model serving.
- Policy-as-code (OPA) and compliance implementation across CIS, NIST, ISO 27001, with automated remediation integrated via CSPM tools (e.g., Wiz).
- Applying AI in CI/CD, observability, and incident response using AWS Bedrock/SageMaker and Model Context Protocol (MCP).
- Hands-on Agile delivery experience, actively participating in stand-ups and sprint ceremonies.
Why Choose Fitch:
- Hybrid Work Environment: 2 to 3 days a week in office required based on your line of business and location.
- A Culture of Learning.
Service Reliability Engineer - Manchester employer: Fitch Group
Fitch Group is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration within its London-based Media Strategy & Communications team. Employees benefit from a hybrid work environment, competitive compensation, and ample opportunities for professional growth, making it an ideal place for those looking to make a meaningful impact in the financial services sector.