AI Research & QualityHybridFull-timeDisability Confident Committed# RLHF ManagerLead Reinforcement Learning from Human Feedback (RLHF) programmes across Coaley Peak's internal AI engine development and external client AI projects, ensuring our models are safe, aligned, and genuinely useful.LocationRemote (UK) / Cheltenham, UKSalary£55,000 – £75,000 per annum (DOE)Positions1 positionReferenceCP-RLF-2025-001Published21 March 2025Closing14 December 2026›AI Research & Quality›RLHF ManagerAbout the roleReinforcement Learning from Human Feedback is one of the most important mechanisms for making AI systems behave well in the real world. As RLHF Manager at Coaley Peak, you will design and manage the feedback pipelines, annotation programmes, and evaluation frameworks that keep our models (and our clients' models) aligned with human values, business objectives, and UK regulatory expectations.This is a dual-facing role. Internally, you will own the RLHF and alignment processes for our proprietary AI engines (Owlpen, Anvil, Flint, Warden), working directly with our data scientists and engineers to define reward models, manage human evaluator programmes, and measure output quality over time. Externally, you will support client AI projects where alignment and quality assurance are a requirement, particularly in regulated sectors such as financial services, healthcare, and government.We are looking for someone with a rigorous mind, a genuine interest in AI safety and alignment, and the project management capability to run structured annotation and evaluation programmes at scale. This role sits within our AI Research & Quality team and reports to the Head of AI.The kind of person we are looking forAt Coaley Peak, the technical work is only half the job. We are looking for people who are genuinely reliable, who do what they say they will, when they said they would, without needing to be chased. People who are friendly and easy to work with, both with colleagues and with clients. People who can sit in a boardroom and explain a complex AI model in plain English, and then go back to their laptop and write clean, well-documented code. People who are hungry, hard-working, and take real pride in their output, not because someone is watching, but because that is simply how they operate.You are meticulous without being slow. You understand that getting AI alignment right requires both technical rigour and human judgement, and you are equally comfortable designing an evaluation rubric, briefing a team of human annotators, and presenting findings to a client's technical steering group. You are intellectually honest, you flag problems before they become failures, and you treat every degradation in model behaviour as a signal worth investigating rather than a number to smooth over. You care about AI being safe and useful, not just impressive.- →Support external client AI projects requiring RLHF, alignment review, or model evaluation services, scoping requirements, managing delivery, and reporting outcomesWhat we are looking forEssential* →Strong understanding of RLHF, preference learning, and alignment concepts, able to explain and operationalise them in a business context* →Experience managing structured data annotation, human evaluation, or labelling programmes* →Proficiency in Python; familiarity with ML frameworks and LLM APIs (OpenAI, Anthropic, HuggingFace, or similar)* →Excellent project management skills: able to run parallel programmes with multiple contributors and tight quality standards* →Clear, precise written and verbal communication, able to document methodology and present findings to both technical and non-technical audiences* →Right to work in the United KingdomDesirable (not essential)* →Direct experience with RLHF pipelines in production (reward modelling, PPO, DPO, or similar)* →Familiarity with Constitutional AI, RLAIF, or other scalable oversight approaches* →Experience in AI safety research or AI governance* →Background in cognitive science, linguistics, psychology, or philosophy (relevant to preference elicitation and evaluation design)* →Experience working in regulated sectors where AI output quality has legal or compliance implicationsWhat we offer* →Salary of £55,000 – £75,000 per annum, dependent on experience, reviewed annually* →28 days annual leave plus bank holidays* →Flexible hybrid working, remote-first with optional Cheltenham HQ access* →Private healthcare and employee assistance programme (EAP)* →£2,000 annual professional development budget plus access to research resources* →The opportunity to shape alignment practice at a fast-growing AI consultancy* →Regular team offsites and an annual international travel programme* →High-trust, outcome-focused culture* →Accredited Living Wage EmployerVetting & pre-employment checksEnhanced DBSAll Coaley Peak roles require a minimum of an Enhanced DBS check as standard. By applying you consent to these checks being conducted in the event of an offer being made. This role involves access to sensitive model outputs and may involve work on client projects in regulated sectors. A standard background and reference check is required in addition to Enhanced DBS.How to applySend your CV and a covering note to careers@coaleypeak.co.uk, quoting reference CP-RLF-2025-001. We are particularly interested in your experience with annotation programme design, evaluation methodology, or alignment work, describe something concrete you have built or managed. We review on a rolling basis and acknowledge all applications within five working days.Applications are accepted electronically via our online form below, or by email to careers@coaleypeak.co.uk. If you require this role information or the application form in an alternative format (including large print, audio, or plain text) please email us before applying and we will provide it promptly.Disability Confident Committed, inclusive hiringDWP Disability Confident Committed employerWe are committed to inclusive and accessible recruitment. As a Disability Confident Committed employer, we will offer an interview to any disabled applicant who meets the essential criteria for this role as set out above. Please indicate in your application if you wish to be considered under this commitment.We anticipate and provide reasonable adjustments throughout every stage of the recruitment process (including the application form, any assessments, and the interview itself. If you need any adjustments) alternative formats, additional preparation time, a different interview setting, British Sign Language interpretation, or anything else, please let us know as early as possible by emailing careers@coaleypeak.co.uk.This job advert and all supporting documents are available in alternative accessible formats on request, including large print and electronic formats. Email us at careers@coaleypeak.co.uk with your preferred format.All requests are handled in confidence and will not affect how your application is assessed. We are committed to supporting any existing employee who acquires a disability or long-term health condition to remain in work and continue contributing their skills and experience.Career progressionTypical entry pointsML EngineerResearch ScientistAI Quality AnalystData ScientistWhere this can leadHead of AI QualityDirector of AI ResearchVP of Machine LearningChief AI OfficerDisclaimer: Career progression paths shown are indicative and based on typical industry trajectories. They are not a guarantee of promotion or role availability at Coaley Peak or any other organisation. Progression depends on individual performance, business requirements, and market conditions.RLHF ManagerRole at a glanceTeamAI Research & QualityLocationRemote (UK) / Cheltenham, UKWork typeHybridContractFull-timeSalary£55,000 – £75,000 per annum (DOE)VettingEnhanced DBSRefCP-RLF-2025-001Our #J-18808-Ljbffr
RLHF Manager in Cheltenham
RLHF Manager in Cheltenham
Cheltenham Full-Time No working from home possible