At a Glance
- Tasks: Measure and improve AI systems by designing evaluations and building datasets.
- Company: Join a leading tech firm focused on cutting-edge AI innovations.
- Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
- Other info: Collaborative environment with a focus on real-world applications and career advancement.
- Why this job: Make a real impact on AI quality and user experience in production.
- Qualifications: Strong Python skills and experience with ML systems required.
The predicted salary is between 60000 - 80000 £ per year.
This role sits at the centre of how we measure and improve AI systems in production. You’ll define what good performance means across LLMs, ASR, TTS, and full speech-to-speech pipelines, and build the datasets, metrics, and evaluation systems that make AI quality measurable and comparable in the real world. You’ll work closely with engineering and product teams to ensure model changes lead to real improvements in user experience, not just better offline benchmarks.
What you’ll do:
- Design and run evaluations across LLM, ASR, TTS, and speech-to-speech systems
- Build real-world datasets and test cases from production behaviour and edge cases
- Define metrics and scorecards for model and system quality
- Benchmark internal models against external and frontier systems
- Evaluate full pipelines (ASR → LLM → TTS), not just individual models
- Build Python tools to automate evaluation workflows
- Create internal leaderboards, red-teaming setups, and regression tests
- Work with engineers and product teams to diagnose system failures
- Turn vague product goals into measurable evaluation frameworks
What this role is about:
- Defining and measuring AI quality in production systems
- Turning real user behaviour into structured evaluation signals
- Ensuring model changes improve real-world performance
- Understanding why AI systems fail, not just whether they do
What good looks like:
- You can translate improved quality into measurable metrics
- You think in terms of system impact (before vs after), not just accuracy
- You’re comfortable working across code, data, and production systems
- You care about real-world behaviour, not just benchmarks
Core skills:
- Strong Python (scripting, data analysis, tooling)
- Experience with ML systems, evaluation, or experimentation
- Understanding of LLMs or speech systems (ASR / TTS)
- Ability to design test cases and structured datasets
- Comfortable working with engineers and product teams
Nice to have:
- Experience with LLM evaluation or benchmarking
- Exposure to speech or multimodal systems
- Familiarity with production APIs or ML systems
- Experience with automated testing or CI-style workflows
AI Evaluations Engineer in Manchester employer: ConnexAI
At ConnexAI, we pride ourselves on being an excellent employer by fostering a collaborative and innovative work culture that empowers our employees to thrive. Our Manchester office offers a dynamic environment where you can engage in meaningful AI projects while benefiting from continuous learning and growth opportunities. Join us to be part of a global team dedicated to shaping the future of language technology, all while enjoying the excitement of working at the forefront of AI advancements.
StudySmarter Expert Advice🤫
We think this is how you could land AI Evaluations Engineer in Manchester
✨Tip Number 1
Network like a pro! Reach out to folks in the AI and ML space, especially those who work with LLMs, ASR, and TTS. Attend meetups or webinars, and don’t be shy about asking questions – it’s all about making connections that could lead to your next opportunity.
✨Tip Number 2
Show off your skills! Create a portfolio showcasing your Python projects, especially those related to AI evaluations or benchmarking. This will not only demonstrate your technical prowess but also give potential employers a taste of what you can bring to the table.
✨Tip Number 3
Prepare for interviews by brushing up on real-world applications of AI systems. Be ready to discuss how you would define metrics and evaluate performance in production environments. We want to see your thought process and how you tackle challenges!
✨Tip Number 4
Don’t forget to apply through our website! It’s the best way to ensure your application gets seen by the right people. Plus, we love seeing candidates who are proactive and engaged with our company.
We think you need these skills to ace AI Evaluations Engineer in Manchester
Some tips for your application 🫡
Tailor Your Application:Make sure to customise your CV and cover letter for the AI Evaluations Engineer role. Highlight your experience with Python, ML systems, and any relevant projects that showcase your skills in evaluation and experimentation.
Showcase Real-World Impact:When describing your past experiences, focus on how your work has led to measurable improvements in AI quality or user experience. We want to see how you think about system impact, not just technical accuracy.
Be Clear and Concise:Keep your application straightforward and to the point. Use clear language to explain your skills and experiences, making it easy for us to see why you’d be a great fit for the team.
Apply Through Our Website:Don’t forget to submit your application through our website! It’s the best way for us to receive your details and ensures you’re considered for the role. We can’t wait to see what you bring to the table!
How to prepare for a job interview at ConnexAI
✨Know Your AI Systems
Make sure you brush up on your knowledge of LLMs, ASR, and TTS systems. Understand how they work together in a speech-to-speech pipeline. Being able to discuss these technologies confidently will show that you're not just familiar with the concepts but can also apply them in real-world scenarios.
✨Prepare Real-World Examples
Think of specific instances where you've designed evaluations or worked with datasets in production. Be ready to share how you turned vague product goals into measurable metrics. This will demonstrate your ability to translate theory into practice, which is crucial for this role.
✨Showcase Your Python Skills
Since strong Python skills are essential, be prepared to discuss your experience with scripting, data analysis, and building tools. If possible, bring examples of your work or even a small project that highlights your ability to automate evaluation workflows.
✨Collaborate and Communicate
This role involves working closely with engineers and product teams, so highlight your teamwork and communication skills. Share experiences where you successfully collaborated on projects, especially those that required diagnosing system failures or improving user experience.