At a Glance
- Tasks: Design and implement evaluation frameworks for cutting-edge AI models.
- Company: Join LILT, a leader in AI-driven translation technology.
- Benefits: Competitive salary, remote work options, and growth opportunities.
- Other info: Dynamic team environment with a focus on collaboration and innovation.
- Why this job: Be the quality gatekeeper for innovative AI solutions that change communication.
- Qualifications: Expertise in Python and modern AI frameworks required.
The predicted salary is between 36000 - 60000 £ per year.
About LILT AI is changing how the world communicates — and LILT is leading that transformation. We are on a mission to make the world's information accessible to everyone, regardless of the language they speak. We use cutting‑edge AI, machine translation, and human‑in‑the‑loop expertise to translate content faster, more accurately, and more cost‑effectively without compromising on brand, voice, or quality. At LILT, we empower our teammates with leading tools, global collaboration, and growth opportunities to do their best work. Our company virtues—Work together, win together; Find a way or make one; Quicker than they expect; Quality is Job 1—guide everything we do.
We are trusted by Intel Corporation, Canva, the United States Department of Defense, the United States Air Force, ASICS, and hundreds of global enterprises. Backed by Sequoia, Intel Capital, and Redpoint, we’re building a category‑defining company in a $50B+ global translation market being redefined by AI.
About The Role As a Research Engineer focused on Model Evaluation, you are the final arbiter of technical quality for our frontier AI deliverables. You will design sophisticated evaluation suites and serve as the lead calibrator, reviewing and refining the contributions of other engineers to ensure our data samples and model outputs meet the exacting standards of the world’s leading AI labs. This is a highly technical role for someone who enjoys getting in the weeds of model behaviour, RAG performance, and RLHF alignment.
- Key Responsibilities
- Eval Architecture & Benchmarking: Design and implement automated and human‑in‑the‑loop evaluation frameworks to measure model performance across multiple modalities (text, code, image, etc.).
- Calibration & Peer Review: Act as the Gold Standard reviewer for other engineers. You will calibrate their data generation and evaluation contributions, providing technical feedback to ensure scientific consistency and high‑fidelity output.
- Frontier Sample Generation: Write and refine complex prompts and golden response pairs for frontier‑model training, specifically focusing on edge cases in reasoning and multilingual contexts.
- Quality Control (End‑to‑End): Develop the logic for multimodal QC checks, ensuring that high‑volume data samples are correct across diverse domains and languages.
- Technical Mentorship: Bring new knowledge and best practices to our established delivery and forward‑deployed engineering teams on model evaluations.
Qualifications
- Education: B.S. in Computer Science, AI, or a related field or 5+ years of relevant experience in a high‑growth AI/Research environment.
- Deep Technical Proficiency: Expert-level Python skills and hands‑on experience with modern AI frameworks (PyTorch, Transformers, LangChain/LlamaIndex).
- Evaluation Experience: Experience building model evaluation suites (e.g., MMLU‑style benchmarks, custom RAG metrics, or human‑preference alignment).
- Domain Expertise: Deep understanding of RAG architectures, vector database retrieval logic, and agentic workflows. Experience with RLHF/RLAIF environments and the mechanics of preference signalling/reward modelling.
- Multimodal & Multilingual Rigor: Experience handling data quality at scale across different languages and modalities (images, video, or audio).
- Precision‑ and Quality‑Orientation: You find bugs in model reasoning that others miss. You are comfortable being the final quality arbiter for technical deliverables that others produce.
Preferred Skills
- Fluency in multiple languages (highly preferred for multilingual model calibration).
- Experience in Frontier Labs or high‑tier AI research environments.
- At least one of: a portfolio of research contributions, an example of evals or “model‑breaking” samples, or use of open‑source AI evaluation tools.
Our Story
Our founders, Spence and John met at Google working on Google Translate. As researchers at Stanford and Berkeley, they both worked on language technology to make information accessible to everyone. While together at Google, they were amazed to learn that Google Translate wasn’t used for enterprise products and services inside the company. The quality just wasn’t there. So they set out to build something better. LILT was born. LILT has been a machine learning company since its founding in 2015. At the time, machine translation didn’t meet the quality standard for enterprise translations, so LILT assembled a cutting‑edge research team tasked with closing that gap. While meeting customer demand for translation services, LILT has prioritised investments in Large Language Models, human‑in‑the‑loop systems, and now agentic AI. With AI innovation accelerating and enterprise demand growing, the next phase of LILT’s journey is just beginning.
What Sets Our Platform Apart
- Brand‑aware AI that learns your voice, tone, and terminology to ensure every translation is accurate and consistent.
- Agentic AI workflows that automate the entire translation process from content ingestion to quality review to publishing.
- 100+ native integrations with systems such as Adobe Experience Manager, Webflow, Salesforce, GitHub, and Google Drive to simplify content translation.
- Human‑in‑the‑loop reviews via our global network of professional linguists, for high‑impact content that requires expert review.
LILT in the News
Featured in The Software Report’s Top 100 Software Companies! LILT makes it onto the Inc. 5000 List. LILT continues to be an intellectual powerhouse, holding numerous patents that help power the most efficient and sophisticated AI and language models in the industry. Check out all our news on our website.
Information collected and processed as part of your application process, including any job applications you choose to submit, is subject to LILT's Privacy Policy at https://lilt.com/legal/privacy. At LILT, we are committed to a fair, inclusive, and transparent hiring process. As part of our recruitment efforts, we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications, including résumé screening, assessment scoring, and interview analysis. These tools are designed to support human decision‑making and help us identify qualified candidates efficiently and objectively. All final hiring decisions are made by people. If you have any concerns, require accommodations, or would like to opt‑out of the use of AI in our hiring process, please let us know at recruiting@lilt.com. LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individual’s race, religion, color, national origin, ancestry, sex, sexual orientation, gender identity, age, physical or mental disability, medical condition, genetic characteristics, veteran or marital status, pregnancy, or any other classification protected by applicable local, state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.
Research Engineer, Evaluations, Applied AI employer: LILT AI
LILT AI is an exceptional employer that values creativity and autonomy, offering Welsh voice talents the chance to work on innovative AI projects from the comfort of their own homes. With competitive rates and prompt payments, our flexible scheduling allows you to balance your passion for language and technology with your personal commitments, making it a rewarding opportunity for those looking to grow in the voice acting field.
StudySmarter Expert Advice🤫
We think this is how you could land Research Engineer, Evaluations, Applied AI
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like LILT AI!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Research Engineer, Evaluations, Applied AI at LILT AI.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like LILT AI.
✨Apply Directly through Our Website
When you find a suitable opening like Research Engineer, Evaluations, Applied AI at LILT AI, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Research Engineer, Evaluations, Applied AI
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at LILT AI, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at LILT AI. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at LILT AI
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at LILT AI!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.