Member of Technical Staff (Language Model Evaluations) in Tipton

Member of Technical Staff (Language Model Evaluations) in Tipton

Tipton Full-Time 80000 - 100000 £ / year (est.) Home office (partial)
Artificial Analysis, Inc.

At a Glance

  • Tasks: Design and build next-gen evaluations for cutting-edge AI models.
  • Company: Join Artificial Analysis, the leading independent AI benchmarking company.
  • Benefits: Competitive salary, equity options, and a chance to shape AI's future.
  • Other info: Dynamic team environment with opportunities for rapid career growth.
  • Why this job: Influence AI development and become an expert in frontier technologies.
  • Qualifications: 3+ years in language model evaluation and strong Python skills required.

The predicted salary is between 80000 - 100000 £ per year.

Location: San Francisco (preferred), Sydney, Melbourne, Brisbane

About Artificial Analysis

Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. Our benchmarks don’t just measure the cutting edge of AI, they are actively shaping the frontier. Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.

The Opportunity

Language model evaluation is the sharpest question in AI: what can these systems actually do? We’re hiring Members of Technical Staff to build the next generation of them. This is a role for people who want to build frontier benchmarks: designing evaluations that stay ahead of frontier capabilities, constructing datasets that resist contamination, and measuring what everyone else has not yet worked out how to measure.

What You’ll Do

  • Design Next-Generation Frontier Evals: Conceive and ship the next generation of frontier evaluations, like AA-Briefcase and AA-Omniscience, across reasoning, knowledge, coding, agentic capability and beyond.
  • Build Evaluation Datasets and Infrastructure: Construct the datasets, harnesses and scoring systems behind our benchmarks, engineered for contamination resistance and repeatability at frontier scale.
  • Shape the Future Intelligence Index: The evaluations you build will contribute to future versions of the Artificial Analysis Intelligence Index and other areas of our platform, defining how the industry measures frontier capability.
  • Publish Influential Analysis: Produce the reports, indexes and data visualizations that shape how the industry understands language model progress.
  • Work with Frontier Labs on Pre-Release Models: Benchmark the leading labs’ systems, including pre-release and newly launched models, working directly with their research teams.
  • Evaluate Every Major Model: Run our evaluation suite across frontier releases as they land, and own the integrity of the results the industry quotes.
  • Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking.

What We’re Looking For

You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail. Backgrounds include: evaluation and benchmarking teams at AI labs; research or engineering roles at evaluation-focused organizations; ML engineers who have built evaluation harnesses and datasets in production; or academic researchers in NLP and ML evaluation with a strong record of published work.

Required:

  • 3+ years of relevant professional experience, across industry or research.
  • Strong analytical and critical thinking skills.
  • Strong Python, with hands‑on experience running evaluation harnesses and building datasets.
  • Deep familiarity with the LLM evaluation landscape: the major benchmarks and their failure modes, contamination, preference‑based methods, and agentic evaluation.
  • Strong statistical grounding: you know when a result is signal and when it is noise.
  • Genuine, demonstrable interest and knowledge of Frontier AI.

Why Artificial Analysis?

  • Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities.
  • Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released.
  • Work with the most important players in AI: You’ll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice.
  • Join at a defining moment: We’re 40+ people, on track to double by end of year, backed by some of the most connected investors in AI.
  • Competitive compensation including equity.

Member of Technical Staff (Language Model Evaluations) in Tipton employer: Artificial Analysis, Inc.

Artificial Analysis is an exceptional employer, offering a unique opportunity to work at the forefront of AI technology in the vibrant city of San Francisco. With a strong focus on employee growth and development, team members are encouraged to become world experts in AI while collaborating with leading labs and enterprises. The company fosters a dynamic work culture that values innovation and reliability, providing competitive compensation and equity as part of its commitment to attracting top talent.

Artificial Analysis, Inc.

Contact Details:

Artificial Analysis, Inc. Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Member of Technical Staff (Language Model Evaluations) in Tipton

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Artificial Analysis, Inc.!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Member of Technical Staff (Language Model Evaluations) at Artificial Analysis, Inc..

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Artificial Analysis, Inc..

Apply Directly through Our Website

When you find a suitable opening like Member of Technical Staff (Language Model Evaluations) at Artificial Analysis, Inc., make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Member of Technical Staff (Language Model Evaluations) in Tipton

Language Model Evaluation
Benchmarking
Dataset Construction
Python Programming
Statistical Analysis
Analytical Skills
Critical Thinking

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at Artificial Analysis, Inc., your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Artificial Analysis, Inc.. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at Artificial Analysis, Inc.

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Artificial Analysis, Inc.!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.