Senior Researcher in London

Senior Researcher in London

London Full-Time 70000 - 90000 £ / year (est.) No working from home possible
C

At a Glance

  • Tasks: Lead innovative research in GPU infrastructure and machine learning to enhance system reliability.
  • Company: CoreWeave, a pioneering cloud platform for AI with a focus on innovation.
  • Benefits: Competitive salary, family-level medical insurance, generous pension contributions, and tuition reimbursement.
  • Other info: Join a dynamic team in a fast-growing company with endless opportunities for growth.
  • Why this job: Make a real impact on AI infrastructure and work at the forefront of technology.
  • Qualifications: 8+ years in machine learning or applied AI, with strong Python skills.

The predicted salary is between 70000 - 90000 £ per year.

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability.

Role Overview

We are looking for a Senior Researcher to join Monolith’s Research team, now part of CoreWeave. This is a high-impact, high-ownership role for a researcher who combines deep technical expertise in machine learning, statistical modelling, optimisation, and large-scale systems data with the ability to take complex, ambiguous problems from first principles through to production. The Monolith Data Science team is building a layered reliability and intelligence platform that shifts CoreWeave from reactive troubleshooting to proactive reliability engineering. The platform spans telemetry ingestion, feature engineering, anomaly detection, failure prediction, distributed straggler detection, performance modelling, workload optimisation, and agentic root cause analysis.

You will work closely with Fleet, Infrastructure, AI Platform, engineering, product, and client-facing teams to improve cluster reliability, increase effective utilisation, reduce MTTR, protect uptime, and turn large-scale GPU infrastructure telemetry into measurable operational and commercial impact. This is not a traditional data science role focused on dashboards, business metrics, or standard forecasting. The role sits at the intersection of applied research, GPU infrastructure, high-performance computing, distributed systems, reliability engineering, telemetry, optimisation, and Physical AI. It demands rigorous scientific thinking, strong execution, and comfort working in a high-ambiguity environment where the right problem framing is often as important as the final model.

What You’ll Do

  • Research Leadership & Strategy: Contribute meaningfully to Monolith and CoreWeave’s research direction by identifying high-leverage problems in GPU infrastructure analytics, cluster reliability, workload performance, scheduling, and utilisation. Originate novel research directions for turning raw infrastructure telemetry into actionable intelligence, rather than simply applying standard machine learning or data science techniques. Evaluate emerging methods across statistical modelling, machine learning, observability, optimisation, simulation, reinforcement learning, anomaly detection, and autonomous diagnostics, providing well-grounded technical judgement on which approaches are most likely to create real-world impact. Champion rigour, reproducibility, and scientific integrity across research outputs, experiments, prototypes, and production validation. Help establish a research foundation for understanding how large-scale GPU systems behave, why workloads underperform, where bottlenecks emerge, and how reliability can be improved proactively.
  • Technical Depth & Execution: Lead the design and development of sophisticated statistical, machine learning, and optimisation systems for large-scale GPU infrastructure telemetry, including compute, networking, storage, workload, and distributed systems data. Develop advanced models and methodologies to optimise GPU utilisation, workload scheduling, infrastructure efficiency, and system reliability. Build models and methods for anomaly detection, failure prediction, distributed straggler detection, degraded workload identification, bottleneck diagnosis, and agentic root cause analysis. Design experiments, analyse large-scale system telemetry, and prototype predictive and optimisation algorithms that directly inform production systems. Drive technical decisions on difficult modelling problems involving noisy time-series data, high-dimensional telemetry, causal inference, uncertainty, robustness, generalisation, and out-of-distribution behaviour. Explore simulation, digital-twin, reinforcement learning, and adaptive scheduling approaches where they can improve understanding or optimisation of GPU clusters and distributed training environments. Take end-to-end ownership of research work from problem framing and exploratory analysis through prototype development, validation, and collaboration with engineering teams on production deployment. Maintain deep personal technical expertise; remain a hands-on contributor in Python and modern scientific computing / machine learning tooling.
  • Organisational Influence & Collaboration: Serve as a strong technical voice within the research organisation, helping shape how Monolith approaches complex infrastructure intelligence problems. Work closely with Fleet, Infrastructure, AI Platform, engineering, product, and customer-facing teams to ensure research work lands with real operational and commercial impact. Translate research findings into production-ready prototypes, deployable solutions, and technical recommendations that improve performance, reliability, utilisation, and cost efficiency. Contribute to research practices and norms that improve how the team handles ambiguous, high-dimensional, real-world systems problems. Communicate complex technical work and its implications clearly to a range of audiences, from close technical collaborators to senior leadership and external stakeholders. Help build a shared understanding of how large-scale AI infrastructure behaves, where it fails, and how it can be made more reliable, efficient, and intelligent.

What We’re Looking For

  • 8+ years of experience, or equivalent research experience, applying statistical modelling, machine learning, optimisation, or applied AI to large-scale datasets.
  • MS or PhD in Computer Science, Statistics, Applied Mathematics, Machine Learning, Physics, Engineering, or a related quantitative field.
  • Strong proficiency in Python and scientific computing libraries such as NumPy, pandas, SciPy, scikit-learn, PyTorch, or TensorFlow.
  • Experience working with large-scale structured datasets, time-series data, infrastructure telemetry, performance data, sensor data, or other complex operational data.
  • Experience designing and analysing controlled experiments, including A/B testing, hypothesis testing, causal inference, or rigorous model validation.
  • Experience building and validating predictive models in production or research environments.
  • Experience with distributed data systems such as Spark, Ray, Dask, or similar.
  • Proficiency in SQL and working with large-scale structured data.
  • Strong understanding of optimisation techniques such as linear programming, convex optimisation, stochastic optimisation, reinforcement learning, or adaptive scheduling.
  • Demonstrated ability to solve ambiguous technical problems where the right approach is not already known.
  • Ability to translate research findings into production-ready prototypes, deployable workflows, or operational tooling.
  • Strong scientific judgement, including experimental design, reproducibility, validation, and awareness of uncertainty.
  • The ability to communicate clearly and influence across research, engineering, product, infrastructure, and leadership audiences.

Preferred Experience

  • PhD with published research in systems optimisation, distributed computing, ML systems, performance modelling, reliability engineering, scientific computing, or a related area.
  • Experience with GPU workloads, distributed training, AI infrastructure, HPC, or large-scale compute environments.
  • Familiarity with Kubernetes, containerised workloads, cloud-native systems, or distributed infrastructure.
  • Experience developing reinforcement learning, adaptive scheduling, autonomous diagnostics, or agentic systems.
  • Background in capacity planning, forecasting, resource allocation modelling, or infrastructure efficiency.
  • Experience with observability, hardware telemetry, performance monitoring, root cause analysis, or failure prediction.
  • Contributions to open-source machine learning, systems, infrastructure, or scientific computing projects.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast. We’re in an exciting stage of hyper-growth, operating at the centre of the demand for large-scale accelerated compute. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: Be Curious at Your Core, Act Like an Owner, Empower Employees, Deliver Best-in-Class Client Experiences, Achieve More Together.

By joining Monolith’s Research team within CoreWeave, you will work on problems that sit directly at the frontier of AI infrastructure: how massive GPU systems behave, why workloads underperform, how they fail, and how they can be made more reliable, efficient, and intelligent. This is an opportunity to help build a new category of infrastructure intelligence — one that moves beyond monitoring and dashboards toward systems that can understand, explain, predict, and optimise the behaviour of large-scale GPU clusters. We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As the organisation continues to grow, the opportunities to shape new technical directions are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too.

Senior Researcher in London employer: CoreWeave Europe

CoreWeave is an exceptional employer that champions innovation and collaboration, making it an ideal place for a Data Centre Construction Project Manager. With a strong focus on employee growth, competitive benefits including family-level medical insurance and generous pension contributions, and a vibrant work culture that embraces curiosity and entrepreneurial thinking, CoreWeave offers a unique opportunity to thrive in the fast-paced AI cloud sector. Located in London, you will be part of a dynamic team dedicated to delivering best-in-class client experiences while enjoying the perks of working in a Living Wage accredited environment.

C

Contact Details:

CoreWeave Europe Recruitment Team

We think you need these skills to ace Senior Researcher in London

Machine Learning
Statistical Modelling
Optimisation
Large-Scale Systems Data Analysis
Anomaly Detection
Failure Prediction
Distributed Systems