Reinforcement Learning Researcher

Reinforcement Learning Researcher

Full-Time 59400 - 72600 £ / year (est.) Home office (partial)
AI Startups UK

At a Glance

  • Tasks: Design and implement cutting-edge RL-based decision systems for energy optimisation.
  • Company: Join a pioneering tech company focused on sustainable energy solutions.
  • Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
  • Other info: Dynamic team environment with a focus on real-world applications and safety.
  • Why this job: Make a real impact in the energy sector with innovative AI technologies.
  • Qualifications: PhD in relevant fields and extensive RL research experience required.

The predicted salary is between 59400 - 72600 £ per year.

Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations.

We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable abundance for a growing planet.

The hydrocarbon industry keeps the world running.

But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data.

We built Orbital to change that.

It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational data and optimising in real time for any metric.

Decisions get faster, operations get safer, and carbon intensity falls.

We’ve raised over $32 million, including one of the largest seed rounds for an AI company in the UK. We’re just getting started

  • What You’ll Own
  • Orbital’s learning-based optimisation and control stack
  • RL + control hybrid systems for industrial processes
  • Safe and constrained policy learning frameworks
  • Simulation environments and digital twin integrations
  • Research → production translation for RL systems
  • Benchmarking standards for decision-making systems
  • Must-Have Qualifications
  • Ph D in Computer Science, Robotics, Control, Applied Mathematics, or related field

• First-author publications in

  • Reinforcement Learning
  • Control systems
  • Sequential decision-making
  • 3+ years of hands‑on RL research experience

Strong foundation in

  • Reinforcement Learning (online + offline)
  • Optimisation and control theory (MPC, dynamic programming, etc.)
  • Deep learning (Py Torch)

Experience with

  • Real‑world deployment of ML systems
  • Simulation environments or digital twins
  • Working with noisy, real‑world data
  • How We Work
  • Research is judged by production impact, not paper count
  • We optimise for real systems, not benchmarks alone
  • We value safe, reliable decision‑making over theoretical elegance
  • Physics, control, and learning are treated asone system
  • What This Role Is Not
  • Not toy RL environments (Atari, Mu Jo Co‑only thinking)
  • Not unconstrainedpolicy learning without safety guarantees
  • Not offline research disconnected from deployment
  • Not a support role; this position owns core optimisation IP
  • Core Responsibilities
  • 1. Design & Implement RL‑Based Decision Systems
  • Process optimisation (yield, efficiency, cost reduction)
  • Control policy learning (setpoint optimisation, constraint handling)
  • Sequential decision‑making under uncertainty

Work across

  • Model‑free RL (policy gradients, actor‑critic, offline RL)
  • Model‑based RL (world models, planning‑based methods)
  • Hybrid approaches combining RL with optimisation / MPC
  • 2. Build Physics‑Constrained RL Systems

Embed domain knowledge into policy learning

  • Hard constraints (safety, operating limits, regulatory bounds)
  • Soft constraints (efficiency, degradation, economic trade‑offs)
  • Physics‑informed reward shaping and transition models

Ensure policies

  • Respect physical feasibility
  • Generalise across operating regimes
  • Remain stable under real‑world disturbances
  • 3. Offline RL, Simulation & Digital Twin Integration
  • Develop RL systems that work in data‑scarce and risk‑sensitive environments:
  • Offline RL from historical plant data
  • Simulation‑based training via digital twins
  • Sim‑to‑real transfer strategies

Handle

  • Distribution shift
  • Partial observability
  • Sparse / delayed rewards
  • 4. Safety, Robustness & Interpretability

Design safe RL systems for production environments

  • Constrained RL / safe exploration
  • Policy validation before deployment
  • Fail‑safe mechanisms and fallback strategies

Ensure outputs are

  • Interpretable to engineers and operators
  • Auditable and explainable
  • Reliable under sensor faults and regime changes
  • 5. Production‑Grade Deployment

Deploy RL systems into real‑world infrastructure

  • Containerised deployment (Docker, AWS / Azure)
  • Integration with control systems (APC, DCS, advisory layers)
  • Real‑time inference and monitoring

Build pipelines for

  • Continuous policy evaluation
  • Safe rollout and rollback
  • Online / batch policy updates
  • 6. Benchmarking & Validation

Define evaluation standards for RL systems

  • Offline policy evaluation
  • Counterfactual analysis
  • Comparison vs MPC, heuristics, and operator baselines

Ensure

  • Measurable economic impact
  • Reproducible results
  • Defensible performance claims
  • #J-18808-Ljbffr

Reinforcement Learning Researcher employer: AI Startups UK

Odyssey is an exceptional employer, offering a dynamic work environment in London where innovation thrives. As a VP of Research, you'll lead a world-class team at the forefront of AI technology, with ample opportunities for professional growth and collaboration across disciplines. Our culture fosters creativity and ambition, making it an ideal place for those passionate about shaping the future of world models and AI.

AI Startups UK

Contact Details:

AI Startups UK Recruitment Team

StudySmarter Expert Advice🤫

We think this is how you could land Reinforcement Learning Researcher

Get Involved in Data Science Meetups

Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like AI Startups UK!

Show Off Your Projects

Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Reinforcement Learning Researcher at AI Startups UK.

Leverage Professional Networks

Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like AI Startups UK.

Apply Directly through Our Website

When you find a suitable opening like Reinforcement Learning Researcher at AI Startups UK, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!

We think you need these skills to ace Reinforcement Learning Researcher

Reinforcement Learning
Control Systems
Optimisation and Control Theory
Deep Learning (PyTorch)
Simulation Environments
Digital Twin Integration
Model-Free RL

Some tips for your application 🫡

Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!

Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!

Craft a Tailored Cover Letter:For a full-time role at AI Startups UK, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.

Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at AI Startups UK. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!

How to prepare for a job interview at AI Startups UK

Brush Up on Your Statistics

For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!

Showcase Your Projects

Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!

Get Comfortable with Python and R

Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at AI Startups UK!

Prepare for Case Studies

Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.