At a Glance
- Tasks: Architect and maintain data pipelines for real-time AI applications in the energy sector.
- Company: Join a pioneering tech company focused on sustainable energy solutions.
- Benefits: Competitive salary, flexible work options, and opportunities for professional growth.
- Other info: Dynamic team environment with exciting challenges and career advancement.
- Why this job: Be at the forefront of AI in energy, making a real-world impact.
- Qualifications: Strong skills in PostgreSQL, Python, AWS, and data engineering.
The predicted salary is between 63000 - 77000 £ per year.
Applied Computing was founded in 2024 to build Orbital, a physics-informed foundation model for energy operations.
We’re live across oil and gas, refineries, and petrochemicals, working towards our mission: sustainable
abundance for a growing planet.
The hydrocarbon industry keeps the world running.
But its complexity has left operators tied to legacy systems, making critical decisions on less than 10% of available data.
We built Orbital to change that.
It’s a foundation model built specifically for energy that lets companies use AI at scale, harnessing all of their operational
data and optimising in real time for any metric.
Decisions get faster, operations get safer, and carbon intensity falls.
We’ve raised over $32 million, including one of the largest seed rounds for an
AI company in the UK. We’re just getting started
The Role
As our Data Engineer, you’ll architect and maintain pipelines that make high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs.
You’ll be working across AWS (EKS, S3, EBS, KMS, Cloud Watch) and Databricks/Py Spark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and real-time LLM workloads.
This isn’t a traditional ETL role, you’ll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement.
- Technical Requirements
- Deep expertise in Postgre SQL (partitioning, indexing, query optimisation, storage design).
- Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
- Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, Cloud Watch, etc.)for secure and scalable data pipelines.
- Proven ability to work with Databricks and Py Spark for large-scale distributed data processing.
- Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
- Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
- Bonus: Experience working as a data engineer in oil and gas or energy environments
- Bonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.
- Core Responsibilities
- 1. Ingest & Contextualise Data
- Ingest from OPC UA servers, process historians, Io T sensors, LIMS systems, alarms/events, and P&IDs.
- Map signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.
- 2. Data Movement & Accessibility
- Build pipelines that handle real-time streaming and batch ingestion into the Lakehouse.
- Manage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).
- Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
- Handle secure, high-throughput transfers between historian archives and sandbox/live environments.
- 3. Change Tracking & Integrity
- Detect and manage schema changes, signal drift, and inconsistencies acrosstime.
- Implement lineage and audit trails across Spark/Databricks and AWS pipelines.
- 4. Data Preparation for AI
• Build and maintaindual pipelines
- Training→ large-scale historical data prep for time-series + LLM training.
- Inference→ low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.
- Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).
- 5. Database Performance & Optimisation
- Tune Postgre SQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation).
- Optimise pipelines for both fast analytical queries and high-efficiency model training.
- Deploy and manage data pipelines in AWS EKS (Kubernetes) with persisten t EBS-backed storage.
- What Success Looks Like
- Live data streams are contextualised, queryable, and AI-ready.
- Schema changes and signal drift are detected and handled without breaking downstream workflows.
- Training and inference pipelines run smoothly in parallel, optimised for scale and latency.
- #J-18808-Ljbffr
Data Engineer, Forward Deployed employer: Applied Computing
Applied Computing is an exceptional employer, offering a dynamic work environment where innovation meets sustainability. As a People Partner, you'll play a crucial role in shaping our People practices while enjoying autonomy and support from a collaborative team. With a focus on employee growth and a commitment to meaningful work, you'll have the opportunity to make a real impact in a fast-paced, scaling organisation dedicated to transforming the energy sector.
StudySmarter Expert Advice🤫
We think this is how you could land Data Engineer, Forward Deployed
✨Get Involved in Data Science Meetups
Tap into local data science meetups or workshops to connect with fellow enthusiasts and professionals. These events are goldmines for networking, and sometimes even lead directly to job openings at companies like Applied Computing!
✨Show Off Your Projects
Start building a public portfolio showcasing your data science projects on platforms like GitHub or personal websites. Highlight unique analyses or models you've developed. This not only demonstrates your skills but also gets your name out there for roles like Data Engineer, Forward Deployed at Applied Computing.
✨Leverage Professional Networks
Join professional bodies related to data science, like the Data Science Society or similar organisations. Getting involved can lead to mentorship opportunities and insider knowledge about full-time positions at companies like Applied Computing.
✨Apply Directly through Our Website
When you find a suitable opening like Data Engineer, Forward Deployed at Applied Computing, make sure to apply directly through our website. It gives you an edge and shows you're keen to join our team. Plus, who doesn’t love a direct application? It’s easier than navigating through job boards!
We think you need these skills to ace Data Engineer, Forward Deployed
Some tips for your application 🫡
Show Off Your Projects:In the world of data science, your projects can speak volumes about your skills. Make sure to showcase a few key projects in your CV or portfolio, especially those that highlight your ability to work with data sets, build models, or use relevant tools like Python, R, or SQL. Don’t forget to include links to any GitHub repositories if applicable!
Quantify Your Achievements:Employers love numbers! When drafting your CV, highlight your achievements with quantifiable results. For instance, mention how your data analysis led to a certain percentage increase in efficiency or revenue at a previous job or project. These details can really make your application pop!
Craft a Tailored Cover Letter:For a full-time role at Applied Computing, your cover letter should reflect your passion for data science and your excitement about the specific projects or values of the company. Dive into why you’re a good fit, how your skills align with their needs, and any unique perspectives you can bring to the team.
Stand Out with Relevant Courses and Certifications:Although experience talks, relevant courses or certifications can be your ticket to impressing hiring managers at Applied Computing. Mention any standout courses you've completed that equipped you with essential skills, such as machine learning certifications or data visualisation courses. This shows your commitment to continuously developing your skills in the field!
How to prepare for a job interview at Applied Computing
✨Brush Up on Your Statistics
For a data science role, we need to seriously sharpen our statistics skills. Get ready to tackle technical questions on probability distributions, hypothesis testing, and regression analysis. These are often the bread and butter of data science interviews, so don't just skim over them!
✨Showcase Your Projects
Prepare a killer portfolio showcasing your data science projects. We should include details about the datasets used, the tools and techniques applied, and the impact of your findings. If we can walk them through a particularly challenging project or a cool visualisation that had real-world implications, it’ll really make us stand out!
✨Get Comfortable with Python and R
Most data science positions require us to be proficient in programming languages like Python and R. We should practice common libraries like pandas, NumPy, and scikit-learn, and be ready for live coding exercises or algorithm questions. Showing off our coding chops can really impress the interviewers at Applied Computing!
✨Prepare for Case Studies
Expect to encounter real-world case studies during the interview. We might be asked how we’d approach a data problem or analyse a dataset to extract insights. It's essential to think out loud and demonstrate our problem-solving process so that the interviewer can see our logical thinking in action.