At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market.What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture.Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform.We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebratingour wins - big and small.Role OverviewWe are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario planning platform.You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency.Your ImpactProvision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.Your SkillsHands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.Demonstrated experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments.Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring.Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive culture strengthens our business. Build your career in a place where diversity, equity, inclusion and belonging aren’t just words on paper - this is what drives our innovation, it’s how we connect, and it contributes to what makes us a market leader. We believe in a hiring and working environment where all people are respected and valued, regardless of gender identity or expression, sexual orientation, religion, ethnicity, age, neurodiversity, disability status, citizenship, or any other aspect which makes people unique. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, perform essential job functions, and receive equitable benefits and all privileges of employment. Please contact us to request accommodation. Fraud Recruitment DisclaimerIt has come to our attention that fraudulent and fictitious job opportunities are being circulated on the Internet. Prospective candidates are being contacted by certain individuals, mainly through telephone calls, emails and correspondence, claiming they are representatives of Anaplan. Extend offers to candidates without an extensive interview process with a member of our recruitment team and a hiring manager via video or in person. Should you have any doubts about the authenticity of an email, letter or telephone communication purportedly from, for, or on behalf of Anaplan, please send an email to people@anaplan.Candidate data processed during our recruitment activities is handled in accordance with our Candidate Privacy Notice. This may include the use of artificial intelligence or automated tools to assist our team in evaluating qualifications.
AWS Engineer - Remote Working in London employer: Anaplan
Anaplan is an exceptional employer that fosters a Winning Culture, where innovation and collaboration thrive. Located in London, we offer a dynamic work environment that champions diversity and provides ample opportunities for professional growth, ensuring that every employee feels valued and inspired to contribute to our mission of transforming business decision-making. With a commitment to inclusivity and a focus on celebrating achievements, Anaplan is the perfect place for those looking to make a meaningful impact while advancing their careers.