At a Glance
- Tasks: Build and optimise AI compute infrastructure for cutting-edge financial technology.
- Company: Join Kraken, a leading crypto platform with a global impact.
- Benefits: Enjoy competitive salary, remote work, and opportunities for professional growth.
- Other info: Dynamic, inclusive culture that values diverse perspectives and skills.
- Why this job: Be part of a high-impact team shaping the future of AI in finance.
- Qualifications: 5+ years in infrastructure engineering with GPU and ML experience.
The predicted salary is between 72000 - 88000 £ per year.
Building the Future of Open Finance. Payward - the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services and CF Benchmarks - has spent the last 15 years building one of the most modern and globally accessible financial infrastructure platforms in the industry, built to advance an open, global financial system.
The team is responsible for GPU and accelerator infrastructure, cluster operations, scheduling, model serving, observability, capacity planning, and cost-efficient compute at scale. This is the backbone that allows Kraken to train, serve, evaluate, and iterate on AI systems in-house where it matters for privacy, latency, reliability, cost, or product differentiation.
You will join a small, senior, high-impact team working directly with AI/ML researchers, platform engineers, security teams, and product teams. The mandate is simple: make Kraken's AI ambitions real by building compute infrastructure that is fast, dependable, efficient, and production-grade.
The opportunity:
- Own and operate GPU and accelerator clusters used for training, inference, evaluation, and experimentation.
- Design infrastructure that enables Kraken teams to run models locally on GPUs.
- Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerator environments.
- Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost.
- Partner with ML engineers and researchers to remove bottlenecks in workflows.
- Build observability for GPU utilization, memory pressure, queue depth, and other metrics.
- Drive reliability, incident response, alerting, runbooks, and post-incident improvements.
- Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
- Build tooling that makes GPU usage visible and easier for internal teams.
- Contribute to long-term architecture decisions that balance performance, cost efficiency, scalability, operational simplicity, and production safety.
What You Bring:
- 5+ years of infrastructure engineering experience, with significant time spent on GPU compute, ML infrastructure, distributed systems, or large-scale production platforms.
- Hands-on experience operating GPU clusters or accelerator-backed infrastructure.
- Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, and production debugging.
- Experience with ML serving frameworks.
- Proficiency in Python for infrastructure automation and operational workflows.
- Practical understanding of performance tradeoffs across various metrics.
- Track record of optimizing compute costs while maintaining performance expectations.
- Experience building observable systems with useful metrics and incident workflows.
- Comfortable working in high-stakes environments.
- Clear communicator who can translate infrastructure tradeoffs.
Nice to haves:
- Experience at a frontier AI lab, hyperscaler, or high-frequency trading firm.
- Familiarity with custom silicon or specialized accelerators.
- Background in capacity planning or GPU fleet cost management.
- Experience with distributed training frameworks.
- Experience debugging performance issues.
- Experience with systems languages used for performance-critical infrastructure.
- Crypto, financial services, or security-sensitive production infrastructure experience.
Unless a specific application deadline is stated in the job posting, applications are accepted on an ongoing basis.
Our commitment: Payward is powered by people from around the world and we celebrate the diverse talents, backgrounds, contributions, and unique perspectives that everyone brings to the table. We hire based on merit, seeking out people with the right abilities, knowledge, and skills for the job.
As an equal opportunity employer, we don’t tolerate discrimination or harassment of any kind.
Senior AI Compute Infrastructure Engineer in Moffat employer: Kraken
At Payward, we pride ourselves on being an exceptional employer, offering a fully remote work environment that promotes flexibility and work-life balance. Our culture is rooted in hyper-transparency and open dialogue, fostering collaboration across diverse teams while providing ample opportunities for professional growth and development in the fast-paced fintech landscape. Join us to be part of a forward-thinking team dedicated to building a resilient financial infrastructure that empowers innovation and inclusivity in the global financial system.