hackajob is partnering directly with Accenture to hire for this role. YOU ARE As a Lead and Principal Infrastructure Architect, youownend-to-end responsibility for designing optimized compute infrastructure for large-scale AI and machine learning systems, including large-scale distributed training environments. You are the authority who translates business goals, SLAs, and client standards into infrastructure architectures that perform at scale while being deliberately engineered for cost-efficiency. Drawing on deep experience, you weigh multipleviablesolutions for any given problem acrosscompute, networking, storage, orchestration, and model serving and make rational, well-justified architectural decisions tailored to each client's situation, constraints, and standards. You architect andoptimizethe full computational stack for performance, power, cost, and scalability; design and tune large-scale GPU clusters and distributed training systems; and ensure infrastructure meets security, compliance, and regulatory requirements. As the recognized AI infrastructure expert in at least onehyperscalercloud (such as AWS, Azure, or Google Cloud), you bring authoritative knowledge of that platform's AI/ML services, accelerators, networking, and cost levers, and apply it to deliver best-in-class solutions. Beyond design, you set technical direction and standards, lead and mentor engineers and architects, partner with clients and stakeholders to shape the infrastructureroadmap, andare ultimately accountable for delivering AI/ML infrastructure that meets business SLAs, controls cost, and scales to enterprise and frontier workloads. THE WORK Own the end-to-end architecture and design of optimized compute infrastructure for large-scale AI/ML systems, including large-scale distributed training environments, from concept through delivery. Develop and evaluate architecture alternatives, weighing trade-offs acrosscompute, networking, storage, orchestration, and model serving to make rational, well-justified decisions tailored to each client's situation and standards. Lead architecture assessments and reviews of existing and proposed environments,identifyinggaps, risks, bottlenecks, and optimization opportunities, and recommending remediation. Drive architectural decision-making, documenting rationale, trade-offs, andassumptionsso decisions are transparent, defensible, and aligned with business SLAs and standards. Define andmaintainthe AI infrastructure roadmap, planning capacity, scaling, and technology evolution in step with business and product goals. Architect andoptimizethe full computational stack for performance, power, cost, and scalability, ensuring infrastructure meets business SLAs while being deliberately engineered for cost-efficiency. Design and tune large-scale GPU clusters and distributed training systems, including acceleratorselection, interconnect/networking, and storage for high-throughput training workloads. Serve as the authoritative AI infrastructure expert in at least onehyperscalercloud (AWS, Azure, or GCP), applying deep knowledge of its AI/ML services, accelerators, networking, and cost levers. Design deployment, automation, and CI/CD strategies for reliable, repeatable, and scalable releases of AI systems, models, and data pipelines into production. Establish AI monitoring and observability strategy acrossInfraOpsandMLOps, defining SLAs, SLOs, alerting, and performance/cost tracking, and driving continuous optimization. Integrate AI/ML systems into enterprise environments, ensuring interoperability, security, compliance, and adherence to regulatory and client standards. Lead capacity planning and cost modeling, forecasting compute needs and engineering cost-efficiency into the architecture without compromising performance. Collaborate with clients, stakeholders, and engineering teams to align infrastructure decisions with business outcomes, translating requirements into actionable architecture and standards. Set technical direction, standards, and best practices, mentoring engineers and architects and leading design and code reviews across the team. EDUCATION Bachelor's Degree in ComputerScience, ComputerEngineering, related Engineering field BASIC (REQUIRED) QUALIFICATION Solid background in coding, building, monitoring,troubleshooting applicationsof AI/ML models; selecting, designing and infrastructurefor deployingand running themon premiseor on public cloud. Strong understanding of AI and machine learning as a subject. Strong understanding of computinginfrastructure asubject, preferred knowledge of AI infrastructure. Good proficiencyin programming languages such as Python, Java, or C++. Experience with data pipeline and workflow management tools (e.g., Apache Airflow, Kubeflow). Strong problem-solving skills and ability to work in a fast-paced environment. Excellent communication and collaboration skills. Significant experience in AI/ML infrastructure engineering or related roles on ahyperscalerplatform for deploying large scale solutions. Proven experience in leading and managing AI projects and teams. Strong project management skills, with the ability to manage multiple projects simultaneously. Demonstrated experience in evaluating and selecting AI technologies and frameworks. Ability to work with cross-functional teams and drive project alignment. About Accenture Accenture is a leading global professional services company that helps the worlds leading businesses, governments and other organizations build their digital core, optimize their operations, accelerate revenue growth and enhance citizen servicescreating tangible value at speed and scale. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of change today, and we are one of the worlds leaders in helping drive that change, with strong ecosystem relationships. We combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and global delivery capability. Our broad range of services, solutions and assets across Strategy
AI Infrastructure Lead Architect employer: Hackajob Ltd
JPMorgan Chase is an exceptional employer, offering a dynamic work environment where innovation and collaboration thrive. As a Lead Site Reliability Engineer, you will not only tackle complex challenges but also benefit from extensive professional development opportunities and a strong commitment to diversity and inclusion. Located in a global financial hub, you'll be part of a team that values your expertise and encourages a culture of continuous improvement and technical excellence.