Fleet Management builds the automation layer for Meta's data center fleet. We own the systems that plan, orchestrate, and execute the operational workflows that keep hundreds of thousands of machines healthy, available, and efficient across every region, and increasingly across public cloud providers as well.Our platforms cover the full lifecycle of work performed on the fleet:Planned maintenance and rollouts: firmware, kernel, OS, and hardware maintenance executed safely and predictably at fleet scale.Unplanned work and repair: detecting failures, deciding what to do about them, and driving the repair workflow through until capacity returns to production.Fleet lifecycle operations: turn-ups, decommissions, server and service moves, rack logistics, and capacity rebalancing.Safety and capacity control: deciding how much of the fleet may be unavailable at once, resolving conflicts between competing operations, rate limiting, graceful cancellation, and emergency stop.Autonomy: replacing human-driven operational decisions with automation and agentic workflows, backed by the observability and quality signals needed to trust them.Practically, this means we build large-scale workflow orchestration engines: systems that model operational intent, schedule it against fleet constraints, execute it across dozens of downstream systems, and decide what to do when steps fail. The engineering problems are distributed systems problems - state machines, scheduling, conflict resolution, idempotency, partial failure, and strong safety guarantees over irreversible physical actions.The fleet is growing fast, and the operational volume with it. We are not going to staff our way through that, so our mandate is to make fleet operations scale by an order of magnitude without a corresponding increase in human effort.BS/MS in Computer Science or equivalent practical experience 8+ years of professional software engineering experience, including significant time as the technical owner of a large production systemDemonstrated staff-level scope: you have led the design and delivery of systems spanning multiple teams, across multiple planning cycles, with measurable impactDeep distributed systems expertise β workflow or orchestration engines, control planes, scheduling, state machines, or resource management systemsProficiency in one or more of Python, C++, Java, Rust or Go, demonstrated through professional experience building and maintaining production systems in one or more of Python, C++, Java, Rust or GoTrack record of operating critical systems: oncall ownership, incident leadership, observability and SLO design, and structural reliability improvementAbility to lead ambiguous, cross-organisational work to completion, and to communicate clearly in writing to both engineers and senior leadershipExperience growing other engineers through mentoring, design review and technical direction-settingExperience with constraint-based scheduling, planning or solver-backed systemsDemonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)Experience applying AI/agentic automation to operational workflows in productionBackground in infrastructure automation, fleet or capacity management, hardware lifecycle, data center operations, or cloud infrastructureExperience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologiesComfort working in a large mature codebase with heavy cross-team dependenciesExperience building workflow/orchestration or job execution engines used by many internal customersExperience consolidating or deprecating legacy systems while keeping them runningExperience integrating with heterogeneous external infrastructure APIs, including public cloud provider lifecycle and maintenance models
#J-18808-Ljbffr
Software Engineer - Fleet Management & Repair Automation in London employer: Meta Careers
Meta is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration among top engineers in the heart of London. As a Production Engineering Intern, you'll not only gain invaluable hands-on experience working on high-impact projects but also benefit from numerous growth opportunities and a supportive environment that encourages learning and development.