Principal Core Infrastructure Engineer

Principal Core Infrastructure Engineer

Full-Time 63000 - 77000 £ / year (est.) No working from home possible
O

At a Glance

  • Tasks: Solve complex problems and build automation for Oracle Cloud Infrastructure services.
  • Company: Join Oracle, a leader in AI and cloud solutions impacting billions.
  • Benefits: Flexible medical, life insurance, retirement options, and community volunteer programs.
  • Other info: Be part of a diverse team committed to innovation and inclusivity.
  • Why this job: Make a real impact on service reliability and customer experience in a fast-paced environment.
  • Qualifications: 8+ years in Site Reliability Engineering with strong coding and troubleshooting skills.

The predicted salary is between 63000 - 77000 £ per year.

Core Infrastructure Engineering within Oracle Cloud Infrastructure (OCI) is seeking a motivated Principal Site Reliability Engineer (SRE) who thrives in a fast-paced, rapidly evolving technology environment. The ideal candidate should have experience supporting cloud-scale, highly distributed storage or database services on major cloud platforms such as OCI, AWS, GCP, or Azure.

You will be focused on improving service reliability, performance and operability of services used by Oracle OCI Tier-0 services and Oracle customers. You will have your hand on the pulse of the services and will play a key role in responding to live service issues. As a hands-on engineer with strong coding skills, you will debug complex production issues, build automation and monitoring tools. You will have the opportunity to create automation and tooling that will allow us to continuously improve our services. You will own the release certification process by determining whether code is ready for production deployment.

In this role, you will be responsible for improving the stability, performance, and reliability of database and storage services which are the backbone of OCI. You will collaborate with multiple development teams to identify and resolve cross-functional operational risks by combining engineering expertise, troubleshooting skills, and operational best practices. The role requires a high degree of independence, excellent communication and organizational skills, and a strong commitment to improving customer experience by enhancing service reliability, reducing support tickets, and delivering scalable operational solutions. You will also define and deliver mission-critical services with a strong focus on security, resiliency, scalability, capacity planning, performance management, deployment, and release engineering.

Qualifications:

  • 8+ years’ experience in Site Reliability Engineering and in storage, networking and database troubleshooting for improving application reliability, scalability, availability
  • Excellent troubleshooting skills for resolving critical production issues in cloud services
  • Expertise in developing scripts (linux scripting), utilities and tools to automate routine or manual intensive tasks
  • Solid experience with CI/CD pipelines, Jenkins and Version Control tools (GitHub, Bit Bucket, GIT)
  • Container administration and development experience utilizing Kubernetes, Docker or similar
  • Solid experience with Configuration Management tools
  • Programming languages development experience using Python, Golang, Terraform
  • Experience with monitoring tools such as Grafana
  • Experience in managing 24/7 high-availability production applications
  • Excellent organizational, verbal, and written communication skills
  • Good understanding of Agile software development principles including using common tools such as JIRA

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Career Level - IC4

Solve complex problems related to Oracle OCI services and build automation to prevent problem recurrence. Identify opportunities and drive the implementation of automation to improve service health, manageability, reliability and telemetry. Configure, design, and script end-to-end service telemetry, alerting and self-healing capabilities for platforms. Author functional and technical documentation. Serve as part of a 24x7 On Call rotation in support of the infrastructure life cycle. Have a professional curiosity and a desire to develop a deep understanding of services and technologies. Educate yourself and others on anything that helps Service Teams more quickly and easily build, test, deploy & run their Services to be more reliable. Identify problems and/or opportunities for improvements that are common across many teams/services and design and implement the solutions.

Principal Core Infrastructure Engineer employer: Oracle Corporation

Oracle is an exceptional employer, offering a dynamic work culture that fosters innovation and collaboration in the healthcare market research sector. With a commitment to employee growth, competitive benefits, and a focus on inclusivity, Oracle empowers its workforce to make impactful contributions while enjoying flexible medical, life insurance, and retirement options. Join us in a role that not only influences high-stakes decisions for leading organisations but also allows you to integrate cutting-edge technology into meaningful research initiatives.

O

Contact Details:

Oracle Corporation Recruitment Team

We think you need these skills to ace Principal Core Infrastructure Engineer

Site Reliability Engineering
Cloud Services Support
Troubleshooting Skills
Linux Scripting
CI/CD Pipelines
Jenkins
Version Control (GitHub, Bit Bucket, GIT)