At a Glance
- Tasks: Drive operational excellence in incident management and service reliability.
- Company: Join a forward-thinking tech company with a collaborative culture.
- Benefits: Enjoy a competitive salary, 401(k) matching, and comprehensive benefits.
- Other info: Remote work with opportunities for professional growth and influence.
- Why this job: Make a real impact on service stability and operational readiness.
- Qualifications: 3+ years in operations or incident coordination; strong analytical skills.
The predicted salary is between 100000 - 100000 £ per year.
About the rol e
We're looking for a Reliability Operations Specialist to help drive operational excellence across incident management, service reliability, observability, and continuous improvement initiatives.
This role serves as a central coordinator and subject matter expert for reliability practices, helping teams improve service stability, reduce operational risk, and strengthen operational readiness across the organization.
The Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate incident response, oversee post-incident reviews, track corrective actions, and provide visibility into the health and reliability of our platforms.
This role does not have direct people management responsibilities but plays a critical role in influencing reliability outcomes through collaboration, process ownership, and data-driven decision making.
Location
Remote
Employment type
Permanent, Full-time
Pay Range
$85,000 - $100,000 Annually, The final compensation offered will be determined based on factors including location, experience, skills, qualifications, and market conditions.
What you'll do
- Incident Management & Operational Excellence
- Participate in major incident response activities and serve as an Incident Commander when assigned.
- Coordinate incident response efforts across multiple teams during service-impacting events.
- Facilitate escalation management, stakeholder communications, and status reporting.
- Support the ongoing improvement of incident management processes, procedures, and operational readiness.
- Drive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
- Identify opportunities to improve operational efficiency, reliability, and service delivery.
- Post-Mortem Management & Continuous Improvement
- Coordinate and facilitate blameless post-mortem reviews following significant incidents.
- Ensure post-mortems are completed accurately, consistently, and within established timelines.
- Analyze incident trends to identify recurring issues, systemic risks, and improvement opportunities.
- Maintain accountability for corrective action tracking and closure.
- Partner with stakeholders to prioritize and drive reliability-focused improvements.
- Foster a culture of learning, accountability, and continuous improvement.
- Reliability & Observability
- Partner with engineering teams to define, maintain, and mature Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Support the development and evolution of observability practices, including monitoring, alerting, dashboards, and telemetry standards.
- Analyze reliability metrics and operational performance data to identify opportunities for improvement.
- Recommend and track initiatives that improve platform stability, resiliency, and service performance.
- Help establish operational best practices that support scalable and reliable platform operations.
- Reporting & Stakeholder Communication
- Develop reliability reporting for engineering leadership and executive stakeholders.
- Maintain incident communication standards and stakeholder notification processes.
- Provide regular reporting on incident performance, corrective actions, reliability trends, and service health.
- Translate technical reliability metrics into actionable business insights and recommendations.
- Present findings and recommendations to technical and non-technical audiences
What you bring
- 3+ years of experience in Product Operations, Platform Operations, Technical Customer Support, Incident Coordination, Site Operations, IT Service Management (ITSM), or a related operational role.
- Experience participating in or coordinating major incident response activities.
- Knowledge of incident management, root cause analysis, problem management, and post-mortem methodologies.
- Experience working with monitoring, alerting, observability, or operational reporting tools.
- Strong analytical and organizational skills with exceptional attention to detail.
- Excellent written and verbal communication skills.
- Ability to work effectively across multiple teams and influence outcomes without direct authority.
- Strong problem-solving skills and the ability to remain calm and organized during high-pressure situations.
- Preferred Qualifications
- Experience working with Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics.
- Familiarity with Linux systems, cloud infrastructure, networking concepts, hosting platforms, or distributed systems.
- Knowledge of ITIL, operational excellence frameworks, or Site Reliability Engineering (SRE) principles.
- Experience supporting high-availability Saa S, hosting, cloud, or infrastructure environments.
- Experience creating executive-level operational reports, dashboards, and presentations.
- Experience using observability and incident management platforms.
What We Offer
- Comprehensive benefits package
- Traditional and Roth 401(k) with company matching
- A collaborative, team-oriented culture
- Consistent and predictable work hours
- Engaging, varied work that keeps each day different
- Opportunities to contribute ideas and influence how work gets done
Disclaimer
This job description is only a summary of the typical functions of the position.
It is not intended to be an exhaustive or comprehensive list of all job responsibilities, tasks, or duties.
Additional duties and tasks may be assigned as part of the job function.
Nexcess reserves the right to modify, interpret, or apply this job description in a way that best supports the organizational needs.
The job description in no way creates or implies an employment contract.
The employment contract remains “at will”.
Equal Employment Opportunity Policy
Nexcess is committed to offering equal employment opportunity without regard to age, color, disability, gender, gender identity, genetic information, marital status, military status, national origin, race, religion, sexual orientation, veteran status, or any other legally protected characteristic.
Reliability Operations Specialist employer: Nexcess
Nexcess is an exceptional employer that fosters a collaborative and team-oriented culture, making it an ideal place for a Reliability Operations Specialist to thrive. With a comprehensive benefits package, including a 401(k) with company matching, and opportunities for professional growth, employees are encouraged to contribute ideas and influence operational practices. The remote work environment ensures consistent and predictable hours, allowing for a healthy work-life balance while engaging in varied and meaningful tasks.
StudySmarter Expert Advice🤫
We think this is how you could land Reliability Operations Specialist
✨Join the IT Consultancy Buzz
Get involved in local or virtual IT consultancy meetups and forums. This is where we can rub shoulders with industry professionals, get insights into what Nexcess values, and even spot unadvertised opportunities. Don't miss out on these chances to make a name for ourselves in the IT world!
✨Show Off Your Skills
Create a personal project or case study relevant to the challenges Nexcess might face. Use platforms like GitHub or Medium to share your findings. This not only demonstrates our consulting skills but shows a proactive attitude, making us stand out from the crowd when applying for that full-time gig.
✨Leverage LinkedIn for Connections
Follow and engage with the relevant thought leaders and influencers in IT consultancy on LinkedIn. Share insightful content and join discussions to gain visibility. A well-placed comment or shared article could catch the attention of someone at Nexcess!
✨Direct Apply to Nexcess
Let's not forget to apply directly through the Nexcess website! Tailor your application to showcase our understanding of their consulting style and how we can contribute to their projects. A personalised approach can make a huge difference in landing that full-time position!
We think you need these skills to ace Reliability Operations Specialist
Some tips for your application 🫡
Showcase Your Problem-Solving Skills:In IT consulting, it's all about problem-solving, so make sure your CV highlights your analytical skills and any relevant projects you've tackled. Mention specific technologies or methodologies you've used to resolve issues or improve processes; this shows you can think critically and deliver results, which is vital for us at Nexcess.
Highlight Relevant Certifications:Certifications like ITIL, PMP, or even specific tech stack qualifications can really make you stand out. Make sure to include these in your CV, as they not only demonstrate your expertise but also your commitment to staying current in the field. We love seeing candidates who are proactive about their professional development!
Tailor Your Cover Letter:Your cover letter is your chance to connect personally with us at Nexcess. Share stories about your experiences in IT consulting, and how they shaped your desire to join our team. Mention why you’re excited about this particular role, and how you see yourself contributing to our projects.
Keep It Clear and Concise:We're all busy, so make sure your application is easy to read. Use bullet points for key achievements, and don’t overload us with jargon. A clean, professional layout goes a long way. Remember, the clearer your application, the more likely we are to invite you in for an interview!
How to prepare for a job interview at Nexcess
✨Brush Up on Your Technical Skills
For an IT consulting role, be ready to demonstrate your technical prowess. You might face questions on systems integration, cloud technologies, or even troubleshooting specific software. If you have experience with tools like AWS, Azure, or even specific programming languages, make sure you can talk about them fluently.
✨Showcase Your Problem-Solving Approach
IT consulting is all about solving problems for clients. Think about how you can illustrate your approach to a past challenge using the STAR method (Situation, Task, Action, Result). It's a great way to show how you tackle complex issues and come up with effective solutions.
✨Know the Business Impact of IT Solutions
When discussing your experiences, focus not just on the tech solutions you implemented, but also on their business impact. Employers want to see that you can connect IT with organisational goals. Prep examples that highlight how your tech contributions improved efficiency or reduced costs for past clients or projects.
✨Prepare for Behavioural Questions
Since IT consulting often involves teamwork and client interactions, expect behavioural questions that assess your interpersonal skills. Be prepared with examples that demonstrate your adaptability, communication skills, and how you handle client feedback. Before the interview, think of situations where you worked closely with clients to create effective IT strategies or changes.